Skip to content

Latest commit

 

History

History
169 lines (139 loc) · 5.88 KB

File metadata and controls

169 lines (139 loc) · 5.88 KB

Config Architecture

BioNeMo Inference Runtime (BioIR) has two pydantic trees. They do not share types.

  • Model configs describe the nn.Module. Every node is a BaseConfig.
  • Pipeline configs describe the five-stage processor. The root is EngineProcessorConfig.

They meet at EngineConfig: the processor puts a model tree (or get_pretrained_config()) next to device and acceleration settings, then FoldingEngine builds the module. How to call that surface is in the API reference.

flowchart TB
    EPC[EngineProcessorConfig] --> EC[EngineConfig]
    BC[BaseConfig tree] --> EC
    EC --> FE[FoldingEngine]
    FE --> MOD[nn.Module]
Loading

Model Configs

bionemo_ir/configs/ holds shared types only. Family composites live in bionemo_ir/models/<family>/config.py next to that family's PRETRAINED_CONFIG_REGISTRY.

Primitives (PairformerConfig, DiffusionTransformerConfig, EvoformerStackConfig) are reusable layers. Family stacks (MSAModuleConfig, ExtraMSAStackConfig, AffinityModuleConfig) are family-specific assemblies — that is why they are not in configs/modules.py.

classDiagram
    BaseConfig <|-- PrimitiveConfig
    BaseConfig <|-- FamilyConfig
    FamilyConfig *-- PrimitiveConfig
    FamilyConfig *-- FamilyStack
Loading

Boltz1Config reuses MSAModuleConfig from Boltz-2. Other family roots (OpenFold3Config, ProtenixConfig, …) follow the same pattern: inherit BaseConfig, compose primitives, keep family stacks in the family file. Pretrained variants (OpenFold2_FT2_Config, AlphaFold2_1_Config, Boltz2AffinityConfig, …) subclass the family root.

set_* helpers (set_dtype, set_triangle_attention_backend, …) walk the tree by value. Class defaults are not what a run uses — get_pretrained_config() in modeling.py fills dtypes and backends. runtime_args (recycling_steps, …) are a processor dict, not fields on this tree.

EngineConfig does not inherit BaseConfig. It wraps one:

classDiagram
    class EngineConfig {
        name
        model : BaseConfig
        device : DeviceConfig
        accelerated : AcceleratedConfig
        postprocessor : PostProcessorConfig
    }
    EngineConfig *-- BaseConfig
    EngineConfig *-- DeviceConfig
    EngineConfig *-- AcceleratedConfig
    EngineConfig *-- PostProcessorConfig
Loading

FoldingEngineWrapper fills EngineConfig from engine_kwargs (config, device, accelerated_configs, postprocessor_config, profile_inference). CUDA-graph wrap is architecture — acceleration.

Folding configs hold each graph region's CUDA-graph policy in a graph_optimization_config field, such as trunk.graph_optimization_config. A region with a policy captures by default. config.disable_cuda_graphs() clears every policy in the tree; setting one field to None keeps that region eager. Both apply to models constructed from the config afterwards; refer to API — CUDA graphs.

Boltz-2 Graph Caches

Each Boltz-2 region keeps a bounded number of exact input shapes, so traffic that rotates through more shapes recaptures graphs. Boltz2Config.with_graph_cache returns a copy of the config with new limits:

from bionemo_ir.models.boltz2 import Boltz2
from bionemo_ir.models.boltz2.config import Boltz2GraphCacheConfig

config = Boltz2.get_pretrained_config("boltz-2")
config = config.with_graph_cache(
    Boltz2GraphCacheConfig(
        max_tokens=2048,
        max_graphs=4,
        budget_bytes=16 << 30,
    )
)

The limits apply to each of the trunk, diffusion_module, and confidence_pairformer regions: max_tokens is the inclusive token limit for exact shapes without padding, max_graphs the maximum number of cached shapes, and budget_bytes the estimated graph memory budget. Retained graphs consume GPU memory alongside model weights and activations; choose limits for the device and workload. Pass the returned config when constructing the model or as engine_kwargs["config"].

Pipeline Configs

ProcessorConfig is the executor (batch size, Ray vs serial). EngineProcessorConfig adds the model key, engine_kwargs, runtime_args, and one field per stage (parser, tokenizer, feature generator, engine, writer). Those five types all inherit _StageConfigBase.

classDiagram
    ProcessorConfig <|-- EngineProcessorConfig
    _StageConfigBase <|-- StageConfig
    EngineProcessorConfig --> StageConfig : five stages
Loading

Each stage field accepts bool, dict, or a typed *StageConfig. True means "run with processor defaults." resolve_stage_config() is the only constructor build_processor uses: copy a typed config, wrap a bool, or parse a dict, then fill None fields from the processor (batch_size, compute, runtime_env, model_source).

build_processor always runs all five stages. enabled is not a public skip switch.

Stage extras: init_context on tokenizer / feature generator (set a default random_seed on the feature-generator stage; a row-level seed overrides it), output_path / format on the writer, parallelism_mode=REPLICA and num_gpus on the engine. Worked examples: API — build_processor.

EngineProcessorConfig
  ├─ *StageConfig          → five stages
  ├─ runtime_args          → model.forward kwargs
  └─ engine_kwargs.config  → BaseConfig → EngineConfig

Related