Conversation
Signed-off-by: seonjinn <sna@nvidia.com>
…erns Signed-off-by: seonjinn <sna@nvidia.com>
Signed-off-by: seonjinn <sna@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
/ok to test c49e043 |
Signed-off-by: seonjinn <sna@nvidia.com>
Signed-off-by: seonjinn <sna@nvidia.com>
|
Updated in d49af6c: explicit BF16 rollout with inherited quantization_ignore_patterns now emits a warning when no quantization configuration is active, then continues normally. Mixed MXFP8 rollout with BF16 first/last layers retains its effective-exclusion log and does not receive this warning. Added coverage for bf16/bfloat16 aliases, missing/None configuration, and mixed MXFP8. GB200 red/green regression validation has been submitted; the new warning tests are not yet claimed passing. |
|
GB200 validation completed (job 7096575, exit 0). With the new warning assertions and the old implementation, 8 regression cases failed as expected. With the warning implementation, the full selected suite passed: 267 tests. This includes BF16 aliases with absent/None quantization config and mixed MXFP8 exclusions without a BF16 warning. Validation used integration source 1802f7c with the same helper/test changes as d49af6c; this is not a claim of every end-to-end configuration passing. |
Signed-off-by: seonjinn <sna@nvidia.com>
Summary
BF16 rollout can inherit
quantization_ignore_patternsfrom an MXFP8 recipe. The scope logger currently raisesKeyError: 'quantization_config'before model initialization because BF16 has no generated quantization config.Skip this diagnostic when its configuration is absent. Existing MXFP8 output and quantization-config validation are unchanged. No changes to weight loading, refit, or kernels.
Tests