Skip to content

fix(vllm): allow BF16 rollout with inherited quantization exclusions - #4110

Draft
seonjinn wants to merge 8 commits into
NVIDIA-NeMo:mainfrom
seonjinn:sna/bf16-scope-log-fix-20260911
Draft

seonjinn wants to merge 8 commits into
NVIDIA-NeMo:mainfrom
seonjinn:sna/bf16-scope-log-fix-20260911

Conversation

@seonjinn

Copy link
Copy Markdown
Contributor

Summary

BF16 rollout can inherit quantization_ignore_patterns from an MXFP8 recipe. The scope logger currently raises KeyError: 'quantization_config' before model initialization because BF16 has no generated quantization config.

Skip this diagnostic when its configuration is absent. Existing MXFP8 output and quantization-config validation are unchanged. No changes to weight loading, refit, or kernels.

Tests

  • Four missing/null-config regression cases fail before the fix.
  • GB200, vLLM 0.25.1: 256 generation/refit tests pass with this fix on the integration branch, including all 14 HF-override tests.
  • Fresh 20-step BF16 rollout reruns are submitted; E2E results are pending.

@seonjinn seonjinn added the CI:L0 Run doctests and unit tests label Sep 12, 2026
@copy-pr-bot

copy-pr-bot Bot commented Sep 12, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@seonjinn

Copy link
Copy Markdown
Contributor Author

/ok to test c49e043

@seonjinn

Copy link
Copy Markdown
Contributor Author

Updated in d49af6c: explicit BF16 rollout with inherited quantization_ignore_patterns now emits a warning when no quantization configuration is active, then continues normally. Mixed MXFP8 rollout with BF16 first/last layers retains its effective-exclusion log and does not receive this warning. Added coverage for bf16/bfloat16 aliases, missing/None configuration, and mixed MXFP8. GB200 red/green regression validation has been submitted; the new warning tests are not yet claimed passing.

@seonjinn

Copy link
Copy Markdown
Contributor Author

GB200 validation completed (job 7096575, exit 0). With the new warning assertions and the old implementation, 8 regression cases failed as expected. With the warning implementation, the full selected suite passed: 267 tests. This includes BF16 aliases with absent/None quantization config and mixed MXFP8 exclusions without a BF16 warning. Validation used integration source 1802f7c with the same helper/test changes as d49af6c; this is not a claim of every end-to-end configuration passing.

@seonjinn seonjinn added CI:Lfast Runs a fast test suite and re-use nightly `main` container (but sync dependencies to PRs version) and removed CI:L0 Run doctests and unit tests labels Sep 17, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI:Lfast Runs a fast test suite and re-use nightly `main` container (but sync dependencies to PRs version)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant