Resolve continuous batching config against the text config for composite models - #48299
Merged
Conversation
…ite models PagedAttentionCache and the CB config resolution read decoder attributes (num_key_value_heads, num_hidden_layers, head_dim, ...) from the top-level config, which does not carry them for composite configs (vision-language models keep them on text_config). Continuous batching on any ForConditionalGeneration model crashed with 'num_key_value_heads or num_attention_heads could not be found in the config'. Resolve against config.get_text_config(), a no-op for text-only configs.
qgallouedec
force-pushed
the
fix-cb-vlm-text-config
branch
from
August 25, 2026 15:34
eb55b42 to
e3bd6f8
Compare
|
The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update. |
Contributor
CI recapDashboard: View test results in Grafana |
remi-or
approved these changes
Aug 26, 2026
remi-or
left a comment
Collaborator
There was a problem hiding this comment.
A bit drastic, but since we don't support multimodal models for now this works fine. Thanks!
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Aug 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes #48298. Part of #48294.
Change: the three call sites in
ContinuousBatchingManagerthat read decoder attributes (resolve_continuous_batching_config,PagedAttentionCache,ContinuousBatchProcessor) passself.model.config.get_text_config()instead ofself.model.config.get_text_config()returnsselffor text-only configs, so the existing paths are untouched.Validation: text-only
generate_batchon Qwen3-VL-4B-Instruct works with this change and crashes without it (Qwen3.5-style hybrids then hit their separate linear-attention limitation, tracked in its own issue).Who can review?
@remi-or (continuous batching)