Repository navigation
feat(qwen3): architecture support - #162
Merged
Merged
Conversation
2 tasks
…mbed) Adds Qwen3 to the conversion pipeline as a reusable architecture, ready for Qwen3-1.7B / 4B / 8B and any QK-normed Qwen-family model. vs Qwen2 the differences are: - per-head RMSNorm on Q and K applied before RoPE (QK-norm) - attention_bias = False on q/k/v projections - otherwise identical: GQA, SwiGLU MLP, RoPE, RMSNorm, tied word embed What lands: - conversion/models/qwen3.py: Qwen3Model (from_pretrained sets has_qk_norm=True and forces attention_bias=False; weight_map adds the q_norm.weight / k_norm.weight per layer) - conversion/base_model.py: ModelConfig.has_qk_norm flag (default False so Qwen2 / Gemma builds are bit-identical), conditional q_norm/k_norm in ANEAttention, and a warn() when rope_scaling is set but YaRN isn't implemented (Qwen3 ships rope_scaling configs we don't honor) - conversion/exporter.py: MonolithicWrapper applies q_norm/k_norm before RoPE when the layer has has_qk_norm set - conversion/convert.py: 'qwen3' architecture routes to Qwen3Model (auto-detect functions already returned 'qwen3'; the loader was the remaining piece) Off-default flag: existing Qwen2 / Gemma 3/4 / LFM2 builds are unaffected. Extracted from feat/qwen3-bonsai-investigation (commit 56ee545). The companion Bonsai post-mortem docs land in a separate PR.
john-rocky
force-pushed
the
feat/qwen3-architecture
branch
from
April 30, 2026 02:54
6f3b353 to
f007344
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds Qwen3 (QK-norm, no attention bias, tied embed) to the conversion pipeline as a reusable architecture, ready for Qwen3-1.7B / 4B / 8B and any QK-normed Qwen-family model.
vs Qwen2:
attention_bias = Falseon q/k/v projectionsFiles
conversion/models/qwen3.py(new):Qwen3Modelwith the QK-norm weight-map and tied-embed loaderconversion/base_model.py:ModelConfig.has_qk_normflag (defaultFalse), conditionalq_norm/k_normmodules inANEAttention, and awarn()whenrope_scalingis set but YaRN isn't implemented (Qwen3 ships scaling configs we don't honor)conversion/exporter.py:MonolithicWrapperappliesq_norm/k_normbefore RoPE when the layer hashas_qk_normset (usesgetattr(..., False)so older architectures stay untouched)conversion/convert.py:'qwen3'architecture routes toQwen3Model(the auto-detect helpers already returned'qwen3'; only the loader was missing)Off-default flag: existing Qwen2 / Gemma 3/4 / LFM2 builds are byte-for-byte unaffected.
Extracted from
feat/qwen3-bonsai-investigation(commit 56ee545). The companion Bonsai post-mortem documentation will land in a separate docs PR.Test plan
python conversion/convert.py --model-id Qwen/Qwen3-1.7B --output /tmp/qwen3-1.7b --quantize int4produces a working monolithic .mlpackage