Skip to content

feat(qwen3): architecture support - #162

Merged
john-rocky merged 1 commit into
mainfrom
feat/qwen3-architecture
Apr 30, 2026
Merged

john-rocky merged 1 commit into
mainfrom
feat/qwen3-architecture

Conversation

@john-rocky

Copy link
Copy Markdown
Owner

Summary

Adds Qwen3 (QK-norm, no attention bias, tied embed) to the conversion pipeline as a reusable architecture, ready for Qwen3-1.7B / 4B / 8B and any QK-normed Qwen-family model.

vs Qwen2:

  • per-head RMSNorm on Q and K applied before RoPE (QK-norm)
  • attention_bias = False on q/k/v projections
  • otherwise identical: GQA, SwiGLU MLP, RoPE, RMSNorm, tied word embed

Files

  • conversion/models/qwen3.py (new): Qwen3Model with the QK-norm weight-map and tied-embed loader
  • conversion/base_model.py: ModelConfig.has_qk_norm flag (default False), conditional q_norm/k_norm modules in ANEAttention, and a warn() when rope_scaling is set but YaRN isn't implemented (Qwen3 ships scaling configs we don't honor)
  • conversion/exporter.py: MonolithicWrapper applies q_norm/k_norm before RoPE when the layer has has_qk_norm set (uses getattr(..., False) so older architectures stay untouched)
  • conversion/convert.py: 'qwen3' architecture routes to Qwen3Model (the auto-detect helpers already returned 'qwen3'; only the loader was missing)

Off-default flag: existing Qwen2 / Gemma 3/4 / LFM2 builds are byte-for-byte unaffected.

Extracted from feat/qwen3-bonsai-investigation (commit 56ee545). The companion Bonsai post-mortem documentation will land in a separate docs PR.

Test plan

  • python conversion/convert.py --model-id Qwen/Qwen3-1.7B --output /tmp/qwen3-1.7b --quantize int4 produces a working monolithic .mlpackage
  • Existing Qwen2 / Gemma 3/4 / LFM2 conversions reproduce bit-identical artifacts (off-by-default flag)
  • Logits parity vs HF reference on a 256-token greedy decode (top-1 ≥ 99%)

…mbed)

Adds Qwen3 to the conversion pipeline as a reusable architecture, ready
for Qwen3-1.7B / 4B / 8B and any QK-normed Qwen-family model.

vs Qwen2 the differences are:
- per-head RMSNorm on Q and K applied before RoPE (QK-norm)
- attention_bias = False on q/k/v projections
- otherwise identical: GQA, SwiGLU MLP, RoPE, RMSNorm, tied word embed

What lands:
- conversion/models/qwen3.py: Qwen3Model (from_pretrained sets
  has_qk_norm=True and forces attention_bias=False; weight_map adds the
  q_norm.weight / k_norm.weight per layer)
- conversion/base_model.py: ModelConfig.has_qk_norm flag (default False
  so Qwen2 / Gemma builds are bit-identical), conditional q_norm/k_norm
  in ANEAttention, and a warn() when rope_scaling is set but YaRN isn't
  implemented (Qwen3 ships rope_scaling configs we don't honor)
- conversion/exporter.py: MonolithicWrapper applies q_norm/k_norm before
  RoPE when the layer has has_qk_norm set
- conversion/convert.py: 'qwen3' architecture routes to Qwen3Model
  (auto-detect functions already returned 'qwen3'; the loader was the
  remaining piece)

Off-default flag: existing Qwen2 / Gemma 3/4 / LFM2 builds are unaffected.

Extracted from feat/qwen3-bonsai-investigation (commit 56ee545). The
companion Bonsai post-mortem docs land in a separate PR.
@john-rocky
john-rocky force-pushed the feat/qwen3-architecture branch from 6f3b353 to f007344 Compare April 30, 2026 02:54
@john-rocky
john-rocky merged commit fe9b73b into main Apr 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant