feat(MODEL-MM-GLM53-FLASH): W7a — the glm5_next GGUF converter, gated byte-for-byte against llama.cpp b10451 (#2011) - #2017
Open
localai-bot wants to merge 5 commits into
Conversation
… every GGUF convention the converter needs (#2011) W7 as written was one wave that needed a GPU and a 300-600 GiB download, which made the converter -- the thing every GPU gate on this row is waiting for -- unreachable until somebody granted a large asset. The two halves have different blockers, so this splits them: W7a is the converter and its synthetic-fixture gate, CPU-only and needing no checkpoint at all; W7b is the artifact and keeps the download authority W7 always required. Issue #2011 owns W7a. Three findings move what a later wave should believe, and all three came from reading rather than guessing. **D6 was too strong.** llama.cpp still implements no `glm5_next` -- re-verified at `origin/master` `539f24529` fetched 2026-08-26 and at our pin: the enumerators are `LLM_ARCH_GLM4`, `LLM_ARCH_GLM4_MOE`, `LLM_ARCH_GLM_DSA` (`src/llama-arch.h:86-88`), and `src/models/glm-dsa.cpp` is GLM-5.2, citing `zai-org/GLM-5.2/blob/main/config.json`, a different model. But "no implementation" is not "no convention", and every convention this converter needs is present AT `b10451`: `class KDA` with `{arch}.kda.head_dim` and `{arch}.kda.gate_lower_bound` (`gguf-py/gguf/constants.py:262-264`), the KDA tensor spellings including the three separate `ssm_conv1d_q/k/v` this checkpoint's packing needs (`src/llama-arch.cpp:465-479`), the HF module paths that map onto them (`gguf-py/gguf/tensor_mapping.py:896-933`, Kimi-Linear's paths being GLM-5.3-Flash's verbatim), and the indexer names including the k-pool compressor (`:626-636`). `KDA.SAFE_GATE` is the only member that is `master`-only, and this model declares no `safe_gate`. **No pin advance is required by any part of W7, and none is owed.** **Upstream Python cannot quantize at all.** `gguf.quants.Q2_K` implements `dequantize_blocks` and no `quantize_blocks`; the encoders live only in `ggml/src/ggml-quants.c`. So W7a ports them and gates them byte-for-byte against the pinned C reference rather than against a tolerance, which is recorded here because "we could not use gguf-py" is otherwise the kind of claim a reader has to re-derive. **The arm arithmetic moved, and the table it replaces was the weaker kind.** The §Hardware table is bits-per-weight times a parameter count. The converter resolves a type per tensor, so running its own resolver over the real topology is the arithmetic that will be written: 1719 tensors carrying 313,890,512,702 parameters, which is 321.32B less the 7.43B MTP block and therefore an independent confirmation that the skip is exactly the 2.31% §Port map measured. The Q2_K arm is **100.35 GiB, not 102.6**, because the old table stated every figure including layer 45. Against ~119.63 GiB at 128K context and one sequence that leaves ~17.7 GiB after 1.43 GiB of KV and 0.14 GiB of KDA state. Still arithmetic and not measurement: #1963 and #1966 record this accounting being wrong by 48x, and W5 re-derives it from the runner. Three debts are added rather than waived. O7, no artifact exists and what producing one needs is named. O8, the Q3_K/Q4_K/Q5_K encoders are not ported and the converter refuses those arms. O9, the emitted file is not loadable by this tree because `glm5next` has no `general.architecture` dispatch entry -- that wiring is W1's, and it is written down because W7a lands a capability a production entry point does not yet reach. Spec and records only. The converter is not in this commit. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Claude:claude-fable-5 [Claude Code]
… byte-for-byte against llama.cpp b10451 (#2011) Every GPU gate on this row is blocked behind an artifact that does not exist. Measured live 2026-08-26 against ~119.63 GiB usable on `dgx:gpu0`: FP8 305.78 GiB, BF16 598.53, the smallest published NVFP4 181.32, and all four repositories named `*-GGUF` contain zero `.gguf` files. No upstream tool can make one, so this authors it. `scripts/convert-glm5-next-gguf.py` reads a safetensors checkpoint and writes a GGUF at arch `glm5next`. It streams -- headers for the plan, `mmap` slices for the data -- so peak resident memory is one tensor rather than one shard, which is what makes a 305.78 GiB source tractable at all. FP8 e4m3 is decoded against the `weight_scale_inv` grid the checkpoint declares (`weight_block_size: [128, 128]`); per-expert tensors are stacked into `ffn_{gate,up,down}_exps`; the layer-45 MTP block and every `shared_head.*` are dropped, following the reference's own `_keys_to_ignore_on_load_unexpected` and the `glm4_moe_lite_registry.cpp:21-26` precedent. The metadata carries the parameters the port hinges on, spelled at the pin. `glm5next.kda.gate_lower_bound` is the load-bearing one: -5.0 selects `-bound * sigmoid(exp(A_log) * (g + dt_bias))`, a different function from the `-exp(A_log) * softplus(g + dt_bias)` our Kimi-Linear KDA implements, and even the sign of `decay_rate` differs. Writing it into the file is what lets a loader take the right branch instead of inheriting Kimi's. Alongside it: the indexer geometry with `index_kpool` = 4 where the config class default is 16, the mHC triple keeping `hc_eps` 1e-6 distinct from `rms_norm_eps` 1e-5, and `layer_types` as the authoritative schedule -- the reference ignores `linear_attn_config.kda_layers` entirely, so only one of those two lists may travel. No `rope.freq_base` is written; the text stack is NoPE end to end. **Q2_K, Q6_K and Q8_0 are ported from `ggml/src/ggml-quants.c` at our pin `b10451` and are byte-identical to it.** They had to be ported rather than called: `gguf.quants.Q2_K` upstream implements `dequantize_blocks` and no `quantize_blocks`, so no upstream Python can produce a k-quant. The gate is a frozen golden captured from that reference compiled `-ffp-contract=off`, and it is bytes rather than a tolerance because an encoder that is close but not exact writes a file that loads, generates fluent text, and is quietly worse than the arm it claims to be. Two traps changed bytes during the port and are recorded in the source so nobody re-finds them: `nearest_int` is the `+12582912.0` add-and-mask trick at `:621` and rounds half to EVEN, not `round`; and C `roundf` in `quantize_row_q8_0_ref` rounds half AWAY FROM ZERO where `np.rint` rounds half to even. The second was caught by a crafted `[0.5, -0.5, 1.5, ...]` case: reference `[1,-1,2,-2,3,-3]`, `np.rint` `[0,0,2,-2,2,-2]`. The arms, from the converter's own per-tensor plan over the real topology (1719 tensors, 313.89B parameters after the MTP block is dropped from 321.32B): q2_k 100.35 GiB at 2.746 mixed bpw, q6_k 239.89, q8_0 310.67, bf16 584.67. Only q2_k fits. Experts are 97% of the model, so the arm name is the expert type and the other 3% rides at Q6_K almost for free. The fallback ladder is `want -> Q8_0 -> F32` rather than `want -> F32`, and that is not tidiness: with an F32 fallback a finer arm can come out LARGER than a coarser one, which is not a property a size table may have. Unimplemented arms are refused by name with the missing part. The i-quants are refused because an imatrix needs a forward pass, a forward pass needs 181 GiB, and the dependency is circular on this fleet -- a boundary, not a to-do. Q3_K/Q4_K/Q5_K are refused because their encoders are unported and shipping an ungated encoder is worse than refusing. So are `--keep-mtp`, a non-`glm5_next` config, and an FP8 tensor whose `weight_scale_inv` companion is missing, which would otherwise write a loadable, wrong file. `tests/scripts/test_convert_glm5_next_gguf.py` is 50 assertions over a SYNTHETIC tiny-shape checkpoint and an independent in-test GGUF reader that shares no code with the writer. No checkpoint, no GPU, no C++ build. RED before this commit was `FAIL scripts/convert-glm5-next-gguf.py is absent`, rc=1. **NOT REACHED, and this is the part to read before reviewing.** `glm5next` has no entry in the `general.architecture` dispatch (`src/vllm/entrypoints/model_loader.cpp:1000`), so the file this converter writes is not loadable by this tree. That wiring is W1's, owned by row `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` and tracked by #1998; the spec lists it under `## Owed` as O9. The converter itself IS reached, as a command-line path, and the gate enters through it as a user does. No artifact has been produced against the real checkpoint either -- that needs staged weights, disk and a box, and is owed as O7. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Claude:claude-fable-5 [Claude Code]
…ding rules it claims to hold (#2011) Mutation found it, which is what mutation is for. Two guarantees the converter states in its own comments -- that `nearest_int` rounds half to EVEN and that C `roundf` rounds half AWAY FROM ZERO -- were not held by the gate. Replacing either rule left the golden green: M1 nearest_int half-to-EVEN -> half-UP gate rc 0, 0 failures GATE BLIND M2 roundf -> np.rint gate rc 0, 0 failures GATE BLIND The cause is the inputs, not the assertion. The golden was captured over weight-like random data, heavy-tailed data and zeros, and none of those ever lands a value exactly on a rounding tie, so the two rules agree everywhere the fixture looks. A mis-encoded `.5` is not a hypothetical here: it is how the `np.rint` defect was found in the first place, by a hand-crafted case that the frozen fixture then failed to preserve. Three super-blocks are added and the golden is recaptured from the same pinned reference. Blocks 3 and 4 were SEARCHED on an eighth-lattice until half-to-even and half-up disagree under Q2_K and under Q6_K respectively -- they are not hand-derived, they are the first blocks found that discriminate. Block 5 is eight Q8_0 sub-blocks whose `amax` is exactly 127, so `x * id` is 0.5, -0.5, 1.5, -1.5, 2.5, -2.5 and the reference emits [1,-1,2,-2,3,-3] where `np.rint` emits [0,0,2,-2,2,-2]. Both mutations now fail, and the other eight in the table were already caught. Also corrects the `.agents/model-matrix.md` rollup, which the row's move to `ACTIVE` invalidated: ACTIVE 10 -> 11, READY 5 -> 4. `scripts/check-model-checklist.py` reports OK. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Claude:claude-fable-5 [Claude Code]
The base branch `row/MODEL-MM-GLM53-FLASH` (pull request #2001) took a merge of `origin/main` while W7a was in flight, so this branch was two commits behind its own base. Merging rather than rebasing keeps #2001's spec commit as the shared merge base of both branches; a rebase would move it under the open pull request. No conflict. The full gate is rerun on this merge commit, not on the pre-merge head, because the base carried #1999 (`KV-GDN-STATE-BUDGET`) and a converter gate that never saw it would be measuring the wrong tree. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Claude:claude-fable-5 [Claude Code]
… and repair the two anchors that edit shifted (#2011) `test_glm5_next_row_is_inside_the_model_ratchet` pins the row's lifecycle state as a literal, and its own docstring gives the premise: `READY` because the spec was committed and no product code had landed. W7a landed product code, so the premise expired and the pin moves with it, in the same change that moves the matrix row. The assertion is not weakened -- it still names one exact state, and a pin that followed the row automatically would assert nothing. What it stops catching is the one transition it was updated for; a rename, a second glm5_next row, and any later state change made without touching this file still red it. Editing that docstring added nine lines to `tests/scripts/test_agent_record.py` and shifted `RecordAnchorRatchet` from :1539 to :1548 and `test_one_good_link_does_not_cover_a_rotted_bare_citation` from :1607 to :1616. `ENG-RECORD-ANCHOR-RATCHET` cites both, so the anchor ratchet went 31 -> 33 and the gate reported a regression. The anchors are repaired rather than the baseline raised, which is what the ratchet's own message demands. It is also the failure that row exists to measure, produced the way that row's record says it gets produced: by an edit to the very file the row cites, inside the pull request that makes it. `scripts/check-agent-record.py --report` now reports stale=31 broken=6 rot=37, unchanged from the base, and `tests/scripts/test_agent_record.py` runs 118 tests OK. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Claude:claude-fable-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
W7a of
MODEL-MM-GLM53-FLASH: the safetensors→GGUF converter forglm5_next, and the gate that holds its k-quant encoders to the byte.Every GPU gate on this row is blocked behind an artifact that does not exist. Measured live 2026-08-26 against ~119.63 GiB usable on
dgx:gpu0: FP8 305.78 GiB, BF16 598.53, the smallest published NVFP4 181.32, and all four repositories named*-GGUFcontain zero.gguffiles. No upstream tool can make one — llama.cpp has noglm5_nextat any revision — so this authors it.What lands
scripts/convert-glm5-next-gguf.pyreads a safetensors checkpoint and writes a GGUF at archglm5next. It streams: headers for the plan,mmapslices for the data, so peak resident memory is one tensor rather than one shard, which is what makes a 305.78 GiB source tractable. FP8 e4m3 is decoded against theweight_scale_invgrid the checkpoint declares (weight_block_size: [128, 128]); per-expert tensors are stacked intoffn_{gate,up,down}_exps; the layer-45 MTP block and everyshared_head.*are dropped, following the reference's own_keys_to_ignore_on_load_unexpectedand theglm4_moe_lite_registry.cpp:21-26precedent.The metadata carries the parameters the port hinges on.
glm5next.kda.gate_lower_boundis the load-bearing one: -5.0 selects-bound * sigmoid(exp(A_log) * (g + dt_bias)), a different function from the-exp(A_log) * softplus(g + dt_bias)our Kimi-Linear KDA implements, and even the sign ofdecay_ratediffers. Writing it into the file is what lets a loader take the right branch instead of inheriting Kimi's. Beside it: the indexer geometry withindex_kpool= 4 where the config class default is 16, the mHC triple keepinghc_eps1e-6 distinct fromrms_norm_eps1e-5, andlayer_typesas the authoritative schedule — the reference ignoreslinear_attn_config.kda_layersentirely, so only one of those two lists may travel. Norope.freq_baseis written; the text stack is NoPE end to end.The encoders are byte-identical to the pinned reference, and that is the point
gguf.quants.Q2_Kupstream implementsdequantize_blocksand noquantize_blocks. Upstream Python cannot produce a k-quant at all, so Q2_K, Q6_K and Q8_0 are ported fromggml/src/ggml-quants.cat our pinb10451(:891,:1869,:276, overmake_qkx2_quants:799,make_qx_quants:628,nearest_int:621) and gated against a frozen golden captured from that reference compiled-ffp-contract=off. Bytes rather than a tolerance, because an encoder that is close but not exact writes a file that loads, generates fluent text, and is quietly worse than the arm it claims to be.Two traps changed bytes during the port and are recorded in the source so nobody re-finds them:
nearest_intis the+12582912.0add-and-mask trick and rounds half to even, notround.roundfinquantize_row_q8_0_refrounds half away from zero wherenp.rintrounds half to even. Caught by a crafted[0.5, -0.5, 1.5, -1.5, 2.5, -2.5]case: reference[1,-1,2,-2,3,-3],np.rint[0,0,2,-2,2,-2].The pin already had everything, so nothing advances
D6said llama.cpp has noglm5_next, and it still does not — re-verified atorigin/master539f24529(fetched 2026-08-26) and at the pin: the enumerators areLLM_ARCH_GLM4,LLM_ARCH_GLM4_MOE,LLM_ARCH_GLM_DSA(src/llama-arch.h:86-88), andsrc/models/glm-dsa.cppis GLM-5.2, citingzai-org/GLM-5.2/blob/main/config.json, a different model. But "no implementation" is not "no convention", and every convention this converter needs is present atb10451:class KDAwith{arch}.kda.head_dimand{arch}.kda.gate_lower_bound(gguf-py/gguf/constants.py:262-264); the KDA tensor spellings including the three separatessm_conv1d_q/k/vthis checkpoint's packing needs (src/llama-arch.cpp:465-479); the HF module paths that map onto them (gguf-py/gguf/tensor_mapping.py:896-933, Kimi-Linear's paths being GLM-5.3-Flash's verbatim); and the indexer names including the k-pool compressor (:626-636).KDA.SAFE_GATEis the onlymaster-only member and this model declares nosafe_gate. No pin advance was taken and none is owed.The arithmetic, from the converter's own plan
Not bits-per-weight times a parameter count — the converter resolves a type per tensor, so its plan over the real topology is what will actually be written: 1719 tensors carrying 313,890,512,702 parameters, which is 321.32B less the 7.43B MTP block and therefore an independent confirmation that the skip is exactly the 2.31% the spec measured.
q2_k(experts Q2_K, rest Q6_K)q6_kq8_0bf16Against ~119.63 GiB at 128K context and one sequence: KV 1.43 GiB (11 MLA layers ×
kv_lora_rank512 × 2 B, plus an indexer side cache ofindex_head_dim128 × 2 B /index_kpool4 = 64 B per layer × 11, so 11,968 B/token), KDA recurrent state 0.14 GiB (64 heads × 128 × 128 × 4 B × 34 layers) plus conv states, leaving ~17.7 GiB. The Q2_K figure is 100.35 and not the spec's 102.6 because that table stated every arm including layer 45. Arithmetic, not measurement: #1963 and #1966 record this accounting being wrong by 48x, and W5 re-derives it from the runner.The fallback ladder is
want → Q8_0 → F32rather thanwant → F32, and that is not tidiness: with an F32 fallback a finer arm can come out larger than a coarser one, which is not a property a size table may have. The gate asserts the ordering.Refusals
Every unimplemented arm is refused by name with the missing part. The i-quants because an imatrix needs a forward pass, a forward pass needs 181 GiB, and the dependency is circular on this fleet — a boundary, not a to-do.
q3_k/q4_k/q5_kbecause their encoders are unported and shipping an ungated encoder is worse than refusing. So are--keep-mtp, a non-glm5_nextconfig, and an FP8 tensor whoseweight_scale_invcompanion is missing, which would otherwise write a loadable, wrong file.Evidence
RED, with the converter absent:
GREEN:
python3 tests/scripts/test_convert_glm5_next_gguf.py→ 55 assertions,All cases passed., rc=0. The suite builds a SYNTHETIC tiny-shape checkpoint and parses the result with an independent in-test GGUF reader that shares no code with the writer. No checkpoint, no GPU, no C++ build — noninjatarget was built by this change and no build directory was touched, which is deliberate given the disk state on this box.Two defects the gate caught during development, both of the write-a-loadable-wrong-file kind: the
np.rintrounding above, and a 5-axis tensor.model.visual.patch_embed.proj.weightis a Conv3d at[1024, 3, 2, 14, 14]and ggml carries at mostGGML_MAX_DIMS = 4, so writing it verbatim produces a header no reader can index. 4-D and 5-D convolution kernels are now flattened to the im2col form[out, prod(rest)]; 3-D shapes are left alone because they mean something in ggml (the depthwise conv is[ch, 1, k], the expert lane is[experts, n, m]). A hard refusal backs it up.Mutation
Every guarantee claimed above was broken in a scratch copy, the break was verified to have LANDED, the gate was rerun, the file was restored and the restore was verified by SHA-256, and the gate was rerun again.
__pycache__is cleared before every run, because a restored file can still execute the mutant's bytecode. Noninjatarget exists in this change, so there is no build step to report an exit code for: the harness runs a Python gate directly,ninjawas never invoked, and no build directory was created or touched.nearest_inthalf-to-EVEN → half-UProundf→np.rintglm5next.kda.gate_lower_boundgate_lower_boundis -5.0weight_scale_inv--arm iq2_xxsis refused and names the imatrixlayer_typeslayer_typesis the authoritative scheduleBaseline rc=0/0 failures; post-restore rc=0/0 failures; both file hashes match.
The first run of this table found a hole, and the third commit closes it. M1 and M2 came back
gate rc 0, 0 failures— GATE BLIND. The assertions were right and the inputs were wrong: the golden was captured over weight-like random data, heavy-tailed data and zeros, none of which ever lands a value exactly on a rounding tie, so half-to-even and half-up agree everywhere the fixture looked. Three super-blocks were added and the golden recaptured from the same pinned reference: blocks 3 and 4 were searched on an eighth-lattice until the two rules disagree under Q2_K and Q6_K, and block 5 is eight Q8_0 sub-blocks withamaxexactly 127 sox * idis±0.5, ±1.5, ±2.5. The table above is the rerun.M9's first form was also a non-mutation: adding
iq2_xxstoARMSchanges nothing whileREFUSED_ARMSis checked first. It is reported here because "a mutation that never applied reads as a passing test", and the fix was to mutate the refusal itself.Not reached — read this before reviewing
glm5nexthas no entry in thegeneral.architecturedispatch (src/vllm/entrypoints/model_loader.cpp:1000), so the file this converter writes is not loadable by this tree. That wiring is W1's, owned by rowMODEL-MM-glm5-next-glm5-next-for-conditional-generationand tracked by #1998; the spec lists it under## Owedas O9. The converter itself IS reached, as a command-line path, and the gate enters through it as a user does.No artifact has been produced against the real checkpoint either: that needs the 300–600 GiB weights staged on local disk, explicit developer authority for the download, and a box with room for source and output at once. Owed as O7; W7b owns it. The Q3_K/Q4_K/Q5_K encoders are owed as O8.
Gate
scripts/agent-preflight.shon the merge commit: 1 gate failed,test_cpu_x86_llamacpp_floor, and it is not attributable. Its failure is #618/#529 verbatim —AssertionError: 4 != 2and, on the isolated re-run,NO_QUIET_WINDOW after 30s (busy=121% builders=0 load=46.91 66.29 65.22). The box carried loadavg 25–85 from other sessions throughout. This branch changes no.cpp,.cuor.hfile, nothing underbenchmarks/, and not the harness itself — the whole diff is one Python converter, its Python gate, a JSON fixture, records and docs. Re-run and reported rather than waved through.Everything else is green, including
test_convert_glm5_next_gguf,check-agent-record(stale=31 broken=6 rot=37, unchanged from the base),check-model-checklist,audit-live-rows,issue-index append-only,commit-trailersandcommit-style.One process note worth recording, because it produced a red that was not one: the worktree was deleted underneath a running preflight by another session's reaper, and the run reported 74 failing gates of the form
python3: can't open file 'scripts/check-issue-index-append-only.py': No such file or directory. That signature is a moved tree, not a defect. The branch survived in the shared repository, the worktree was recreated at the same commit, and the verdict above is from the clean re-run.Closes #2011.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Claude:claude-fable-5 [Claude Code]