Gemma4 ov int4 and stateful fixes - #260
Open
cavusmustafa wants to merge 11 commits into
Open
Conversation
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
from
July 21, 2026 19:26
530aa62 to
cbb8242
Compare
wine99
force-pushed
the
dev_backend_openvino
branch
from
July 27, 2026 04:33
4250de7 to
175f422
Compare
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
from
August 11, 2026 21:16
53d099b to
61c08f4
Compare
wine99
force-pushed
the
dev_backend_openvino
branch
2 times, most recently
from
August 14, 2026 05:24
37b164f to
2bacf9e
Compare
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
from
August 14, 2026 18:07
61c08f4 to
4acdec3
Compare
wine99
force-pushed
the
dev_backend_openvino
branch
from
August 18, 2026 01:01
678ba21 to
16f69ac
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
from
August 18, 2026 14:43
eb54a50 to
04b5691
Compare
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
3 times, most recently
from
August 22, 2026 00:46
8636129 to
c985875
Compare
cavusmustafa
marked this pull request as ready for review
August 22, 2026 01:10
ravi9
force-pushed
the
dev_backend_openvino
branch
from
August 24, 2026 16:49
8ba9348 to
d691121
Compare
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
2 times, most recently
from
August 24, 2026 22:36
4dfe10d to
28743bb
Compare
Assisted-by: Claude Sonnet
cavusmustafa
force-pushed
the
gemma4_ov_int4_and_stateful_fixes
branch
from
August 24, 2026 22:58
28743bb to
6f670df
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
from
August 25, 2026 16:53
d691121 to
407e209
Compare
cavusmustafa
added a commit
to cavusmustafa/llama.cpp
that referenced
this pull request
Aug 25, 2026
…state The stateful path seeds its KV state from ggml's cache when the decode position is ahead of what the state holds. That only works when ggml's cache is a plain prefix, where cell i holds position i. A sliding-window layer keeps just the last n_swa positions and drops the rest, so past the window cell i no longer holds position i and the seeded state is wrong. Slicing the state to the decode position also had no bounds check, so a position past the end surfaced as a bare ov::Exception from the ROI constructor (llama_decode ret = -3, with no reason given at default verbosity). Refuse both cases with a clear message instead, and refuse on the compile path too, where a new model starts with an empty state and so can only serve a sequence from its beginning. Reproducible with llama-bench -d, which restores a saved sequence state rather than recomputing the depth prefill. Assisted-by: Claude Opus 5
The stateful path reinterprets ggml's KV buffer [1, 1, seq, n_heads_kv * head_size] as [1, seq, n_heads_kv, head_size]. The head size is already taken from the tensor's own combined dim, because gemma-4 varies it per layer type, but the head count still came from a model-level scalar that compute_llm_params() overwrites per attention node, so it ended up holding whatever the last layer said. gemma-4 varies the head count per layer too: 12B has 8 x 256 sliding layers and 1 x 512 full layers, 31B has 16 x 256 and 4 x 512. So 40 of 12B's 48 layers were split as 1 x 2048 instead of 8 x 256, and attention read the state with the wrong head split - both models decoded garbage on CPU and GPU. E2B is unaffected, its head count is 1 everywhere. Record the count per layer instead and look it up by the cache_k_l<N> leaf name. Key it by layer, not by layer type: the sliding/full classification comes from cache extents, which tie at a small -c, while the head count does not. The stateful state trim now derives its sequence axis per state for the same reason, since pass::KVStateSeqAxis matches per state on the head count. Assisted-by: Claude Opus 5
pass::KVStateSeqAxis was limited to states with a single KV head, where moving the sequence axis from dim 1 to dim 2 is a pure metadata change. The limit was also based on a measurement showing no gain for a multi-head model, but that was taken at depth 0, which is the one depth where this change does nothing. With several heads the pass does more than move metadata: it drops the reader side transpose of the whole accumulated state, which the graph otherwise redoes every token at a cost that grows with the context length, and replaces it with a transpose of the single new row. Measured on GPU, tg128, alternating arms: gemma-4-12B 6.27 -> 9.11 t/s at depth 8192 (stateless is 7.69, so stateful now wins at depth instead of losing), Llama-3.2-1B 47.8 -> 59.6 t/s. Both are within noise at depth 0, which is why the earlier check saw nothing. The state refill needs the rows copied rather than reinterpreted now: ggml stores [seq][n_heads_kv * head_size], and a relayout state with several heads is a different element order. Without that, a refill would seed wrong data - it is reachable today through llama-bench -d. Assisted-by: Claude Opus 5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.