Skip to content

Gemma4 ov int4 and stateful fixes - #260

Open
cavusmustafa wants to merge 11 commits into
ravi9:dev_backend_openvinofrom
cavusmustafa:gemma4_ov_int4_and_stateful_fixes
Open

Gemma4 ov int4 and stateful fixes#260
cavusmustafa wants to merge 11 commits into
ravi9:dev_backend_openvinofrom
cavusmustafa:gemma4_ov_int4_and_stateful_fixes

Conversation

@cavusmustafa

@cavusmustafa cavusmustafa commented Jul 18, 2026

Copy link
Copy Markdown
Collaborator
  • enables gemma4 dense stateful support (limited support, up to 1k context)
  • optional int4 downscale for requant: sacrifice accuracy for performance (env to enable - GGML_OPENVINO_INT4_REQUANT)
  • optional disk spill when weight loading to reduce initial RSS peak for low memory systems (env to enable: GGML_OPENVINO_SPILL_DIR)

@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch from 530aa62 to cbb8242 Compare July 21, 2026 19:26
@wine99
wine99 force-pushed the dev_backend_openvino branch from 4250de7 to 175f422 Compare July 27, 2026 04:33
@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch from 53d099b to 61c08f4 Compare August 11, 2026 21:16
@wine99
wine99 force-pushed the dev_backend_openvino branch 2 times, most recently from 37b164f to 2bacf9e Compare August 14, 2026 05:24
@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch from 61c08f4 to 4acdec3 Compare August 14, 2026 18:07
@wine99
wine99 force-pushed the dev_backend_openvino branch from 678ba21 to 16f69ac Compare August 18, 2026 01:01
@ravi9
ravi9 force-pushed the dev_backend_openvino branch from eb54a50 to 04b5691 Compare August 18, 2026 14:43
@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch 3 times, most recently from 8636129 to c985875 Compare August 22, 2026 00:46
@cavusmustafa
cavusmustafa marked this pull request as ready for review August 22, 2026 01:10
@cavusmustafa
cavusmustafa requested a review from wine99 as a code owner August 22, 2026 01:10
@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch 2 times, most recently from 4dfe10d to 28743bb Compare August 24, 2026 22:36
@cavusmustafa
cavusmustafa force-pushed the gemma4_ov_int4_and_stateful_fixes branch from 28743bb to 6f670df Compare August 24, 2026 22:58
@ravi9
ravi9 force-pushed the dev_backend_openvino branch from d691121 to 407e209 Compare August 25, 2026 16:53
cavusmustafa added a commit to cavusmustafa/llama.cpp that referenced this pull request Aug 25, 2026
…uant target

Cherry-picked from PR ravi9#260 (21b5e13); the MoE fused path needs a symmetric or
integer zero-point expert layout.
…state

The stateful path seeds its KV state from ggml's cache when the decode position
is ahead of what the state holds. That only works when ggml's cache is a plain
prefix, where cell i holds position i. A sliding-window layer keeps just the last
n_swa positions and drops the rest, so past the window cell i no longer holds
position i and the seeded state is wrong.

Slicing the state to the decode position also had no bounds check, so a position
past the end surfaced as a bare ov::Exception from the ROI constructor
(llama_decode ret = -3, with no reason given at default verbosity).

Refuse both cases with a clear message instead, and refuse on the compile path
too, where a new model starts with an empty state and so can only serve a
sequence from its beginning. Reproducible with llama-bench -d, which restores a
saved sequence state rather than recomputing the depth prefill.

Assisted-by: Claude Opus 5
The stateful path reinterprets ggml's KV buffer [1, 1, seq, n_heads_kv * head_size]
as [1, seq, n_heads_kv, head_size]. The head size is already taken from the
tensor's own combined dim, because gemma-4 varies it per layer type, but the head
count still came from a model-level scalar that compute_llm_params() overwrites
per attention node, so it ended up holding whatever the last layer said.

gemma-4 varies the head count per layer too: 12B has 8 x 256 sliding layers and
1 x 512 full layers, 31B has 16 x 256 and 4 x 512. So 40 of 12B's 48 layers were
split as 1 x 2048 instead of 8 x 256, and attention read the state with the wrong
head split - both models decoded garbage on CPU and GPU. E2B is unaffected, its
head count is 1 everywhere.

Record the count per layer instead and look it up by the cache_k_l<N> leaf name.
Key it by layer, not by layer type: the sliding/full classification comes from
cache extents, which tie at a small -c, while the head count does not.

The stateful state trim now derives its sequence axis per state for the same
reason, since pass::KVStateSeqAxis matches per state on the head count.

Assisted-by: Claude Opus 5
pass::KVStateSeqAxis was limited to states with a single KV head, where moving
the sequence axis from dim 1 to dim 2 is a pure metadata change. The limit was
also based on a measurement showing no gain for a multi-head model, but that was
taken at depth 0, which is the one depth where this change does nothing.

With several heads the pass does more than move metadata: it drops the reader
side transpose of the whole accumulated state, which the graph otherwise redoes
every token at a cost that grows with the context length, and replaces it with a
transpose of the single new row. Measured on GPU, tg128, alternating arms:
gemma-4-12B 6.27 -> 9.11 t/s at depth 8192 (stateless is 7.69, so stateful now
wins at depth instead of losing), Llama-3.2-1B 47.8 -> 59.6 t/s. Both are within
noise at depth 0, which is why the earlier check saw nothing.

The state refill needs the rows copied rather than reinterpreted now: ggml stores
[seq][n_heads_kv * head_size], and a relayout state with several heads is a
different element order. Without that, a refill would seed wrong data - it is
reachable today through llama-bench -d.

Assisted-by: Claude Opus 5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants