ggml-org#27206
ggml-openvino: split-boundary tensors break compute on hybrid MoE models (three defects)
Target repo: ggml-org/llama.cpp (OpenVINO backend, ggml/src/ggml-openvino/)
Status: DRAFT (2026-08-16, TASK-30). gh is NOT authenticated on this machine —
filing needs Sander's GitHub. Local fix verified in the kernel project's
patches/ (llamacpp-openvino-split-shape-fixes.patch, applied on top of
llamacpp-openvino-deterministic-tensor-names.patch at commit 885c5bb).
Related existing issues (searched 2026-08-16)
Title
ggml-openvino: hybrid-attention MoE models (qwen35moe, e.g. Ornith 1.0 35B) fail to compute — three defects in split-boundary tensor handling (reshape view-src extra sharing, boundary view/reshape outputs never emitted, shape-blind compiled-model cache)
Symptom chain (one model, three defects in the same split path)
Model: qwen35moe 35B (Ornith 1.0 35B, 20G GGUF, 40 layers, 256 experts,
GatedDeltaNet conv-state layers), -ngl 99 -dev OPENVINO0, build 10428
(commit 885c5bb), OpenVINO 2026.4.0-git, intel_gpu plugin, iGPU sharing
system DRAM. ornith:9b (same architecture) computes fine — the 35B graph gets
split across OV/CPU subgraphs (MoE experts run on CPU via q4_K_8x8 repack) and
only the split path is affected.
Without any fix, the first warmup compute fails at input binding:
[GPU] The tensor size is not equal to model, can't set input tensor with index: 7,
because model input (shape=[1,2,8,1]) and tensor (shape=[1,1,2,8]) are incompatible
After fixing that, the compute proceeds further and aborts in the CPU MoE
repack with garbage expert ids; after fixing that, the compiled-model cache
reuses a wrong-shape blob. Each defect is described below.
Defect 1 — RESHAPE results share the source tensor's ov::Tensor (wrong shape at split boundary)
ggml_backend_openvino_buffer_init_tensor (ggml-openvino.cpp) shares
view_src->extra for ANY tensor with view_src != nullptr, including RESHAPE
results (ggml_new_tensor_impl sets view_src for reshape results too).
When a RESHAPE result crosses a graph-split boundary and becomes a model
input, the input tensor's extra is the SOURCE tensor's ov::Tensor (source
shape), so binding rejects the model's baked (reshape) shape:
- Model input
ffn_moe_weights_norm-0 (reshaped): baked [1,2,8,1] (OV order,
ggml ne=[1,8,2,1]); runtime tensor bound via the shared extra: [1,1,2,8]
(ggml ne=[8,2,1,1] — the 2D div result [8,N]).
Fix: only share the extra when tensor->op == GGML_OP_VIEW; RESHAPE falls
through and gets a fresh extra with its own shape (data is a memory alias, so
no copy cost). One-line change, no behavior change for real views.
Defect 2 — split-boundary VIEW/RESHAPE results are never emitted as model outputs
GgmlOvDecoder::compute_model_outputs (ggml-decoder.cpp) skips VIEW/RESHAPE
nodes unconditionally, and suppresses parents whose only consumers are inside
the graph (input_use_count == use_count). For a scheduler split graph whose
boundary tensor IS a view/reshape (e.g. the MoE top-k view
ffn_moe_topk-N = view of the argsort output ffn_moe_argsort-N, consumed by
a CPU-backend MUL_MAT_ID), the data is therefore never written to the ggml
buffer. The consumer split copies uninitialized/stale buffer memory:
GGML_ASSERT(i02 >= 0 && i02 < n_as) failed (ggml-cpu/repack.cpp, MUL_MAT_ID ids)
ids_name=CPU#ffn_moe_topk-1#0 src0_name=blk.1.ffn_gate_exps.weight
first_ids: -1073594368 -1129906176 ... (stale F16 data read as int32)
Fix: for VIEW/RESHAPE nodes whose data is consumed outside the split
(use_count > uses inside the graph — use_counts are shared with the full
graph), emit the PARENT tensor (view_src, resolved through nested views) as a
model output. The plugin writes the parent's full data into the shared buffer;
consumers of the view read the same memory. Whole-model (non-split) graphs are
unaffected (no external consumers).
Defect 3 — compiled-model cache ignores input shapes; split models are compiled static
Split-graph boundary inputs are compiled with the concrete token count baked
in (get_graph_input_shape only keeps the tracked dynamic dim when
m_model_is_splitted, which is only set with the GGML_OPENVINO_ENABLE_FALLBACK
env). The frontend compiled-model cache key
(ggml_openvino_model_fingerprint, model-cache.cpp) hashes topology + weights
- config but NOT shapes, so a blob compiled for n_tokens=1 is imported for an
n_tokens=2 decode and the static shape rejects the bind:
BIND-FAIL input[1] name=attn_post_norm-2 runtime_shape=[1,1,2,2048] model_shape=[1,1,1,2048]
Fix: fold the shapes of the statically-compiled input leafs into the
fingerprint (tokens/positions/masks stay dynamic and are excluded; weights are
already hashed; caches/scalars have constant shapes). Whole-model graphs keep
a single reusable blob; split graphs get one blob per token count. Cache
format version bumped (v2 → v5 in the local patch: v3 added the shape scan,
v4 switched to a graph-input scan because split subgraphs have an empty leafs
array, v5 resolves VIEW chains like compute_model_inputs()) so stale blobs MISS.
Note: making the boundary inputs actually dynamic instead (removing the
m_model_is_splitted gate) was also tried — the intel_gpu plugin then
segfaults in the IGC kernel JIT (SIGSEGV in libopencl-clang.so.17 during
compile_model, first conv-state split). That may be a separate plugin bug
worth its own report.
Environment
- llama.cpp build 10428 (885c5bb), OpenVINO 2026.4.0-git, intel_gpu plugin,
X1 Carbon Gen 14 (PTL-H484), 32G RAM, iGPU shares system DRAM.
- GatedDeltaNet conv-state warnings (
cannot determine dynamic dim for RESHAPE node conv_states_reshaped-N etc.) appear for BOTH the working 9B and the
broken 35B and are NOT the trigger — they are noise; the trigger is the
split-boundary handling above.
- Cache-independent: Defects 1 and 2 reproduce without the cache env.
Local fix
patches/llamacpp-openvino-split-shape-fixes.patch (kernel project, applied on
top of the TASK-28 deterministic-tensor-names patch):
- ggml-openvino.cpp: extra sharing only for GGML_OP_VIEW.
- ggml-decoder.cpp: boundary view/reshape parents emitted as model outputs.
- model-cache.cpp: fingerprint includes static input leaf shapes.
- utils.cpp: cache format version v5; input-binding debug/diagnostics guarded
by env GGML_OPENVINO_DEBUG_INPUT (shape-only print, no tensor data access).
Verification — 2026-08-16, TASK-30 wrap-up (end-to-end, fixed binary)
Fixed binary built at 885c5bb + the two patches; scoped cgroup
(MemoryHigh=29G/MemoryMax=55G/SwapMax=16G, CPUQuota=400%), -lv 5,
-ngl 99 -dev OPENVINO0 --no-cache-prompt, env GGML_OPENVINO_DEVICE=GPU
GGML_OPENVINO_MEMORY_OPTIMIZE=1, model = Ornith 1.0 35B (20G GGUF), -c 16384.
- Cold compile (21:21 run): completes — 574 blobs written — and the first
warmup compute SUCCEEDS. Full 8-token completion: stopped by limit, n_gen = 8, n_predict = 8 (tokens 2, <think>, Here, 's, a,
thinking...). No ov::Exception, no Compute error, no repack assert
(grep count 0).
- Warm load (evening acceptance run): /health ok in 37s; 89 blob HITs;
direct curl /completion → HTTP 200
{"content":"2\n\n<think>\nHere's a thinking","stop":true, "tokens_predicted":8,"tokens_evaluated":4} and the server log confirms
stopped by limit, n_gen = 8, n_predict = 8.
- Memory profile (common_memory_breakdown_print, OPENVINO0): model
1452 MiB retained + context 571 + compute 60 on the 31586 MiB device pool;
41/41 layers offloaded to GPU. No runaway.
- Regression check: ornith:9b and qwen3.5:4b still compute fine.
- Patch hygiene: both patches
git apply --check cleanly (exit 0) on a
fresh clone at 885c5bb; the combined diff is byte-identical to the
verified build workdir.
Non-blocking observations for maintainers:
- A minority of blob imports fail at plugin compile with
could not create a primitive descriptor for the reduction primitive (intel_gpu plugin,
plugin.cpp:80) where the fresh-compile path succeeds — backend falls back to
recompile (deterministic per blob; warm loads still ~37s to /health).
- Per-shape blob fan-out: split graphs compile static shapes, so each
distinct token count exercised gets its own blob (by design of defect-3's
fix).
ggml-org#27206
ggml-openvino: split-boundary tensors break compute on hybrid MoE models (three defects)
Target repo: ggml-org/llama.cpp (OpenVINO backend,
ggml/src/ggml-openvino/)Status: DRAFT (2026-08-16, TASK-30). gh is NOT authenticated on this machine —
filing needs Sander's GitHub. Local fix verified in the kernel project's
patches/ (llamacpp-openvino-split-shape-fixes.patch, applied on top of
llamacpp-openvino-deterministic-tensor-names.patch at commit 885c5bb).
Related existing issues (searched 2026-08-16)
[GPU] The tensor size is not equal to modelbind rejection forpassthrough VIEW outputs (
create_ov_output_tensorop_case=0). Thisreport covers the split-boundary RESHAPE-input case (defect 1), the
boundary-view-output case (defect 2 — distinct code path: outputs never
EMITTED, not wrong shape) and the shape-blind cache (defect 3). Verified
fixes for all three (below).
backend/error detail; cannot be matched to these defects.
Title
ggml-openvino: hybrid-attention MoE models (qwen35moe, e.g. Ornith 1.0 35B) fail to compute — three defects in split-boundary tensor handling (reshape view-src extra sharing, boundary view/reshape outputs never emitted, shape-blind compiled-model cache)
Symptom chain (one model, three defects in the same split path)
Model:
qwen35moe35B (Ornith 1.0 35B, 20G GGUF, 40 layers, 256 experts,GatedDeltaNet conv-state layers),
-ngl 99 -dev OPENVINO0, build 10428(commit 885c5bb), OpenVINO 2026.4.0-git, intel_gpu plugin, iGPU sharing
system DRAM. ornith:9b (same architecture) computes fine — the 35B graph gets
split across OV/CPU subgraphs (MoE experts run on CPU via q4_K_8x8 repack) and
only the split path is affected.
Without any fix, the first warmup compute fails at input binding:
After fixing that, the compute proceeds further and aborts in the CPU MoE
repack with garbage expert ids; after fixing that, the compiled-model cache
reuses a wrong-shape blob. Each defect is described below.
Defect 1 — RESHAPE results share the source tensor's ov::Tensor (wrong shape at split boundary)
ggml_backend_openvino_buffer_init_tensor(ggml-openvino.cpp) sharesview_src->extrafor ANY tensor withview_src != nullptr, including RESHAPEresults (
ggml_new_tensor_implsetsview_srcfor reshape results too).When a RESHAPE result crosses a graph-split boundary and becomes a model
input, the input tensor's extra is the SOURCE tensor's ov::Tensor (source
shape), so binding rejects the model's baked (reshape) shape:
ffn_moe_weights_norm-0 (reshaped): baked[1,2,8,1](OV order,ggml ne=[1,8,2,1]); runtime tensor bound via the shared extra:
[1,1,2,8](ggml ne=[8,2,1,1] — the 2D div result
[8,N]).Fix: only share the extra when
tensor->op == GGML_OP_VIEW; RESHAPE fallsthrough and gets a fresh extra with its own shape (data is a memory alias, so
no copy cost). One-line change, no behavior change for real views.
Defect 2 — split-boundary VIEW/RESHAPE results are never emitted as model outputs
GgmlOvDecoder::compute_model_outputs(ggml-decoder.cpp) skips VIEW/RESHAPEnodes unconditionally, and suppresses parents whose only consumers are inside
the graph (
input_use_count == use_count). For a scheduler split graph whoseboundary tensor IS a view/reshape (e.g. the MoE top-k view
ffn_moe_topk-N= view of the argsort outputffn_moe_argsort-N, consumed bya CPU-backend MUL_MAT_ID), the data is therefore never written to the ggml
buffer. The consumer split copies uninitialized/stale buffer memory:
Fix: for VIEW/RESHAPE nodes whose data is consumed outside the split
(use_count > uses inside the graph — use_counts are shared with the full
graph), emit the PARENT tensor (view_src, resolved through nested views) as a
model output. The plugin writes the parent's full data into the shared buffer;
consumers of the view read the same memory. Whole-model (non-split) graphs are
unaffected (no external consumers).
Defect 3 — compiled-model cache ignores input shapes; split models are compiled static
Split-graph boundary inputs are compiled with the concrete token count baked
in (
get_graph_input_shapeonly keeps the tracked dynamic dim whenm_model_is_splitted, which is only set with the GGML_OPENVINO_ENABLE_FALLBACKenv). The frontend compiled-model cache key
(
ggml_openvino_model_fingerprint, model-cache.cpp) hashes topology + weightsn_tokens=2 decode and the static shape rejects the bind:
Fix: fold the shapes of the statically-compiled input leafs into the
fingerprint (tokens/positions/masks stay dynamic and are excluded; weights are
already hashed; caches/scalars have constant shapes). Whole-model graphs keep
a single reusable blob; split graphs get one blob per token count. Cache
format version bumped (v2 → v5 in the local patch: v3 added the shape scan,
v4 switched to a graph-input scan because split subgraphs have an empty leafs
array, v5 resolves VIEW chains like compute_model_inputs()) so stale blobs MISS.
Note: making the boundary inputs actually dynamic instead (removing the
m_model_is_splittedgate) was also tried — the intel_gpu plugin thensegfaults in the IGC kernel JIT (SIGSEGV in libopencl-clang.so.17 during
compile_model, first conv-state split). That may be a separate plugin bugworth its own report.
Environment
X1 Carbon Gen 14 (PTL-H484), 32G RAM, iGPU shares system DRAM.
cannot determine dynamic dim for RESHAPE node conv_states_reshaped-Netc.) appear for BOTH the working 9B and thebroken 35B and are NOT the trigger — they are noise; the trigger is the
split-boundary handling above.
Local fix
patches/llamacpp-openvino-split-shape-fixes.patch (kernel project, applied on
top of the TASK-28 deterministic-tensor-names patch):
by env GGML_OPENVINO_DEBUG_INPUT (shape-only print, no tensor data access).
Verification — 2026-08-16, TASK-30 wrap-up (end-to-end, fixed binary)
Fixed binary built at 885c5bb + the two patches; scoped cgroup
(MemoryHigh=29G/MemoryMax=55G/SwapMax=16G, CPUQuota=400%), -lv 5,
-ngl 99 -dev OPENVINO0 --no-cache-prompt, env GGML_OPENVINO_DEVICE=GPUGGML_OPENVINO_MEMORY_OPTIMIZE=1, model = Ornith 1.0 35B (20G GGUF), -c 16384.
warmup compute SUCCEEDS. Full 8-token completion:
stopped by limit, n_gen = 8, n_predict = 8(tokens2,<think>,Here,'s,a,thinking...). Noov::Exception, noCompute error, no repack assert(grep count 0).
direct curl /completion → HTTP 200
{"content":"2\n\n<think>\nHere's a thinking","stop":true, "tokens_predicted":8,"tokens_evaluated":4}and the server log confirmsstopped by limit, n_gen = 8, n_predict = 8.1452 MiB retained + context 571 + compute 60 on the 31586 MiB device pool;
41/41 layers offloaded to GPU. No runaway.
git apply --checkcleanly (exit 0) on afresh clone at 885c5bb; the combined diff is byte-identical to the
verified build workdir.
Non-blocking observations for maintainers:
could not create a primitive descriptor for the reduction primitive(intel_gpu plugin,plugin.cpp:80) where the fresh-compile path succeeds — backend falls back to
recompile (deterministic per blob; warm loads still ~37s to /health).
distinct token count exercised gets its own blob (by design of defect-3's
fix).