Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
96 commits
Select commit Hold shift + click to select a range
f445fa8
spec(KERNEL-QUANT-CIQ-GEMM-ROCM): commit the keep-quant provider spec
ghazni101 Aug 21, 2026
6236e9e
feat(KERNEL-QUANT-CIQ-GEMM-ROCM): land the W1 keep-quant providers on…
ghazni101 Aug 21, 2026
2578c9b
spec(ROCM-QUANT-GEMM-BW): commit the keep-quant bandwidth spec
ghazni101 Aug 21, 2026
8e78dfa
perf(ROCM-QUANT-GEMM-BW): split each super-block across the warp's lanes
ghazni101 Aug 22, 2026
0783930
spec(GFX1100-TG150): commit the 150 tok/s campaign spec
ghazni101 Aug 22, 2026
41d060b
perf(ROCM-QUANT-GEMM-BW): branch-free Q6K scalar decode + gated qg4 a…
ghazni101 Aug 22, 2026
eb9e46f
perf(ROCM-QUANT-GEMM-BW): split-K decode arm for the keep-quant GEMM
ghazni101 Aug 22, 2026
c112d88
perf(BACKEND-ROCM): f32-query decode-GQA arm behind VT_ATTN_DECODE_GQ…
ghazni101 Aug 22, 2026
094f603
perf(BACKEND-ROCM): row-permuted keep-quant in_proj + tiny-N f32-out …
ghazni101 Aug 22, 2026
592afd3
spec(GFX1100-TG200): commit the 200 tok/s campaign spec
ghazni101 Aug 22, 2026
767d369
measure(GFX1100-TG200): T1 re-prices the tip -- 40.65 tok/s median, b…
ghazni101 Aug 22, 2026
f758643
measure(GFX1100-TG200): T2a splits the GdnPostConv symbol -- the grid…
ghazni101 Aug 22, 2026
5e57df7
research(GFX1100-TG200): rank vLLM/SGLang mechanisms against our meas…
ghazni101 Aug 22, 2026
2034e17
record(GFX1100-TG200): reject the pointer-keyed quant cache -- alloca…
ghazni101 Aug 23, 2026
bdc8114
perf(GFX1100-TG200): merged keep-quant gate_up -- one quant GEMM per …
ghazni101 Aug 23, 2026
64f38f1
perf(GFX1100-TG200): T2b flips ROCm support_static_graph_mode -- deco…
ghazni101 Aug 23, 2026
68316bd
record(GFX1100-TG200): session-state note appended to t2b evidence (h…
ghazni101 Aug 23, 2026
8a2ee61
perf(GFX1100-TG200): T3a ports the f32-query decode-GQA attention arm…
ghazni101 Aug 23, 2026
239c106
record(GFX1100-TG200): t3a evidence session-state note (hindsight 500s)
ghazni101 Aug 23, 2026
9df8f0a
test(GFX1100-TG200): T4a lands the red-first ROCm quant-dot gate
ghazni101 Aug 23, 2026
205df07
perf(GFX1100-TG200): T4a adds the VT_GEMV_MMVQ=1 K-quant decode GEMV …
ghazni101 Aug 23, 2026
9ce2055
perf(GFX1100-TG200): T4a folds activation quant into the MMVQ GEMV pr…
ghazni101 Aug 23, 2026
554e01d
record(GFX1100-TG200): T4a evidence -- decode GEMV lever closed negat…
ghazni101 Aug 23, 2026
60dac6c
fix(GFX1100-TG200): T4a repairs the MMVQ arm -- m-gates the whole dis…
ghazni101 Aug 23, 2026
baf97fe
record(GFX1100-TG200): T4a repair-round evidence -- lever adopted at …
ghazni101 Aug 23, 2026
0ebe869
test(GFX1100-TG200): T4a repair-2 adds host-side dispatch-route count…
ghazni101 Aug 23, 2026
28cbc6d
perf(GFX1100-TG200): T4a lever-B1 makes the fold crossover tunable be…
ghazni101 Aug 23, 2026
26e14bb
record(GFX1100-TG200): T4a lever-B1 evidence -- fold-crossover re-tun…
ghazni101 Aug 23, 2026
5140bd1
record(GFX1100-TG200): T4a lever-B2 attributes all 21.6 Cijk calls/to…
ghazni101 Aug 23, 2026
6c55cdf
test(GFX1100-TG200): T4a lever-B2 adds the red-first f32-out decode-s…
ghazni101 Aug 23, 2026
8dcac46
perf(GFX1100-TG200): T4a lever-B2 adds the VT_SKINNY_BF16=1 f32-out d…
ghazni101 Aug 23, 2026
6218d9b
test(GFX1100-TG200): T4a lever-B2 restores the routing-witness case p…
ghazni101 Aug 24, 2026
9f5d9da
test(GFX1100-TG200): T4a lever-B2 corrects the routing-witness expect…
ghazni101 Aug 24, 2026
cecceb7
record(GFX1100-TG200): T4a lever-B2 adopts the VT_SKINNY_BF16=1 f32-o…
ghazni101 Aug 24, 2026
a17a209
record(GFX1100-TG200): T4a lever-B2 closes the loop -- ON-arm capture…
ghazni101 Aug 24, 2026
aab871f
test(GFX1100-TG200): T4a lever-B2 makes the TRUE-unset routing window…
ghazni101 Aug 24, 2026
242e8da
record(GFX1100-TG200): T4a lever-B2 re-pins the default-routing witne…
ghazni101 Aug 24, 2026
c574c7c
attribution(GFX1100-TG200): lever-C maps all 97 QuantizeQ8KK decode l…
ghazni101 Aug 24, 2026
cfbba00
test(GFX1100-TG200): lever-C adds red-first witnesses for the fused n…
ghazni101 Aug 24, 2026
25f511f
feat(GFX1100-TG200): lever-C fuses Q8_K activation quant into the Rms…
ghazni101 Aug 24, 2026
79a992d
record(GFX1100-TG200): lever-C adopts VT_NORM_QUANT_FUSED=1 at +7.3% …
ghazni101 Aug 24, 2026
65d8555
perf(GFX1100-TG200): T5a vectorizes the shared Q8_K quant superblock …
ghazni101 Aug 25, 2026
53eb36a
perf(GFX1100-TG200): T5b extends the f32-Q DecodeGqa arm to head_dim 128
ghazni101 Aug 25, 2026
bd6551f
record(GFX1100-TG200): T5b evidence — attention routing hole closed, …
ghazni101 Aug 25, 2026
580c185
record(GFX1100-TG200): T5c closed negative — MMVQ nontemporal weight …
ghazni101 Aug 25, 2026
314dd01
perf(GFX1100-TG200): T6a adds the warp-per-row cooperative GDN scan arm
ghazni101 Aug 25, 2026
b8608a2
record(GFX1100-TG200): T6a evidence — cooperative scan adopted at +4.…
ghazni101 Aug 25, 2026
0a7ac39
perf(GFX1100-TG200): T6b adds the warp-per-item cooperative attn prea…
ghazni101 Aug 25, 2026
3945b7f
record(GFX1100-TG200): T6b evidence — cooperative preamble adopted at…
ghazni101 Aug 25, 2026
1cee023
record(GFX1100-TG200): session-close attribution — 76.6 tok/s median …
ghazni101 Aug 25, 2026
9d028b7
record(GFX1100-TG200): classify the campaign env knobs kernel-internal
ghazni101 Aug 25, 2026
4793e87
Merge upstream main into row/GFX1100-TG200
ghazni101 Aug 25, 2026
bfaff5b
record(GFX1100-TG200): T7 evidence — COALK load topology closed wash,…
ghazni101 Aug 25, 2026
2c30b3e
perf(GFX1100-TG200): T8 adds a cooperative single-row rmsnorm arm
ghazni101 Aug 25, 2026
90ae363
record(GFX1100-TG200): T8 evidence — cooperative rmsnorm adopted at +…
ghazni101 Aug 25, 2026
36a6b87
perf(GFX1100-TG200): T9 gives the gated norm a per-row block
ghazni101 Aug 25, 2026
23f1858
record(GFX1100-TG200): T9 evidence — cooperative gated norm adopted a…
ghazni101 Aug 25, 2026
7c518f6
perf(GFX1100-TG200): T10 and T11 add warp postconv and row-split scan…
ghazni101 Aug 26, 2026
c5827b3
record(GFX1100-TG200): T12 evidence — gated-quant fusion not adopted,…
ghazni101 Aug 26, 2026
cce72ca
record(GFX1100-TG200): dispatch-gap audit names the sampling round trip
ghazni101 Aug 26, 2026
38b08ed
record(GFX1100-TG200): retract T10/T11 engine claims — corrupted outp…
ghazni101 Aug 26, 2026
35fc458
test(GFX1100-TG200): close the gate gap that let the T10 stride bug ship
ghazni101 Aug 26, 2026
b078310
record(GFX1100-TG200): add mechanical decision rules for the T10/T11 …
ghazni101 Aug 26, 2026
920994b
record(GFX1100-TG200): refine dispatch-gap into three measured sub-ta…
ghazni101 Aug 26, 2026
9b75706
perf(GFX1100-TG200): T14 adds a row-split greedy argmax arm
ghazni101 Aug 26, 2026
1033485
record(GFX1100-TG200): verify at ISA level that the dp4a core uses v_…
ghazni101 Aug 26, 2026
317b198
record(GFX1100-TG200): T10/T11 re-measured in a clean window — adopte…
ghazni101 Aug 26, 2026
af50d9d
record(GFX1100-TG200): full-config verification — 92.7-92.9 tok/s can…
ghazni101 Aug 26, 2026
92392bb
record(GFX1100-TG200): async-serving A/B is a wash under HTTP overhea…
ghazni101 Aug 26, 2026
7972238
record(GFX1100-TG200): LDS epilogue closed negative; host-load sensit…
ghazni101 Aug 26, 2026
3be8d2f
record(GFX1100-TG200): position-resolved wvSplitKSml audit corrects t…
ghazni101 Aug 26, 2026
0081ed9
perf(GFX1100-TG200): T16 adds wvSplitK launch-config sweep knobs
ghazni101 Aug 26, 2026
cf03a41
record(GFX1100-TG200): T14 stacked engine A/B closed — adopted at +0.9%
ghazni101 Aug 26, 2026
2ae5496
record(GFX1100-TG200): core pinning does not isolate host-memory cont…
ghazni101 Aug 26, 2026
acfc6ee
record(GFX1100-TG200): T13 implementation plan scoped with file:line …
ghazni101 Aug 26, 2026
39b8b32
adopt(GFX1100-TG200): T16 YTILE=4 default — wins 5/5 paired, bit-iden…
ghazni101 Aug 26, 2026
5d0bc50
record(GFX1100-TG200): idle-window gate 100.4 tok/s, T13 wash, copy s…
ghazni101 Aug 26, 2026
4f3c87e
record(GFX1100-TG200): T17 v_dot2_f32_bf16 closed not-adopted — memor…
ghazni101 Aug 26, 2026
96c523d
feat(GFX1100-TG200): T18 v_dot4 instruction selection in KQuantGemvMm…
ghazni101 Aug 26, 2026
e3551d0
record(GFX1100-TG200): T20 full-warp cooperative GEMV closed not-adop…
ghazni101 Aug 26, 2026
0b787af
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN proje…
ghazni101 Aug 26, 2026
6836c11
fix(GFX1100-TG200): T22 NormQuant bridge token survives non-matching …
ghazni101 Aug 26, 2026
485386a
feat(GFX1100-TG200): T24 LDS-buffered quant epilogue in RmsNormRowCoo…
ghazni101 Aug 26, 2026
65675eb
T25: keep ssm_out as Q5_K with runtime input permutation
ghazni101 Aug 26, 2026
05455b6
T27: warp-cooperative QuantizeQ8KK for decode (+2.06%, byte-identical)
ghazni101 Aug 27, 2026
f36ba60
feat(GFX1100-TG150): fuse Q6_K bias correction into single dot product
ghazni101 Aug 27, 2026
82d65ae
spec(GFX1100-TG150): add Outcome, update Now after attempt cap
ghazni101 Aug 27, 2026
e98ed73
merge(GFX1100-TG200): integrate TG150-SPEC campaign spec
ghazni101 Aug 27, 2026
70fb405
merge(GFX1100-TG200): integrate ROCM-QUANT-GEMM-BW keep-quant work
ghazni101 Aug 27, 2026
87c6f25
cherry-pick(KV-FP8): W6 ROCm fp8-e4m3 KV cache store and read onto TG200
ghazni101 Aug 27, 2026
eca1377
spec(rocm): fp8 KV cache decode attention for PagedAttnDecodeGqaF32Q …
ghazni101 Aug 27, 2026
e43a39d
feat(rocm): fp8 KV cache support in PagedAttnDecodeGqaF32Q (#7)
ghazni101 Aug 27, 2026
0617b3f
perf(ROCm): parallel random sample with shared primitives
ghazni101 Aug 27, 2026
08d0df5
feat(rocm): advertise fp8 KV cache dtype support and add GGUF chat te…
ghazni101 Aug 27, 2026
e1ea27c
merge(GFX1100-TG200): integrate upstream main to clear #1936
ghazni101 Aug 28, 2026
b058bb7
fix(GFX1100-TG200): repair the record and env gates the branch carrie…
ghazni101 Aug 28, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/backend-matrix.md

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions .agents/engine-matrix.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/feature-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,7 +82,7 @@ Confirmed NON-gap: vLLM has removed prompt adapters.
| SGLang RadixAttention behavior parity (fuse-or-flag) | SGLang v0.5.15 `f63458b` `mem_cache/radix_cache.py`, `managers/schedule_policy.py`, `constrained/outlines_jump_forward.py` | `ACTIVE` T2 | **Scoped 2026-07-27 (`CLAIM-SGLANG-RADIX-SCOPE`); W1+W2 IMPLEMENTED 2026-07-27 (`CLAIM-SGLANG-IMPL`, rows now `ACTIVE`).** VERDICT: SGLang's radix TREE == our block-hash APC ⇒ RadixAttention is **already FUSED**, `--enable-radix-attention` is an ALIAS for the APC toggle (LANDED: server alias + C-ABI `enable_prefix_caching` tri-state, ABI v7). Genuinely-distinct behavior = cache-aware **LPM scheduling `--schedule-policy=lpm`** (LANDED: `SchedulerPolicy::kLPM` reorders the FCFS waiting deque by APC longest-match, ported FROM `schedule_policy.py:205,229`, output-neutral; gate `test_scheduler_lpm` 6/6). **SW2 in-batch prefix-collision de-prioritization LANDED 2026-07-27 (`CLAIM-SGLANG-SW2`)** inside the `kLPM` reorder (block-hash APC keys, no second trie; ported FROM `schedule_policy.py:253-301,311`), output-neutral; its throughput lever is NOT-APPLICABLE — our APC caches at allocation time so the 2nd same-step collider already hits (within-step dedup subsumes it). overlap scheduler == `ENG-ASYNC-SCHED` (fused). **SW3 jump-forward decoding — safe TOKEN-UNIQUE subset LANDED 2026-07-28 (`CLAIM-SGLANG-SW3`)**: forced-token detection hook `StructuredOutputGrammar::forced_token()` + opt-in driver `DrainForcedTokens` (env `VT_ENABLE_JUMP_FORWARD`, default OFF), provably byte-identical to per-token constrained decode (jumps only where the grammar leaves exactly one valid token — no re-tokenization); gate `test_jump_forward` 5/5 (RED-first). Residual: SW4 + the general re-tokenization span + production scheduler splice (named). **ABI/API/flag EXPOSURE 2026-07-28 (`CLAIM-SGLANG-ABI-DOCS`, reconciled to ABI v10):** LPM + jump-forward made first-class DOCUMENTED knobs on ALL THREE surfaces (were server-only / env-only) — LPM via the concurrent session's C-ABI **string** field `vllm_model_params.scheduling_policy="lpm"` (ABI v9; NO duplicate int knob) + C++ `EngineParams::policy=kLPM` + server `--scheduling-policy lpm`; jump-forward via new C-ABI `vllm_model_params.enable_jump_forward` (tri-state int, ABI **v10** appended after the v9 fields) + C++ `EngineParams::enable_jump_forward` + server `--[enable\|disable]-jump-forward`. `VT_ENABLE_JUMP_FORWARD` retained as env override. User docs [docs/SGLANG-COMPAT.md](../docs/SGLANG-COMPAT.md) + spec [sglang-enablement.md](specs/sglang-enablement.md); ABI e2e `tests/capi/test_capi.cpp` (2 v10 jump-forward cases; `vllm_abi_version()`==10). Default-inert (all-zero ⇒ byte-identical). Rows `KV-SGLANG-RADIX-CACHE` + `ENG-SGLANG-BEHAVIOR-FLAG`. Sibling benchmark track = `BACKEND-GATE-CUDA-SGLANG*` (unchanged) | [sglang-radixattention.md](specs/sglang-radixattention.md) |
| **SGLang parity PROGRAM** (whole-surface inventory + oracle) | SGLang v0.5.15 `f63458b` — full runtime surface | `SPIKE` T2 | **Elevated 2026-07-27 (`CLAIM-SGLANG-PARITY-PROGRAM`).** The vLLM-parity approach replicated for SGLang: a tabular whole-surface inventory (44 rows) classifying every SGLang capability **FUSED (23) / SGLANG-DISTINCT (8) / INVENTORIED (5) / OUT-OF-SCOPE (8)**, plus SGLang stood up as a correctness + performance ORACLE (dgx GB10 via the arm64 cu130 image — no from-source build needed). SGLang is a competitor perf FLOOR + correctness cross-check, NOT the mirror source (vLLM stays behavior truth). Headline SGLANG-DISTINCT opt-ins: LPM scheduling, in-batch prefix de-prioritization, radix eviction strategies, jump-forward, custom logit processors, batch-invariant determinism, PD disaggregation, two-batch EP overlap. Full map + ranked plan in the matrix. Sibling benchmark rows `BACKEND-GATE-CUDA-SGLANG*` unchanged | [sglang-matrix.md](sglang-matrix.md); [sglang-parity-oracle.md](specs/sglang-parity-oracle.md) |
| SlidingWindowSpec + ChunkedLocalAttentionSpec | `v1/kv_cache_interface.py` | `PARTIAL` T1 | Both execution leaves are implemented: W1 sliding-window and W3 chunked-local sizing, registry/grouping, manager prefix/recycling policy, admission and hybrid-disabled conversion pass their ported CPU/property/sanitizer gates (G1/G2). The compute-locality consumers are now GPU-gated (2026-07-27 `CLAIM-ROADMAP-C5`, dgx GB10: Gemma-2/Gemma-3 sliding-window model gates 48/48; `test_chunked_local_attention` 5/5). The KV memory-OPTIMIZATION path (optimized-manager held-block cap vs the full-allocation fallback the current model gates use) still needs a model-level hybrid-manager memory gate (G8) — kept `PARTIAL` honestly | [sliding-local-yarn-long-context.md](specs/sliding-local-yarn-long-context.md) |
| fp8 KV cache (`cache_dtype=fp8*`) | `layers/quantization/kv_cache.py` | ◐ T1 | `KV-FP8` ACTIVE / `QUANT-KV-FP8` PARTIAL — W1 CPU fp8-e4m3 store+read+config-parse landed; CUDA + memory-halving e2e later | [fp8-kv-cache](specs/fp8-kv-cache.md) |
| fp8 KV cache (`cache_dtype=fp8*`) | `layers/quantization/kv_cache.py` | ◐ T1 | `KV-FP8` ACTIVE / `QUANT-KV-FP8` PARTIAL — W1 CPU fp8-e4m3 store+read+config-parse landed; W2 CUDA store+read landed; W6 ROCm store+read landed; memory-halving e2e later | [fp8-kv-cache](specs/fp8-kv-cache.md) |
| nvfp4 / per-token-head / turboquant KV | `config/cache.py` | ☐ T2 | | `planned: specs/nvfp4-kv-cache.md` |
| KV offload (CPU tiering, LRU/ARC) | `v1/kv_offload/` | ☐ T2 | | `planned: specs/kv-offload.md` |
| External KV-cache provider ABI + LMCache (MP service and in-process connectors) | `config/kv_transfer.py`, `distributed/kv_transfer/kv_connector/v1/{base,lmcache_connector,lmcache_mp_connector}.py` | ☐ T2 | explicit roadmap outcome `KV-EXTERNAL-CACHE`: mirror `kv_producer`/`kv_consumer`/`kv_both`, scheduler/worker metadata, async layer load/store, dynamic external connector modules, failure policy, metrics and cache-lifecycle ownership; gate the official LMCache shared-prefix quickstart plus Qwen3.6 hybrid behavior | `planned: specs/external-kv-cache-lmcache.md` |
Expand Down
7 changes: 7 additions & 0 deletions .agents/issue-index.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/quantization-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,7 +157,7 @@ Pinned vLLM source: `vllm/config/cache.py:19-36`.

| ID | Item | Upstream | Our code | Tests/evidence | Spike/spec | State | Owner |
|---|---|---|---|---|---|---|---|
| `QUANT-KV-FP8` | fp8, fp8_e4m3, fp8_e5m2 | `vllm/config/cache.py:19-25`; `vllm/model_executor/layers/quantization/kv_cache.py:42-191`; store `cache_kernels.cu:241-252`; scale convention `quant_utils.cuh:296-308` | **W1 CPU fp8-e4m3 store+read LANDED**: [codec](../include/vt/fp8_kv.h#L39), [store kernel](../src/vt/cpu/cpu_cache.cpp#L143), [read dequant](../src/vt/cpu/cpu_paged_attn.cpp#L82), [config parse](../include/vllm/v1/kv_cache_dtype.h#L37). **W2 CUDA fp8-e4m3 store+read LANDED** ([#1593](https://github.com/mudler/vllm.cpp/issues/1593)): [store kernel](../src/vt/cuda/cuda_cache.cu), [read dequant](../src/vt/cuda/cuda_paged_attn.cu) -- gate [test_cuda_fp8_kv_cache](../tests/vt/test_cuda_fp8_kv_cache.cpp), whose DEVICE cases are UNEXECUTED; the CUDA TUs COMPILE (CI `cuda-fat-build`, ten architectures, run 32495320287 on `4d71e776e`) but that job builds with tests OFF, so none has been executed (spec `## Owed`). **W3 runner integration LANDED** ([#1593](https://github.com/mudler/vllm.cpp/issues/1593)): [`--kv-cache-dtype`](../src/vllm/entrypoints/openai/server_main.cpp) reaches `EngineParams`, the checkpoint's `kv_cache_quant_algo` is honoured when no flag is typed, [`ApplyCacheDType`](../src/vllm/v1/kv_cache_interface.cpp) retypes every attention spec (group and heterogeneous per-layer alike) so one byte budget buys exactly 2x the blocks, and [`kv_cache_route.h`](../include/vllm/model_executor/models/kv_cache_route.h) is the ONE place the store and the read decide float versus fp8. An fp8 cache DISABLES FA-2 prefill and decode, the WMMA ladder and the vectorized decode kernels, which are bf16-native; the resulting throughput cost is unmeasured and recorded as such. e5m2 compute, per-head scales, the Metal/ROCm arms, the C-ABI field, the 16 unrouted architectures (15 of which refuse with a message that names neither fp8 nor the flag) and the `k_scale`/`v_scale` weight-loader read are named later bricks (see spec) | [test_ops_fp8_kv_cache](../tests/vt/test_ops_fp8_kv_cache.cpp#L1) — 8 cases / 511 assertions, round-trip within the e4m3 band + fp8-vs-bf16 NMSE<1% + paged-attention e2e; RED-first (wrong store direction fails 3/480); W3 [test_kv_cache_fp8_wiring](../tests/vllm/entrypoints/test_kv_cache_fp8_wiring.cpp) — 31 cases, G1-G12, entering through `LoadedEngine` rather than by building a spec by hand, and bounding the fp8 pages against a bf16 run inside e4m3's own round-trip envelope, and [test_serve_kv_cache_dtype](../tests/vllm/entrypoints/openai/test_serve_kv_cache_dtype.cpp) — the flag through the REAL `VllmServerMain` | [fp8-kv-cache](specs/fp8-kv-cache.md) | `PARTIAL` | - |
| `QUANT-KV-FP8` | fp8, fp8_e4m3, fp8_e5m2 | `vllm/config/cache.py:19-25`; `vllm/model_executor/layers/quantization/kv_cache.py:42-191`; store `cache_kernels.cu:241-252`; scale convention `quant_utils.cuh:296-308` | **W1 CPU fp8-e4m3 store+read LANDED**: [codec](../include/vt/fp8_kv.h#L39), [store kernel](../src/vt/cpu/cpu_cache.cpp#L143), [read dequant](../src/vt/cpu/cpu_paged_attn.cpp#L82), [config parse](../include/vllm/v1/kv_cache_dtype.h#L37). **W2 CUDA fp8-e4m3 store+read LANDED** ([#1593](https://github.com/mudler/vllm.cpp/issues/1593)): [store kernel](../src/vt/cuda/cuda_cache.cu), [read dequant](../src/vt/cuda/cuda_paged_attn.cu) -- gate [test_cuda_fp8_kv_cache](../tests/vt/test_cuda_fp8_kv_cache.cpp), whose DEVICE cases are UNEXECUTED; the CUDA TUs COMPILE (CI `cuda-fat-build`, ten architectures, run 32495320287 on `4d71e776e`) but that job builds with tests OFF, so none has been executed (spec `## Owed`). **W3 runner integration LANDED** ([#1593](https://github.com/mudler/vllm.cpp/issues/1593)): [`--kv-cache-dtype`](../src/vllm/entrypoints/openai/server_main.cpp) reaches `EngineParams`, the checkpoint's `kv_cache_quant_algo` is honoured when no flag is typed, [`ApplyCacheDType`](../src/vllm/v1/kv_cache_interface.cpp) retypes every attention spec (group and heterogeneous per-layer alike) so one byte budget buys exactly 2x the blocks, and [`kv_cache_route.h`](../include/vllm/model_executor/models/kv_cache_route.h) is the ONE place the store and the read decide float versus fp8. An fp8 cache DISABLES FA-2 prefill and decode, the WMMA ladder and the vectorized decode kernels, which are bf16-native; the resulting throughput cost is unmeasured and recorded as such. e5m2 compute, per-head scales, the Metal/ROCm arms, the C-ABI field, the 16 unrouted architectures (15 of which refuse with a message that names neither fp8 nor the flag) and the `k_scale`/`v_scale` weight-loader read are named later bricks (see spec) | [test_ops_fp8_kv_cache](../tests/vt/test_ops_fp8_kv_cache.cpp#L1) — 8 cases / 511 assertions, round-trip within the e4m3 band + fp8-vs-bf16 NMSE<1% + paged-attention e2e; RED-first (wrong store direction fails 3/480); W3 [test_kv_cache_fp8_wiring](../tests/vllm/entrypoints/test_kv_cache_fp8_wiring.cpp) — 31 cases, G1-G12, entering through `LoadedEngine` rather than by building a spec by hand, and bounding the fp8 pages against a bf16 run inside e4m3's own round-trip envelope, and [test_serve_kv_cache_dtype](../tests/vllm/entrypoints/openai/test_serve_kv_cache_dtype.cpp) — the flag through the REAL `VllmServerMain`. **W6 ROCm fp8-e4m3 store+read LANDED** ([#2065](https://github.com/mudler/vllm.cpp/issues/2065)): [store kernel](../src/vt/rocm/rocm_dense_basic.hip), [read dequant](../src/vt/rocm/rocm_paged_attn.hip) -- gate [test_rocm_fp8_kv_cache](../tests/vt/test_rocm_fp8_kv_cache.cpp), 7/7 cases 28/28 assertions on gfx1100 | [fp8-kv-cache](specs/fp8-kv-cache.md) | `PARTIAL` | - |
| `QUANT-KV-FP8-VENDOR` | fp8_inc, fp8_ds_mla | `vllm/config/cache.py:24-25`; vendor KV implementations selected by attention backend | - | no quantized KV cache | `planned: specs/vendor-fp8-kv-cache.md` | `INVENTORIED` | - |
| `QUANT-KV-TURBO` | k8v4, 4bit_nc, k3v4_nc, 3bit_nc | `vllm/config/cache.py:28-33`; TurboQuant dependency path | - | no quantized KV cache | `planned: specs/turboquant-kv-cache.md` | `INVENTORIED` | - |
| `QUANT-KV-PER-HEAD` | int4/int8/fp8 per-token-head | `vllm/config/cache.py:34`; quantized cache kernels selected by backend | - | no quantized KV cache | `planned: specs/per-head-kv-cache.md` | `INVENTORIED` | - |
Expand Down
Loading