From b92af0d312a9ef0984f8e50e259d1d9b8642d9d9 Mon Sep 17 00:00:00 2001 From: Luca Barbato Date: Wed, 26 Aug 2026 19:53:02 +0200 Subject: [PATCH] =?UTF-8?q?record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD):?= =?UTF-8?q?=20post-W2c=20P150=20refresh=20=E2=80=94=20the=20hybrid=20opt-o?= =?UTF-8?q?ut=20BEATS=20the=20eager=20default=201.24x=20(#2003);=20Mistral?= =?UTF-8?q?-7B=20warm=20eager=209.8=20tok/s=20reproduced?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every Tenstorrent figure on record predates the BACKEND-TENSTORRENT-QWEN35 W2a–W2c landing (#1715), and this row's headline number was one of them. This change re-measures on P150 thalia at 21fe11cf1, one $HOME/gpu.lock hold across all legs, tt-smi -r first, an idle leftover vllm-server stopped before work began. L1 Qwen3-0.6B b1 greedy (order-alternated pairs x3, --repeat 5, in-process run 1 discarded): host-free eager DEFAULT median 10.822 tok/s (n=12) vs VT_TT_HOST_FREE_DECODE=0 median 13.369 — #1604's flip premise is inverted; the default arm is unchanged while the opt-out improved ~2.5x unattributed. Filed as #2003, owned by this row under ## Owed in the same change. L2 Mistral-7B-v0.3 warm eager median 9.817 tok/s (n=15) replaces the single-run 4.26 anecdote as the row's standing figure — a reproduction-class upgrade, not an attributable delta (different prompt). L3 kGdnDecode op bench: composed 1.148 ms/step vs chunked 3.356 (2.92x), h2d=0 d2h=0 both arms; op-level only, production-unreached until the #1715 wiring row lands, so it is recorded without a gap-row entry. Stated plainly: NO clock window exists for any figure here (gpu_clock_state.py is NVIDIA-only); ordering alternation cancels drift but nothing attributes clocks, so every number quotes as clock-unattributed. The captured arm was not retested (#1625 hang, #1627 readback). Records move together: benchmark-record entry, both open-gaps speed rows, issue-index row for #2003, and the host-free spec's ## Owed gains the inversion with its next traceable step. Evidence verbatim at docs/bench-evidence/tt-p150-refresh-20260826.log. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai-glm-5.3-flash [maki] --- .agents/benchmark-record.md | 56 ++ .agents/issue-index.md | 1 + .../specs/tenstorrent-host-free-forward.md | 9 + .../tt-p150-refresh-20260826.log | 518 ++++++++++++++++++ docs/benchmarks/open-gaps.md | 2 +- 5 files changed, 585 insertions(+), 1 deletion(-) create mode 100644 docs/bench-evidence/tt-p150-refresh-20260826.log diff --git a/.agents/benchmark-record.md b/.agents/benchmark-record.md index 7b1ae27fb..aa79dbcf3 100644 --- a/.agents/benchmark-record.md +++ b/.agents/benchmark-record.md @@ -28752,6 +28752,62 @@ the repository: `docs/bench-evidence/tower-skip-rss-qwen3vl-thor-20260824.log` `docs/bench-evidence/tower-skip-rss-qwen3vl-thor-20260824.legs.log` (five `/usr/bin/time -v` records, four server logs, the cmake configure). +## TT P150 refresh at post-W2c main: the host-hybrid opt-out OUTPERFORMS the host-free eager DEFAULT 1.24x — the #1604 flip premise inverted (#2003); Mistral-7B eager warm reproduced at 9.8 tok/s (2026-08-26, `bench/tenstorrent-p150-refresh`, P150 `thalia`) + +Base `21fe11cf1` (origin/main, post-BACKEND-TENSTORRENT-QWEN35 W2a–W2c, +#1715). TT Release build against +`/home/lu_zero/Sources/tt/tt-metal/build_Release/lib64/cmake` + +`.../share/cmake` — **the #1604-era `-DCMAKE_PREFIX_PATH=.../build_Release` +recipe is STALE**: after the tt-metal rebase (`a3d33028975`, still carrying our +`copy_default_tilized` patches) the exported configs moved to install layout +and the stale top-level `tt-metalium-config.cmake` shadows `find_package`. +One `$HOME/gpu.lock` hold across every leg, `tt-smi -r` inside it first; an +idle leftover `vllm-server` (Qwen3.5-0.8B b32 on :8123) held that lock for +3h08m and was stopped before any leg. Evidence: +[`../docs/bench-evidence/tt-p150-refresh-20260826.log`](../docs/bench-evidence/tt-p150-refresh-20260826.log) +(verbatim per-leg stderr). Workload. Workload: greedy, batch 1, +prompt = "Write a short story about a robot learning to paint.". + +**L1 Qwen3-0.6B host-free A/B (`--repeat 5`, in-process run 1 discarded; +order-alternated pairs ×3):** default host-free eager **median 10.822 tok/s** +(n=12, 10.51–11.03) vs `VT_TT_HOST_FREE_DECODE=0` host-hybrid +**median 13.369** (n=12; 13.09–13.55 across 11 of 12 legs, one 11.49 outlier). +**The opt-out wins 1.24×** — the #1604 default-rate result read the other way. +The DEFAULT arm is unchanged against its 2026-08-21 figures (10.94–11.06), +within what no-clock-attribution can resolve, so the movement since then is a +~2.5× improvement of the OPT-OUT arm with mechanism UNATTRIBUTED; +[#2003](https://github.com/mudler/vllm.cpp/issues/2003) owns it and its next +traceable step is a per-op delta of the host-hybrid path from `b86e3705f` to +here. No clock window was sampled for ANY figure in this entry +(`tools/bench/gpu_clock_state.py` is NVIDIA-only); ordering was alternated to +cancel drift, but every number below is clock-unattributed and quotable only +as such. + +**L2 Mistral-7B-v0.3 default arm (`--repeat 6`, whole-process JIT leg +discarded per process, ~128 s each cold proc): warm median 9.817 tok/s** +(n=15, 9.46–9.94), coherent greedy prose throughout. This REPLACES the +2026-08-12 single-run 4.26 tok/s anecdote as the row's standing figure: that +run had a different prompt and no lock or repetition discipline, so the ~2.3× +difference is a reproduction-class upgrade rather than an attributable delta. +No vLLM ratio exists or can: vLLM has no Tenstorrent backend. +Checkpoint pin: `mistralai/Mistral-7B-v0.3` snapshot +`caa1feb0e54d415e2df31207e5f4e273e33509b1`, 3 shards totalling +14,496,080,928 bytes (bf16 arm; correctness standing on this line is the +refreshed golden gate green at `c31cad9c1`). Qwen checkpoint = +the documented `docs/USAGE.md` pin `c1899de289a04d12100db370d81485cdf75e47ca`. + +**L3 `kGdnDecode` op microbench (op-level ONLY — production-unreached until +the #1715 wiring row): composed 1.148 ms/step vs +`VT_TT_GDN_DECODE=chunked` 3.356** (B=8, GQA 2:8, Dk=Dv=128, 50 steps) — +composed stays the right default at **2.92×**, agreeing with the W2 decision +measured 2026-08-23 (1.139 vs 3.164). Steady-state traffic h2d=0 d2h=0 on BOTH +arms (state shadow residency holds under load). Not published to the speed-gap +rows: not an end-to-end number. + +Not retested here: the captured opt-in arm (27.1 tok/s single-request on the +old tree) — multi-request capture still hangs +([#1625](https://github.com/mudler/vllm.cpp/issues/1625)), TT async readback +remains [#1627](https://github.com/mudler/vllm.cpp/issues/1627). ## TT P150 #2003 RE-ADJUDICATED CLOCK-ATTRIBUTED: the inversion stands at governor parity — every busy sample of both arms at the 1350 MHz cap; tt_clock_state lands as the TT sibling of gpu_clock_state (#2005) (2026-08-26, `bench/tt-clock-state`, P150 `thalia`) Same binary (`21fe11cf1` bench build), workload, order-alternation, and lock diff --git a/.agents/issue-index.md b/.agents/issue-index.md index 82c1fcbbd..c17902374 100644 --- a/.agents/issue-index.md +++ b/.agents/issue-index.md @@ -797,3 +797,4 @@ rather than merged. `scripts/check-agent-record.py` gates both. | [#2042](https://github.com/mudler/vllm.cpp/issues/2042) | `SPEC-DFLASH2` | **`--enable-prefix-caching` with a DFlash2 draft kills EngineCore on the first request that takes a cache hit, at concurrency 1, and that makes the SGLang-compat `lpm` scheduler unreachable.** Measured on `3d895a202`, `sm_121a`, dgx:gpu0 under an `rc` lease, DFlash2 k=8, 1024 in / 512 out: the same binary serves 8/8 with `--no-enable-prefix-caching` and reads `ok=0 failed=8` with it on, throwing `propose_drafts_block: context position discontinuity` from inside the EngineCore step, after which every later request returns `[request submitted to a stopped AsyncLLM]`. **The invariant is the DETECTOR, not the defect**, and the three facts that settle it are: the scheduler admits a cache-hit request with `num_computed_tokens` already equal to the cached prefix (`sched/scheduler.cpp`, the waiting-admission `get_computed_blocks` arm) and the worker turns that straight into absolute positions (`prepare_inputs.cpp`, `positions[t] = num_computed_tokens_cpu[r] + query_pos[t]`); the target is served from cache and never produces the aux hidden states the draft projects, so the private store genuinely holds ZERO context rows while the target has committed N, which the second `VT_CHECK` (`L == DeviceKVNumCtx`) confirms rather than contradicts; and **upstream never reaches that state because it keeps no private store at all** — its DFlash draft writes the context K/V into the engine's own paged KV cache through `attn.impl.do_kv_cache_update(...)` (`vllm/model_executor/models/qwen3_dflash.py:601-619` at pin `5559679229`) on a slot mapping built from the TARGET's block table (`vllm/v1/spec_decode/dflash.py:145-153`), so a prefix hit hands it the draft context for free. FIXED by mirroring upstream's OTHER answer, the one for a proposer that cannot serve a request: an EMPTY draft and the target running alone (`vllm/v1/spec_decode/ngram_proposer.py:150-159`, `suffix_decoding.py:55-62`, both `continue` and neither raises), which is [#1919](https://github.com/mudler/vllm.cpp/issues/1919)'s `disabled` fallback reached from a second place. **STACKED ON [#2010](https://github.com/mudler/vllm.cpp/pull/2010) ([#2008](https://github.com/mudler/vllm.cpp/issues/2008)) AND CANNOT LAND FIRST, for correctness rather than tidiness:** the classification keys on #2010's `first_sight` predicate ("this runner has never held context for this request"), and under the pre-#2010 row-indexed state that question could not be asked, because a request the batch had MOVED presented identically to a never-seen one — so the same fallback would have swallowed #2008's crash and turned it into a silent acceptance loss. Measured, not argued: mutation M3 drops the freshness gate and reddens exactly that assertion. #2010 does NOT fix this — on its head the engine still throws the discontinuity on the second request, seven times in one run. **What it buys and costs is stated rather than implied, and it is not a free win:** prefix caching's TTFT half is kept because the target still skips the cached prefill, and a hit request stops speculating for its life, so on a shared-system-prompt workload prefix caching and DFlash2 become mutually exclusive in effect and output throughput can fall; what is unambiguously fixed is that the configuration is currently a CRASH. The repair that removes the trade is the paged context store owed under [dflash2-ctx-store-capacity.md](specs/dflash2-ctx-store-capacity.md) and tracked by #1919; a cheaper partial that keeps speculation over a TRUNCATED draft context anchored at the cache boundary is recorded under `## Owed` and deliberately not taken, because it moves draft acceptance and acceptance cannot be measured without a device. Gated by `tests/vllm/v1/spec_decode/test_dflash2_prefix_cache.cpp` (5 cases, 69 assertions, CPU, through the production `AsyncLLM` front): red-before 3/5 cases fail with the engine dead, green-after 5/5, with G3 and G5 green on both sides as controls. Wave spec [dflash2-prefix-cache.md](specs/dflash2-prefix-cache.md) | bug | | [#2067](https://github.com/mudler/vllm.cpp/issues/2067) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **The `glm5_next` converter writes a file nothing in this tree can open: register the architecture, and give it its `general.architecture` dispatch row (O9).** W1 of [#1998](https://github.com/mudler/vllm.cpp/issues/1998); spec [`specs/glm5-next-flash.md`](specs/glm5-next-flash.md) §W1. W7a ([#2011](https://github.com/mudler/vllm.cpp/issues/2011)) authored `scripts/convert-glm5-next-gguf.py`, which emits `general.architecture = glm5next`; that key had no row in `kGgufArchArms` and `Glm5NextForConditionalGeneration` was registered by no translation unit, so both entry points refused the model by name as unrecognized and every downstream wave (W3, W5, W6, W7b) had nothing to load. **One parser, two sources.** `Glm5NextHfConfigFromGguf` reads the converter's metadata and synthesizes an HF-shaped `text_config`/`vision_config` under the *same key spellings* `config.json` uses, so a GGUF descends through the SAME `ParseGlm5NextParams` a `config.json` does — one validation surface, not two that can drift. `ParseGlm5NextParams` mirrors `Glm5NextTextConfig.__post_init__` and all five `validate_architecture` rejections at transformers **v5.16.1** (`eb4d9e2a64`, the first release carrying `glm5_next`; `v5.16.0` is 404): the `full_attention` -> `deepseek_sparse_attention` layer-kind rewrite (so `Glm5NextLayerKind` has no `kFullAttention` enumerator at all and the checkpoint's spelling is unrepresentable rather than merely unused); the `linear_attn_config` -> `linear_{head_dim,num_heads,conv_kernel_dim,lower_bound}` remap together with its `safe_gate`-defaults-True rule, and the deliberate IGNORING of that dict's `kda_layers`/`full_attn_layers` index lists, which the reference never reads; the `mlp_layer_types` default `[dense]*min(3,L) + [sparse]*(L-3)`; the `indexer_types` freq/offset schedule; and the forced `head_dim = qk_rope_head_dim`, `qk_head_dim = qk_rope_head_dim + qk_nope_head_dim` overrides. **The two validators are exact complements, and that is the structural finding.** Upstream RAISES when `qk_rope_head_dim > 0` ("Expecting NoPE for the DSA attention layers"); our `MlaBlockDims::Validate` RAISES when it is not `> 0` (`mla_attention.cpp:90-93`). No value satisfies both. W1 mirrors upstream and accepts `0`; the relaxation is W3's and is recorded as **O11**, pinned by a test so W3 cannot land the geometry without moving the pin. The HF->GGUF tensor name map is enumerated structurally per layer KIND, and the config builder uses it for one reachable, shard-safe check: a `blk.N` that carries KDA tensors while the metadata declares that layer `deepseek_sparse_attention` (or the converse) is refused, because absence proves nothing on a sharded file but a CONTRADICTION is a wrong model loading quietly. **Scope honesty.** This makes the architecture RESOLVE and its config PARSE and VALIDATE. It does NOT make the model load and it does NOT make it forward: the loader, the forward and the KV-cache spec each refuse by name, naming the missing primitive and the wave that owes it (**O10**). No token, no speed, no artifact — O1 holds unchanged, and no oracle can execute this model on any device this fleet reaches. vLLM implements `glm5_next` at NO revision, so no pin was advanced and none is owed; the sole admissible reference is transformers, and **W0's lane pin for `v5.16.1` is still unwritten** — this wave cites the revision it read without recording a pin, which stays W0's deliverable | feature | | [#2070](https://github.com/mudler/vllm.cpp/issues/2070) | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` | **The shared config reader synthesizes `layer_types` from `linear_attn_config.kda_layers` as ONE-INDEXED, and GLM-5.3-Flash's list is ZERO-INDEXED.** Found while implementing W1 of [#1998](https://github.com/mudler/vllm.cpp/issues/1998) ([#2067](https://github.com/mudler/vllm.cpp/issues/2067)) by a test that erased `layer_types` to check the port reproduces upstream's default schedule; it produced a schedule off by one and every assertion about which layer is KDA failed. `src/vllm/transformers_utils/hf_config.cpp` synthesizes `cfg.layer_types` from `text_config.linear_attn_config.kda_layers` when `layer_types` is absent, resolving it as `is_kda[one_indexed - 1] = true` and dropping any entry below 1 — correct for Kimi-Linear, whose upstream defines `is_kda_layer(l) := (l+1) in kda_layers`. `zai-org/GLM-5.3-Flash`'s list is ZERO-indexed, and the checkpoint settles it two ways: it contains `0`, which a one-indexed list of 45 layers cannot, and its maximum is `44` on `num_hidden_layers: 45`. Read through the one-indexed rule the `0` is dropped and everything else shifts down, so layer 2 comes out `full_attention` where the checkpoint calls it `linear_attention` — a wrong attention kind on a third of the stack, chosen silently. **Worse than an ordinary off-by-one:** the transformers reference IGNORES `kda_layers` entirely for `glm5_next`. `Glm5NextTextConfig.__post_init__` reads only `head_dim`, `num_heads`, `short_conv_kernel_size` and `gate_lower_bound` out of that dict and derives the schedule from the top-level `layer_types` or from `idx % 4 != 3`, so the shared reader would be deriving a load-bearing schedule from a list upstream never consults, under another family's indexing convention. **Not live on `main` today**, and that is the only reason this is not a shipped defect: no `glm5_next` reached `ParseHfConfig` at all until #2067 registered it, and every published `glm5_next` config carries an explicit `layer_types`, which the synthesis is guarded behind (`cfg.layer_types.empty()`). It is a trap set for the first wave to hand this model a config without one — which is what a converter, a hand-written test fixture, or a text-only variant produces. **REPAIRED IN FLOW by #2067:** `ParseGlm5NextParams` resolves `layer_types` from its own `text_config` and from upstream's `idx % 4 != 3` default, never from `cfg.layer_types`, so this model's schedule cannot be decided by a heuristic written for another family; it additionally cross-checks `kda_layers` / `full_attn_layers` against the resolved schedule AS ZERO-INDEXED and refuses on a disagreement rather than picking a winner. The shared reader's Kimi-Linear branch is left exactly as it is — it is correct for the family it was written for, and narrowing it is a change to Kimi-Linear's behaviour this row has no gate for. Pinned by `test_glm5_next_scaffold.cpp`'s `kda_layers is ZERO-indexed, and the schedule ignores it` case; mutation M5, taking `layer_types` from the shared reader again, reds it | bug | +| [#2003](https://github.com/mudler/vllm.cpp/issues/2003) | `BACKEND-TENSTORRENT-HOST-FREE-FORWARD` | **The #1604 flip premise inverted at post-W2c `21fe11cf1`: `VT_TT_HOST_FREE_DECODE=0` (host-hybrid) outperforms the shipped eager DEFAULT 1.24x on the P150** — Qwen3-0.6B b1 greedy, order-alternated pairs ×3, in-process run 1 discarded, one `$HOME/gpu.lock` hold, `tt-smi -r` first: default median 10.822 tok/s (n=12, 10.51–11.03) vs opt-out median 13.369 (n=12; ≥13.09 on 11 of 12). The default arm is UNCHANGED against its 2026-08-21 figures (10.94–11.06 at `b86e3705f`), so what moved is a ~2.5x improvement of the opt-out arm whose mechanism is unattributed; the next traceable step is a per-op delta of the host-hybrid path from `b86e3705f` to `21fe11cf1`. Stated rather than implied: NO clock window was sampled (`tools/bench/gpu_clock_state.py` is NVIDIA-only), so every figure including the record entry that cites this issue is clock-unattributed and quotable only as such; one model shape, one board (Blackhole P150, aarch64 host, tt-metal `a3d33028975`); the captured opt-in arm was NOT retested (#1625 still blocks multi-request capture, #1627 still open). The shipped default now serves the slower of the two eager arms, which any gate using the default as denominator inherits | perf | diff --git a/.agents/specs/tenstorrent-host-free-forward.md b/.agents/specs/tenstorrent-host-free-forward.md index 25d5a0df5..e414a4f0d 100644 --- a/.agents/specs/tenstorrent-host-free-forward.md +++ b/.agents/specs/tenstorrent-host-free-forward.md @@ -326,6 +326,15 @@ investigation row but MUST be addressed by the item-5 port: ## Owed +- **The default-polarity question reopened by + [#2003](https://github.com/mudler/vllm.cpp/issues/2003).** At post-W2c + `21fe11cf1` the host-hybrid opt-out outperforms the shipped eager default + 1.24x on the P150 (Qwen3-0.6B b1, order-alternated pairs ×3, clock + unattributed — see `.agents/benchmark-record.md`, 2026-08-26 entry); the + default arm is unchanged against its #1604 figures and the opt-out improved + ~2.5x unattributed. Owed: a per-op delta of the host-hybrid path from + `b86e3705f` to current main, then a polarity decision that carries a clock + window per arm. - **No case pins `HostFreeDecodeEnabled()`'s no-caching contract on the RAC path ([#1688](https://github.com/mudler/vllm.cpp/issues/1688)).** The R5 fresh review found `ReshapeAndCacheKernel` still latching the flag in a diff --git a/docs/bench-evidence/tt-p150-refresh-20260826.log b/docs/bench-evidence/tt-p150-refresh-20260826.log new file mode 100644 index 000000000..91dd19f20 --- /dev/null +++ b/docs/bench-evidence/tt-p150-refresh-20260826.log @@ -0,0 +1,518 @@ +# TT P150 refresh legs, 2026-08-26 — thalia, base 21fe11cf1 +# build: cmake Release -DVLLM_CPP_TENSTORRENT=ON -DCMAKE_PREFIX_PATH=~/Sources/tt/tt-metal/build_Release/lib64/cmake:~/Sources/tt/tt-metal/build_Release/share/cmake +# lock: ~/gpu.lock held across ALL legs; tt-smi -r before L0; workload greedy b1 prompt='Write a short story about a robot learning to paint.' + +===== smi-before.log ===== +{ + "time": "2026-08-26T19:19:53.765427", + "host_info": { + "OS": "Linux", + "Distro": "Gentoo Linux", + "Kernel": "7.0.1-gentoo", + "Hostname": "thalia", + "Platform": "aarch64", + "Python": "3.13.12", + "Memory": "255.18 GB", + "Driver": "TT-KMD 2.10.1-pre" + }, + "host_sw_vers": { + "tt_smi": "6.2.1", + "pyluwen": "0.9.0", + "tt_umd": "0.9.9" + }, + "device_info": [ + { + "smbus_telem": { + "BOARD_ID_HIGH": "0x403", + "BOARD_ID_LOW": "0x3191406b", + "ASIC_ID": null, + "HARVESTING_STATE": "0x0", + "UPDATE_TELEM_SPEED": "0x64", + "VCORE": "0x2d6", + "TDP": "0x13", + "TDC": "0x1b", + "VDD_LIMITS": "0x38402bc", + "THM_LIMIT_SHUTDOWN": "0x6e", + "ASIC_TEMPERATURE": "0x3474da", + "VREG_TEMPERATURE": "0x0", + "BOARD_TEMPERATURE": "0x0", + "AICLK": "0x320", + "AXICLK": "0x3c0", + "ARCCLK": "0x320", + "L2CPUCLK0": "0x0", + "L2CPUCLK1": "0x0", + "L2CPUCLK2": "0x0", + "L2CPUCLK3": "0x0", + "ETH_LIVE_STATUS": "0x0", + "DDR_STATUS": "0x55555555", + "DDR_SPEED": "0x3e80", + "ETH_FW_VERSION": "0x10900", + "GDDR_FW_VERSION": "0x2000d", + "DM_APP_FW_VERSION": "0x170100", + "DM_BL_FW_VERSION": "0x0", + "FLASH_BUNDLE_VERSION": "0x13070100", + "CM_FW_VERSION": "0x1d0100", + "L2CPU_FW_VERSION": "0x0", + "FAN_SPEED": "0x2a", + "TIMER_HEARTBEAT": "0x194", + "TELEMETRY_ENUM_COUNT": "0x46", + "ENABLED_TENSIX_COL": "0xfff", + "ENABLED_ETH": "0x3edf", + "ENABLED_GDDR": "0xff", + "ENABLED_L2CPU": "0xf", + "PCIE_USAGE": "0x1", + "NOC_TRANSLATION": "0x1", + "FAN_RPM": "0x839", + "GDDR_0_1_TEMP": "0x34342e32", + "GDDR_2_3_TEMP": "0x32343234", + "GDDR_4_5_TEMP": "0x32343438", + "GDDR_6_7_TEMP": "0x32343232", + "GDDR_0_1_CORR_ERRS": "0x0", + "GDDR_2_3_CORR_ERRS": "0x0", + "GDDR_4_5_CORR_ERRS": "0x0", + "GDDR_6_7_CORR_ERRS": "0x0", + "GDDR_UNCORR_ERRS": "0x0", + "MAX_GDDR_TEMP": "0x38", + "ASIC_LOCATION": "0x0", + "BOARD_POWER_LIMIT": "0x12c", + "TDC_LIMIT_MAX": "0xc8", + "THM_LIMIT_THROTTLE": "0x5a", + "TT_FLASH_VERSION": null, + "THERM_TRIP_COUNT": "0x0", + "ASIC_ID_HIGH": "0x94712c24", + "ASIC_ID_LOW": "0x111071e4", + "AICLK_LIMIT_MAX": "0x546", + "TDP_LIMIT_MAX": "0x96", + "AICLK_ARB_MIN": "0x10320", + "AICLK_ARB_MAX": "0x80546", + "ENABLED_MIN_ARB": "0x3", + "ENABLED_MAX_ARB": "0x1d1", + "NUMBER_OF_TAGS": "0x1", + "INPUT_POWER": "0x54" + }, + "board_info": { + "bus_id": "0002:01:00.0", + "board_type": "p150a", + "board_id": "000004033191406b", + "coords": "N/A", + "dram_status": true, + "dram_speed": "16G", + "pcie_speed": 5, + "pcie_width": "16" + }, + "telemetry": { + "voltage": "0.73", + "current": " 27.0", + "power": " 19.0", + "board_power": " 84.0", + "aiclk": " 800", + "asic_temperature": "52.5", + "fan_speed": "2105", + "heartbeat": "67" + }, + "firmwares": { + "fw_bundle_version": "19.7.1.0", + "tt_flash_version": "N/A", + "cm_fw": "0.29.1.0", + "cm_fw_date": "2020-00-29", + "eth_fw": "1.9.0", + "dm_bl_fw": "0.0.0.0", + "dm_app_fw": "0.23.1.0", + "gddr_fw": "2.13" + }, + "limits": { + "vdd_min": "0.70", + "vdd_max": "0.90", + "tdp_limit": "150", + "tdc_limit": "200", + "asic_fmax": "1350", + "therm_trip_l1_limit": "90", + "thm_limit": "110", + "bus_peak_limit": 0, + "fan_rpm_limit": 0, + "board_power_limit": "300" + } + } + ], + "processes": [ + { + "pid": 102111, + "user": "lu_zero", + "device": 0, + "cmdline": "/home/lu_zero/Sources/tt/.venv/bin/python3 /home/lu_zero/Sources/tt/.venv/bin/tt-smi -s" + } + ] +} + +===== smoke-on.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/1 finish_reason=length prompt_tokens=1 completion_tokens=8 secs=11.456 tok_s=0.698 +vllm-cli: run=1/1 generate_start_unix=1787764795.209766 generate_end_unix=1787764806.665906 + +===== q-on-r1.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=21.367 tok_s=4.493 +vllm-cli: run=1/5 generate_start_unix=1787764808.011362 generate_end_unix=1787764829.378264 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.880 tok_s=10.811 +vllm-cli: run=2/5 generate_start_unix=1787764829.378327 generate_end_unix=1787764838.258551 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.862 tok_s=10.832 +vllm-cli: run=3/5 generate_start_unix=1787764838.258612 generate_end_unix=1787764847.121091 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=9.117 tok_s=10.530 +vllm-cli: run=4/5 generate_start_unix=1787764847.121151 generate_end_unix=1787764856.238090 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=9.134 tok_s=10.510 +vllm-cli: run=5/5 generate_start_unix=1787764856.238143 generate_end_unix=1787764865.372149 + +===== q-off-r1.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=17.832 tok_s=5.383 +vllm-cli: run=1/5 generate_start_unix=1787764866.732498 generate_end_unix=1787764884.564886 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.087 tok_s=13.546 +vllm-cli: run=2/5 generate_start_unix=1787764884.564959 generate_end_unix=1787764891.652121 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.262 tok_s=13.219 +vllm-cli: run=3/5 generate_start_unix=1787764891.652166 generate_end_unix=1787764898.914662 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.357 tok_s=11.487 +vllm-cli: run=4/5 generate_start_unix=1787764898.914716 generate_end_unix=1787764907.272206 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.145 tok_s=13.436 +vllm-cli: run=5/5 generate_start_unix=1787764907.272237 generate_end_unix=1787764914.417300 + +===== q-on-r2.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=19.909 tok_s=4.822 +vllm-cli: run=1/5 generate_start_unix=1787764964.220820 generate_end_unix=1787764984.129511 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.775 tok_s=10.940 +vllm-cli: run=2/5 generate_start_unix=1787764984.129565 generate_end_unix=1787764992.904493 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.969 tok_s=10.703 +vllm-cli: run=3/5 generate_start_unix=1787764992.904549 generate_end_unix=1787765001.874044 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=9.038 tok_s=10.622 +vllm-cli: run=4/5 generate_start_unix=1787765001.874111 generate_end_unix=1787765010.912150 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.987 tok_s=10.682 +vllm-cli: run=5/5 generate_start_unix=1787765010.912203 generate_end_unix=1787765019.899298 + +===== q-off-r2.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=17.911 tok_s=5.360 +vllm-cli: run=1/5 generate_start_unix=1787764915.749585 generate_end_unix=1787764933.660826 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.264 tok_s=13.216 +vllm-cli: run=2/5 generate_start_unix=1787764933.660878 generate_end_unix=1787764940.924786 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.304 tok_s=13.144 +vllm-cli: run=3/5 generate_start_unix=1787764940.924846 generate_end_unix=1787764948.228534 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.334 tok_s=13.090 +vllm-cli: run=4/5 generate_start_unix=1787764948.228585 generate_end_unix=1787764955.562520 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.218 tok_s=13.301 +vllm-cli: run=5/5 generate_start_unix=1787764955.562586 generate_end_unix=1787764962.780270 + +===== q-on-r3.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=19.533 tok_s=4.915 +vllm-cli: run=1/5 generate_start_unix=1787765021.283051 generate_end_unix=1787765040.816005 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.769 tok_s=10.947 +vllm-cli: run=2/5 generate_start_unix=1787765040.816052 generate_end_unix=1787765049.585325 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.739 tok_s=10.985 +vllm-cli: run=3/5 generate_start_unix=1787765049.585362 generate_end_unix=1787765058.324843 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.700 tok_s=11.034 +vllm-cli: run=4/5 generate_start_unix=1787765058.324906 generate_end_unix=1787765067.025193 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=8.740 tok_s=10.984 +vllm-cli: run=5/5 generate_start_unix=1787765067.025245 generate_end_unix=1787765075.765621 + +===== q-off-r3.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca +INFO auto-fit max_model_len: reduced from 40960 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=17.741 tok_s=5.411 +vllm-cli: run=1/5 generate_start_unix=1787765077.130031 generate_end_unix=1787765094.871366 +vllm-cli: run=2/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.119 tok_s=13.485 +vllm-cli: run=2/5 generate_start_unix=1787765094.871427 generate_end_unix=1787765101.990206 +vllm-cli: run=3/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.105 tok_s=13.511 +vllm-cli: run=3/5 generate_start_unix=1787765101.990252 generate_end_unix=1787765109.095389 +vllm-cli: run=4/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.131 tok_s=13.462 +vllm-cli: run=4/5 generate_start_unix=1787765109.095462 generate_end_unix=1787765116.226850 +vllm-cli: run=5/5 finish_reason=length prompt_tokens=11 completion_tokens=96 secs=7.110 tok_s=13.502 +vllm-cli: run=5/5 generate_start_unix=1787765116.226894 generate_end_unix=1787765123.337158 + +===== m-discard.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--mistralai--Mistral-7B-v0.3/snapshots/caa1feb0e54d415e2df31207e5f4e273e33509b1 +INFO auto-fit max_model_len: reduced from 32768 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/1 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=131.243 tok_s=0.244 +vllm-cli: run=1/1 generate_start_unix=1787765133.218057 generate_end_unix=1787765264.461066 + +===== m-r1.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--mistralai--Mistral-7B-v0.3/snapshots/caa1feb0e54d415e2df31207e5f4e273e33509b1 +INFO auto-fit max_model_len: reduced from 32768 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=128.848 tok_s=0.248 +vllm-cli: run=1/6 generate_start_unix=1787765271.755639 generate_end_unix=1787765400.603165 +vllm-cli: run=2/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.250 tok_s=9.845 +vllm-cli: run=2/6 generate_start_unix=1787765400.603240 generate_end_unix=1787765403.853574 +vllm-cli: run=3/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.266 tok_s=9.799 +vllm-cli: run=3/6 generate_start_unix=1787765403.853623 generate_end_unix=1787765407.119388 +vllm-cli: run=4/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.358 tok_s=9.529 +vllm-cli: run=4/6 generate_start_unix=1787765407.119444 generate_end_unix=1787765410.477542 +vllm-cli: run=5/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.383 tok_s=9.460 +vllm-cli: run=5/6 generate_start_unix=1787765410.477597 generate_end_unix=1787765413.860281 +vllm-cli: run=6/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.258 tok_s=9.823 +vllm-cli: run=6/6 generate_start_unix=1787765413.860329 generate_end_unix=1787765417.118105 + +===== m-r2.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--mistralai--Mistral-7B-v0.3/snapshots/caa1feb0e54d415e2df31207e5f4e273e33509b1 +INFO auto-fit max_model_len: reduced from 32768 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=127.745 tok_s=0.250 +vllm-cli: run=1/6 generate_start_unix=1787765424.458715 generate_end_unix=1787765552.203376 +vllm-cli: run=2/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.241 tok_s=9.873 +vllm-cli: run=2/6 generate_start_unix=1787765552.203447 generate_end_unix=1787765555.444648 +vllm-cli: run=3/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.317 tok_s=9.648 +vllm-cli: run=3/6 generate_start_unix=1787765555.444678 generate_end_unix=1787765558.761368 +vllm-cli: run=4/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.230 tok_s=9.907 +vllm-cli: run=4/6 generate_start_unix=1787765558.761398 generate_end_unix=1787765561.991320 +vllm-cli: run=5/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.227 tok_s=9.916 +vllm-cli: run=5/6 generate_start_unix=1787765561.991346 generate_end_unix=1787765565.218464 +vllm-cli: run=6/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.220 tok_s=9.937 +vllm-cli: run=6/6 generate_start_unix=1787765565.218491 generate_end_unix=1787765568.438927 + +===== m-r3.err ===== +vllm-cli: loading model from /home/lu_zero/.cache/huggingface/hub/models--mistralai--Mistral-7B-v0.3/snapshots/caa1feb0e54d415e2df31207e5f4e273e33509b1 +INFO auto-fit max_model_len: reduced from 32768 to 8192 to fit the KV cache (256 blocks x 32 tokens). Raise --num-blocks / --kv-cache-memory for a longer context. +vllm.cpp: Asynchronous scheduling is disabled (max_concurrent_batches=1) +vllm-cli: run=1/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=126.583 tok_s=0.253 +vllm-cli: run=1/6 generate_start_unix=1787765575.694321 generate_end_unix=1787765702.276993 +vllm-cli: run=2/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.256 tok_s=9.827 +vllm-cli: run=2/6 generate_start_unix=1787765702.277075 generate_end_unix=1787765705.533554 +vllm-cli: run=3/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.316 tok_s=9.649 +vllm-cli: run=3/6 generate_start_unix=1787765705.533606 generate_end_unix=1787765708.849873 +vllm-cli: run=4/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.276 tok_s=9.769 +vllm-cli: run=4/6 generate_start_unix=1787765708.849921 generate_end_unix=1787765712.125552 +vllm-cli: run=5/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.260 tok_s=9.817 +vllm-cli: run=5/6 generate_start_unix=1787765712.125592 generate_end_unix=1787765715.385285 +vllm-cli: run=6/6 finish_reason=length prompt_tokens=12 completion_tokens=32 secs=3.263 tok_s=9.808 +vllm-cli: run=6/6 generate_start_unix=1787765715.385326 generate_end_unix=1787765718.647979 + +===== gdn-composed.log ===== +2026-08-26 19:37:29.235 | info | Device | Opening user mode device driver (tt_cluster.cpp:228) +2026-08-26 19:37:29.235 | info | UMD | Cluster constructor started. (cluster.cpp:389) +2026-08-26 19:37:29.235 | info | UMD | Creating TopologyDiscovery for architecture: blackhole (topology_discovery.cpp:96) +2026-08-26 19:37:29.235 | info | UMD | Starting topology discovery. (topology_discovery.cpp:115) +2026-08-26 19:37:29.239 | info | UMD | Established firmware bundle version: 19.7.1 (topology_discovery.cpp:575) +2026-08-26 19:37:29.239 | info | UMD | Completed topology discovery. (topology_discovery.cpp:119) +2026-08-26 19:37:29.280 | warning | UMD | Sysmem (0x40000000 bytes) using regular pages; pre-allocate hugepages for better DMA performance (e.g. 1 × 1GB or 512 × 2MB; on AArch64 also 2 × 512MB). (silicon_sysmem_manager.cpp:106) +2026-08-26 19:37:29.281 | info | UMD | Opening local chip ids/PCIe ids: {0}/[0] and remote chip ids {} (cluster.cpp:173) +2026-08-26 19:37:29.281 | info | UMD | IOMMU: enabled (cluster.cpp:147) +2026-08-26 19:37:29.281 | info | UMD | KMD version: 2.10.1 (cluster.cpp:150) +2026-08-26 19:37:29.281 | info | UMD | Cluster constructor completed. (cluster.cpp:648) +2026-08-26 19:37:29.282 | info | UMD | Starting devices in cluster (cluster.cpp:1184) +2026-08-26 19:37:29.350 | info | UMD | Starting devices in cluster completed. (cluster.cpp:1192) +[doctest] doctest version is "2.5.2" +[doctest] run with "--help" for options +2026-08-26 19:37:29.494 | info | Distributed | Using auto discovery to generate mesh graph. (metal_env.cpp:437) +2026-08-26 19:37:29.494 | info | Distributed | Constructing control plane using auto-discovery (no mesh graph descriptor). (metal_env.cpp:499) +2026-08-26 19:37:29.494 | warning | Always | Unknown motherboard 'AMPONED8-2T/BCM' for chip_id=0 (bus_id=0x1) — falling back to bus_id as tray_id. Add this motherboard and its bus IDs to mobo_to_bus_ids in physical_system_discovery.cpp. (physical_system_discovery.cpp:119) +2026-08-26 19:37:29.494 | info | Fabric | Logical multi-mesh adjacency: intermesh degree histogram {0:1}; intra-mesh degree histograms mesh0 {0:1} (topology_mapper_utils.cpp:2410) +2026-08-26 19:37:29.494 | info | Fabric | Physical multi-mesh adjacency: intermesh degree histogram {0:1}; intra-mesh degree histograms mesh0 {0:1} (topology_mapper_utils.cpp:2418) +2026-08-26 19:37:30.220 | info | Metal | Enabling program cache on MeshDevice 1 (mesh_device.cpp:1066) +=============================================================================== +/home/lu_zero/Sources/vllmcpp-tt-bench/tests/vt/test_tenstorrent_backend.cpp:3094: +TEST CASE: kTENSTORRENT kGdnDecode step microbench (opt-in) +/home/lu_zero/Sources/vllmcpp-tt-bench/tests/vt/test_tenstorrent_backend.cpp:3180: MESSAGE: kGdnDecode step microbench composed B=8 Hk=2 Hv=8 Dk=128 Dv=128: 1.14812 ms/step over 50 steps; traffic h2d=0 d2h=0 (state bytes=4194304) +=============================================================================== +[doctest] test cases: 1 | 1 passed | 0 failed | 38 skipped +[doctest] assertions: 4 | 4 passed | 0 failed | +[doctest] Status: SUCCESS! +2026-08-26 19:37:30.609 | info | BuildKernels | JIT cache stats: 58/58 hits (100.0%) [58 cached, 2 build-once dedup, 0 merged artifacts, 0 merged genfiles] (build_cache_telemetry.cpp:207) +2026-08-26 19:37:30.609 | info | BuildKernels | JIT telemetry: 3 registered TelemetryTokens (build_cache_telemetry.cpp:230) +2026-08-26 19:37:30.615 | info | Device | Closing user mode device drivers (tt_cluster.cpp:513) +2026-08-26 19:37:30.615 | info | UMD | Closing devices in cluster (cluster.cpp:1197) +2026-08-26 19:37:30.683 | info | UMD | Closing devices in cluster completed. (cluster.cpp:1206) +2026-08-26 19:37:30.683 | info | UMD | Cluster destructor started. (cluster.cpp:912) +2026-08-26 19:37:30.683 | info | UMD | Cluster destructor completed. (cluster.cpp:915) + +===== gdn-chunked.log ===== +2026-08-26 19:37:30.819 | info | Device | Opening user mode device driver (tt_cluster.cpp:228) +2026-08-26 19:37:30.819 | info | UMD | Cluster constructor started. (cluster.cpp:389) +2026-08-26 19:37:30.819 | info | UMD | Creating TopologyDiscovery for architecture: blackhole (topology_discovery.cpp:96) +2026-08-26 19:37:30.819 | info | UMD | Starting topology discovery. (topology_discovery.cpp:115) +2026-08-26 19:37:30.823 | info | UMD | Established firmware bundle version: 19.7.1 (topology_discovery.cpp:575) +2026-08-26 19:37:30.823 | info | UMD | Completed topology discovery. (topology_discovery.cpp:119) +2026-08-26 19:37:30.867 | warning | UMD | Sysmem (0x40000000 bytes) using regular pages; pre-allocate hugepages for better DMA performance (e.g. 1 × 1GB or 512 × 2MB; on AArch64 also 2 × 512MB). (silicon_sysmem_manager.cpp:106) +2026-08-26 19:37:30.867 | info | UMD | Opening local chip ids/PCIe ids: {0}/[0] and remote chip ids {} (cluster.cpp:173) +2026-08-26 19:37:30.867 | info | UMD | IOMMU: enabled (cluster.cpp:147) +2026-08-26 19:37:30.867 | info | UMD | KMD version: 2.10.1 (cluster.cpp:150) +2026-08-26 19:37:30.867 | info | UMD | Cluster constructor completed. (cluster.cpp:648) +2026-08-26 19:37:30.868 | info | UMD | Starting devices in cluster (cluster.cpp:1184) +2026-08-26 19:37:30.950 | info | UMD | Starting devices in cluster completed. (cluster.cpp:1192) +[doctest] doctest version is "2.5.2" +[doctest] run with "--help" for options +2026-08-26 19:37:31.095 | info | Distributed | Using auto discovery to generate mesh graph. (metal_env.cpp:437) +2026-08-26 19:37:31.095 | info | Distributed | Constructing control plane using auto-discovery (no mesh graph descriptor). (metal_env.cpp:499) +2026-08-26 19:37:31.095 | warning | Always | Unknown motherboard 'AMPONED8-2T/BCM' for chip_id=0 (bus_id=0x1) — falling back to bus_id as tray_id. Add this motherboard and its bus IDs to mobo_to_bus_ids in physical_system_discovery.cpp. (physical_system_discovery.cpp:119) +2026-08-26 19:37:31.095 | info | Fabric | Logical multi-mesh adjacency: intermesh degree histogram {0:1}; intra-mesh degree histograms mesh0 {0:1} (topology_mapper_utils.cpp:2410) +2026-08-26 19:37:31.095 | info | Fabric | Physical multi-mesh adjacency: intermesh degree histogram {0:1}; intra-mesh degree histograms mesh0 {0:1} (topology_mapper_utils.cpp:2418) +2026-08-26 19:37:31.805 | info | Metal | Enabling program cache on MeshDevice 1 (mesh_device.cpp:1066) +=============================================================================== +/home/lu_zero/Sources/vllmcpp-tt-bench/tests/vt/test_tenstorrent_backend.cpp:3094: +TEST CASE: kTENSTORRENT kGdnDecode step microbench (opt-in) +/home/lu_zero/Sources/vllmcpp-tt-bench/tests/vt/test_tenstorrent_backend.cpp:3180: MESSAGE: kGdnDecode step microbench chunked B=8 Hk=2 Hv=8 Dk=128 Dv=128: 3.35639 ms/step over 50 steps; traffic h2d=0 d2h=0 (state bytes=4194304) +=============================================================================== +[doctest] test cases: 1 | 1 passed | 0 failed | 38 skipped +[doctest] assertions: 4 | 4 passed | 0 failed | +[doctest] Status: SUCCESS! +2026-08-26 19:37:32.474 | info | BuildKernels | JIT cache stats: 97/97 hits (100.0%) [97 cached, 22 build-once dedup, 0 merged artifacts, 0 merged genfiles] (build_cache_telemetry.cpp:207) +2026-08-26 19:37:32.474 | info | BuildKernels | JIT telemetry: 3 registered TelemetryTokens (build_cache_telemetry.cpp:230) +2026-08-26 19:37:32.480 | info | Device | Closing user mode device drivers (tt_cluster.cpp:513) +2026-08-26 19:37:32.480 | info | UMD | Closing devices in cluster (cluster.cpp:1197) +2026-08-26 19:37:32.589 | info | UMD | Closing devices in cluster completed. (cluster.cpp:1206) +2026-08-26 19:37:32.589 | info | UMD | Cluster destructor started. (cluster.cpp:912) +2026-08-26 19:37:32.589 | info | UMD | Cluster destructor completed. (cluster.cpp:915) + +===== smi-after.log ===== +{ + "time": "2026-08-26T19:35:19.393079", + "host_info": { + "OS": "Linux", + "Distro": "Gentoo Linux", + "Kernel": "7.0.1-gentoo", + "Hostname": "thalia", + "Platform": "aarch64", + "Python": "3.13.12", + "Memory": "255.18 GB", + "Driver": "TT-KMD 2.10.1-pre" + }, + "host_sw_vers": { + "tt_smi": "6.2.1", + "pyluwen": "0.9.0", + "tt_umd": "0.9.9" + }, + "device_info": [ + { + "smbus_telem": { + "BOARD_ID_HIGH": "0x403", + "BOARD_ID_LOW": "0x3191406b", + "ASIC_ID": null, + "HARVESTING_STATE": "0x0", + "UPDATE_TELEM_SPEED": "0x64", + "VCORE": "0x2d5", + "TDP": "0x15", + "TDC": "0x1d", + "VDD_LIMITS": "0x38402bc", + "THM_LIMIT_SHUTDOWN": "0x6e", + "ASIC_TEMPERATURE": "0x3daf5a", + "VREG_TEMPERATURE": "0x0", + "BOARD_TEMPERATURE": "0x0", + "AICLK": "0x320", + "AXICLK": "0x3c0", + "ARCCLK": "0x320", + "L2CPUCLK0": "0x0", + "L2CPUCLK1": "0x0", + "L2CPUCLK2": "0x0", + "L2CPUCLK3": "0x0", + "ETH_LIVE_STATUS": "0x0", + "DDR_STATUS": "0x55555555", + "DDR_SPEED": "0x3e80", + "ETH_FW_VERSION": "0x10900", + "GDDR_FW_VERSION": "0x2000d", + "DM_APP_FW_VERSION": "0x170100", + "DM_BL_FW_VERSION": "0x0", + "FLASH_BUNDLE_VERSION": "0x13070100", + "CM_FW_VERSION": "0x1d0100", + "L2CPU_FW_VERSION": "0x0", + "FAN_SPEED": "0x39", + "TIMER_HEARTBEAT": "0x25bb", + "TELEMETRY_ENUM_COUNT": "0x46", + "ENABLED_TENSIX_COL": "0xfff", + "ENABLED_ETH": "0x3edf", + "ENABLED_GDDR": "0xff", + "ENABLED_L2CPU": "0xf", + "PCIE_USAGE": "0x1", + "NOC_TRANSLATION": "0x1", + "FAN_RPM": "0xa97", + "GDDR_0_1_TEMP": "0x3e3e3c3e", + "GDDR_2_3_TEMP": "0x3e403c40", + "GDDR_4_5_TEMP": "0x3c403c42", + "GDDR_6_7_TEMP": "0x3a3e3a3e", + "GDDR_0_1_CORR_ERRS": "0x0", + "GDDR_2_3_CORR_ERRS": "0x0", + "GDDR_4_5_CORR_ERRS": "0x0", + "GDDR_6_7_CORR_ERRS": "0x0", + "GDDR_UNCORR_ERRS": "0x0", + "MAX_GDDR_TEMP": "0x42", + "ASIC_LOCATION": "0x0", + "BOARD_POWER_LIMIT": "0x12c", + "TDC_LIMIT_MAX": "0xc8", + "THM_LIMIT_THROTTLE": "0x5a", + "TT_FLASH_VERSION": null, + "THERM_TRIP_COUNT": "0x0", + "ASIC_ID_HIGH": "0x94712c24", + "ASIC_ID_LOW": "0x111071e4", + "AICLK_LIMIT_MAX": "0x546", + "TDP_LIMIT_MAX": "0x96", + "AICLK_ARB_MIN": "0x10320", + "AICLK_ARB_MAX": "0x80546", + "ENABLED_MIN_ARB": "0x3", + "ENABLED_MAX_ARB": "0x1d1", + "NUMBER_OF_TAGS": "0x1", + "INPUT_POWER": "0x5e" + }, + "board_info": { + "bus_id": "0002:01:00.0", + "board_type": "p150a", + "board_id": "000004033191406b", + "coords": "N/A", + "dram_status": true, + "dram_speed": "16G", + "pcie_speed": 5, + "pcie_width": "16" + }, + "telemetry": { + "voltage": "0.72", + "current": " 29.0", + "power": " 21.0", + "board_power": " 94.0", + "aiclk": " 800", + "asic_temperature": "61.7", + "fan_speed": "2711", + "heartbeat": "1609" + }, + "firmwares": { + "fw_bundle_version": "19.7.1.0", + "tt_flash_version": "N/A", + "cm_fw": "0.29.1.0", + "cm_fw_date": "2020-00-29", + "eth_fw": "1.9.0", + "dm_bl_fw": "0.0.0.0", + "dm_app_fw": "0.23.1.0", + "gddr_fw": "2.13" + }, + "limits": { + "vdd_min": "0.70", + "vdd_max": "0.90", + "tdp_limit": "150", + "tdc_limit": "200", + "asic_fmax": "1350", + "therm_trip_l1_limit": "90", + "thm_limit": "110", + "bus_peak_limit": 0, + "fan_rpm_limit": 0, + "board_power_limit": "300" + } + } + ], + "processes": [ + { + "pid": 45228, + "user": "lu_zero", + "device": 0, + "cmdline": "/home/lu_zero/Sources/tt/.venv/bin/python3 /home/lu_zero/Sources/tt/.venv/bin/tt-smi -s" + } + ] +} diff --git a/docs/benchmarks/open-gaps.md b/docs/benchmarks/open-gaps.md index aeadff61c..2d9a6a833 100644 --- a/docs/benchmarks/open-gaps.md +++ b/docs/benchmarks/open-gaps.md @@ -61,7 +61,7 @@ | ROCm `d=128` decode arm (`BACKEND-ROCM`, [#382](https://github.com/mudler/vllm.cpp/issues/382)) | **DIRECTIONAL, not binding.** gfx1200, both sides in the pinned oracle container, 1024/128 c1, 8 prompts, 3 reps: TPOT 42.40 -> **11.67 ms** with `VT_ATTN_DECODE_D128=1`; vLLM `555967922` 6.68 ms, so **6.35x -> 1.75x** | Harnesses differ (oracle over HTTP, ours in-process) and no same-tool per-call trace exists, so [#488](https://github.com/mudler/vllm.cpp/issues/488) stays open. Owed: decode-windowed `rocprofv3` both sides | | Gemma4 prefill-peer Finish-barrier cost (#1047 item 3) | **Attribution GREEN at T=2029, not a product ship number.** A (wait-only BEFORE) 1122.10 tok/s vs B (AFTER) 1094.24; +2.55% / 46.05 ms/req. Wait deletion is not authorized to land. | T=19 not run. Overlap/async retirement unmeasured. Detail: [benchmark-record](../../.agents/benchmark-record.md) | | Tenstorrent Blackhole (`BACKEND-TENSTORRENT`) | **NOT APPLICABLE (speed).** Correctness: OPT-125m STRICT 6/6 e2e on real hardware. Qwen3-0.6B has a device-specific golden and short 4-token warm smoke (~0.28 tok/s), not a completed speed run | Full 16x16 Qwen3 gate, then device-resident tensors + `ttnn::sdpa_decode` before any performance comparison. [Spec](../../.agents/specs/tenstorrent-backend.md) | -| Mistral-7B-v0.3 on Tenstorrent (`BACKEND-TENSTORRENT-MISTRAL`) | **PENDING (speed).** First data point: 4.26 tok/s warm, batch 1, 32 tok, single run on a P150. Not a gate, not reproduced. No vLLM ratio exists or can (no TT backend). Correctness 16/16 | Reproduce idle with a same-binary A/B before quoting. [Record](../../.agents/benchmark-record.md), [spec](../../.agents/specs/tenstorrent-mistral.md) | +| Mistral-7B-v0.3 on Tenstorrent (`BACKEND-TENSTORRENT-MISTRAL`) | **PENDING (speed gate), figure reproduced: warm eager median 9.817 tok/s** (b1, 32-tok greedy, n=15 under one lock, 2026-08-26, post-W2c `21fe11cf1`), replacing the single-run 4.26 anecdote. No vLLM ratio exists or can (no TT backend); clock-unattributed | A binding speed gate still needs the same-binary paired method with a sampled board window; [Record](../../.agents/benchmark-record.md), [spec](../../.agents/specs/tenstorrent-mistral.md) | | Host-free decode graph (`BACKEND-TENSTORRENT-HOST-FREE-FORWARD`) | **INVERTED vs the #1604 flip AND CLOCK-ATTRIBUTED at post-W2c `21fe11cf1`: default eager median 10.88 tok/s vs median 13.645 for `VT_TT_HOST_FREE_DECODE=0` (Qwen3-0.6B b1, order-alternated pairs x3, ratio 1.254) ([#2003](https://github.com/mudler/vllm.cpp/issues/2003)).** The 2026-08-21 flip evidence (10.94-11.06 vs 5.34 at `b86e3705f`) was real; by 2026-08-26 the OPT-OUT arm improved ~2.5x while the default is unchanged. Wired through the [#2005](https://github.com/mudler/vllm.cpp/issues/2005) `tt_clock_state.py` harness: raw windows refuse on within-run spread because the P150 AICLK governor is two-state (800 idle / pegged cap), and the busy-slice refold shows **every busy sample of both arms at exactly 1350 MHz** — six windows spread 0.00%, cross-arm offsets 0/0%, judge PASS — so the slower default is a real path difference at clock parity, not an excursion | Owed on [#2003](https://github.com/mudler/vllm.cpp/issues/2003): per-op delta of the host-hybrid path from `b86e3705f`; claimed-max 1350 pin UNVERIFIED. Captured-arm hang [#1625](https://github.com/mudler/vllm.cpp/issues/1625), #1627 TT async readback. [Record](../../.agents/benchmark-record.md), [spec](../../.agents/specs/tenstorrent-host-free-forward.md) | | Qwen3.5-0.8B on Tenstorrent (`BACKEND-TENSTORRENT-QWEN35`, [#1715](https://github.com/mudler/vllm.cpp/issues/1715)) | **FIRST TT numbers for the family, and they are ~100x under a 7B dense on the same board:** warm 0.089 tok/s (~11.2 s/token), reproduced exactly across runs, arms and a 96-token generation (P150, single lock hold, production entry point). The host-free opt-out arm is IDENTICAL (0.080/0.089/0.089 vs 0.080/0.090/0.089), so #1604/#2003 polarity is irrelevant here. Output coherent. Unattributed by sampler; minute-scale duty cycles make cap-pegging near certain, which only sharpens the gap. Leading hypothesis: a per-layer host/reference fallback inside the hybrid-GDN stack (~400 ms/layer-token over ~28 layers) | Profile ONE eager decode step and name the dominant op BEFORE registering more kernels blindly. Captured tracing stays blocked behind [#1625](https://github.com/mudler/vllm.cpp/issues/1625), so the first pass is eager-side. [Record](../../.agents/benchmark-record.md), evidence `docs/bench-evidence/tt-qwen35-first-speed-20260826.log` | | Prompt logprobs (`SAMPLE-PROMPT-LOGPROBS`, #223) | **NO number measured, claimed or owed.** Correctness-only, CPU. Upstream ships this path explicitly unoptimized (`gpu_model_runner.py:5622-5623`); a step where no request asks is unchanged | Floor if one is ever wanted: vLLM's own `prompt_logprobs=k`, same model and prompt |