record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): post-W2c P150 refresh — the hybrid opt-out BEATS the eager default 1.24x (#2003); Mistral-7B warm eager 9.8 tok/s reproduced - #2004
Open
lu-zero wants to merge 1 commit into
Conversation
…— the hybrid opt-out BEATS the eager default 1.24x (mudler#2003); Mistral-7B warm eager 9.8 tok/s reproduced Every Tenstorrent figure on record predates the BACKEND-TENSTORRENT-QWEN35 W2a–W2c landing (mudler#1715), and this row's headline number was one of them. This change re-measures on P150 thalia at 21fe11c, one $HOME/gpu.lock hold across all legs, tt-smi -r first, an idle leftover vllm-server stopped before work began. L1 Qwen3-0.6B b1 greedy (order-alternated pairs x3, --repeat 5, in-process run 1 discarded): host-free eager DEFAULT median 10.822 tok/s (n=12) vs VT_TT_HOST_FREE_DECODE=0 median 13.369 — mudler#1604's flip premise is inverted; the default arm is unchanged while the opt-out improved ~2.5x unattributed. Filed as mudler#2003, owned by this row under ## Owed in the same change. L2 Mistral-7B-v0.3 warm eager median 9.817 tok/s (n=15) replaces the single-run 4.26 anecdote as the row's standing figure — a reproduction-class upgrade, not an attributable delta (different prompt). L3 kGdnDecode op bench: composed 1.148 ms/step vs chunked 3.356 (2.92x), h2d=0 d2h=0 both arms; op-level only, production-unreached until the mudler#1715 wiring row lands, so it is recorded without a gap-row entry. Stated plainly: NO clock window exists for any figure here (gpu_clock_state.py is NVIDIA-only); ordering alternation cancels drift but nothing attributes clocks, so every number quotes as clock-unattributed. The captured arm was not retested (mudler#1625 hang, mudler#1627 readback). Records move together: benchmark-record entry, both open-gaps speed rows, issue-index row for mudler#2003, and the host-free spec's ## Owed gains the inversion with its next traceable step. Evidence verbatim at docs/bench-evidence/tt-p150-refresh-20260826.log. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai-glm-5.3-flash [maki]
lu-zero
added a commit
to lu-zero/vllm.cpp
that referenced
this pull request
Aug 26, 2026
…ted — inversion stands at governor parity; tt_clock_state first wired use (mudler#2005) The tool landed same-day and its first consumer was the morning's unattributed result. Same binary, workload, alternation and lock discipline; one sampler window per arm around six order-alternated legs. Throughput: default host-free eager median 10.880 tok/s vs opt-out 13.645 (ratio 1.254) — reproduces the morning's 10.822 vs 13.369. Attribution finding: the Blackhole P150 AICLK governor is two-state on this telemetry, 800 idle vs pegged under load, so all six raw windows refuse within-run spread at 40.74% — kept verbatim in the evidence log, because loosening the NVIDIA rule to fit one board is exactly the failure mode the protocol forbids. Attribution instead rides the live-recorded busy column: pid-held /dev/tenstorrent fds per interval, recorded before any outcome value existed. tt_refold_busy.py rebuilds those slices and every window shows ONE distinct AICLK {1350}, spread 0.00%, cross-arm offsets 0/0, judge PASS rc=0. Both arms computed at identical cap clocks; the inverted ratio is a real path difference and mudler#2003 stays open for the per-op delta. Also on this branch: the consolidated host-free gap row (supersedes either single-session wording regardless of merge order against mudler#2004), the mudler#2005 index row naming its owed items — verified claimed-max pin, in-process pyluwen sampling, and the spread-scope policy decision — and the spec's Evidence/Owed sections carrying both findings and the explicit never-loosen-a-threshold note. Evidence verbatim: docs/bench-evidence/tt-p150-clock-attributed-20260826.log FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai-glm-5.3-flash [maki]
Collaborator
Author
|
Reviewed (records-only scope): discipline, evidence traceability (all figures traced to docs/bench-evidence/tt-p150-refresh-20260826.log), honesty markers, trailers, issue-index append-only — all PASS. One finding: the host-free row and benchmark-record entries conflict textually with the later clock-attributed re-adjudication on #2006 in BOTH merge orders. Resolution agreed there: land this PR first, #2006 rebases over it and wins the superseded rows. No change requested on this branch. |
lu-zero
added a commit
that referenced
this pull request
Aug 26, 2026
…inversion stands at governor parity; tt_clock_state first wired use (#2005) The tool landed same-day and its first consumer was the morning's unattributed result. Same binary, workload, alternation and lock discipline; one sampler window per arm around six order-alternated legs. Throughput: default host-free eager median 10.880 tok/s vs opt-out 13.645 (ratio 1.254) — reproduces the morning's 10.822 vs 13.369. Attribution finding: the Blackhole P150 AICLK governor is two-state on this telemetry, 800 idle vs pegged under load, so all six raw windows refuse within-run spread at 40.74% — kept verbatim in the evidence log, because loosening the NVIDIA rule to fit one board is exactly the failure mode the protocol forbids. Attribution instead rides the live-recorded busy column: pid-held /dev/tenstorrent fds per interval, recorded before any outcome value existed. tt_refold_busy.py rebuilds those slices and every window shows ONE distinct AICLK {1350}, spread 0.00%, cross-arm offsets 0/0, judge PASS rc=0. Both arms computed at identical cap clocks; the inverted ratio is a real path difference and #2003 stays open for the per-op delta. Also on this branch: the consolidated host-free gap row (supersedes either single-session wording regardless of merge order against #2004), the #2005 index row naming its owed items — verified claimed-max pin, in-process pyluwen sampling, and the spread-scope policy decision — and the spec's Evidence/Owed sections carrying both findings and the explicit never-loosen-a-threshold note. Evidence verbatim: docs/bench-evidence/tt-p150-clock-attributed-20260826.log FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai-glm-5.3-flash [maki]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): post-W2c P150 refresh — the hybrid opt-out BEATS the eager default 1.24x (#2003); Mistral-7B warm eager 9.8 tok/s reproduced
Every Tenstorrent figure on record predates the
BACKEND-TENSTORRENT-QWEN35 W2a–W2c landing (#1715), and this row's headline
number was one of them. This change re-measures on P150 thalia at 21fe11c,
one $HOME/gpu.lock hold across all legs, tt-smi -r first, an idle leftover
vllm-server stopped before work began.
L1 Qwen3-0.6B b1 greedy (order-alternated pairs x3, --repeat 5, in-process
run 1 discarded): host-free eager DEFAULT median 10.822 tok/s (n=12) vs
VT_TT_HOST_FREE_DECODE=0 median 13.369 — #1604's flip premise is inverted;
the default arm is unchanged while the opt-out improved ~2.5x unattributed.
Filed as #2003, owned by this row under ## Owed in the same change.
L2 Mistral-7B-v0.3 warm eager median 9.817 tok/s (n=15) replaces the
single-run 4.26 anecdote as the row's standing figure — a reproduction-class
upgrade, not an attributable delta (different prompt).
L3 kGdnDecode op bench: composed 1.148 ms/step vs chunked 3.356 (2.92x),
h2d=0 d2h=0 both arms; op-level only, production-unreached until the #1715
wiring row lands, so it is recorded without a gap-row entry.
Stated plainly: NO clock window exists for any figure here
(gpu_clock_state.py is NVIDIA-only); ordering alternation cancels drift but
nothing attributes clocks, so every number quotes as clock-unattributed.
The captured arm was not retested (#1625 hang, #1627 readback).
Records move together: benchmark-record entry, both open-gaps speed rows,
issue-index row for #2003, and the host-free spec's ## Owed gains the
inversion with its next traceable step. Evidence verbatim at
docs/bench-evidence/tt-p150-refresh-20260826.log.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]