Skip to content

record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): post-W2c P150 refresh — the hybrid opt-out BEATS the eager default 1.24x (#2003); Mistral-7B warm eager 9.8 tok/s reproduced - #2004

Open
lu-zero wants to merge 1 commit into
mudler:mainfrom
lu-zero:bench/tenstorrent-p150-refresh

Conversation

@lu-zero

@lu-zero lu-zero commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): post-W2c P150 refresh — the hybrid opt-out BEATS the eager default 1.24x (#2003); Mistral-7B warm eager 9.8 tok/s reproduced

Every Tenstorrent figure on record predates the
BACKEND-TENSTORRENT-QWEN35 W2a–W2c landing (#1715), and this row's headline
number was one of them. This change re-measures on P150 thalia at 21fe11c,
one $HOME/gpu.lock hold across all legs, tt-smi -r first, an idle leftover
vllm-server stopped before work began.

L1 Qwen3-0.6B b1 greedy (order-alternated pairs x3, --repeat 5, in-process
run 1 discarded): host-free eager DEFAULT median 10.822 tok/s (n=12) vs
VT_TT_HOST_FREE_DECODE=0 median 13.369 — #1604's flip premise is inverted;
the default arm is unchanged while the opt-out improved ~2.5x unattributed.
Filed as #2003, owned by this row under ## Owed in the same change.

L2 Mistral-7B-v0.3 warm eager median 9.817 tok/s (n=15) replaces the
single-run 4.26 anecdote as the row's standing figure — a reproduction-class
upgrade, not an attributable delta (different prompt).

L3 kGdnDecode op bench: composed 1.148 ms/step vs chunked 3.356 (2.92x),
h2d=0 d2h=0 both arms; op-level only, production-unreached until the #1715
wiring row lands, so it is recorded without a gap-row entry.

Stated plainly: NO clock window exists for any figure here
(gpu_clock_state.py is NVIDIA-only); ordering alternation cancels drift but
nothing attributes clocks, so every number quotes as clock-unattributed.
The captured arm was not retested (#1625 hang, #1627 readback).

Records move together: benchmark-record entry, both open-gaps speed rows,
issue-index row for #2003, and the host-free spec's ## Owed gains the
inversion with its next traceable step. Evidence verbatim at
docs/bench-evidence/tt-p150-refresh-20260826.log.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]

…— the hybrid opt-out BEATS the eager default 1.24x (mudler#2003); Mistral-7B warm eager 9.8 tok/s reproduced

Every Tenstorrent figure on record predates the
BACKEND-TENSTORRENT-QWEN35 W2a–W2c landing (mudler#1715), and this row's headline
number was one of them. This change re-measures on P150 thalia at 21fe11c,
one $HOME/gpu.lock hold across all legs, tt-smi -r first, an idle leftover
vllm-server stopped before work began.

L1 Qwen3-0.6B b1 greedy (order-alternated pairs x3, --repeat 5, in-process
run 1 discarded): host-free eager DEFAULT median 10.822 tok/s (n=12) vs
VT_TT_HOST_FREE_DECODE=0 median 13.369 — mudler#1604's flip premise is inverted;
the default arm is unchanged while the opt-out improved ~2.5x unattributed.
Filed as mudler#2003, owned by this row under ## Owed in the same change.

L2 Mistral-7B-v0.3 warm eager median 9.817 tok/s (n=15) replaces the
single-run 4.26 anecdote as the row's standing figure — a reproduction-class
upgrade, not an attributable delta (different prompt).

L3 kGdnDecode op bench: composed 1.148 ms/step vs chunked 3.356 (2.92x),
h2d=0 d2h=0 both arms; op-level only, production-unreached until the mudler#1715
wiring row lands, so it is recorded without a gap-row entry.

Stated plainly: NO clock window exists for any figure here
(gpu_clock_state.py is NVIDIA-only); ordering alternation cancels drift but
nothing attributes clocks, so every number quotes as clock-unattributed.
The captured arm was not retested (mudler#1625 hang, mudler#1627 readback).

Records move together: benchmark-record entry, both open-gaps speed rows,
issue-index row for mudler#2003, and the host-free spec's ## Owed gains the
inversion with its next traceable step. Evidence verbatim at
docs/bench-evidence/tt-p150-refresh-20260826.log.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]
lu-zero added a commit to lu-zero/vllm.cpp that referenced this pull request Aug 26, 2026
…ted — inversion stands at governor parity; tt_clock_state first wired use (mudler#2005)

The tool landed same-day and its first consumer was the morning's
unattributed result. Same binary, workload, alternation and lock discipline;
one sampler window per arm around six order-alternated legs.

Throughput: default host-free eager median 10.880 tok/s vs opt-out 13.645
(ratio 1.254) — reproduces the morning's 10.822 vs 13.369.

Attribution finding: the Blackhole P150 AICLK governor is two-state on this
telemetry, 800 idle vs pegged under load, so all six raw windows refuse
within-run spread at 40.74% — kept verbatim in the evidence log, because
loosening the NVIDIA rule to fit one board is exactly the failure mode the
protocol forbids. Attribution instead rides the live-recorded busy column:
pid-held /dev/tenstorrent fds per interval, recorded before any outcome
value existed. tt_refold_busy.py rebuilds those slices and every window
shows ONE distinct AICLK {1350}, spread 0.00%, cross-arm offsets 0/0,
judge PASS rc=0. Both arms computed at identical cap clocks; the inverted
ratio is a real path difference and mudler#2003 stays open for the per-op delta.

Also on this branch: the consolidated host-free gap row (supersedes either
single-session wording regardless of merge order against mudler#2004), the mudler#2005
index row naming its owed items — verified claimed-max pin, in-process
pyluwen sampling, and the spread-scope policy decision — and the spec's
Evidence/Owed sections carrying both findings and the explicit
never-loosen-a-threshold note.

Evidence verbatim: docs/bench-evidence/tt-p150-clock-attributed-20260826.log

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]
@lu-zero

lu-zero commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator Author

Reviewed (records-only scope): discipline, evidence traceability (all figures traced to docs/bench-evidence/tt-p150-refresh-20260826.log), honesty markers, trailers, issue-index append-only — all PASS. One finding: the host-free row and benchmark-record entries conflict textually with the later clock-attributed re-adjudication on #2006 in BOTH merge orders. Resolution agreed there: land this PR first, #2006 rebases over it and wins the superseded rows. No change requested on this branch.

lu-zero added a commit that referenced this pull request Aug 26, 2026
…inversion stands at governor parity; tt_clock_state first wired use (#2005)

The tool landed same-day and its first consumer was the morning's
unattributed result. Same binary, workload, alternation and lock discipline;
one sampler window per arm around six order-alternated legs.

Throughput: default host-free eager median 10.880 tok/s vs opt-out 13.645
(ratio 1.254) — reproduces the morning's 10.822 vs 13.369.

Attribution finding: the Blackhole P150 AICLK governor is two-state on this
telemetry, 800 idle vs pegged under load, so all six raw windows refuse
within-run spread at 40.74% — kept verbatim in the evidence log, because
loosening the NVIDIA rule to fit one board is exactly the failure mode the
protocol forbids. Attribution instead rides the live-recorded busy column:
pid-held /dev/tenstorrent fds per interval, recorded before any outcome
value existed. tt_refold_busy.py rebuilds those slices and every window
shows ONE distinct AICLK {1350}, spread 0.00%, cross-arm offsets 0/0,
judge PASS rc=0. Both arms computed at identical cap clocks; the inverted
ratio is a real path difference and #2003 stays open for the per-op delta.

Also on this branch: the consolidated host-free gap row (supersedes either
single-session wording regardless of merge order against #2004), the #2005
index row naming its owed items — verified claimed-max pin, in-process
pyluwen sampling, and the spread-scope policy decision — and the spec's
Evidence/Owed sections carrying both findings and the explicit
never-loosen-a-threshold note.

Evidence verbatim: docs/bench-evidence/tt-p150-clock-attributed-20260826.log

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant