Repository navigation
perf(codegen): recover guarded ordinary-parameter specialization lost by #8033 #8079
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Aug 14, 2026 Status: the in-flight branch works, and it recovers nearly all of the regression soundly
The
perf/issue-8079-guarded-param-specializationworktree had ~2,900 lines of new code across 34 files sitting uncommitted with no active owner. I committed it as-is to the branch first (2f70cd7bb, unvalidated preservation commit —git reset --soft HEAD~1restores the exact working-tree state) and then validated it. Nothing was rebased or edited.Acceptance test: passes byte-exact
test-files/test_gap_specabi_ordinary_param_guards.tsagainst the pinned oracle (nodev26.5.1, matching.node-version):PASS test_gap_specabi_ordinary_param_guards Parity Pass: 1 Parity Fail: 0 Compile Fail: 0 Crashed: 0 Skipped: 0 Parity Rate: 100.0% (harness exit 0)The test covers every case this issue asked for: annotation lies on array elements, array container, scalar, bare
number, object fields, class instance, and both flat and nested discriminated-union positions; an escaped/indirect call throughconst escapedArrayTotal = arrayTotal; an accessor-backed property with a getter-hit count; a parameter the body mutates; a recursive generic (treeTotal); and a 50,000-allocation loop exercised on both the clone and the failed-guard fallback.The clone is actually live — a pass alone would be vacuous
A guard test passes trivially if the optimization never engages and everything falls back to generic. From
--trace llvmon that fixture:- 7
js_param_type_guardcall sites emitted. - 6 functions carry a
$spec_*/$genericpair:arrayTotal$spec_b,addFirst$spec_b_b,choose$spec_i32_b,render$spec_b,treeTotal$spec_b,surviveMovingGc$spec_b— each beside its$genericsibling. mutatePayloadhas no clone, which is the ineligibility rule firing on the parameter whose value the body invalidates. That negative is as important as the positives.
Instruction recovery: 55–93%, matching the unsound control
A/B against the preserved sweep binaries at
601a02d235af2d9af1841089426a1da379edac0a— the exact base of this branch, and the regressed state this issue was filed about. Same sources, same--no-cache/PERRY_NO_AUTO_OPTIMIZE=1invocation, same output basenames.instructions retired, which is load-independent, so a contended box does not distort it:benchmark base (regressed) guarded recovered unsound control (for comparison) pipeline 52,171,319,894 3,721,524,902 92.9% 92.9% asyncpipe 17,060,951,945 1,369,231,377 92.0% 92.5% retain_wide 18,446,669,808 2,471,890,150 86.6% — interp 64,502,116,097 12,725,435,786 80.3% 80.3% iso_miss 72,582,797,852 15,357,602,597 78.8% 78.9% shapes 5,141,727,141 1,291,612,836 74.9% — tree 31,779,662,666 14,261,587,660 55.1% 72.0% All 14 runs exited 0, and all 14 stdouts are byte-identical to the node-derived expected outputs.
The right-hand column is the diagnostic-only, deliberately unsound parameter-hint control recorded when this issue was filed. On
pipeline,interpandiso_missthe sound guarded implementation matches it to a tenth of a percent, and onasyncpipeto half a percent. That is the outcome you want: the guard is not costing measurable work relative to simply trusting the annotation.treeis the exception — 55.1% against the control's 72.0%. Worth understanding before this lands; it is the one row where guarding is visibly leaving recovery on the table.What is NOT done
I am reporting, not merging. Still outstanding:
- Full gap-suite regression. Only the one new test was run. A 28-file codegen change needs the whole suite.
- Unit tests for the new modules —
ordinary_param_guard_tests.rsand theperry-codegensuite have not been executed. - Wall clock on the quiet mini. Instructions measure work, not time, and are blind to idle and understate stalls. The mini numbers are what decide whether the standings recover.
- Forced-moving-GC instruments.
surviveMovingGcallocates hard, but it has not been run underPERRY_GC_ZEAL=1 PERRY_GC_PROTECT_FROMSPACE=1. A guarded clone that keeps a proof-bearing parameter live across 50k allocations is exactly the shape that wants that pairing. - Rebase. The branch is 10 commits behind
main(506f4ab11) butgit merge-treereports a clean merge. - No PR is open.
cargo check -p perry-codegen -p perry-runtimeis clean (exit 0, zero errors).Note on the artifacts
/private/tmp/perry-current-sweep-artifacts/is what made this A/B possible in minutes rather than a 30-minute baseline rebuild — it holds the corpus sources, the expected outputs and the baseline-compiledp_*binaries. Keep it until the mini re-sweep is done.- 7
- added a commit that references this issue
on Aug 14, 2026 Not ready to close — #8094 needs a design change before it can land, and I'd rather say so
than have it merge on the strength of the instruction-count numbers alone.The performance result is real and reproduces. What does not hold is the soundness argument.
js_param_type_guardvalidates the argument at entry, but the proof it licenses is
stamped for the entire clone body, and nothing re-establishes it after a call that can write
through an alias.collectors::has_any_mutation— the demotion gate atcodegen/mod.rs:2468
— only sees syntactic writes through the parameter, becauseexpr_has_mutationhas no
LocalGetarm. Sob.v = xinside the body demotes the parameter, and the identical write
inside a callee does not.This is reachable from ordinary TypeScript with no casts, because
anyis assignable to
number:interface Box { v: number; } const poison: any = "lie"; function assign(b: Box): void { b.v = poison; } function compute(b: Box): string { const before = b.v + 1; assign(b); const after = b.v + 1; return "before=" + before + " after=" + after + " typeof=" + typeof b.v; } console.log(compute({ v: 41 }));
Same runtime archive for every arm, exit 0 everywhere:
- node v26.5.1 —
before=42 after=lie1 typeof=string - base
601a02d23—before=42 after=lie1 typeof=string - the branch with
PERRY_SPECIALIZED_ABI=0—before=42 after=lie1 typeof=string - the branch, default —
before=42 after=lie typeof=string
An array-of-interface parameter reproduces the same way. The IR shows a bare
fadd double %field, 1.0inside the clone with no tag check; per #7773 anfaddon a NaN-box
propagates the payload, so the wrong-typed value survives the arithmetic intact and the program
prints a plausible-looking wrong answer instead ofNaN.Full evidence, the three-way control, the IR, and three candidate fixes (one of which is
already implemented in this same PR, inspec_return_proof.rs, just not applied to the codegen
body path) are in the PR comment.Separately,
lint— a required context — fails on the branch with four findings that
git diff origin/main..HEADshows are introduced by it. Every other script in thelintjob
passes at514ff7c9a, including the three #8092 reported red onmain.The
treerow from the sweep is also explained:param_guard.rs::build_namedrefuses class
types outright, so intree.tsonlybuild(depth: number)is guardable —count(t: Tree),
the hot recursive one, is not, and theTreeconstructor is refused at one remove because
Tree | nullis a union containing a class. That is a scope boundary, not an inherent limit;
the runtime already implements the class-identity check (class_chain_reaches,
param_type_guard.rs:399) and codegen simply never emits a non-zeroclass_id.- node v26.5.1 —
- added 2 commits that reference this issue
on Aug 14, 2026 Update: the soundness hole reported above is fixed on the PR branch, and the performance result
survives it.The fix is not the shape that was first proposed. Treating "the parameter was handed to a callee"
as the trigger closes only the route where we pass the reference; it does not close this one,
which I measured:let stash: any = null; function poison(): void { stash.v = "lie"; } function victim(b: Box): string { const before = b.v + 1; poison(); // b is NEVER passed anywhere return "before=" + before + " after=" + (b.v + 1); }
bnever appears in an argument list, so an escape analysis structurally cannot see it, yet the
callee reaches the same object through the global the caller stashed it in first. So the rule is
keyed on "did unknown code run", not "did the reference escape":guard_blocked[i] = body_contains_call(&f.body) && is_reference_like(&aliases, &p.ty, 0)
Primitive parameters stay eligible — a callee has no route to the caller's copy of a number or a
string — and that is why most of the win survives.Post-fix instruction recovery vs
601a02d23, all stdout byte-exact againstexpected/, exit 0:bench post-fix previously published delta pipeline 93.0% 92.9% +0.1 asyncpipe 92.4% 92.0% +0.4 retain_wide 85.6% 86.6% −1.0 shapes 75.2% 74.9% +0.3 interp 75.6% 80.3% −4.7 iso_miss 74.8% 78.8% −4.0 tree 55.1% 55.1% 0.0 The cost lands exactly where the mechanism predicts —
interpandiso_missare the two programs
whose hot parameters are object-typed (Parser,Env,Value) and passed into call-bearing
bodies. Everything else is within noise.Three regression rows are in
test_gap_specabi_ordinary_param_guards.tscovering the
passed-reference, array-element and never-passed shapes, sabotage-verified green → red → green
with a real rebuild each way. Details, including a gate that had quietly stopped testing anything
and how it was restored, are in the PR.Still open before this closes: the gap suite and the full
cargo testruns, which I stopped
rather than run on a box that was down to 8.8 GiB of disk. The class-typed-parameter gap behind
thetreerow is split out as #8099.Status: the mechanism you asked for is delivered; the outcome is not yet verified.
Two PRs have landed against this:
- perf(codegen): guarded ordinary-parameter specialization #8094 (merged) — the guarded ordinary-parameter specialized clone with the unchanged generic body as fallback. That is this issue's literal ask, and it explicitly does not revert fix(codegen): require runtime evidence for local binding types #8033 or seed the generic body from declared types.
- perf(codegen): let a specialized entry re-enter itself #8167 (merged) — lets a specialized entry re-enter itself. Without it the clone was reachable from one call site per module, so a recursive function took the fast path exactly once and every recursive step paid dynamic dispatch. On
fib(40): 211.54 G → 4.67 G instructions, 45.3x.
Why this stays open
A three-way sweep on the quiet mini (best-of-5, two independent runs, both
VERDICT: CLEAN) measured beats node 7/19 at38cf15336. That is better than the 2/19 this issue reports, and worse than the 10/19 it wants — but it was measured before #8167 landed, so the current number is unknown.The sweep's most useful result is that the oracle side is clean: every node cell is within ±3.2% of the 2026-08-12 table (16 of 19 within 1.3%) and scriptc 0.0.23 within ±1.5%. So the entire remaining gap is perry-side, not host drift or a toolchain change.
Nine rows are still above their 08-12 level, and they are NOT one mechanism
fib40,tree,tree_wide,churn,shapes,interp,cycles,deeplist,pipeline.Dynamic-arith call-site counts at
38cf15336are not uniform —fib40total,iso_miss67,interp64,shapes14,churn4,cycles4,tree3, anddeeplisthas 0 dynamic sites yet is +17.2% anyway. So at least one of those rows regressed for a reason unrelated to #8033's parameter-proof removal, and "fix the cliff" will not sweep them all up.Two reclassifications from the same sweep, both perry-side:
shapesbroke out of the vs-scriptc GC cluster: 1.54 → 2.27 while scriptc stayed at 0.038. It is now a specialization problem, not a high-survival-GC one.asyncpipeleft the vs-node cluster: 1.63 → 1.06, and it now runs 1 copying minor rather than zero, so the long-standing "pure mutator, GC levers irrelevant" characterisation is no longer exactly true.
What would close this
- Re-measure the perry arm post-perf(codegen): let a specialized entry re-enter itself #8167. Node and scriptc do not need re-running —
~/perry-sweep-0815/on the mini has their binaries,sources/,expected/and both results files; only newp_*binaries are needed. - Then attribute whatever remains per row rather than as a single cliff.
Related and deliberately separate: #8169 (the Tier-B half of the self-recursion defect, which is the shape most code hits —
tree,interp,shapeshavenumber-typed hot params with no literal call site), and #8171 (#8167's residual 22%, a range test that is provably unnecessary where a branch already narrows the parameter).Re-measured at
499e29627. #8167 is a single-row fix; the cliff is not closed.Quiet mini, best-of-5, two independent interleaved runs, both
VERDICT: CLEAN, 74/74 cells, zero failures.metric 08-12 38cf15336now beats node 9/19 7/19 8/19 beats scriptc31 14/19 12/19 13/19 within 1.3x node 15/19 11/19 12/19 Every one of those three gains is
fib40. No other row changed category.The nine regressed rows
bench vs 38cffib40−92.8% recovered most of it — but still +93.2% vs 08-12 tree,tree_wide,churn,shapes,interp,cycles,deeplist,pipeline−0.3% … +0.1% unmoved, all inside run-to-run noise (median spread 0.10%) So the guarded-clone work (#8094) plus self-recursion (#8167) has recovered exactly the shape whose entire hot body is the specialized entry. The other eight are untouched.
A negative result that should redirect the next attempt
Dynamic-arith call sites were re-counted at both arms:
bench @ 38cfnow iso_miss 67 45 interp 64 42 shapes 14 13 churn / cycles / tree / deeplist 4 / 4 / 3 / 0 unchanged interpandiso_misseach lost 22 dynamic-arith sites and got no faster at all — 0.8528 → 0.8524 and 0.9896 → 0.9901. Site count is not tracking cost on those rows.The residual is also highly concentrated:
parseApplyalone holds 30 of the ~43 remaining sites in both programs. Anyone attacking this cluster should start there rather than sweeping sites generally, and should not use site count as the success metric.And instructions are not tracking cost either
fib40's instructions fell 45.5x but wall time only 13.9x, because IPC collapsed 6.30 → 1.92 — the lowest in the corpus by a wide margin (others 4.5–6.5; node'sfib40is 4.98). Perry now retires 3.5x fewer instructions than node on that row and is only 1.36x faster. Filed separately as #8175; it is plausibly the same range test #8171 describes.RSS
Every row moved ≤0.2 MB. Nothing in this batch changes peak footprint, so #8122's held status is unaffected by it. Perry's peak RSS is below node on all 19 rows —
tree_wide63.5 vs 643.0 MB,tree31.4 vs 287.8 MB.Method notes worth keeping
- The oracle side re-reproduced as a control: node within ±1.9% (mean −0.04%, 17/19 within 1.5%), scriptc within ±3.6%. The delta column is perry-side.
- Version strings are useless for freshness — both arms report
0.5.1510. Discrimination came from commit plus archive sha256. ~/perry-sweep-0815/on the mini now holds the499e29627binaries with the previous arm preserved atperry-38cf/, so either can be re-run without a rebuild.
This stays open: eight of nine rows are unrecovered and the mechanism for them is not yet identified.
Re-measured at
bfb0707be, and the guard-blocked rule is measured at zero: recommend closing the rule-lifting direction1. The base re-sweep
bfb0707beis instruction-identical to499e29627on all 19 rows (max |Δ| 0.54%, most 0.0%; runs 5+6 on the mini, run6VERDICT: CLEAN, oracle cells within ±3.7% of runs 3/4). Beats-node still 8/19. The 16 intervening commits — #8165, #8183, #8188 included — moved nothing on this corpus. The eight regressed rows stand exactly where the last comment left them.2. Lifting
guard_blockedis worth 0.0% — measured, not arguedFour one-compiler diagnostic arms (env-var-gated at TS-compile time, never shipped; all 19 binaries byte-exact vs
expected/in every arm; instructions retired, best-of-5):bench base SKIPVAL (validators free) UNGUARD (rule lifted + validators free) SEED (declared types proven everywhere — the causal-control shape) tree 15.821B −33.82% −33.80% −0.25% tree_wide 31.587B −17.00% −17.00% +0.02% interp 15.053B −19.59% −19.58% −8.75% iso_miss 17.698B −16.35% −16.33% −7.29% asyncpipe 1.293B −0.01% −1.01% −1.09% churn / shapes / cycles / deeplist / pipeline / fib40 noise noise noise UNGUARD − SKIPVAL ≈ +0.0% on every regressed row: even with hypothetically free guards, the proofs the rule refuses buy nothing the generic body doesn't already have. (Only asyncpipe −1.0% — a row already at 1.05x node.) And SEED — the absolute ceiling of parameter evidence, the shape of the original
601a02d23control — is worse than SKIPVAL on interp/iso_miss and flat on tree: what that control appeared to buy on these rows is now mostly guard elision via static routing, not better lowering. The 2026-08-12-era value of reference-param evidence has since been absorbed by #8094's primitives plus the generic body's own class-field paths (#8165's identity-only revert measured the same thing from the other side).The census agrees with the zero: the whole corpus contains 21 guard-blocked params over 17 functions in 5 programs (
count(t)×2, the interp/iso_miss parser+eval family ×8 each, asyncpipevalidate/enrich×3). cycles, deeplist, churn have none; pipeline's hot callables are capture-bearing arrows (the #8103 gate, not this rule) and generic-class methods; shapes' hot path is virtual method dispatch. Those five rows regressed for reasons this issue's mechanism cannot reach.The soundness side is equally closed: escape analysis is structurally dead for parameters (the
poison()global-alias case, now a comment incollectors/mutation.rs); the call-surviving "identity-only" split was measured and reverted in #8165 (tree1.089 → 1.646 s for nothing); per-callsite invalidation is L-sized (53 flow-insensitive consumer sites across 13 files plus a pre-pass) with this measured ceiling of zero; and the field-bearing descriptors the rule refuses are unpayable anyway —js_param_type_guarddeep-walks the value, O(reachable heap) per call onTree/Node/Env. The rule is correct as written.3. Where the instructions actually were: the guards Perry already emits
The SKIPVAL column is the real finding: the interpretive validator itself was 34% of
tree, 20% ofinterp, 17% oftree_wide, 16% ofiso_miss— ~450 instructions of fixed cost per call (descriptor parse + 768-byteGuardStateinit) on every unproven call through the public wrapper, including a Tier-B clone's own recursion (build$spec_bre-enters via the wrapper — #8169's exact shape, now priced).- PR perf(codegen): decide scalar parameter descriptors with the typed-abi leaf guards #8201 (validated, ready): scalar descriptors decided by the predicate-identical typed-abi leaf guards.
tree−31.8% instr / −46.4% wall,tree_wide−15.9%/−23.1%,interp−8.3%/−9.1%,iso_miss−6.9%/−6.5%, RSS flat, everything else ±0.4%; confirmed on the mini (run7). - A guarded $spec_b clone never re-enters itself either — the Tier-B half of #8167 #8169: the remaining tree/tree_wide share (wrapper round-trip on Tier-B self-recursion, ~2pp on tree after perf(codegen): decide scalar parameter descriptors with the typed-abi leaf guards #8201).
- perf(runtime): js_param_type_guard's interpretive walk dominates the remaining ~11pp (fixed per-call overhead now largely removed) #8202 (new, priced): the structural share —
asNum's union /peek's object walks, −11.3pp interp / −9.4pp iso_miss still available; candidate shapes listed there.
4. What this issue should become
The literal ask ("recover guarded ordinary-parameter specialization") is delivered as far as it can pay: #8094+#8167 recovered the primitive cliff, and the reference-param remainder is measured at zero against a free-guard ceiling. I recommend closing this issue in favor of the per-row residuals: #8169 + #8202 (tree, tree_wide, interp, iso_miss), and fresh per-mechanism attribution for shapes / cycles / deeplist / churn / pipeline, which no ordinary-parameter mechanism touches — deeplist and pipeline compile to byte-identical binaries even under the SEED ceiling.
- PR perf(codegen): decide scalar parameter descriptors with the typed-abi leaf guards #8201 (validated, ready): scalar descriptors decided by the predicate-identical typed-abi leaf guards.
- added a commit that references this issue
on Aug 16, 2026 Closing: the
guard_blockedrule is correct as written. Lifting it is worth 0.0%, and the instructions this issue was chasing were in the guards Perry already emits — now fixed in #8201.Step 1 — re-sweep, because every number here was stale
bfb0707beis instruction-identical to499e29627on all 19 rows (max |Δ| 0.54%). The 16 intervening commits — #8165, #8183, #8188 included — moved nothing. The owner's earlier verdict ("every gain is fib40; the other eight rows are inside noise") holds at today's base. Beats-node is still 8/19.Step 2 — the discriminating measurement
Four diagnostic arms from one compiler, env-gated, all byte-exact output, instructions retired best-of-5:
bench base SKIPVAL (validators → true)UNGUARD (rule lifted + guards made free) SEED (declared types proven everywhere) tree15.82 B −33.8% −33.8% −0.3% tree_wide31.59 B −17.0% −17.0% +0.0% interp15.05 B −19.6% −19.6% −8.8% iso_miss17.70 B −16.3% −16.3% −7.3% others noise noise noise UNGUARD − SKIPVAL = 0.0%on every regressed row. Even with the rule lifted and the guards made free, the proofs it refuses buy nothing. AndSEED— the causal-control shape, the ceiling of any conceivable scheme — is worse thanSKIPVALoninterp/iso_miss, which means that control's apparent value was mostly guard elision, not better lowering.Census: 21 blocked params across 17 functions in 5 programs corpus-wide, and zero in
cycles/deeplist/churn.pipelineis closures (#8103) plus generic methods;shapesis virtual dispatch;deeplistandpipelinecompile byte-identical even under SEED.So all three candidate directions are dead against a measured ceiling of zero: escape/mutation analysis on the parameter, identity-only descriptors (#8165 — already measured at −51% and still not the lever here), and per-callsite invalidation (L-sized: 53 flow-insensitive consumer sites plus a pre-pass at
codegen/function.rs:897). None of them should be built.Step 3 — what was actually costing 16–34%
js_param_type_guarditself: ~450 instructions of fixed cost per call — descriptor parse plus a 768-byteGuardStateinit (crates/perry-runtime/src/param_type_guard.rs) — on every unproven call through the public wrapper, including Tier-B clones re-entering themselves through it (#8169's shape, now priced).#8201 decides single-node scalar descriptors with the predicate-identical typed-abi leaf guards (
codegen/param_guard.rs::scalar_descriptor_rep). Sound by bit-for-bit predicate equality, so routing is unchanged for honest and lying callers alike — verified with a lying-caller probe and the fulltest_gap_specabi_ordinary_param_guards.ts, both byte-exact against Node 26.5.1.Measured, local and confirmed on the quiet mini:
row instructions wall peak RSS tree−31.8% −46.4% unchanged tree_wide−15.9% −23.1% unchanged interp−8.3% −9.1% unchanged iso_miss−6.9% −6.5% unchanged all others ±0.4% unchanged No unrelated-row movement, which matters here — this repo has repeatedly found that "more type visibility" changes silently un-gate latent fast paths, so a win somewhere unexpected would have been a red flag rather than a bonus.
19/19 corpus byte-exact on every arm; codegen suite reproduces exactly the 9 pre-existing failures by name against a baseline re-run on pristine
bfb0707be;perry-runtime --lib2499/0;perry --bin perry987/0; all 12 script gates green. Diagnostic knobs were stripped before the PR.What remains, priced
- A guarded $spec_b clone never re-enters itself either — the Tier-B half of #8167 #8169 — Tier-B self-recursion wrapper round-trip, the remaining ~2 pp on
tree/tree_wide. - perf(runtime): js_param_type_guard's interpretive walk dominates the remaining ~11pp (fixed per-call overhead now largely removed) #8202 — structural-validator fixed overhead (
asNumunion,peekobject): −11.3 pp oninterp, −9.4 pp oniso_missstill available. shapes/cycles/deeplist/churn/pipelineneed per-mechanism attribution outside the ordinary-parameter machinery entirely.
Closing this issue: its stated direction is refuted, and the instructions it was after have been recovered by a different and much smaller change.
- A guarded $spec_b clone never re-enters itself either — the Tier-B half of #8167 #8169 — Tier-B self-recursion wrapper round-trip, the remaining ~2 pp on
Summary
PR #8033 correctly stopped treating erased TypeScript binding annotations as runtime representation proofs, but it also removed ordinary function-parameter evidence from generic function bodies. That causes a large, reproducible specialization cliff across the current performance corpus.
Do not revert #8033 or seed the generic body from declared parameter types. The correctness constraint from #7846 is real. We need a guarded parameter-specialized clone (or equivalent entry guard) with the unchanged generic body as fallback.
Current-main regression
Measured at exact main
601a02d235af2d9af1841089426a1da379edac0a.The clean M1 mini sweep regressed from the previously recorded
843ef621fsweep as follows:Largest wall-time changes on the clean mini:
843ef621fThe current timings also reproduce the older conservative-codegen
regbaseacross 15 common rows: geomean current/regbase is1.00003x, with a maximum row difference of 1.24%.Causal A/B
I built a diagnostic-only compiler from exact current main that changed only ordinary
compile_functionadmission:local_typesremained unchanged;proven_local_typeswas seeded fromf.paramsonly;This control is intentionally unsound for lying annotations and is not a proposed patch. It exists only to identify the lost optimization surface.
Every tested benchmark retained byte-exact output. Same-host retired instructions were:
interpwas repeated three times in each arm. Current was 64.455-64.478B instructions; the parameter-only control was 12.684-12.699B. The control returns to the previously recorded roughly 11.5-14.1B instruction band.This proves that lost ordinary-parameter type evidence is the dominant cause of the headline regression.
Mechanism
In current
codegen/function.rs:f.paramsstill populatelocal_types.FnCtxinitializesproven_local_typesempty.stable_local_type_proof()deliberately reads onlyproven_local_typesand rejects reassigned locals.This is especially costly for the interpreter corpus. Hot recursive functions such as
evalNode(n: Node, env: Env),lookup(env: Env, name: string), parser functions takingParser, andlex(src: string)lose typed property, array, string, and numeric lowering at every use.Current
interp --opt-report=jsonrecords 8 selections and 64 denials, including:The diagnostic control does not repair pointer-shape clone routing (it remains 0/49). Its performance recovery comes from letting the ordinary body lowering consume the parameter type again.
Ruled out
PERRY_RS4GC=0shadow-root control did not recover performance.asyncpiperetired 17.08B vs 17.07B instructions;interpworsened to 71.30B from 64.63B.Required soundness boundary
TypeScript annotations are erased and may lie. #7846 demonstrates that a declared scalar, class, array, or capture type cannot by itself select an unguarded representation or behavior. The generic body and GC rooting must remain conservative.
A sound recovery should therefore build on guarded clone-and-route machinery:
proven_local_typesonly inside the successfully guarded clone.Node/Envshapes here.This likely belongs with the clone-and-route work already referenced by the pointer-shape report as #7034 section 1.
Acceptance
test_gap_7846_local_binding_type_proofs.ts, pass through the generic fallback with Node-equivalent output.interp,iso_miss,pipeline,asyncpipe, andtreeare remeasured on the clean mini with exact-output gates.Measurement provenance
Perry was built with:
Compiles pinned the exact release directory with
PERRY_RUNTIME_DIR,PERRY_NO_AUTO_OPTIMIZE=1,PERRY_NO_CACHE=1, and--no-cache. Nothing was built on the mini. The accepted mini cells were warmed and then measured best-of-five in interleaved order with exit/output validation on every run.