Skip to content

perf(codegen): recover guarded ordinary-parameter specialization lost by #8033 #8079

Description

@proggeramlug

Summary

PR #8033 correctly stopped treating erased TypeScript binding annotations as runtime representation proofs, but it also removed ordinary function-parameter evidence from generic function bodies. That causes a large, reproducible specialization cliff across the current performance corpus.

Do not revert #8033 or seed the generic body from declared parameter types. The correctness constraint from #7846 is real. We need a guarded parameter-specialized clone (or equivalent entry guard) with the unchanged generic body as fallback.

Current-main regression

Measured at exact main 601a02d235af2d9af1841089426a1da379edac0a.

The clean M1 mini sweep regressed from the previously recorded 843ef621f sweep as follows:

  • Perry/Node wins: 10/19 -> 2/19
  • Perry/scriptc wins: 14/19 -> 8/18
  • Perry geomean across all 19 rows: 3.480x slower
  • clean and separately dirty corroboration agreed within 0.2%

Largest wall-time changes on the clean mini:

benchmark 843ef621f current main change
pipeline 0.1630 s 3.3585 s +1960%
asyncpipe 0.0980 s 1.1216 s +1045%
retain_wide 0.1840 s 1.0826 s +488%
shapes 0.0550 s 0.3183 s +479%
interp 0.8090 s 3.9931 s +394%
iso_miss 0.9380 s 4.4743 s +377%

The current timings also reproduce the older conservative-codegen regbase across 15 common rows: geomean current/regbase is 1.00003x, with a maximum row difference of 1.24%.

Causal A/B

I built a diagnostic-only compiler from exact current main that changed only ordinary compile_function admission:

  • local_types remained unchanged;
  • proven_local_types was seeded from f.params only;
  • module globals and local initializers were excluded;
  • runtime, stdlib, benchmark input, compile flags, and execution environment were otherwise held fixed.

This control is intentionally unsound for lying annotations and is not a proposed patch. It exists only to identify the lost optimization surface.

Every tested benchmark retained byte-exact output. Same-host retired instructions were:

benchmark current main parameter-hint control reduction current/control
asyncpipe 17,054,037,641 1,279,430,024 92.5% 13.329x
interp 64,455,747,802 12,683,767,308 80.3% 5.082x
iso_miss 72,572,449,776 15,332,837,352 78.9% 4.733x
pipeline 52,150,730,952 3,714,983,318 92.9% 14.038x
tree 31,782,383,093 8,896,702,263 72.0% 3.572x

interp was repeated three times in each arm. Current was 64.455-64.478B instructions; the parameter-only control was 12.684-12.699B. The control returns to the previously recorded roughly 11.5-14.1B instruction band.

This proves that lost ordinary-parameter type evidence is the dominant cause of the headline regression.

Mechanism

In current codegen/function.rs:

  1. f.params still populate local_types.
  2. The generic FnCtx initializes proven_local_types empty.
  3. stable_local_type_proof() deliberately reads only proven_local_types and rejects reassigned locals.
  4. Therefore an ordinary typed parameter is unproven throughout the generic body unless a separate guarded clone supplies evidence.

This is especially costly for the interpreter corpus. Hot recursive functions such as evalNode(n: Node, env: Env), lookup(env: Env, name: string), parser functions taking Parser, and lex(src: string) lose typed property, array, string, and numeric lowering at every use.

Current interp --opt-report=json records 8 selections and 64 denials, including:

  • pointer-shape: 0 selected / 49 denied
  • canonical-slot: 7 selected / 3 denied
  • specialized ABI: 1 selected / 12 denied

The diagnostic control does not repair pointer-shape clone routing (it remains 0/49). Its performance recovery comes from letting the ordinary body lowering consume the parameter type again.

Ruled out

  • Mini noise: independent clean/dirty sweeps agree within 0.2%.
  • Native-root GC overhead / Propagate native-root decisions to codegen-unit workers #8071: a compile-time PERRY_RS4GC=0 shadow-root control did not recover performance. asyncpipe retired 17.08B vs 17.07B instructions; interp worsened to 71.30B from 64.63B.
  • Output/correctness differences: Perry remained byte-exact on all 19 sweep rows; the five causal-control programs above were byte-exact in both arms.

Required soundness boundary

TypeScript annotations are erased and may lie. #7846 demonstrates that a declared scalar, class, array, or capture type cannot by itself select an unguarded representation or behavior. The generic body and GC rooting must remain conservative.

A sound recovery should therefore build on guarded clone-and-route machinery:

  1. Select eligible ordinary parameter tuples without changing the generic entry/body semantics.
  2. Emit runtime guards for the complete tuple before entering a specialized body clone.
  3. Populate proven_local_types only inside the successfully guarded clone.
  4. Preserve the generic boxed fallback for failed guards and indirect/unknown callers.
  5. Route direct call sites only when their current argument facts prove the same tuple.
  6. Support recursive calls from a guarded clone without reintroducing annotation trust.
  7. Add object/array/discriminated-union parameter coverage; existing specialized ABI coverage is primarily scalar/string/typed-array and does not cover the hot Node/Env shapes here.

This likely belongs with the clone-and-route work already referenced by the pointer-shape report as #7034 section 1.

Acceptance

  • Lying-annotation regressions, including test_gap_7846_local_binding_type_proofs.ts, pass through the generic fallback with Node-equivalent output.
  • Positive tests prove the runtime guard is present and the specialized clone receives parameter proofs only after that guard.
  • Negative tests prove a wrong scalar/object/array/union value cannot enter the clone.
  • Direct, indirect, and recursive call paths retain a valid generic fallback.
  • Precise-root/forced-moving verification remains green for both clone and fallback paths.
  • interp, iso_miss, pipeline, asyncpipe, and tree are remeasured on the clean mini with exact-output gates.
  • The recovered instruction count is structurally pinned or otherwise guarded so this cliff cannot silently recur.

Measurement provenance

Perry was built with:

cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static

Compiles pinned the exact release directory with PERRY_RUNTIME_DIR, PERRY_NO_AUTO_OPTIMIZE=1, PERRY_NO_CACHE=1, and --no-cache. Nothing was built on the mini. The accepted mini cells were warmed and then measured best-of-five in interleaved order with exit/output validation on every run.

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Aug 14, 2026
  2. self-assigned this
    on Aug 14, 2026
  3. proggeramlug commented on Aug 14, 2026

    @proggeramlug
    ContributorAuthor

    Status: the in-flight branch works, and it recovers nearly all of the regression soundly

    The perf/issue-8079-guarded-param-specialization worktree had ~2,900 lines of new code across 34 files sitting uncommitted with no active owner. I committed it as-is to the branch first (2f70cd7bb, unvalidated preservation commit — git reset --soft HEAD~1 restores the exact working-tree state) and then validated it. Nothing was rebased or edited.

    Acceptance test: passes byte-exact

    test-files/test_gap_specabi_ordinary_param_guards.ts against the pinned oracle (node v26.5.1, matching .node-version):

    PASS  test_gap_specabi_ordinary_param_guards
    Parity Pass: 1   Parity Fail: 0   Compile Fail: 0   Crashed: 0   Skipped: 0
    Parity Rate: 100.0%      (harness exit 0)
    

    The test covers every case this issue asked for: annotation lies on array elements, array container, scalar, bare number, object fields, class instance, and both flat and nested discriminated-union positions; an escaped/indirect call through const escapedArrayTotal = arrayTotal; an accessor-backed property with a getter-hit count; a parameter the body mutates; a recursive generic (treeTotal); and a 50,000-allocation loop exercised on both the clone and the failed-guard fallback.

    The clone is actually live — a pass alone would be vacuous

    A guard test passes trivially if the optimization never engages and everything falls back to generic. From --trace llvm on that fixture:

    • 7 js_param_type_guard call sites emitted.
    • 6 functions carry a $spec_* / $generic pair: arrayTotal$spec_b, addFirst$spec_b_b, choose$spec_i32_b, render$spec_b, treeTotal$spec_b, surviveMovingGc$spec_b — each beside its $generic sibling.
    • mutatePayload has no clone, which is the ineligibility rule firing on the parameter whose value the body invalidates. That negative is as important as the positives.

    Instruction recovery: 55–93%, matching the unsound control

    A/B against the preserved sweep binaries at 601a02d235af2d9af1841089426a1da379edac0a — the exact base of this branch, and the regressed state this issue was filed about. Same sources, same --no-cache / PERRY_NO_AUTO_OPTIMIZE=1 invocation, same output basenames. instructions retired, which is load-independent, so a contended box does not distort it:

    benchmark base (regressed) guarded recovered unsound control (for comparison)
    pipeline 52,171,319,894 3,721,524,902 92.9% 92.9%
    asyncpipe 17,060,951,945 1,369,231,377 92.0% 92.5%
    retain_wide 18,446,669,808 2,471,890,150 86.6% —
    interp 64,502,116,097 12,725,435,786 80.3% 80.3%
    iso_miss 72,582,797,852 15,357,602,597 78.8% 78.9%
    shapes 5,141,727,141 1,291,612,836 74.9% —
    tree 31,779,662,666 14,261,587,660 55.1% 72.0%

    All 14 runs exited 0, and all 14 stdouts are byte-identical to the node-derived expected outputs.

    The right-hand column is the diagnostic-only, deliberately unsound parameter-hint control recorded when this issue was filed. On pipeline, interp and iso_miss the sound guarded implementation matches it to a tenth of a percent, and on asyncpipe to half a percent. That is the outcome you want: the guard is not costing measurable work relative to simply trusting the annotation.

    tree is the exception — 55.1% against the control's 72.0%. Worth understanding before this lands; it is the one row where guarding is visibly leaving recovery on the table.

    What is NOT done

    I am reporting, not merging. Still outstanding:

    • Full gap-suite regression. Only the one new test was run. A 28-file codegen change needs the whole suite.
    • Unit tests for the new modules — ordinary_param_guard_tests.rs and the perry-codegen suite have not been executed.
    • Wall clock on the quiet mini. Instructions measure work, not time, and are blind to idle and understate stalls. The mini numbers are what decide whether the standings recover.
    • Forced-moving-GC instruments. surviveMovingGc allocates hard, but it has not been run under PERRY_GC_ZEAL=1 PERRY_GC_PROTECT_FROMSPACE=1. A guarded clone that keeps a proof-bearing parameter live across 50k allocations is exactly the shape that wants that pairing.
    • Rebase. The branch is 10 commits behind main (506f4ab11) but git merge-tree reports a clean merge.
    • No PR is open.

    cargo check -p perry-codegen -p perry-runtime is clean (exit 0, zero errors).

    Note on the artifacts

    /private/tmp/perry-current-sweep-artifacts/ is what made this A/B possible in minutes rather than a 30-minute baseline rebuild — it holds the corpus sources, the expected outputs and the baseline-compiled p_* binaries. Keep it until the mini re-sweep is done.

  4. proggeramlug commented on Aug 14, 2026

    @proggeramlug
    ContributorAuthor

    Not ready to close — #8094 needs a design change before it can land, and I'd rather say so
    than have it merge on the strength of the instruction-count numbers alone.

    The performance result is real and reproduces. What does not hold is the soundness argument.
    js_param_type_guard validates the argument at entry, but the proof it licenses is
    stamped for the entire clone body, and nothing re-establishes it after a call that can write
    through an alias. collectors::has_any_mutation — the demotion gate at codegen/mod.rs:2468
    — only sees syntactic writes through the parameter, because expr_has_mutation has no
    LocalGet arm. So b.v = x inside the body demotes the parameter, and the identical write
    inside a callee does not.

    This is reachable from ordinary TypeScript with no casts, because any is assignable to
    number:

    interface Box { v: number; }
    const poison: any = "lie";
    function assign(b: Box): void { b.v = poison; }
    function compute(b: Box): string {
      const before = b.v + 1;
      assign(b);
      const after = b.v + 1;
      return "before=" + before + " after=" + after + " typeof=" + typeof b.v;
    }
    console.log(compute({ v: 41 }));

    Same runtime archive for every arm, exit 0 everywhere:

    • node v26.5.1 — before=42 after=lie1 typeof=string
    • base 601a02d23 — before=42 after=lie1 typeof=string
    • the branch with PERRY_SPECIALIZED_ABI=0 — before=42 after=lie1 typeof=string
    • the branch, default — before=42 after=lie typeof=string

    An array-of-interface parameter reproduces the same way. The IR shows a bare
    fadd double %field, 1.0 inside the clone with no tag check; per #7773 an fadd on a NaN-box
    propagates the payload, so the wrong-typed value survives the arithmetic intact and the program
    prints a plausible-looking wrong answer instead of NaN.

    Full evidence, the three-way control, the IR, and three candidate fixes (one of which is
    already implemented in this same PR, in spec_return_proof.rs, just not applied to the codegen
    body path) are in the PR comment.

    Separately, lint — a required context — fails on the branch with four findings that
    git diff origin/main..HEAD shows are introduced by it. Every other script in the lint job
    passes at 514ff7c9a, including the three #8092 reported red on main.

    The tree row from the sweep is also explained: param_guard.rs::build_named refuses class
    types outright, so in tree.ts only build(depth: number) is guardable — count(t: Tree),
    the hot recursive one, is not, and the Tree constructor is refused at one remove because
    Tree | null is a union containing a class. That is a scope boundary, not an inherent limit;
    the runtime already implements the class-identity check (class_chain_reaches,
    param_type_guard.rs:399) and codegen simply never emits a non-zero class_id.

  5. proggeramlug commented on Aug 14, 2026

    @proggeramlug
    ContributorAuthor

    Update: the soundness hole reported above is fixed on the PR branch, and the performance result
    survives it.

    The fix is not the shape that was first proposed. Treating "the parameter was handed to a callee"
    as the trigger closes only the route where we pass the reference; it does not close this one,
    which I measured:

    let stash: any = null;
    function poison(): void { stash.v = "lie"; }
    function victim(b: Box): string {
      const before = b.v + 1;
      poison();                 // b is NEVER passed anywhere
      return "before=" + before + " after=" + (b.v + 1);
    }

    b never appears in an argument list, so an escape analysis structurally cannot see it, yet the
    callee reaches the same object through the global the caller stashed it in first. So the rule is
    keyed on "did unknown code run", not "did the reference escape":

    guard_blocked[i] = body_contains_call(&f.body) && is_reference_like(&aliases, &p.ty, 0)

    Primitive parameters stay eligible — a callee has no route to the caller's copy of a number or a
    string — and that is why most of the win survives.

    Post-fix instruction recovery vs 601a02d23, all stdout byte-exact against expected/, exit 0:

    bench post-fix previously published delta
    pipeline 93.0% 92.9% +0.1
    asyncpipe 92.4% 92.0% +0.4
    retain_wide 85.6% 86.6% −1.0
    shapes 75.2% 74.9% +0.3
    interp 75.6% 80.3% −4.7
    iso_miss 74.8% 78.8% −4.0
    tree 55.1% 55.1% 0.0

    The cost lands exactly where the mechanism predicts — interp and iso_miss are the two programs
    whose hot parameters are object-typed (Parser, Env, Value) and passed into call-bearing
    bodies. Everything else is within noise.

    Three regression rows are in test_gap_specabi_ordinary_param_guards.ts covering the
    passed-reference, array-element and never-passed shapes, sabotage-verified green → red → green
    with a real rebuild each way. Details, including a gate that had quietly stopped testing anything
    and how it was restored, are in the PR.

    Still open before this closes: the gap suite and the full cargo test runs, which I stopped
    rather than run on a box that was down to 8.8 GiB of disk. The class-typed-parameter gap behind
    the tree row is split out as #8099.

  6. proggeramlug commented on Aug 15, 2026

    @proggeramlug
    ContributorAuthor

    Status: the mechanism you asked for is delivered; the outcome is not yet verified.

    Two PRs have landed against this:

    Why this stays open

    A three-way sweep on the quiet mini (best-of-5, two independent runs, both VERDICT: CLEAN) measured beats node 7/19 at 38cf15336. That is better than the 2/19 this issue reports, and worse than the 10/19 it wants — but it was measured before #8167 landed, so the current number is unknown.

    The sweep's most useful result is that the oracle side is clean: every node cell is within ±3.2% of the 2026-08-12 table (16 of 19 within 1.3%) and scriptc 0.0.23 within ±1.5%. So the entire remaining gap is perry-side, not host drift or a toolchain change.

    Nine rows are still above their 08-12 level, and they are NOT one mechanism

    fib40, tree, tree_wide, churn, shapes, interp, cycles, deeplist, pipeline.

    Dynamic-arith call-site counts at 38cf15336 are not uniform — fib40 total, iso_miss 67, interp 64, shapes 14, churn 4, cycles 4, tree 3, and deeplist has 0 dynamic sites yet is +17.2% anyway. So at least one of those rows regressed for a reason unrelated to #8033's parameter-proof removal, and "fix the cliff" will not sweep them all up.

    Two reclassifications from the same sweep, both perry-side:

    • shapes broke out of the vs-scriptc GC cluster: 1.54 → 2.27 while scriptc stayed at 0.038. It is now a specialization problem, not a high-survival-GC one.
    • asyncpipe left the vs-node cluster: 1.63 → 1.06, and it now runs 1 copying minor rather than zero, so the long-standing "pure mutator, GC levers irrelevant" characterisation is no longer exactly true.

    What would close this

    1. Re-measure the perry arm post-perf(codegen): let a specialized entry re-enter itself #8167. Node and scriptc do not need re-running — ~/perry-sweep-0815/ on the mini has their binaries, sources/, expected/ and both results files; only new p_* binaries are needed.
    2. Then attribute whatever remains per row rather than as a single cliff.

    Related and deliberately separate: #8169 (the Tier-B half of the self-recursion defect, which is the shape most code hits — tree, interp, shapes have number-typed hot params with no literal call site), and #8171 (#8167's residual 22%, a range test that is provably unnecessary where a branch already narrows the parameter).

  7. proggeramlug commented on Aug 15, 2026

    @proggeramlug
    ContributorAuthor

    Re-measured at 499e29627. #8167 is a single-row fix; the cliff is not closed.

    Quiet mini, best-of-5, two independent interleaved runs, both VERDICT: CLEAN, 74/74 cells, zero failures.

    metric 08-12 38cf15336 now
    beats node 9/19 7/19 8/19
    beats scriptc31 14/19 12/19 13/19
    within 1.3x node 15/19 11/19 12/19

    Every one of those three gains is fib40. No other row changed category.

    The nine regressed rows

    bench vs 38cf
    fib40 −92.8% recovered most of it — but still +93.2% vs 08-12
    tree, tree_wide, churn, shapes, interp, cycles, deeplist, pipeline −0.3% … +0.1% unmoved, all inside run-to-run noise (median spread 0.10%)

    So the guarded-clone work (#8094) plus self-recursion (#8167) has recovered exactly the shape whose entire hot body is the specialized entry. The other eight are untouched.

    A negative result that should redirect the next attempt

    Dynamic-arith call sites were re-counted at both arms:

    bench @38cf now
    iso_miss 67 45
    interp 64 42
    shapes 14 13
    churn / cycles / tree / deeplist 4 / 4 / 3 / 0 unchanged

    interp and iso_miss each lost 22 dynamic-arith sites and got no faster at all — 0.8528 → 0.8524 and 0.9896 → 0.9901. Site count is not tracking cost on those rows.

    The residual is also highly concentrated: parseApply alone holds 30 of the ~43 remaining sites in both programs. Anyone attacking this cluster should start there rather than sweeping sites generally, and should not use site count as the success metric.

    And instructions are not tracking cost either

    fib40's instructions fell 45.5x but wall time only 13.9x, because IPC collapsed 6.30 → 1.92 — the lowest in the corpus by a wide margin (others 4.5–6.5; node's fib40 is 4.98). Perry now retires 3.5x fewer instructions than node on that row and is only 1.36x faster. Filed separately as #8175; it is plausibly the same range test #8171 describes.

    RSS

    Every row moved ≤0.2 MB. Nothing in this batch changes peak footprint, so #8122's held status is unaffected by it. Perry's peak RSS is below node on all 19 rows — tree_wide 63.5 vs 643.0 MB, tree 31.4 vs 287.8 MB.

    Method notes worth keeping

    • The oracle side re-reproduced as a control: node within ±1.9% (mean −0.04%, 17/19 within 1.5%), scriptc within ±3.6%. The delta column is perry-side.
    • Version strings are useless for freshness — both arms report 0.5.1510. Discrimination came from commit plus archive sha256.
    • ~/perry-sweep-0815/ on the mini now holds the 499e29627 binaries with the previous arm preserved at perry-38cf/, so either can be re-run without a rebuild.

    This stays open: eight of nine rows are unrecovered and the mechanism for them is not yet identified.

  8. proggeramlug commented on Aug 16, 2026

    @proggeramlug
    ContributorAuthor

    Re-measured at bfb0707be, and the guard-blocked rule is measured at zero: recommend closing the rule-lifting direction

    1. The base re-sweep

    bfb0707be is instruction-identical to 499e29627 on all 19 rows (max |Δ| 0.54%, most 0.0%; runs 5+6 on the mini, run6 VERDICT: CLEAN, oracle cells within ±3.7% of runs 3/4). Beats-node still 8/19. The 16 intervening commits — #8165, #8183, #8188 included — moved nothing on this corpus. The eight regressed rows stand exactly where the last comment left them.

    2. Lifting guard_blocked is worth 0.0% — measured, not argued

    Four one-compiler diagnostic arms (env-var-gated at TS-compile time, never shipped; all 19 binaries byte-exact vs expected/ in every arm; instructions retired, best-of-5):

    bench base SKIPVAL (validators free) UNGUARD (rule lifted + validators free) SEED (declared types proven everywhere — the causal-control shape)
    tree 15.821B −33.82% −33.80% −0.25%
    tree_wide 31.587B −17.00% −17.00% +0.02%
    interp 15.053B −19.59% −19.58% −8.75%
    iso_miss 17.698B −16.35% −16.33% −7.29%
    asyncpipe 1.293B −0.01% −1.01% −1.09%
    churn / shapes / cycles / deeplist / pipeline / fib40 noise noise noise

    UNGUARD − SKIPVAL ≈ +0.0% on every regressed row: even with hypothetically free guards, the proofs the rule refuses buy nothing the generic body doesn't already have. (Only asyncpipe −1.0% — a row already at 1.05x node.) And SEED — the absolute ceiling of parameter evidence, the shape of the original 601a02d23 control — is worse than SKIPVAL on interp/iso_miss and flat on tree: what that control appeared to buy on these rows is now mostly guard elision via static routing, not better lowering. The 2026-08-12-era value of reference-param evidence has since been absorbed by #8094's primitives plus the generic body's own class-field paths (#8165's identity-only revert measured the same thing from the other side).

    The census agrees with the zero: the whole corpus contains 21 guard-blocked params over 17 functions in 5 programs (count(t) ×2, the interp/iso_miss parser+eval family ×8 each, asyncpipe validate/enrich ×3). cycles, deeplist, churn have none; pipeline's hot callables are capture-bearing arrows (the #8103 gate, not this rule) and generic-class methods; shapes' hot path is virtual method dispatch. Those five rows regressed for reasons this issue's mechanism cannot reach.

    The soundness side is equally closed: escape analysis is structurally dead for parameters (the poison() global-alias case, now a comment in collectors/mutation.rs); the call-surviving "identity-only" split was measured and reverted in #8165 (tree 1.089 → 1.646 s for nothing); per-callsite invalidation is L-sized (53 flow-insensitive consumer sites across 13 files plus a pre-pass) with this measured ceiling of zero; and the field-bearing descriptors the rule refuses are unpayable anyway — js_param_type_guard deep-walks the value, O(reachable heap) per call on Tree/Node/Env. The rule is correct as written.

    3. Where the instructions actually were: the guards Perry already emits

    The SKIPVAL column is the real finding: the interpretive validator itself was 34% of tree, 20% of interp, 17% of tree_wide, 16% of iso_miss — ~450 instructions of fixed cost per call (descriptor parse + 768-byte GuardState init) on every unproven call through the public wrapper, including a Tier-B clone's own recursion (build$spec_b re-enters via the wrapper — #8169's exact shape, now priced).

    4. What this issue should become

    The literal ask ("recover guarded ordinary-parameter specialization") is delivered as far as it can pay: #8094+#8167 recovered the primitive cliff, and the reference-param remainder is measured at zero against a free-guard ceiling. I recommend closing this issue in favor of the per-row residuals: #8169 + #8202 (tree, tree_wide, interp, iso_miss), and fresh per-mechanism attribution for shapes / cycles / deeplist / churn / pipeline, which no ordinary-parameter mechanism touches — deeplist and pipeline compile to byte-identical binaries even under the SEED ceiling.

  9. proggeramlug commented on Aug 16, 2026

    @proggeramlug
    ContributorAuthor

    Closing: the guard_blocked rule is correct as written. Lifting it is worth 0.0%, and the instructions this issue was chasing were in the guards Perry already emits — now fixed in #8201.

    Step 1 — re-sweep, because every number here was stale

    bfb0707be is instruction-identical to 499e29627 on all 19 rows (max |Δ| 0.54%). The 16 intervening commits — #8165, #8183, #8188 included — moved nothing. The owner's earlier verdict ("every gain is fib40; the other eight rows are inside noise") holds at today's base. Beats-node is still 8/19.

    Step 2 — the discriminating measurement

    Four diagnostic arms from one compiler, env-gated, all byte-exact output, instructions retired best-of-5:

    bench base SKIPVAL (validators → true) UNGUARD (rule lifted + guards made free) SEED (declared types proven everywhere)
    tree 15.82 B −33.8% −33.8% −0.3%
    tree_wide 31.59 B −17.0% −17.0% +0.0%
    interp 15.05 B −19.6% −19.6% −8.8%
    iso_miss 17.70 B −16.3% −16.3% −7.3%
    others noise noise noise

    UNGUARD − SKIPVAL = 0.0% on every regressed row. Even with the rule lifted and the guards made free, the proofs it refuses buy nothing. And SEED — the causal-control shape, the ceiling of any conceivable scheme — is worse than SKIPVAL on interp/iso_miss, which means that control's apparent value was mostly guard elision, not better lowering.

    Census: 21 blocked params across 17 functions in 5 programs corpus-wide, and zero in cycles/deeplist/churn. pipeline is closures (#8103) plus generic methods; shapes is virtual dispatch; deeplist and pipeline compile byte-identical even under SEED.

    So all three candidate directions are dead against a measured ceiling of zero: escape/mutation analysis on the parameter, identity-only descriptors (#8165 — already measured at −51% and still not the lever here), and per-callsite invalidation (L-sized: 53 flow-insensitive consumer sites plus a pre-pass at codegen/function.rs:897). None of them should be built.

    Step 3 — what was actually costing 16–34%

    js_param_type_guard itself: ~450 instructions of fixed cost per call — descriptor parse plus a 768-byte GuardState init (crates/perry-runtime/src/param_type_guard.rs) — on every unproven call through the public wrapper, including Tier-B clones re-entering themselves through it (#8169's shape, now priced).

    #8201 decides single-node scalar descriptors with the predicate-identical typed-abi leaf guards (codegen/param_guard.rs::scalar_descriptor_rep). Sound by bit-for-bit predicate equality, so routing is unchanged for honest and lying callers alike — verified with a lying-caller probe and the full test_gap_specabi_ordinary_param_guards.ts, both byte-exact against Node 26.5.1.

    Measured, local and confirmed on the quiet mini:

    row instructions wall peak RSS
    tree −31.8% −46.4% unchanged
    tree_wide −15.9% −23.1% unchanged
    interp −8.3% −9.1% unchanged
    iso_miss −6.9% −6.5% unchanged
    all others ±0.4% unchanged

    No unrelated-row movement, which matters here — this repo has repeatedly found that "more type visibility" changes silently un-gate latent fast paths, so a win somewhere unexpected would have been a red flag rather than a bonus.

    19/19 corpus byte-exact on every arm; codegen suite reproduces exactly the 9 pre-existing failures by name against a baseline re-run on pristine bfb0707be; perry-runtime --lib 2499/0; perry --bin perry 987/0; all 12 script gates green. Diagnostic knobs were stripped before the PR.

    What remains, priced

    Closing this issue: its stated direction is refuted, and the instructions it was after have been recovered by a different and much smaller change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

performanceRuntime, compile-time, build-size, or memory performance

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions