Skip to content

perf(gc): allocation-site pretenuring — long-lived cohorts are copied twice (Eden→survivor→old) #7598

Description

@proggeramlug

Follow-up to #7592 (fixed by #7594 + #7596 up to a constant factor). The quadratic is gone, but the dominant remaining cost on promote-heavy workloads is structural: every long-lived object is copied twice — Eden→survivor by the first copying minor, survivor→old by the next.

Measured shape (json_pipeline, 500k records, after #7596)

 #  kind          trigger      old_before  eden_live  promoted   pause
 3  minor     arena_bytes         4.2 MB     268 MB        0    1,658 ms   <- 268 MB Eden->survivor, ALL of it long-lived
 4  minor     arena_bytes         4.2 MB       1 MB     269 MB  2,223 ms   <- the same bytes, survivor->old

3.9 s of the remaining 5.1 s build_out is these two copies of the same 268 MB. A collection-free run of the same loop is 764 ns/record vs the current ~23,000, so this is the bulk of what is left on #7592's workload.

Why the cheap fixes do not work (all measured or ruled out this campaign)

  • The survival-rate lock is a cycle too late by construction — it keys on the previous cycle's survivor round-trip, so the first big minor always pays the wasted copy. It did engage here (cycle 4 promoted), which is exactly why the waste is confined to cycle 3's copy.
  • An Eden-side lock (≥90 % of Eden survives ⇒ S=1) measured 1.0× — it also only takes effect on the next copying minor, and a one-burst workload has no next minor that matters.
  • A mid-cycle valve (promote directly once the to-survivor space overflows) is ruled out by gc: adaptive tenuring + young-scoped scavenge cap (fixes the large-live-set scavenge regression) #7432: which objects promote would then depend on root traversal order, breaking the bit-identical copied/promoted counters the gc-ratchet's determinism contract requires. The per-object decision must stay a pure function of (flags, age, cycle-start threshold).
  • Eden in-use at cycle start is not a survival signal — churn workloads also overshoot Eden, with ~0 % survival; gating S on it would mis-promote garbage.

The structural fix: allocation-site pretenuring

HotSpot-style: per-allocation-site survival feedback (site → promoted-fraction), fed back either as a runtime site tag on the allocation header or as a codegen hint (arena_alloc_gc_old directly for hot sites whose cohort demonstrably tenures). Deterministic per-cycle (the decision is made at allocation, not during the collection), so it composes with #7432.

Perry has unusually good static material for this compared to a JIT: out.push({...})-into-an-accumulator is visible in HIR (the #7469 loop-site work already classifies allocation sites by loop context, and collectors/all_pointer_arrays.rs proves which locals only receive freshly-allocated pushes). A static "allocated in a loop AND stored into an accumulator that outlives the loop" proof may cover the common case without any runtime feedback.

Acceptance

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Aug 7, 2026
  2. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    The design note's P1 is implemented and measured — #7613

    Route: P1, in the smallest sound form. P3 was rejected on this workload
    (these records are ~100–200 byte objects, so any PretenureSizeThreshold low
    enough to catch them catches almost everything, and its stated failure mode is
    unbounded); P2 is a refinement of P1 rather than an alternative, and the problem
    here is latency, not precision.

    One correction to P1 as drafted: the note scopes it to the full collection.
    The census is available on every cycle that reaches the mark-sweep path,
    which is fulls and non-copying minor fallbacks — both of the blind spots
    this issue named — so the seed reads it there. It costs nothing extra: that walk
    already classifies every Eden header live or dead.

    Self-reference. The note asked for a paragraph proving the chosen signal is
    not suppressed by the state it is meant to leave. The stronger claim holds here:
    the Eden live/dead split is produced by the mark-sweep's own arena walk, whose
    marks come from reachability, and neither the mark phase nor the sweep reads
    tenuring_survivals() — the threshold is consulted in exactly one place,
    copying.rs's per-object move, which this path does not run. It is not merely
    observed to survive at S=4; it cannot be a function of S. Measured while S was
    still 4: eden_live_bytes=279,964,968 eden_dead_bytes=896 live_pct=99 seeds=true.

    The copy-halving signature, and why it is not cadence. json_pipeline
    500k, per cycle, main → this:

    # kind old_before eden_live copied promoted S
    1 full 112,695,904 0 → 0 0 → 0 0 → 0 —
    2 full 121,002,720 0 → 0 0 → 0 0 → 0 —
    3 minor 4,440,144 280,997,080 → 280,997,080 280,997,080 → 0 0 → 280,997,080 4 → 1
    4 minor 4,440,144 1,048,576 → (gone) 0 282,045,656 → — 1 → —

    Cycles 1–3 are unchanged in kind, trigger, old_before and eden_live; cycle 3
    receives the same 280,997,080 bytes to the byte and sends them to old-gen
    instead of the survivor space. Total bytes moved 563,042,736 → 280,997,080
    (0.499×) at 500k and 0.498× at 200k. The collection count does drop 4 → 3,
    and the note is right to treat that as suspicious — but the cycle that
    disappears is the old cycle 4, whose own Eden influx was 1.0 MB and whose entire
    content was the second copy. That is the waste being removed, not the interval
    being lengthened.

    Anti-vacuity. The changed arm cannot satisfy copied_objects > 0 because
    eliminating that copy is the change. What the clause exists to catch is
    refuted directly: 3 collections, one copying minor that moved 4,117,015
    objects / 280,997,080 bytes, eligible=true fallback=none. Reported as
    moved_objects > 0.

    RSS goes down, not up — contrary to the note's expectation. Pinned quiet
    host, 5 interleaved reps, cmp-identical output: 200k 1.85 s → 1.44 s (−22.2%),
    peak RSS 608.7 MB → 485.7 MB (−20.2%); 500k 5.12 s → 3.86 s (−24.6%), peak RSS
    1,404.5 MB → 1,109.8 MB (−21.0%). Promoting earlier does raise the old-gen
    high-water mark, but it removes a larger term: at S=4 the 268 MB cohort exists
    twice at once at the peak, as Eden from-space plus survivor to-space.

    gc-ratchet (--check, pinned_host, on #7609's re-pinned baseline): OK, run
    as two arms in one session so the two commits of drift are attributed
    separately. Every semantic cell on all 12 probes is bit-identical between a
    main build and this one. The seed does not fire on any probe, and the
    diagnostic shows that is a refusal rather than an absence:
    12_large_live_set prints live_pct=36 seeds=false. 12_large_live_set.wall_ms
    — #7610's flag — reads −13.11% on the main arm and −13.14% on this one, so
    this change does not move it. Separately for #7610: the +13.58% did not
    reproduce in this session at all; a main build measures 3,016 ms against the
    3,471 ms pinned three commits earlier on the same host.

    What remains for #7598's headline (allocation-site pretenuring): this removes
    the double copy for cohorts a mark-sweep can observe first. A workload whose
    first collection is already a copying minor still pays one wasted copy, and that
    is what the static "allocated in a loop and stored into an accumulator that
    outlives it" proof would cover.

  3. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Built, measured: site-level static pretenuring is a NET WALL LOSS as-is — branch parked, do not merge

    Implementation on perf/7598-static-pretenure (commit 6e73919): the loop-position collector (accumulator let at loop depth 0, every push at depth ≥ 1, on top of the all-pointer admission terms), a mem::taken per-site flag threaded to both allocation tiers (lower_object_literal's shaped path and lower_call/new.rs's outlined arm — object literals reach codegen as synthesized AnonShape classes, so the new path is the one that actually fires), and born-tenured runtime twins (js_object_alloc_with_shape_pretenured, js_object_alloc_class_inline_keys_pretenured) built on arena_alloc_gc_old_born_tenured — so the #7602 Old ⟹ TENURED contract holds by construction, and field-store correctness is inherited from the live-header barriers.

    The mechanism works exactly as designed (IR-verified: json emits the pretenured call, push_bench's per-round accumulator is refused, correctly). json_pipeline 200k, hash-identical:

    base (main) pretenure
    GC pause 2,887 ms 967 ms
    bytes copied by minors 108 MB 0.0 MB
    peak RSS 570 MB 458 MB (−111 MB)
    wall 2.73 s 4.19 s ✗

    The double copy is gone — and the mutator pays more than the collector saved. Symbolicated profile names the two costs:

    1. Old-gen allocation is not a bump path. Per object: arena_alloc_gc_old (free-list probe) + register_old_object_pages + old_object_page_overlaps + old_page_meta_for_object — together the top of the profile. The young inline bump this replaced is ~13 cycles.
    2. A remembered-set insert per young child edge. A pretenured record holds ~10 young field values (fresh toUpperCase string, tags array, addr object, …), so 100k records ⇒ ~1M remember_old_to_young_slot inserts (old_to_young_slot/old_to_young_remembered_set prominent in the profile). Pretenuring the parent without its subtree converts every field initializer into an old→young edge.

    What the real fix needs (the v2 design constraints this measurement establishes)

    • A bump-allocated pretenure region: born-tenured cohort allocation must be O(bump), with page registration amortized per block, not per object — the current old-gen allocator's per-object bookkeeping alone erases the win.
    • Subtree placement, not site placement: the object AND its by-construction-fresh children (the AnonShape ctor's freshly allocated arguments) must be born in the same generation, or the remembered set eats the win. That is cohort/region pretenuring (HotSpot-style), not a per-site allocator swap.
    • The static admission logic (loop-position collector) and the codegen threading are reusable as-is for v2 — they are the which sites half; what's missing is the where do the bytes go half.

    Adversarial case (function-local depth-0 accumulator dropped every call) was built (adversarial_bench.ts in the session scratchpad) but is moot until the primary case wins.

    Numbers note: the base arm above is slightly stale main (pre-#7600/#7601/#7602 runtime); current main's 200k baseline is ~2.6 s, which makes the pretenure arm's 4.19 s worse still. The verdict does not depend on the arm choice.

  4. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Cross-referencing from #7623's audit: the static-pretenure measured win was a #7613-seed confound between arms — full diagnosis at #7623 (comment) (latest). Key fact for the next attempt: on json_pipeline the moved cohort is the RUNTIME-allocated parse tree (~113 MB), not codegen-visible literals (~12 MB, ~1 MB live at minor time). Static site-based pretenure structurally cannot reach it; the two live routes are dynamic feedback or runtime-side allocation-context pretenure in the JSON materialiser. The deferred-page-registration fix inside #7623 is independently valuable (113 MB/run promotes through the quadratic register_old_object_pages path on current main) and should be extracted.

  5. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Status after #7623's audit cycle, so this issue carries the complete state:

    • The static-pretenure attempt is closed (gc: pretenure-accumulator admission infrastructure (#7598, scope-reduced per audit) #7623, unmerged by author choice). The mechanism was verified correct; the target was wrong: json_pipeline's minor-moved cohort (~113 MB) is the runtime-allocated parse tree, and codegen-visible literals are ~1 MB live at minor time. The measured "win" was a confound between arms that differed in perf(gc): seed promote-on-first-copy from a completed mark-sweep (#7598) #7613's promote-on-first-copy seed.
    • The admission infrastructure is preserved on branch perf/7598-static-pretenure (head eaa7057b7: loop-position collector, refusal tests, region_runs_once threading, polarity test — all green). It lands only together with a consumer that actually wants static admission, per the audit's drift argument.
    • The deferred-page-registration fix (register_old_object_pages quadratic per allocation, taxing perf(gc): seed promote-on-first-copy from a completed mark-sweep (#7598) #7613's promote path at ~113 MB/run today) is being extracted separately on perf/old-page-registration-deferral.
    • The two live routes for this issue: dynamic feedback, or allocation-context pretenure inside the JSON materialiser — the moved cohort is runtime-allocated, and no static admission refinement can reach it.
  6. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Step-zero measurement baseline for whatever comes next here, taken per the process rules — current main, pinned mini, best-of-5, hash-stable:

    json_pipeline 200k: total ~1.3 s (readFile 22 / parse 252 / build_out 877 / stringify 106 / write 11 / fnv1a 41 — #7601 already killed fnv1a). Peak RSS 451 MB. Census: 3 cycles, 816 ms pause, minor promotes 101 MB with 0.0 MB survivor-copied — the #7613 seed single-hops the parse cohort on current main, confirming the #7623 audit end to end.

    Profile ranking (symbolicated): per-slot layout machinery ~52 samples (now #7630), old-page/promote family ~36 (the in-flight perf/old-page-registration-deferral target — profile-confirmed), remembered-set inserts ~15.

    Sequencing for this issue's materialiser-context route: after the deferral fix lands (born-old materialisation would push 101 MB/run through the per-object registration quadratic otherwise), and #7630's declare-at-birth is complementary groundwork in the same code region.

  7. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    The materialiser birth window: built, verified, and parked — main got there first

    The route this issue recorded (allocation-context pretenure inside the materialiser) was implemented on perf/7598-materialiser-pretenure and verified end to end: a nesting-counted ParseBirthOldScope armed at the four doors parses actually use (found by counting redirected births — the two lazy-tape batch entries, the ≥64 KB eager arms, and lazy_get's per-element materialisation, which is the door sequential access actually walks since the batch crossover never trips), whole-subtree born-old placement, hash-identical output, 1,909/0 runtime tests, ratchet semantically identical across 144 medians (one +0.00 % noise cell), planted birth-window test.

    Mid-verification it measured build_out 627 → 352 ms (−44 %), census promoted 101 MB → 0.0, RSS −88 MB against its contemporary main.

    Then main caught up while the final A/B was running. Rebased onto v0.5.1368-era main with both arms matched: wall identical (250/250 ms at 200k, 631/632 ms at 500k), RSS a wash — and the cycle-kind breakdown shows why: current main runs this workload with 2 full cycles and zero copying minors (#7650's double-traversal skip + #7646's census-based membership changed the cycle economics), so the single-hop promotion this window eliminates no longer happens. The parse-cohort GC cost this issue tracked is gone on main, reached by a different road.

    Two measurement rules this cycle burned into the record:

    1. A tree-swap A/B against origin/main silently measures main's progress as your regression unless you rebase first — this repo merges several times a day; the ratchet caught it as impossible semantic deltas on probes that never parse JSON.
    2. copied/promoted census counters count copying minors only — before claiming "zero-hop", break down cycle kinds, or a workload that shifted to fulls reads as a vacuous 0.0.

    The branch stays parked (perf/7598-materialiser-pretenure, head 7e5a600eb) for if the cycle economics ever shift back; per the #7623 standard it does not merge while it optimizes nothing.

    The arc this issue opened is complete: build_out on this workload has gone from 57.4 s (#7592 filing) to ~250 ms per 200k on current main — #7594, #7596, #7613, #7624, #7633, #7602, #7646, #7650 collectively. Recommend closing.

  8. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Closing: the architectural thesis was tested and does not survive the evidence

    The last open question was whether birth-time placement (the parked window) beats main's emergent cycle economics on the shape that should discriminate: parse-then-serve — a big dataset parsed at startup, held live, with sustained request allocation after it. If the young-resident cohort were fragile, serving-phase copying minors would re-confront it and the born-old arm would win.

    Measured on the pinned mini, matched bases, hash-identical, at two loads:

    • 400k requests: zero minors in either arm (2 fulls) — the proportional bands absorb the churn entirely. Wall, RSS, pause identical.
    • 4M requests: minors appear (4 per arm) — and they are non-copying with copied=0.0 / promoted=0.0 in both arms, pause 224 vs 221 ms, serve wall 2,475 vs 2,524 ms. Identical within noise.

    The predicted re-confrontation never happens: #7613 bounds the cohort to at most a single hop even when copying minors run, and current pacing (#7596/#7646/#7650) produces none at all on these shapes. Birth-time placement remains the cleaner invariant on paper, but it is empirically indistinguishable at parse-exit, parse-serve, and 10× parse-serve — and per the standard this campaign set (#7623): nothing merges while it optimizes nothing measurable.

    Branch perf/7598-materialiser-pretenure (head 7e5a600eb) stays as the historical artifact — window mechanism, planted tests, and the door map (eager arms past LAZY_MAX_BLOB_BYTES, lazy_get's per-element path) are all documented there if collector economics ever change in a way that revives the question. The re-open condition is concrete: a census showing copying minors repeatedly moving a parse cohort.

    The arc closes: build_out 57.4 s → ~250 ms per 200k records, via #7594, #7596, #7613, #7624, #7633, #7602, #7646, #7650 — three parallel workstreams, every claimed win either verified on the pinned host or caught and corrected by audit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions