Repository navigation
perf(gc): allocation-site pretenuring — long-lived cohorts are copied twice (Eden→survivor→old) #7598
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Aug 7, 2026 The design note's P1 is implemented and measured — #7613
Route: P1, in the smallest sound form. P3 was rejected on this workload
(these records are ~100–200 byte objects, so anyPretenureSizeThresholdlow
enough to catch them catches almost everything, and its stated failure mode is
unbounded); P2 is a refinement of P1 rather than an alternative, and the problem
here is latency, not precision.One correction to P1 as drafted: the note scopes it to the full collection.
The census is available on every cycle that reaches the mark-sweep path,
which is fulls and non-copying minor fallbacks — both of the blind spots
this issue named — so the seed reads it there. It costs nothing extra: that walk
already classifies every Eden header live or dead.Self-reference. The note asked for a paragraph proving the chosen signal is
not suppressed by the state it is meant to leave. The stronger claim holds here:
the Eden live/dead split is produced by the mark-sweep's own arena walk, whose
marks come from reachability, and neither the mark phase nor the sweep reads
tenuring_survivals()— the threshold is consulted in exactly one place,
copying.rs's per-object move, which this path does not run. It is not merely
observed to survive at S=4; it cannot be a function of S. Measured while S was
still 4:eden_live_bytes=279,964,968 eden_dead_bytes=896 live_pct=99 seeds=true.The copy-halving signature, and why it is not cadence.
json_pipeline
500k, per cycle,main→ this:# kind old_before eden_live copied promoted S 1 full 112,695,904 0 → 0 0 → 0 0 → 0 — 2 full 121,002,720 0 → 0 0 → 0 0 → 0 — 3 minor 4,440,144 280,997,080 → 280,997,080 280,997,080 → 0 0 → 280,997,080 4 → 1 4 minor 4,440,144 1,048,576 → (gone) 0 282,045,656 → — 1 → — Cycles 1–3 are unchanged in kind, trigger,
old_beforeandeden_live; cycle 3
receives the same 280,997,080 bytes to the byte and sends them to old-gen
instead of the survivor space. Total bytes moved 563,042,736 → 280,997,080
(0.499×) at 500k and 0.498× at 200k. The collection count does drop 4 → 3,
and the note is right to treat that as suspicious — but the cycle that
disappears is the old cycle 4, whose own Eden influx was 1.0 MB and whose entire
content was the second copy. That is the waste being removed, not the interval
being lengthened.Anti-vacuity. The changed arm cannot satisfy
copied_objects > 0because
eliminating that copy is the change. What the clause exists to catch is
refuted directly: 3 collections, one copying minor that moved 4,117,015
objects / 280,997,080 bytes,eligible=true fallback=none. Reported as
moved_objects > 0.RSS goes down, not up — contrary to the note's expectation. Pinned quiet
host, 5 interleaved reps,cmp-identical output: 200k 1.85 s → 1.44 s (−22.2%),
peak RSS 608.7 MB → 485.7 MB (−20.2%); 500k 5.12 s → 3.86 s (−24.6%), peak RSS
1,404.5 MB → 1,109.8 MB (−21.0%). Promoting earlier does raise the old-gen
high-water mark, but it removes a larger term: at S=4 the 268 MB cohort exists
twice at once at the peak, as Eden from-space plus survivor to-space.gc-ratchet (
--check,pinned_host, on #7609's re-pinned baseline): OK, run
as two arms in one session so the two commits of drift are attributed
separately. Every semantic cell on all 12 probes is bit-identical between a
mainbuild and this one. The seed does not fire on any probe, and the
diagnostic shows that is a refusal rather than an absence:
12_large_live_setprintslive_pct=36 seeds=false.12_large_live_set.wall_ms
— #7610's flag — reads −13.11% on themainarm and −13.14% on this one, so
this change does not move it. Separately for #7610: the +13.58% did not
reproduce in this session at all; amainbuild measures 3,016 ms against the
3,471 ms pinned three commits earlier on the same host.What remains for #7598's headline (allocation-site pretenuring): this removes
the double copy for cohorts a mark-sweep can observe first. A workload whose
first collection is already a copying minor still pays one wasted copy, and that
is what the static "allocated in a loop and stored into an accumulator that
outlives it" proof would cover.Built, measured: site-level static pretenuring is a NET WALL LOSS as-is — branch parked, do not merge
Implementation on
perf/7598-static-pretenure(commit 6e73919): the loop-position collector (accumulatorletat loop depth 0, every push at depth ≥ 1, on top of the all-pointer admission terms), amem::taken per-site flag threaded to both allocation tiers (lower_object_literal's shaped path andlower_call/new.rs's outlined arm — object literals reach codegen as synthesized AnonShape classes, so thenewpath is the one that actually fires), and born-tenured runtime twins (js_object_alloc_with_shape_pretenured,js_object_alloc_class_inline_keys_pretenured) built onarena_alloc_gc_old_born_tenured— so the #7602Old ⟹ TENUREDcontract holds by construction, and field-store correctness is inherited from the live-header barriers.The mechanism works exactly as designed (IR-verified: json emits the pretenured call, push_bench's per-round accumulator is refused, correctly). json_pipeline 200k, hash-identical:
base (main) pretenure GC pause 2,887 ms 967 ms bytes copied by minors 108 MB 0.0 MB peak RSS 570 MB 458 MB (−111 MB) wall 2.73 s 4.19 s ✗ The double copy is gone — and the mutator pays more than the collector saved. Symbolicated profile names the two costs:
- Old-gen allocation is not a bump path. Per object:
arena_alloc_gc_old(free-list probe) +register_old_object_pages+old_object_page_overlaps+old_page_meta_for_object— together the top of the profile. The young inline bump this replaced is ~13 cycles. - A remembered-set insert per young child edge. A pretenured record holds ~10 young field values (fresh
toUpperCasestring, tags array, addr object, …), so 100k records ⇒ ~1Mremember_old_to_young_slotinserts (old_to_young_slot/old_to_young_remembered_setprominent in the profile). Pretenuring the parent without its subtree converts every field initializer into an old→young edge.
What the real fix needs (the v2 design constraints this measurement establishes)
- A bump-allocated pretenure region: born-tenured cohort allocation must be O(bump), with page registration amortized per block, not per object — the current old-gen allocator's per-object bookkeeping alone erases the win.
- Subtree placement, not site placement: the object AND its by-construction-fresh children (the AnonShape ctor's freshly allocated arguments) must be born in the same generation, or the remembered set eats the win. That is cohort/region pretenuring (HotSpot-style), not a per-site allocator swap.
- The static admission logic (loop-position collector) and the codegen threading are reusable as-is for v2 — they are the which sites half; what's missing is the where do the bytes go half.
Adversarial case (function-local depth-0 accumulator dropped every call) was built (
adversarial_bench.tsin the session scratchpad) but is moot until the primary case wins.Numbers note: the base arm above is slightly stale main (pre-#7600/#7601/#7602 runtime); current main's 200k baseline is ~2.6 s, which makes the pretenure arm's 4.19 s worse still. The verdict does not depend on the arm choice.
- Old-gen allocation is not a bump path. Per object:
- added a commit that references this issue
on Aug 8, 2026 Cross-referencing from #7623's audit: the static-pretenure measured win was a #7613-seed confound between arms — full diagnosis at #7623 (comment) (latest). Key fact for the next attempt: on json_pipeline the moved cohort is the RUNTIME-allocated parse tree (~113 MB), not codegen-visible literals (~12 MB, ~1 MB live at minor time). Static site-based pretenure structurally cannot reach it; the two live routes are dynamic feedback or runtime-side allocation-context pretenure in the JSON materialiser. The deferred-page-registration fix inside #7623 is independently valuable (113 MB/run promotes through the quadratic register_old_object_pages path on current main) and should be extracted.
Status after #7623's audit cycle, so this issue carries the complete state:
- The static-pretenure attempt is closed (gc: pretenure-accumulator admission infrastructure (#7598, scope-reduced per audit) #7623, unmerged by author choice). The mechanism was verified correct; the target was wrong: json_pipeline's minor-moved cohort (~113 MB) is the runtime-allocated parse tree, and codegen-visible literals are ~1 MB live at minor time. The measured "win" was a confound between arms that differed in perf(gc): seed promote-on-first-copy from a completed mark-sweep (#7598) #7613's promote-on-first-copy seed.
- The admission infrastructure is preserved on branch
perf/7598-static-pretenure(headeaa7057b7: loop-position collector, refusal tests,region_runs_oncethreading, polarity test — all green). It lands only together with a consumer that actually wants static admission, per the audit's drift argument. - The deferred-page-registration fix (
register_old_object_pagesquadratic per allocation, taxing perf(gc): seed promote-on-first-copy from a completed mark-sweep (#7598) #7613's promote path at ~113 MB/run today) is being extracted separately onperf/old-page-registration-deferral. - The two live routes for this issue: dynamic feedback, or allocation-context pretenure inside the JSON materialiser — the moved cohort is runtime-allocated, and no static admission refinement can reach it.
Step-zero measurement baseline for whatever comes next here, taken per the process rules — current main, pinned mini, best-of-5, hash-stable:
json_pipeline 200k: total ~1.3 s (readFile 22 / parse 252 / build_out 877 / stringify 106 / write 11 / fnv1a 41 — #7601 already killed fnv1a). Peak RSS 451 MB. Census: 3 cycles, 816 ms pause, minor promotes 101 MB with 0.0 MB survivor-copied — the #7613 seed single-hops the parse cohort on current main, confirming the #7623 audit end to end.
Profile ranking (symbolicated): per-slot layout machinery ~52 samples (now #7630), old-page/promote family ~36 (the in-flight
perf/old-page-registration-deferraltarget — profile-confirmed), remembered-set inserts ~15.Sequencing for this issue's materialiser-context route: after the deferral fix lands (born-old materialisation would push 101 MB/run through the per-object registration quadratic otherwise), and #7630's declare-at-birth is complementary groundwork in the same code region.
- added 7 commits that reference this issue
on Aug 8, 2026 The materialiser birth window: built, verified, and parked — main got there first
The route this issue recorded (allocation-context pretenure inside the materialiser) was implemented on
perf/7598-materialiser-pretenureand verified end to end: a nesting-countedParseBirthOldScopearmed at the four doors parses actually use (found by counting redirected births — the two lazy-tape batch entries, the ≥64 KB eager arms, andlazy_get's per-element materialisation, which is the door sequential access actually walks since the batch crossover never trips), whole-subtree born-old placement, hash-identical output, 1,909/0 runtime tests, ratchet semantically identical across 144 medians (one +0.00 % noise cell), planted birth-window test.Mid-verification it measured build_out 627 → 352 ms (−44 %), census promoted 101 MB → 0.0, RSS −88 MB against its contemporary main.
Then main caught up while the final A/B was running. Rebased onto v0.5.1368-era main with both arms matched: wall identical (250/250 ms at 200k, 631/632 ms at 500k), RSS a wash — and the cycle-kind breakdown shows why: current main runs this workload with 2 full cycles and zero copying minors (#7650's double-traversal skip + #7646's census-based membership changed the cycle economics), so the single-hop promotion this window eliminates no longer happens. The parse-cohort GC cost this issue tracked is gone on main, reached by a different road.
Two measurement rules this cycle burned into the record:
- A tree-swap A/B against
origin/mainsilently measures main's progress as your regression unless you rebase first — this repo merges several times a day; the ratchet caught it as impossible semantic deltas on probes that never parse JSON. copied/promotedcensus counters count copying minors only — before claiming "zero-hop", break down cycle kinds, or a workload that shifted to fulls reads as a vacuous 0.0.
The branch stays parked (
perf/7598-materialiser-pretenure, head7e5a600eb) for if the cycle economics ever shift back; per the #7623 standard it does not merge while it optimizes nothing.The arc this issue opened is complete:
build_outon this workload has gone from 57.4 s (#7592 filing) to ~250 ms per 200k on current main — #7594, #7596, #7613, #7624, #7633, #7602, #7646, #7650 collectively. Recommend closing.- A tree-swap A/B against
Closing: the architectural thesis was tested and does not survive the evidence
The last open question was whether birth-time placement (the parked window) beats main's emergent cycle economics on the shape that should discriminate: parse-then-serve — a big dataset parsed at startup, held live, with sustained request allocation after it. If the young-resident cohort were fragile, serving-phase copying minors would re-confront it and the born-old arm would win.
Measured on the pinned mini, matched bases, hash-identical, at two loads:
- 400k requests: zero minors in either arm (2 fulls) — the proportional bands absorb the churn entirely. Wall, RSS, pause identical.
- 4M requests: minors appear (4 per arm) — and they are non-copying with
copied=0.0 / promoted=0.0in both arms, pause 224 vs 221 ms, serve wall 2,475 vs 2,524 ms. Identical within noise.
The predicted re-confrontation never happens: #7613 bounds the cohort to at most a single hop even when copying minors run, and current pacing (#7596/#7646/#7650) produces none at all on these shapes. Birth-time placement remains the cleaner invariant on paper, but it is empirically indistinguishable at parse-exit, parse-serve, and 10× parse-serve — and per the standard this campaign set (#7623): nothing merges while it optimizes nothing measurable.
Branch
perf/7598-materialiser-pretenure(head7e5a600eb) stays as the historical artifact — window mechanism, planted tests, and the door map (eager arms pastLAZY_MAX_BLOB_BYTES,lazy_get's per-element path) are all documented there if collector economics ever change in a way that revives the question. The re-open condition is concrete: a census showing copying minors repeatedly moving a parse cohort.The arc closes:
build_out57.4 s → ~250 ms per 200k records, via #7594, #7596, #7613, #7624, #7633, #7602, #7646, #7650 — three parallel workstreams, every claimed win either verified on the pinned host or caught and corrected by audit.
Follow-up to #7592 (fixed by #7594 + #7596 up to a constant factor). The quadratic is gone, but the dominant remaining cost on promote-heavy workloads is structural: every long-lived object is copied twice — Eden→survivor by the first copying minor, survivor→old by the next.
Measured shape (json_pipeline, 500k records, after #7596)
3.9 s of the remaining 5.1 s
build_outis these two copies of the same 268 MB. A collection-free run of the same loop is 764 ns/record vs the current ~23,000, so this is the bulk of what is left on #7592's workload.Why the cheap fixes do not work (all measured or ruled out this campaign)
The structural fix: allocation-site pretenuring
HotSpot-style: per-allocation-site survival feedback (site → promoted-fraction), fed back either as a runtime site tag on the allocation header or as a codegen hint (
arena_alloc_gc_olddirectly for hot sites whose cohort demonstrably tenures). Deterministic per-cycle (the decision is made at allocation, not during the collection), so it composes with #7432.Perry has unusually good static material for this compared to a JIT:
out.push({...})-into-an-accumulator is visible in HIR (the #7469 loop-site work already classifies allocation sites by loop context, andcollectors/all_pointer_arrays.rsproves which locals only receive freshly-allocated pushes). A static "allocated in a loop AND stored into an accumulator that outlives the loop" proof may cover the common case without any runtime feedback.Acceptance
build_out≤ ~2.5 s (one copy of the live cohort instead of two) with output hash unchanged.