Skip to content

perf: json_pipeline at 500k records is 97.6x bun (60.4s vs 618ms) while the same workload at 100 records BEATS bun — a scaling cliff, not a constant factor #7592

Description

@proggeramlug

Measured in today's public-baseline regeneration (v0.5.1335, 2ba59501b, pinned quiet M1 Max)

honest_bench's JSON pipeline, 20 runs per cell, output verified against the
Bun reference 20/20 for every language
— so this is a pure performance
finding, not a correctness one.

workload perry bun node rust zig perry RSS
json_pipeline_small (100 records) 77.0 ms 90.8 128.1 73.4 72.8 8.5 MB
json_pipeline_full (500,000 records, 107.5 MB) 60,358.5 ms 618.1 1,004.3 650.0 875.5 1,064 MB
image_convolution (4K, 5×5) 302.5 ms 947.2 1,284.8 429.8 279.8 56.2 MB

At 100 records Perry beats bun and node and lands within 5% of Rust and Zig,
on a quarter of bun's memory. At 500k records it is 97.6× bun.
Same source
file, same binary, same harness — only the fixture size changes.

That is the shape of the finding: not slowness, a scaling cliff. Across the
5,000× increase in input, bun's time grows 6.8× (startup-dominated at the small
end); Perry's grows 784×. Rust 8.9×, Zig 12.0×, node 7.8×. Perry is the only
implementation whose growth is superlinear in the data, which points at an
algorithmic or GC-pathological term rather than a constant-factor gap.

RSS is 1,064 MB against bun's 579 MB — 1.8×, and notably not 97×. The memory
is not exploding proportionally to the time, which argues against "it simply
swapped" and for real work being repeated.

Where to look

The workload is small and does five things (workloads/1_json_pipeline/perry/json_pipeline.ts):
fs.readFileSync a 107.5 MB string → JSON.parse to 500k objects → a filter +
field-derive loop building out → JSON.stringify(out) → fs.writeFileSync,
then an FNV-1a hash over the serialized string.

Each of those is separately suspect at this scale and they should be timed
individually before anything is optimised — this campaign has repeatedly paid
for working from an unattributed headline number. Candidates worth ruling in or
out first:

A leaf profile (PERRY_DEBUG_SYMBOLS=1 + sample) will separate these in one
run; at 60 s per iteration there is plenty of signal.

Two stale claims in the workload source, now disproved

json_pipeline.ts carries a comment block from v0.5.29 stating that (a) the
driver "runs this binary on the 100-record fixture only", and (b) iterating a
large JSON.parse result "triggers a GC-scan issue at scale — records allocated
by the JSON parser get swept mid-iteration… above [~200 records] output is
non-deterministic".

Both are now false. The driver does run Perry on the full 500k fixture
(run.sh:268), and the output matched the Bun reference on 20 of 20 runs.
The correctness half was fixed at some point without the comment being updated.
Worth deleting so the next reader does not attribute the slowness to a
corruption bug that no longer exists.

Why this matters beyond the number

README.md cites this harness by name for its JSON row and quotes the
100-record result. The same report's 500k row is 97.6× bun. Whatever the
intent, the README also says "We publish everything, including the workloads
where V8's JIT still beats us — no cherry-picked table can survive an open
harness."
Those two things need to be reconciled; tracked separately.

Activity

  1. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    Root cause: this is GC pacing, not JSON

    Per the issue's own instruction, the 60 s was split before anything was touched. An instrumented copy of the workload (Date.now() between the five phases, nothing else changed), 500k records:

    phase time share
    readFileSync 89 ms 0.1%
    JSON.parse 742 ms 1.2%
    build_out 57,409 ms 94.1%
    JSON.stringify 1,451 ms 2.4%
    writeFileSync 90 ms 0.1%
    fnv1a 1,250 ms 2.0%

    JSON.parse is not the problem — it handles 107 MB in 742 ms and scales linearly (16 ms at 25k → 742 ms at 500k, 20× data for 46× time). The cliff is entirely the out.push({...}) loop.

    It is superlinear, and it is GC

    build_out ns/record roughly doubles every time N doubles:

    records build_out ns/record vs prev
    25,000 41 ms 3,246 —
    50,000 455 ms 18,156 5.59×
    100,000 2,043 ms 40,769 2.25×
    200,000 8,451 ms 84,341 2.07×
    500,000 56,833 ms 227,354 2.70×

    From PERRY_GC_TRACE, build_out is ~100 % GC pause — 8,840 ms of pause against 8,633 ms of build_out at 100k. Bytes copied per record are flat (~1.1 KB), so this is not a copying/promotion problem; it is collection frequency against a growing live set.

    The mechanism

    effective_next_arena_trigger() is min(adaptive_base, scavenge_nursery_cap_effective_bytes()), and that cap is PERRY_GC_SCAVENGE_NURSERY_MB (default 16 MB) × NURSERY_CAP_SCALE, where gc/tenuring.rs sets NURSERY_CAP_SCALE_MAX = 4. So the cap has a hard ceiling of 64 MB.

    Collection cadence is therefore a fixed number of allocated bytes, while each collection costs O(live). Total GC work is (alloc / cap) × O(live) = O(N²). Work per allocated byte is c·L/cap, which is bounded only if cap ∝ L.

    That means no constant ceiling fixes this — it only moves the cliff to a larger N, which is exactly what a fixed-N cap sweep shows:

    records 64 MB 256 MB 1024 MB
    100,000 559 ns/rec 619 559
    200,000 85,749 ✗ 5,190 5,469
    500,000 229,786 ✗ 233,634 ✗ 764 ✓

    With a live-set-sized nursery, 500k build_out goes 57,441 ms → 191 ms (300×) and ns/record goes flat — the quadratic is gone, not merely reduced.

    Why the existing adaptive scale does not save it

    NURSERY_CAP_SCALE was introduced for exactly this shape, but it cannot reach its own ceiling here, for two independent reasons:

    1. It is fed by copying minors only. The sole production caller of retune_after_scavenge is gc/copying.rs:1280. Full collections and non-copying minor fallbacks feed it nothing. That is self-reinforcing: once the workload escalates to fulls, the feedback loop goes blind and the cap stays pinned, which produces more fulls.
    2. The debounce outruns the workload. Growth is ×2 per two consecutive qualifying cycles, and two steps are needed to reach ×4 — so ≥4 feeding cycles. The 50k run had 3 cycles total, 2 of them feeding.

    Proposed fix

    Make the collection budget proportional to the live set rather than a bounded multiple of a fixed base — cap = max(base, live_bytes / K), the standard Go-GOGC / V8 heap-growth shape. K trades RSS for GC time, with peak ≈ L(1 + 1/K).

    The RSS cost is the real decision here, since this necessarily raises the high-water mark on live-set-bound workloads — so it needs to go through the RSS ratchet rather than be argued from the wall-clock win alone. I am measuring that next.

    Also worth fixing independently of the pacing change: the feedback loop should be fed by every collection that observes a live set, not only copying minors, so it cannot go blind exactly when it is needed most.

  2. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    Two probes, and a correction to the fix I proposed above

    I built the live-proportional cap I described and measured it. It is the wrong fix, and the reason is instructive.

    Probe 1 — live-proportional nursery cap: 5.9×, but structurally broken

    cap = max(base, live_bytes) where live_bytes is arena in-use recorded at the end of every cycle. 500k build_out 57.4 s → 9.7 s, but ns/record still climbed with N. The diagnostic shows why:

    [gc-live] live= 44,748,304  cap= 44,748,304  from_space=      2,656
    [gc-live] live=114,343,400  cap=114,343,400  from_space=112,195,992
    [gc-live] live=118,978,680  cap=118,978,680  from_space=116,372,472
    

    The cap gates young_scavenge_cap_due(), which compares from-space occupancy against it. But arena in-use includes the young generation, and on this workload from-space is nearly all of live — so cap = live is a fixed point that from-space can never cross. Scavenging stopped entirely (0 cycles at 200k), the nursery was never evacuated, and the residual 2.7 s is the mutator paying for a huge un-evacuated young space.

    Any live-proportional cap has to be proportional to the tenured live set, not to total arena in-use, or it self-references. That also means it cannot work on its own — see below.

    Probe 2 — promote on first copy: 10.3×, and the quadratic is gone

    Forcing tenuring_survivals() == 1:

    records main ns/rec promote-on-first-copy speedup
    25,000 3,167 3,087 1.0×
    50,000 19,513 14,246 1.4×
    100,000 41,866 18,798 2.2×
    200,000 85,529 19,591 4.4×
    500,000 231,570 22,534 10.3×

    500k build_out 57.4 s → 5.6 s, and ns/record is flat from 100k on (18.8k → 19.6k → 22.5k). That is the quadratic gone, not merely reduced.

    The mechanism is the one the module header already describes: these objects are long-lived (they are what out accumulates), but at S=4 they stay young and get fully re-copied on every cycle.

    Why the existing survival-rate lock does not do this

    It is not missing the condition, it is a cycle too late. The lock keys on prev_copied — it needs a previous copying minor to have filled the survivor space, so it cannot engage until the second one, and it only changes the threshold for the cycle after that. This workload has ~6 cycles, most of them fulls (which never call retune_after_scavenge at all).

    I tried adding the Eden-side form of the same proof — lock when ≥90% of Eden survives, which is available on the first copying minor rather than the second. Measured 1.0× across the board: still a no-op, because it too only takes effect on the next copying minor, and on this workload that cycle mostly never comes. Reverted rather than shipped.

    Where that leaves the fix

    The two levers are complementary and neither works alone:

    1. Promotion has to engage on the first copy, not one or two copying minors later. The signal has to be acted on within the cycle that observes it, or read from a source that survives across cycles — the current loop is fed only by copying minors, so it goes blind exactly when it is needed.
    2. Then a tenured-live-proportional collection budget becomes well-defined, because old-gen starts tracking the real live set instead of sitting at ~2 MB while everything piles up in from-space.

    Both change promotion rates and heap high-water marks, so this needs to go through the RSS ratchet before it is a candidate — the 300× available here is a wall-clock number, and #7438 is a standing reminder that this collector's RSS behaviour is the binding constraint. Note also #7432's finding that a mid-cycle valve breaks ratchet counter determinism, which rules out the most obvious way of acting within the cycle.

    No code from these probes is proposed for merge; the tree is back at main.

  3. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    Found it — and my probe result above was an artifact

    Correction first. The "promote on first copy = 10.3×" measurement in my previous comment does not mean what I said it did. tenuring_survivals() also feeds copied_minor_promotable_active_survivor_bytes(), which drives the promotion-handoff pacing — so forcing it to 1 changed triggering, not just promotion. That arm ran zero collections. It was fast because it did not collect, not because it promoted early.

    Two measurement traps produced that, both worth recording:

    • The auto-optimize path relinks the runtime with --no-default-features, which silently drops the diagnostics feature. PERRY_GC_TRACE then prints diagnostics feature disabled instead of JSON, and my cycle counts read 0 — for three different arms in a row. It also meant the baseline binary (built before I cleared that cache) had diagnostics while the new one did not, so the arms were not feature-identical. The fix is to build -p perry -p perry-runtime-static -p perry-stdlib-static and link with a pinned PERRY_RUNTIME_DIR and PERRY_NO_AUTO_OPTIMIZE=1.
    • PERRY_GC_TRACE lines are enormous; piping them SIGPIPEs and truncates the cycle list, which is where an earlier "21 fulls" came from.

    The actual root cause: a livelock

    With working diagnostics, the 200k trace is unambiguous — 22 collections:

     19 x full/survivor_promotion_bytes     each freeing 0.0 MB at ~400 ms
      1 x full/old_gen_bytes
      1 x full/arena_bytes
      1 x minor/arena_bytes                 promoted: 0.0 MB
    

    copied_minor_promotion_handoff_due replaces a minor with a full mark-sweep to make room in old-gen for survivors about to be promoted. But a full mark-sweep is non-moving — it promotes nothing. It cannot relieve the pressure that scheduled it: the survivor space still holds the same 108 MB, the reclaim baseline it resets does not count those bytes, and the predicate is true again at the next minor. The copying minor that would have done the promotion never runs.

    That is 7.6 s of an 8.6 s phase spent freeing nothing. Peak RSS is the same whether those 19 collections happen or not, which is the tell: they were pure loss.

    Fix — #7594

    Latch it: one handoff per copying minor. The handoff makes room; the copying minor performs the promotion that consumes it.

    cycles pause survivor_promotion fulls promoted
    main 22 8,782 ms 19 0.0 MB
    #7594 6 2,828 ms 1 110.0 MB

    build_out at 500k: 57,242 ms → 10,551 ms (5.4×), output hash identical. RSS is up 17–24 % at the large sizes because the run now collects 6 times instead of 22 — a real trade, stated in the PR.

    On the ratchet, all 12 probes agree with main on every semantic counter across 144 metric medians (measured as a back-to-back A/B — the pinned baseline is 0.5.1315 on another host and cannot separate this change from drift).

    What is left

    build_out is still ~28,500 ns/record and not flat. With a nursery large enough to avoid collecting during the loop the same phase runs at 764 ns/record, so roughly another order of magnitude is available, and it is the collection-budget question: the cap is a constant (16 MB × NURSERY_CAP_SCALE_MAX), so cadence is independent of live-set size.

    The obvious repair does not work as written, and it is worth writing down why: a live-proportional cap gates on from-space occupancy, and on this workload from-space is nearly all of live, so cap = live is a fixed point that from-space can never cross — scavenging stops entirely. Any such budget has to key on the tenured live set, which in turn only becomes meaningful once promotion actually happens. Leaving this issue open for that half.

  4. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    Resolved: the cliff is gone (#7594 + #7596, both merged)

    Final three-arm measurement, identically built, hash-verified per row:

    records main (before) #7594 #7594+#7596 speedup ns/rec
    25,000 42 ms 42 42 1.0× 3,325
    100,000 2,257 ms 1,499 863 2.6× 17,221
    200,000 8,889 ms 3,505 2,617 3.4× 26,118
    500,000 61,589 ms 12,139 5,784 10.6× 23,138

    ns/record is now flat within ~30 % across a 20× size range; before, it grew 70×. The two fixes:

    • perf(gc): break the survivor-promotion handoff livelock (#7592) #7594 — the survivor-promotion handoff livelock: 19 consecutive full collections each freeing 0.0 MB, because a non-moving full can never relieve the survivor pressure that scheduled it. Latched to one handoff per copying minor.
    • perf(gc): live-proportional collection budgets at both generations (#7592) #7596 — no constant band may pace a collector whose per-cycle cost is O(live): the handoff now requires current old-gen pressure (not predicted), promoted bytes credit the reclaim baseline (they are live by construction), and both the old-reclaim growth band and the scavenge nursery cap are live-proportional (the nursery cap keyed on tenured occupancy — total arena in-use is a fixed point from-space can never cross).

    What remains, tracked elsewhere

    Closing as fixed — the reported scaling cliff (97.6× bun at 500k, beating bun at 100 records) was the quadratic, and the quadratic is dead.

  5. proggeramlug commented on Aug 7, 2026

    @proggeramlug
    ContributorAuthor

    Design note: getting promote-on-first-copy to engage in time

    #7596's stated follow-up is that surviving bytes are copied twice —
    Eden → survivor space → old-gen — 268 MB each way, ~3.9 s of the remaining
    ~5.1 s of build_out. Promoting on the first copy halves that. This is the
    design for it; no code is proposed here.

    The obstacle is latency, not blindness — and it is worth being precise

    Two failed probes are already recorded above, and they failed for different
    reasons. Conflating them sends the next attempt at the wrong target.

    • The survival-rate lock (gc/tenuring.rs) already computes exactly the
      right condition: a substantial survivor intake of which ≥90% comes back out
      alive means the aging round filters nothing. It is not missing the
      condition. It keys on prev_copied, so it needs a previous copying minor
      to have filled the survivor space, and it changes the threshold for the
      cycle after that. Two cycles of latency on a run that has about six.
    • The Eden-side variant measured 1.0x, and the note above attributes that
      to the same latency. Half true. The deeper reason is that the quantity it
      reached for is not observable at the threshold it is trying to leave:
      promoted_bytes is zero by construction at S=4, so "was the promotion
      rate high last cycle?" can never answer yes while S=4 holds. That is a
      fixed point of the same family as perf(gc): break the survivor-promotion handoff livelock (#7592) #7594's and the from-space cap's
      — a
      signal suppressed by the state it is supposed to leave. Any redesign must
      state, for its chosen signal, why it stays measurable at S=4.

    The existing lock's signal is threshold-invariant (live Eden bytes get moved
    somewhere at any S; survivor-space survival is measured over the previous
    intake, which exists at any S ≥ 2). So the signal is sound and the remaining
    problem is purely: decide sooner.

    The determinism constraint (#7432), and what it actually forbids

    #7432 found that a valve flipped part-way through a cycle breaks
    gc-ratchet counter determinism: the semantic counters then depend on how far
    the cycle had progressed when the flip happened, which varies run to run.

    That forbids exactly one thing — re-deciding S while objects are being moved.
    It does not forbid deciding S earlier, as long as the decision is fixed at
    cycle entry and derived from state written by a completed prior collection.
    Then every object in the cycle sees one threshold and the counters are a pure
    function of (heap state at entry, threshold at entry), which is what the
    ratchet requires. All three proposals below respect that.

    P1 — seed the decision from the last FULL collection

    The issue already records that the feedback loop is fed only by copying
    minors
    (retune_after_scavenge's sole production caller is
    gc/copying.rs), and that this is self-reinforcing: once a workload escalates
    to fulls the loop goes blind exactly when it is needed. On this workload a
    full runs before the first copying minor.

    A full mark-sweep already walks the whole live set. Have it record the young
    generation's live fraction — young_live_bytes / young_in_use_bytes at the
    full — and let a fraction ≥ 90% seed S=1 for the next copying minor.

    • Measurable at S=4: a full's marking does not depend on the tenuring
      threshold at all, so the signal is not suppressed by the state it exits.
    • Deterministic: written at the end of a completed collection, read at the
      entry of the next.
    • Kills the latency: the decision is available before the first copying
      minor rather than after the second.
    • Failure mode, stated: a program whose nursery happens to be mostly-live
      at one full but is normally churn-heavy promotes one cycle's worth of
      short-lived objects into old-gen, and pays an old-gen reclaim to get them
      back. Bounded by the existing unlock path (influx below desired/4 for two
      cycles), but the exposure is one cycle of Eden, i.e. up to the nursery cap.

    P2 — decide per age band at entry, not one number per cycle

    S is currently one number for the whole cycle. It could instead be a per-band
    decision fixed at entry: age-0 objects promote iff the entry-time measurement
    says so, older cohorts keep the ladder. Still not a mid-cycle valve — the band
    boundaries and their thresholds are all fixed before the first object moves.

    This mostly buys precision rather than latency, so it is a refinement of P1
    rather than an alternative to it.

    P3 — pretenure by size, with no measurement at all

    Promote any Eden survivor above N bytes directly to old-gen. Decidable per
    object from the object alone: no cross-cycle state, no signal that can be
    suppressed, and trivially deterministic. HotSpot's PretenureSizeThreshold is
    the same idea. The records in this workload are ~1 KB, and large objects are
    also the ones that cost the most to copy twice.

    • Failure mode, stated: a workload of large short-lived objects — a
      parse-and-discard loop over big buffers — promotes pure garbage into old-gen
      and converts a cheap scavenge into an old-gen reclaim. This is the classic
      pretenuring regression and it is why N must be set from measurement across
      the whole bench suite, not from this one workload.

    How to measure it without repeating the two traps this issue has already hit

    1. Assert the subject was live. The "10.3x promote-on-first-copy" probe
      above was an artifact: forcing tenuring_survivals() == 1 also drives
      copied_minor_promotable_active_survivor_bytes(), so it changed
      triggering, and the fast arm ran zero collections. Any candidate must
      gate its verdict on copied_objects > 0 && promoted_bytes > 0 (gc-matrix: --pressure disables the very path #7019 added — the 'default' arm runs ZERO copying minors on all 22 corpus rows #7024/gc-matrix: the 'moved' liveness counter sums the C4b mark-sweep evacuation with the copying minor, so a move-arm can be green with zero scavenged objects #7025).
    2. The signature that separates a real win from that artifact is specific:
      promote-on-first-copy should roughly halve total bytes copied while
      leaving the collection count unchanged
      . If the collection count drops
      instead, the arm is not promoting early — it is not collecting. Report both
      counters side by side, per arm.
    3. RSS is the binding constraint (GC: tree.ts scavenge-on peak RSS (221 MB) still loses to scavenge-off (103 MB) — old-gen churn + young capacity high-water #7438). Promoting earlier moves bytes
      into old-gen sooner, which raises the old-gen high-water mark even when
      total live is identical; it must go through the ratchet, and the
      12_large_live_set heap_total_bytes row is the one to watch — perf(gc): live-proportional collection budgets at both generations (#7592) #7596
      already spent +36% there.
  6. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Re-derived the next lever on current main — #7630's ordering could not survive #7624 + #7633, and the remaining time is two specific pieces of pure overhead

    main @ 88e0812a7 (v0.5.1363), pinned quiet mini, canonical fixture (w7592/fix/input_{200k,500k}.json; the 500k one is byte-identical to benchmarks/honest_bench/assets/input.json), PERRY_NO_AUTO_OPTIMIZE=1, runtime pinned via PERRY_RUNTIME_DIR.

    Phase split — 5 rounds, spread ≤ 2 ms per cell

    phase 200k 500k
    readFileSync 24 ms 2.1% 58 ms 1.9%
    JSON.parse 237 ms 20.5% 594 ms 19.7%
    build_out 731 ms 63.2% 1,956 ms 64.9%
    JSON.stringify 111 ms 9.6% 270 ms 9.0%
    writeFileSync 11 ms 1.0% 28 ms 0.9%
    fnv1a 43 ms 3.7% 108 ms 3.6%
    total 1,156 ms 3,015 ms
    peak RSS 433 MB 977 MB

    GC census, attributed per phase (PERRY_GC_TRACE, markers interleaved into the same stream)

    phase cycles pause copied promoted
    200k parse 1 (full/old_gen_bytes) 0.9 ms 0 0
    200k build_out 2 (full/arena_bytes + minor/arena_bytes) 695.8 ms = 95.2% of the phase 0 B / 0 obj 113,226,592 B / 1,657,965 obj
    500k parse 1 (full/old_gen_bytes) 1.9 ms 0 0
    500k build_out 2 (same kinds) 1,871.8 ms = 95.7% of the phase 0 B / 0 obj 280,996,456 B / 4,117,014 obj

    copied_bytes = 0 with the whole cohort promoted is #7613 holding: the long-lived cohort is copied once, not twice. Every other phase runs zero collections.

    So the cadence work is done — there are only two collections in the hot phase, and the question is what those two cost.

    Where those two collections go (the GC's own phase_us, 500k)

    full  748.3 ms : trace_worklist 388.3 | build_valid_pointer_set 245.5 | remembered_set_marking 44.4
                     | atomic_finalize 37.2 | sweep 22.4
    minor 1123.5 ms: copying_nursery 666.2 recorded ... 457 ms NOT inside any recorded phase
    mutator        : ~85 ms   <- the actual out.push({...}) loop
    

    Fresh symbolicated leaf profile

    Sample window, stated: 500k, sample <pid> 2 1 fired the instant the workload's PHASE_BEGIN build_out stderr marker appeared. Marker at epoch …574903; sampler start …574905 (2 ms attach lag); sampler end …576977; build_out ended at …576973. The window is [+2 ms, +2074 ms] against a phase spanning [0, +2070 ms] — 99.9% inside build_out, 4 ms of tail. 1504 leaf samples on the main thread. (200k cannot be windowed this way: sample's duration truncates below 1 s and build_out at 200k is 775 ms, so ~20% of any window falls outside the phase. The 500k window is the one to cite.)

    Leaf (self) samples, each std/hashbrown/libc leaf charged to its nearest enclosing perry-runtime module so container costs land on the family that called them:

    family self samples share
    gc::trace (mark + rewrite + census) 392 26.1%
    gc::copying (evacuation + eligibility preflight) 356 23.7%
    gc::layout (trace-side slot descriptors) 191 12.7%
    arena::page_meta 133 8.8%
    _tlv_get_addr 101 6.7%
    gc::barrier + remembered set 101 6.7%
    mimalloc / memmove 76 5.1%
    everything else 154 10.2%

    Top single leaves: trace_heap_rewrite_slots 220 (14.6%), BTreeSet::insert 139 (9.2%), _tlv_get_addr 101, classify_heap_space 90, PtrHashSet::insert 79, GcMutableSlotDescriptor::visit_slots 58.

    But the family table is the wrong unit for picking a lever. The subtree decomposition under gc_check_trigger (1419 of 1504 samples — 94.3% of build_out) is:

    work item samples share of build_out
    copying minor: evacuation + RS 542 36.0%
    copying minor: eligibility preflight (2 reachability walks) 334 22.2%
    full: mark propagation 271 18.0%
    full: valid-pointer census (BTreeSet build) 141 9.4%
    full: root scan + atomic finalize + sweep 87 5.8%
    the out.push({...}) loop itself ~85 5.7%

    The two bolded rows are 702 ms of the 1,956 ms build_out — 36% of the phase spent on work that produces no collection result at all. The preflight number also reconciles independently with the 457 ms of minor pause that sits outside any recorded phase_us.

    Both were probed, and both are real

    Env-gated probes on one binary, 5 interleaved rounds, output SHA identical on every row:

    arm build_out 200k build_out 500k total 200k total 500k peak RSS 200k
    main 734 ms 1,963 ms 1,158 ms 3,024 ms 440 MB
    census from runs 669 (−8.9%) 1,840 (−6.3%) 1,092 (−5.7%) 2,886 (−4.6%) 429 MB
    skip preflight 602 (−18.0%) 1,536 (−21.8%) 1,025 (−11.5%) 2,586 (−14.5%) 439 MB
    both 537 (−26.8%) 1,416 (−27.9%) 958 (−17.3%) 2,465 (−18.5%) 428 MB

    Each probe's subject was verified live rather than merely "nothing threw" (#7024/#7025) — per-cycle, 500k:

    main census arm preflight arm
    full phase_us.build_valid_pointer_set 245.1 ms 22.8 ms 245.1 ms
    full phase_us.trace_worklist 393.8 ms 493.1 ms (+99) 393.8 ms
    full pause 754.3 ms 633.4 ms (−16.0%) 754.3 ms
    minor pause 1,121.3 ms 1,121.3 ms 684.7 ms (−38.9%)
    minor layout_scans.pointer_slots_read 22,041,026 22,041,026 13,827,513
    promoted_bytes / freed_bytes / cycle kinds — bit-identical bit-identical

    Verdict

    The top lever is the copying-minor eligibility preflight (21.8% of build_out, 14.5% of total wall). It is a full extra transitive walk of the young graph, plus a PtrHashSet sized to it, to answer two booleans — and one of those booleans (MallocRegistryUnavailable) is already O(1), while the other's producers are three countable GC_FLAG_PINNED sites. I am not shipping it here: it is a guard on the moving collector, and removing it needs a completeness gate over the pin sites, a sabotage test, and a deliberate ratchet counter shift. Filed separately with the full measurement and design.

    Shipping the second one, because it needs no new invariant: the full's valid-pointer census maintained a BTreeSet shadowing the address-ordered run vector it was already building. Same membership answer, same set, one B-tree insert per live arena object deleted. 245.5 ms → 22.8 ms on a 748.3 ms full, net −16.0% on that collection after paying +99 ms back in lookups (recovered by a contiguous run-fence array — see the PR), and peak RSS never regresses. Every semantic counter is bit-identical by construction.

    Two notes on the previous step-zero profile (#7630)

    Its fixture was not the canonical one — different generator, variable-length tags (including empty) vs always-3, float score, string zip, 12.9% fewer objects for the same record count. Its "101 MB promoted" is 105,888,096 B / 1,467,591 obj on that fixture; the canonical figure is 113,226,592 B / 1,657,965 obj, which reproduces the maintainer's #7624 audit number to within one object. And its sample window (sleep 0.5; sample $PID 3 on a 1.3 s run) started inside JSON.parse and ran to process exit, so its leaf list mixes build_out with stringify and fnv1a. Details on #7630.

  7. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Second lever taken: #7646 — the full collection's valid-pointer census kept a BTreeSet shadowing the address-ordered run vector the census walk was already building, one insert per live arena object for a question the runs could already answer.

    build_out −15.2% / −15.4% (200k / 500k), total wall −9.9% / −10.1%, the full collection's pause 294.4 → 182.8 ms and 750.3 → 458.3 ms; phase_us.build_valid_pointer_set 247.0 → 23.4 ms at 500k. Output SHA identical, every non-timing gc_cycle field identical, perry-runtime 1902/1902.

    Two details worth carrying forward:

    • The obvious version of this change is a regression. Searching arena_runs: Vec<Vec<usize>> directly (calling run.first() per probe) moved trace_worklist the wrong way, +99 ms over ~9.6M lookups, netting only −6.3%. The B-tree was not buying a better probe, it was buying a contiguous one — mirroring each run's first key into a flat arena_run_firsts fence array turned +99 ms into −62 ms and took the change from −6.3% to −15.4%.
    • The one gc-ratchet cell that moved is 100% conservative-scan false-root residue, and classify proves it to the byte. 08_map_set_sidetables.heap_used_bytes fell 1,548,960 → 1,512,456; the base arm's precise (scan-off) reading is exactly 1,512,456, i.e. the fix's conservative reading equals the base's precise one. Precise retention is byte-identical on all twelve probes; the excess column went 36,504 B → 0 B.

    The larger lever on this profile — the copying minor's second traversal of the young graph, 21.8% of build_out — is #7645, deliberately not taken here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions