Repository navigation
perf: json_pipeline at 500k records is 97.6x bun (60.4s vs 618ms) while the same workload at 100 records BEATS bun — a scaling cliff, not a constant factor #7592
Description
Activity
Root cause: this is GC pacing, not JSON
Per the issue's own instruction, the 60 s was split before anything was touched. An instrumented copy of the workload (
Date.now()between the five phases, nothing else changed), 500k records:phase time share readFileSync 89 ms 0.1% JSON.parse 742 ms 1.2% build_out 57,409 ms 94.1% JSON.stringify 1,451 ms 2.4% writeFileSync 90 ms 0.1% fnv1a 1,250 ms 2.0% JSON.parseis not the problem — it handles 107 MB in 742 ms and scales linearly (16 ms at 25k → 742 ms at 500k, 20× data for 46× time). The cliff is entirely theout.push({...})loop.It is superlinear, and it is GC
build_outns/record roughly doubles every time N doubles:records build_out ns/record vs prev 25,000 41 ms 3,246 — 50,000 455 ms 18,156 5.59× 100,000 2,043 ms 40,769 2.25× 200,000 8,451 ms 84,341 2.07× 500,000 56,833 ms 227,354 2.70× From
PERRY_GC_TRACE,build_outis ~100 % GC pause — 8,840 ms of pause against 8,633 ms ofbuild_outat 100k. Bytes copied per record are flat (~1.1 KB), so this is not a copying/promotion problem; it is collection frequency against a growing live set.The mechanism
effective_next_arena_trigger()ismin(adaptive_base, scavenge_nursery_cap_effective_bytes()), and that cap isPERRY_GC_SCAVENGE_NURSERY_MB(default 16 MB) ×NURSERY_CAP_SCALE, wheregc/tenuring.rssetsNURSERY_CAP_SCALE_MAX = 4. So the cap has a hard ceiling of 64 MB.Collection cadence is therefore a fixed number of allocated bytes, while each collection costs O(live). Total GC work is
(alloc / cap) × O(live)= O(N²). Work per allocated byte isc·L/cap, which is bounded only ifcap ∝ L.That means no constant ceiling fixes this — it only moves the cliff to a larger N, which is exactly what a fixed-N cap sweep shows:
records 64 MB 256 MB 1024 MB 100,000 559 ns/rec 619 559 200,000 85,749 ✗ 5,190 5,469 500,000 229,786 ✗ 233,634 ✗ 764 ✓ With a live-set-sized nursery, 500k
build_outgoes 57,441 ms → 191 ms (300×) and ns/record goes flat — the quadratic is gone, not merely reduced.Why the existing adaptive scale does not save it
NURSERY_CAP_SCALEwas introduced for exactly this shape, but it cannot reach its own ceiling here, for two independent reasons:- It is fed by copying minors only. The sole production caller of
retune_after_scavengeisgc/copying.rs:1280. Full collections and non-copying minor fallbacks feed it nothing. That is self-reinforcing: once the workload escalates to fulls, the feedback loop goes blind and the cap stays pinned, which produces more fulls. - The debounce outruns the workload. Growth is ×2 per two consecutive qualifying cycles, and two steps are needed to reach ×4 — so ≥4 feeding cycles. The 50k run had 3 cycles total, 2 of them feeding.
Proposed fix
Make the collection budget proportional to the live set rather than a bounded multiple of a fixed base —
cap = max(base, live_bytes / K), the standard Go-GOGC / V8 heap-growth shape.Ktrades RSS for GC time, with peak ≈L(1 + 1/K).The RSS cost is the real decision here, since this necessarily raises the high-water mark on live-set-bound workloads — so it needs to go through the RSS ratchet rather than be argued from the wall-clock win alone. I am measuring that next.
Also worth fixing independently of the pacing change: the feedback loop should be fed by every collection that observes a live set, not only copying minors, so it cannot go blind exactly when it is needed most.
- It is fed by copying minors only. The sole production caller of
Two probes, and a correction to the fix I proposed above
I built the live-proportional cap I described and measured it. It is the wrong fix, and the reason is instructive.
Probe 1 — live-proportional nursery cap: 5.9×, but structurally broken
cap = max(base, live_bytes)wherelive_bytesis arena in-use recorded at the end of every cycle. 500kbuild_out57.4 s → 9.7 s, but ns/record still climbed with N. The diagnostic shows why:[gc-live] live= 44,748,304 cap= 44,748,304 from_space= 2,656 [gc-live] live=114,343,400 cap=114,343,400 from_space=112,195,992 [gc-live] live=118,978,680 cap=118,978,680 from_space=116,372,472The cap gates
young_scavenge_cap_due(), which compares from-space occupancy against it. But arena in-use includes the young generation, and on this workload from-space is nearly all of live — socap = liveis a fixed point that from-space can never cross. Scavenging stopped entirely (0 cycles at 200k), the nursery was never evacuated, and the residual 2.7 s is the mutator paying for a huge un-evacuated young space.Any live-proportional cap has to be proportional to the tenured live set, not to total arena in-use, or it self-references. That also means it cannot work on its own — see below.
Probe 2 — promote on first copy: 10.3×, and the quadratic is gone
Forcing
tenuring_survivals() == 1:records main ns/rec promote-on-first-copy speedup 25,000 3,167 3,087 1.0× 50,000 19,513 14,246 1.4× 100,000 41,866 18,798 2.2× 200,000 85,529 19,591 4.4× 500,000 231,570 22,534 10.3× 500k
build_out57.4 s → 5.6 s, and ns/record is flat from 100k on (18.8k → 19.6k → 22.5k). That is the quadratic gone, not merely reduced.The mechanism is the one the module header already describes: these objects are long-lived (they are what
outaccumulates), but at S=4 they stay young and get fully re-copied on every cycle.Why the existing survival-rate lock does not do this
It is not missing the condition, it is a cycle too late. The lock keys on
prev_copied— it needs a previous copying minor to have filled the survivor space, so it cannot engage until the second one, and it only changes the threshold for the cycle after that. This workload has ~6 cycles, most of them fulls (which never callretune_after_scavengeat all).I tried adding the Eden-side form of the same proof — lock when ≥90% of Eden survives, which is available on the first copying minor rather than the second. Measured 1.0× across the board: still a no-op, because it too only takes effect on the next copying minor, and on this workload that cycle mostly never comes. Reverted rather than shipped.
Where that leaves the fix
The two levers are complementary and neither works alone:
- Promotion has to engage on the first copy, not one or two copying minors later. The signal has to be acted on within the cycle that observes it, or read from a source that survives across cycles — the current loop is fed only by copying minors, so it goes blind exactly when it is needed.
- Then a tenured-live-proportional collection budget becomes well-defined, because old-gen starts tracking the real live set instead of sitting at ~2 MB while everything piles up in from-space.
Both change promotion rates and heap high-water marks, so this needs to go through the RSS ratchet before it is a candidate — the 300× available here is a wall-clock number, and #7438 is a standing reminder that this collector's RSS behaviour is the binding constraint. Note also #7432's finding that a mid-cycle valve breaks ratchet counter determinism, which rules out the most obvious way of acting within the cycle.
No code from these probes is proposed for merge; the tree is back at
main.Found it — and my probe result above was an artifact
Correction first. The "promote on first copy = 10.3×" measurement in my previous comment does not mean what I said it did.
tenuring_survivals()also feedscopied_minor_promotable_active_survivor_bytes(), which drives the promotion-handoff pacing — so forcing it to 1 changed triggering, not just promotion. That arm ran zero collections. It was fast because it did not collect, not because it promoted early.Two measurement traps produced that, both worth recording:
- The auto-optimize path relinks the runtime with
--no-default-features, which silently drops thediagnosticsfeature.PERRY_GC_TRACEthen printsdiagnostics feature disabledinstead of JSON, and my cycle counts read 0 — for three different arms in a row. It also meant the baseline binary (built before I cleared that cache) had diagnostics while the new one did not, so the arms were not feature-identical. The fix is to build-p perry -p perry-runtime-static -p perry-stdlib-staticand link with a pinnedPERRY_RUNTIME_DIRandPERRY_NO_AUTO_OPTIMIZE=1. PERRY_GC_TRACElines are enormous; piping them SIGPIPEs and truncates the cycle list, which is where an earlier "21 fulls" came from.
The actual root cause: a livelock
With working diagnostics, the 200k trace is unambiguous — 22 collections:
19 x full/survivor_promotion_bytes each freeing 0.0 MB at ~400 ms 1 x full/old_gen_bytes 1 x full/arena_bytes 1 x minor/arena_bytes promoted: 0.0 MBcopied_minor_promotion_handoff_duereplaces a minor with a full mark-sweep to make room in old-gen for survivors about to be promoted. But a full mark-sweep is non-moving — it promotes nothing. It cannot relieve the pressure that scheduled it: the survivor space still holds the same 108 MB, the reclaim baseline it resets does not count those bytes, and the predicate is true again at the next minor. The copying minor that would have done the promotion never runs.That is 7.6 s of an 8.6 s phase spent freeing nothing. Peak RSS is the same whether those 19 collections happen or not, which is the tell: they were pure loss.
Fix — #7594
Latch it: one handoff per copying minor. The handoff makes room; the copying minor performs the promotion that consumes it.
cycles pause survivor_promotionfullspromoted main 22 8,782 ms 19 0.0 MB #7594 6 2,828 ms 1 110.0 MB build_outat 500k: 57,242 ms → 10,551 ms (5.4×), output hash identical. RSS is up 17–24 % at the large sizes because the run now collects 6 times instead of 22 — a real trade, stated in the PR.On the ratchet, all 12 probes agree with main on every semantic counter across 144 metric medians (measured as a back-to-back A/B — the pinned baseline is 0.5.1315 on another host and cannot separate this change from drift).
What is left
build_outis still ~28,500 ns/record and not flat. With a nursery large enough to avoid collecting during the loop the same phase runs at 764 ns/record, so roughly another order of magnitude is available, and it is the collection-budget question: the cap is a constant (16 MB ×NURSERY_CAP_SCALE_MAX), so cadence is independent of live-set size.The obvious repair does not work as written, and it is worth writing down why: a live-proportional cap gates on from-space occupancy, and on this workload from-space is nearly all of live, so
cap = liveis a fixed point that from-space can never cross — scavenging stops entirely. Any such budget has to key on the tenured live set, which in turn only becomes meaningful once promotion actually happens. Leaving this issue open for that half.- The auto-optimize path relinks the runtime with
Resolved: the cliff is gone (#7594 + #7596, both merged)
Final three-arm measurement, identically built, hash-verified per row:
records main (before) #7594 #7594+#7596 speedup ns/rec 25,000 42 ms 42 42 1.0× 3,325 100,000 2,257 ms 1,499 863 2.6× 17,221 200,000 8,889 ms 3,505 2,617 3.4× 26,118 500,000 61,589 ms 12,139 5,784 10.6× 23,138 ns/record is now flat within ~30 % across a 20× size range; before, it grew 70×. The two fixes:
- perf(gc): break the survivor-promotion handoff livelock (#7592) #7594 — the survivor-promotion handoff livelock: 19 consecutive full collections each freeing 0.0 MB, because a non-moving full can never relieve the survivor pressure that scheduled it. Latched to one handoff per copying minor.
- perf(gc): live-proportional collection budgets at both generations (#7592) #7596 — no constant band may pace a collector whose per-cycle cost is O(live): the handoff now requires current old-gen pressure (not predicted), promoted bytes credit the reclaim baseline (they are live by construction), and both the old-reclaim growth band and the scavenge nursery cap are live-proportional (the nursery cap keyed on tenured occupancy — total arena in-use is a fixed point from-space can never cross).
What remains, tracked elsewhere
- perf(gc): allocation-site pretenuring — long-lived cohorts are copied twice (Eden→survivor→old) #7598 — the long-lived cohort is still copied twice (Eden→survivor→old, 3.9 s of the remaining 5.1 s). Needs allocation-site pretenuring; every cheaper fix was measured or ruled out (details in that issue).
- The workload's non-GC phases (
fnv1a1,250 ms,JSON.stringify1,451 ms) are the integer-primitives/repsel story (codegen: type-directed specialization — emit native ops where the type is statically known (#5334 Tier 3 / Lever E) #5497/Perf: untyped typed-array element access ~6x slower (per-access thread-local kind lookup) — bcryptjs cost-12 ~28s vs ~250ms #5525), not GC.
Closing as fixed — the reported scaling cliff (97.6× bun at 500k, beating bun at 100 records) was the quadratic, and the quadratic is dead.
- added a commit that references this issue
on Aug 7, 2026 Design note: getting promote-on-first-copy to engage in time
#7596's stated follow-up is that surviving bytes are copied twice —
Eden → survivor space → old-gen — 268 MB each way, ~3.9 s of the remaining
~5.1 s ofbuild_out. Promoting on the first copy halves that. This is the
design for it; no code is proposed here.The obstacle is latency, not blindness — and it is worth being precise
Two failed probes are already recorded above, and they failed for different
reasons. Conflating them sends the next attempt at the wrong target.- The survival-rate lock (
gc/tenuring.rs) already computes exactly the
right condition: a substantial survivor intake of which ≥90% comes back out
alive means the aging round filters nothing. It is not missing the
condition. It keys onprev_copied, so it needs a previous copying minor
to have filled the survivor space, and it changes the threshold for the
cycle after that. Two cycles of latency on a run that has about six. - The Eden-side variant measured 1.0x, and the note above attributes that
to the same latency. Half true. The deeper reason is that the quantity it
reached for is not observable at the threshold it is trying to leave:
promoted_bytesis zero by construction at S=4, so "was the promotion
rate high last cycle?" can never answer yes while S=4 holds. That is a
fixed point of the same family as perf(gc): break the survivor-promotion handoff livelock (#7592) #7594's and the from-space cap's — a
signal suppressed by the state it is supposed to leave. Any redesign must
state, for its chosen signal, why it stays measurable at S=4.
The existing lock's signal is threshold-invariant (live Eden bytes get moved
somewhere at any S; survivor-space survival is measured over the previous
intake, which exists at any S ≥ 2). So the signal is sound and the remaining
problem is purely: decide sooner.The determinism constraint (#7432), and what it actually forbids
#7432 found that a valve flipped part-way through a cycle breaks
gc-ratchet counter determinism: the semantic counters then depend on how far
the cycle had progressed when the flip happened, which varies run to run.That forbids exactly one thing — re-deciding S while objects are being moved.
It does not forbid deciding S earlier, as long as the decision is fixed at
cycle entry and derived from state written by a completed prior collection.
Then every object in the cycle sees one threshold and the counters are a pure
function of (heap state at entry, threshold at entry), which is what the
ratchet requires. All three proposals below respect that.P1 — seed the decision from the last FULL collection
The issue already records that the feedback loop is fed only by copying
minors (retune_after_scavenge's sole production caller is
gc/copying.rs), and that this is self-reinforcing: once a workload escalates
to fulls the loop goes blind exactly when it is needed. On this workload a
full runs before the first copying minor.A full mark-sweep already walks the whole live set. Have it record the young
generation's live fraction —young_live_bytes / young_in_use_bytesat the
full — and let a fraction ≥ 90% seed S=1 for the next copying minor.- Measurable at S=4: a full's marking does not depend on the tenuring
threshold at all, so the signal is not suppressed by the state it exits. - Deterministic: written at the end of a completed collection, read at the
entry of the next. - Kills the latency: the decision is available before the first copying
minor rather than after the second. - Failure mode, stated: a program whose nursery happens to be mostly-live
at one full but is normally churn-heavy promotes one cycle's worth of
short-lived objects into old-gen, and pays an old-gen reclaim to get them
back. Bounded by the existing unlock path (influx belowdesired/4for two
cycles), but the exposure is one cycle of Eden, i.e. up to the nursery cap.
P2 — decide per age band at entry, not one number per cycle
S is currently one number for the whole cycle. It could instead be a per-band
decision fixed at entry: age-0 objects promote iff the entry-time measurement
says so, older cohorts keep the ladder. Still not a mid-cycle valve — the band
boundaries and their thresholds are all fixed before the first object moves.This mostly buys precision rather than latency, so it is a refinement of P1
rather than an alternative to it.P3 — pretenure by size, with no measurement at all
Promote any Eden survivor above N bytes directly to old-gen. Decidable per
object from the object alone: no cross-cycle state, no signal that can be
suppressed, and trivially deterministic. HotSpot'sPretenureSizeThresholdis
the same idea. The records in this workload are ~1 KB, and large objects are
also the ones that cost the most to copy twice.- Failure mode, stated: a workload of large short-lived objects — a
parse-and-discard loop over big buffers — promotes pure garbage into old-gen
and converts a cheap scavenge into an old-gen reclaim. This is the classic
pretenuring regression and it is why N must be set from measurement across
the whole bench suite, not from this one workload.
How to measure it without repeating the two traps this issue has already hit
- Assert the subject was live. The "10.3x promote-on-first-copy" probe
above was an artifact: forcingtenuring_survivals() == 1also drives
copied_minor_promotable_active_survivor_bytes(), so it changed
triggering, and the fast arm ran zero collections. Any candidate must
gate its verdict oncopied_objects > 0 && promoted_bytes > 0(gc-matrix: --pressure disables the very path #7019 added — the 'default' arm runs ZERO copying minors on all 22 corpus rows #7024/gc-matrix: the 'moved' liveness counter sums the C4b mark-sweep evacuation with the copying minor, so a move-arm can be green with zero scavenged objects #7025). - The signature that separates a real win from that artifact is specific:
promote-on-first-copy should roughly halve total bytes copied while
leaving the collection count unchanged. If the collection count drops
instead, the arm is not promoting early — it is not collecting. Report both
counters side by side, per arm. - RSS is the binding constraint (GC: tree.ts scavenge-on peak RSS (221 MB) still loses to scavenge-off (103 MB) — old-gen churn + young capacity high-water #7438). Promoting earlier moves bytes
into old-gen sooner, which raises the old-gen high-water mark even when
total live is identical; it must go through the ratchet, and the
12_large_live_setheap_total_bytesrow is the one to watch — perf(gc): live-proportional collection budgets at both generations (#7592) #7596
already spent +36% there.
- The survival-rate lock (
Re-derived the next lever on current
main— #7630's ordering could not survive #7624 + #7633, and the remaining time is two specific pieces of pure overheadmain@88e0812a7(v0.5.1363), pinned quiet mini, canonical fixture (w7592/fix/input_{200k,500k}.json; the 500k one is byte-identical tobenchmarks/honest_bench/assets/input.json),PERRY_NO_AUTO_OPTIMIZE=1, runtime pinned viaPERRY_RUNTIME_DIR.Phase split — 5 rounds, spread ≤ 2 ms per cell
phase 200k 500k readFileSync24 ms 2.1% 58 ms 1.9% JSON.parse237 ms 20.5% 594 ms 19.7% build_out731 ms 63.2% 1,956 ms 64.9% JSON.stringify111 ms 9.6% 270 ms 9.0% writeFileSync11 ms 1.0% 28 ms 0.9% fnv1a43 ms 3.7% 108 ms 3.6% total 1,156 ms 3,015 ms peak RSS 433 MB 977 MB GC census, attributed per phase (
PERRY_GC_TRACE, markers interleaved into the same stream)phase cycles pause copied promoted 200k parse1 (full/ old_gen_bytes)0.9 ms 0 0 200k build_out2 (full/ arena_bytes+ minor/arena_bytes)695.8 ms = 95.2% of the phase 0 B / 0 obj 113,226,592 B / 1,657,965 obj 500k parse1 (full/ old_gen_bytes)1.9 ms 0 0 500k build_out2 (same kinds) 1,871.8 ms = 95.7% of the phase 0 B / 0 obj 280,996,456 B / 4,117,014 obj copied_bytes = 0with the whole cohort promoted is #7613 holding: the long-lived cohort is copied once, not twice. Every other phase runs zero collections.So the cadence work is done — there are only two collections in the hot phase, and the question is what those two cost.
Where those two collections go (the GC's own
phase_us, 500k)full 748.3 ms : trace_worklist 388.3 | build_valid_pointer_set 245.5 | remembered_set_marking 44.4 | atomic_finalize 37.2 | sweep 22.4 minor 1123.5 ms: copying_nursery 666.2 recorded ... 457 ms NOT inside any recorded phase mutator : ~85 ms <- the actual out.push({...}) loopFresh symbolicated leaf profile
Sample window, stated: 500k,
sample <pid> 2 1fired the instant the workload'sPHASE_BEGIN build_outstderr marker appeared. Marker at epoch…574903; sampler start…574905(2 ms attach lag); sampler end…576977;build_outended at…576973. The window is [+2 ms, +2074 ms] against a phase spanning [0, +2070 ms] — 99.9% insidebuild_out, 4 ms of tail. 1504 leaf samples on the main thread. (200k cannot be windowed this way:sample's duration truncates below 1 s andbuild_outat 200k is 775 ms, so ~20% of any window falls outside the phase. The 500k window is the one to cite.)Leaf (self) samples, each std/hashbrown/libc leaf charged to its nearest enclosing perry-runtime module so container costs land on the family that called them:
family self samples share gc::trace(mark + rewrite + census)392 26.1% gc::copying(evacuation + eligibility preflight)356 23.7% gc::layout(trace-side slot descriptors)191 12.7% arena::page_meta133 8.8% _tlv_get_addr101 6.7% gc::barrier+ remembered set101 6.7% mimalloc / memmove 76 5.1% everything else 154 10.2% Top single leaves:
trace_heap_rewrite_slots220 (14.6%),BTreeSet::insert139 (9.2%),_tlv_get_addr101,classify_heap_space90,PtrHashSet::insert79,GcMutableSlotDescriptor::visit_slots58.But the family table is the wrong unit for picking a lever. The subtree decomposition under
gc_check_trigger(1419 of 1504 samples — 94.3% ofbuild_out) is:work item samples share of build_outcopying minor: evacuation + RS 542 36.0% copying minor: eligibility preflight (2 reachability walks) 334 22.2% full: mark propagation 271 18.0% full: valid-pointer census ( BTreeSetbuild)141 9.4% full: root scan + atomic finalize + sweep 87 5.8% the out.push({...})loop itself~85 5.7% The two bolded rows are 702 ms of the 1,956 ms
build_out— 36% of the phase spent on work that produces no collection result at all. The preflight number also reconciles independently with the 457 ms of minor pause that sits outside any recordedphase_us.Both were probed, and both are real
Env-gated probes on one binary, 5 interleaved rounds, output SHA identical on every row:
arm build_out200kbuild_out500ktotal 200k total 500k peak RSS 200k main 734 ms 1,963 ms 1,158 ms 3,024 ms 440 MB census from runs 669 (−8.9%) 1,840 (−6.3%) 1,092 (−5.7%) 2,886 (−4.6%) 429 MB skip preflight 602 (−18.0%) 1,536 (−21.8%) 1,025 (−11.5%) 2,586 (−14.5%) 439 MB both 537 (−26.8%) 1,416 (−27.9%) 958 (−17.3%) 2,465 (−18.5%) 428 MB Each probe's subject was verified live rather than merely "nothing threw" (#7024/#7025) — per-cycle, 500k:
main census arm preflight arm full phase_us.build_valid_pointer_set245.1 ms 22.8 ms 245.1 ms full phase_us.trace_worklist393.8 ms 493.1 ms (+99) 393.8 ms full pause 754.3 ms 633.4 ms (−16.0%) 754.3 ms minor pause 1,121.3 ms 1,121.3 ms 684.7 ms (−38.9%) minor layout_scans.pointer_slots_read22,041,026 22,041,026 13,827,513 promoted_bytes/freed_bytes/ cycle kinds— bit-identical bit-identical Verdict
The top lever is the copying-minor eligibility preflight (21.8% of
build_out, 14.5% of total wall). It is a full extra transitive walk of the young graph, plus aPtrHashSetsized to it, to answer two booleans — and one of those booleans (MallocRegistryUnavailable) is already O(1), while the other's producers are three countableGC_FLAG_PINNEDsites. I am not shipping it here: it is a guard on the moving collector, and removing it needs a completeness gate over the pin sites, a sabotage test, and a deliberate ratchet counter shift. Filed separately with the full measurement and design.Shipping the second one, because it needs no new invariant: the full's valid-pointer census maintained a
BTreeSetshadowing the address-ordered run vector it was already building. Same membership answer, same set, one B-tree insert per live arena object deleted. 245.5 ms → 22.8 ms on a 748.3 ms full, net −16.0% on that collection after paying +99 ms back in lookups (recovered by a contiguous run-fence array — see the PR), and peak RSS never regresses. Every semantic counter is bit-identical by construction.Two notes on the previous step-zero profile (#7630)
Its fixture was not the canonical one — different generator, variable-length
tags(including empty) vs always-3, floatscore, stringzip, 12.9% fewer objects for the same record count. Its "101 MB promoted" is105,888,096 B / 1,467,591 objon that fixture; the canonical figure is113,226,592 B / 1,657,965 obj, which reproduces the maintainer's #7624 audit number to within one object. And its sample window (sleep 0.5; sample $PID 3on a 1.3 s run) started insideJSON.parseand ran to process exit, so its leaf list mixesbuild_outwithstringifyandfnv1a. Details on #7630.Second lever taken: #7646 — the full collection's valid-pointer census kept a
BTreeSetshadowing the address-ordered run vector the census walk was already building, one insert per live arena object for a question the runs could already answer.build_out−15.2% / −15.4% (200k / 500k), total wall −9.9% / −10.1%, the full collection's pause 294.4 → 182.8 ms and 750.3 → 458.3 ms;phase_us.build_valid_pointer_set247.0 → 23.4 ms at 500k. Output SHA identical, every non-timinggc_cyclefield identical,perry-runtime1902/1902.Two details worth carrying forward:
- The obvious version of this change is a regression. Searching
arena_runs: Vec<Vec<usize>>directly (callingrun.first()per probe) movedtrace_worklistthe wrong way, +99 ms over ~9.6M lookups, netting only −6.3%. The B-tree was not buying a better probe, it was buying a contiguous one — mirroring each run's first key into a flatarena_run_firstsfence array turned +99 ms into −62 ms and took the change from −6.3% to −15.4%. - The one gc-ratchet cell that moved is 100% conservative-scan false-root residue, and
classifyproves it to the byte.08_map_set_sidetables.heap_used_bytesfell 1,548,960 → 1,512,456; the base arm's precise (scan-off) reading is exactly 1,512,456, i.e. the fix's conservative reading equals the base's precise one. Precise retention is byte-identical on all twelve probes; the excess column went 36,504 B → 0 B.
The larger lever on this profile — the copying minor's second traversal of the young graph, 21.8% of
build_out— is #7645, deliberately not taken here.- The obvious version of this change is a regression. Searching
- added 4 commits that reference this issue
on Aug 8, 2026
Measured in today's public-baseline regeneration (v0.5.1335,
2ba59501b, pinned quiet M1 Max)honest_bench's JSON pipeline, 20 runs per cell, output verified against theBun reference 20/20 for every language — so this is a pure performance
finding, not a correctness one.
json_pipeline_small(100 records)json_pipeline_full(500,000 records, 107.5 MB)image_convolution(4K, 5×5)At 100 records Perry beats bun and node and lands within 5% of Rust and Zig,
on a quarter of bun's memory. At 500k records it is 97.6× bun. Same source
file, same binary, same harness — only the fixture size changes.
That is the shape of the finding: not slowness, a scaling cliff. Across the
5,000× increase in input, bun's time grows 6.8× (startup-dominated at the small
end); Perry's grows 784×. Rust 8.9×, Zig 12.0×, node 7.8×. Perry is the only
implementation whose growth is superlinear in the data, which points at an
algorithmic or GC-pathological term rather than a constant-factor gap.
RSS is 1,064 MB against bun's 579 MB — 1.8×, and notably not 97×. The memory
is not exploding proportionally to the time, which argues against "it simply
swapped" and for real work being repeated.
Where to look
The workload is small and does five things (
workloads/1_json_pipeline/perry/json_pipeline.ts):fs.readFileSynca 107.5 MB string →JSON.parseto 500k objects → a filter +field-derive loop building
out→JSON.stringify(out)→fs.writeFileSync,then an FNV-1a hash over the serialized string.
Each of those is separately suspect at this scale and they should be timed
individually before anything is optimised — this campaign has repeatedly paid
for working from an unattributed headline number. Candidates worth ruling in or
out first:
JSON.parseon 500k objects, which is json tape: full-array scans pay 2.3× the direct parser (field_access 2981ms vs bun 223) — decomposition + roadmap #7478/gc: the JSON tape's large per-parse block interacts badly with the generational collector (3.2x RSS, 25x variance vs mark-sweep) #7539's tape territory. The tapework was measured on
json_polyglot, a different and much smaller shape.out.push({...})loop — 500k object literals with 11 fields each, i.e.the perf(runtime): allocation path spends 34% of self time in _tlv_get_addr — 24× behind Node on object churn with the collector already idle #7469 construction path at a scale nothing else in the suite reaches.
JSON.stringifyof a 500k-element array.charCodeAtover a multi-megabyte string.A leaf profile (
PERRY_DEBUG_SYMBOLS=1+sample) will separate these in onerun; at 60 s per iteration there is plenty of signal.
Two stale claims in the workload source, now disproved
json_pipeline.tscarries a comment block from v0.5.29 stating that (a) thedriver "runs this binary on the 100-record fixture only", and (b) iterating a
large
JSON.parseresult "triggers a GC-scan issue at scale — records allocatedby the JSON parser get swept mid-iteration… above [~200 records] output is
non-deterministic".
Both are now false. The driver does run Perry on the full 500k fixture
(
run.sh:268), and the output matched the Bun reference on 20 of 20 runs.The correctness half was fixed at some point without the comment being updated.
Worth deleting so the next reader does not attribute the slowness to a
corruption bug that no longer exists.
Why this matters beyond the number
README.mdcites this harness by name for its JSON row and quotes the100-record result. The same report's 500k row is 97.6× bun. Whatever the
intent, the README also says "We publish everything, including the workloads
where V8's JIT still beats us — no cherry-picked table can survive an open
harness." Those two things need to be reconciled; tracked separately.