Repository navigation
perf: object representation is now the binding constraint on the retain cluster (72 bytes per 2-field literal; 216 MB written to store 48 MB) #7916
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Aug 12, 2026 The ≥4-field header shrink, priced before building it — and the answer is "yes, but not before the hidden-class work"
Agent
hdr, worktreewt-hdr, basea769fafc6(0.5.1495). Full write-up and
harnesses:gc-handoff/HDR-NOTES.md,m0810/build_hdr.sh,m0810/hdr_metrics.py,
m0810/hdr_pad_patch.sh.1. First, a correction to the target numbers
Removing the three derivable header members (
object_type,field_count,
keys_array) is 16 B, which lands at 40 B for{a,b}and 88 B for
retain_wide— not the 32 / 80 quoted fromREPR-NOTES§4. Those are the
endgame: they additionally require deletingmeta(#6759 Phase B put it in the
header to remove a side-table probe) and mergingclass_idwith
parent_class_id. Also, only removing bothu32s saves anything: dropping
one leaves 12 bytes that re-pad to 16 for the following pointer.2. Measured, not extrapolated: a padding probe measures the shrink's mirror
The removal moves every codegen offset across ~40 files, so before building it I
priced it with its exact mirror — growObjectHeaderby 16 B and read
pad16 → baseas the shrink (72→56 onretain, 120→104 onretain_wide), at the
same absolute sizes, with #7961 in place. 11/11 byte-exact, exit 0 on both arms.bench Δ instr Δ RSS minors pad→base objects/minor ratio retain −25.63 % −26.19 % 6 → 4 1.084 retain1 +3.64 % −18.72 % 3 → 2 0.992 retain_wide +0.43 % −10.52 % 8 → 6 1.133 retain_wide1 −1.30 % −13.42 % 4 → 3 0.945 deeplist +6.25 % −13.77 % 3 → 2 0.877 tree −0.06 % −18.64 % 40 → 40 — tree_wide +0.02 % −10.66 % 40 → 40 — churn / churn_alloc −0.6 % −0.3 % 105 → 88 1.010 shapes / pipeline ≤ +0.4 % ≤ −0.2 % 1→1 / 8→7 — Best-of-3; the
basearm's own spread was up to 3.5 % (dev box at load 200), so
|Δinstr| ≲ 3 % is noise.Cycle counts go DOWN, not up. Every program's minor count falls or is
unchanged — smaller objects allocate fewer total bytes, so a byte-denominated
trigger fires less often.retain's −25.6 % is that: 2.93 M promotions → 2.11 M
inside the exit-bounded window.3. ★ #7961's compensation is clamped OFF for exactly these objects — and the tax is still small
nursery_cap_object_scale_permilleisclamp(mean_surviving_object_bytes / 72, 500, 1000). Measured[gc-tenuring]factors:retain/retain1/deeplist56 B →
777‰;retain_wide/retain_wide1104 B → 1000‰ (clamped);tree_wide
32 B → 500‰ (floored). A 16 B shrink moves neither the clamp nor the floor, so
a ≥4-field object gets no pacing compensation at all.The
objects/minorcolumn is that prediction landing:retain10.992,retain1.084 — the compensated class. Un-compensated this
would be 72/56 = 1.286; fix(gc): denominate the nursery constant band in objects (#7929) #7961 removes essentially all of it.retain_wide1.133 — un-compensated. Predicted 120/104 = 1.154.
But it is 13 % more work spread over 25 % fewer cycles, and it costs +0.43 %
instructions. So the #7929 tax is real, is exactly where the clamp says it is,
and is not a reason to hold the shrink. I predicted it would be, from the
clamp alone, before running the probe; the probe says otherwise.4. ★ The real blocker: the three fields ARE the inline-cache guard
REPR-NOTES§4 prices this as "every codegen offset has to move in lockstep".
That understates it.perry-codegen/src/expr/class_field_inline_guard.rs— the
precheck admitting the proven-receiver clone, the hottest property path in the
compiler — emits in ONE basic block off the receiver:object_type @0 == OBJECT_TYPE_REGULAR class_id @4 == expected_class_id field_count @12 > max_field_index keys_array @16 == @perry_class_keys_* <- "the load-bearing one"Three of those four are the fields being deleted. And
keys_arrayidentity
cannot become a ShapeId compare as things stand:delete inst.bpreserves
class_idwhile compacting inline slots and installing a freshly cloned keys
array (object/delete_rest.rs), and the check is deliberately dynamic because the
deleteshape barrier is module-scoped while receivers alias across modules
(#7143). A ShapeId compare is equivalent only once deletion mints a new
ShapeId — i.e. once shape transitions are real.So the 16 B is not a second
INLINE_SLOT_FLOOR. It is not separable from the
hidden-class transition machinery.5. Recommendation
The prize is real and it is footprint: −10.5 % … −26.2 % peak RSS,
{a,b}at
2.50× payload andretain_wideat 1.38×,retainwriting 120 MB instead of
168 MB — at zero measurable instruction cost on 9 of 11 programs.Do it inside the hidden-class project, not before it. There the same change
removes three guard loads instead of trading them for dependent probes.Two follow-ups worth filing separately:
- fix(gc): denominate the nursery constant band in objects (#7929) #7961's denominator for the tails. The clamp exists because the estimator is
a byte-weighted mean (push_num's is 3600 B). Pacing on the object count the
last minor moved makes both the 1000‰ clamp and the 500‰ floor unnecessary and
takesretain_wide's 1.133 to 1.000. A second-order win (~13 % of one term), not
a prerequisite. deeplist+6.3 % is the one program where a smaller object costs
instructions while doing less GC (3→2 minors, 961 k→562 k promotions). Not
explained; worth a look before the real shrink lands.
- added 4 commits that reference this issue
on Aug 12, 2026 - added a commit that references this issue
on Sep 21, 2026
Summary
Established while decomposing the retain cluster's GC cost (
gc-handoff/RETAIN4-NOTES.md): the collector is no longer the binding constraint on these benchmarks — object representation is.A 2-field object literal costs 72 bytes.
retainwrites 216 MB to store 48 MB of numbers — a 4.5x write amplification.The decisive evidence is that the remaining gap survives a hypothetically free collector:
retain_wide(206.2 − 75.4 = 130.8 ms of pure mutator) andshapes(60.2 ms of mutator) still lose to scriptc with a zero-cost collector.retainwould need exactly zero GC to reach parity. So further collector work cannot close this cluster.Why this is the representation problem, not a GC problem
scriptc's advantage here is not its refcounting — it is that it stores these objects more compactly. 72 bytes for two fields means header + slot overhead dominates payload. The write amplification is then paid twice: once in allocation bandwidth, and again in everything downstream that must touch those bytes (promotion, copying, cache pressure).
This is the concrete, measured instance of the direction recorded in the representation-selection RFC: type-driven representation selection rather than uniform NaN-boxed slots for every field.
Corollaries
shapesshould leave the "high-survival GC" cluster. At 7% GC pause it is a mutator benchmark that was mis-grouped; GC levers will not move it.Acceptance