Skip to content

perf: object representation is now the binding constraint on the retain cluster (72 bytes per 2-field literal; 216 MB written to store 48 MB) #7916

Description

@proggeramlug

Summary

Established while decomposing the retain cluster's GC cost (gc-handoff/RETAIN4-NOTES.md): the collector is no longer the binding constraint on these benchmarks — object representation is.

A 2-field object literal costs 72 bytes. retain writes 216 MB to store 48 MB of numbers — a 4.5x write amplification.

The decisive evidence is that the remaining gap survives a hypothetically free collector:

program wall GC pause pause frac scriptc
retain 159.5 ms 52.0 33% 109
retain1 71.4 38.7 54% 38
retain_wide 206.2 75.4 37% 122
retain_wide1 73.9 32.3 44% 43
shapes 64.8 4.6 7% 38

retain_wide (206.2 − 75.4 = 130.8 ms of pure mutator) and shapes (60.2 ms of mutator) still lose to scriptc with a zero-cost collector. retain would need exactly zero GC to reach parity. So further collector work cannot close this cluster.

Why this is the representation problem, not a GC problem

scriptc's advantage here is not its refcounting — it is that it stores these objects more compactly. 72 bytes for two fields means header + slot overhead dominates payload. The write amplification is then paid twice: once in allocation bandwidth, and again in everything downstream that must touch those bytes (promotion, copying, cache pressure).

This is the concrete, measured instance of the direction recorded in the representation-selection RFC: type-driven representation selection rather than uniform NaN-boxed slots for every field.

Corollaries

  • shapes should leave the "high-survival GC" cluster. At 7% GC pause it is a mutator benchmark that was mis-grouped; GC levers will not move it.
  • Parallel GC is not indicated for this cluster (independently confirmed): the per-object visits split roughly 50/50 into work already deleted by perf(gc): describe a promoted block's page-object list instead of storing it #7914 and ~9.6 ms of a 159 ms program. There is not enough collector time left to parallelise, which matches the standing directive that parallelism is a last resort.

Acceptance

  • A measured proposal for compacting at least the common shapes here (small fixed-field object literals with numeric fields).
  • Report bytes-per-object and total write volume alongside any speedup — those are the primary metrics for this issue, not wall clock.

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Aug 12, 2026
  2. proggeramlug commented on Aug 12, 2026

    @proggeramlug
    ContributorAuthor

    The ≥4-field header shrink, priced before building it — and the answer is "yes, but not before the hidden-class work"

    Agent hdr, worktree wt-hdr, base a769fafc6 (0.5.1495). Full write-up and
    harnesses: gc-handoff/HDR-NOTES.md, m0810/build_hdr.sh, m0810/hdr_metrics.py,
    m0810/hdr_pad_patch.sh.

    1. First, a correction to the target numbers

    Removing the three derivable header members (object_type, field_count,
    keys_array) is 16 B, which lands at 40 B for {a,b} and 88 B for
    retain_wide — not the 32 / 80 quoted from REPR-NOTES §4. Those are the
    endgame: they additionally require deleting meta (#6759 Phase B put it in the
    header to remove a side-table probe) and merging class_id with
    parent_class_id. Also, only removing both u32s saves anything: dropping
    one leaves 12 bytes that re-pad to 16 for the following pointer.

    2. Measured, not extrapolated: a padding probe measures the shrink's mirror

    The removal moves every codegen offset across ~40 files, so before building it I
    priced it with its exact mirror — grow ObjectHeader by 16 B and read
    pad16 → base as the shrink (72→56 on retain, 120→104 on retain_wide), at the
    same absolute sizes, with #7961 in place. 11/11 byte-exact, exit 0 on both arms.

    bench Δ instr Δ RSS minors pad→base objects/minor ratio
    retain −25.63 % −26.19 % 6 → 4 1.084
    retain1 +3.64 % −18.72 % 3 → 2 0.992
    retain_wide +0.43 % −10.52 % 8 → 6 1.133
    retain_wide1 −1.30 % −13.42 % 4 → 3 0.945
    deeplist +6.25 % −13.77 % 3 → 2 0.877
    tree −0.06 % −18.64 % 40 → 40 —
    tree_wide +0.02 % −10.66 % 40 → 40 —
    churn / churn_alloc −0.6 % −0.3 % 105 → 88 1.010
    shapes / pipeline ≤ +0.4 % ≤ −0.2 % 1→1 / 8→7 —

    Best-of-3; the base arm's own spread was up to 3.5 % (dev box at load 200), so
    |Δinstr| ≲ 3 % is noise.

    Cycle counts go DOWN, not up. Every program's minor count falls or is
    unchanged — smaller objects allocate fewer total bytes, so a byte-denominated
    trigger fires less often. retain's −25.6 % is that: 2.93 M promotions → 2.11 M
    inside the exit-bounded window.

    3. ★ #7961's compensation is clamped OFF for exactly these objects — and the tax is still small

    nursery_cap_object_scale_permille is clamp(mean_surviving_object_bytes / 72, 500, 1000). Measured [gc-tenuring] factors: retain/retain1/deeplist 56 B →
    777‰; retain_wide/retain_wide1 104 B → 1000‰ (clamped); tree_wide
    32 B → 500‰ (floored). A 16 B shrink moves neither the clamp nor the floor, so
    a ≥4-field object gets no pacing compensation at all.

    The objects/minor column is that prediction landing:

    But it is 13 % more work spread over 25 % fewer cycles, and it costs +0.43 %
    instructions.
    So the #7929 tax is real, is exactly where the clamp says it is,
    and is not a reason to hold the shrink. I predicted it would be, from the
    clamp alone, before running the probe; the probe says otherwise.

    4. ★ The real blocker: the three fields ARE the inline-cache guard

    REPR-NOTES §4 prices this as "every codegen offset has to move in lockstep".
    That understates it. perry-codegen/src/expr/class_field_inline_guard.rs — the
    precheck admitting the proven-receiver clone, the hottest property path in the
    compiler — emits in ONE basic block off the receiver:

    object_type  @0  == OBJECT_TYPE_REGULAR
    class_id     @4  == expected_class_id
    field_count  @12 >  max_field_index
    keys_array   @16 == @perry_class_keys_*      <- "the load-bearing one"
    

    Three of those four are the fields being deleted. And keys_array identity
    cannot become a ShapeId compare as things stand: delete inst.b preserves
    class_id while compacting inline slots and installing a freshly cloned keys
    array (object/delete_rest.rs), and the check is deliberately dynamic because the
    delete shape barrier is module-scoped while receivers alias across modules
    (#7143). A ShapeId compare is equivalent only once deletion mints a new
    ShapeId
    — i.e. once shape transitions are real.

    So the 16 B is not a second INLINE_SLOT_FLOOR. It is not separable from the
    hidden-class transition machinery.

    5. Recommendation

    The prize is real and it is footprint: −10.5 % … −26.2 % peak RSS, {a,b} at
    2.50× payload and retain_wide at 1.38×, retain writing 120 MB instead of
    168 MB — at zero measurable instruction cost on 9 of 11 programs.

    Do it inside the hidden-class project, not before it. There the same change
    removes three guard loads instead of trading them for dependent probes.

    Two follow-ups worth filing separately:

    • fix(gc): denominate the nursery constant band in objects (#7929) #7961's denominator for the tails. The clamp exists because the estimator is
      a byte-weighted mean (push_num's is 3600 B). Pacing on the object count the
      last minor moved makes both the 1000‰ clamp and the 500‰ floor unnecessary and
      takes retain_wide's 1.133 to 1.000. A second-order win (~13 % of one term), not
      a prerequisite.
    • deeplist +6.3 % is the one program where a smaller object costs
      instructions while doing less GC (3→2 minors, 961 k→562 k promotions). Not
      explained; worth a look before the real shrink lands.
  3. added 4 commits that reference this issue on Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions