Skip to content

perf(gc): per-slot layout bookkeeping is the top cost family on json_pipeline — declare-at-birth for the JSON materialiser's cohort #7630

Description

@proggeramlug

Step-zero profile of json_pipeline 200k on current main (v0.5.1360-era), pinned mini (quiet, best-of-5, hash-stable): total ~1.3 s, build_out 877 ms of which 816 ms is GC pause (census: 3 cycles — 1 minor promoting 101 MB with 0.0 MB survivor-copied, 2 fulls). The largest cost family in the symbolicated 3 s sample is now the per-slot layout machinery, ahead of the old-page/promote family that perf/old-page-registration-deferral is already extracting:

17  layout_note_slot                          (store-side, per field store)
10  layout_slot_visit::visit_gc_layout_slot_descriptors   (trace-side)
 7  layout_slot_visit::fixed_slot
 7  layout_only
 6  layout_transfer                           (per object, during promote/copy)
 5  layout_tables::layout_forget_object
——
52 samples — vs 36 for the old-page family, 15 for remembered-set inserts

This is the trace/promote/store half of the layout story — #7510/#7525/#7532 addressed the construction half (side-table emptiness, shape-declared-at-allocation for class ctors and literals). The payer here is the JSON parse materialiser's cohort: 200k records × ~13 slots (record + addr + tags) ⇒ millions of runtime_store_jsvalue_slot → layout_note_slot calls at materialisation, then per-object layout_transfer for 101 MB of promotions, then descriptor-driven slot visits at trace time.

The likely lever, mirroring the proven #7525/#7532 pattern: the materialiser knows the complete object layout at materialisation time (it builds each object from the tape in one pass, and every slot it writes is a NaN-boxed JSValue — no raw-f64 slots). Declaring GC_LAYOUT_ALL_POINTERS at birth (header bits only, no side table) and routing its stores through a note-free path would eliminate the per-store notes and collapse the trace visits to the all-pointers fast path — a pointer mask only ever pays for itself by skipping raw-f64 slots, which this cohort does not have.

Soundness note for whoever takes this: an elided store into a POINTER_FREE-published object is a stranded live child (the #7469 lesson) — the elision must be gated on the live header state exactly as expr/array_push.rs's admission test is, or confined to a path where the materialiser owns the object end-to-end before publication.

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Aug 8, 2026
  2. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Closing the census data point from the #7633 merge comment: the 101 MB vs your 113,227,216 B is fixture provenance, not binary provenance. My step-zero binary was built at origin/main (38ff7eccc) — but the 200k fixture was regenerated from a re-created generator after the original was lost with a deleted worktree (benchmarks/honest_bench/assets/ is not tracked). Same record count and field names, different string contents → different promoted-bytes total. The "v0.5.1360-era" wording was sloppy; the commit is the fact. Within-session A/Bs are unaffected (both arms always ran the same fixture), but my absolute census numbers are not comparable to yours — yours, on the canonical asset, are the ones to cite.

    On the attribution ordering (layout family ahead of old-page family): both were measured on the same run so the ordering was real for that binary, but with #7624 merged the question is moot — the old-page family is gone, and the post-#7633 profile should be re-taken on current main before anything new is ranked.

  3. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Fixture provenance settled — and it is bigger than "different string contents"

    Re-measured on current main (v0.5.1363, 88e0812a7) on the pinned mini, against the canonical asset (benchmarks/honest_bench/assets/input.json, md5 ab95265ea2a53186026123fe9d485b8d, and its 200k prefix — the fixture ab_mini.sh picks up as w7592/fix/input_200k.json):

    fixture input bytes @200k promoted bytes promoted objects
    this issue's step-zero (~/stepzero/jp200k.json) re-created generator 41,669,851 105,888,096 ("101 MB") 1,467,591
    maintainer's #7624 audit canonical 44,745,110 113,227,216 1,657,966
    my re-measure, current main canonical 44,745,110 113,226,592 1,657,965

    So the 113 MB figure is the right one to cite — I reproduce it to within one object / 624 bytes (an output-path-length-dependent allocation), three independent measurements now agreeing.

    The correction worth recording is that the two fixtures are not the same workload with different strings. #7630's generator emits

    {"id":0,"name":"user_0_0","email":"user0@example0.com","age":18,"country":"US",
     "tags":[],"score":0.0,"active":true,"addr":{...,"zip":"10000"}}
    

    versus the canonical

    {"id":0,"name":"Frank Garcia","email":"frank.garcia0@example.com","age":42,"country":"CA",
     "tags":["user","pro","basic"],"score":68831,"active":true,"addr":{...,"zip":22626}}
    

    tags is variable length including empty rather than always 3; score is a float rather than an integer; zip is a string rather than an integer. Same record count, 12.9% fewer objects (1.47M vs 1.66M) and a different slot-type mix per record.

    That matters specifically for this issue's conclusion: the ranking it published put the per-slot layout family first (52 samples) and the old-page family second (36). Both of those families are priced per slot and per pointer-bearing slot, so a cohort with ~1 fewer heap object per record and integers where the canonical has strings is not a proxy for the canonical one. The census numbers were correctly flagged as non-comparable; the attribution ordering was measured on the same non-canonical workload and should carry the same caveat.

    Separately, the step-zero sample window was sleep 0.5; sample $PID 3 on a ~1.3 s run — from the sample header, Launch Time 00:45:15.716 / Date/Time 00:45:16.269, 795 main-thread samples. That window starts inside JSON.parse and runs to process exit, so it covers parse + build_out + stringify + write + fnv1a, not build_out. fnv1a32 at 35 samples and write_escaped_string at 17 in that leaf list are the tell.

    None of this changes that #7633 was a real win (it was an interleaved A/B, both arms on the same fixture). It only means the ordering it left behind cannot be inherited. Fresh post-#7624/#7633 decomposition on the canonical fixture is on #7592.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions