Repository navigation
perf(object): remove the derivable object_type and field_count header words — 56 B -> 48 B #8113
Description
Activity
- added 15 commits that reference this issue
on Aug 15, 2026 13 remaining items
- added 7 commits that reference this issue
on Aug 16, 2026 Done — landed as #8204 (
bf8fd868e), and it is a win on both axes.The shrink alone was a net instruction loss (0 of 19 rows faster,
deeplist+9.03%). What made it landable was the recovery work, and two thirds of the regression turned out not to be the shrink's semantics at all:push_cls/churn_alloc/churn(+5.5%, zero RSS) — an LLVM store-merging artefact. Withobject_typegone the two header words no longer folded into one 16-byte constant store, so the 40-bitGcHeaderconstant was rematerialised (mov+movk+movk) at every inlinenew, ~+4.5 instructions per allocation. IR identical modulo offsets; the machine code was not. Composing the header prefix once per function as a<2 x i64>took those rows to +0.1–0.2%.deeplist/retain1/retain(the "SIZE" rows) — GC pacing, not probes. The first copying minor fired on the 16 MB byte cap before any census existed, so the shrunk arm traced 17% more objects in the one traced cycle. A one-time header-walk census before minor #0 plus one descriptor lookup per traced object tookdeeplistto −17.8% instructions / −17.8% RSS versus main.interp/iso_miss/pipelinewere a third thing again: perf(codegen): guarded ordinary-parameter specialization #8094 landed a hot new consumer of both removed words after this branch's base, andpipeline's residual was LTO foldingshape_install_shared+recordintoinit_typed_shape_layout(474 → 811 instructions on the memo-hit path) — fixed with a#[cold] #[inline(never)]split acting as an inlining-stability device.
Final measured result, corroborated independently on a second arm set: corpus SUM −0.17% instructions / −7.28% peak RSS,
deeplist−17.8%/−17.8%,tree−12.8% RSS, every stdout byte-equal, all 19 binariescmp-verified different.Two things worth carrying forward from this:
A latent memory-safety bug on
mainthat the layout change turned fatal.fs::extract_string_ptraccepted any non-finite NaN-box with a plausible payload — noSTRING_TAGtest — somkdir_mode_from_options'sstring_value(options)read aStringHeaderoff the options object. Onmainthat misreadbyte_lenfromObjectHeader::class_id(a small number: a harmless one-byte garbage string thatparse_mode_stringrejected). Under the 48-byte layout the same read lands on the ShapeId (0x8000_0000+),from_utf8_lossywalks 2 GB, and everyfs.mkdirSync(dir, { recursive: true })segfaulted. Fixed at the source — the pointer read is now preceded by the tag that says what it points at. Found by the gap suite, not by any GC instrument.The pacing half is representation-independent.
PERRY_GC_SCAVENGE_NURSERY_MB=12on stock main binaries already boughtdeeplist−11.2% / −7.1% with no code change, because main's two-field literal is 56 B against the 72 BNURSERY_CAP_REFERENCE_OBJECT_BYTESanchor — main was running its first cycle ~29% oversized by its own calibration. So the census + threshold move would have been worth landing even without the shrink.Two follow-ons deliberately not folded in: the untraced-promotion predicate compares a ratio whose denominator the policy itself moves, and
maintoday sits ~2‰ from a cliff worth +20–27% instructions (re-denominating in absolute implied-dead bytes decouples it permanently); and the ratchet could record each probe's survival reading and fail when a gating probe sits within a few ‰ of a threshold — a margin assertion, in the spirit of "a gate must assert its subject".#8047's 56→40 rung remains blocked on #8112, and its headline prize (
retain−25.63%/−26.19%) was found to be a padding mirror-probe that does not reproduce — re-derive it before scheduling.- added 4 commits that reference this issue
on Aug 30, 2026
Summary
Remove
ObjectHeader::object_typeandfield_count, shrinking a common two-slot object from 56 B to 48 B and the eight-slot case from 104 B to 96 B. This is an independent rung of #8047 that needs none of the GC descriptor-rooting work in #8112.Why this is a separate, landable rung
Measured with
rustc -Oon the exact#[repr(C)]shapes (LP64):size_of::<ObjectHeader>(){a,b}object_typeonlyfield_countonlykeys_arraytoo (#8047)Neither field alone buys anything — the struct re-pads. Together they are a clean 8 bytes, exactly half of #8047's measured prize, and they do not touch the keys edge that #8112 has to redesign.
Work
1. Replace the seven raw-offset-0 Error discriminators
ObjectHeader.object_typeis prefix-punned againsterror::ErrorHeader, whose own firstu32is alsoobject_type. These read raw offset 0 on an untyped pointer:error.rs:750(js_error_is_error),error.rs:1542(gates(*error).errors)exception.rs:452,exception.rs:492(print_uncaught)value/dynamic_object.rs:531,548and:728,731Delete the word and offset 0 becomes
class_id; an ordinary object whoseclass_idequalsOBJECT_TYPE_ERRORis then read atErrorHeader's offsets.ErrorHeaderis already allocatedGC_TYPE_ERROR(error.rs:198-203), soGcHeader.obj_typeis the replacement — the same move #8086 made forobject_is_regular.2.
proxy.rs:1523needs the descriptor kind, notobject_is_regulargates
plan_eligible. It deliberately excludesOBJECT_TYPE_CLASS, and the comment above it records what breaks otherwise (#6595 — bundled zod'sZodX.createvanishing from ClassRef static dispatch).object_is_regularis true for class objects, so it is not a valid substitution. Useobject_kind == ShapeObjectKind::Ordinary.3. Make the live-slot publication atomic
object/mod.rs:1909 set_object_live_slot_countclears the stamp, writes the header word, then re-mints. Consumers falling back to the already-widened header word is what makes that window safe today, andsynchronize_object_shape_descriptor_fromallocates (HashMap insert), so a collection can land inside it. With no header word, a collection there sees the old live bound and the newly exposed slot is invisible to tracing and rewriting — a fresh #7154/#7164. Restructure to mint-then-stamp.4. Give the two capacity consumers a real fact
field_get_set/field_ops.rs:127computesalloc_limit = max((*obj).field_count, INLINE_SLOT_FLOOR)for its OOB bound and says so ("a generous limit … to avoid false positives").dyn_eval/env.rsuses the same idiom. Both want physical capacity, whichfield_countonly approximates.5. Codegen
lower_call/new_alloc.rs— the inlinenewwritesobject_typeatraw+8andfield_countpacked atraw+16; collapse to oneclass_id ‖ shape_idstore. Also fix the three stale comments there (:250-254says the header is 24 bytes withoutmeta;:606says slots start atraw+32when it israw+40;:380-383says the ILP32 header is 20 when it is 24).class_id @+4/ ShapeId@+8to+0/+4:class_field_inline_guard.rs:279,283,400,495,499;element_shape_guard.rs:359;generic_dispatch.rs:389;proxy_reflect.rs:550,557,867,872. GcHeader-relative offsets (-8/-7/-6) are unaffected.object_header_size_bytesheader-skip geps need no edit.expr/proxy_reflect.rs:922andstmt/loops.rs:3207divide the header size by 8 for a word index.24 / 8is exact; the new ILP32 header is 16 (see below) so16 / 8stays exact — but convert them to byte geps anyway, since perf(object/GC): finish the common-object header shrink from 56 B to 40 B after shape-transition migration #8047 makes the value 12 and12 / 8 == 1silently. Note a{u32, u32, *mut}ILP32 header is 12 bytes with align 4, which would put 8-byte JSValue slots at a 4-aligned offset and violate thei64:64arm64_32 ABInew_alloc.rs:588-591warns about; this rung keepskeys_array, so ILP32 stays 16 and the hazard is deferred to perf(object/GC): finish the common-object header shrink from 56 B to 40 B after shape-transition migration #8047.6. FFI and out-of-runtime
perry-ffi/src/types.rs:37-53mirror + theoffset_of!asserts at:149-175.perry-ext-ws/src/lib.rs:847(let n = (*ptr).field_count;) — needs a C accessor orjs_object_keys+js_array_length.perry-stdlib/src/worker_threads.rs:677-683— structured-clone walk.worker_options.rs:114already models the replacement, but thekeys_array.is_null()branch is deliberate (class instances take thefield_countarm); verifyjs_object_keysreturns the same set forclass_id != 0before collapsing, since it filters private#xkeys.perry-ui-android/src/json.rs:487,490,496,598— likely just delete the module: every function is private with no callers and its own trailing comment saysjs_json_*now lives inperry-runtime/json.rs.7. Gates
scripts/shape_descriptor_census.py'sFIELDStuple, self-test fixture, sabotage fixture, and the baseline JSON. (ci(object): make the ObjectHeader shape-descriptor census a real gate #8110 wires this script intolint; land that first.)scripts/raw_handle_debt_files.txt:115andscripts/addr_class_ratchet_baseline.txt:138,257— an entry matching nothing fails, andaddr_class_inventory.pyis a requiredlintstep.perry-ffi'sobject_header_matches_runtimehas never executed — it is#[cfg(all(test, feature = "runtime-link"))]andruntime-linkis enabled nowhere in.github/. Field deletion still goes red viarustc-warnings, but a size/padding divergence is invisible. Wire it up as part of this change.perry-ffiis published to crates.io — an out-of-tree wrapper on the old mirror linked against a new runtime readsclass_idout of the deletedobject_typeslot with no compile error. Needs a deliberate semver decision.docs/src/platforms/watchos.md:35,106-111is the only user-facing statement of the LP64/ILP32 pair. Four codegen doc comments already say "24 on 64-bit, 20 on ILP32" and have been wrong sincemetalanded (property_get.rs:1724,generic_dispatch.rs:435,property_set.rs:1443,lower_call/new.rs:693).Acceptance
size_of::<ObjectHeader>() == 24(LP64);{a,b}totals 48 B includingGcHeaderand two slots; the wide case reaches 96 B. Report both.GcHeaderstays 8 bytes.copied_objects > 0orpromoted_objects > 0, not merely that nothing threw.Refs #8047, #8067, #8086, #8110, #8112, #6595, #7154, #7164, #7916.