Repository navigation
Process-global side tables are cleared by test guards that readers are not required to take — three flakes in two days #7672
Description
Activity
- added a commit that references this issue
on Aug 9, 2026 Design and fix in #7674, with the survey extended: 23 tables across every family the guards clear, converted to per-thread storage in test builds via
guard_cleared_global!.Option 1 as filed is not enough on its own — a reader-side assertion is still an opt-in, and the issue's own second requirement is that a future reader must not have to remember anything. So the storage moves instead: the accessor and every call site are unchanged (
PerThread<T>derefs toT), and outside a test build the macro expands to the plainstaticbyte for byte. A lock genuinely cannot close this — the damage window is between a test's write and its read, and only the test knows that span.Option 4 stays rejected for the reason you gave. What replaces the ~180-reader burden is a
lintgate (scripts/global_sink_isolation.py) that derives the clear list fromreset_copying_nursery_runtime_test_state's own source and fails on any barestaticleft behind — the burden lands on the ~20 table authors, mechanically.Shown able to fail: reverting the macro's
#[cfg(test)]arm to the pre-fix bare static — one edit, all 23 tables — fails 7 of the 9 new tests, each naming its table. A canary declared as a plainstaticruns the same probe procedure and requires the wipe to be observed, so a green file means the detector works.Fourth instance: found, three of them, all with the split-lock-domain signature and all outside the guards' clear list, so this PR's gate is blind to them — filed as #7680. The sharpest is
async_hooks'NEXT_ASYNC_IDunder four disjoint domains withresource_ids_are_monotonic_even_without_hooksassertingb.async_id == a.async_id + 1;tui::state::SLOTSunder three, withalloc_returns_sequential_handlesassertingh0==0, h1==1, h2==2.Evidence: 25 consecutive
cargo test -p perry-runtime --lib --no-fail-fastruns green with 0 vacuous (each checked forRunning unittestsanderror[before being counted),cargo check --all-targetsclean, all 24 lint gates green.- added 6 commits that reference this issue
on Aug 9, 2026 - added 6 commits that reference this issue
on Aug 9, 2026 - added 9 commits that reference this issue
on Sep 22, 2026
Three flaky tests in two days, all the same architecture: a process-global side table, plus per-suite guards that clear it, plus tests that read it without taking the clearing lock.
opt_report'sMutex<Vec<Entry>>rows.len() == 2failed at 3 — a neighbour's entry in the snapshotext_registry'sUSED_PROVIDERSioredisstill presentclosure/dynamic_props.rs'sCLOSURE_PROPSTAG_UNDEFINEDIn every case the lock existed and was correct for the tests that took it. The failure is that nothing requires a reader to take it, so the defence is opt-in and the opt-in is invisible at the read site.
Each was exposed by an unrelated PR adding tests and changing the parallel schedule — never introduced by it. That is the expensive part: the author of the exposing PR spends the diagnosis, and the natural conclusion ("my change broke something") is wrong.
The remaining exposure, surveyed
Files whose tests read closure dynamic props, and whether they take
crate::gc::global_side_table_test_lock():closure/dynamic_props.rsobject/global_this_webassembly.rsobject/native_module_stream.rs(after #7671)array/tests.rs(65),node_stream_tests.rs(42),node_submodules/tests.rs(33),object/instanceof.rs(5),object/native_module/constants.rs(4),value/to_string.rs(4),node_stream_state_tests.rs(4),tls.rs(3),object/native_module/callable_export_arity_table.rs(3),perf_hooks.rs(2),object/prototype_chain.rs(2),object/native_call_method.rs(2),value/dynamic_object.rs(2), and 6 more with 1 eachRoughly 180 unguarded tests. Most will never bite — the hazard needs a test that reads a persisted prop across a window a guard can land in — but "most" is doing real work in that sentence, and the three we found were each a surprise.
Options, in rough order of preference
test_clear_closure_side_tables()record that it ran, and give readers a cheap assertion that no clear happened during their read. Turns a silent wrong value into a named failure at the point of damage, without serialising anything.CLOSURE_PROPSis keyed by heap address with a GC scanner attached, so this needs care.Option 1 is the one that pays for itself even if the class never bites again, because it converts "intermittent, diagnosed by luck" into "fails with a message naming the clearer".
Context: #7665, #7671, and
gc/tests/support.rs's own comment, which documents the hazard but does not enforce anything.