Skip to content

fix(developer-rust): bound the build pool against free space, not just the clock - #77

Merged
catinspace-au merged 1 commit into
mainfrom
fix/rust-cache-free-space-guard
Sep 2, 2026
Merged

catinspace-au merged 1 commit into
mainfrom
fix/rust-cache-free-space-guard

Conversation

@catinspace-au

Copy link
Copy Markdown
Contributor

Dragonfly hit 89% with 289G of pooled Rust build artefacts against a 115G ceiling. The prune was not broken - it ran on Aug 30, freed 175.3G and left the pool at 97.4G. It just runs weekly, and the pool grows about 55G a day, so it crossed its own ceiling eight hours later and had six days with nothing enforcing anything.

The README already called that out as deliberate, the escape hatch being "check free space before reaching for a shorter schedule". Nobody made that check. This automates it.

What changes:

  • --if-free-below SIZE|PERCENT on the prune tool. Above the floor it exits after one statvfs, before walking the pool and before prompting, so it is cheap enough to schedule hourly.
  • Full prune goes weekly to daily. Set rust_cache_prune_schedule_weekday to put it back.
  • An hourly guard runs the same prune to the same ceiling once free space drops under rust_cache_prune_free_floor (20%). systemd timer on Linux, launchd agent on StartInterval for macOS.
  • A guard run that ends inside the ceiling reports that the pool is not what filled the disk, rather than cutting below its own bound to look useful.

There is a real bug fixed in here too. remove() warned and returned on a failed rmtree while both callers subtracted the size regardless, so a pool whose evictions all failed reported under the ceiling at 0B with every byte still on disk - the exact number the guard's verdict reads. remove() now reports whether the bytes went and prune_pool returns what is left. Reproduced against the old code, covered by two regression tests.

The macOS load sequence was five tasks and a second agent would have duplicated all of it, so it moved to tasks/launchd_agent.yml, included once per agent to keep each one's register scope.

Verified on dragonfly: both timers live, guard reports 371.3G free of 691.5G, above the 138.3G floor and returns in under a second, daily next fires 03:00. The role runs twice in that playbook and the second pass is all ok. 65 tests pass, ansible-lint clean for this change.

Worth knowing before this goes near other boxes: the macOS half is rendered and plist-validated but not run on a Mac from here. And a box with the pool on a dedicated volume (desktop-derek uses /cache) will refuse to prune at all - main rejects a pool outside $XDG_CACHE_HOME/~/.cache, and a systemd unit never loads the profile that would set XDG_CACHE_HOME. Both units would exit 1 every run. Follow-up, not in here.

…t the clock

The pool ceiling only binds while the prune is running, and the prune ran
weekly. A box that grows the pool faster than that spends the gap over its
ceiling, which is how dragonfly reached 89% with 289G pooled against a 115G
ceiling. The README described the overshoot as deliberate and asked whoever
owned the box to watch free space themselves. Nobody did.

So automate that check rather than only shortening the schedule:

- hyperi-rust-cache-prune takes --if-free-below SIZE|PERCENT. Above the floor
  it exits after one statvfs, before walking the pool and before prompting, so
  it can be scheduled hourly and still cost nothing almost every time.
- The full prune moves from weekly to daily. Setting
  rust_cache_prune_schedule_weekday puts it back.
- An hourly guard runs the same prune to the same ceiling once free space drops
  below rust_cache_prune_free_floor, 20% by default. systemd timer on Linux, a
  launchd agent on StartInterval for macOS.
- A guard run that ends inside the ceiling says so instead of cutting deeper.
  The space went somewhere the prune does not own, and naming that beats
  evicting artefacts that were not the cause.

Also fixes the accounting the guard's verdict rests on. remove() warned and
returned on a failed rmtree while both callers subtracted the size anyway, so a
pool whose evictions all failed reported "under the ceiling at 0B" with every
byte still on disk. remove() now reports whether the bytes went, and prune_pool
returns what the pool still holds.

The macOS load sequence was five tasks and a second agent would have duplicated
all of it, so it moves to tasks/launchd_agent.yml, included once per agent to
keep each one's register scope.

Verified on dragonfly: both timers live, the guard reports 371.3G free of
691.5G against a 138.3G floor and returns in under a second, and the role's
second pass in the same run is all ok. 65 tests pass.
@catinspace-au
catinspace-au force-pushed the fix/rust-cache-free-space-guard branch from f552533 to aa1c2e7 Compare September 2, 2026 23:29
@catinspace-au
catinspace-au merged commit ee099d7 into main Sep 2, 2026
16 checks passed
@catinspace-au
catinspace-au deleted the fix/rust-cache-free-space-guard branch September 2, 2026 23:33
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant