Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 14 additions & 8 deletions ansible/roles/developer-rust/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,8 @@ silently stops being cached while `sccache` is still installed and configured.
`hyperi-rust-cache-prune` reports it rather than showing an empty ceiling, and
`sccache --stop-server` clears it.

## Bounding the build-artefact pool

**Build artefacts have no upstream cap at all.** Cargo does not track them
(rust-lang/cargo#13136), so nothing reclaims a `target/` ever, and one per repo
across a tree full of them is what actually fills a disk.
Expand All @@ -82,7 +84,7 @@ binaries, so `cargo run`, IDEs and anything globbing for a built artefact are
unaffected.

`hyperi-rust-cache-prune` then bounds that pool on a schedule -- a systemd timer
on Linux, a launchd agent on macOS, weekly and at idle IO priority. It drops
on Linux, a launchd agent on macOS, daily and at idle IO priority. It drops
workspaces not built for `rust_cache_max_age_days`, then evicts
least-recently-built ones until the pool is under `rust_cache_build_dir_max`.
It touches no project `target/`, and reports the self-capping caches without
Expand All @@ -94,14 +96,18 @@ derived from it chases itself downward. That puts a 692G build box at 115G and a
256G laptop at the floor, so one default suits both. Set
`rust_cache_build_dir_max` to an explicit size to override it.

**The ceiling is a weekly reset, not a live cap.** Nothing enforces it between
runs, so a busy build box spends most of the week above it -- one added 37G in
two days and another 37G in a single afternoon. That is the design working, not
a prune that failed. It only matters if the peak, not the floor, would fill the
disk; check free space before reaching for a shorter schedule, and change
`rust_cache_prune_schedule_weekday` if the peak is genuinely too high.
**The ceiling binds while the tool runs, not between runs.** A pool that grows
faster than the schedule spends the gap above it, so the prune runs daily and an
hourly guard backs it up -- one statvfs while the disk has room, a prune to the
same ceiling once free space falls below `rust_cache_prune_free_floor` (20%).

A guard run that finds the pool already under its ceiling stops and says so. The
space went somewhere the prune does not own, and naming that is more use than
evicting artefacts that were not the cause. Set
`rust_cache_prune_guard_enabled: false` to drop the guard, or
`rust_cache_prune_schedule_weekday` to go back to weekly.

sccache and ccache keep fixed ceilings; both enforce their own.
## Which sccache builds actually use

**sccache comes from upstream's release, not the distro.** Ubuntu ships 0.13.0
against an upstream on 0.17.x, and a wrapper four versions behind is what a
Expand Down
19 changes: 18 additions & 1 deletion ansible/roles/developer-rust/defaults/main.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,12 +37,29 @@ rust_cache_build_dir_max: "auto"
rust_cache_max_age_days: 14

# Off-hours, so a prune is unlikely to land mid-build.
rust_cache_prune_schedule_weekday: "Sun"
#
# Daily rather than weekly: the ceiling binds only while the tool runs, so the
# gap between runs is time the pool spends unbounded. Set a weekday for weekly.
rust_cache_prune_schedule_weekday: ""
rust_cache_prune_schedule_hour: 3

# Backstop for growth that outruns the daily prune. One statvfs an hour, and a
# prune only below the floor.
#
# A percentage because one size cannot suit a laptop and a build box.
#
# The guard prunes to the ordinary ceiling and no further. A pool already inside
# the ceiling means the space went elsewhere, and the run reports that.
rust_cache_prune_guard_enabled: true
rust_cache_prune_free_floor: "20%"
rust_cache_prune_guard_schedule: "hourly"
rust_cache_prune_guard_interval_seconds: 3600

# macOS launchd only; systemd takes its unit name from the file.
rust_cache_prune_label: "io.hyperi.rust-cache-prune"
rust_cache_prune_plist: "{{ user_home }}/Library/LaunchAgents/{{ rust_cache_prune_label }}.plist"
rust_cache_prune_guard_label: "io.hyperi.rust-cache-prune-guard"
rust_cache_prune_guard_plist: "{{ user_home }}/Library/LaunchAgents/{{ rust_cache_prune_guard_label }}.plist"

# Start the sccache server from a systemd user unit rather than leaving it to
# whichever process compiles first, which decides the server's cgroup and
Expand Down
126 changes: 115 additions & 11 deletions ansible/roles/developer-rust/files/hyperi-rust-cache-prune
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Run it yourself:
hyperi-rust-cache-prune # report, then ask before deleting
hyperi-rust-cache-prune --check # report only, delete nothing
hyperi-rust-cache-prune --yes # no prompt (how the timer invokes it)
hyperi-rust-cache-prune --if-free-below 20% # only when the disk is tight

Why this exists
---------------
Expand Down Expand Up @@ -33,6 +34,25 @@ a workspace you compile against daily and one you only link against look the
same. Age is measured from the newest thing cargo wrote, which is the closest
honest proxy.

When it runs at all
-------------------
The ceiling is enforced when the tool runs and at no other moment, so a pool
that grows faster than the schedule spends the gap unbounded. Shortening the
schedule alone trades that for a walk of the whole pool every time.

`--if-free-below` is the answer to that. It costs one statvfs, so the tool can
be scheduled often enough to matter and still do nothing almost every time: if
the filesystem holding the pool has more free space than the floor, it exits
before walking anything.

The floor takes a percentage as well as a byte count, because the same number
cannot be right on a 256G laptop and a 692G build box.

A guarded run that finds the pool already inside its ceiling does not go
looking for more to delete. It says the pool is not where the space went and
stops. Reclaiming artefacts that were not the problem would be the wrong
answer to a disk filled by something else.

What it does NOT touch
----------------------
Project `target/` directories, anywhere. `build.build-dir` moves intermediates
Expand Down Expand Up @@ -73,6 +93,7 @@ AUTO_FLOOR = "40G"
POOL_DEPTH = 2

_SIZE = re.compile(r"^\s*(\d+(?:\.\d+)?)\s*([KMGT]?)i?B?\s*$", re.IGNORECASE)
_PERCENT = re.compile(r"^\s*(\d+(?:\.\d+)?)\s*%\s*$")
_MULTIPLIER = {"": 1, "K": 1024, "M": 1024**2, "G": 1024**3, "T": 1024**4}


Expand Down Expand Up @@ -122,6 +143,49 @@ def parse_size_or_auto(text: str):
return parse_size(text)


def parse_size_or_percent(text: str):
"""A byte count, or a share of the filesystem written as a percentage.

Returned as a (kind, value) pair rather than a number because a percentage
cannot be resolved until the pool path names a filesystem to take it of.
"""
match = _PERCENT.match(text)
if match:
share = float(match.group(1))
if not 0 < share < 100:
raise argparse.ArgumentTypeError(
f"not a usable percentage: {text!r} -- give something above 0 and below 100"
)
return ("percent", share)
return ("bytes", parse_size(text))


def above_free_floor(pool: Path, floor, rep: Reporter) -> bool:
"""Has the filesystem holding the pool got more free space than the floor?

One statvfs, called before anything walks the pool, so a guarded run that
has nothing to do costs nothing. `free`, not `total - used`: reserved
blocks are not ours to spend and counting them would let the guard sit
quiet while writes are already failing.
"""
kind, value = floor
try:
usage = shutil.disk_usage(pool)
except OSError as exc:
# Unmeasurable is not the same as fine. Skipping the prune here would
# turn a broken statvfs into a silently unbounded cache.
rep.warn(f"could not measure free space at {pool} ({exc}) -- pruning anyway")
return False

threshold = int(usage.total * value / 100) if kind == "percent" else value
where = f"{human(usage.free)} free of {human(usage.total)}"
if usage.free > threshold:
rep.info(f"{where}, above the {human(threshold)} floor -- nothing to do")
return True
rep.info(f"{where}, below the {human(threshold)} floor")
return False


def resolve_max_size(value, pool: Path, rep: Reporter) -> int:
"""Turn `auto` into bytes for the filesystem that actually holds the pool."""
if value != "auto":
Expand Down Expand Up @@ -277,27 +341,38 @@ def find_workspaces(pool: Path) -> list[Workspace]:
return [Workspace(leaf) for leaf in leaves]


def remove(workspace: Workspace, reason: str, dry_run: bool, rep: Reporter) -> None:
def remove(workspace: Workspace, reason: str, dry_run: bool, rep: Reporter) -> bool:
"""Drop one workspace. False means the bytes are still on disk.

The caller has to know, because counting a failed removal against the pool
total reports a cache back under its ceiling while it is still over.
"""
if dry_run:
rep.info(f"[check] would drop {workspace.path.name} ({human(workspace.size)}, {reason})")
rep.freed += workspace.size
return
return True
try:
shutil.rmtree(workspace.path)
except OSError as exc:
rep.warn(f"could not remove {workspace.path}: {exc}")
return
return False
rep.change(f"dropped {workspace.path.name} ({human(workspace.size)}, {reason})")
rep.freed += workspace.size
return True


def prune_pool(
pool: Path, dry_run: bool, rep: Reporter, *, max_size: int, max_age_days: int
) -> None:
) -> int:
"""Bound the pool, and return what it still occupies.

Callers want that rather than the freed count, which reads the same whether
there was nothing to free or every eviction failed.
"""
workspaces = find_workspaces(pool)
if not workspaces:
rep.info(f"{pool}: empty, nothing to prune")
return
return 0

total = sum(item.size for item in workspaces)
rep.info(
Expand All @@ -306,21 +381,21 @@ def prune_pool(

stale = [item for item in workspaces if item.age_days > max_age_days]
for item in stale:
remove(item, f"idle {item.age_days:.0f}d", dry_run, rep)
if remove(item, f"idle {item.age_days:.0f}d", dry_run, rep):
total -= item.size

remaining = [item for item in workspaces if item not in stale]
total -= sum(item.size for item in stale)

if total <= max_size:
rep.info(f"under the ceiling at {human(total)}")
return
return total

# Oldest first: the least-recently-built workspace is the cheapest to lose.
for item in sorted(remaining, key=lambda entry: entry.mtime):
if total <= max_size:
break
remove(item, f"over ceiling, {human(total)} used", dry_run, rep)
total -= item.size
if remove(item, f"over ceiling, {human(total)} used", dry_run, rep):
total -= item.size

if total > max_size:
rep.warn(f"still {human(total)} after pruning everything eligible")
Expand All @@ -330,6 +405,8 @@ def prune_pool(
if not dry_run:
drop_empty_shards(pool)

return total


def drop_empty_shards(pool: Path) -> None:
"""Remove shard directories emptied by eviction, so the pool does not silt up."""
Expand Down Expand Up @@ -433,6 +510,16 @@ def main() -> int:
metavar="DAYS",
help=f"drop workspaces not built in this long (default: {DEFAULT_MAX_AGE_DAYS})",
)
parser.add_argument(
"--if-free-below",
type=parse_size_or_percent,
default=None,
metavar="SIZE|PERCENT",
help=(
"do nothing unless the filesystem holding the pool has less than this "
"free (e.g. 20%% or 80G) -- how the hourly guard invokes it"
),
)
parser.add_argument(
"--pool",
type=Path,
Expand Down Expand Up @@ -470,6 +557,12 @@ def main() -> int:

print(f"hyperi-rust-cache-prune: {pool}")

# Before the walk and before the prompt: a guarded run with room to spare
# must cost one statvfs and must never ask the user anything.
guarded = args.if_free_below is not None
if guarded and above_free_floor(pool, args.if_free_below, rep):
return 0

max_size = resolve_max_size(args.max_size, pool, rep)

if not args.check and not args.yes:
Expand All @@ -485,14 +578,25 @@ def main() -> int:
return 0

rep.step("Pooled build artefacts")
prune_pool(
remaining = prune_pool(
pool,
args.check,
rep,
max_size=max_size,
max_age_days=args.max_age_days,
)

# A guarded run ending inside the ceiling means the pool is not what filled
# the disk. Tested on what the pool still holds, not on what was freed -- a
# failed eviction also frees nothing.
if guarded and remaining <= max_size:
rep.warn(
"below the free floor with the pool already inside its ceiling -- "
"nothing here to reclaim. Whatever filled the disk is outside what "
"this tool bounds, and the caches reported below are the only other "
"ones it can see."
)

rep.step("Self-capping caches (reported, not pruned)")
try:
report_self_capping_caches(rep)
Expand Down
Loading
Loading