Skip to content

Size the prep stages from the clip; 3 GB headroom on every estimate - #25

Merged
pgodlews merged 1 commit into
mainfrom
disk/prep-stage-estimates
Sep 30, 2026
Merged

pgodlews merged 1 commit into
mainfrom
disk/prep-stage-estimates

Conversation

@pgodlews

Copy link
Copy Markdown
Owner

Closes #24.

The prep stages were held to the flat 20 GB floor (QUEUE_MIN_FREE_GB), so a prep-only box on the split pipeline (2026-09-30: 21 GB disk, job peak 13.7 GB) had to be launched with the floor set by hand. #13 did this for train and export; this does it for the prep half.

  • frames, select, mask, sfm are sized from the job: the frames they handle (fps x the clip's span; then select's panoramas, or candidates / window before select has run) x the clip's pixels per frame (both lenses of a fisheye rig) x what each stage wrote per pixel on 0141 (frames 0.133, mask 0.248, sfm 0.118 B/px; select only hardlinks, 0.005), with a 1.5 margin since it is one clip.
  • The handoff bundle is checked before its tar is written: it is a copy of what it packs and came after the last stage's check (it left that box 6.5 GB free).
  • 3 GB headroom on every estimate: an estimated stage (train and export included) wants estimate + QUEUE_ESTIMATED_SLACK_GB (default 3, was a 2 GB minimum), or its history if larger.
  • A clip that does not probe keeps the 20 GB floor. QUEUE_MIN_FREE_GB set by hand still holds for every stage and turns the estimates off.
  • ffprobe moves from main.py to queue/app/probe.py so the worker can use it (the API behaves as before).
  • No cache key changes: no stage writes anything different.

Against the measured run: estimates for 0141 are frames 8.3 GB, mask 5.2, sfm 2.5, select 0.2 (measured 5.56 / 3.47 / 1.65 / ~0), each plus 3 GB; on that 21 GB box every stage and the tar pass without a floor set.

Tests: queue/test_stages.py checks the 0141 numbers, trims, select's record winning once it ran, the no-probe fallback and the headroom rule; the full list in AGENTS.md passes locally. Docs: queue/README.md "Disk retention", docs/troubleshooting.md #36.

frames, select, mask and sfm were held to the flat 20 GB floor
(QUEUE_MIN_FREE_GB), whatever they would write, so a prep-only box
(2026-09-30, 21 GB disk, job peak 13.7 GB) had to be launched with the
floor set by hand. They are now sized from the job, like train and
export: the frames they handle (fps x the clip's span, then select's
panoramas) times the clip's pixels per frame (both lenses of a fisheye
rig) times what each stage wrote per pixel on 0141 (frames 0.133, mask
0.248, sfm 0.118 B/px; select only links), with a 1.5 margin. The
handoff bundle is checked the same way before its tar is written; it
came after the last stage's check and left that box 6.5 GB free.

An estimated stage now wants its estimate plus QUEUE_ESTIMATED_SLACK_GB
(default 3, was a 2 GB minimum), or its history if larger. A clip that
does not probe keeps the 20 GB floor; QUEUE_MIN_FREE_GB set by hand
still holds for every stage and turns the estimates off.

ffprobe moves from main.py to probe.py so the worker can use it. No
cache keys change: no stage writes anything different.
@pgodlews
pgodlews merged commit 4a49497 into main Sep 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Disk preflight: size the prep stages from the clip instead of the flat 20 GB floor

1 participant