Skip to content

Prep on Apple silicon, with the queue - #27

Merged
pgodlews merged 1 commit into
mainfrom
mac-prep
Oct 2, 2026
Merged

pgodlews merged 1 commit into
mainfrom
mac-prep

Conversation

@pgodlews

@pgodlews pgodlews commented Oct 2, 2026

Copy link
Copy Markdown
Owner

The stages before training (frames, select, mask, sfm) run on an Apple silicon Mac, and the queue runs there as a prep install that writes a handoff bundle for a CUDA box. Training stays on CUDA.

What changes

  • Scripts: sift_backend.py picks the SIFT backend (cuda, metal, cpu); the stitch and Mask R-CNN run on MPS; frames decode with VideoToolbox. 30_run_sfm.py (stitched input) refuses to run without CUDA.
  • Tools: scripts/setup_mac.sh builds pycolmap from pgodlews/colmap branch metal-sift-lanxinger, pinned by commit, and a venv_gs with torch 2.14.1.
  • Queue: PREP_BACKEND (cuda, or apple on macOS). On a Mac the queue schedules on one device without nvidia-smi, requires run_until, and refuses stitched input. queue/run_mac.sh starts it.
  • Cache keys: JobConfig.prep_backend, set by the host. Apple jobs carry PREP_APPLE in every key; CUDA keys are unchanged (the pins in test_stages.py pass untouched). A host refuses to build a prep stage of a job made for the other backend.

Measured

Trained on an RTX 3090 (30,000 iterations, 3M splats, 0.2.0-rc3 train image), one run per arm, against a CUDA prep of the same job:

Clip Mac prep, PSNR / SSIM / LPIPS CUDA prep
0141 (Avata 360, 473 rig frames), SfM from the Mac 27.109 / 0.8342 / 0.1240 27.010 / 0.8316 / 0.1265
0005 (Osmo 360, 957 rig frames), every stage from the Mac 21.990 / 0.6770 / 0.2155 22.027 / 0.6775 / 0.2156

Through the queue on an M5 Max, 0141 preps in 28.2 min and writes a 2.27 GB bundle, which a second queue instance set to cuda imported under the same keys.

Not done

  • The bundle made by the queue on the Mac was imported but not trained; the trained bundles above were made from the same stages run by hand.
  • Images built before this change refuse a Mac's bundle (unknown config field prep_backend), so training one needs a train image built from this.
  • No guard against other programs using the Mac's GPU, SAM 3 on MPS is untested, and nothing smaller than an M5 Max was measured.

Details and all figures: docs/how-it-works.md, "Prep on Apple silicon".

The queue now runs on a Mac as a prep install: a job with run_until "sfm"
goes through frames, select, mask and sfm and writes a handoff bundle.

- config.PREP_BACKEND ("cuda", or "apple" on macOS; QUEUE_PREP_BACKEND
  overrides, which the tests use to model a CUDA host wherever they run).
  On apple the queue schedules on one device without nvidia-smi and is a
  prep install (IMAGE_VARIANT defaults to "prep").
- JobConfig.prep_backend, set by the host that creates the job. An apple job
  carries PREP_APPLE in every cache key; cuda adds no term, so existing keys
  and the pins in test_stages.py are unchanged. The term travels in the
  bundle's config, so a CUDA box computes the same keys on import.
- A host refuses to build a prep stage of a job made for the other backend.
- Stitched input is refused on a Mac (30_run_sfm.py needs CUDA).
- queue/run_mac.sh: the launcher, deploy.sh without systemd.
- Telemetry fills the CPU model and memory size on macOS.
- queue/test_prep_backend.py, added to CI.

Checked on clip 0141 through the queue on an M5 Max: 28.2 min, 473/473 rig
frames, 510,745 points, 0.807 px, a 2.27 GB bundle; a second instance set to
cuda imported it under the same keys. That bundle was not trained. Images
without the field refuse a Mac's bundle as an unknown config field.
@pgodlews
pgodlews merged commit 2962264 into main Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant