Skip to content

Prep on Apple silicon: Metal SIFT, MPS masks, VideoToolbox decode - #26

Merged
pgodlews merged 1 commit into
mainfrom
mac-prep
Oct 2, 2026
Merged

pgodlews merged 1 commit into
mainfrom
mac-prep

Conversation

@pgodlews

@pgodlews pgodlews commented Oct 2, 2026

Copy link
Copy Markdown
Owner

The stages before training (frames, select, mask, sfm) now run on an Apple silicon Mac, by hand and outside the queue. Training stays on CUDA.

  • scripts/sift_backend.py picks the SIFT backend (cuda, metal, cpu; SPLAT_SIFT overrides). The CUDA call is unchanged, so no cache key moves. On the Metal build extraction runs on the GPU and matching on the CPU.
  • scripts/setup_mac.sh builds pycolmap from pgodlews/colmap, branch metal-sift-lanxinger, pinned by commit, and a venv_gs with torch 2.14.1: torchvision 0.24.1's roi_align is unusable on MPS.
  • 81_fisheye_stitch.py builds its float64 grids on the CPU for MPS; 70_person_masks.py runs on MPS; frame decoding uses VideoToolbox on macOS (SPLAT_HWACCEL overrides). 30_run_sfm.py refuses to run without CUDA.
  • Docs: "Prep on Apple silicon" in how-it-works.md with the measurements (clips 0141 and 0005, trained against CUDA preps), troubleshooting #37 and #38, install and README pointers.

Measured: trained PSNR 27.109 vs 27.010 (0141) and 21.990 vs 22.027 (0005) against the CUDA prep of the same job. Not done: queue scheduling on a Mac, and writing a handoff bundle from stages run by hand.

The stages before training (frames, select, mask, sfm) now run on an Apple
silicon Mac, by hand and outside the queue. Training stays on CUDA.

- scripts/sift_backend.py picks the SIFT backend (cuda, metal, cpu;
  SPLAT_SIFT overrides). The CUDA call is unchanged, so no cache key moves.
  On the Metal build extraction runs on the GPU and matching on the CPU.
- scripts/setup_mac.sh builds pycolmap from pgodlews/colmap, branch
  metal-sift-lanxinger, pinned by commit, and a venv_gs with torch 2.14.1:
  torchvision 0.24.1's roi_align is unusable on MPS.
- 81_fisheye_stitch.py builds its float64 grids on the CPU for MPS;
  70_person_masks.py runs on MPS; frame decoding uses VideoToolbox on macOS
  (SPLAT_HWACCEL overrides). 30_run_sfm.py refuses to run without CUDA.
- Docs: "Prep on Apple silicon" in how-it-works.md with the measurements
  (clips 0141 and 0005, trained against CUDA preps), troubleshooting #37 and
  #38, install and README pointers.

Measured: trained PSNR 27.109 vs 27.010 (0141) and 21.990 vs 22.027 (0005)
against the CUDA prep of the same job. Not done: queue scheduling on a Mac,
and writing a handoff bundle from stages run by hand.
@pgodlews
pgodlews merged commit 7522c1c into main Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant