Skip to content

Spot GPUs: checkpoint snapshots, restore points, reclaim notices and resume (#10) - #23

Merged
pgodlews merged 3 commits into
mainfrom
trainer/lfs-snapshot-patch
Sep 30, 2026
Merged

pgodlews merged 3 commits into
mainfrom
trainer/lfs-snapshot-patch

Conversation

@pgodlews

Copy link
Copy Markdown
Owner

Stream training checkpoints out and resume elsewhere, for spot and other interruptible GPUs. The investigation, measurements and design are in #10 (two comments).

Closes #10

LichtFeld patch (scripts/lichtfeld-patches/)

Applied by setup_lichtfeld.sh, which stops the build if it no longer applies. It adds options and changes no default, so TRAINER does not move. Intended for upstream later.

  • --save-steps a,b,...: project snapshot iterations, the same parser and step scaling as --eval-steps. Until now save_steps could only be set through --config, which replaces the whole optimisation parameter set.
  • SIGUSR1 in plain headless mode requests a snapshot without stopping training. The request runs on its own worker thread, so it never delays SIGTERM, and it is only wired during train().

Tested on an RTX 3090, sm_86 image built from this branch:

  • --save-steps works without --config.
  • A SIGUSR1 snapshot was prepared 0.15–0.3 s after the signal, early in training and at steady state.
  • SIGUSR1 followed by SIGTERM still exits in about 10 s.
  • The exports are written.
  • PSNR was 19.958 with snapshots and 19.952 without, against 19.952–19.959 before the patch.

Queue

  • train.checkpoint_every: a snapshot every N training steps (0 off, at least 500). It becomes --save-steps in LichtFeld's unscaled units. It is not a cache key term: the keys are identical to main's for existing configs, with or without it.
  • Restore points (checkpoints.py, scripts/licht_restore_point.py):
    • After each committed snapshot, LichtFeld's clean_project_file keeps only the newest checkpoint (1.56 GB to 467 MB measured) and it must verify.
    • It is packed with restore.json (config, keys, TRAINER, version terms, cache root, step, sha256) as job<id>-restore.tar.
    • The tar goes to CHECKPOINT_UPLOAD_URL, replacing the previous one. S3, R2 and GCS replace an object only on a complete upload. QUEUE_RESTORE_POINTS=1 keeps them locally instead.
    • A failure is recorded and training goes on.
  • Reclaim notices (preempt.py, QUEUE_PREEMPT_WATCH=aws|gcp):
    • On a reclaim notice, SIGUSR1 goes to each trainer that has reported itself up, and the queue is paused.
    • A rebalance recommendation snapshots only.
    • POST /api/snapshot does the same by hand.
  • Resume: RESUME_URL downloads into QUEUE_ROOT/resume/, and POST /api/resume/import imports with or without a handoff bundle. It is refused unless the trainer, version terms, cache root and keys match. The job's train stage runs --resume … --export=….
  • Telemetry: checkpoints, resumed and transfers.resume, with no URLs or paths. Documented in docs/job-telemetry.md, and a "Spot GPUs" section in docs/cloud.md.

Tests

queue/test_checkpoint.py: 26 CPU-only tests, added to CI and AGENTS.md. They use a fake trainer and a stub lichtfeld.io module. The log parser was also replayed against real patched-trainer logs: every snapshot was found at its real step.

Not yet run: licht_restore_point.py against the real lichtfeld.io module in an image (the lab box went down). It should be exercised on the next rc.

)

Both are needed to stream training checkpoints out for spot GPUs (#10).
save_steps could only be set through --config, which replaces the whole
optimisation parameter set; SIGUSR1 lets a spot-reclaim watcher ask for a
snapshot without stopping training. setup_lichtfeld.sh applies every patch
in scripts/lichtfeld-patches/ and stops if one no longer applies.
Measured on a 3090: a SIGUSR1 at step ~550 was served only at step 700,
while one at step ~1556 was served in 1.3 s. Whether the handler itself
blocked is not known, but if it did, the monitor thread could not notice a
SIGTERM meanwhile. The request now runs on its own worker thread and logs how
long queueing took; train() waits for it before returning.
…cher, resume (#10)

- train.checkpoint_every: a snapshot every N training steps, passed to the
  patched trainer as --save-steps in its unscaled units. Not a cache key
  term: keys are unchanged for every existing config, with or without it.
- Restore points (checkpoints.py, scripts/licht_restore_point.py): after each
  committed snapshot, LichtFeld's clean_project_file keeps only the newest
  checkpoint; packed with restore.json (config, keys, TRAINER, version terms,
  cache root, step, sha256) as job<id>-restore.tar and sent to
  CHECKPOINT_UPLOAD_URL, replacing the previous one. Failures are recorded
  and training goes on.
- QUEUE_PREEMPT_WATCH=aws|gcp (preempt.py): on a reclaim notice, SIGUSR1 to
  every trainer that has reported itself up, then pause the queue; a
  rebalance recommendation snapshots only. POST /api/snapshot by hand.
- Resume: RESUME_URL into QUEUE_ROOT/resume/, POST /api/resume/import with
  or without a handoff bundle. Refused unless trainer, version terms, cache
  root and keys match; the job's train stage runs LichtFeld --resume.
- Telemetry: checkpoints, resumed, transfers.resume (no URLs or paths).
- queue/test_checkpoint.py: 26 CPU-only tests, in CI.
@pgodlews
pgodlews merged commit 776e250 into main Sep 30, 2026
3 checks passed
@pgodlews
pgodlews deleted the trainer/lfs-snapshot-patch branch September 30, 2026 16:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Investigate: stream training checkpoints out and resume elsewhere (for interruptible/spot GPUs)

1 participant