Repository navigation
Spot GPUs: checkpoint snapshots, restore points, reclaim notices and resume (#10) - #23
Merged
Merged
Conversation
) Both are needed to stream training checkpoints out for spot GPUs (#10). save_steps could only be set through --config, which replaces the whole optimisation parameter set; SIGUSR1 lets a spot-reclaim watcher ask for a snapshot without stopping training. setup_lichtfeld.sh applies every patch in scripts/lichtfeld-patches/ and stops if one no longer applies.
Measured on a 3090: a SIGUSR1 at step ~550 was served only at step 700, while one at step ~1556 was served in 1.3 s. Whether the handler itself blocked is not known, but if it did, the monitor thread could not notice a SIGTERM meanwhile. The request now runs on its own worker thread and logs how long queueing took; train() waits for it before returning.
…cher, resume (#10) - train.checkpoint_every: a snapshot every N training steps, passed to the patched trainer as --save-steps in its unscaled units. Not a cache key term: keys are unchanged for every existing config, with or without it. - Restore points (checkpoints.py, scripts/licht_restore_point.py): after each committed snapshot, LichtFeld's clean_project_file keeps only the newest checkpoint; packed with restore.json (config, keys, TRAINER, version terms, cache root, step, sha256) as job<id>-restore.tar and sent to CHECKPOINT_UPLOAD_URL, replacing the previous one. Failures are recorded and training goes on. - QUEUE_PREEMPT_WATCH=aws|gcp (preempt.py): on a reclaim notice, SIGUSR1 to every trainer that has reported itself up, then pause the queue; a rebalance recommendation snapshots only. POST /api/snapshot by hand. - Resume: RESUME_URL into QUEUE_ROOT/resume/, POST /api/resume/import with or without a handoff bundle. Refused unless trainer, version terms, cache root and keys match; the job's train stage runs LichtFeld --resume. - Telemetry: checkpoints, resumed, transfers.resume (no URLs or paths). - queue/test_checkpoint.py: 26 CPU-only tests, in CI.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stream training checkpoints out and resume elsewhere, for spot and other interruptible GPUs. The investigation, measurements and design are in #10 (two comments).
Closes #10
LichtFeld patch (
scripts/lichtfeld-patches/)Applied by
setup_lichtfeld.sh, which stops the build if it no longer applies. It adds options and changes no default, soTRAINERdoes not move. Intended for upstream later.--save-steps a,b,...: project snapshot iterations, the same parser and step scaling as--eval-steps. Until nowsave_stepscould only be set through--config, which replaces the whole optimisation parameter set.train().Tested on an RTX 3090, sm_86 image built from this branch:
--save-stepsworks without--config.Queue
train.checkpoint_every: a snapshot every N training steps (0 off, at least 500). It becomes--save-stepsin LichtFeld's unscaled units. It is not a cache key term: the keys are identical tomain's for existing configs, with or without it.checkpoints.py,scripts/licht_restore_point.py):clean_project_filekeeps only the newest checkpoint (1.56 GB to 467 MB measured) and it must verify.restore.json(config, keys, TRAINER, version terms, cache root, step, sha256) asjob<id>-restore.tar.CHECKPOINT_UPLOAD_URL, replacing the previous one. S3, R2 and GCS replace an object only on a complete upload.QUEUE_RESTORE_POINTS=1keeps them locally instead.preempt.py,QUEUE_PREEMPT_WATCH=aws|gcp):POST /api/snapshotdoes the same by hand.RESUME_URLdownloads intoQUEUE_ROOT/resume/, andPOST /api/resume/importimports with or without a handoff bundle. It is refused unless the trainer, version terms, cache root and keys match. The job's train stage runs--resume … --export=….checkpoints,resumedandtransfers.resume, with no URLs or paths. Documented indocs/job-telemetry.md, and a "Spot GPUs" section indocs/cloud.md.Tests
queue/test_checkpoint.py: 26 CPU-only tests, added to CI and AGENTS.md. They use a fake trainer and a stublichtfeld.iomodule. The log parser was also replayed against real patched-trainer logs: every snapshot was found at its real step.Not yet run:
licht_restore_point.pyagainst the reallichtfeld.iomodule in an image (the lab box went down). It should be exercised on the next rc.