What happened (2026-09-20)
Env creates failed with containerd no space left on device. The root disk was at 100% (248 MB free) while GET /host reported 142 GB free — because health.js measures config.dataDir, which is the 300 GB DigitalOcean block volume (env bind mounts only), while Docker images + build cache live on the 120 GB root disk under /var/lib/containerd (Docker 29, containerd image store).
| Volume |
Size |
Used |
Holds |
| vda1 root disk |
116 G |
106 G → 97 G after cleanup |
/var/lib/containerd (images 84 GB + build cache 43 GB, overlapping), /root/.npm 9 GB, OS |
| sda block volume |
298 G |
151 G |
DEVBOX_DATA_DIR (env data dirs) only |
Manual fix applied: untagged 2,125 images whose env dir no longer exists (3 per destroyed env: <name>-workspace, <name>-dev, <name>-wordpress, with and without the katalyst- prefix). Reclaimed only ~9 GB because the orphans shared layer chains with live envs. Nothing else touched (no builder prune, no container prune).
Why it fills up
- Destroy never removes images. Teardown is
npm run down → docker compose down with no --rmi local. Every destroyed env leaves its three built image tags behind forever (~1,090 destroyed envs → 2,125 tags).
- Generations, not envs, are the real cost. Live workspace images (60) resolve to only 12 distinct layer chains; dev → 9; wordpress → 5. Each build-cache miss creates a new ~3.4 GB chain that every env built afterwards shares until the next miss. Misses come from:
docker compose build --pull in initial-setup.sh (base image updates),
ADD https://raw.githubusercontent.com/wp-cli/builds/gh-pages/phar/wp-cli.phar in workspace.Dockerfile — BuildKit re-checks the remote on every build; a new phar invalidates everything after it, including the 1.67 GB npm install -g layer and the 585 MB Cursor layer,
- unpinned
npm install -g of the agent CLIs, and the curl cursor.com/install | bash step.
Twelve generations × three images ≈ the 84 GB.
start = up -d --build, so a stopped env restarted after a cache miss gets yet another generation and its old tag goes dangling.
- Build cache is never pruned (43 GB, 377 entries, 0 in use).
- Health/capacity checks look at the wrong disk, so creates run until containerd fails mid-create instead of being refused up front.
Ideas (pick any subset)
Safe-to-do-now notes
docker builder prune is safe for space at any time; the only cost is a longer next cache-miss build.
docker system prune -f (no -a) loses no env data (all bind mounts) but removes stopped envs' containers; they're recreated by compose up from the same images. -a would drop the image generations that stopped envs still need → 10-min rebuilds on start.
🤖 Generated with Claude Code
https://claude.ai/code/session_018puUXTJeSWdbG8ZMCFVqQL
What happened (2026-09-20)
Env creates failed with containerd
no space left on device. The root disk was at 100% (248 MB free) whileGET /hostreported 142 GB free — becausehealth.jsmeasuresconfig.dataDir, which is the 300 GB DigitalOcean block volume (env bind mounts only), while Docker images + build cache live on the 120 GB root disk under/var/lib/containerd(Docker 29, containerd image store)./var/lib/containerd(images 84 GB + build cache 43 GB, overlapping),/root/.npm9 GB, OSDEVBOX_DATA_DIR(env data dirs) onlyManual fix applied: untagged 2,125 images whose env dir no longer exists (3 per destroyed env:
<name>-workspace,<name>-dev,<name>-wordpress, with and without thekatalyst-prefix). Reclaimed only ~9 GB because the orphans shared layer chains with live envs. Nothing else touched (no builder prune, no container prune).Why it fills up
npm run down→docker compose downwith no--rmi local. Every destroyed env leaves its three built image tags behind forever (~1,090 destroyed envs → 2,125 tags).docker compose build --pullininitial-setup.sh(base image updates),ADD https://raw.githubusercontent.com/wp-cli/builds/gh-pages/phar/wp-cli.pharinworkspace.Dockerfile— BuildKit re-checks the remote on every build; a new phar invalidates everything after it, including the 1.67 GBnpm install -glayer and the 585 MB Cursor layer,npm install -gof the agent CLIs, and thecurl cursor.com/install | bashstep.Twelve generations × three images ≈ the 84 GB.
start=up -d --build, so a stopped env restarted after a cache miss gets yet another generation and its old tag goes dangling.Ideas (pick any subset)
root=in/etc/containerd/config.toml,data-rootin/etc/docker/daemon.json; stop server/docker/containerd → rsync → repoint → start). DO volumes resize online, so disk becomes "make it as big as needed". Needs a maintenance window (~20–40 min for 82 GB; all env containers down during the copy).--rmi local(compose down --rmi localremoves exactly the project's built images) + an orphan sweep on server boot: untag any env-shaped image whose name isn't in the registry (what was done by hand today).katalystwp/workspace:<hash of Dockerfile+args>) and reference it viaimage:from each env's compose instead of per-envbuild:. One chain shared by all envs; a new generation only when the template changes; old generations GC'd when no env references them.--pullfrom initial setup in favour of an explicit "refresh base" action.--buildon start / warm claim. Build at create and on explicit rebuild only.docker image prune(dangling) +docker builder prune --keep-storage <cap>.docker info/ containerd root) alongsidedataDir, and refuse creates with a clear503below a free-space threshold.Safe-to-do-now notes
docker builder pruneis safe for space at any time; the only cost is a longer next cache-miss build.docker system prune -f(no-a) loses no env data (all bind mounts) but removes stopped envs' containers; they're recreated bycompose upfrom the same images.-awould drop the image generations that stopped envs still need → 10-min rebuilds on start.🤖 Generated with Claude Code
https://claude.ai/code/session_018puUXTJeSWdbG8ZMCFVqQL