Skip to content

Docker disk exhaustion: orphan images, layer generations, and health reporting the wrong volume #84

Description

@louisreingold

What happened (2026-09-20)

Env creates failed with containerd no space left on device. The root disk was at 100% (248 MB free) while GET /host reported 142 GB free — because health.js measures config.dataDir, which is the 300 GB DigitalOcean block volume (env bind mounts only), while Docker images + build cache live on the 120 GB root disk under /var/lib/containerd (Docker 29, containerd image store).

Volume Size Used Holds
vda1 root disk 116 G 106 G → 97 G after cleanup /var/lib/containerd (images 84 GB + build cache 43 GB, overlapping), /root/.npm 9 GB, OS
sda block volume 298 G 151 G DEVBOX_DATA_DIR (env data dirs) only

Manual fix applied: untagged 2,125 images whose env dir no longer exists (3 per destroyed env: <name>-workspace, <name>-dev, <name>-wordpress, with and without the katalyst- prefix). Reclaimed only ~9 GB because the orphans shared layer chains with live envs. Nothing else touched (no builder prune, no container prune).

Why it fills up

  1. Destroy never removes images. Teardown is npm run down → docker compose down with no --rmi local. Every destroyed env leaves its three built image tags behind forever (~1,090 destroyed envs → 2,125 tags).
  2. Generations, not envs, are the real cost. Live workspace images (60) resolve to only 12 distinct layer chains; dev → 9; wordpress → 5. Each build-cache miss creates a new ~3.4 GB chain that every env built afterwards shares until the next miss. Misses come from:
    • docker compose build --pull in initial-setup.sh (base image updates),
    • ADD https://raw.githubusercontent.com/wp-cli/builds/gh-pages/phar/wp-cli.phar in workspace.Dockerfile — BuildKit re-checks the remote on every build; a new phar invalidates everything after it, including the 1.67 GB npm install -g layer and the 585 MB Cursor layer,
    • unpinned npm install -g of the agent CLIs, and the curl cursor.com/install | bash step.
      Twelve generations × three images ≈ the 84 GB.
  3. start = up -d --build, so a stopped env restarted after a cache miss gets yet another generation and its old tag goes dangling.
  4. Build cache is never pruned (43 GB, 377 entries, 0 in use).
  5. Health/capacity checks look at the wrong disk, so creates run until containerd fails mid-create instead of being refused up front.

Ideas (pick any subset)

  • Host option, no code: move Docker + containerd roots to the block volume (root= in /etc/containerd/config.toml, data-root in /etc/docker/daemon.json; stop server/docker/containerd → rsync → repoint → start). DO volumes resize online, so disk becomes "make it as big as needed". Needs a maintenance window (~20–40 min for 82 GB; all env containers down during the copy).
  • Teardown with --rmi local (compose down --rmi local removes exactly the project's built images) + an orphan sweep on server boot: untag any env-shaped image whose name isn't in the registry (what was done by hand today).
  • Build the workspace image once per template version on the server (katalystwp/workspace:<hash of Dockerfile+args>) and reference it via image: from each env's compose instead of per-env build:. One chain shared by all envs; a new generation only when the template changes; old generations GC'd when no env references them.
  • Cache-stable Dockerfile: pin the wp-cli phar to a versioned release URL, pin npm package versions, move volatile steps last, drop --pull from initial setup in favour of an explicit "refresh base" action.
  • No --build on start / warm claim. Build at create and on explicit rebuild only.
  • Scheduled GC in the server: docker image prune (dangling) + docker builder prune --keep-storage <cap>.
  • Health: report the Docker root disk (from docker info / containerd root) alongside dataDir, and refuse creates with a clear 503 below a free-space threshold.

Safe-to-do-now notes

  • docker builder prune is safe for space at any time; the only cost is a longer next cache-miss build.
  • docker system prune -f (no -a) loses no env data (all bind mounts) but removes stopped envs' containers; they're recreated by compose up from the same images. -a would drop the image generations that stopped envs still need → 10-min rebuilds on start.

🤖 Generated with Claude Code

https://claude.ai/code/session_018puUXTJeSWdbG8ZMCFVqQL

Activity

  1. louisreingold commented on Sep 26, 2026

    @louisreingold
    MemberAuthor

    Host option done (2026-09-26): Docker + containerd roots moved to the block volume per docs/docker-data-root-migration.md (PR to follow).

    • /etc/containerd/config.toml → root = "/mnt/volume_nyc1_1783631841059/containerd"; /etc/docker/daemon.json → data-root = "/mnt/volume_nyc1_1783631841059/docker" (backups *.bak.pre-move).
    • Verified: identical image/container counts (201 / 301) and docker system df totals after the move; a probe build wrote 579 MB to the volume and 0 to the root disk; an env restarted and a Claude session ran fine. Outage ≈ 4 min (warm rsync first).
    • Still pending: delete /var/lib/containerd.old + /var/lib/docker.old after a day (frees ~80 GB on root); resize the volume (now 87 % full).

    The code-side items above (teardown --rmi local, orphan sweep, shared base image, cache-stable Dockerfile, no --build on start, scheduled GC, health reporting the Docker disk) remain open. Note health's diskFor(dataDir) now happens to measure the right disk, since images and env data share the volume.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions