Skip to content

fix(leashd): degrade to proxy-only enforcement when cgroup path unavailable - #74

Open
sixtoad wants to merge 1 commit into
strongdm:mainfrom
sixtoad:fix/cgroup-degraded-mode
Open

sixtoad wants to merge 1 commit into
strongdm:mainfrom
sixtoad:fix/cgroup-degraded-mode

Conversation

@sixtoad

@sixtoad sixtoad commented Jun 19, 2026

Copy link
Copy Markdown

Summary

On Docker Desktop with Kubernetes the agent container runs in a private cgroup namespace where /proc/self/cgroup reports 0::/. leash-entry's emitCgroupPath() skips / as invalid, so /leash/cgroup-path is never written, and leashd then FATALs at startup with cgroup path required (set --cgroup) — before any enforcement starts.

This is separate from #66 (which fixed the xt_cgroup network path): even with that fix, leashd can't start on Docker Desktop K8s because cgroup-path discovery fails first.

This change makes leash degrade gracefully instead of failing, following the same philosophy as #66 (cgroup unavailable → fall back + WARNING, never FATAL, preserve the boundary):

  • cmd/leash-entry/main.go — when no scopable cgroup is found, log a WARNING and continue (don't abort the container) so leashd can start.
  • internal/leashd/runtime.go — preFlight allows an empty cgroup path (proxy-only, warns); a non-empty path is still validated as before.
  • internal/lsm/manager.go — UpdateRuntimeRules no-ops when there is no cgroup, so no cgroup-scoped BPF-LSM programs are attached (they'd have nothing valid to scope to).
  • internal/assets/apply-{iptables,ip6tables,nftables}.sh — when there is no target cgroup, still isolate the control plane with a namespace-wide block (the same boundary as the existing network: fall back to blanket port block when cgroup isolation unavailable #66 cgroup fallback), so degraded mode stays safe.

Net: the kernel LSM layer (file/exec/connect) is disabled when the cgroup is undiscoverable, but the L7 MITM proxy (hostname/header/MCP) keeps enforcing and the control plane stays isolated — defense-in-depth minus the kernel layer, with a clear warning, rather than a hard failure.

Design note

This follows #66's pattern: automatic fallback + WARNING, no new flag. Since degraded mode is a reduction in enforcement, I'm happy to instead gate it behind an explicit opt-in flag (e.g. --allow-no-cgroup) if you'd prefer that posture.

Validation

  • go build + go vet clean for internal/lsm, internal/leashd, cmd/leash-entry.
  • New unit tests pass: LSM-manager no-op without a cgroup; preFlight allows a missing cgroup. All 14 existing preFlight tests still pass (non-empty invalid cgroup is still rejected).
  • sh -n clean on all three netfilter scripts.
  • Not yet validated end-to-end on Docker Desktop Kubernetes (no access to that environment). Would appreciate maintainer/CI verification there, plus a no-regression check on bare Linux. Happy to iterate.

Closes #67

…ilable

Some environments (notably Docker Desktop with Kubernetes) run the agent
container in a private cgroup namespace where /proc/self/cgroup reports
"0::/". leash-entry then cannot emit a scopable /leash/cgroup-path, and
leashd previously FATALed at startup with "cgroup path required (set
--cgroup)" before any enforcement could start.

Mirror the cgroup-unavailable network fallback (#66): degrade gracefully
instead of failing.

- cmd/leash-entry/main.go: when no scopable cgroup is found, log a WARNING
  and continue (do not abort the container) so leashd can start.
- internal/leashd/runtime.go: preFlight allows an empty cgroup path (warns,
  proxy-only); a non-empty path is still validated.
- internal/lsm/manager.go: UpdateRuntimeRules no-ops without a cgroup, so no
  cgroup-scoped BPF-LSM programs are attached (kernel layer disabled).
- internal/assets/apply-{iptables,ip6tables,nftables}.sh: with no target
  cgroup, still isolate the control plane via a namespace-wide block (same
  boundary as the cgroup fallback), so degraded mode stays safe.

The L7 MITM proxy (hostname/header/MCP) is unaffected and keeps enforcing,
preserving defense-in-depth minus the kernel layer.

Closes #67
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cgroup path discovery fails on Docker Desktop Kubernetes (path resolves to /)

1 participant