Skip to content

Sever an allowlist run this worker cannot filter - #2048

Merged
ppXD merged 1 commit into
mainfrom
fix/sever-an-allowlist-this-worker-cannot-build
Sep 29, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/sever-an-allowlist-this-worker-cannot-build

Conversation

@ppXD

@ppXD ppXD commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Summary

  • On a confining worker that cannot filter, an allowlist run is severed. Before, a worker with ip and nft installed planned every allowlist run Filtered (FilteredEgressNetns.IsSupported). The shipped worker (uid 1654, no capabilities) then admitted the run, spent on it, and aborted it at ip netns add ("Filtered-egress netns setup failed").
    • A root worker whose /proc/sys is read-only did worse. procps' sysctl -w net.ipv4.ip_forward=1 warns and exits 0 there, so the run launched as filtered while none of its allowed hosts was reachable.
    • Now, where bubblewrap confines, the plan keys on FilteredEgressNetns.CanFilter (renamed from CanSeal). That is the tools, one throwaway namespace, and ip_forward reading 1 or passing access(W_OK), a check that changes nothing.
    • Where that does not hold, the allowlist fails closed to no egress, as SandboxEgressPolicy documents. bubblewrap severs the run, its record says NetworkSevered, and a brokered run still reaches its broker through the relay.
  • Unconfined hosts behave exactly as before. Where bubblewrap does not confine (the worker as docker-compose.yml ships it), nothing would enforce a severance, so a severed allowlist would share the worker's network. There the binaries alone still plan the run Filtered: its setup filters it or aborts the launch, and it never reaches the worker's network.
    • The launch and the admission both read that one host fact: LocalProcessRunner.FiltersAllowlist / HostFiltersAllowlist in LocalProcessRunner.Durable.cs.
    • Dockerfile.worker describes both postures.
  • Runs in flight keep what they had. Tearing a namespace down by name and the legacy re-bind still key on the tools alone.
  • Boot line. The posture line reports CanFilter and FilterUnavailableReason. FilteredEgressNetns.FilterProbe(path) is the probe CanFilter runs, with the forwarding file passed in, so the root lane can pin the whole probe and not only its parts.

Test plan

  • Unit, full suite, macOS: 11499/11500 (1 skipped).
  • Unit, touched classes, on Linux in a privileged container, run as root, as uid 1654 where bubblewrap confines, and as uid 1654 unconfined: 427 each. The one failure, LocalProcessDurableRunnerTests.Supervisor_seals_logs_only_after_a_background_descendant_closes_its_output_writer, fails identically on the base.
    • AllowlistFilterProbeTests covers:
      • the probe truth table (tools, namespace, forwarding);
      • the posture table (confines, canFilter, tools);
      • forwarding against real files.
    • SealedEgressAdmissionTests pins the public admission against this host's probes.
  • Integration (AgentRunExecutorTests, macOS): 144/144.
  • Sandbox, root, with the real claude and codex CLIs: 77/77.
    • The allowlist run is still filtered.
    • A file reading 0, bound read-only, is named by FilterProbe, and a writable one is not.
  • Sandbox, non-root (uid 1654, no capabilities, no_new_privs): 10/10.
    • The allowlist arm is admitted and launched with no namespace, is recorded NetworkSevered, and reaches its broker through the relay. The allowlisted IP stays unreachable.
    • The fixture asserts that its wall is the namespace, not forwarding.
  • Sandbox, unconfined (new Category=SandboxUnconfined: uid 1654, CODESPACE_BWRAP_PATH pointing at no binary, RequireConfinement unset): 3/3.
    • The allowlist run is aborted at setup, durable or not, and never launched on the worker's network.
    • The admission predicts its namespace.
  • The workflow's assertion blocks pass against these trx files. The new step and its assertion pass when replayed verbatim.
  • Mutations turn red, then pass once restored:
    • forwarding step dropped from FilterProbe;
    • CanFilter everywhere (the unconfined lane launches unfiltered);
    • IsSupported everywhere (the non-root run aborts again);
    • the public admission reads CanFilter;
    • the non-root wall assertion removed (a staged uid that can build a namespace passes the guard);
    • the reads-1 branch dropped.

@ppXD
ppXD force-pushed the fix/guard-an-allowlist-runs-veth-both-ways branch from 1c54271 to bb72607 Compare September 29, 2026 07:21
@ppXD
ppXD force-pushed the fix/sever-an-allowlist-this-worker-cannot-build branch from b7619db to 2372eb7 Compare September 29, 2026 07:21
@ppXD
ppXD changed the base branch from fix/guard-an-allowlist-runs-veth-both-ways to main September 29, 2026 08:16
@ppXD
ppXD force-pushed the fix/sever-an-allowlist-this-worker-cannot-build branch from 2372eb7 to 847bb38 Compare September 29, 2026 08:16
An allowlist run was planned Filtered wherever ip and nft were
installed, which is all FilteredEgressNetns.IsSupported checks. The
shipped worker is uid 1654 with no capabilities and ships both, so
every allowlist run there was admitted, spent, and then aborted at
`ip netns add` ("Filtered-egress netns setup failed"). A root worker
whose /proc/sys is read-only fared worse: procps' `sysctl -w
net.ipv4.ip_forward=1` warns and exits 0 there, so with forwarding at
0 the setup succeeded and the run launched as filtered with its
allowed hosts unreachable.

Where bubblewrap confines, the runner's egress derivation, and the
admission that predicts it, now key on FilteredEgressNetns.CanFilter,
renamed from CanSeal: the tools, the throwaway-namespace probe, and
net.ipv4.ip_forward reading 1 or passing access(W_OK), which changes
nothing. Where it does not hold, an allowlist fails closed to no
egress, the contract SandboxEgressPolicy has always stated: bubblewrap
severs the run, its record says NetworkSevered, and a brokered run
still reaches its broker through the relay. The boot posture line
reports CanFilter and FilterUnavailableReason, and the Dockerfile says
what an allowlist needs and what a worker without it does.

Where bubblewrap does not confine, as docker-compose.yml ships the
worker, nothing would enforce that severance: the run would share the
worker's network. There the binaries alone still plan it Filtered, so
its setup filters it or aborts the launch, exactly as before. The
launch and the admission read that one host fact
(LocalProcessRunner.FiltersAllowlist). A worker that can filter still
filters. Tearing a namespace down by name and the legacy re-bind still
key on the tools alone, so runs in flight keep what they have.

The non-root sandbox lane gains an allowlist arm that is admitted,
launched with no namespace, recorded severed, and reaches its broker
while the allowlisted IP stays unreachable; keyed back on the tools it
aborts at setup again. A new unconfined lane (uid 1654, with
CODESPACE_BWRAP_PATH naming no binary) pins that the same run there is
aborted at setup, durable or not, and that the admission predicts its
namespace; keyed on CanFilter everywhere, it launches on the worker's
network. The root lane asserts its allowlist run is still filtered, and
binds a file reading 0 read-only to show that the whole filter probe,
not only its forwarding step, names forwarding root may not turn on.
The non-root fixture also asserts its wall is the namespace, not
forwarding, so a lane that can build one cannot pass as the shipped
posture.
@ppXD
ppXD merged commit 4d28091 into main Sep 29, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant