Skip to content

Seal a network-off run to its model broker - #2035

Merged
ppXD merged 1 commit into
mainfrom
fix/seal-a-network-off-run-to-its-broker
Sep 27, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/seal-a-network-off-run-to-its-broker

Conversation

@ppXD

@ppXD ppXD commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • On a confining host, a network-off run got --unshare-net. That severed it from its model broker along with everything else. Standard (network-off) is the tier every recipe recommends, and research mode forces the network off as well, so on such a host no default-tier brokered run reached a model. Pin whether a read-only reviewer can read its diff with git #2032 pinned this: the reviewer's model saw zero requests.
  • The executor now stamps the broker lease port onto a network-off run's spec, as SandboxSpec.ModelBrokerPort. AgentRunExecutor.ApplySealedEgress does this inside HardenSpec, which takes an input record so it stays under the parameter cap.
    • Two conditions must hold: bubblewrap confines, and FilteredEgressNetns.CanSeal has proved this process can build a namespace by creating and deleting a throwaway one.
    • When both hold, the runner builds a sealed namespace (FilteredEgressPlan.BuildSealed) instead of severing: no route, no NAT, no DNS, and an inet input filter on the host veth that admits only tcp dport <lease port> on the gateway.
  • Reusing the allowlist plan with no IPs would have left three holes:
    • it accepts DNS to any address;
    • it routes and NATs out;
    • it has no input filter, so Kestrel and every other run's broker port are reachable through the gateway.
  • A host that cannot seal keeps severing, exactly as before. Unconfined hosts are unchanged. That includes production as shipped: a non-root pod has the binaries but cannot build a namespace, which is why CanSeal is a real probe and not IsSupported.
  • The launch records SandboxConfinement.EgressSealedToBroker, which implies NetworkSevered. The journal then reads "Network: off (Standard) — confined: egress sealed to the run's model broker". That wording is pinned in the shared posture fixture and checked by both stacks.
  • Teardown also deletes the inet table. Every reaper is keyed on the run id alone, so none of them needed changes.
  • Review fixes, now that sealing is the default for network-off brokered runs:
    • The allocator skips any /30 the host already routes (HostRoutedPrefixes, read from ip -j -4 route show table all). The lock alone sees neither the worker's own network, where a /30 would shadow real peers, nor a run whose namespace outlived its worker. Re-issuing that run's /30 split two runs' broker replies between two veths.
    • The sealed ruleset replaces a leftover table of the same name instead of appending to it. Every revise round reuses the name.
    • A native launch whose bootstrap committed a rejected receipt now tears down the namespace and cgroup it was given. That receipt proves nothing was released, and before this no reaper could find those resources.
    • The seal probe now:
      • runs before the bootstrap's admission window opens;
      • retries a failure after a minute instead of caching it for the process lifetime;
      • keeps the failed step as SealUnavailableReason;
      • leaves at most one probe namespace behind per worker.
    • The Room ranks a sealed agent below a severed one, so the least-confined fold is deterministic, and names it ConfinedEgressSealedToBroker ("confined · egress sealed to its model broker") on both stacks.
  • Second review round:
    • A routed candidate does not count against the 4096 the walk may try. A broad route over the first range (a 10.1.0.0/16 LAN, say) moves the walk on through 10.0.0.0/8 instead of refusing. Routes broader than a /8 (a VPN's 0.0.0.0/1 pair) are treated as defaults.
    • Running out of /30s is a failed setup, so a sealed run's refusal can type it. A host whose reservation directory breaks is still told it is the mount, even when routes are in play.
    • A sealed launch drops the HTTP(S)_PROXY/ALL_PROXY variables it inherits. Its one destination is the broker on its gateway; a proxy is unreachable from the namespace and would turn every model call into a failure that looks like the provider's.
    • The seal probe is timed on a monotonic clock. The first probe is waited for by every caller; a re-probe is single-flight, and other launches keep the standing answer instead of queueing behind it.
    • A launch failure waits while the receipt is still committed, so a rejection that lands a moment later is still torn down.
  • Third review round:
    • A routed range is jumped over whole, so a 10/8 route costs one step instead of 4.1M. IndexAtOrAbove is pinned as the inverse of the candidate order at every octet boundary.
    • Null routes (blackhole/unreachable/prohibit/throw) are ignored.
    • This worker's own /30s are counted apart, so an all-routed host is never misreported as "4096 live runs".
    • The tools probe retries like the seal probe (CapabilityProbe).
    • A committed receipt whose broker is dead also proves nothing was released (the ready receipt precedes release), and is torn down.
    • Every brokered launch exempts its broker's address from NO_PROXY. This covers the allowlist netns gateway and a loopback that a NO_PROXY omits.
  • Fifth review round:
    • A * in either NO_PROXY spelling had collapsed both to a lone *. Codex's reqwest matches no IP address against *, so its brokered calls still went to the proxy. Now NO_PROXY keeps the broker's address beside the * (for Go and reqwest), and no_proxy stays a lone * (for curl, Python and undici, which honour it only as the whole value).
    • A child that was handed no proxy keeps its environment. A NO_PROXY written there exempts nothing, and it turned off Python's OS-proxy lookup on a macOS or Windows worker.
    • The allocator's route listing cannot see a null route in a table that a policy rule consults before main. Such a route wins by rule order, not prefix length. Every setup now runs ip route get <ns> from <gateway> after its setup commands (FilteredEgressPlan.RouteCheckArgv) and fails closed unless the answer is the run's own veth. The check also catches rule-level blackhole/prohibit actions, which no route listing shows. It runs for the allowlist plan too, but it is an output lookup: an allowlist run's NAT'd replies are routed on input, so a rule keyed on the uplink (iif) still escapes it. That gap is not new; the allowlist plan had no check before.
  • Sixth review round (every fix confirmed, minors only):
    • The route check reads ip route get's text answer (the word after its one dev). route get learned -j only in iproute2 5.0, three releases after the route listing, so on Debian 10's 4.20 every setup was refused. Reviewers checked it on a real kernel against Ubuntu 20.04/22.04/24.04, Debian 12, Rocky 8/9, Amazon Linux 2/2023 and Alpine. They also checked common policy routing: wg-quick, Tailscale with and without an exit node, Calico, Cilium, AWS VPC CNI, a VRF, rp_filter=1 and linkdown. Every verdict matched a real TCP exchange.
    • The proxy gate counts a worker variable only if the env scrub keeps it. A worker's ALL_PROXY never reaches the child, so it no longer writes NO_PROXY into a child with no proxy.
    • The broker-death proof takes its liveness read as a seam, so a test can move the receipt at the moment liveness is sampled. That pins the order itself, not only the re-read.
    • The sealed E2E asserts the child carries no proxy variable. The broker answering was not enough: the NO_PROXY exemption alone lets urllib past a proxy.
    • The policy-route E2E scopes its rule to its own /30 and first clears any stale rule a killed run left for that /30.
  • Known residual: a repo-committed .claude/settings.json env proxy still reaches an unpinned Claude run. That is the existing "repo settings load when --setting-sources is not pinned" exposure. Pinning every Claude run first needs a real-CLI check that project CLAUDE.md still loads.
  • Known residual, for the owner: the /30s come from 10.0.0.0/8, starting at 10.1.x. The route check sees the worker's connected and local networks and any surviving namespace. It cannot see a 10.x peer reached through a gateway (for example, a database on the host's LAN behind a Docker bridge). Moving the range to 198.18.0.0/15 (RFC 2544) would avoid this, but the broker's source-address gate and the allowlist resolver's SSRF guard key on 10/8, so it is left as a separate decision.

Test plan

  • Unit tests:
    • SandboxEgressPolicyTests: Sealed only for network-off plus a sealable port.
    • FilteredEgressPlanTests:
      • a sealed plan has no route, sysctl, DNS, masquerade or subnet key;
      • the table is inet;
      • the input filter carries the exact port;
      • teardown matches the run-id-only commands, including the inet table.
    • BubblewrapSandboxTests: the sealed record is severed and sealed; an unconfined record claims neither.
    • AgentRunExecutorEgressTests: the port is stamped only for network-off brokered runs.
    • NativeLaunchRegistryTests: the port binds the hash and is absent from an unbrokered spec's JSON.
    • NetworkPostureWordingDriftTests and launchInput.test.ts (84/84): pass on the new fixture case.
  • Integration tests:
    • A Standard brokered launch hands the runner its lease port.
    • A Trusted launch and an unbrokered launch carry none.
    • The sealed record survives the agent_run.sandbox_confinement jsonb column and reads back as sealed.
    • Mutation check: dropping ApplySealedEgress from HardenSpec turns the Standard row red.
  • Unit suite: full run green locally.
  • Sandbox lane (root, real bwrap, ip and nft):
    • New SealedEgressE2ETests, durable and non-durable, running the real runner and real broker, with a python3 probe inside the chain:
      • the broker answers 200;
      • 1.1.1.1:80 and 8.8.8.8:53 (TCP and UDP) stay shut;
      • a listener the test opens on the gateway stays shut;
      • the durable handle carries the run-keyed netns and a sealed record;
      • the namespace is gone after the run;
      • a netns-level arm shows the host veth's IPv6 link-local address is dropped too, or prints that IPv6 is off on the host.
    • Pin whether a read-only reviewer can read its diff with git #2032's finding test flips into A_network_off_reviewer_reaches_its_model_through_the_sealed_namespace: real Claude and Codex, a durable launch, a sealed record, and the diff reaching the model.
    • Every reviewer arm drops its network-on deviation and runs its tier's production posture.
  • Unit:
    • EgressSubnetAllocatorTests: a host-routed /30 is skipped, including a survivor's /30 and a /30 holding a host address. A host that routes every candidate is refused with a message naming its routes, not "4096 live runs". Plus the overlap theory.
    • The rejected-receipt predicate: only rejected tears down.
    • The sealed ruleset is pinned whole.
    • Room: sealed vs severed in either arrival order, and the producer word.
    • The sealed subnet caveat.
  • Frontend: vitest (106), tsc --noEmit -p tsconfig.app.json, eslint.
  • Sandbox lane:
    • The IPv6 arm now waits out duplicate-address detection and proves its probe with a table-deleted control. A skip no longer satisfies the lane.
    • The durable arm checks that the host inet table is gone.
    • New A_30_still_held_by_a_run_that_outlived_its_worker_is_not_handed_to_the_next_run.
    • ModelCredentialBrokerNetnsE2ETests covers both the sealed and the allowlist namespace.
  • Unit:
    • EgressSubnetAllocatorTests: a 10.1.0.0/16 route moves the walk to 10.2.1.0/30; /1 routes are defaults; the remount-ro refusal names the mount with routes present.
    • SealProbeCacheTests: retry interval, kept reason, a proof kept for good, and a non-blocking re-probe.
    • The proxy strip: sealed only.
    • A receipt still committed is waited for.
  • Sandbox lane: the sealed probe carries an unreachable HTTP_PROXY, and it reports the proxy variables its child actually has; the lane requires none, on both durable and non-durable launches.
  • Unit (third round): the inverse drift detector, routed-range jumps, null routes, the own-hold count, both-probe retries including concurrent first callers and repeated failures, the committed-and-dead-broker receipt with atomically published fixtures, the NO_PROXY exemption, and the Reserve retype.
  • Integration: the full suite on the stack top (5217) before this round; the touched classes after it.
  • Unit (fifth round):
    • The NO_PROXY theory with per-spelling expectations, including both * placements.
    • A child with no proxy is left untouched.
    • The route check: argv, a route through the run's veth, another device, a local shadow, blackhole and unreachable exits, and malformed answers.
    • A dead broker whose receipt is already ready/indeterminate, or committed by another broker, proves nothing.
    • Mutation-checked: dropping the proxy gate, or collapsing the death proof to a liveness check, turns the new tests red.
    • Full unit suite green (11299).
  • Unit (sixth round):
    • Route-check rows in text form: the veth, a gateway hop, a local shadow, no device, two answers, JSON.
    • A worker ALL_PROXY/all_proxy alone writes nothing, while a worker https_proxy is exempted.
    • The broker-death order, through the seam.
    • Mutation-checked: reading the receipt before liveness, or falling back to every worker variable, turns them red.
  • Sandbox classes on a real kernel, locally (privileged Ubuntu 24.04 container: bwrap, iproute2 6.1, nft 1.0.9): SealedEgressE2ETests, FilteredEgressNetnsE2ETests, ModelCredentialBrokerNetnsE2ETests, DurableLaunchEgressE2ETests.
  • Sandbox lane: new A_host_whose_policy_rule_discards_the_run_s_replies_fails_the_setup_and_leaks_nothing. It uses a TEST-NET-1 /30 that the allocator never hands out. The same plan first sets up cleanly as a control. It then fails under unreachable 192.0.2.0/24 in a table consulted before main, names the lookup that found it, and leaves no namespace or veth behind. The kernel answers were checked by hand in a privileged container: the veth on a clean host, and exit 2 Host is unreachable under the rule.

@ppXD
ppXD force-pushed the fix/stand-codex-sandbox-down-under-ours branch from 2cb0cfc to 0e96f1f Compare September 26, 2026 23:35
@ppXD
ppXD force-pushed the fix/seal-a-network-off-run-to-its-broker branch from eb9be4b to 66df2db Compare September 26, 2026 23:54
@ppXD
ppXD force-pushed the fix/stand-codex-sandbox-down-under-ours branch from 0e96f1f to 6a6a9b4 Compare September 27, 2026 00:32
@ppXD
ppXD force-pushed the fix/seal-a-network-off-run-to-its-broker branch 4 times, most recently from f71bbf6 to 3aaa4bb Compare September 27, 2026 02:54
@ppXD
ppXD force-pushed the fix/stand-codex-sandbox-down-under-ours branch from 6a6a9b4 to f8bcc43 Compare September 27, 2026 08:15
@ppXD
ppXD force-pushed the fix/seal-a-network-off-run-to-its-broker branch from 3aaa4bb to 0baf9a0 Compare September 27, 2026 08:16
@ppXD
ppXD changed the base branch from fix/stand-codex-sandbox-down-under-ours to main September 27, 2026 08:16
Under bubblewrap a network-off run got --unshare-net: an empty
namespace, which severed it from its model broker along with everything
else. Standard, the tier every recipe recommends, is network-off, so on
a confining host no default-tier brokered run reached a model.

The executor now stamps a network-off run's broker lease port onto its
spec. Where bubblewrap confines and the host has proved it can build a
namespace, the runner runs such a run in a sealed one instead: no
route, no NAT, no DNS, and an inet input filter on the host veth that
admits only that port on the gateway. The allowlist plan with no IPs
would not do: it accepts DNS anywhere, routes out, and has no input
filter, so the worker's own listeners and every other run's broker were
reachable through the gateway.

A host that cannot seal keeps severing, as before. The launch records
the seal, and the journal says the run was sealed to its broker.

Sealing makes the per-run /30 the default for network-off brokered
runs, so the allocator now also skips any /30 the host already routes
(ip -j -4 route show table all). Its lock sees neither the worker's own
network nor a run whose namespace outlived the worker that reserved
it; handing either out shadowed real peers or split two runs' broker
replies between two veths. A routed range is passed over whole without
counting against the 4096 the walk may try, so a broad route over the
first range moves the walk on through 10.0.0.0/8 rather than refusing;
null routes broader than a /30 are ignored, and running out of /30s is
a failed setup like any other. A null route in a table that a policy
rule consults before main wins by rule order, not prefix length, which
no route listing shows, so every setup now asks the kernel how it
reaches the namespace (ip route get <ns> from <gateway>, read as text:
route get learned -j only in iproute2 5.0) and fails unless the answer
is the run's own veth. The sealed ruleset replaces a
leftover table of the same name instead of appending to it, since every
revise round reuses the name. A launch that provably released nothing (a rejected
receipt, or a committed one whose broker died) now tears down the
namespace and cgroup it was given. The tools and seal probes retry a
failure after a minute instead of caching it for the process, keep the
reason, and the seal probe runs before the bootstrap's admission window
opens. The Room ranks and names
a sealed agent as its own posture, between confined and severed. A
sealed launch drops the proxy variables it inherits, since its one
destination is the broker on its gateway, which a proxy cannot reach.
Every brokered launch that was handed a proxy (by its task, or by a
worker variable the env scrub keeps) also exempts its broker's address
from NO_PROXY. Beside a "*" the address stays in NO_PROXY,
because reqwest (Codex) matches no IP address against "*", while
no_proxy stays a lone "*", which curl, Python and undici honour only as
the whole value.
@ppXD
ppXD force-pushed the fix/seal-a-network-off-run-to-its-broker branch from 0baf9a0 to d3c6729 Compare September 27, 2026 09:10
@ppXD
ppXD merged commit 908ee2f into main Sep 27, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant