Skip to content

Grade with the producing run's network and resource posture - #2080

Merged
ppXD merged 1 commit into
mainfrom
fix/grade-with-the-producing-runs-posture
Oct 7, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/grade-with-the-producing-runs-posture

Conversation

@ppXD

@ppXD ppXD commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Summary

  • Before this change, an acceptance setup step ran with AllowNetwork=true hard-coded (host network under bubblewrap), and neither setup nor check had cgroup ceilings. Benchmark and qualification grading ran the fixture's tests over the agent's workspace on the raw local runner.
    • Every lane now grades under an AcceptanceGradingPosture taken from the run whose bytes it executes: its network grant and egress allowlist, plus its tier's memory and CPU row, clamped by Sandbox:MaxAutonomy and narrowed by Sandbox:AgentMemoryCeilingMb.
    • AcceptanceGradingPosturePolicy.Bind wraps the grading runner once, narrow-only, for the setup and every oracle.
    • A missing posture fails closed to network off and the Confined ceilings.
  • What a grade reports follows what the sandbox did, not what the spec asked.
    • Runners say which egress they enforce (ISandboxEgressEnforcement; LocalProcessRunner.EnforcedEgress).
    • A notice (setup network: …) heads the evidence only when the setup was actually severed or filtered.
    • A setup the sandbox severed fails as setup-failed-network-severed: …. That is infra (no revise round), and agent.code does not respawn it, because the posture comes from the same stored task every time.
  • A check killed at the grade's memory ceiling (ResourceExhausted) is tests-resource-exhausted, class Environment, like tests-timed-out. Before, it was a Genuine tests-failed-exit-137.
  • Lanes:
    • Executor branch, patch, multi-repo and local lanes: the run's own task.
    • Supervisor per-unit, baseline, captured, resolve (branch and patch) and branchless-stop lanes: the unit's stored task.
    • Stop over the integrated head: the run profile's tier.
    • BenchmarkRunner: its agent's task.
    • TaskLaunch benchmark cells: the posture all attempts agree on, otherwise fail-closed.

Owner-visible behaviour (setup network applies where bubblewrap confines; ceilings apply where Sandbox:CgroupRoot is delegated):

Producing run Setup network Setup + check ceilings
Confined severed (was host network) 1024 MiB, 1 core (was uncapped)
Standard (default for AgentTask, agent.code, supervisor units) severed (was host network) 4096 MiB, 4 cores (was uncapped)
Trusted / Unleashed host network (unchanged) 6144 MiB, 4 cores (was uncapped)
Egress allowlist only the operator's EgressAllowHosts, filtered where the host filters, severed where it cannot; no hosts → severed the tier's row
Unknown producer / positional overload severed Confined row
  • An operator setup that downloads on a network-off run no longer does where the sandbox confines. The fix is an egress allowlist or a higher tier.
  • The operator stop floor over the integrated head now grades under the run profile's tier (default Standard) inside the 300 s default window. A CPU-heavy floor is likelier to hit tests-timed-out.
  • Benchmark and qualification checks now run under the agent's own ceilings, and the TaskLaunch fallback is Confined.
  • Baselines are memoized per posture.
  • EvaluatorVersion is supervisor-acceptance/v9.

Test plan

  • Unit:
    • derivation table
    • For() read through real configuration (Sandbox:MaxAutonomy, AgentMemoryCeilingMb)
    • EnforcedEgress table and notice table
    • severed / unconfined / filtered setup failures through the real grader + TestsPassGrader
    • a check OOM under Standard and Trusted postures
    • pinned literals (tests-resource-exhausted, setup-failed-network-severed:)
    • agent.code retry rows
    • BenchmarkTaskGrading posture and fail-closed
    • TaskLaunch GradedPosture
    • inventory of grade lanes and oracle-runner bindings
  • Integration:
    • resolve posture Theory over branch and patch arms
    • executor flow under the real LocalProcessRunner
    • real BenchmarkRunner hands the oracle a runner bound to its agent's posture
    • supervisor per-unit fold flows
  • Regression: full unit suite (11950 passed, 1 skipped); integration Acceptance|OracleRestore|NonCodingOracle|LocalAcceptance|Benchmark|Qualification|AgentNodeFlow|ReviseLoop (338 passed); .Supervisor|Oracle|PlanMap|UnderClaim|Resolve (679 passed)
  • Mutation: Posture = null on the resolve patch arm fails 2 rows; For() without the host budget fails the configuration test
  • Sandbox lane (Linux bwrap): AcceptanceGradingPostureE2ETests, covering a real severed setup that reports setup-failed-network-severed: and a host that reports the severance it enforces

@ppXD
ppXD changed the base branch from fix/keep-model-authored-acceptance-from-carrying-setup to main October 7, 2026 04:41
@ppXD
ppXD force-pushed the fix/grade-with-the-producing-runs-posture branch from 44bb768 to 19467a2 Compare October 7, 2026 04:41
An acceptance grade's setup step was hard-coded to AllowNetwork=true,
so under bubblewrap it shared the host network whatever the producing
run's tier, and neither the setup nor the check had a memory or CPU
ceiling. A setup like `npm ci` or `pip install` runs manifests the
agent wrote, after the agent's own sandbox is gone. So a network-off
Standard run could have its planted code executed with full egress,
and a runaway install had no cgroup bound. Benchmark and qualification
grading ran the fixture's tests over the agent's workspace on the raw
local runner, with no ceiling either.

Every lane now carries an AcceptanceGradingPosture, derived once from
the run whose bytes it grades. The posture is that run's network grant
and egress allowlist plus its tier's resource ceilings, clamped by
Sandbox:MaxAutonomy the way RunCommandService.BuildSpec clamps
agent.run_command. The executor lanes (branch, patch, multi-repo,
local) take it from the run's own task. The supervisor's per-unit,
baseline, captured, resolve and branchless-stop lanes read the unit's
stored task. The stop over the integrated head uses the run profile's
tier, which is the tier every unit is clamped to. BenchmarkRunner uses
its agent's task; a TaskLaunch cell uses the posture its attempts agree
on, and fails closed when they do not.

AcceptanceGradingPosturePolicy.Bind wraps the grading runner once, so
the setup and every oracle's command run narrowed, narrow-only. The
setup keeps the network only when the producer had it. An allowlist
producer is filtered to its operator-configured hosts, and an
allowlist with no host severs. Both steps run under the tier's
ceilings. A request with no posture grades with network off under the
Confined ceilings.

What the grade reports follows what the sandbox did, not what the spec
asked. The runner now says which egress it enforces for a spec
(ISandboxEgressEnforcement). Only when it severed or filtered the setup
does one notice head the grade's evidence, so an unconfined host never
reads "off" for a setup that kept its network. A setup the sandbox
severed fails as "setup-failed-network-severed:": still infra, but the
posture comes from the same stored task on every attempt, so agent.code
no longer re-buys an agent run to sever it again.

Grades can now hit a memory ceiling. A check the runner kills there
(ResourceExhausted) is "tests-resource-exhausted", an Environment fact
like tests-timed-out, instead of a genuine failure that bought revise
rounds and recorded a verdict on the code.

This changes what operators see: an operator setup on a network-off
run no longer downloads where the sandbox confines. The fix is an
egress allowlist or a higher tier. Trusted and Unleashed runs keep
full egress. Baselines are now memoized per posture, so two units of
different tiers off one base are no longer compared across two
different sandboxes. EvaluatorVersion moves to v9.
@ppXD
ppXD merged commit 4cc86d7 into main Oct 7, 2026
6 checks passed
@ppXD
ppXD deleted the fix/grade-with-the-producing-runs-posture branch October 7, 2026 06:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant