Skip to content

Recite the operator's acceptance floor to the supervisor decider - #2060

Merged
ppXD merged 1 commit into
mainfrom
fix/show-the-supervisor-its-acceptance-floor
Sep 30, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/show-the-supervisor-its-acceptance-floor

Conversation

@ppXD

@ppXD ppXD commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • The stop path runs the operator's floor argv (SupervisorTurnContext.AcceptanceChecks) on every branch a stop ships, 300 s each. A failing floor withholds those branches and ends the run as AcceptanceFailed, and a stop is final, so no turn is left to fix it. The decider never read the field (it rendered only the free-text criteria, and the system prompt mentions "the operator's floor" without showing it), so the brain drove the work blind to the check that decides the run. BuildUserPrompt now renders this block right after the criteria block, only when a floor is set, shown here for ["dotnet","test","--filter","Category=Unit|Category=Smoke"]:

    Operator acceptance floor (the server runs this argv on every branch a stop ships, 300 s timeout each; if it fails, those branches are withheld and the run ends as AcceptanceFailed — a stop is final, so no turn is left to fix it; drive the work until it passes before declaring success):
      ["dotnet","test","--filter","Category=Unit|Category=Smoke"]
    Do not repeat it as the stop payload's acceptance: the stop runs it regardless, so a copy only runs it twice on every branch. A subtask's own acceptance is a separate per-unit check.
    
  • The copy is chosen so it never contradicts another instruction. It ends on "before declaring success", not "before stopping": the model cannot see the floor's verdict before it stops, and gave_up and ask_human stay open on a run that cannot pass. "Every branch a stop ships" is deliberate, because a branchless stop records operator-floor not graded (no head). The closing line names the stop payload's acceptance because a subtask's own acceptance is a different field, and a per-unit check that runs the same command is the one way the model can see a verdict early. The 300 s and the AcceptanceFailed word are read from SupervisorLane.AcceptanceGradeTimeoutSeconds and SupervisorOutcome.AcceptanceFailedOutcome, not retyped. The header is a static readonly string rather than a const only because C# cannot fold an int constant into a constant string; the closing line is a const. The tests use both, so they never retype the copy.

  • The argv is written as a JSON array of strings, not as shell text. The grader spawns it with no shell (TestsPassGrader hands the first element to the runner as the program and the rest as arguments), so a shell-looking line misleads: --filter Category=Unit|Category=Smoke reads as a pipe, a joined line cannot tell ["sh", " ", "check.sh"] from two arguments, and BoundOneLine flattens a newline into a space. The array is written with WorkflowJson.InterpolatedText, the repo's own way to put an array into prompt text, so quote, pipe, ampersand, angle brackets and plus stay as themselves while a newline stays a visible escape. The line is bounded with BoundOneLine(..., 400) like the per-subtask check line, and a longer one is cut and marked with the class's ellipsis. A floor-free prompt is byte-identical to before.

  • Untouched on purpose: the system prompt (its digest gates paid qualification through QualificationRuntimeGate), the stopped-now recital (precomposed from the tape and mirrored by the golden fixture), SupervisorQualityFacts (the floor is a run-level gate, not a unit-level CheckDeclared), and the per-subtask check line, which still joins its argv with spaces.

  • Exposure, flagged and not changed: the floor is operator-configured argv and now goes to the model provider verbatim. No redaction is added. The recording decorator masks the run's resolved credential values and declared secret paths only in what it stores (RecordingLLMClientDecorator.cs:233-236; PersistenceSecretRedactor masks "only at persistence boundaries"), so it does not stop a literal in the argv from reaching the provider. The same argv is already persisted verbatim in the frozen supervisor config, the launch contract snapshot and the grader's evidence, so this adds the provider hop rather than a new store.

  • Code versus docs: the doc comment on SupervisorTurnContext.AcceptanceChecks says "blank entries dropped", but NormalizeCommand keeps blank elements after the executable and refuses a blank executable or a NUL anywhere (pinned by Rehydrate_and_the_stop_grader_preserve_exact_operator_argv, and by TaskLaunchFlowTests for ["sh", " ", "check.sh"]). The renderer follows the code and the comment is left as it is.

Test plan

  • New tests, 9 methods and 18 cases. SupervisorDeciderTests +6 methods / 11 cases: the whole block for the pipe-bearing argv above, the timeout and outcome word read from their constants, the ["sh"," ","check.sh"] boundary, a JSON-fidelity theory (pipe, it's & <b> +1, a newline versus a space, the rehydrate's empty and blank elements), the 400 bound (exactly 400 verbatim, 401 cut with …), and the position right after the criteria. SupervisorTurnServiceTests +2 methods / 6 cases, both through the real RehydrateFromDecisionLogAsync: the operator's argv reaches the prompt with its boundaries, and no block renders for null, [], ["", "x"], [" ", "x"] or ["x", "a\0b"], each with a fixture check that the rehydrate really refuses it. Golden receipt +1.
  • Golden digest: GoldenPromptDigest moves from 2dafcdc7e52c3d22d1d2d209bb05bbe701191131defeb9e11f26360ef70bf5fb to 6b9c621bf49c7e5858a1289f62581319c3974d8f86cab3272c54d1574084c893. Only repeat-failure-under-a-declared-check (floor ["dotnet", "test"]) carries a floor, so it is the only prompt that moves. Only_the_tape_that_carries_an_operator_floor_gains_the_floor_block derives that: the scenarios carrying a floor must equal the scenarios whose prompt differs with the floor withheld; that scenario's prompt minus the block's own lines (taken from the decider's constants) must equal its floor-withheld prompt, so the block is the only thing that moved; and the corpus rendered with every floor withheld must digest to the previous pin, kept as PreOperatorFloorCorpusDigest, so the other 28 scenarios are byte-identical. The three older anchors, the exclusion sets and the scenario-count pin (29) are unchanged, since they already exclude that scenario. The remarks name the moved block.
  • Mutation, on the rebased head. Dropping the call to AppendOperatorAcceptanceFloor turns 13 cases red: 10 in SupervisorDeciderTests, the rehydrate seam test, the golden receipt and the digest pin (the constants test and the five omit cases stay green, since they do not depend on the block rendering). Others, each caught where expected: shell-style join instead of JSON (10 unit and 2 golden red), the default JSON encoder (1 red, the punctuation row), bound 400 to 399 (2), block moved above the criteria (1, only the position test sees it), a timeout literal instead of the constant (the constants test plus the digest), and the header reworded back to the old "no second chance ... before stopping" copy (only the digest, since the tests read the constants). One mutant survives by design: a guard that admits an empty list, because the rehydrate never hands the decider one.
  • Green on the rebased head (1a9055a91 plus this commit): SupervisorDeciderTests 210/210, SupervisorTurnServiceTests 74/74, SupervisorGoldenPromptFidelityTests 29/29, SupervisorDecisionEvalDeciderTests 5/5, SupervisorDecisionEvalTests 6/6, SupervisorTrajectoryEvalTests 49/49, and the full solution builds with 0 errors and no new warnings from the touched files.
  • Full unit project on the rebased head, run once: 11636 passed, 1 skipped (LocalRwxAtomicCreateTests.Cross_device_publication_is_typed_unsupported_and_never_copies_into_the_destination, an environment skip), 1 failed: ModelCredentialBrokerTests.A_closing_lease_gives_up_its_socket_before_its_port_so_a_worker_rebinding_mid_close_is_refused_and_its_socket_survives, an OS port-rebind race in the broker that this change does not touch. It passes 108/108 in isolation, five runs in a row, and an earlier full run of this branch (before the latest rework and rebase) had 0 failures.
  • After merge: the supervisor decision eval runs on main and scores the golden corpus under the new digest. No live signal is needed for this block's correctness, since the corpus does not depend on its presence.

The stop path runs the operator's floor argv
(SupervisorTurnContext.AcceptanceChecks) on every branch a stop ships,
300 s each. A failing floor withholds those branches and ends the run as
AcceptanceFailed, and a stop is final, so no turn is left to fix it. The
decider never read the field: it rendered only the free-text acceptance
criteria, and the system prompt mentions "the operator's floor" without
showing it. The brain drove the work blind to the one check that decides
the run.

BuildUserPrompt now renders the floor right after the criteria block,
only when the context carries one. The header says what runs and on
what, the timeout, and what a failure costs. It ends on "before
declaring success" rather than "before stopping": the model cannot see
the floor's verdict before it stops, and gave_up and ask_human stay open
on a run that cannot pass. "Every branch a stop ships" is deliberate, a
branchless stop records the floor as not graded. The closing line says
that repeating the floor as the stop payload's acceptance only runs it
twice, and that a subtask's own acceptance is a separate check, which is
the one way the model can see a verdict early. The timeout and the
outcome word are read from the grader's own constants so the copy cannot
drift, and the header and closing line are constants the tests use.

The argv is written as a JSON array of strings, not as shell text. The
grader spawns it with no shell, so a shell-looking line misleads:
`--filter Category=Unit|Category=Smoke` reads as a pipe, and a joined
line cannot tell ["sh", " ", "check.sh"] from two arguments. The array
uses WorkflowJson.InterpolatedText, the repo's way to write an array
into prompt text, so quote, pipe, ampersand, angle brackets and plus
stay as themselves while a newline stays a visible escape. The line is
bounded at 400 chars like the per-subtask check line. No redaction is
added: the floor is operator-configured and now reaches the model
provider verbatim, and the recording decorator masks the run's resolved
credential values only in what it stores.

The system prompt is untouched (its digest is frozen into the
qualification runtime manifest), as are the stopped-now recital
(precomposed from the tape and mirrored by the golden fixture) and
SupervisorQualityFacts (the floor is a run-level gate, not a unit-level
declared check).

GoldenPromptDigest is re-pinned on purpose (2dafcdc7 -> 6b9c621b). Only
repeat-failure-under-a-declared-check carries a floor, so it is the only
prompt that moves and the other 28 scenarios are byte-identical. The
receipt derives the mover set, requires the carrier's prompt minus the
block's own lines to equal its floor-withheld prompt, and requires the
corpus rendered with every floor withheld to digest to the previous pin.
The three older anchors already exclude that scenario and are unchanged.
@ppXD
ppXD force-pushed the fix/show-the-supervisor-its-acceptance-floor branch from cc7226a to f98f98d Compare September 30, 2026 13:43
@ppXD
ppXD merged commit 6450a52 into main Sep 30, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant