Skip to content

Port Manager Agent Gym to native Ergon composition - #134

Draft
cm2435 wants to merge 33 commits into
devfrom
feature/manager-gym-native
Draft

cm2435 wants to merge 33 commits into
devfrom
feature/manager-gym-native

Conversation

@cm2435

@cm2435 cm2435 commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Ports all 20 pinned Manager Agent Gym scenarios into native Ergon samples. Managers, AI workers, simulated human workers and stakeholders use Ergon task assignment, dependency propagation, messaging, E2B sandboxes and evaluator jobs. All model roles and judges use the standing internal Qwen deployment.

The runtime changes extend existing owners for task admission and pending edits, actor binding, durable continuation, message idempotency, cancellation, sandbox cleanup and incomplete evaluations. Large checkpoints use existing resource/blob retention, and the dashboard reloads complete saved evaluations after small update notifications. There is no separate benchmark execution scheduler. Blob publication preserves normal file permissions, and graph locking permits concurrent evaluation writes. CI uses a shared, bounded Inngest /health readiness probe instead of an event lookup.

Start review with the annotated implementation map, architecture and acceptance report. The runbook covers reproduction, frozen-snapshot reevaluation, export and cleanup. The implementation record documents fidelity decisions for scheduling, clocks and observations.

Validation:

  • All 20 scenarios passed integration acceptance on one frozen seed-0, 50-decision profile: 1,230 terminal criteria, 505 native tasks, 1,562 messages and 6,380 resources. Nine bounded model-work failures remain native failures with complete numeric evaluations; 11 samples are Completed and nine are Failed. Acceptance does not impose a utility threshold.
  • Live native contract and delayed-cancellation proofs cover dependency order, repeated-human fatigue, pending edits, messages, failure propagation, artifacts, no late spawn and sandbox cleanup.
  • Export verification covered 15,042 files and 14,331 referenced blobs. An isolated PostgreSQL restore passed. Owned sandboxes and temporary VM resources were removed; the standing model deployment remained unchanged.
  • Fresh local checks at 4e0be42a: 1,077 Python unit tests passed (one skipped, one expected failure), 138 dashboard tests passed, no generated-contract drift, and backend/frontend lint and type checks passed. Complexity and suppression budgets passed; slopcop reported zero errors and 661 warnings.
  • GitHub validation at c86baec4: CI Fast passed, including all 46 PostgreSQL/Inngest integration tests and dashboard end-to-end tests. All three canonical benchmark smoke tests passed: ResearchRubrics, MiniF2F and SWE-bench Verified.

The live benchmark evidence covers this evaluation profile. It does not establish full-timeline coverage, upstream statistical parity, RL readiness or zero CPU/EBS/E2B billing.

@cm2435
cm2435 changed the base branch from main to dev September 13, 2026 20:50
@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

No PNG screenshots were uploaded for this leg. See screenshots/pr-134 for the uploaded placeholder.

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

Screenshots pushed to screenshots/pr-134.

minif2f/3218eed0-1309-46f1-a58f-af5f1f0d8345-activity-stack.png
minif2f/3218eed0-1309-46f1-a58f-af5f1f0d8345-sad.png
minif2f/3218eed0-1309-46f1-a58f-af5f1f0d8345-visual-debugger-full.png
minif2f/5c7a000f-b51c-4538-84fd-b11826f12eeb-activity-stack.png
minif2f/5c7a000f-b51c-4538-84fd-b11826f12eeb-happy.png
minif2f/5c7a000f-b51c-4538-84fd-b11826f12eeb-visual-debugger-full.png
minif2f/600dad12-1524-4347-9fcb-aa6d49b11d9e-activity-stack.png
minif2f/600dad12-1524-4347-9fcb-aa6d49b11d9e-happy.png
minif2f/600dad12-1524-4347-9fcb-aa6d49b11d9e-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

Screenshots pushed to screenshots/pr-134.

swebench-verified/56e3c923-f644-4696-b423-160d74ecf3b3-activity-stack.png
swebench-verified/56e3c923-f644-4696-b423-160d74ecf3b3-sad.png
swebench-verified/56e3c923-f644-4696-b423-160d74ecf3b3-visual-debugger-full.png
swebench-verified/5f368f8e-373f-4127-a91c-a8185d09d77a-activity-stack.png
swebench-verified/5f368f8e-373f-4127-a91c-a8185d09d77a-happy.png
swebench-verified/5f368f8e-373f-4127-a91c-a8185d09d77a-visual-debugger-full.png
swebench-verified/d0108856-ede9-410e-860d-422a62d0089d-activity-stack.png
swebench-verified/d0108856-ede9-410e-860d-422a62d0089d-happy.png
swebench-verified/d0108856-ede9-410e-860d-422a62d0089d-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

Screenshots pushed to screenshots/pr-134.

researchrubrics/317da642-1def-4244-8616-64500d53883f-activity-stack.png
researchrubrics/317da642-1def-4244-8616-64500d53883f-sad.png
researchrubrics/317da642-1def-4244-8616-64500d53883f-visual-debugger-full.png
researchrubrics/c5cbd9d1-d12a-4495-8f33-ce0f1f0a9073-activity-stack.png
researchrubrics/c5cbd9d1-d12a-4495-8f33-ce0f1f0a9073-happy.png
researchrubrics/c5cbd9d1-d12a-4495-8f33-ce0f1f0a9073-visual-debugger-full.png
researchrubrics/c84f4658-61bc-4776-8bf6-08150dd00faf-activity-stack.png
researchrubrics/c84f4658-61bc-4776-8bf6-08150dd00faf-happy.png
researchrubrics/c84f4658-61bc-4776-8bf6-08150dd00faf-visual-debugger-full.png

The evidence directory recorded internal infrastructure identifiers and
iteration artefacts that do not belong in the public repository. The
catalog test now reads a trimmed rubric inventory (id, name, max score,
definition hash) stored next to the test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cm2435
cm2435 force-pushed the feature/manager-gym-native branch from c86baec to 4bf4ff1 Compare September 28, 2026 19:25
@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

Screenshots pushed to screenshots/pr-134.

swebench-verified/09564537-5ba7-4c5f-b148-c021f25ef494-activity-stack.png
swebench-verified/09564537-5ba7-4c5f-b148-c021f25ef494-happy.png
swebench-verified/09564537-5ba7-4c5f-b148-c021f25ef494-visual-debugger-full.png
swebench-verified/288bbcda-97e1-48a3-84b0-8aae87ccf7df-activity-stack.png
swebench-verified/288bbcda-97e1-48a3-84b0-8aae87ccf7df-sad.png
swebench-verified/288bbcda-97e1-48a3-84b0-8aae87ccf7df-visual-debugger-full.png
swebench-verified/90a04e11-8d4e-43dd-a058-7349153f2aaa-activity-stack.png
swebench-verified/90a04e11-8d4e-43dd-a058-7349153f2aaa-happy.png
swebench-verified/90a04e11-8d4e-43dd-a058-7349153f2aaa-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

Screenshots pushed to screenshots/pr-134.

researchrubrics/1b7caeaf-46a1-4881-a493-2a7f93796c18-activity-stack.png
researchrubrics/1b7caeaf-46a1-4881-a493-2a7f93796c18-sad.png
researchrubrics/1b7caeaf-46a1-4881-a493-2a7f93796c18-visual-debugger-full.png
researchrubrics/5a57d335-5329-4c40-b8db-3fdf7b490a81-activity-stack.png
researchrubrics/5a57d335-5329-4c40-b8db-3fdf7b490a81-happy.png
researchrubrics/5a57d335-5329-4c40-b8db-3fdf7b490a81-visual-debugger-full.png
researchrubrics/8c78264e-0ab8-4c90-8fff-336675f0971e-activity-stack.png
researchrubrics/8c78264e-0ab8-4c90-8fff-336675f0971e-happy.png
researchrubrics/8c78264e-0ab8-4c90-8fff-336675f0971e-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

Screenshots pushed to screenshots/pr-134.

minif2f/73e6fc65-ab8a-4f88-a22d-a02aa948adf2-activity-stack.png
minif2f/73e6fc65-ab8a-4f88-a22d-a02aa948adf2-happy.png
minif2f/73e6fc65-ab8a-4f88-a22d-a02aa948adf2-visual-debugger-full.png
minif2f/b2d71fda-8ee7-4fb4-80cb-c24cc45e1cfc-activity-stack.png
minif2f/b2d71fda-8ee7-4fb4-80cb-c24cc45e1cfc-sad.png
minif2f/b2d71fda-8ee7-4fb4-80cb-c24cc45e1cfc-visual-debugger-full.png
minif2f/cffbb835-18da-4605-b709-e0ee13eddc00-activity-stack.png
minif2f/cffbb835-18da-4605-b709-e0ee13eddc00-happy.png
minif2f/cffbb835-18da-4605-b709-e0ee13eddc00-visual-debugger-full.png

cm2435-hcomp and others added 9 commits September 28, 2026 20:53
The previous copy of the MAG scenarios, schemas, evaluators and prompts had
been machine-normalised (flattened strings, 300-character lines, a stray
Workflow docstring), so it matched neither upstream nor a readable style,
and its manifest hashed upstream paths that no local file corresponded to.

scripts/vendor_mag.py now regenerates manager_gym/_vendor/mag from a pinned
upstream checkout. Files keep their upstream paths and bytes; only import
statements change, by three mechanical rules. Five small shims stand in for
MAG runtime modules that Ergon implements natively. MANIFEST.json records
every file's kind and hash, and native code reaches upstream names only
through manager_gym/upstream.py.

Behaviour is unchanged: 92 of 93 old/new file pairs have identical ASTs once
imports are ignored (the exception is the stakeholder factory, now a shim),
and all 1281 rubric definitions, including callable sources, hash-match the
upstream inventory. The judge prompt moves to a readable native template
with a golden test, and the catalog test now checks callable rubric hashes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MAG previously accepted only one internal gateway host and defaulted every
role to it. The sample factories now take a required model target, and
inference goes straight through Ergon's normal model resolution.

Inference limits move from a free-form metadata dict and scattered literals
into a typed InferenceProfile carried on EpisodeConfig, so the manager,
every work role and the judge read the same limits and each sample records
them. Callers name their role instead of passing a temperature. The vLLM
thinking-token budget is now opt-in, because other providers reject unknown
request fields.

The OpenAI-compatible API key is read through Settings rather than
os.environ, and the example scripts take --model (or $ERGON_MAG_MODEL) and
--thinking-token-budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
examples/manager_gym now holds only what a user runs: preflight, submit and
reevaluate, with a README covering prerequisites, model targets, options,
scoring, limits and the upstream citation. The live acceptance drivers
(contract, pilot and catalog stages, cancellation and step-error checks) and
their compose overlay move to tests/real_llm/manager_gym, where their
imports of tests.fixtures are ordinary, and acceptance.main is split into
helpers.

docs/architecture/09_manager_gym.md is rewritten as the design reference:
purpose and catalog size, abstractions, control flow, invariants, a parity
section listing kept upstream quirks, deliberate native differences and
rubric versions, model-call failure handling, and limits. It absorbs what
was worth keeping from the port RFC, which is removed. Other architecture
docs lose infrastructure-specific wording, and the four bugs this branch
fixes move to docs/bugs/fixed with their run identifiers removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Actors are typed upstream configs (AIAgentConfig, HumanAgentConfig,
  StakeholderConfig) discriminated on agent_type, instead of dicts indexed
  by string keys; bindings and message maps are keyed by UUID.
- Shared helpers replace copies: inherited_dependencies/dependency_graph
  (three dependency walks), JsonDecoded (five JSON-string validators),
  persist_message (two message writers) and snapshot_output (two snapshot
  writers). constants.py names the shared ids, paths and timeouts.
- Inline prompt text moves to prompts.py as named templates, each marked as
  upstream or native wording.
- Large functions are split: workers._work into noise draws, prompt builders
  and result assembly; project_native_state into task, message and roll-up
  projections; apply_action into a match with a case per action.
- Loop closures are built by small factories, statuses compare against
  TaskStatus and Ergon's runtime status constants, scenario timelines are
  cached, EpisodeConfig validates its scenario and persona.
- baseline_infer has an explicit signature and records only model errors;
  other failures now raise, so its slopcop exemption is gone.
- Dead code removed: RequestEndWorkflowAction, unused ActionResult values,
  WorkInputs.completed_tasks, always-empty chunk returns, an unused model
  parameter and two unused imports.
- Public classes and functions have docstrings. Output-schema models keep
  none, because the model sees them; their JSON schemas are unchanged.

Checked against the previous commit with a harness over all 20 scenarios:
snapshot hashes, manager observations across ticks, rubric validation
contexts, every output schema, RandomV2 schemas, and 212 end-to-end work-role
runs (system prompt, prompt, role and result) are identical. Only the work
task payload's shape changes (the stakeholder's current weights now travel
as their own field).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The native work roles had drifted from upstream MAG's prompts: human task
prompts dropped the "sharp and focused" note, the expertise line and the
step-by-step instruction; misunderstandings were a one-line paraphrase
appended to the normal prompt instead of upstream's separate prompt; the
time estimator used a shortened prompt; input resources were listed in a
different format with a different empty message; and human execution notes
recorded the model's process text instead of upstream's worker, style,
experience, fatigue and quality lines plus challenges.

prompts.py now holds these templates extracted byte for byte from upstream
(core/workflow_agents/human_agent.py and ai_agent.py). This changes what the
simulated workers see, so it is a behaviour change, kept separate from the
preceding refactor: across all 20 scenarios only work-role prompts (212 of
212 runs) and human execution notes (50 runs) differ from the refactor's
output. The remaining deliberate differences are listed in the architecture
doc's parity section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rubric version 1 reproduces upstream MAG's scoring. Version 2, the new
default, corrects three upstream bugs through a small overlay in
rubric_versions.py, leaving the vendored rubrics untouched:

- ICAAP quality/seeking_sourcing counts web-search tool calls that no MAG
  runner gives workers, never requests tool usage, and scores a 0-1 fraction
  against a maximum of 3, so it always scored 0. Version 2 drops it; this is
  the only change that affects utility.
- stakeholder_management/response_latency_adherence declares a maximum of
  20 for a function that caps at 8; version 2 uses 8.
- operational_efficiency/agent_utilization_efficiency had a copied
  description and never requested the agent states it reads.

Slugs are stable across versions. The version is a field of MAGRubric and
MAGCriterion, a make_manager_gym_sample/make_snapshot_reevaluation_sample
argument, a --rubric-version flag on the examples, and is recorded in
criterion, evaluation and sample metadata.

rubric.py also gets docstrings, named upstream thresholds (PASS_FRACTION,
CATEGORICAL_SCORE), a cached slug lookup that raises a clear error instead
of StopIteration, a documented dispatcher for upstream's three rubric
signatures, explicit errors for a missing snapshot hash, float-only score
normalisation, and a note that upstream's declared aggregation strategies
are deliberately not applied.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- RuntimeGraphRepository: lock_sample, dependencies_complete,
  cancelled_for_invalidation and replace_dependencies get docstrings and use
  the runtime status constants. replace_dependencies reports a missing
  target task as NodeNotFoundError instead of a DanglingEdgeError with an
  invented edge id. Unused imports and an unused local go.
- WorkerContext.wait_for_task waits on NON_AUTONOMOUS_STATUSES and
  TaskCompletedEvent.name instead of repeating their values, and it and
  refine_task document what they do now.
- The engine's NullPool comment and the per-sample lock note read as
  design notes, and worker_execute no longer imports SampleContextEvent twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Shared helpers move into ergon_core.test_support.runtime_harness:
  SessionContext, make_session, seed_parent, make_task, task_service and a
  RuntimeHarness handle with spawn/node/complete. A graph_runtime fixture in
  ergon_core's runtime tests and the MAG tests replaces the preport fixture,
  so no test module imports another's private helpers or fixtures.
- The "pre-port proof" module becomes test_dynamic_task_runtime.py with
  behaviour names. The strict xfail about uncheckpointed SDK work is gone,
  and the test pinning non-idempotent message retries becomes a positive
  test that a retried message with an idempotency key is stored once.
- MAG unit tests move to ergon_builtins/tests/unit/benchmarks/manager_gym,
  next to the other benchmarks.
- The contract fixture names its tasks and roles instead of indexing plans,
  raises when cancellation fails to stop the manager, and no longer calls
  remove_work with a stale argument. The step-error driver checks the
  exception type and the empty-message fallback separately.
- tests/fixtures/mag_preport.py is removed: its continuation worker was
  unused and called infer with an outdated signature; the evaluation test
  that borrowed its criterion defines a small one, without lazy imports.
- The Postgres integration test is renamed test_runtime_concurrency_postgres
  and its comments describe the guarantees rather than their history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Files added or changed by this branch are now clean under ruff's pyflakes,
pycodestyle, bugbear, isort and pyupgrade rules (UP040 aside: turning the
Annotated dependency aliases into PEP 695 `type` aliases changes how
pydantic resolves them). That removes unused imports, sorts imports, wraps
long lines, chains four re-raised ValueErrors to their cause, drops a
Python 3.10 StrEnum fallback, and makes two job contract modules' event
re-exports explicit through __all__.

CI, package.json and CLAUDE.md now run ruff over examples/ and scripts/ as
well. Turning these rules on for the whole repository is left to a
separate change, since most remaining findings are in code outside this one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

E2E smoke — minif2f

Screenshots pushed to screenshots/pr-134.

minif2f/1f5e4055-b4fc-492a-92d0-bd223af97f9c-activity-stack.png
minif2f/1f5e4055-b4fc-492a-92d0-bd223af97f9c-sad.png
minif2f/1f5e4055-b4fc-492a-92d0-bd223af97f9c-visual-debugger-full.png
minif2f/4b31a970-d0da-450e-9f01-27fdeeade973-activity-stack.png
minif2f/4b31a970-d0da-450e-9f01-27fdeeade973-happy.png
minif2f/4b31a970-d0da-450e-9f01-27fdeeade973-visual-debugger-full.png
minif2f/e432b971-850d-49c3-8a8d-1433c92a119f-activity-stack.png
minif2f/e432b971-850d-49c3-8a8d-1433c92a119f-happy.png
minif2f/e432b971-850d-49c3-8a8d-1433c92a119f-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — swebench-verified

Screenshots pushed to screenshots/pr-134.

swebench-verified/9c9892e1-94e6-4dc4-80aa-6ef57b229ae8-activity-stack.png
swebench-verified/9c9892e1-94e6-4dc4-80aa-6ef57b229ae8-sad.png
swebench-verified/9c9892e1-94e6-4dc4-80aa-6ef57b229ae8-visual-debugger-full.png
swebench-verified/e79409a2-afb7-4a56-a1ca-f9c1799e9e9b-activity-stack.png
swebench-verified/e79409a2-afb7-4a56-a1ca-f9c1799e9e9b-happy.png
swebench-verified/e79409a2-afb7-4a56-a1ca-f9c1799e9e9b-visual-debugger-full.png
swebench-verified/f5d0785f-3101-4eab-a846-310e89ce2728-activity-stack.png
swebench-verified/f5d0785f-3101-4eab-a846-310e89ce2728-happy.png
swebench-verified/f5d0785f-3101-4eab-a846-310e89ce2728-visual-debugger-full.png

@github-actions

Copy link
Copy Markdown

E2E smoke — researchrubrics

Screenshots pushed to screenshots/pr-134.

researchrubrics/770adbbc-0f2e-40f3-a157-eb55ea7e67d7-activity-stack.png
researchrubrics/770adbbc-0f2e-40f3-a157-eb55ea7e67d7-sad.png
researchrubrics/770adbbc-0f2e-40f3-a157-eb55ea7e67d7-visual-debugger-full.png
researchrubrics/871c2070-6431-4494-9985-3bcf1ad03c58-activity-stack.png
researchrubrics/871c2070-6431-4494-9985-3bcf1ad03c58-happy.png
researchrubrics/871c2070-6431-4494-9985-3bcf1ad03c58-visual-debugger-full.png
researchrubrics/caa5c3d7-e052-489c-8128-904567c0f21d-activity-stack.png
researchrubrics/caa5c3d7-e052-489c-8128-904567c0f21d-happy.png
researchrubrics/caa5c3d7-e052-489c-8128-904567c0f21d-visual-debugger-full.png

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants