Conversation
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
The evidence directory recorded internal infrastructure identifiers and iteration artefacts that do not belong in the public repository. The catalog test now reads a trimmed rubric inventory (id, name, max score, definition hash) stored next to the test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
c86baec to
4bf4ff1
Compare
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|
The previous copy of the MAG scenarios, schemas, evaluators and prompts had been machine-normalised (flattened strings, 300-character lines, a stray Workflow docstring), so it matched neither upstream nor a readable style, and its manifest hashed upstream paths that no local file corresponded to. scripts/vendor_mag.py now regenerates manager_gym/_vendor/mag from a pinned upstream checkout. Files keep their upstream paths and bytes; only import statements change, by three mechanical rules. Five small shims stand in for MAG runtime modules that Ergon implements natively. MANIFEST.json records every file's kind and hash, and native code reaches upstream names only through manager_gym/upstream.py. Behaviour is unchanged: 92 of 93 old/new file pairs have identical ASTs once imports are ignored (the exception is the stakeholder factory, now a shim), and all 1281 rubric definitions, including callable sources, hash-match the upstream inventory. The judge prompt moves to a readable native template with a golden test, and the catalog test now checks callable rubric hashes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MAG previously accepted only one internal gateway host and defaulted every role to it. The sample factories now take a required model target, and inference goes straight through Ergon's normal model resolution. Inference limits move from a free-form metadata dict and scattered literals into a typed InferenceProfile carried on EpisodeConfig, so the manager, every work role and the judge read the same limits and each sample records them. Callers name their role instead of passing a temperature. The vLLM thinking-token budget is now opt-in, because other providers reject unknown request fields. The OpenAI-compatible API key is read through Settings rather than os.environ, and the example scripts take --model (or $ERGON_MAG_MODEL) and --thinking-token-budget. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
examples/manager_gym now holds only what a user runs: preflight, submit and reevaluate, with a README covering prerequisites, model targets, options, scoring, limits and the upstream citation. The live acceptance drivers (contract, pilot and catalog stages, cancellation and step-error checks) and their compose overlay move to tests/real_llm/manager_gym, where their imports of tests.fixtures are ordinary, and acceptance.main is split into helpers. docs/architecture/09_manager_gym.md is rewritten as the design reference: purpose and catalog size, abstractions, control flow, invariants, a parity section listing kept upstream quirks, deliberate native differences and rubric versions, model-call failure handling, and limits. It absorbs what was worth keeping from the port RFC, which is removed. Other architecture docs lose infrastructure-specific wording, and the four bugs this branch fixes move to docs/bugs/fixed with their run identifiers removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Actors are typed upstream configs (AIAgentConfig, HumanAgentConfig, StakeholderConfig) discriminated on agent_type, instead of dicts indexed by string keys; bindings and message maps are keyed by UUID. - Shared helpers replace copies: inherited_dependencies/dependency_graph (three dependency walks), JsonDecoded (five JSON-string validators), persist_message (two message writers) and snapshot_output (two snapshot writers). constants.py names the shared ids, paths and timeouts. - Inline prompt text moves to prompts.py as named templates, each marked as upstream or native wording. - Large functions are split: workers._work into noise draws, prompt builders and result assembly; project_native_state into task, message and roll-up projections; apply_action into a match with a case per action. - Loop closures are built by small factories, statuses compare against TaskStatus and Ergon's runtime status constants, scenario timelines are cached, EpisodeConfig validates its scenario and persona. - baseline_infer has an explicit signature and records only model errors; other failures now raise, so its slopcop exemption is gone. - Dead code removed: RequestEndWorkflowAction, unused ActionResult values, WorkInputs.completed_tasks, always-empty chunk returns, an unused model parameter and two unused imports. - Public classes and functions have docstrings. Output-schema models keep none, because the model sees them; their JSON schemas are unchanged. Checked against the previous commit with a harness over all 20 scenarios: snapshot hashes, manager observations across ticks, rubric validation contexts, every output schema, RandomV2 schemas, and 212 end-to-end work-role runs (system prompt, prompt, role and result) are identical. Only the work task payload's shape changes (the stakeholder's current weights now travel as their own field). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The native work roles had drifted from upstream MAG's prompts: human task prompts dropped the "sharp and focused" note, the expertise line and the step-by-step instruction; misunderstandings were a one-line paraphrase appended to the normal prompt instead of upstream's separate prompt; the time estimator used a shortened prompt; input resources were listed in a different format with a different empty message; and human execution notes recorded the model's process text instead of upstream's worker, style, experience, fatigue and quality lines plus challenges. prompts.py now holds these templates extracted byte for byte from upstream (core/workflow_agents/human_agent.py and ai_agent.py). This changes what the simulated workers see, so it is a behaviour change, kept separate from the preceding refactor: across all 20 scenarios only work-role prompts (212 of 212 runs) and human execution notes (50 runs) differ from the refactor's output. The remaining deliberate differences are listed in the architecture doc's parity section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rubric version 1 reproduces upstream MAG's scoring. Version 2, the new default, corrects three upstream bugs through a small overlay in rubric_versions.py, leaving the vendored rubrics untouched: - ICAAP quality/seeking_sourcing counts web-search tool calls that no MAG runner gives workers, never requests tool usage, and scores a 0-1 fraction against a maximum of 3, so it always scored 0. Version 2 drops it; this is the only change that affects utility. - stakeholder_management/response_latency_adherence declares a maximum of 20 for a function that caps at 8; version 2 uses 8. - operational_efficiency/agent_utilization_efficiency had a copied description and never requested the agent states it reads. Slugs are stable across versions. The version is a field of MAGRubric and MAGCriterion, a make_manager_gym_sample/make_snapshot_reevaluation_sample argument, a --rubric-version flag on the examples, and is recorded in criterion, evaluation and sample metadata. rubric.py also gets docstrings, named upstream thresholds (PASS_FRACTION, CATEGORICAL_SCORE), a cached slug lookup that raises a clear error instead of StopIteration, a documented dispatcher for upstream's three rubric signatures, explicit errors for a missing snapshot hash, float-only score normalisation, and a note that upstream's declared aggregation strategies are deliberately not applied. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- RuntimeGraphRepository: lock_sample, dependencies_complete, cancelled_for_invalidation and replace_dependencies get docstrings and use the runtime status constants. replace_dependencies reports a missing target task as NodeNotFoundError instead of a DanglingEdgeError with an invented edge id. Unused imports and an unused local go. - WorkerContext.wait_for_task waits on NON_AUTONOMOUS_STATUSES and TaskCompletedEvent.name instead of repeating their values, and it and refine_task document what they do now. - The engine's NullPool comment and the per-sample lock note read as design notes, and worker_execute no longer imports SampleContextEvent twice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Shared helpers move into ergon_core.test_support.runtime_harness: SessionContext, make_session, seed_parent, make_task, task_service and a RuntimeHarness handle with spawn/node/complete. A graph_runtime fixture in ergon_core's runtime tests and the MAG tests replaces the preport fixture, so no test module imports another's private helpers or fixtures. - The "pre-port proof" module becomes test_dynamic_task_runtime.py with behaviour names. The strict xfail about uncheckpointed SDK work is gone, and the test pinning non-idempotent message retries becomes a positive test that a retried message with an idempotency key is stored once. - MAG unit tests move to ergon_builtins/tests/unit/benchmarks/manager_gym, next to the other benchmarks. - The contract fixture names its tasks and roles instead of indexing plans, raises when cancellation fails to stop the manager, and no longer calls remove_work with a stale argument. The step-error driver checks the exception type and the empty-message fallback separately. - tests/fixtures/mag_preport.py is removed: its continuation worker was unused and called infer with an outdated signature; the evaluation test that borrowed its criterion defines a small one, without lazy imports. - The Postgres integration test is renamed test_runtime_concurrency_postgres and its comments describe the guarantees rather than their history. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Files added or changed by this branch are now clean under ruff's pyflakes, pycodestyle, bugbear, isort and pyupgrade rules (UP040 aside: turning the Annotated dependency aliases into PEP 695 `type` aliases changes how pydantic resolves them). That removes unused imports, sorts imports, wraps long lines, chains four re-raised ValueErrors to their cause, drops a Python 3.10 StrEnum fallback, and makes two job contract modules' event re-exports explicit through __all__. CI, package.json and CLAUDE.md now run ruff over examples/ and scripts/ as well. Turning these rules on for the whole repository is left to a separate change, since most remaining findings are in code outside this one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
E2E smoke —
|
E2E smoke —
|
E2E smoke —
|

















































































Ports all 20 pinned Manager Agent Gym scenarios into native Ergon samples. Managers, AI workers, simulated human workers and stakeholders use Ergon task assignment, dependency propagation, messaging, E2B sandboxes and evaluator jobs. All model roles and judges use the standing internal Qwen deployment.
The runtime changes extend existing owners for task admission and pending edits, actor binding, durable continuation, message idempotency, cancellation, sandbox cleanup and incomplete evaluations. Large checkpoints use existing resource/blob retention, and the dashboard reloads complete saved evaluations after small update notifications. There is no separate benchmark execution scheduler. Blob publication preserves normal file permissions, and graph locking permits concurrent evaluation writes. CI uses a shared, bounded Inngest
/healthreadiness probe instead of an event lookup.Start review with the annotated implementation map, architecture and acceptance report. The runbook covers reproduction, frozen-snapshot reevaluation, export and cleanup. The implementation record documents fidelity decisions for scheduling, clocks and observations.
Validation:
4e0be42a: 1,077 Python unit tests passed (one skipped, one expected failure), 138 dashboard tests passed, no generated-contract drift, and backend/frontend lint and type checks passed. Complexity and suppression budgets passed; slopcop reported zero errors and 661 warnings.c86baec4: CI Fast passed, including all 46 PostgreSQL/Inngest integration tests and dashboard end-to-end tests. All three canonical benchmark smoke tests passed: ResearchRubrics, MiniF2F and SWE-bench Verified.The live benchmark evidence covers this evaluation profile. It does not establish full-timeline coverage, upstream statistical parity, RL readiness or zero CPU/EBS/E2B billing.