Discover capabilities. Verify execution. Control collaboration. A verified orchestration layer over coding-agent CLIs — your agents work with receipts.
简体中文 → README.zh-CN.md
dual-agent is an engineering layer between your application and the coding-agent
CLIs it drives (Claude Code, Codex CLI, Gemini CLI, …). It discovers what is
installed, verifies what each runtime can actually prove, admits it to a
verified pool, and orchestrates architect → coder → tester → reviewer work
under explicit budgets and loop protection — with provenance on every result.
Agent runtime ≠ agent orchestration. Runtimes execute; this project engineers the layer above them. It is not a chatbot, a model provider, a single-runtime wrapper, or a distributed agent network. No network transport, no credentials touched, zero runtime dependencies (pure standard library).
Name map: GitHub repository
runtime-neutral-agent-engineering· PyPI distributiondual-agent-development· importdual_agent· console scriptdual-agent. One product, one version truth (dual_agent.__version__).
No runtime, no login, no API key, no network. A fresh clone is enough:
git clone https://github.com/Tsubasa-Kaede/runtime-neutral-agent-engineering.git
cd runtime-neutral-agent-engineering
python examples/offline_mock_run.pyExpected output — one closed, secret-free JSON summary:
{"path": "FOUR_STAGE", "status": "SUCCESS", "stages": ["architect", "coder", "tester", "reviewer"], ...}This runs the real production facade end to end with mock adapters — the same
engine, honestly labeled OFFLINE. Terminal demo GIF (render once with
vhs assets/demo.tape):
Python >= 3.10, zero runtime dependencies, no clone needed:
pip install dual-agent-development
dual-agent --version
# or run without installing (uv):
uvx --from dual-agent-development dual-agent --versionOther ways in: editable install (pip install -e . from a checkout) or the
one-command bootstrap (python scripts/bootstrap.py — installs this project
only, never touches runtimes, secrets, or system config). Examples like the
offline demo ship with the repository, not the wheel.
- Runtime Discovery — is a runtime present at all?
- Runtime Validation — gated qualification runs (G1–G14) producing real evidence
- Capability-based Selection — selection by proven capability, never by name
- Agent Orchestration — architect → coder → tester → reviewer stage chains
- Structured Collaboration — validated packets over an append-only ledger
- Budget Control — invocation slots reserved before every call
- LoopGuard — duplicate / repeated-failure / cycle protection before spend
- Provenance — every validation result carries
OFFLINEorREALevidence - Security Boundary — no-secrets contract, content scanning, protected paths
1. Cross-runtime second opinion. You live in Claude Code but want Codex CLI or Gemini CLI on the same task. Adapters normalize every runtime to one contract, so the four-stage pipeline runs over whatever you qualified — your orchestration code never names a vendor.
2. Agent work with receipts. You need to know the work was bounded and
verified, not just told it succeeded. Every invocation is budgeted before it
happens, LoopGuard rejects duplicates and cycles before any spend, the ledger
is append-only, and every result carries provenance — REAL only with
real-call evidence, OFFLINE honestly labeled otherwise. No silent fallbacks,
no fabricated success words.
3. Vendor-neutral agent tooling. You are building a tool and refuse to
couple it to one runtime. Implement the six-method ExternalAgentAdapter
contract and your runtime plugs into discovery, qualification, and
orchestration without touching the engine.
dual-agent qualify # the ONLY command that qualifies
dual-agent run "Add a slug helper and its test" # reads persisted evidence
dual-agent run --observe "Add a slug helper and its test" # + execution events on stderrqualifyruns the gated G1–G14 qualification over discovered runtimes and persistsVERIFIED+REALevidence under~/.dual-agent/qualification/. Real model calls requireRUN_REAL_PROVIDER_TESTS=1; Offline results are reported honestly and never persisted — Offline validation is not REAL validation. Per-call progress streams to stderr; each call defaults to a 300-second bound (--timeout-seconds).runonly reads persisted evidence. With no evidence it exits2with a machine-readable reason (NO_EVIDENCE_NO_QUALIFIER) — it never auto-qualifies and never falls back.
Modes (--mode): OFF returns the delegated empty result — never silently
runs; AUTO (default) classifies the task and routes SIMPLE/MEDIUM to the
single-agent path, COMPLEX to the dual-agent path; ON forces the dual-agent
path. Embedding applications can inject a pre-configured facade directly
(cli.main._facade = my_facade).
Multi-Agent Collaboration Cockpit (V3.2) — dual-agent cockpit TASK --step ROLE=RUNTIME_ID [--step ...] composes a fixed sequential multi-agent run;
every referenced runtime must hold persisted VERIFIED qualification
evidence. Exit contract: 0 COMPLETED / 2 FAILED or user error / 3
ABORTED / 4 PARKED; exactly one machine JSON line on stdout, human
diagnostics on stderr.
First-run funnel (2.6.0) — dual-agent cockpit with no arguments, run
on an interactive terminal with the [tui] extra installed
(pip install dual-agent-development[tui]), opens the composition funnel
instead of erroring: type the task, review the default collaboration plan
(the first composition binds verified runtimes in canonical sorted
runtime-id order; later default compositions in the same session may order
seats deterministically from observed invocation counts, with the canonical
tie-break, and every reordering is disclosed for confirmation — 2.10.0;
2 runtimes map to architect + coder, 3 add a tester, 4 add a reviewer),
press Enter to start.
With fewer than two verified runtimes the screen states the blocked reason
and the qualify command to run. The explicit --step developer path
keeps its exact semantics and never merges with the funnel: piped or
redirected output, --json, or a missing Textual install all keep the
pre-2.6.0 behavior and byte-identical errors.
Live Collaboration Cockpit (2.7.0) — the TUI cockpit is now a live
control surface for a running collaboration: a horizontal pipeline renders
every agent slot with true connection semantics (──→ observed HANDOFF vs
┄┄→ planned adjacency), arrow keys select, Enter expands one agent's
bounded detail window, and L switches the whole display between English
and Chinese. While the run is live you can steer it directly: typing any
text opens the persistent steering composer (no prefix key), tab toggles
the revision target (next invocation vs submission prompt), Enter submits
the revision through the same control plane (p pause / r resume /
a abort / t trace / q quit), and the composer stays open for
back-to-back revisions — every submission answers with an honest receipt
(accepted ≠ applied: applied facts land in the trace view, never claimed
on the main screen). t opens the read-only observation trace: the
control journal, revision queue, and per-invocation events of the whole
run. Rendering performance kept pace with the richer surface — projection
is change-detected and zone-scoped, and typing touches only the dock line
(zero full projections per keypress).
Conversational Collaboration Cockpit (2.8.0) — the cockpit is now a
multi-run conversation: when a collaboration finishes, type the next task
and the same session starts the next run (divider lines separate runs in
the log; each run gets fresh ids, facts, and control state — quitting
without typing behaves exactly as before, and the exit code still maps to
the last run's outcome). New session commands: /again reloads the last
task through the same pool and fingerprint gates, /compose reopens
runtime selection, /runs lists every run this session, and /new
resets the display history (engine facts stay untouched). Compositions
with multiple groups render side by side with the same true connection
markers (──→ observed vs ┄┄→ planned), g on the selection screen
assigns members to groups, and H engages an explicit horizontal-scroll
mode for oversized layouts. Agent focus deepened: d opens a bounded
right-hand detail sidebar on wide terminals, and x (or Enter) on an
expanded slot jumps into the trace pre-focused on that agent. /runs
also appends a per-run detail block — composition, outcome, a result
preview, and a usage summary that sums only reported-known tokens while
counting unknown/unsupported honestly. The run boundary is pinned by
tests: no text from an earlier run ever enters a later run's prompts.
Experimental foundations: Context models & deterministic routing (2.9.0) — the stable product core is unchanged: runtime-neutral collaboration orchestration over coding-agent CLIs, the qualification chain and verified runtime pool, remote collaboration across a process boundary, the collaboration cockpit (TUI and the non-TTY contract), and the existing execution and control behavior all ship exactly as in 2.8.0. This version adds two experimental capability lines, delivered as frozen semantic modules with zero production wiring. The context trilogy: a scoped context item model (validity and provenance pinned), a collaboration memory record model (completed runs only, content-addressed), and a deterministic context compiler (character budgets, honest token tri-stating) — semantic models delivered; production execution wiring remains deferred. ORCH-5: deterministic default-composition routing, a pure routing projection that consumes per-runtime usage facts with a canonical tie-break — integrated behind an explicitly disabled switch in 2.9.0, activated by default in 2.10.0. Deferred on purpose: context production wiring, memory persistence, query/target/capabilities surfaces, and token/cost benchmarking. No cost or token savings are claimed; usage telemetry stays honestly three-state (KNOWN / UNKNOWN / UNSUPPORTED).
Deterministic default-composition routing, active (2.10.0) — the
ORCH-5 routing projection now drives default compositions by default. It
is deterministic and verified-runtime-aware: the candidate set is always
the VERIFIED pool (qualification unchanged), roles stay on their template
positions, and ordering consumes only invocation facts already observed
in the current session (usage-fact-driven, invocation-aware). The first
composition in a session is identical to the canonical sorted order —
empty evidence equals canonical ordering — and later default compositions
may reorder seats by observed invocation counts with the canonical
tie-break. Every reordering is disclosure-gated: the composition
fingerprint gate surfaces any change with honest reasons and requires
explicit confirmation before a run starts. Explicit compositions
(--step, the selection screen) never route — user authority passes
straight through. Usage telemetry stays honestly three-state
(KNOWN / UNKNOWN / UNSUPPORTED); UNKNOWN is never treated as zero, and no
token, price, or latency is ever inferred. No cost or token savings are
claimed; this is deterministic routing, not an optimizer.
They are strong tools for building LLM-chaining applications. This project solves a different problem — engineering discipline over coding-agent CLIs that already exist on your machine:
| Typical orchestration frameworks | dual-agent | |
|---|---|---|
| What is orchestrated | LLM API calls you wire up yourself | external coding-agent CLIs, via adapters |
| Runtime coupling | often one provider or SDK | runtime-neutral: nothing names a vendor |
| Admission | configure and go | gated G1–G14 qualification, VERIFIED + REAL evidence only |
| Result claims | framework-reported | provenance on every envelope; REAL refused without real-call evidence |
| Failure behavior | fallbacks and retries are common features | no fallback, no silent success — closed failure vocabulary |
| Dependencies | heavy SDK stacks | pure standard library, zero runtime dependencies |
| Transport | often cloud/network | local process boundary only; no network transport |
Use them together if you like: this layer does not replace your app framework — it sits between your application and the agent CLIs.
flowchart TD
T[Task] --> MG["Mode Gate: OFF / AUTO / ON"]
MG --> CL["Classifier: SIMPLE / MEDIUM / COMPLEX / UNRESOLVED"]
subgraph VP["Verified path (production stack)"]
D["Runtime Discovery"] --> H["Runtime Health"]
H --> Q["Qualification G1-G14 (gated)"]
Q --> V["Verification: VERIFIED + REAL"]
V --> ADM["Verified Runtime Pool admission"]
ADM --> SEL["Verified selection (score-less)"]
end
subgraph RP["ReadyPool path (classic engine)"]
H2["Runtime Health"] --> CAP["Capability Registry"]
CAP --> POOL["ReadyPool"]
POOL --> SSE["Scored selection"]
end
CL --> VP
CL --> RP
SEL --> EX["Execution: architect - coder - tester - reviewer"]
SSE --> EX
EX --> G["Per-invoke gates: Handoff - LoopGuard - Budget reserve - Invoke"]
G --> OUT["Closed, secret-free summary"]
Load-bearing invariant: the verified path never silently borrows the ReadyPool.
An empty verified selection normalizes to NO_CAPABLE_AGENT instead of
consulting the ready-pool registry. The five distinctions the engine never
blurs: Discovery ≠ Health, Health ≠ Qualification, Qualification ≠
Verification, Verification ≠ Admission, READY ≠ VERIFIED.
Task lifecycle: one ProductionFacade owns exactly one task; budget, guard,
and ledger are per-task. SINGLE path: at most 1 real invocation; four-stage
path: at most 4 (each role exactly once). Failures are structured and
terminal — *_INVOKE_FAILED, *_PACKET_INVALID, MISSING_HANDOFF,
BUDGET_EXHAUSTED, LOOP_GUARD_REJECTED, NO_CAPABLE_AGENT,
NO_VERIFICATION_CAPABILITY. Honest retries require a new task_id.
Deeper architecture: docs/architecture/ (overview, collaboration, execution, ready-vs-verified, runtime lifecycle).
Support is reported at exactly two levels — REAL VERIFIED (gated
qualification produced evidence and pool admission) and adapter
implemented (offline-tested, not yet REAL-verified in this repository).
Treat adapter-implemented runtimes as unverified until you run
dual-agent qualify in your own environment.
| Agent Runtime | Adapter | Offline Tests | REAL Verification |
|---|---|---|---|
| Claude Code CLI | claude_code_adapter.py |
✅ | ✅ REAL VERIFIED — full chain + REAL dual-agent collaboration (Claude Code CLI 2.1.227) |
| Codex CLI | codex_adapter.py |
✅ | ✅ REAL VERIFIED — audited multi-runtime four-stage E2E (2026-09) |
| Pi | pi_adapter.py |
✅ | ✅ REAL VERIFIED — audited multi-runtime four-stage E2E (2026-09) |
| Gemini CLI | gemini_adapter.py |
✅ | ❌ Not performed — gated REAL assets ship in the suite |
| Qwen Code | qwen_adapter.py |
✅ | ❌ Not performed |
| OpenCode | opencode_adapter.py |
✅ | ❌ Not performed |
| Cline | cline_adapter.py |
✅ | ❌ Not performed |
| tiny-agents (Hugging Face) | tiny_agents_adapter.py |
✅ | ❌ Not performed |
Prerequisites are runtime-level, never this package's: the CLI is on PATH and
logged in through its own flow (tiny-agents needs TINY_AGENTS_AGENT_PATH +
TINY_AGENTS_COMMAND). The engine never installs, logs in to, or configures a
runtime, and never reads credentials.
Help REAL-verify the remaining adapters — it is the highest-value
contribution right now: install the CLI, run dual-agent qualify with
RUN_REAL_PROVIDER_TESTS=1, and report your evidence. See
CONTRIBUTING.md.
Declare an agent (identity + role + runtime binding), compose a remote
session with one call, and exchange verified task packets with an agent
running in its own process on your machine — under the same packet contract
as the local pipeline. The boundary carries packets only — never
conversations, never credentials — and a DELIVERED receipt never claims
execution.
python examples/remote_offline_demo.py # offline, scripted adapter
python examples/remote_real_claude.py # with Claude Code CLI installed + logged inFull flow, agent addressing (agent:{agent-id}:{role}), and the common
failures table: docs/architecture/collaboration.md.
Implement the six-method ExternalAgentAdapter protocol — three core
invocation methods (discover, invoke, cancel) plus three health methods
(check_authentication, check_provider_model, minimal_health_check).
Adapters own all runtime specifics (executable resolution, auth state,
subprocess environment whitelist); the orchestrator only sees the protocol,
so adding a runtime never means modifying the orchestrator. Full contract:
dual-agent-development/references/adapter-contract.md;
hand probe: adapter_probe.py. A step-by-step guide is in
CONTRIBUTING.md.
- No-secrets contract: raw output, secrets, and model reasoning never
enter packets, the ledger, traces, or results;
content_safetyis the single scan authority. - Protected paths: REAL validation snapshots caller-declared credential files; any change during the run fails gate G13.
- Minimal environment: adapter subprocesses start with a whitelist env
(
PATH/HOME/USERPROFILE/SYSTEMROOT) — credential-bearing variables are never forwarded. - Real runtime calls are opt-in and off by default (
RUN_REAL_PROVIDER_TESTS=1). - The engine never reads, stores, prints, or modifies credentials. See SECURITY.md for reporting policy.
python -m pytest tests/ -q # offline suite + gated skips
python -m compileall -q dual-agent-developmentOffline suite: 3400+ tests green in CI (a handful of entries are opt-in
REAL-gated skips). REAL tests invoke real runtimes and require
RUN_REAL_PROVIDER_TESTS=1 plus a logged-in CLI — see
docs/development/testing.md. The honest status
of every layer (what is offline-tested vs REAL-verified) is tracked in
docs/architecture/ and the release notes.
Published to PyPI via Trusted Publishing (OIDC only — no tokens, no secrets)
on a pushed vX.Y.Z tag; the tag must equal dual_agent.__version__, enforced
by scripts/version_gate.py before any build. Every vX.Y.Z
tag push also creates a GitHub Release with attached artifacts and generated
notes — see the Releases page.
Package maturity: pre-1.0; the installed-CLI surface may still change shape. Remaining limitations (point-in-time qualification, closed keyword classifier, single-machine remote boundary) are listed honestly in docs/roadmap/v2-to-v3.md.
Fork → branch → keep the offline suite green → PR. The two highest-value contributions right now are adding a runtime adapter and REAL-verifying an existing one. Details and the code of conduct: CONTRIBUTING.md · CODE_OF_CONDUCT.md.
MIT — see LICENSE.