An agentic SDLC that runs a ticket from a business request all the way to production.
The full shape of M0 works — webhook → transactional dedup → job in queue → clone into a containerized sandbox → install/lint/typecheck/build/test → comment via outbox → Jira (REST v3 + ADF) — plus, beyond the original scope: a heuristic manifest proposal (repo propose) and a first UI screen (/repos + the doctor report — an explicit exception for M0).
Requirements: Node 22+; Docker Desktop for Postgres, the containerized sandbox, and integration tests (some commands work without it — marked below).
npm install
cd web; npm install; npm run build; cd .. # frontend → web/dist (once; then again after changes under web/)Without a database and without Docker:
npx tsx src/cli.ts repo doctor fixtures/demo-repo # gap report (demo with deliberate holes)
npx tsx src/cli.ts repo propose <path> --stdout # heuristic forge.yml proposal
npx tsx src/cli.ts run --ticket SHOP-1 --repo fixtures/green-repo --sandbox local # M0 pipeline on the host (demo, no isolation)Postgres + migrations — one command (compose with a healthcheck):
npm run db:up # postgres:16 container + wait for healthy + migrations; idempotent
$URL = "postgres://postgres:forge@127.0.0.1:5433/forge" # 5433, to avoid colliding with a local PG on 5432
npx tsx src/cli.ts repo doctor <path-to-a-git-repo> --exec --db $URL # first history pointnpm run db:down stops the container (data survives on the forge-pg-data volume). The repo you seed with should be the root of a git repository — without that the report has no SHA, and the ⚖ action from the UI won't clone it (the worker clones from git_url).
UI:
npm run ui # → http://127.0.0.1:4000Why forge itself doesn't run inside a container: the sandbox spawns
docker run -v <clone-dir>, and the-vpath is interpreted by the host daemon — forge boxed inside a container would create clones the daemon can't see (Docker-in-Docker is rejected by design; local repo paths would also stop existing). So we containerize the infrastructure (Postgres,docker-compose.yml), not the process that talks to Docker itself.
uirequires a running database at startup (built-in worker + LISTEN).- The
⚖ Run doctoraction: static mode works immediately;--execruns the gates in the containerized sandbox (Docker). An unclonable repo ends in a reddoctor run failedbanner with an attempt counter — this is a designed failure path, not a crash. - Frontend dev with hot-reload:
uistays on :4000, alongsidecd web; npm run dev→ http://localhost:5173 (proxying/api).
Full M0 — sandbox / webhook:
npx tsx src/cli.ts run --ticket SHOP-482 --repo fixtures/green-repo # container: --internal network + squid proxy (first run: image pulls)
npx tsx src/cli.ts serve --repo <repo> --db $URL # Jira webhook on :3000
# the Jira channel goes live when JIRA_BASE_URL / JIRA_EMAIL / JIRA_API_TOKEN are set in env; without them it falls back to the console.
# serve uses the containerized sandbox BY DEFAULT (a webhook triggers someone else's content); --sandbox local is demo-only.Gates before changing code:
npm run lint; npm run typecheck; npm test # root (integration tests need Docker; without it, only those fail)
cd web; npm run lint; npm test; npm run build # frontend — the same gates the doctor requires from other reposThree architectural decisions land directly in SQL:
- dedup + run creation + queue job in one transaction (
registerWebhookAndEnqueue) — no dual-write: the database is the broker; the same transaction writesruns.repo_id, so per-repo metrics have something to stand on, - per-repo serialization through a Graphile Worker named queue — jobs on the same queue run strictly one at a time; finished Graphile jobs are removed from the table (housekeeping),
- messages through an outbox with a UNIQUE idempotency key +
FOR UPDATE SKIP LOCKED— a resumed pipeline won't duplicate a message, and two flushers won't send the same one twice. The flush is per row, in its own transaction:send()is an external effect (HTTP) that a rollback can't undo, so one row's failure doesn't roll back thesent_atof rows already sent (that used to cause duplicate Jira comments) and doesn't block the rows behind it; failures land inattempts/last_error, and after 5 tries a row is parked instead of retried every minute forever.
The human channel has one method — deliver(target, message, idempotencyKey) — and takes a structured message (HumanMessage: result | failure; more variants are added additively in M1–M3). Rendering belongs to the adapter: the console and outbox use renderText, the Jira adapter (human/jiraChannel.ts) renders through renderAdf into ADF and POSTs a REST v3 comment, and a UI adapter would store the message without rendering it to text at all.
Pluggability — a new channel without touching the core. flushOutbox calls OutboxSender.send(channel, payload), and createChannelRouter({ jira, ui, notification }) (human/channelRouter.ts) routes the message to an adapter by ChannelId. Adding Slack/Teams means a new HumanChannel implementation plus an entry in the router's map; the pipeline, the outbox, and the graph don't change. The Jira transport (HttpPost) is injected — the adapter is tested without a network (a fake), and the production fetch can be swapped in. Comment idempotency lives on the outbox side (UNIQUE + sent_at + FOR UPDATE SKIP LOCKED), because Jira REST v3 has no native idempotency key of its own.
The reasoning is architectural: the previous signature, postComment(ticketKey, text, key), baked in three Jira assumptions — the recipient is a ticket, the content is text, the unit is a comment. A UI adapter satisfies none of them, so a graph built on that signature would need its contract rebuilt rather than just gaining an adapter.
A consequence that comes for free: createOutboxChannel(pool, runId, ["jira", "ui"]) writes the same message once per adapter, keyed by <key>:<adapter> — a mirrored Jira comment for a decision made in the UI needs no separate mechanism.
Integration tests run against a real Postgres 16 via Testcontainers (this needs Docker; npm test without Docker still runs the unit tests — the integration ones fail at container startup).
Untrusted code from someone else's repo (install/lint/typecheck/build/test) runs in a container, not on the host (src/sandbox/). Two network phases:
- install — the container sits on a Docker
--internalnetwork (no route out to the internet); egress goes only through a proxy (squid) with an allow-list of hosts from forge's own policy (SandboxPolicy, not the repo's manifest — otherwise an agent could add a host to its own allow-list via a PR). A raw socket bypassingHTTP_PROXYhas nowhere to go: the--internalnetwork has no external route. The enforcement point is the same one the connection actually goes through. - verification (
lint/typecheck/build/test) —--network none, hard zero egress.
Per-container hardening: --cap-drop ALL, no-new-privileges, limits from policy (--memory/--pids-limit/--cpus — a fork bomb in a gate command runs into those limits) and a command timeout enforced by killing the container (rm -f) — a hung command doesn't hang the worker (per-repo serialization means a hung command would block that repo's entire queue). A stdout+stderr tail (~4 KB) comes back in ExecResult.outputTail — a red gate shows a human why, without SSH access to the runner.
Cleanup (down -v) always, including after a crash: step containers (tracked by name — a process killed mid docker run leaves a live container behind), the proxy, the network, the directories. Measured on green-repo: a full green run in containers; code that tries to bypass the proxy is blocked in both phases; the proxy passes an allow-listed host (200) and rejects anything else. --sandbox local mode (host, no isolation) is kept for demos without Docker, explicitly marked as unsafe. Docker with the host daemon, no Docker-in-Docker.
By design, the doctor was built before the rest of the foundation (the webhook, Postgres, the queue): it answers "is this repo even fit to plug in" with a single command, before anyone invests in the orchestrator. The doctor's report is the basis for pricing the work on the target-repo side — the actual build-vs-buy decision.
npx tsx src/cli.ts repo doctor <path-to-repo> # static checks
npx tsx src/cli.ts repo doctor <path-to-repo> --exec # + real install/lint/typecheck/build/test
npx tsx src/cli.ts repo doctor <path-to-repo> --json # machine-readable output
npx tsx src/cli.ts repo doctor <path> --exec --db <url> # + write to history
npx tsx src/cli.ts repo history <path> --db <url> # timeline of progress on a repo
# deliberate repo hookup (the doctor as a gate; registration only at 0 blockers):
npx tsx src/cli.ts repo add <path> --db <url> # register: repo + manifest + first history point
npx tsx src/cli.ts repo add <path> --db <url> --scaffold # when there's no manifest: write an empty scaffold to fill in
npx tsx src/cli.ts repo add <path> --db <url> --exec # run the gates for real before hooking upExit code: 1 on any blocker or any red command under --exec — CI-friendly.
repo add = hooking up a repo as a deliberate act: it runs the doctor and registers the repo only at zero blockers; with blockers it ends in a REFUSAL (exit 1) and never touches the database. A registered repo gets the manifest it passed with written into the repos.manifest column — this is what tells a repo hooked up on purpose apart from one merely seen in passing via doctor --db (there, manifest stays NULL). --scaffold writes an empty .forge/forge.yml scaffold and never overwrites an existing manifest; the scaffold doesn't guess commands — guessing is what repo propose is for.
npx tsx src/cli.ts repo propose <path> # write to .forge/forge.yml (refuses if a manifest already exists)
npx tsx src/cli.ts repo propose <path> --stdout # just print the proposalThe only sanctioned place where forge guesses build commands — and for that reason, never silently: every proposed value carries a # PROPOSED (source, confidence) — verify marker, and anything the heuristic can't detect with confidence (e.g. typecheck with no matching script, risk.always_human_paths) is left commented out, never invented. Detected ecosystems: npm (a lockfile → install, scripts → commands) and .NET (*.sln/*.csproj → dotnet). Monorepos with multiple ecosystems get their commands printed out commented, for manual assembly — stitching build chains together is a human decision, not a heuristic one. forbidden_paths is written in directly as an invariant ("the agent doesn't modify the environment that verifies it"). The result is measurable: Detection[] with confidence/source fields, shown in the CLI as ● written in / ○ commented out. A proposal doesn't register the repo — review it, fix it, then run repo doctor / repo add.
The plan puts the UI in M4, with one exception: the doctor report already has data ready before the graph exists. Stack: React + Vite + TanStack Query + Tailwind, served through Fastify.
- Read-only API (
/api/repos,/api/repos/:slug,/api/repos/:slug/reports,/api/reports/:id; TypeBox, schema-first) — the UI is a mirror of state: the manifest and policy change by PR / outside forge, never from the screen. - The only write:
POST /api/repos/:slug/doctor→ enqueues onto the repo's named queue (serialized with M0 runs). Work never happens inside the HTTP handler —--execruns the repo's real commands in the sandbox. - Live over SSE (
/api/events):saveDoctorReport→pg_notify('doctor_reports')→ LISTEN → adoctor.report_saved|doctor.run_failedevent — the database is the broker here too, no separate bus. The client invalidates its queries on every reconnect (the server doesn't replay) and reconnects itself after the EventSource permanently closes — the "live" indicator never promises a reconnect it won't deliver. - Failure path
⚖: a job error (e.g. an unclonable repo) → NOTIFYrun_failedwithattempt/maxAttempts→ a red banner on the repo screen; a later successful report for that repo clears it on its own. - Shared FE/BE domain:
src/shared/doctorDomain.ts— one definition for finding types, severity/milestone ordering, andshortSha(rule: never truncate the-dirtymarker); the backend re-exports it, the web app imports straight from source, and tests exist on both sides of the boundary. web/has its own gates: eslint (the same standard as root plus hook rules), vitest,tsc + vite build.
Every finding has a severity (blocker/warning/info) and a milestone tag whose missing precondition blocks that milestone:
| Check | Blocks | Behavior |
|---|---|---|
.forge/forge.yml manifest present and valid |
M0 | missing/invalid = blocker; forge doesn't guess by heuristic |
install/lint/typecheck/build/test commands |
M0 | each missing one is its own blocker; --exec runs them for real and stops at the first red one |
| CI resistant to "pwn request" | M0 | pull_request_target/workflow_run + checkout of the PR head or interpolation of PR fields into run: = blocker (the root cause behind the s1ngularity supply-chain attack) |
forbidden_paths declared |
M0 | blocker — the agent must not be able to modify the environment that verifies it |
risk.always_human_paths |
M0 | warning |
role profiles without a persona, skills with a live example |
M0 | warning — SKILL.md format |
test_integration |
M3 | warning |
| runtime up/down + healthchecks | M4 | warning |
external_services declarations |
M4 | info (hosts and test keys → forge policy) |
selector strategy + data-testid heuristic across workspaces |
M5 | warning |
release triggers + observation window with min_samples/max_window_minutes |
M7 | info/warning |
- Pure functions with tests:
parseManifest,auditWorkflow,checkManifest,proposeManifest,summarizeReport… — control logic never touches disk. - Injected dependencies (
DoctorDeps,ProposeDeps): fs/exec come in from outside; tests use in-memory implementations. Zero module-level instances. - The manifest parser is lenient (sections are optional) — the severity of a gap is decided by checks with milestone tags, not by the schema.
- Domain rules have a single source (
src/shared/doctorDomain.ts) with tests on both sides of the FE/BE boundary — an untested copy drifts silently. - Code and comments in English; the CLI's exit code is CI-friendly.
src/
cli.ts # commander: repo doctor|add|propose|history · run · serve · ui · db migrate
app.ts # composition root — the only place adapters get assembled
systemDeps.ts # production implementation of DoctorDeps (fs, spawn)
localAdapters.ts # M0 adapters: local sandbox (host, unsafe) + containerized, git clone, stdout channel
shared/
doctorDomain.ts # shared FE/BE domain: Finding/Severity/Milestone, finding order, shortSha
api/
uiServer.ts # Fastify: /api/* (TypeBox) + static web/dist + SSE /api/events
reportNotifier.ts # LISTEN doctor_reports → fan-out to SSE subscribers
sandbox/
policy.ts # SandboxPolicy: image + network allow-lists (forge's own policy, not the repo manifest)
dockerArgs.ts # pure builders: squid.conf (allow-list), docker run args, phases
containerSandbox.ts # the sandbox: container + --internal network + proxy, two phases, always cleans up
human/
message.ts # HumanChannel + HumanMessage — the human-channel contract
renderText.ts # text renderer (console/outbox adapter)
renderAdf.ts # ADF renderer (Jira adapter) — same HumanMessage, different format
jiraChannel.ts # Jira adapter: HumanMessage → ADF → POST REST v3 comment (injected transport)
channelRouter.ts # OutboxSender routing by ChannelId — the seam for pluggable new channels
rubberStamp.ts # rubber-stamping detector — measures, doesn't block
doctor/
doctor.ts # check orchestration
manifest.ts # forge.yml schema (zod) + parser
report.ts # sorting, rendering, exit code
summary.ts # counts + "ready for M?" — a pure function, feeds history and the UI
types.ts # DoctorReport, DoctorDeps (finding types re-exported from shared/)
checks/
ciPwn.ts # GH Actions workflow audit (pwn request, script injection)
manifestChecks.ts # preconditions over the manifest
dotForge.ts # skills (SKILL.md, example) and role profiles (no personas)
repo/
paths.ts # forbidden_paths + disjoint agent scopes (pure functions)
repoClient.ts # the only path to writing and to git; git runs sequentially (queued)
manifestFile.ts # scaffold/propose-write/load .forge/forge.yml
proposeManifest.ts # `repo propose` heuristic: npm/.NET → Detection[] with per-field confidence
webhook/
events.ts # Jira event parsing, dedup key (sanitized), DedupStore
server.ts # Fastify: 202 immediately, atomic ingest, filters out forge's own comments
pipeline/
m0.ts # clone → doctor --exec → comment; always cleans up
doctorRun.ts # the UI's `⚖` action job: clone → doctor → save (NOTIFY run_failed in the catch)
db/
migrate.ts # sql/ runner + Graphile Worker migrations
stores.ts # registerWebhookAndEnqueue (1 tx), outbox, run_events
doctorReports.ts # report history + pg_notify('doctor_reports') on save
repos.ts # read model for the /repos screen: repo + latest report (LATERAL)
humanGates.ts # human gates + the two-channel race rule in SQL
queue/
worker.ts # m0_run and outbox_flush tasks (cron every minute)
events/
catalog.ts # closed dictionary of run_events.type + typed payloads
timeline.ts # lead-time breakdown: queue / machine / waiting on a human
sql/
001_init.sql # repos, webhook_events (UNIQUE), runs, run_events (append-only), outbox
002_doctor_reports.sql # doctor report history — the timeline of progress on a repo
003_human_gates.sql # human gates: answered_via + the invariant "complete knows its channel"
004_outbox_attempts.sql # attempt counter + parking poison rows
web/ # UI: React + Vite + TanStack Query + Tailwind; its own gates (eslint, vitest, build)
src/lib/ # api (types from shared/), queryKeys, live (SSE provider), router, format, findings
src/components/report/ # RunDoctorAction (⚖), ProgressBar, MilestoneGroup, PwnBanner, ExecSidebar, ReportView
src/pages/ # ReposPage, RepoDoctorPage
fixtures/
demo-repo/ # a repo with deliberate gaps (doctor demo)
green-repo/ # a git repo meeting M0 (end-to-end pipeline demo)
M0 is closed and measured: a full green run in the sandbox; a real fetch → router → Jira channel → ADF comment lands on …/issue/<key>/comment; the UI's ⚖ action has a verified round trip (save → SSE → refresh; failure → doctor.run_failed).
The next milestone is M1 — the graph (LangGraph.js: checkpointer + interrupt/resume), landing in the same composition root without touching existing contracts (M0Deps, HumanChannel, the run_events dictionary). Deliberately deferred: the full SSE contract over run_events (a Last-Event-ID cursor) — arrives with the run screens; verification-phase egress (sandbox endpoints: Stripe test mode, etc.) — M3+ scope (test_integration/e2e); a non-empty verifyAllow is rejected for now, so as not to fake support that isn't there; dockerode instead of the CLI, and git worktree instead of a plain clone; status transitions and Jira fields — M1+ (that's when the questions/plan variants land in HumanMessage).