Skip to content

Repository files navigation

forge

An agentic SDLC that runs a ticket from a business request all the way to production.

Status: M0 closed · repo propose · UI v1

The full shape of M0 works — webhook → transactional dedup → job in queue → clone into a containerized sandbox → install/lint/typecheck/build/test → comment via outbox → Jira (REST v3 + ADF) — plus, beyond the original scope: a heuristic manifest proposal (repo propose) and a first UI screen (/repos + the doctor report — an explicit exception for M0).

Quick start

Requirements: Node 22+; Docker Desktop for Postgres, the containerized sandbox, and integration tests (some commands work without it — marked below).

npm install
cd web; npm install; npm run build; cd ..    # frontend → web/dist (once; then again after changes under web/)

Without a database and without Docker:

npx tsx src/cli.ts repo doctor fixtures/demo-repo          # gap report (demo with deliberate holes)
npx tsx src/cli.ts repo propose <path> --stdout            # heuristic forge.yml proposal
npx tsx src/cli.ts run --ticket SHOP-1 --repo fixtures/green-repo --sandbox local   # M0 pipeline on the host (demo, no isolation)

Postgres + migrations — one command (compose with a healthcheck):

npm run db:up        # postgres:16 container + wait for healthy + migrations; idempotent
$URL = "postgres://postgres:forge@127.0.0.1:5433/forge"    # 5433, to avoid colliding with a local PG on 5432
npx tsx src/cli.ts repo doctor <path-to-a-git-repo> --exec --db $URL   # first history point

npm run db:down stops the container (data survives on the forge-pg-data volume). The repo you seed with should be the root of a git repository — without that the report has no SHA, and the ⚖ action from the UI won't clone it (the worker clones from git_url).

UI:

npm run ui                                   # → http://127.0.0.1:4000

Why forge itself doesn't run inside a container: the sandbox spawns docker run -v <clone-dir>, and the -v path is interpreted by the host daemon — forge boxed inside a container would create clones the daemon can't see (Docker-in-Docker is rejected by design; local repo paths would also stop existing). So we containerize the infrastructure (Postgres, docker-compose.yml), not the process that talks to Docker itself.

  • ui requires a running database at startup (built-in worker + LISTEN).
  • The ⚖ Run doctor action: static mode works immediately; --exec runs the gates in the containerized sandbox (Docker). An unclonable repo ends in a red doctor run failed banner with an attempt counter — this is a designed failure path, not a crash.
  • Frontend dev with hot-reload: ui stays on :4000, alongside cd web; npm run dev → http://localhost:5173 (proxying /api).

Full M0 — sandbox / webhook:

npx tsx src/cli.ts run --ticket SHOP-482 --repo fixtures/green-repo   # container: --internal network + squid proxy (first run: image pulls)
npx tsx src/cli.ts serve --repo <repo> --db $URL                      # Jira webhook on :3000
# the Jira channel goes live when JIRA_BASE_URL / JIRA_EMAIL / JIRA_API_TOKEN are set in env; without them it falls back to the console.
# serve uses the containerized sandbox BY DEFAULT (a webhook triggers someone else's content); --sandbox local is demo-only.

Gates before changing code:

npm run lint; npm run typecheck; npm test          # root (integration tests need Docker; without it, only those fail)
cd web; npm run lint; npm test; npm run build      # frontend — the same gates the doctor requires from other repos

Postgres layer

Three architectural decisions land directly in SQL:

  • dedup + run creation + queue job in one transaction (registerWebhookAndEnqueue) — no dual-write: the database is the broker; the same transaction writes runs.repo_id, so per-repo metrics have something to stand on,
  • per-repo serialization through a Graphile Worker named queue — jobs on the same queue run strictly one at a time; finished Graphile jobs are removed from the table (housekeeping),
  • messages through an outbox with a UNIQUE idempotency key + FOR UPDATE SKIP LOCKED — a resumed pipeline won't duplicate a message, and two flushers won't send the same one twice. The flush is per row, in its own transaction: send() is an external effect (HTTP) that a rollback can't undo, so one row's failure doesn't roll back the sent_at of rows already sent (that used to cause duplicate Jira comments) and doesn't block the rows behind it; failures land in attempts/last_error, and after 5 tries a row is parked instead of retried every minute forever.

HumanChannel carries messages, not text

The human channel has one method — deliver(target, message, idempotencyKey) — and takes a structured message (HumanMessage: result | failure; more variants are added additively in M1–M3). Rendering belongs to the adapter: the console and outbox use renderText, the Jira adapter (human/jiraChannel.ts) renders through renderAdf into ADF and POSTs a REST v3 comment, and a UI adapter would store the message without rendering it to text at all.

Pluggability — a new channel without touching the core. flushOutbox calls OutboxSender.send(channel, payload), and createChannelRouter({ jira, ui, notification }) (human/channelRouter.ts) routes the message to an adapter by ChannelId. Adding Slack/Teams means a new HumanChannel implementation plus an entry in the router's map; the pipeline, the outbox, and the graph don't change. The Jira transport (HttpPost) is injected — the adapter is tested without a network (a fake), and the production fetch can be swapped in. Comment idempotency lives on the outbox side (UNIQUE + sent_at + FOR UPDATE SKIP LOCKED), because Jira REST v3 has no native idempotency key of its own.

The reasoning is architectural: the previous signature, postComment(ticketKey, text, key), baked in three Jira assumptions — the recipient is a ticket, the content is text, the unit is a comment. A UI adapter satisfies none of them, so a graph built on that signature would need its contract rebuilt rather than just gaining an adapter.

A consequence that comes for free: createOutboxChannel(pool, runId, ["jira", "ui"]) writes the same message once per adapter, keyed by <key>:<adapter> — a mirrored Jira comment for a decision made in the UI needs no separate mechanism.

Integration tests run against a real Postgres 16 via Testcontainers (this needs Docker; npm test without Docker still runs the unit tests — the integration ones fail at container startup).

Sandbox: container, two-phase network, proxy allow-list

Untrusted code from someone else's repo (install/lint/typecheck/build/test) runs in a container, not on the host (src/sandbox/). Two network phases:

  • install — the container sits on a Docker --internal network (no route out to the internet); egress goes only through a proxy (squid) with an allow-list of hosts from forge's own policy (SandboxPolicy, not the repo's manifest — otherwise an agent could add a host to its own allow-list via a PR). A raw socket bypassing HTTP_PROXY has nowhere to go: the --internal network has no external route. The enforcement point is the same one the connection actually goes through.
  • verification (lint/typecheck/build/test) — --network none, hard zero egress.

Per-container hardening: --cap-drop ALL, no-new-privileges, limits from policy (--memory/--pids-limit/--cpus — a fork bomb in a gate command runs into those limits) and a command timeout enforced by killing the container (rm -f) — a hung command doesn't hang the worker (per-repo serialization means a hung command would block that repo's entire queue). A stdout+stderr tail (~4 KB) comes back in ExecResult.outputTail — a red gate shows a human why, without SSH access to the runner.

Cleanup (down -v) always, including after a crash: step containers (tracked by name — a process killed mid docker run leaves a live container behind), the proxy, the network, the directories. Measured on green-repo: a full green run in containers; code that tries to bypass the proxy is blocked in both phases; the proxy passes an allow-listed host (200) and rejects anything else. --sandbox local mode (host, no isolation) is kept for demos without Docker, explicitly marked as unsafe. Docker with the host daemon, no Docker-in-Docker.

forge repo doctor (the first artifact)

By design, the doctor was built before the rest of the foundation (the webhook, Postgres, the queue): it answers "is this repo even fit to plug in" with a single command, before anyone invests in the orchestrator. The doctor's report is the basis for pricing the work on the target-repo side — the actual build-vs-buy decision.

npx tsx src/cli.ts repo doctor <path-to-repo>          # static checks
npx tsx src/cli.ts repo doctor <path-to-repo> --exec   # + real install/lint/typecheck/build/test
npx tsx src/cli.ts repo doctor <path-to-repo> --json   # machine-readable output
npx tsx src/cli.ts repo doctor <path> --exec --db <url>  # + write to history
npx tsx src/cli.ts repo history <path> --db <url>        # timeline of progress on a repo

# deliberate repo hookup (the doctor as a gate; registration only at 0 blockers):
npx tsx src/cli.ts repo add <path> --db <url>            # register: repo + manifest + first history point
npx tsx src/cli.ts repo add <path> --db <url> --scaffold # when there's no manifest: write an empty scaffold to fill in
npx tsx src/cli.ts repo add <path> --db <url> --exec     # run the gates for real before hooking up

Exit code: 1 on any blocker or any red command under --exec — CI-friendly.

repo add = hooking up a repo as a deliberate act: it runs the doctor and registers the repo only at zero blockers; with blockers it ends in a REFUSAL (exit 1) and never touches the database. A registered repo gets the manifest it passed with written into the repos.manifest column — this is what tells a repo hooked up on purpose apart from one merely seen in passing via doctor --db (there, manifest stays NULL). --scaffold writes an empty .forge/forge.yml scaffold and never overwrites an existing manifest; the scaffold doesn't guess commands — guessing is what repo propose is for.

forge repo propose — a heuristic manifest proposal

npx tsx src/cli.ts repo propose <path>            # write to .forge/forge.yml (refuses if a manifest already exists)
npx tsx src/cli.ts repo propose <path> --stdout   # just print the proposal

The only sanctioned place where forge guesses build commands — and for that reason, never silently: every proposed value carries a # PROPOSED (source, confidence) — verify marker, and anything the heuristic can't detect with confidence (e.g. typecheck with no matching script, risk.always_human_paths) is left commented out, never invented. Detected ecosystems: npm (a lockfile → install, scripts → commands) and .NET (*.sln/*.csproj → dotnet). Monorepos with multiple ecosystems get their commands printed out commented, for manual assembly — stitching build chains together is a human decision, not a heuristic one. forbidden_paths is written in directly as an invariant ("the agent doesn't modify the environment that verifies it"). The result is measurable: Detection[] with confidence/source fields, shown in the CLI as ● written in / ○ commented out. A proposal doesn't register the repo — review it, fix it, then run repo doctor / repo add.

forge ui — the /repos screen + the doctor report (an M0 exception)

The plan puts the UI in M4, with one exception: the doctor report already has data ready before the graph exists. Stack: React + Vite + TanStack Query + Tailwind, served through Fastify.

  • Read-only API (/api/repos, /api/repos/:slug, /api/repos/:slug/reports, /api/reports/:id; TypeBox, schema-first) — the UI is a mirror of state: the manifest and policy change by PR / outside forge, never from the screen.
  • The only write: POST /api/repos/:slug/doctor → enqueues onto the repo's named queue (serialized with M0 runs). Work never happens inside the HTTP handler — --exec runs the repo's real commands in the sandbox.
  • Live over SSE (/api/events): saveDoctorReport → pg_notify('doctor_reports') → LISTEN → a doctor.report_saved | doctor.run_failed event — the database is the broker here too, no separate bus. The client invalidates its queries on every reconnect (the server doesn't replay) and reconnects itself after the EventSource permanently closes — the "live" indicator never promises a reconnect it won't deliver.
  • Failure path ⚖: a job error (e.g. an unclonable repo) → NOTIFY run_failed with attempt/maxAttempts → a red banner on the repo screen; a later successful report for that repo clears it on its own.
  • Shared FE/BE domain: src/shared/doctorDomain.ts — one definition for finding types, severity/milestone ordering, and shortSha (rule: never truncate the -dirty marker); the backend re-exports it, the web app imports straight from source, and tests exist on both sides of the boundary.
  • web/ has its own gates: eslint (the same standard as root plus hook rules), vitest, tsc + vite build.

What the doctor checks (and what it blocks)

Every finding has a severity (blocker/warning/info) and a milestone tag whose missing precondition blocks that milestone:

Check Blocks Behavior
.forge/forge.yml manifest present and valid M0 missing/invalid = blocker; forge doesn't guess by heuristic
install/lint/typecheck/build/test commands M0 each missing one is its own blocker; --exec runs them for real and stops at the first red one
CI resistant to "pwn request" M0 pull_request_target/workflow_run + checkout of the PR head or interpolation of PR fields into run: = blocker (the root cause behind the s1ngularity supply-chain attack)
forbidden_paths declared M0 blocker — the agent must not be able to modify the environment that verifies it
risk.always_human_paths M0 warning
role profiles without a persona, skills with a live example M0 warning — SKILL.md format
test_integration M3 warning
runtime up/down + healthchecks M4 warning
external_services declarations M4 info (hosts and test keys → forge policy)
selector strategy + data-testid heuristic across workspaces M5 warning
release triggers + observation window with min_samples/max_window_minutes M7 info/warning

Conventions (from day one)

  • Pure functions with tests: parseManifest, auditWorkflow, checkManifest, proposeManifest, summarizeReport… — control logic never touches disk.
  • Injected dependencies (DoctorDeps, ProposeDeps): fs/exec come in from outside; tests use in-memory implementations. Zero module-level instances.
  • The manifest parser is lenient (sections are optional) — the severity of a gap is decided by checks with milestone tags, not by the schema.
  • Domain rules have a single source (src/shared/doctorDomain.ts) with tests on both sides of the FE/BE boundary — an untested copy drifts silently.
  • Code and comments in English; the CLI's exit code is CI-friendly.

Structure

src/
  cli.ts                      # commander: repo doctor|add|propose|history · run · serve · ui · db migrate
  app.ts                      # composition root — the only place adapters get assembled
  systemDeps.ts               # production implementation of DoctorDeps (fs, spawn)
  localAdapters.ts            # M0 adapters: local sandbox (host, unsafe) + containerized, git clone, stdout channel
  shared/
    doctorDomain.ts           # shared FE/BE domain: Finding/Severity/Milestone, finding order, shortSha
  api/
    uiServer.ts               # Fastify: /api/* (TypeBox) + static web/dist + SSE /api/events
    reportNotifier.ts         # LISTEN doctor_reports → fan-out to SSE subscribers
  sandbox/
    policy.ts                 # SandboxPolicy: image + network allow-lists (forge's own policy, not the repo manifest)
    dockerArgs.ts             # pure builders: squid.conf (allow-list), docker run args, phases
    containerSandbox.ts       # the sandbox: container + --internal network + proxy, two phases, always cleans up
  human/
    message.ts                # HumanChannel + HumanMessage — the human-channel contract
    renderText.ts             # text renderer (console/outbox adapter)
    renderAdf.ts              # ADF renderer (Jira adapter) — same HumanMessage, different format
    jiraChannel.ts            # Jira adapter: HumanMessage → ADF → POST REST v3 comment (injected transport)
    channelRouter.ts           # OutboxSender routing by ChannelId — the seam for pluggable new channels
    rubberStamp.ts             # rubber-stamping detector — measures, doesn't block
  doctor/
    doctor.ts                 # check orchestration
    manifest.ts               # forge.yml schema (zod) + parser
    report.ts                 # sorting, rendering, exit code
    summary.ts                # counts + "ready for M?" — a pure function, feeds history and the UI
    types.ts                  # DoctorReport, DoctorDeps (finding types re-exported from shared/)
    checks/
      ciPwn.ts                # GH Actions workflow audit (pwn request, script injection)
      manifestChecks.ts       # preconditions over the manifest
      dotForge.ts             # skills (SKILL.md, example) and role profiles (no personas)
  repo/
    paths.ts                  # forbidden_paths + disjoint agent scopes (pure functions)
    repoClient.ts             # the only path to writing and to git; git runs sequentially (queued)
    manifestFile.ts           # scaffold/propose-write/load .forge/forge.yml
    proposeManifest.ts        # `repo propose` heuristic: npm/.NET → Detection[] with per-field confidence
  webhook/
    events.ts                 # Jira event parsing, dedup key (sanitized), DedupStore
    server.ts                 # Fastify: 202 immediately, atomic ingest, filters out forge's own comments
  pipeline/
    m0.ts                     # clone → doctor --exec → comment; always cleans up
    doctorRun.ts              # the UI's `⚖` action job: clone → doctor → save (NOTIFY run_failed in the catch)
  db/
    migrate.ts                # sql/ runner + Graphile Worker migrations
    stores.ts                 # registerWebhookAndEnqueue (1 tx), outbox, run_events
    doctorReports.ts          # report history + pg_notify('doctor_reports') on save
    repos.ts                  # read model for the /repos screen: repo + latest report (LATERAL)
    humanGates.ts             # human gates + the two-channel race rule in SQL
  queue/
    worker.ts                 # m0_run and outbox_flush tasks (cron every minute)
  events/
    catalog.ts                # closed dictionary of run_events.type + typed payloads
    timeline.ts               # lead-time breakdown: queue / machine / waiting on a human
sql/
  001_init.sql                # repos, webhook_events (UNIQUE), runs, run_events (append-only), outbox
  002_doctor_reports.sql      # doctor report history — the timeline of progress on a repo
  003_human_gates.sql         # human gates: answered_via + the invariant "complete knows its channel"
  004_outbox_attempts.sql     # attempt counter + parking poison rows
web/                          # UI: React + Vite + TanStack Query + Tailwind; its own gates (eslint, vitest, build)
  src/lib/                    # api (types from shared/), queryKeys, live (SSE provider), router, format, findings
  src/components/report/      # RunDoctorAction (⚖), ProgressBar, MilestoneGroup, PwnBanner, ExecSidebar, ReportView
  src/pages/                  # ReposPage, RepoDoctorPage
fixtures/
  demo-repo/                  # a repo with deliberate gaps (doctor demo)
  green-repo/                 # a git repo meeting M0 (end-to-end pipeline demo)

What's next

M0 is closed and measured: a full green run in the sandbox; a real fetch → router → Jira channel → ADF comment lands on …/issue/<key>/comment; the UI's ⚖ action has a verified round trip (save → SSE → refresh; failure → doctor.run_failed).

The next milestone is M1 — the graph (LangGraph.js: checkpointer + interrupt/resume), landing in the same composition root without touching existing contracts (M0Deps, HumanChannel, the run_events dictionary). Deliberately deferred: the full SSE contract over run_events (a Last-Event-ID cursor) — arrives with the run screens; verification-phase egress (sandbox endpoints: Stripe test mode, etc.) — M3+ scope (test_integration/e2e); a non-empty verifyAllow is rejected for now, so as not to fake support that isn't there; dockerode instead of the CLI, and git worktree instead of a plain clone; status transitions and Jira fields — M1+ (that's when the questions/plan variants land in HumanMessage).

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages