An event-driven incident-response room where specialist AI agents investigate, challenge, and mitigate simulated production failures—with a human approval gate before any action is taken.
PagerSwarm uses Mozaik to coordinate independent participants on one shared runtime. It is deliberately not a sequential agent pipeline: agents react to events as they arrive, work in parallel, and use a typed shared state to converge on an evidence-backed remedy.
- A deterministic production-system simulator injects a fault and emits telemetry.
- Sentinel opens an incident when an anomaly is detected.
- Loghound, Delta, and Probe investigate concurrently from logs, changes, topology, and metrics.
- Skeptic tests the evidence and confirms a root-cause hypothesis only when independent signals agree.
- Medic proposes a remedy; its operation is paused until a human approves or denies it.
- The simulator validates the change, proves recovery through telemetry, and Scribe produces an incident summary.
flowchart LR
SIM[Production simulator] --> SENT[Sentinel]
SENT --> BUS[Mozaik shared runtime and event bus]
BUS --> LOG[Loghound]
BUS --> DELTA[Delta]
BUS --> PROBE[Probe]
LOG --> SKEP[Skeptic]
DELTA --> SKEP
PROBE --> SKEP
SKEP --> MED[Medic]
MED --> APPROVAL{Human approval}
APPROVAL -->|approve| SIM
APPROVAL -->|deny| MED
SIM --> SCRIBE[Scribe]
BUS --> UI[Live web console]
The runtime state records incidents, findings, hypotheses, approvals, and a human-readable timeline. The dashboard is a dependency-free static UI served through an Express/SSE bridge.
| Participant | Responsibility |
|---|---|
| Sentinel | Detects anomalies and opens or resolves incidents. |
| Loghound | Searches service logs for error patterns. |
| Delta | Correlates incidents with deploys and configuration changes. |
| Probe | Traces dependencies and metric deviations. |
| Skeptic | Challenges weak evidence and confirms root cause. |
| Medic | Selects a remediation and waits for human approval. |
| Scribe | Produces timeline updates and incident summaries. |
| Commander | Human operator who can steer the room and approve actions. |
Prerequisites: Node.js 18 or later.
npm ci
npm run demonpm run demo runs the complete deterministic simulation locally, with no API key or network model dependency. To use the browser console instead:
npm startThen open http://localhost:3000.
Copy the safe template, then add credentials only to your local .env file:
cp .env.example .envOn PowerShell:
Copy-Item .env.example .envPagerSwarm supports direct Mozaik providers (Gemini, OpenAI, and Anthropic) and any OpenAI-compatible endpoint. Gemini is not required.
For a split-provider topology, route the four high-volume investigation and reporting roles through Groq, while Skeptic and Medic use Cerebras for independent judgment:
PAGER_SWARM_FREE_BASE_URL=https://api.groq.com/openai/v1
PAGER_SWARM_FREE_MODEL=openai/gpt-oss-120b
PAGER_SWARM_FREE_API_KEY=your_groq_key
PAGER_SWARM_JUDGE_BASE_URL=https://api.cerebras.ai/v1
PAGER_SWARM_JUDGE_MODEL=gpt-oss-120b
PAGER_SWARM_JUDGE_API_KEY=your_cerebras_keyThe runtime applies pacing and retry logic separately to each provider. This keeps the concurrent swarm intact while avoiding a single shared provider queue. If the Cerebras judgment route cannot complete a request, Skeptic and Medic automatically retry through the Groq volume route. The Groq and Cerebras endpoints are both OpenAI-compatible; see their Groq compatibility documentation and Cerebras compatibility documentation.
Never put real keys in the repository. .env, .env.*, event logs, and generated output are ignored; .env.example is the only configuration file intended for source control.
| Command | Description |
|---|---|
npm run demo |
Deterministic console demonstration; no key required. |
npm run demo:console |
Console simulation using configured providers when available. |
npm start |
Start the live web situation room. |
npm run dev |
Start the web server in watch mode. |
npm run typecheck |
Validate TypeScript without emitting files. |
npm run build |
Compile the TypeScript project. |
The built-in world includes four deterministic failure modes: memory leak, bad deploy, cache stampede, and poisoned cache. Each creates distinctive telemetry and logs, and only the correct remediation causes recovery. This makes the agent collaboration observable and repeatable without access to production infrastructure.
src/
agents/ Mozaik participants and situation handlers
bridge/ SSE, API, snapshots, and audit observer
runtime/ Runtime definition, state, events, provider routing
session/ War-room bootstrapping
sim/ Deterministic production-system simulator
tools/ Read-only investigation and approval-gated operations
ui/ Static incident-console interface
PagerSwarm is a simulator. It does not connect to or modify real infrastructure. In the simulated environment, all remediation tools are intercepted and require an explicit approval decision before execution.