AgentOS is a wrapper-first execution and trace layer for existing automation.
It is designed to help teams progressively move from repeated agentic/manual decisions to deterministic rules, while keeping safe fallback behavior.
In MVP, AgentOS focuses on:
- Wrapping existing scripts, CI jobs, and headless agent commands.
- Tracing runs, decisions, and outcomes.
- Detecting repeated patterns.
- Backtesting deterministic rule candidates.
- Promoting conservative rules with metrics.
- Running rule-first with fallback preserved by default.
Core operating rule:
Run capture is automatic.
Decision capture is declarative.
Compilation requires validated decisions and outcomes.
Core loop:
wrap existing process
→ trace decisions
→ detect repeated patterns
→ backtest deterministic rules
→ promote rules
→ run rule-first with fallback
AgentOS MVP is not:
- A coding agent.
- A chatbot.
- A full workflow orchestrator.
- A replacement for your existing scripts, MCP tools, or CI system.
- Wrapper-first adoption: no forced rewrite of existing automation.
- Fallback by default: promoted/compiled rules must not remove fallback.
- Conservative compilation: rules abstain when uncertain.
- Transparent evidence: promotion requires backtest metrics.
- Simple local primitives: CLI, SQLite, JSONL, filesystem artifacts.
AgentOS sits around existing workers:
- Existing scripts / CI jobs / Codex / Claude / custom tooling perform tasks.
- AgentOS records execution context and decisions.
- AgentOS analyzes repetition and proposes deterministic candidates.
- AgentOS backtests candidates before promotion.
- Runtime can apply rule-first routing, then fallback to original process.
AgentOS starts by wrapping existing automation:
agentos wrap --intent gitlab.fix_ci -- ./run-existing-agent.shThis captures the run.
To capture LLM decisions for compilation, the process must declare operational decisions through a decision file, stdout markers, or explicit instrumentation.
agentos wrap \
--intent gitlab.fix_ci \
--decision-file agentos-artifacts/decisions.json \
-- ./run-existing-agent.shAgentOS does not infer hidden model reasoning. It compiles only validated declared decisions with outcomes.
A declared decision MAY include a features field — a flat dictionary of structured input attributes used for feature-conditioned pattern mining. This lets the rule extractor discover deterministic shortcuts like "when is_root=true AND has_mention=true, the LLM always picks feedback", instead of just "this decision_key is usually X".
Schema:
featuresis an optionaldict[str, str | bool | int | float].- Keys are non-empty strings. Values must be primitives (no nested dicts/lists).
- Decisions without features still work — they're handled by the (key-only)
patterns list.
Example payload:
{
"step_id": "teams.classify_thread",
"decision_type": "llm",
"input_fingerprint": "sha256:...",
"output": {"chosen": "feedback", "confidence": 0.95},
"evidence": ["sender=Pierre Derrier", "channel=devops-claude"],
"features": {
"is_root": true,
"is_reply": false,
"has_mention": true,
"sender_is_devops": true,
"channel_id": "19:devops@thread.tacv2"
},
"compilation_candidate": true
}Recording from the CLI:
agentos decision record \
--run-id "$RUN_ID" \
--key teams.classify_thread \
--step teams.classify_thread \
--type llm \
--input-fingerprint "$FP" \
--output-json '{"chosen":"feedback","confidence":0.95}' \
--evidence-json '[]' \
--features-json '{"is_root":true,"has_mention":true,"sender_is_devops":true}' \
--candidate trueMining feature-conditioned patterns:
# Bucket decisions by (decision_key, features) and report dominant choice per bucket.
# Output: one JSON line per (decision_key, features) combo with confirmed outcomes.
agentos patterns list --by-features --min-support 30
# Filter to a single decision_key:
agentos patterns list --by-features --decision-key teams.classify_thread --min-support 30Feature-conditioned buckets with confidence == 1.0 and high support are candidate deterministic rules.
patterns list --by-features reports exact-match feature buckets — useful when every feature is observed and bucketed identically across runs. For a more powerful mining loop that finds subset rules (e.g. "the LLM picks feedback whenever is_root=true AND has_mention=true, regardless of other features"), AgentOS ships a greedy decision-tree extractor:
agentos rules extract \
--decision-key teams.classify_thread \
--min-coverage 20 \
--min-precision 0.95 \
--max-depth 4The extractor:
- Pulls all valid+candidate decisions for
--decision-keywith confirmed outcomes. - Builds a greedy CART-like decision tree on the structured
featuresfield, splitting on the highest Gini-impurity reduction at each node. - Walks high-precision leaves and emits one JSON-line rule proposal per qualifying leaf.
Output shape (one JSON object per line):
{
"decision_key": "teams.classify_thread",
"predicate": [
{"feature": "is_root", "op": "==", "value": true},
{"feature": "has_mention", "op": "==", "value": true}
],
"chosen": "feedback",
"coverage": 47,
"precision": 1.0,
"support_total": 200,
"support_share": 0.235
}Rules are proposals, not promotions. Promotion to the rule store stays an explicit, human-reviewed step. The extractor is intentionally pure-Python with no ML dependency — Gini impurity, greedy single-feature splits, and equality-only predicates keep proposals simple and reviewable.
After reviewing the proposal, persist it with:
agentos rules promote-extracted \
--decision-key teams.classify_thread \
--predicate-json '[{"feature":"is_root","op":"==","value":true},
{"feature":"has_mention","op":"==","value":true}]' \
--chosen feedback \
--metrics-json '{"coverage":47,"precision":1.0,"support_total":200}'Promoted feature-rules live in the same promoted_rules table as key-only rules — they are distinguished by the predicate_json column.
Wrapped workers consume promoted rules via the runtime SDK before invoking their LLM/agent — if a rule matches, the worker returns the deterministic answer and skips the model call entirely:
from agentos.runtime import check_rule
decision = check_rule("teams.classify_thread", {
"is_root": True,
"has_mention": True,
"sender_is_devops": True,
})
if decision is not None:
# Rule fired — skip the LLM, use the deterministic answer.
return decision["chosen"] # → "feedback"
# No rule matched — fall through to the model.
return llm_classify(...)Returns None when no rule matches — the caller decides what fallback to invoke (LLM, manual flow, etc.). The most specific rule (longest predicate) wins ties, mirroring how a decision tree's deeper leaves carry more discriminating information.
This is the loop's payoff: when a feature combination has been observed enough times with a single deterministic outcome, the LLM call is replaced by a dict lookup. Cost drops, latency drops, accuracy rises (no model variance), and the fallback path is preserved by default.
When a prompt is revised to fix a misclassification bug, decisions made under the old prompt are now biased data — promoting a rule mined from them would lock the bug deterministically. AgentOS supports tagging each decision with a prompt_version (free-form non-empty string, conventionally a short content hash of the prompt source):
agentos decision record \
--key teams.classify_thread \
--step teams.classify_thread \
--type llm \
--input-fingerprint "$FP" \
--output-json '{"chosen":"feedback"}' \
--evidence-json '[]' \
--features-json '{"is_root": true, "has_mention": true}' \
--prompt-version "$PROMPT_HASH" \
--candidate trueMining and rule extraction then accept a matching --prompt-version filter so old data is excluded:
# Only consider decisions made under the current prompt:
agentos patterns list --by-features --prompt-version "$PROMPT_HASH"
agentos rules extract --decision-key teams.classify_thread --prompt-version "$PROMPT_HASH"When the prompt changes, bump the version. Mined data accumulated under the old version stays in the DB (useful for forensics) but never poisons the next round of rule promotion.
LLM non-determinism produces a recurring failure mode: the same input_fingerprint produces different chosen values across runs. Treating each decision as truth would seed contradictory training data. AgentOS provides a sweep command that finds these divergent groups and records latest-wins outcomes:
agentos outcome auto-correct \
--decision-key teams.classify_thread \
--within 24h \
[--prompt-version "$PROMPT_HASH"] \
[--dry-run]For each (decision_key, input_fingerprint) group within the window with at least two distinct chosen values:
- The most recent decision is recorded as
outcome=accepted. - Earlier decisions whose
chosendiffered from the latest are recorded asoutcome=rejected. - Rule mining queries already filter to
outcomes.status IN ('success', 'accepted'), so rejected decisions are automatically excluded.
The command is idempotent (re-running on the same data does nothing) via a per-correction marker stored in the outcome payload. Use --dry-run to preview corrections without writing.
--prompt-version filters the sweep to a single prompt version: cross-prompt-version divergences are usually expected (the prompt was revised) and shouldn't be auto-corrected.
This is a cheap signal — no human labelling, no external ground truth — that turns LLM disagreement-with-itself into training-data hygiene. Wire it as a periodic job alongside the rule miner.
This example shows a wrapper-first integration where you keep your existing headless flows and only add explicit decision declarations.
run-claude-headless.sh (existing worker script):
#!/usr/bin/env bash
set -euo pipefail
mkdir -p agentos-artifacts
# your existing headless invocation (placeholder)
claude-code --headless --input ci_failure.txt --output claude_output.json
# explicit declared decision for AgentOS compilation (not inferred)
cat > agentos-artifacts/decisions.json <<'JSON'
{
"decisions": [
{
"step_id": "route.fix_ci",
"decision_type": "llm",
"input_refs": ["ci_failure.txt"],
"output": {"chosen": "retry", "confidence": 0.93},
"evidence": ["failure_signature:no-unused-vars"],
"compilation_candidate": true
}
],
"outcome": {"status": "success", "pipeline": "green"}
}
JSONWrap it with AgentOS:
python -m agentos wrap \
--intent ci.fix_with_claude \
--decision-file agentos-artifacts/decisions.json \
--strict-decisions \
-- ./run-claude-headless.sh--strict-decisions ensures malformed declarations fail fast instead of silently becoming trusted compilation input.
run-codex-headless.sh (existing worker script):
#!/usr/bin/env bash
set -euo pipefail
# your existing headless invocation (placeholder)
codex exec --task "fix failing CI"
# explicit declared decision marker for AgentOS ingestion
cat <<'MARKER'
===AGENTOS_DECISION_START===
{"step_id":"route.fix_ci","decision_type":"llm","input_refs":["ci_failure.txt"],"output":{"chosen":"retry","confidence":0.91},"evidence":["matched known flaky test pattern"],"compilation_candidate":true}
===AGENTOS_DECISION_END===
MARKERWrap it with AgentOS:
python -m agentos wrap \
--intent ci.fix_with_codex \
--parse-decision-markers \
--strict-decisions \
-- ./run-codex-headless.shIf you do not want to modify model prompts, use a tiny adapter that converts existing tool output into a declaration file:
# adapter sketch (inside your existing script)
python extract_decision.py --from codex_output.json --to agentos-artifacts/decisions.json
python -m agentos wrap --intent ci.fix_with_codex --decision-file agentos-artifacts/decisions.json --strict-decisions -- ./run-codex-headless.shShort answer: yes, minimally—if you want trusted compilation candidates.
- AgentOS can capture full run traces/logs for debugging.
- But MVP intentionally does not trust passive log inference as a declared LLM decision.
- For compilation, the worker must emit an explicit declaration (decision file, stdout marker, CLI/SDK instrumentation).
Why: a free-form log line like "I think retry might work" is ambiguous and can be misread. MVP requires explicit structure to avoid guessing hidden reasoning.
Practical options:
- Prompt contract in headless mode (ask Claude/Codex to emit one strict JSON object or marker block).
- Adapter script (keep prompt unchanged, parse your tool output, then write
agentos-artifacts/decisions.json). - Direct instrumentation (
agentos decision record ...) from your existing script after each operational decision.
If your existing headless prompt is "Fix the CI failure", adapt it to include a strict decision declaration contract:
Task: Fix the CI failure.
After producing your normal output, you MUST emit exactly one AgentOS decision marker block:
===AGENTOS_DECISION_START===
{"step_id":"route.fix_ci","decision_type":"llm","input_refs":["ci_failure.txt"],"output":{"chosen":"retry|escalate|rollback","confidence":0.0},"evidence":["short evidence item"],"compilation_candidate":true}
===AGENTOS_DECISION_END===
Rules:
- output valid JSON (single object)
- confidence in [0,1]
- include at least one input reference
- do not include extra text inside the marker block
Minimal Claude Code headless shape:
claude-code --headless --prompt-file prompts/fix_ci_with_agentos_contract.txtMinimal Codex CLI headless shape:
codex exec --task-file prompts/fix_ci_with_agentos_contract.txtThis preserves wrapper-first adoption: you keep the same worker and just tighten the output contract so AgentOS captures declared decisions safely.
python -m agentos decision list --limit 20
python -m agentos runs trace <run_id>For trusted compilation candidates in MVP, verify:
decision_sourceis one of:decision_file,stdout_marker,cli_record,sdk_record.decision_validityisvalid.compilation_candidateistrue.- an associated outcome is recorded for the run.
If a decision is only visible in passive logs and was never declared through one of the explicit channels above, treat it as debug-only and do not compile it.
agentos_mvp_v0_3.md— canonical MVP product/technical specification.BACKLOG_MVP_PRIORISE_ROADMAP_6_SEMAINES.md— six-week delivery roadmap and prioritized backlog.POSITIONING.md— positioning document covering prior art, non-goals, and project rationale.VERTICAL_SLICE_MVP_RELEASE.md— release-grade end-to-end MVP walkthrough.RELEASE_MVP_CHECKLIST.md— MVP release checklist and anti-drift gates.
This repository currently documents the MVP direction and delivery plan.
Suggested first steps:
- Read the canonical spec in
agentos_mvp_v0_3.md. - Review the roadmap in
BACKLOG_MVP_PRIORISE_ROADMAP_6_SEMAINES.md. - Align implementation work with MVP scope and anti-drift constraints.
- Read
POSITIONING.mdfor positioning details and prior-art boundaries.
For MVP release hardening and sign-off, use:
VERTICAL_SLICE_MVP_RELEASE.mdfor end-to-end execution proof.RELEASE_MVP_CHECKLIST.mdfor anti-drift + quality gates.python -m agentos release checklist --jsonfor executable gate evaluation.
A first MVP bootstrap is now available in this repository:
- Python CLI scaffold (
agentos). agentos wrap --intent ... -- <command>to execute existing scripts without rewrite.- Local persistence with SQLite + JSONL traces under
.agentos/. - Basic inspection commands:
agentos runs list,agentos runs show,agentos runs trace. - Instrumentation commands:
agentos decision record|list|showandagentos outcome record. - Pattern detection command:
agentos patterns listto identify repeated decisions and compute conservative rule metrics. - Walk-forward backtest command:
agentos backtest runfor deterministic candidates before promotion. - Rule promotion command:
agentos rules promotewith fallback preserved and evidence persisted locally. - Rule rejection command:
agentos rules rejectfor explicit non-promotion decisions with recorded evidence. - MVP spec aliases are also available via
agentos compile candidates|backtest|promote|reject. - Optional config file support:
agentos.yamlwithwrap.intent,wrap.source,wrap.capture_stdout,wrap.capture_stderr, andwrap.rule_first. - Release gate command:
agentos release checklist(supports--strictand--jsonfor CI-friendly checks).
Quick local run:
python -m agentos wrap --intent demo.echo -- echo "hello"
python -m agentos decision record --run-id <run_id> --key route.fix_ci --data-json '{"chosen":"retry"}'
python -m agentos outcome record --run-id <run_id> --status success --data-json '{"ci_pipeline":"green"}'
python -m agentos runs listPattern detection example:
# record repeated decisions with a "chosen" field
python -m agentos decision record --run-id <run_id> --key route.fix_ci --data-json '{"chosen":"retry"}'
python -m agentos decision record --run-id <run_id> --key route.fix_ci --data-json '{"chosen":"retry"}'
python -m agentos decision record --run-id <run_id> --key route.fix_ci --data-json '{"chosen":"escalate"}'
# detect repeated patterns with support and abstention-aware metrics
python -m agentos patterns list --min-support 2 --limit 20
# equivalent alias aligned with MVP spec wording
python -m agentos compile candidates --min-support 2 --limit 20
# backtest one deterministic candidate with abstention constraints
python -m agentos backtest run --decision-key route.fix_ci --min-history 3 --min-confidence 0.8
# equivalent alias aligned with MVP spec wording
python -m agentos compile backtest --decision-key route.fix_ci --min-history 3 --min-confidence 0.8
# promote only if backtest metrics satisfy your threshold; fallback stays enabled
python -m agentos rules promote --decision-key route.fix_ci --min-history 3 --min-confidence 0.8 --min-accuracy 1.0
# equivalent alias aligned with MVP spec wording
python -m agentos compile promote --decision-key route.fix_ci --min-history 3 --min-confidence 0.8 --min-accuracy 1.0
# explicit rejection (records rationale + metrics snapshot)
python -m agentos rules reject --decision-key route.fix_ci --reason "manual_review_required"
python -m agentos compile reject --decision-key route.fix_ci --reason "manual_review_required"
python -m agentos rules list --limit 20Rule-first runtime (conservative by default):
# check promoted rule first, but still run fallback process by default ("observe")
python -m agentos wrap --intent demo.rule_first --rule-first --decision-key route.fix_ci -- echo "still runs fallback"
# explicit opt-in to skip fallback if a promoted rule matches
python -m agentos wrap --intent demo.rule_first --rule-first --decision-key route.fix_ci --on-rule-match skip-fallback -- echo "fallback skipped on match"Expected output shape:
{
"decision_key": "route.fix_ci",
"dominant_choice": "retry",
"support": 3,
"dominant_count": 2,
"confidence": 0.666667,
"abstain_rate": 0.333333,
"promote_ready": false
}Interpretation in MVP:
confidenceis empirical dominance of the top deterministic choice in past traces.abstain_rate(1 - confidence) quantifies uncertainty; higher values indicate conservative abstention is required.promote_ready=trueonly when confidence is exactly1.0for observed traces (fallback remains required by default).
Backtest output shape:
{
"decision_key": "route.fix_ci",
"total_observations": 6,
"min_history": 3,
"min_confidence": 0.8,
"candidate_choice": "retry",
"candidate_confidence": 0.833333,
"predictions": 2,
"abstentions": 4,
"correct_predictions": 2,
"accuracy": 1.0,
"abstain_rate": 0.666667,
"coverage_rate": 0.333333,
"promote_ready": true
}Interpretation in MVP:
predictionsandcoverage_rateshow how often a deterministic rule would actually fire.abstentionsandabstain_ratequantify safe fallback usage when confidence is insufficient.promote_ready=truemeans the rule was perfect on predicted samples; fallback still remains mandatory in MVP.
Promotion output shape:
{
"promoted": true,
"status": "promoted",
"rule_id": "rule_123abc",
"decision_key": "route.fix_ci",
"candidate_choice": "retry",
"min_accuracy": 1.0,
"fallback_enabled": true,
"metrics": {
"accuracy": 1.0,
"coverage_rate": 0.333333
}
}Interpretation in MVP:
- Promotion uses backtest metrics as evidence and refuses promotion when thresholds are not met.
fallback_enabledis always true for promoted rules in MVP.- Use
agentos rules listto review promoted rule snapshots and stored metrics before any runtime wiring.
Recommended local QA sequence:
python -m unittest -v
coverage run -m unittest
coverage report -mQuality gates for MVP implementation work:
- Include automated tests for every behavior change.
- Keep touched-module coverage high (target: >=85%).
- Review CLI examples against the current command behavior before merging.
Please read CONTRIBUTING.md before opening issues or pull requests. For future implementations, automated tests are mandatory and coverage should stay high (target: >= 85% on touched modules).
This project is licensed under the MIT License. See LICENSE.