Skip to content

Latest commit

 

History

36 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AgentSwitch Web Agent

An autonomous agent for the Website seat of AgentSwitch, built for the EAG V3 capstone (The School of AI). The agent drives a live, multi-tenant business platform over MCP to do real work, and ships with its own harness that proves the work happened by reading the database, not by trusting the agent's words.

Team A9 · Seat 09 · "RA9 Web Agent: Publishing and Page Analytics"


The one job

The platform judges our agent on a single seat-specific request:

"Publish a post about the new fixture line, and tell me which pages nobody reads."

Two evaluator goals back it: web.publish_post and web.pages_without_traffic. Grading reads database state after the agent runs — a reply that sounds right but writes the wrong row fails. Several steps, a judgement call, and a state change (publishing) in the middle.


How the pieces connect

flowchart LR
  subgraph local["Your machine (this repo)"]
    agent["Web Agent<br/>our code + prompts"]
    harness["Harness<br/>tasks + DB-reading verifiers"]
    probe["probe.py<br/>read-only smoke test"]
    bruno["Bruno collection<br/>manual API calls"]
  end
  model[("Your LLM provider<br/>Gemini / OpenAI / Anthropic")]
  subgraph hosted["AgentSwitch — hosted by the course"]
    mcp["POST /api/mcp<br/>JSON-RPC 2.0"]
    rest["REST + app endpoints"]
    db[("Shared database<br/>all 35 seats, one tenant")]
    eval["Evaluator predicates"]
  end

  agent -->|"reasoning"| model
  agent -->|"login -> Bearer token, then tools/call"| mcp
  harness -->|"runs"| agent
  harness -->|"reads state to verify"| rest
  harness --> eval
  probe -->|"read-only"| mcp
  bruno -->|"by hand"| mcp
  mcp --> db
  rest --> db
  eval --> db
Loading

Full diagrams (auth flow + the graded-task data flow) are in docs/ARCHITECTURE.md.


The agent — a small hierarchy

The request flows through a bounded-determinism graph: a supervisor routes the planned goals to a roster of workers; each real goal is a sub-agent (a subgraph) that gets only a scoped slice of state, runs its own internal DAG, and returns only its declared outputs — its scratch never leaks back. Hand-rolled on our own DAG (stdlib + the Anthropic SDK, no framework).

respond()
  → plan (LLM → goals)
  → Supervisor routes the roster (independents run in parallel):
       publish     (subgraph)  perceive → decide → act → re-read   [idempotent + compensation]
       dead_pages  (subgraph)  fetch pages ‖ menus → detect → assemble   [scoped, isolated]
       refuse      (function)  honest refusal when the data can't support the ask
  → verify (re-read the DB — truth, not the agent's words)
  → compose
File Role
agent/supervisor.py the "daddy": routes goals to workers, then verify; bounded to routing + filling
agent/subgraph.py + agent/subgraphs.py sub-agent primitive + the publish / dead_pages sub-agents (scoped I/O)
agent/state.py typed blackboard — named channels + reducers + scope() / absorb()
agent/dag.py the engine — topological run, parallel independents, per-node timing trace
agent/reliability.py retries · JSON-repair · per-run budget + circuit breaker
agent/verify.py re-read the DB and confirm; LLM-judge fallback; refusal
agent/memory.py 3-tier memory backed by the platform AgentMemory
harness/ scored task runner · adapter.py integration seam · daily_hunt.py (read-only probe-and-draft)

Repository layout

agentswitch-web-agent/
├── README.md                     # you are here
├── docs/
│   ├── SETUP.md                  # dot-by-dot of everything we set up locally
│   ├── ARCHITECTURE.md           # the graph: components, auth flow, data flow
│   └── PLAN.md                   # the four-week plan, week by week
├── notes/
│   └── website-capabilities.md   # live capability map (236 tools, our 36, the gaps)
├── bruno/                        # git-friendly API client collection (.bru files)
│   ├── bruno.json
│   ├── environments/             # Suryodaya (India) + Keystone (US)
│   └── *.bru                     # login, auth-me, MCP handshake, tools/call examples
├── probe.py                      # zero-dep read-only connection test + capability dump
├── agent/                        # the agent: supervisor, subgraphs, state, dag, verify, memory…
├── harness/                      # scored tasks + DB-reading verifiers + adapter + daily_hunt
├── .env.example                  # copy to .env (git-ignored) and fill locally
└── .gitignore

Quickstart

# 1. Config (secrets stay local — .env is git-ignored)
cp .env.example .env          # then paste your team password into AS_PASSWORD=

# 2. Prove the connection (read-only; never prints your token)
set -a; source .env; set +a
python3 probe.py

# 3. Poke it by hand: open the bruno/ folder in Bruno, pick the Suryodaya
#    environment, set the `password` secret, run "01 Login", then anything else.

See docs/SETUP.md for the full walk-through.


What we build vs. what already exists

Already built by the course (we do NOT build it) We build
The platform: 424 entity types, CRUD, auth, approval engine, audit trail, job runner, UI The agent for our seat
The MCP + REST interfaces over all of it The harness: tasks + verifiers that read the DB
The shared database and the evaluator predicates Our own prompts, tools, planning, model choice

Effort split to remember: the agent is the smaller, known part (a perceive -> decide -> act loop); the harness is the bigger, graded part that most teams underbuild.


Security

  • No credentials live in this repo, ever. Passwords and tokens are per-team secrets.
  • .env (shell) and Bruno secret variables are git-ignored / stored outside the repo.
  • Every write to the platform is attributed to whoever is signed in, so the login is guarded like a key.

About

Autonomous agent for the AgentSwitch Website seat (EAG V3 capstone, Team A9): drives a live business platform over MCP, with a database-verifying harness.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages