Repository navigation
feat: tool-calling chatbot agent (LangChain + LangGraph) - #30
Merged
Merged
Conversation
Converts docs/superpowers/specs/chatbot-agent.md into the project's spec format and deletes it. Cancels the RAG design: dev/rag-chat was never merged and has been deleted (tip 16ccfd4). Decisions: LangChain v1 create_agent over Gemini, 8 tool families across all 39 METRIC_FIELDS metrics, sessions via langgraph-checkpoint-mongodb.
This line existed only on dev/rag-chat and would have been lost with rescue/rag-chat-tip. It is directly relevant to the chatbot agent, whose LLM choice is constrained to the Gemini free tier. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017qLA6eXyvqZcnyo5umxF3Z
… fallback Records three decisions taken at the Task 0 approval gate. - 7.1 call budget: one tool call by default; escalate only on four structural triggers (zero rows, two entities, two metric families, ambiguous name). Escalation is never driven by model self-assessment, which costs a turn and reports only the model's own confidence. - 7.2 correctness: no runtime self-grading. Three layers instead — groundedness by construction, a numeric-citation check with no extra model call, and an offline eval set whose ground truth is a direct Mongo query. States plainly that judgement questions have no ground truth and are covered only by the citation check. - 8: the agent never refuses. A genuine miss goes to the web and the answer is prefixed with its source by the module, not the model. The higher grounded-call volume is noted as a cost, not hidden. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017qLA6eXyvqZcnyo5umxF3Z
Done alone, ahead of any agent code, so a driver regression cannot be
mistaken for an agent bug. Test count is unchanged at 154 across the
bump; the 155th is the new config test.
- pymongo 4.10.1 -> >=4.12,<4.17 (resolves 4.16.0), required by
langgraph-checkpoint-mongodb. mongomock 4.2.0.post1 needed no change.
- Two deviations from the plan, both forced by the resolver:
- httpx[socks] pinned at ==0.28.0 silently held langchain-google-genai
back to 3.2.0, which uses the legacy google-ai-generativelanguage
SDK rather than google-genai. Relaxed to >=0.28.1,<0.29.
- Added an explicit langchain-google-genai>=4.4,<5 floor. The
langchain[google-genai] extra carries no floor, so an already-
installed 3.x satisfied it and the legacy SDK came back.
- Agent settings are config-driven, so the model names can be corrected
in one line once Task 13 checks them against the live free tier.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The PreToolUse gate carried `if: Bash(gh pr create*)` but fired on every Bash call, blocking unrelated reads. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Every family tool is built from one schema and one executor, so the METRIC_FIELDS allowlist is enforced in a single place. A metric outside the calling tool's family is rejected before any Mongo call is made. - MetricQuery carries the filters the model may set; min/max become allowlisted gte/lte clauses rather than raw query fragments. - build_metric_tool narrows `metric` to a Literal of the family's own names, so an out-of-family metric is a schema error, not a runtime one. - limit is capped at MAX_ROWS regardless of what the model asks for. - Also allow ruff commands without a prompt. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Groups METRIC_FIELDS into seven tools rather than one tool per metric, so the model picks a family and a metric name instead of choosing between 39 tool descriptions. - A coverage test asserts the families partition METRIC_FIELDS exactly: every metric reachable, none in two families. It fails if a metric is added to the domain and not routed to a family. - defending states the data gap in its own guidance: tackles, interceptions and clearances are not collected, so the model returns nothing and lets the web fallback answer, rather than substituting clean_sheets or writing a refusal. - Rows now carry sleeper_flag and low_sample_size, so the composite_scores guidance about undervalued players is true of what the tool actually returns, and unreliable rows can be flagged in an answer. - Also scope the PostToolUse doc hook: its `if` was Bash(gh pr create*), which matched every Bash call. Uses the documented prefix form now, and the prompt re-checks tool_input.command so a matcher miss cannot block an unrelated command. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The guidance said tackles, interceptions and clearances are not in the database. They are: raw_stats carries 76 Sofascore columns per competition entry, including tackles, tacklesWon, interceptions, clearances, ballRecovery, blockedShots and dribbledPast, populated in all 1256 player_stats docs. The real limit is narrower. Only 35 of those columns are mapped into the typed Stats dataclass, and METRIC_FIELDS is derived from Stats, so the defensive columns are stored but not rankable or filterable by any tool. The instruction to return nothing is unchanged; only the reason is now accurate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
# Conflicts: # .claude/settings.json
…ng tool Master promoted 13 defensive metrics into Stats, so METRIC_FIELDS grew to 52 and the family coverage test failed with 13 unrouted names — the guard working as intended. defending goes from 3 metrics to 16. The old guidance told the model these were not typed metrics and to return nothing so the web fallback could answer. That reason no longer holds: they are queryable now, and falling back to the web for data the app holds would be wrong. Rewritten to describe what the tool covers. Two things the model would otherwise get wrong are stated explicitly: dribbled_past and the two error counts are bad for the player, so low is better; and the two _pct rates have no minimum-volume guard, so a player with very few attempts can top a percentage ranking. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Adds find_player, compare_players and data_coverage, completing the tool set at 10 tools over 8 families. - An unmatched name returns an empty list, never a fabricated row, and a repository failure returns TOOL_ERROR rather than raising. Both are pinned by tests, since a confident invented player is the worst failure this agent can produce. - compare_players keeps the half it found when only one name resolves, rather than discarding both. - The guidance tells the model NOT to call find_player just to resolve a name, because every metric tool already takes player_name. That is the main way a one-call question would turn into two, which the call budget in spec 7.1 exists to prevent. - identity declares a family-level DESCRIPTION alongside its per-tool ones, so the Task 3 contract test still holds for every family rather than being weakened to accommodate this package. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The tool list is built by iterating FAMILIES, so a new family reaches the prompt without anyone editing it. A test asserts every family's GUIDANCE appears, so a family added and not described fails the suite rather than leaving the model to guess what the tool does. Encodes the three decisions from spec 7.1 and 8: - one call by default, with the four structural triggers that justify a second, and an explicit instruction not to call identity just to resolve a name — the metric tools already take player_name - never write a refusal; return empty-handed so the web fallback can answer and label its source - never reveal reasoning, tool names, arguments or raw rows Also carries the volume caveat for the new rate metrics: a percentage must be reported with the attempt count behind it, since tackles_won_pct and aerial_duels_won_pct have no minimum-volume guard. About 1080 tokens, sent on every turn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Measured the real per-turn cost first: ~3979 tokens, of which the tool JSON schemas were ~2901 (73%) and the system prompt ~1078 (27%). - MetricQuery declared `operation: Literal["rank","filter","player_value"]` which is never read anywhere. run_metric_query derives behaviour from min_value/max_value/player_name instead. It cost tokens in all seven metric schemas and invited the model to reason about a parameter with no effect. Removed. - Each family's routing text was sent twice: once as the tool description in the schema, and again as GUIDANCE in the system prompt. GUIDANCE is gone; the text now lives only in the description, which the API sends anyway. The system prompt keeps role and answering policy. Net ~3281 tokens per turn, saving ~698. Almost all of it is the system prompt shrinking from ~1078 to ~414; the schemas are flat because the descriptions absorbed the text that left the prompt. The point is that it is now sent once instead of twice. Two tests keep it that way: no family may declare GUIDANCE again, and no family DESCRIPTION may appear inside SYSTEM_PROMPT. The families are unchanged. Collapsing the seven metric tools into one would save a further ~2100 tokens but was declined: it would reverse the family grouping and drop the per-family allowlist guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
llm.py is the only module that names a provider. Everything downstream depends on BaseChatModel and bind_tools, so a future local Ollama means changing this file alone. - temperature 0, so the same question routes to the same tool rather than varying between runs - max_retries 3, using the SDK's own backoff for 429s, which the free tier will produce - an unset key is passed as None, not "". An empty string is sent as a real credential and fails with a confusing 400; None lets the SDK fall back to its own environment lookup - construction stays lazy and makes no network call, so the service still boots when Gemini is unreachable Six tests patch the constructor, so none of them would notice if an upgrade renamed a kwarg. A seventh checks the installed class directly: langchain-google-genai 4.4.0 accepts all four, subclasses BaseChatModel and exposes bind_tools. That is the Task 1 version pin being exercised rather than assumed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
ChatAgent wraps create_agent and keeps sessions as LangGraph checkpoints keyed by thread_id = session_id, so a conversation survives a restart without any session code of our own. Verified against the installed versions rather than assumed, since the plan flagged all three as uncertain: create_agent takes system_prompt (not prompt), MongoDBSaver takes client/db_name/checkpoint_collection_name, and delete_thread exists on both MongoDBSaver and InMemorySaver, so clear() needs no fallback. - A failure returns GENERIC_ERROR with degraded=True and logs the exception server-side. The API never sees a stack trace and never 500s. - history() rebuilds turns from HumanMessage and non-empty AIMessage only, so tool calls, their arguments and raw rows can never reach the user. A test asserts a tool name does not appear in replayed history. - recursion_limit is MAX_TOOL_ITERATIONS * 2, a runaway guard rather than a budget; the call budget itself lives in the system prompt. Tests use a hand-written FakeToolCallingModel because the built-in LangChain fakes raise NotImplementedError from bind_tools, which create_agent calls. InMemorySaver stands in for Mongo, so no test needs a database or a network call. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The agent never refuses. When no tool produced data the answer is
ungrounded, so a labelled web answer is preferred over the model's own
guess. The label is prepended by the module, never asked of the model, so
it cannot be dropped or reworded.
Stays on the LangChain path: langchain-google-genai 4.4.0 accepts
{"google_search": {}} as a special tool dict and converts it to
types.Tool(google_search=GoogleSearch()). Verified against the installed
library, so this module does NOT need the provider-SDK exception the plan
allowed for. A test asserts that conversion still works, because an
upgrade changing the shape would break every fallback at runtime only.
- allow_web=False disables it entirely, for callers that must stay
inside app data
- a failed or empty grounded reply returns None and the graph's own
answer is kept, so a bare label with nothing after it is impossible
- a web answer is marked degraded, so the API can tell it apart from an
answer backed by the database
Tests no longer depend on the environment: an autouse fixture disables
the fallback by default. Without it, tests asserting degraded is False
passed only because no GEMINI_API_KEY was set, and would have started
making real network calls on a machine that has one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YPKdzzdT4kZaGvTgxcKruZ
Spec 7.2 layer 2. Catches the worst failure this agent can produce, a confident invented statistic, with no extra model call. uncited_numbers compares every number in the answer against the rows the tools returned. It ignores numbers that sit inside a row's text, since a player called "Player 750" is not citing 750, and whole numbers up to 10, since "the top 5" describes the query rather than the data. Rounding is matched by distance, not by round(). Python rounds a half to even, so round(7.25, 1) is 7.2, and a model writing 7.3 for a stored 7.25 would be wrongly accused. The allowed distance follows the precision the model chose: 0.05 for one decimal, 0.5 for none. 7.5 is still flagged, because it is not a way of writing 7.25. Booleans are excluded from citable values. True is also 1 in Python, and rows carry low_sample_size, which would otherwise make 1 citable. A flag logs a warning and sets degraded; the answer is still returned. This is a signal, not a gate — withholding an answer on a false positive would be worse than flagging a real one. Tool rows come back as JSON in ToolMessage.content. Parsed with a pydantic TypeAdapter over list[dict] | dict rather than json.loads plus isinstance checks. Rows are heterogeneous by design (metric row, identity profile, coverage object, error row), so the shape is validated, not a schema. Unreadable output returns None and skips the check, because missing rows would look like missing citations and flag a correct answer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YPKdzzdT4kZaGvTgxcKruZ
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
gemini-2.5-flash and 2.0-flash are no longer served to new API keys. 3.x returns content as a list of parts, so answers are read through .text. The fallback model was configured but never used; it now runs through ModelFallbackMiddleware. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
POST /v1/chat, GET and DELETE /v1/chat/sessions/{session_id}. A failure to build the
agent (no GEMINI_API_KEY) no longer stops the app: chat returns the generic message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
… week Each graph step used to save a full copy of the conversation. Turns now save once (durability=exit) and older checkpoints are pruned, so a session is one document. The library's prune() is a stub, so pruning is our own delete, guarded by a test on the real MongoDBSaver. Documents expire 7 days after the last message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
…thread The graph state holds every message in the session, so from the second turn on a tool call or row from an earlier turn skipped the web fallback and supported citations. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
apiFetch now accepts 204 responses, so session deletion reuses it instead of a second fetch path. The session id survives blocked storage and non-secure contexts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vqcRbCNpjKLZmMaYe7RS6
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Gemini answers in markdown, so bold and lists showed as raw ** and run-on text. react-markdown escapes HTML, and only assistant turns are parsed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The model writes "3,380 minutes"; the citation check read 3 and 380, so a fully grounded answer was marked as unverified in the UI. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
App splits into a route switch and the dashboard, so the early return does not call hooks conditionally. The full-screen page carries a link back and no floating widget. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The provider is a strategy chosen by LLM_PROVIDER, with gemini's free tier as the default and LLM_MODEL/LLM_API_KEY for anything else; openai and anthropic load their langchain package lazily and say what to install. Google Search grounding has no free quota, so the web fallback is gone: the prompt now answers outside questions from the model's own knowledge, labelled. Default model is gemini-3.5-flash because 3.6-flash allows only 20 free requests per day. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- limit must be >= 1: Mongo reads limit(0) as no limit, bypassing MAX_ROWS - add a sort direction, so "fewest" questions on goals_conceded, dribbled_past and the error metrics no longer return the worst players as the answer - escape the name filter: it is user text, so "." must not match every player - read numbers nested in lists and dicts, so data_coverage answers stop being flagged as ungrounded, and ignore markdown list markers past 10 - mark an answer no tool backed as not from the app's data - hide model preamble that came with a tool call from replayed history - clearing a session no longer 500s when storage fails - the eval no longer passes [None] as a fallback model - keep an in-flight turn when the history request resolves late - drop the unused NO_DATA constant Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements
docs/superpowers/specs/2026-09-02-chatbot-agent-design.md. Replaces the cancelled RAG design: no embeddings, no vector store — the model calls typed tools that read this app's own MongoDB.What this adds
Backend
app/agent/agent.py— LangGraphcreate_agentloop, one checkpointed session persession_id.providers.py— model provider strategy + registry (gemini,openai,anthropic).openai/anthropicimport their LangChain package lazily and require an explicitLLM_MODEL.tools/— 8 metric-family tools plusfind_player,compare_players,data_coverage, all built over theMETRIC_FIELDSallowlist (52 metrics), which is also the injection guard.answer_check.py— flags figures in an answer that no tool row supports (degraded), without blocking the answer.checkpoints.py— one document per session, pruned after each turn, TTL-expired 7 days after the last message.API —
POST /v1/chat,GET/DELETE /v1/chat/sessions/{session_id}. No field can carry a tool trace. A missing API key disables the agent but does not stop the app from starting.Frontend — floating chat widget on every tab, full-screen view at
?chat=1, markdown rendering (react-markdown). No navbar tab, by design.Configuration
Free by default: with only
GEMINI_API_KEYset, the agent uses Gemini's free tier.LLM_PROVIDER,LLM_MODEL,LLM_API_KEYandLLM_FALLBACK_MODEL(empty = no fallback) select anything else. See README → Chat Agent Configuration.Verification
tsc --noEmitclean.AGENT_EVAL=1, real model + real DB, ground truth computed from the repository): 6/7. The one failure was a free-tier rate limit, and that case passes when re-run alone. Ground truth is derived fromrepo.get_players, never hand-copied.DELETEclearing a thread, and continuity between the widget and?chat=1./code-review(high effort) on this whole diff: 11 findings, all fixed in2a2c004— includinglimit: 0bypassing the row cap (Mongo reads it as "no limit"), a missing sort direction that made "fewest" questions return the worst players, and an unescaped$regexon the name filter.Decisions worth knowing
LLM_FALLBACK_MODELor a paidLLM_MODELlifts this.After merging
Rebuild the containers once:
docker compose build backend, thendocker compose up -d --force-recreate --renew-anon-volumes frontend(itsnode_moduleslives in an anonymous volume).🤖 Generated with Claude Code