Skip to content

feat: tool-calling chatbot agent (LangChain + LangGraph) - #30

Merged
YonatanHen merged 32 commits into
masterfrom
dev/chatbot-agent
Sep 22, 2026
Merged

YonatanHen merged 32 commits into
masterfrom
dev/chatbot-agent

Conversation

@YonatanHen

Copy link
Copy Markdown
Owner

Implements docs/superpowers/specs/2026-09-02-chatbot-agent-design.md. Replaces the cancelled RAG design: no embeddings, no vector store — the model calls typed tools that read this app's own MongoDB.

What this adds

Backend app/agent/

  • agent.py — LangGraph create_agent loop, one checkpointed session per session_id.
  • providers.py — model provider strategy + registry (gemini, openai, anthropic). openai/anthropic import their LangChain package lazily and require an explicit LLM_MODEL.
  • tools/ — 8 metric-family tools plus find_player, compare_players, data_coverage, all built over the METRIC_FIELDS allowlist (52 metrics), which is also the injection guard.
  • answer_check.py — flags figures in an answer that no tool row supports (degraded), without blocking the answer.
  • checkpoints.py — one document per session, pruned after each turn, TTL-expired 7 days after the last message.

API — POST /v1/chat, GET/DELETE /v1/chat/sessions/{session_id}. No field can carry a tool trace. A missing API key disables the agent but does not stop the app from starting.

Frontend — floating chat widget on every tab, full-screen view at ?chat=1, markdown rendering (react-markdown). No navbar tab, by design.

Configuration

Free by default: with only GEMINI_API_KEY set, the agent uses Gemini's free tier. LLM_PROVIDER, LLM_MODEL, LLM_API_KEY and LLM_FALLBACK_MODEL (empty = no fallback) select anything else. See README → Chat Agent Configuration.

Verification

  • Tests: 273 pass, 7 skipped (the opt-in eval). Ruff lint + format clean, tsc --noEmit clean.
  • Offline eval (AGENT_EVAL=1, real model + real DB, ground truth computed from the repository): 6/7. The one failure was a free-tier rate limit, and that case passes when re-run alone. Ground truth is derived from repo.get_players, never hand-copied.
  • Live through the UI and the API: top-5 forwards by goals, most clean sheets (David Raya, 28), most tackles by a defender (Truffert, 103), most assists at Arsenal (Trossard, 10), a two-player comparison, a "fewest dribbled past" question, session history restored after a page reload, DELETE clearing a thread, and continuity between the widget and ?chat=1.
  • Outside the database ("Which country won the 2018 FIFA World Cup?") → answered from the model's own knowledge, prefixed "Not from the app's data:", never a refusal.
  • /code-review (high effort) on this whole diff: 11 findings, all fixed in 2a2c004 — including limit: 0 bypassing the row cap (Mongo reads it as "no limit"), a missing sort direction that made "fewest" questions return the worst players, and an unescaped $regex on the name filter.

Decisions worth knowing

  • No web fallback. Google Search grounding returns 429 on the free tier, so the module was removed; outside questions are answered from model knowledge with a label instead.
  • Free-tier limit: each Gemini free model allows about 20 requests per day, which is roughly 10–20 chat questions. LLM_FALLBACK_MODEL or a paid LLM_MODEL lifts this.
  • Session storage: one document per session instead of one per graph step, which was about 3.5x smaller in a measured two-turn chat.

After merging

Rebuild the containers once: docker compose build backend, then docker compose up -d --force-recreate --renew-anon-volumes frontend (its node_modules lives in an anonymous volume).

🤖 Generated with Claude Code

YonatanHen and others added 30 commits September 2, 2026 16:19
Converts docs/superpowers/specs/chatbot-agent.md into the project's spec
format and deletes it. Cancels the RAG design: dev/rag-chat was never
merged and has been deleted (tip 16ccfd4).

Decisions: LangChain v1 create_agent over Gemini, 8 tool families across
all 39 METRIC_FIELDS metrics, sessions via langgraph-checkpoint-mongodb.
This line existed only on dev/rag-chat and would have been lost with
rescue/rag-chat-tip. It is directly relevant to the chatbot agent, whose
LLM choice is constrained to the Gemini free tier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qLA6eXyvqZcnyo5umxF3Z
… fallback

Records three decisions taken at the Task 0 approval gate.

- 7.1 call budget: one tool call by default; escalate only on four
  structural triggers (zero rows, two entities, two metric families,
  ambiguous name). Escalation is never driven by model self-assessment,
  which costs a turn and reports only the model's own confidence.
- 7.2 correctness: no runtime self-grading. Three layers instead —
  groundedness by construction, a numeric-citation check with no extra
  model call, and an offline eval set whose ground truth is a direct
  Mongo query. States plainly that judgement questions have no ground
  truth and are covered only by the citation check.
- 8: the agent never refuses. A genuine miss goes to the web and the
  answer is prefixed with its source by the module, not the model. The
  higher grounded-call volume is noted as a cost, not hidden.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qLA6eXyvqZcnyo5umxF3Z
Done alone, ahead of any agent code, so a driver regression cannot be
mistaken for an agent bug. Test count is unchanged at 154 across the
bump; the 155th is the new config test.

- pymongo 4.10.1 -> >=4.12,<4.17 (resolves 4.16.0), required by
  langgraph-checkpoint-mongodb. mongomock 4.2.0.post1 needed no change.
- Two deviations from the plan, both forced by the resolver:
  - httpx[socks] pinned at ==0.28.0 silently held langchain-google-genai
    back to 3.2.0, which uses the legacy google-ai-generativelanguage
    SDK rather than google-genai. Relaxed to >=0.28.1,<0.29.
  - Added an explicit langchain-google-genai>=4.4,<5 floor. The
    langchain[google-genai] extra carries no floor, so an already-
    installed 3.x satisfied it and the legacy SDK came back.
- Agent settings are config-driven, so the model names can be corrected
  in one line once Task 13 checks them against the live free tier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The PreToolUse gate carried `if: Bash(gh pr create*)` but fired on every
Bash call, blocking unrelated reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Every family tool is built from one schema and one executor, so the
METRIC_FIELDS allowlist is enforced in a single place. A metric outside
the calling tool's family is rejected before any Mongo call is made.

- MetricQuery carries the filters the model may set; min/max become
  allowlisted gte/lte clauses rather than raw query fragments.
- build_metric_tool narrows `metric` to a Literal of the family's own
  names, so an out-of-family metric is a schema error, not a runtime one.
- limit is capped at MAX_ROWS regardless of what the model asks for.
- Also allow ruff commands without a prompt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Groups METRIC_FIELDS into seven tools rather than one tool per metric, so
the model picks a family and a metric name instead of choosing between 39
tool descriptions.

- A coverage test asserts the families partition METRIC_FIELDS exactly:
  every metric reachable, none in two families. It fails if a metric is
  added to the domain and not routed to a family.
- defending states the data gap in its own guidance: tackles,
  interceptions and clearances are not collected, so the model returns
  nothing and lets the web fallback answer, rather than substituting
  clean_sheets or writing a refusal.
- Rows now carry sleeper_flag and low_sample_size, so the
  composite_scores guidance about undervalued players is true of what the
  tool actually returns, and unreliable rows can be flagged in an answer.
- Also scope the PostToolUse doc hook: its `if` was Bash(gh pr create*),
  which matched every Bash call. Uses the documented prefix form now, and
  the prompt re-checks tool_input.command so a matcher miss cannot block
  an unrelated command.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The guidance said tackles, interceptions and clearances are not in the
database. They are: raw_stats carries 76 Sofascore columns per
competition entry, including tackles, tacklesWon, interceptions,
clearances, ballRecovery, blockedShots and dribbledPast, populated in all
1256 player_stats docs.

The real limit is narrower. Only 35 of those columns are mapped into the
typed Stats dataclass, and METRIC_FIELDS is derived from Stats, so the
defensive columns are stored but not rankable or filterable by any tool.
The instruction to return nothing is unchanged; only the reason is now
accurate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
# Conflicts:
#	.claude/settings.json
…ng tool

Master promoted 13 defensive metrics into Stats, so METRIC_FIELDS grew to
52 and the family coverage test failed with 13 unrouted names — the guard
working as intended. defending goes from 3 metrics to 16.

The old guidance told the model these were not typed metrics and to
return nothing so the web fallback could answer. That reason no longer
holds: they are queryable now, and falling back to the web for data the
app holds would be wrong. Rewritten to describe what the tool covers.

Two things the model would otherwise get wrong are stated explicitly:
dribbled_past and the two error counts are bad for the player, so low is
better; and the two _pct rates have no minimum-volume guard, so a player
with very few attempts can top a percentage ranking.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Adds find_player, compare_players and data_coverage, completing the tool
set at 10 tools over 8 families.

- An unmatched name returns an empty list, never a fabricated row, and a
  repository failure returns TOOL_ERROR rather than raising. Both are
  pinned by tests, since a confident invented player is the worst
  failure this agent can produce.
- compare_players keeps the half it found when only one name resolves,
  rather than discarding both.
- The guidance tells the model NOT to call find_player just to resolve a
  name, because every metric tool already takes player_name. That is the
  main way a one-call question would turn into two, which the call budget
  in spec 7.1 exists to prevent.
- identity declares a family-level DESCRIPTION alongside its per-tool
  ones, so the Task 3 contract test still holds for every family rather
  than being weakened to accommodate this package.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The tool list is built by iterating FAMILIES, so a new family reaches the
prompt without anyone editing it. A test asserts every family's GUIDANCE
appears, so a family added and not described fails the suite rather than
leaving the model to guess what the tool does.

Encodes the three decisions from spec 7.1 and 8:
- one call by default, with the four structural triggers that justify a
  second, and an explicit instruction not to call identity just to
  resolve a name — the metric tools already take player_name
- never write a refusal; return empty-handed so the web fallback can
  answer and label its source
- never reveal reasoning, tool names, arguments or raw rows

Also carries the volume caveat for the new rate metrics: a percentage
must be reported with the attempt count behind it, since tackles_won_pct
and aerial_duels_won_pct have no minimum-volume guard.

About 1080 tokens, sent on every turn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
Measured the real per-turn cost first: ~3979 tokens, of which the tool
JSON schemas were ~2901 (73%) and the system prompt ~1078 (27%).

- MetricQuery declared `operation: Literal["rank","filter","player_value"]`
  which is never read anywhere. run_metric_query derives behaviour from
  min_value/max_value/player_name instead. It cost tokens in all seven
  metric schemas and invited the model to reason about a parameter with
  no effect. Removed.
- Each family's routing text was sent twice: once as the tool description
  in the schema, and again as GUIDANCE in the system prompt. GUIDANCE is
  gone; the text now lives only in the description, which the API sends
  anyway. The system prompt keeps role and answering policy.

Net ~3281 tokens per turn, saving ~698. Almost all of it is the system
prompt shrinking from ~1078 to ~414; the schemas are flat because the
descriptions absorbed the text that left the prompt. The point is that it
is now sent once instead of twice.

Two tests keep it that way: no family may declare GUIDANCE again, and no
family DESCRIPTION may appear inside SYSTEM_PROMPT.

The families are unchanged. Collapsing the seven metric tools into one
would save a further ~2100 tokens but was declined: it would reverse the
family grouping and drop the per-family allowlist guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
llm.py is the only module that names a provider. Everything downstream
depends on BaseChatModel and bind_tools, so a future local Ollama means
changing this file alone.

- temperature 0, so the same question routes to the same tool rather than
  varying between runs
- max_retries 3, using the SDK's own backoff for 429s, which the free
  tier will produce
- an unset key is passed as None, not "". An empty string is sent as a
  real credential and fails with a confusing 400; None lets the SDK fall
  back to its own environment lookup
- construction stays lazy and makes no network call, so the service still
  boots when Gemini is unreachable

Six tests patch the constructor, so none of them would notice if an
upgrade renamed a kwarg. A seventh checks the installed class directly:
langchain-google-genai 4.4.0 accepts all four, subclasses BaseChatModel
and exposes bind_tools. That is the Task 1 version pin being exercised
rather than assumed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
ChatAgent wraps create_agent and keeps sessions as LangGraph checkpoints
keyed by thread_id = session_id, so a conversation survives a restart
without any session code of our own.

Verified against the installed versions rather than assumed, since the
plan flagged all three as uncertain: create_agent takes system_prompt
(not prompt), MongoDBSaver takes client/db_name/checkpoint_collection_name,
and delete_thread exists on both MongoDBSaver and InMemorySaver, so
clear() needs no fallback.

- A failure returns GENERIC_ERROR with degraded=True and logs the
  exception server-side. The API never sees a stack trace and never 500s.
- history() rebuilds turns from HumanMessage and non-empty AIMessage
  only, so tool calls, their arguments and raw rows can never reach the
  user. A test asserts a tool name does not appear in replayed history.
- recursion_limit is MAX_TOOL_ITERATIONS * 2, a runaway guard rather than
  a budget; the call budget itself lives in the system prompt.

Tests use a hand-written FakeToolCallingModel because the built-in
LangChain fakes raise NotImplementedError from bind_tools, which
create_agent calls. InMemorySaver stands in for Mongo, so no test needs a
database or a network call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKAXmoCoJ3gAu4yyD3zkCL
The agent never refuses. When no tool produced data the answer is
ungrounded, so a labelled web answer is preferred over the model's own
guess. The label is prepended by the module, never asked of the model, so
it cannot be dropped or reworded.

Stays on the LangChain path: langchain-google-genai 4.4.0 accepts
{"google_search": {}} as a special tool dict and converts it to
types.Tool(google_search=GoogleSearch()). Verified against the installed
library, so this module does NOT need the provider-SDK exception the plan
allowed for. A test asserts that conversion still works, because an
upgrade changing the shape would break every fallback at runtime only.

- allow_web=False disables it entirely, for callers that must stay
  inside app data
- a failed or empty grounded reply returns None and the graph's own
  answer is kept, so a bare label with nothing after it is impossible
- a web answer is marked degraded, so the API can tell it apart from an
  answer backed by the database

Tests no longer depend on the environment: an autouse fixture disables
the fallback by default. Without it, tests asserting degraded is False
passed only because no GEMINI_API_KEY was set, and would have started
making real network calls on a machine that has one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YPKdzzdT4kZaGvTgxcKruZ
Spec 7.2 layer 2. Catches the worst failure this agent can produce, a
confident invented statistic, with no extra model call.

uncited_numbers compares every number in the answer against the rows the
tools returned. It ignores numbers that sit inside a row's text, since a
player called "Player 750" is not citing 750, and whole numbers up to 10,
since "the top 5" describes the query rather than the data.

Rounding is matched by distance, not by round(). Python rounds a half to
even, so round(7.25, 1) is 7.2, and a model writing 7.3 for a stored 7.25
would be wrongly accused. The allowed distance follows the precision the
model chose: 0.05 for one decimal, 0.5 for none. 7.5 is still flagged,
because it is not a way of writing 7.25.

Booleans are excluded from citable values. True is also 1 in Python, and
rows carry low_sample_size, which would otherwise make 1 citable.

A flag logs a warning and sets degraded; the answer is still returned.
This is a signal, not a gate — withholding an answer on a false positive
would be worse than flagging a real one.

Tool rows come back as JSON in ToolMessage.content. Parsed with a pydantic
TypeAdapter over list[dict] | dict rather than json.loads plus isinstance
checks. Rows are heterogeneous by design (metric row, identity profile,
coverage object, error row), so the shape is validated, not a schema.
Unreadable output returns None and skips the check, because missing rows
would look like missing citations and flag a correct answer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YPKdzzdT4kZaGvTgxcKruZ
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
gemini-2.5-flash and 2.0-flash are no longer served to new API keys. 3.x returns
content as a list of parts, so answers are read through .text. The fallback model
was configured but never used; it now runs through ModelFallbackMiddleware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
POST /v1/chat, GET and DELETE /v1/chat/sessions/{session_id}. A failure to build the
agent (no GEMINI_API_KEY) no longer stops the app: chat returns the generic message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
… week

Each graph step used to save a full copy of the conversation. Turns now save once
(durability=exit) and older checkpoints are pruned, so a session is one document.
The library's prune() is a stub, so pruning is our own delete, guarded by a test on
the real MongoDBSaver. Documents expire 7 days after the last message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
…thread

The graph state holds every message in the session, so from the second turn on a tool
call or row from an earlier turn skipped the web fallback and supported citations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eei5ktbTcs7MNFq5TzoiYb
apiFetch now accepts 204 responses, so session deletion reuses it instead of a second
fetch path. The session id survives blocked storage and non-secure contexts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vqcRbCNpjKLZmMaYe7RS6
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Gemini answers in markdown, so bold and lists showed as raw ** and run-on text.
react-markdown escapes HTML, and only assistant turns are parsed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The model writes "3,380 minutes"; the citation check read 3 and 380, so a fully
grounded answer was marked as unverified in the UI.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
App splits into a route switch and the dashboard, so the early return does not call
hooks conditionally. The full-screen page carries a link back and no floating widget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The provider is a strategy chosen by LLM_PROVIDER, with gemini's free tier as the
default and LLM_MODEL/LLM_API_KEY for anything else; openai and anthropic load their
langchain package lazily and say what to install. Google Search grounding has no free
quota, so the web fallback is gone: the prompt now answers outside questions from the
model's own knowledge, labelled. Default model is gemini-3.5-flash because 3.6-flash
allows only 20 free requests per day.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
YonatanHen and others added 2 commits September 22, 2026 15:09
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- limit must be >= 1: Mongo reads limit(0) as no limit, bypassing MAX_ROWS
- add a sort direction, so "fewest" questions on goals_conceded, dribbled_past
  and the error metrics no longer return the worst players as the answer
- escape the name filter: it is user text, so "." must not match every player
- read numbers nested in lists and dicts, so data_coverage answers stop being
  flagged as ungrounded, and ignore markdown list markers past 10
- mark an answer no tool backed as not from the app's data
- hide model preamble that came with a tool call from replayed history
- clearing a session no longer 500s when storage fails
- the eval no longer passes [None] as a fallback model
- keep an in-flight turn when the history request resolves late
- drop the unused NO_DATA constant

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@YonatanHen
YonatanHen merged commit 9a9e2ca into master Sep 22, 2026
5 checks passed
@YonatanHen
YonatanHen deleted the dev/chatbot-agent branch September 23, 2026 12:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant