Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
32 commits
Select commit Hold shift + click to select a range
39e72eb
docs(chatbot-agent): design spec replacing the hand-written note
YonatanHen Sep 2, 2026
c2875d8
chore: merge master (scoring overhaul) into dev/chatbot-agent
YonatanHen Sep 3, 2026
b1c24a7
docs: restore the free-tier constraint from the deleted RAG branch
YonatanHen Sep 3, 2026
b38fe5a
docs(chatbot-agent): call budget, correctness measurement, no-refusal…
YonatanHen Sep 3, 2026
0689cc3
chore(chatbot-agent): bump pymongo, add LangChain deps and agent config
YonatanHen Sep 6, 2026
d80050c
chore(claude): allow pytest without a prompt, drop the misfiring PR hook
YonatanHen Sep 6, 2026
9007a3c
feat(chatbot-agent): shared metric tool schema and query executor
YonatanHen Sep 6, 2026
2831114
feat(chatbot-agent): seven metric-family tools covering all 39 metrics
YonatanHen Sep 6, 2026
4ed7518
fix(chatbot-agent): correct the defending data-gap claim
YonatanHen Sep 6, 2026
9d16afe
Merge branch 'master' into dev/chatbot-agent
YonatanHen Sep 6, 2026
55d55e4
feat(chatbot-agent): route the new defensive metrics into the defendi…
YonatanHen Sep 6, 2026
b52a3eb
feat(chatbot-agent): identity tools for lookup, compare and coverage
YonatanHen Sep 6, 2026
c1001b6
feat(chatbot-agent): system prompt generated from the tool registry
YonatanHen Sep 6, 2026
f616358
perf(chatbot-agent): stop sending the same routing text twice per turn
YonatanHen Sep 6, 2026
65521e2
feat(chatbot-agent): Gemini chat model with retry and fallback chain
YonatanHen Sep 6, 2026
499dab1
feat(chatbot-agent): LangGraph agent with Mongo-checkpointed sessions
YonatanHen Sep 6, 2026
df2d247
feat(chatbot-agent): web-grounded fallback when no tool answers
YonatanHen Sep 12, 2026
bb4c244
feat(chatbot-agent): flag answer figures that no tool row supports
YonatanHen Sep 12, 2026
31bcb75
test(chatbot-agent): offline eval set with computed ground truth
YonatanHen Sep 14, 2026
33e3527
fix(chatbot-agent): move to Gemini 3.x and wire the fallback model
YonatanHen Sep 14, 2026
d6c32a1
feat(chatbot-agent): chat API, session endpoints and DI wiring
YonatanHen Sep 14, 2026
4f8fe07
feat(chatbot-agent): one checkpoint per chat session, expired after a…
YonatanHen Sep 14, 2026
3999cc6
fix(chatbot-agent): judge each answer by its own turn, not the whole …
YonatanHen Sep 14, 2026
83fded9
feat(chatbot-agent): frontend chat api client
YonatanHen Sep 15, 2026
6913d5f
feat(chatbot-agent): floating collapsible chat widget
YonatanHen Sep 22, 2026
cd9f819
feat(chatbot-agent): render markdown answers in the chat panel
YonatanHen Sep 22, 2026
343a89d
fix(chatbot-agent): read thousand separators as one number
YonatanHen Sep 22, 2026
8a77db9
feat(chatbot-agent): full-screen chat view via ?chat=1
YonatanHen Sep 22, 2026
d27cfa7
feat(chatbot-agent): configurable model provider, no web fallback
YonatanHen Sep 22, 2026
ef96df8
chore: drop the local pytest and ruff permission allowlist
YonatanHen Sep 22, 2026
f018307
docs: document the chatbot agent, its endpoints and configuration
YonatanHen Sep 22, 2026
2a2c004
fix(chatbot-agent): address code review findings
YonatanHen Sep 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 0 additions & 16 deletions .claude/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,22 +3,6 @@
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
},
"teammateMode": "in-process",
"permissions": {
"allow": [
"Bash(pytest:*)",
"Bash(python -m pytest:*)",
"Bash(.venv/Scripts/python -m pytest:*)",
"Bash(.venv\\Scripts\\python -m pytest:*)",
"Bash(backend/.venv/Scripts/python -m pytest:*)",
"Bash(backend\\.venv\\Scripts\\python -m pytest:*)",
"Bash(ruff:*)",
"Bash(python -m ruff:*)",
"Bash(.venv/Scripts/python -m ruff:*)",
"Bash(.venv\\Scripts\\python -m ruff:*)",
"Bash(backend/.venv/Scripts/python -m ruff:*)",
"Bash(backend\\.venv\\Scripts\\python -m ruff:*)"
]
},
"hooks": {
"PreToolUse": [
{
Expand Down
37 changes: 33 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
- **Frontend**: React 18 + TypeScript + Vite + Tailwind CSS, runs on port 5173
- **Database**: MongoDB 7 (`football_analytics` db)
- **Scraping**: botasaurus + Chrome (inside Docker) for Sofascore; Tor is present but Sofascore's Cloudflare 403s it, so Chrome scrapes run without the Tor proxy
- **Chatbot agent**: LangChain + LangGraph (`create_agent`), configurable LLM provider (`app/agent/providers.py`); Gemini free tier by default

## Running the project

Expand All @@ -22,10 +23,12 @@ uvicorn app.main:app --reload
npm run dev
```

Required env file: `secrets.env` in project root (loaded by Docker). Backend also reads `.env` for local dev. Key variables: `MONGO_URI`, `CORS_ORIGINS`.
Required env file: `secrets.env` in project root (loaded by Docker). Backend also reads `.env` for local dev. Key variables: `MONGO_URI`, `CORS_ORIGINS`, `GEMINI_API_KEY` (chat agent; missing key disables the agent but not the rest of the API). Chat agent tuning: `LLM_PROVIDER`, `LLM_MODEL`, `LLM_API_KEY`, `LLM_FALLBACK_MODEL`, `AGENT_MAX_TOOL_ITERATIONS`, `AGENT_MAX_ROWS`, `CHECKPOINT_COLLECTION`, `CHAT_SESSION_TTL_SECONDS` — see `backend/app/config.py`.

Data loading is done via the `tools/fetch_cli` developer CLI, not the UI — see `tools/fetch_cli/README.md`.

After pulling changes touching backend or frontend dependencies: `docker compose build backend`, and `docker compose up -d --force-recreate --renew-anon-volumes frontend` (frontend `node_modules` lives in an anonymous volume).

## Tests

```bash
Expand All @@ -48,6 +51,8 @@ No frontend tests currently.
.venv/Scripts/python -m pytest fetch_cli/tests
```

Agent tests: `backend/tests/agent/`. The offline eval in `backend/tests/agent/eval/` is skipped by default (needs a live model + populated DB) — run with `AGENT_EVAL=1 pytest tests/agent/eval -v -s`.

## DB snapshots

Run inside the backend container (or locally with the stack up):
Expand All @@ -70,6 +75,9 @@ app/
# GET /v1/fetch/competitions, /seasons, /fetched — catalog + fetched-state reads for tools/fetch_cli
players.py # GET /v1/players, GET /v1/players/{id}
analysis.py # GET /v1/analysis/scatter
chat.py # POST /v1/chat, GET/DELETE /v1/chat/sessions/{session_id}
modals/
chat_modals.py # ChatRequest, ChatResponse, ChatTurn, ChatHistory
modes/ # Strategy pattern
base.py # AnalysisMode ABC: fetch_data(), process()
factory.py # ModeFactory.create("fantasy")
Expand All @@ -79,14 +87,29 @@ app/
models.py # PlayerDTO, Stats, Score, CompetitionEntry, AggregatedScores
scoring_engine.py
sleeper_detector.py
player_assembler.py # build_player(), merge(), aggregate_stats()
competitions.py # canonical_competition() — normalizes competition names
defensive_stats.py # Defensive metric computation, feeds Stats
player_assembler.py # build_player(), merge(), aggregate_stats()
competitions.py # canonical_competition() — normalizes competition names
metric_fields.py # METRIC_FIELDS — field registry consumed by agent tools
infrastructure/
mongo_repository.py # All MongoDB I/O; serializes/deserializes domain models
sofascore_client.py # Fetches from Sofascore via ScraperFC (Chrome/botasaurus)
text_utils.py # normalize_text() for name/team fuzzy matching
agent/ # Chatbot: LangGraph tool-calling loop over the DB
agent.py # ChatAgent, build_agent() — create_agent loop, checkpointed per session
providers.py # ModelProvider strategy + PROVIDERS registry (gemini/openai/anthropic)
llm.py # build_chat_model(), build_fallback_model()
system_prompt.py # SYSTEM_PROMPT
answer_check.py # uncited_numbers() — flags figures not backed by a tool row
checkpoints.py # MongoDBSaver wiring; keep_latest_checkpoint() prunes old checkpoints
constants.py # MAX_TOOL_ITERATIONS, MAX_ROWS, error strings
tools/
base.py # MetricQuery schema, run_metric_query(), build_metric_tool()
identity/ # find_player, compare_players, data_coverage — non-metric lookups
attacking/, shots/, defending/, goalkeeping/, discipline/, playing_time/, composite_scores/
# one metric-family tool each, built over METRIC_FIELDS
config.py # Pydantic Settings (env vars)
dependencies.py # FastAPI DI: get_repo(), get_mode_factory()
dependencies.py # FastAPI DI: get_repo(), get_mode_factory(), get_agent()
logging_config.py
main.py # App wiring: lifespan, CORS, router registration
```
Expand All @@ -97,12 +120,15 @@ app/

**Read path**: `GET /v1/players` → `MongoRepository.get_players()` → paginates + filters by position/team/nationality/sleeper_flag.

**Chat path**: `POST /v1/chat` → `ChatAgent.answer()` → LangGraph `create_agent` loop calls metric-family tools (each wraps `MongoRepository.get_players()`) → `answer_check.uncited_numbers()` flags any figure in the reply not present in a tool row (sets `degraded=True`, does not block the reply) → session state checkpointed to MongoDB by `session_id` (thread id). If the agent could not be built at startup (e.g. no `GEMINI_API_KEY`), `get_agent()` returns `None` and `/v1/chat` responds with a degraded generic answer instead of failing.

### MongoDB collections

- `player_bios` — one doc per player (identity/bio: `name`, `norm_name`, `sofascore_player_id`, `position`, `nationality`, `photo_url`). Indexed on `sofascore_player_id` (sparse unique) and `norm_name`.
- `player_stats` — one doc per `(player_bio_id, season)`. Contains `competitions[]`, `aggregated_stats`, `aggregated_scores`, `team`, `low_sample_size`. Each competition entry includes `stats`, `scores`, `raw_stats` (untyped ScraperFC columns), and `total_matches`.
- `fetch_log` — audit trail for each `fetch_data()` call.
- `league_meta` — one doc per `(competition, season)` tracking `total_matches` played so far. Populated at fetch time; stored per competition entry and exposed by the API. No longer feeds `s_final`.
- `chat_checkpoints` / `checkpoint_writes` — one document per chat session (thread), written by LangGraph's `MongoDBSaver`; TTL-expired `chat_session_ttl_seconds` (default 7 days) after the last write. `keep_latest_checkpoint()` also prunes all but the newest checkpoint after each turn.

### Key domain concepts

Expand All @@ -123,6 +149,8 @@ app/

Data loading is developer-driven via `tools/fetch_cli` (see its README) — the frontend has no fetch-triggering UI; when the DB is empty it just points to the CLI.

`ChatWidget` — floating chat bubble rendered on every tab (`src/components/ChatWidget.tsx`); opens `ChatPanel`. `?chat=1` renders `ChatFullScreen` instead of the tabbed `Dashboard` (`src/App.tsx`), reusing the same `ChatPanel`. Deliberately no navbar tab. Session id is generated client-side (`src/api/chat.ts`) and persisted for `GET/DELETE /v1/chat/sessions/{session_id}`.

### tools/fetch_cli

A standalone developer CLI, deliberately outside `backend/` — not part of the shipped
Expand Down Expand Up @@ -155,6 +183,7 @@ running backend's HTTP API (`http://localhost:8000` by default). Three commands:
- Always run python backend modules (like Pytest) via `.venv\Scripts\python`.
- Never push from `dev/*` to `master` without PR, ask user to approve merge only if CI passed.
- Delete the feature branch right after the changes were merge to the master branch.
- Advise only on free-tier technologies, this project should not cost any money.

## Never do these

Expand Down
44 changes: 42 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,8 +18,9 @@ This project is a **free, open-source, educational tool** built for football ent

| Layer | Technology |
|---|---|
| Frontend | React 18 · TypeScript · Vite · Tailwind CSS · Recharts |
| Frontend | React 18 · TypeScript · Vite · Tailwind CSS · Recharts · react-markdown |
| Backend | FastAPI · Python 3.12 · PyMongo · Pydantic Settings |
| Chatbot Agent | LangChain · LangGraph (`create_agent`), configurable LLM provider (Gemini free tier by default) |
| Database | MongoDB 7 |
| Data Fetching | ScraperFC · botasaurus · Chromium |
| Infrastructure | Docker Compose |
Expand All @@ -41,18 +42,23 @@ flowchart LR
PA["Player Assembler\nbuild · merge · aggregate"]
SC["Stats Client\nScraperFC + Chrome"]
MR["Mongo Repository"]
CA["Chat Agent\nLangGraph tool loop"]
end
DB[("MongoDB 7\n:27017")]
end

EXT["Live Football\nData Source"]
LLM["LLM Provider"]

User --> FE
FE -- REST --> API
API --> SE
API --> PA
API --> CA
PA --> SC
PA --> MR
CA --> MR
CA -- "chat tools" --> LLM
MR --> DB
SC -- "ScraperFC / botasaurus" --> EXT
```
Expand All @@ -61,6 +67,8 @@ flowchart LR

**Read path:** React SPA → API routers → `MongoRepository.get_players()` → paginated and filterable by position, team, nationality, or sleeper flag.

**Chat path:** floating chat widget (or `?chat=1` full-screen view) → `POST /v1/chat` → `ChatAgent` (LangGraph `create_agent` loop over per-metric-family DB query tools) → answer built from the rows those tools returned. Session history is a MongoDB checkpoint per thread, expiring 7 days after the last message.

---

## Core Features
Expand All @@ -72,6 +80,7 @@ flowchart LR
- **Player Detail** — per-competition stat breakdown and aggregated scores for any player, including those without a linked external ID
- **Head-to-Head Compare** — side-by-side comparison of exactly two players across all stat dimensions
- **Scatter Plot** — interactive xG+xA vs G+A chart (Recharts) across the full dataset
- **Chat Agent** — floating widget on every tab (also a full-screen view at `?chat=1`) answers natural-language questions about players and metrics from live DB tool calls; figures with no supporting row are flagged, and questions the database cannot answer are answered from the model's own knowledge and labelled as such; no navbar tab by design
- **Developer Data Loading** — `tools/fetch_cli`, a standalone CLI for browsing available competitions/seasons and loading data into MongoDB, with live per-task fetch progress
- **DB Snapshots** — JSON dump/restore scripts (`backend/scripts/DB/`) for safe local dev iteration

Expand Down Expand Up @@ -134,14 +143,24 @@ Player data is split across two MongoDB collections:
```env
MONGO_URI=mongodb://mongodb:27017/football_analytics
CORS_ORIGINS=["http://localhost:5173"]
GEMINI_API_KEY=your-key-here
```

`GEMINI_API_KEY` powers the chat agent (default provider, free tier). Without it the rest of the API still starts — the agent is just disabled and `/v1/chat` returns a degraded response. See [Chat Agent Configuration](#chat-agent-configuration) below for other providers.

### Full stack

```bash
docker compose up
```

After pulling changes to the backend or frontend dependencies, rebuild:

```bash
docker compose build backend
docker compose up -d --force-recreate --renew-anon-volumes frontend # node_modules lives in an anonymous volume
```

| Service | URL |
|---|---|
| Frontend | http://localhost:5173 |
Expand Down Expand Up @@ -173,7 +192,28 @@ pytest tests/domain/test_scoring_engine.py
pytest tests/domain/test_scoring_engine.py::test_name
```

Tests use `mongomock` — no running MongoDB required.
Tests use `mongomock` — no running MongoDB required. Agent tests live in `backend/tests/agent/`; the offline eval under `backend/tests/agent/eval/` is skipped by default and needs `AGENT_EVAL=1`, a live model, and a populated DB:

```bash
AGENT_EVAL=1 pytest tests/agent/eval -v -s
```

### Chat Agent Configuration

Set in `secrets.env` / `.env`:

| Variable | Default | Notes |
|---|---|---|
| `LLM_PROVIDER` | `gemini` | `gemini`, `openai`, or `anthropic` |
| `LLM_MODEL` | provider default | `openai`/`anthropic` have no default — must be set explicitly |
| `LLM_API_KEY` | unset | falls back to the provider SDK's own env var (e.g. `GEMINI_API_KEY`) |
| `LLM_FALLBACK_MODEL` | unset | model to retry with on failure; empty means no fallback |
| `AGENT_MAX_TOOL_ITERATIONS` | `8` | tool-call loop limit per turn |
| `AGENT_MAX_ROWS` | `25` | max rows a DB query tool can return |
| `CHECKPOINT_COLLECTION` | `chat_checkpoints` | MongoDB collection for session state |
| `CHAT_SESSION_TTL_SECONDS` | `604800` (7 days) | session expiry after the last message |

`openai` and `anthropic` require installing their LangChain package (`langchain-openai` / `langchain-anthropic`) and setting `LLM_MODEL` explicitly.

### Loading Data

Expand Down
Empty file added backend/app/agent/__init__.py
Empty file.
161 changes: 161 additions & 0 deletions backend/app/agent/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,161 @@
"""The chatbot agent: a LangGraph tool-calling loop over the database tools."""

import asyncio
import logging
from collections.abc import Callable
from dataclasses import dataclass
from functools import partial

from langchain.agents import create_agent
from langchain.agents.middleware import ModelFallbackMiddleware
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import AIMessage, HumanMessage, ToolMessage
from pydantic import TypeAdapter, ValidationError

from app.agent.answer_check import uncited_numbers
from app.agent.checkpoints import build_checkpointer, keep_latest_checkpoint
from app.agent.constants import GENERIC_ERROR, MAX_TOOL_ITERATIONS
from app.agent.llm import build_chat_model, build_fallback_model
from app.agent.system_prompt import SYSTEM_PROMPT
from app.agent.tools import build_tools

logger = logging.getLogger(__name__)


@dataclass
class ChatResult:
answer: str
used_tools: bool
degraded: bool = False


class ChatAgent:
def __init__(
self,
model: BaseChatModel,
repo,
checkpointer,
fallback_models: list[BaseChatModel] | None = None,
prune: Callable[[str], None] | None = None,
) -> None:
self._prune = prune
self._graph = create_agent(
model=model,
tools=build_tools(repo),
system_prompt=SYSTEM_PROMPT,
checkpointer=checkpointer,
middleware=[ModelFallbackMiddleware(*fallback_models)] if fallback_models else [],
)

def _config(self, session_id: str) -> dict:
return {
"configurable": {"thread_id": session_id},
"recursion_limit": MAX_TOOL_ITERATIONS * 2,
}

async def answer(self, message: str, session_id: str) -> ChatResult:
try:
# "exit" saves one checkpoint when the turn ends, not one per graph step.
state = await self._graph.ainvoke(
{"messages": [HumanMessage(content=message)]},
config=self._config(session_id),
durability="exit",
)
except Exception:
logger.exception("Agent failed for session %s", session_id)
return ChatResult(answer=GENERIC_ERROR, used_tools=False, degraded=True)
await self._prune_session(session_id)

messages = _current_turn(state["messages"])
used_tools = any(isinstance(m, ToolMessage) for m in messages)
answer = messages[-1].text if messages else ""

if used_tools and answer:
rows = _tool_rows(messages)
uncited = uncited_numbers(answer, rows) if rows is not None else []
if uncited:
# A signal, not a gate: withholding the answer would be worse than flagging it.
logger.warning(
"Ungrounded figures %s in answer for session %s; rows=%s",
uncited,
session_id,
rows,
)
return ChatResult(answer=answer, used_tools=True, degraded=True)

# No tool behind the answer means it is the model's own knowledge, not this app's data.
return ChatResult(
answer=answer or GENERIC_ERROR,
used_tools=used_tools,
degraded=not (answer and used_tools),
)

async def _prune_session(self, session_id: str) -> None:
if self._prune is None:
return
try:
await asyncio.to_thread(self._prune, session_id)
except Exception:
# Old checkpoints only cost storage, and the TTL removes them anyway.
logger.exception("Could not prune checkpoints for session %s", session_id)

def history(self, session_id: str) -> list[dict]:
"""Replay the thread as {role, content} turns. Tool messages are never exposed."""
try:
state = self._graph.get_state(self._config(session_id))
except Exception:
logger.exception("Could not read history for session %s", session_id)
return []
turns = []
for m in (state.values or {}).get("messages", []):
if isinstance(m, HumanMessage):
turns.append({"role": "user", "content": m.text})
# Text sent alongside a tool call is the model's preamble; it was never shown.
elif isinstance(m, AIMessage) and m.text and not m.tool_calls:
turns.append({"role": "assistant", "content": m.text})
return turns

def clear(self, session_id: str) -> None:
try:
self._graph.checkpointer.delete_thread(session_id)
except Exception:
# The caller gets 204 either way; the TTL removes the thread later.
logger.exception("Could not clear session %s", session_id)


def _current_turn(messages: list) -> list:
"""Messages after the latest user message; the state holds the whole thread."""
for i in range(len(messages) - 1, -1, -1):
if isinstance(messages[i], HumanMessage):
return messages[i + 1 :]
return messages


# Rows are heterogeneous by design: a metric row, an identity profile, a coverage object or
# an error row. Validate the shape, not a schema.
_ToolPayload = TypeAdapter(list[dict] | dict)


def _tool_rows(messages) -> list[dict] | None:
"""Rows every tool returned this turn, or None if any output could not be read."""
rows: list[dict] = []
for m in messages:
if not isinstance(m, ToolMessage):
continue
try:
payload = _ToolPayload.validate_json(m.content)
except ValidationError:
return None # unreadable rows would look like missing citations
rows.extend(payload if isinstance(payload, list) else [payload])
return rows


def build_agent(repo, mongo_client) -> ChatAgent:
checkpointer = build_checkpointer(mongo_client)
return ChatAgent(
model=build_chat_model(),
repo=repo,
checkpointer=checkpointer,
fallback_models=[m for m in [build_fallback_model()] if m],
prune=partial(keep_latest_checkpoint, checkpointer),
)
Loading
Loading