Skip to content

Repository files navigation

Football Analytics

A football analytics platform for observing and analyzing player statistics across leagues worldwide, with dedicated support for fantasy league decision-making.

Screenshots

The GIFs below show the dark-theme UI running against real 2025-26 season data.

Player Details — Ranked by the 0-10 Fantasy Score, with live filters and sorting.

Player details — ranked by the 0-10 Fantasy Score, with live filters and sorting

Player Modal — Fantasy Score and 0-10 pillars (offensive, defensive, tactical), per-competition breakdown, goal and shot bars.

Player modal — Fantasy Score and 0-10 pillars, per-competition breakdown, goal and shot bars

Compare — Head-to-head with mirrored bars and a Δ column.

Compare — head-to-head with mirrored bars and a Δ column

xGI Outliers — "Due to score" and "Overperforming" players.

xGI Outliers — Due to score and Overperforming players

Scatter Plot — xG+xA vs G+A with outlier highlighting.

Scatter plot — xG+xA vs G+A with outlier highlighting

Ask AI — Answers built from database tool calls, with a tool trace.

Ask AI — answers built from database tool calls, with a tool trace

Disclaimer & Intended Use

This project is a free, open-source, educational tool built for football enthusiasts exploring soccer data and fantasy players seeking personal analytical insights.

  • No Affiliation: This project is not affiliated with, endorsed by, or sponsored by any football league, club, fantasy sports platform, or data provider. All trademarks, data, and intellectual property belong to their respective owners.
  • Third-Party Data: The software may retrieve data from third-party sources. The authors do not own, control, or warrant this data.
  • User Responsibility: You are solely responsible for ensuring that your use of this software—including accessing or fetching third-party platforms—complies with applicable laws, regulations, and the respective third-party Terms of Service.

The software is provided "as is", without warranty of any kind. The authors accept no liability for its use or for the accuracy of any generated insights. See the LICENSE file for full terms.


Tech Stack

Layer Technology
Frontend React 18 · TypeScript · Vite · Tailwind CSS · Recharts · react-markdown · lucide-react icons · Inter/JetBrains Mono fonts
Backend FastAPI · Python 3.12 · PyMongo · Pydantic Settings
Chatbot Agent LangChain · LangGraph (create_agent), configurable LLM provider (Gemini free tier by default)
Database MongoDB 7
Data Fetching ScraperFC
Infrastructure Docker Compose
Testing pytest · mongomock

Architecture

flowchart LR
    User["Browser"]

    subgraph Docker["Docker Compose"]
        FE["React SPA\nVite :5173"]
        subgraph Backend["FastAPI :8000"]
            API["API Routers"]
            SE["Scoring Engine\nFantasy Score = clamp(10 x raw / 8, 0, 10)"]
            PA["Player Assembler\nbuild · merge · aggregate"]
            SC["Stats Client\nScraperFC"]
            MR["Mongo Repository"]
            CA["Chat Agent\nLangGraph tool loop"]
        end
        DB[("MongoDB 7\n:27017")]
    end

    EXT["Live Football\nData Source"]
    LLM["LLM Provider"]

    User --> FE
    FE -- REST --> API
    API --> SE
    API --> PA
    API --> CA
    PA --> SC
    PA --> MR
    CA --> MR
    CA -- "chat tools" --> LLM
    MR --> DB
    SC -- "ScraperFC" --> EXT
Loading

Fetch path: developer runs tools/fetch_cli → POST /v1/fetch/ → FantasyMode → FetchRunner pulls stats per competition (concurrent, no fetch rate limit) → PlayerAssembler scores via ScoringEngine and classifies sleepers → MongoRepository upserts to player_bios / player_stats.

Read path: React SPA → API routers → MongoRepository.get_players() → paginated and filterable by name, team, position, nationality, or sleeper flag; name and team are case- and accent-insensitive substring matches, combined with AND when both are set. Sorting also accepts sort_by=xratio (the sleeper ratio), which is sort-only — it is not a filter metric and not an agent tool metric.

Chat path: floating chat widget (or ?chat=1 full-screen view) → POST /v1/chat → ChatAgent (LangGraph create_agent loop over per-metric-family DB query tools) → answer built from the rows those tools returned. The response also carries tool_calls (tool name + row count only, never arguments or row contents) and uncited (figures in the answer with no backing row), which the UI renders as a trace and a warning. Session history is a MongoDB checkpoint per thread, expiring 7 days after the last message.

Meta path: GET /v1/meta?season= → player count, last fetch time, stored seasons, and the active chat model/tool-call cap. Powers the app header's dataset status, season selector, and the Ask AI tab's model line.


Core Features

  • Fantasy Scoring — composite Fantasy Score on a strict 0-10 scale, computed only from the player's own stats: raw = (raw_per90 x starter_bonus + playing_time_bonus) x confidence, then 10 x raw / 8, clamped to [0, 10]. raw_per90 is (Offensive + Defensive + Tactical) / max(minutes / 90, 1) with position-specific goal/assist weights (GK goals worth 10 pts, FW goals worth 4 pts). starter_bonus rewards regular starters, confidence discounts small samples (the lower of the appearance tier and the tier of 60-minute games played), and playing_time_bonus (max 1) rewards average minutes per game. The offensive, defensive and tactical pillar scores are also 0-10 (tactical 5 = neutral discipline) — see Mathematical_Specification.md
  • Player Details — paginated player table sorted by Fantasy Score by default; live-filterable (250ms debounce) by name, team, position, nationality, and sleeper flag, plus a removable-chip "Add metric filter" popover for any allowlisted metric; click a row for the per-competition stat breakdown and aggregated scores, including players without a linked external ID
  • xGI Outliers — flags a player "Due to score" (HIGH_VALUE) when xG+xA is more than 1.20x G+A (also when G+A is 0 and xG+xA is above 0), and "Overperforming" (OVERPERFORMING) when G+A is more than 1.25x xG+xA; only applies to players with more than 450 minutes. The xGI Outliers page has one tab per flag, sorted by the xG+xA/G+A ratio (highest first for "Due to score", lowest first for "Overperforming"), and shows the xGI gap for each player
  • Head-to-Head Compare — side-by-side comparison of exactly two players across all stat dimensions, with mirrored bars and a per-row delta
  • Scatter Plot — interactive xG+xA vs G+A chart (Recharts) with selectable X/Y metrics, a position filter, a minimum-minutes filter, an outlier-highlight toggle, and a detail panel for the selected point
  • Dataset Status — the app header shows player count, last fetch time, and a season selector, backed by GET /v1/meta; the Ask AI tab shows the active chat model and its tool-call cap instead
  • Chat Agent — "Ask AI" navbar tab plus a floating widget on the other tabs (also a full-screen view at ?chat=1) answers natural-language questions about players and metrics from live DB tool calls; each answer shows which tools ran and how many rows they returned, figures with no supporting row are flagged with a warning, and questions the database cannot answer are answered from the model's own knowledge and labelled as such
  • Developer Data Loading — tools/fetch_cli, a standalone CLI for browsing available competitions/seasons and loading data into MongoDB, with live per-task fetch progress
  • DB Snapshots — JSON dump/restore scripts (backend/scripts/DB/) for safe local dev iteration

Data Schema

Player data is split across two MongoDB collections:

player_bios — one doc per player (identity, stable across seasons):

{
  "_id": "ObjectId(...)",
  "sofascore_player_id": "277174",
  "name": "Harry Kane",
  "norm_name": "harry kane",
  "position": "FW",
  "position_exact": "ST",
  "nationality": "England",
  "photo_url": "https://photo-url"
}

player_stats — one doc per (player, season):

{
  "_id": "ObjectId(...)",
  "player_bio_id": "ObjectId(...)",
  "season": "2025-2026",
  "team": "Bayern Munich",
  "competitions": [
    {
      "competition": "Germany Bundesliga",
      "competition_type": "club",
      "total_matches": 34,
      "stats": { "goals": 28, "assists": 8, "minutes": 2880, "xg": 24.1, ... },
      "scores": { "offensive": 88.2, "defensive": 2.5, "tactical": 4.1, "s_final": 4.63 },
      "raw_stats": { "totwAppearances": 9 }
    }
  ],
  "aggregated_stats": { "goals": 28, "assists": 8, "minutes": 2880, "xg": 24.1 },
  "aggregated_scores": {
    "offensive": 88.2, "defensive": 2.5, "tactical": 4.1, "s_final": 4.63,
    "underpredicted_flag": null, "underpredicted_ratio": 1.16
  },
  "low_sample_size": false,
  "last_updated": "2026-06-22T10:00:00+00:00"
}

competition_type is "club" for domestic leagues and cups, "national" for international tournaments (World Cup, European Championship, etc.). The stats-view filter on the Player Details page uses this field to let you see club-only or national-team-only aggregated scores.

s_final is the stored field name for what the UI calls the Fantasy Score.

Running the Project

Prerequisites

  • Docker + Docker Compose
  • secrets.env in the project root:
MONGO_URI=mongodb://mongodb:27017/football_analytics
CORS_ORIGINS=["http://localhost:5173"]
GEMINI_API_KEY=your-key-here

GEMINI_API_KEY powers the chat agent (default provider, free tier). Without it the rest of the API still starts — the agent is just disabled and /v1/chat returns a generic error message. To use OpenAI or Anthropic instead, see Chat Agent LLM Providers.

Full stack

docker compose up

After pulling changes to the backend or frontend dependencies, rebuild:

docker compose build backend
docker compose up -d --force-recreate --renew-anon-volumes frontend  # node_modules lives in an anonymous volume
Service URL
Frontend http://localhost:5173
Backend http://localhost:8000
API docs http://localhost:8000/docs

On first run the database is empty — use tools/fetch_cli (see its README) to load data, developer-driven.

Local development (without Docker)

# Backend — from backend/
pip install -r requirements.txt
uvicorn app.main:app --reload

# Frontend — from frontend/
npm install
npm run dev

Tests

# All tests — from backend/
pytest

# Single file or test
pytest tests/domain/test_scoring_engine.py
pytest tests/domain/test_scoring_engine.py::test_name

Tests use mongomock — no running MongoDB required. Agent tests live in backend/tests/agent/; the offline eval under backend/tests/agent/eval/ is skipped by default and needs AGENT_EVAL=1, a live model, and a populated DB:

AGENT_EVAL=1 pytest tests/agent/eval -v -s

Chat agent settings (provider, model, limits) are documented in backend/app/agent/README.md.

Chat Agent LLM Providers

The chat agent supports three LLM providers: Gemini (default), OpenAI and Anthropic. You choose one with environment variables. No code change is needed.

Where to put the variables

How you run the backend File
Docker (docker compose up) secrets.env in the project root
Local (uvicorn from backend/) backend/.env

Both files are gitignored. Never commit an API key.

Variables

Variable Purpose
LLM_PROVIDER gemini (default), openai or anthropic
LLM_MODEL Model id. Optional for Gemini (default gemini-3.5-flash). Required for OpenAI and Anthropic.
LLM_API_KEY API key for the selected provider. Takes priority over the provider-specific variables below.
GEMINI_API_KEY / OPENAI_API_KEY / ANTHROPIC_API_KEY Provider-specific key. Used when LLM_API_KEY is not set.
LLM_FALLBACK_MODEL Optional. Model to retry with when the main model fails (for example, an HTTP 503). It must be different from LLM_MODEL. Gemini defaults to gemini-3.5-flash-lite. OpenAI and Anthropic have no fallback unless you set one.

The selected model must support tool calling, because the agent answers by calling database tools.

Gemini (default, free tier)

  1. Create a free API key in Google AI Studio.

  2. Add it to your env file:

    GEMINI_API_KEY=your-gemini-key
    # Optional overrides:
    # LLM_MODEL=gemini-3.5-flash
    # LLM_FALLBACK_MODEL=gemini-3.5-flash-lite
  3. Restart the backend (see Apply the changes).

No extra package is needed. langchain-google-genai is already in backend/requirements.txt.

OpenAI (paid)

  1. Create an API key in the OpenAI dashboard. The OpenAI API is billed per token, so the account needs credit.

  2. Install the LangChain OpenAI package. It is not installed by default:

    • Docker: add the line langchain-openai to backend/requirements.txt, then rebuild the image:

      docker compose build backend
    • Local: from backend/, run pip install langchain-openai.

  3. Add the provider settings to your env file:

    LLM_PROVIDER=openai
    LLM_MODEL=gpt-5-mini              # any OpenAI chat model with tool calling
    LLM_API_KEY=sk-...                # or OPENAI_API_KEY=sk-...
    # Optional:
    # LLM_FALLBACK_MODEL=gpt-5-nano
  4. Restart the backend (see Apply the changes).

Check the OpenAI models page for current model ids.

Anthropic (paid)

  1. Create an API key in the Anthropic Console. The Anthropic API is billed per token, so the account needs credit.

  2. Install the LangChain Anthropic package. It is not installed by default:

    • Docker: add the line langchain-anthropic to backend/requirements.txt, then rebuild the image:

      docker compose build backend
    • Local: from backend/, run pip install langchain-anthropic.

  3. Add the provider settings to your env file:

    LLM_PROVIDER=anthropic
    LLM_MODEL=claude-sonnet-5         # or claude-haiku-4-5-20251001 (cheaper)
    LLM_API_KEY=sk-ant-...            # or ANTHROPIC_API_KEY=sk-ant-...
    # Optional:
    # LLM_FALLBACK_MODEL=claude-haiku-4-5-20251001
  4. Restart the backend (see Apply the changes).

Check the Anthropic models page for current model ids.

Apply the changes

The backend reads these settings once, at startup. After you edit the env file, restart it:

# Docker: recreate the container so it reads the new secrets.env
docker compose up -d --force-recreate backend

# Local: stop uvicorn (Ctrl+C) and start it again
uvicorn app.main:app --reload

Verify

  • GET http://localhost:8000/v1/meta returns the active model in agent.model. The UI shows the same value.
  • Ask a question in the Ask AI tab.

If the agent does not answer, check the backend logs (docker compose logs backend). Common causes:

Log message Fix
Provider 'anthropic' needs a package: pip install langchain-anthropic (or openai) Install the package (step 2) and rebuild the image.
Provider 'openai' has no default model: set LLM_MODEL. Set LLM_MODEL.
Unknown LLM_PROVIDER ... Use gemini, openai or anthropic (lowercase).
Authentication error (401) The API key is wrong, or it belongs to a different provider than LLM_PROVIDER.

When the agent cannot start, the rest of the app keeps working. Only /v1/chat returns a generic error message.

Loading Data

The app has no fetch-triggering UI — data loading is developer-driven via tools/fetch_cli, a standalone CLI (its own lightweight venv, separate from backend/) that talks to the running backend over HTTP:

cd tools
python -m venv .venv
.venv/Scripts/pip install -r requirements.txt

.venv/Scripts/python -m fetch_cli.cli refresh   # pull the competition/season catalog
.venv/Scripts/python -m fetch_cli.cli browse    # see what's available
.venv/Scripts/python -m fetch_cli.cli fetch     # pick a league + season, load it into MongoDB

See tools/fetch_cli/README.md for full details.

DB Snapshots

# From backend/ with the stack up
python scripts/DB/snapshot_dump.py   # → backend/scripts/snapshots/cl-2025-2026.json
python scripts/DB/snapshot_load.py   # restore

Claude Code users: the db-snapshot skill (.claude/skills/db-snapshot/SKILL.md) wraps these scripts for on-demand snapshots.

About

Football analytics platform: load player data, explore stats, compare players, and check their Fantasy Score. Includes an AI chat that answers questions from the data.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages