diff --git a/.env.example b/.env.example new file mode 100644 index 0000000..7153669 --- /dev/null +++ b/.env.example @@ -0,0 +1,24 @@ +# Copy this to .env.prod and fill in real values — never reuse dev placeholder values here. +# Used by: +# docker compose --env-file .env.prod up -d --build +# +# This stack deploys against Postgres/Redis that already exist elsewhere (no postgres/redis +# service blocks in docker-compose.yml) — point these at your real instances. + +# Full connection strings for your existing Postgres/Redis — not split into user/password/host +# parts, since docker-compose.yml passes these straight through as-is to crawler/api/ws. +DATABASE_URL=postgres://USER:PASSWORD@HOST:5432/DBNAME?sslmode=disable +REDIS_ADDR=redis://default:PASSWORD@HOST:6379 + +# The real, deployed frontend URL — not localhost. This is the one origin the api and ws +# services will both accept cross-origin / WebSocket-upgrade requests from. (api's own copy of +# this is effectively vestigial now that it has no public domain — see PRODUCTION.md's +# "Isolating api behind the frontend" — kept only as harmless defense-in-depth.) +CORS_ALLOWED_ORIGIN=https://your-frontend-domain.example.com + +# ws's real public wss:// domain — a full URL, not a path under the frontend's or api's own +# domain. Baked into the frontend's client bundle at container startup (see +# frontend/docker-entrypoint.sh); the browser connects to this directly. There's no equivalent +# var for api anymore — it's reached only through the frontend's own internal proxy, configured +# via API_INTERNAL_URL inside docker-compose.yml itself, not a secret worth putting here. +VITE_WS_URL=wss://your-ws-domain.example.com diff --git a/.gitignore b/.gitignore index bc5973a..ad3f3fc 100644 --- a/.gitignore +++ b/.gitignore @@ -26,7 +26,12 @@ frontend/dist/ logs/ *.log -# Go (scraper) +# Go (scraper + api + ws) +# NOTE: go.sum must NEVER be ignored — it locks dependency checksums and is required for +# reproducible builds. A prior "*.sum" rule here silently excluded both backend/scraper/go.sum +# and backend/api/go.sum from every commit ever made in this repo; removed, and both files are +# now tracked for real. backend/scraper/vendor/ -*.sum *.exe +backend/api/quotes-api +backend/ws/quotes-ws diff --git a/PRODUCTION.md b/PRODUCTION.md new file mode 100644 index 0000000..edb6d19 --- /dev/null +++ b/PRODUCTION.md @@ -0,0 +1,372 @@ +# Production deployment + +One file, `docker-compose.yml` at the repo root, deploys `crawler` (`backend/scraper`), `api` +(`backend/api`), `ws` (`backend/ws`), and `frontend` (TanStack Start, via `frontend/Dockerfile`) +together. It deliberately has **no** `postgres`/`redis` service blocks — this deploys against +Postgres and Redis that already exist elsewhere (e.g. Dokploy Database resources you manage and +back up separately), via `DATABASE_URL`/`REDIS_ADDR`. The frontend is bundled in with the +backend, not deployed to its own separate host, for a specific reason: see **"Isolating api +behind the frontend"** below. + +**`api` has no public domain at all.** The browser never talks to it directly — only the +frontend's own server does, over this compose network, via a `createServerFn`-based proxy (see +`frontend/src/lib/typing/sentences.ts`) configured with the plain runtime env var +`API_INTERNAL_URL=http://api:8080`. `ws` is still reachable at its own public `wss://` domain for +now (the browser connects to it directly) — proxying real-time WebSocket traffic through the +frontend's server is a separate, more involved piece of work than api's plain HTTP proxy, and +hasn't shipped yet. + +`frontend` itself takes one build-time-looking env var the same way `ws`'s own domain does: +`VITE_WS_URL`, a full `wss://` URL. Vite only ever inlines `import.meta.env.VITE_*` at build +time, not runtime, but `VITE_WS_URL` is set here as an **ordinary runtime environment +variable** — `frontend/Dockerfile`'s build stage bakes in a fixed placeholder token instead of a +real value, and `frontend/docker-entrypoint.sh` substitutes the real runtime env var for that +placeholder across every built file the instant the container starts, before the server boots. +This exists specifically because plenty of PaaS UIs (Dokploy included) only ever expose +_runtime_ environment variables with no way to pass a real Docker build-arg through to +`docker build`, which would otherwise leave the dev fallback (`localhost`) silently baked in with +no error anywhere. `API_INTERNAL_URL` needs none of this — it's read fresh from `process.env` +inside a server-only `createServerFn` handler, which never runs in the browser, so there's +nothing to bake at build time in the first place. + +Everything here was built and verified against a real, isolated instance of this exact stack — +not assumed to work from the compose file alone. Several real bugs were caught this way (see +"Findings from live testing" below), and the api proxy above was verified by running a real +mock `api` service with zero host-published port, reachable only by its container name on an +isolated Docker network, and confirming a real race actually fetched a quote through the +frontend's proxy with no direct network path from the browser to that container at all. + +## Isolating api behind the frontend + +Why this needs the frontend _in_ this compose file, not deployed separately: there is no way to +make a public backend service reachable "only by the frontend" while the **browser** still talks +to it directly — whatever URL a browser can reach, anyone can reach (copy it straight out of +devtools; no secret is involved). The only real fix is having the frontend's own server be the +one thing that ever talks to `api`, with `api` itself given no public domain at all. + +That in turn requires `api` and `frontend` to be able to reach each other by plain internal +DNS (`http://api:8080`), which **only works when they're deployed together as one compose +project** — see "Deploying via Dokploy" below for why that's a hard platform requirement, not +just a preference. + +## 1. Required environment variables + +``` +DATABASE_URL — your existing Postgres connection string +REDIS_ADDR — your existing Redis connection string (redis://... form is fine + directly — see "Deploying via Dokploy" below) +CORS_ALLOWED_ORIGIN — the frontend's real public domain (e.g. https://vroomy.example.com) +VITE_WS_URL — ws's real public wss:// domain (e.g. wss://ws.vroomy.example.com) +PUBLIC_URL — same as CORS_ALLOWED_ORIGIN, no trailing slash — builds the OIDC + redirect_uri (.../auth/callback); must exactly match a Redirect URI + registered on the identity provider's client — see "Authentication" below +OIDC_ISSUER — the identity provider's own URL, including any realm/org path it needs + (e.g. https://auth.example.com/realms/apps for Keycloak) — see below +OIDC_CLIENT_ID — Vroomy's client/application ID in the identity provider +OIDC_CLIENT_SECRET — Vroomy's client secret in the identity provider +SESSION_SECRET — random, >=32 chars (e.g. `openssl rand -base64 48`) — seals Vroomy's own + login-session cookie; never reuse across environments +``` + +`.env.example` documents the same four vars — copy it to `.env.prod` and fill in real values if +deploying with plain `docker compose` rather than Dokploy (`.env.prod` is gitignored, never +committed). **Never reuse dev placeholder values** here; `backend/scraper`'s own local-dev +`docker-compose.yml` is a completely separate, unrelated Postgres/Redis instance from whatever +you point `DATABASE_URL`/`REDIS_ADDR` at here. + +## 2. Bring up the stack (plain `docker compose`, no Dokploy) + +```bash +docker compose --env-file .env.prod up -d --build +docker compose --env-file .env.prod logs -f +``` + +All four containers run with `restart: unless-stopped`, so a crash or host reboot brings them +back automatically. `api` has no host-published port at all — only reachable from other +containers on this compose network, by service name. `frontend`/`ws` also publish nothing to the +host by default (see "Deploying via Dokploy" below for why) — if running this outside Dokploy, +front them yourself with a reverse proxy (see `backend/ws/Caddyfile.example` and +`frontend/Caddyfile.example` for two-line Caddy configs that get real, auto-renewing Let's +Encrypt certs) proxying to each container's own address on this compose network. + +The crawler self-migrates the database schema on startup (same as it does in dev); neither the +API nor `ws` ever migrates or writes, so make sure the crawler has started at least once against +a fresh database before relying on either of them. + +## 3. Set up backups + +**If Postgres/Redis are managed by an orchestrator with its own backup scheduler (e.g. Dokploy's +built-in per-database scheduled backups to S3-compatible storage), use that instead of the +script below** — it already handles off-host storage and scheduling, which the script does not. +The script here is for a Postgres you're managing yourself with no orchestrator backing it: + +```bash +export POSTGRES_USER=quotes POSTGRES_DB=quotes POSTGRES_CONTAINER=quotes-postgres +./backend/scraper/scripts/backup.sh +``` + +Dumps to `./backups/quotes_.sql.gz` and prunes anything older than 14 days +(`RETENTION_DAYS`, `BACKUP_DIR` are both overridable). Schedule it — a cron entry works fine: + +```cron +0 3 * * * cd /path/to/vroomy && POSTGRES_USER=quotes POSTGRES_DB=quotes POSTGRES_CONTAINER=quotes-postgres ./backend/scraper/scripts/backup.sh >> /var/log/quotes-backup.log 2>&1 +``` + +This writes to local disk only — for real disaster recovery (surviving the host itself being +lost), `BACKUP_DIR` should point somewhere that's itself synced offsite, or the script extended +to push each dump to object storage. + +## 4. Deploying via Dokploy + +**Deploy `docker-compose.yml` as a single Dokploy "Compose" project — not as separate +Dockerfile-type Applications, one per service.** This is a hard requirement, not a style +preference: Dokploy only gives automatic internal service-name DNS resolution (what lets +`frontend` reach `http://api:8080`) to services deployed together as one Compose project. +Separately-deployed standalone Applications each get their own isolated Docker network and +**cannot reach each other by name at all** — a currently-open Dokploy limitation +([Dokploy#3670](https://github.com/Dokploy/dokploy/issues/3670)). Deployed any other way, `api`'s +isolation from "Isolating api behind the frontend" above simply doesn't work: `frontend` would +fail to resolve `api` and every quote fetch would fall back to the local sentence pool. + +**Postgres/Redis Database resources run as Docker Swarm services on Dokploy's shared +`dokploy-network` overlay network — not whatever network a plain Compose deploy creates by +default.** `docker-compose.yml`'s `networks.default` is pinned to `dokploy-network` (`external: +true`) specifically so this project lands on the same network your existing Database resources +are already on; without this, `DATABASE_URL`/`REDIS_ADDR` hostnames fail to resolve at all (a +DNS lookup failure, not an auth/connection-refused error — easy to misdiagnose as a credentials +problem when it's actually a networking one). Find your own Database resource's network with +`docker service ls` and `docker network ls` on the host if you're ever unsure it's really +`dokploy-network`. + +Within the Compose project, Dokploy lets you configure a public **Domain** per service, +independently of each other: + +- **`frontend` needs a Domain**, container **Port 3000**, health-check path `/` (it has no + dedicated `/health` route — the root route responding is the signal). +- **`ws` needs a Domain** too, for now — container **Port 8081**, health-check path `/health`. + This goes away once ws is proxied through the frontend the same way `api` already is (not yet + built — see this file's top section). +- **`api` and `crawler` get no Domain at all.** `crawler` never did (it's a background worker + with no HTTP server — nothing for a reverse proxy to route to or poll); `api` is deliberately + unexposed, the whole point of this setup. +- `docker-compose.yml` deliberately has no `ports:` host-publish on `frontend`/`ws` — Dokploy's + Traefik routes to them directly over `dokploy-network` based on the Domain config above, so + publishing a host port isn't needed the way it is for the plain-`docker compose`+Caddy path in + step 2. **Port 3000 specifically must stay unpublished** — it collides with Dokploy's own + dashboard, which already owns that host port. +- Set the four env vars from "Required environment variables" above as this Compose project's + environment variables, under Dokploy's **Environment Settings** tab. Make sure **"Create + Environment File"** is enabled there — if it's off, the values you type are saved in Dokploy's + own database but never written to an actual `.env` file next to the compose file, so every + `${VAR}` in it silently resolves to an empty string with no error anywhere (confirmed live: + `api` crash-looped with `user=root database=` — pgx's fallback for a genuinely empty connection + string — until this was switched on and redeployed). Editing/saving these values alone doesn't + retroactively change already-running containers either — trigger an actual redeploy after + saving, since `restart: unless-stopped` just keeps recreating the _same_ crashed container with + its old baked-in env otherwise. +- **`REDIS_ADDR` accepts Dokploy's internal Redis connection URL directly** (`redis://...`) — + `ConnectRedis` (see finding #1 below) switches to URL parsing for anything containing `://`, + so `REDIS_PASSWORD`/`REDIS_DB` env vars are unused in that case; only set `REDIS_ADDR`. Only + the crawler uses Redis at all — api and ws only ever need `DATABASE_URL`. +- **Memory limits are already encoded in `docker-compose.yml` itself** (each service's own + `deploy.resources.limits.memory`), and Dokploy's Compose deployment reads that file directly — + no need to re-enter these per service in Dokploy's UI. The table below is just what's already + in the file, for reference (crawler sized for its real measured ~1.25GB working set — see + finding #2 below; api/ws/frontend are far lighter, no language models loaded — **frontend's + figure specifically is an unmeasured starting point**, worth checking with `docker stats` + under real traffic): + + | Service | Limit | Also set | + | -------- | ------ | -------------------- | + | crawler | 2 GB | `GOMEMLIMIT=1536MiB` | + | api | 128 MB | `GOMEMLIMIT=100MiB` | + | ws | 128 MB | `GOMEMLIMIT=100MiB` | + | frontend | 256 MB | — | + + If Dokploy's UI also exposes a separate Reservation (soft) field for Compose-deployed services, + a reasonable soft target is roughly half of each hard Limit above — a bare single hard limit + with no slack turns an ordinary transient spike into an unnecessary SIGKILL. + +- **A memory-limit kill is a blip here, not a real incident, by design**: `WarmSimhashCache` and + `WarmFrontierCache` fully rehydrate Redis from Postgres on every restart, and + `db.RequeueStuckInProgress` recovers any URL a killed process left stuck mid-fetch — both + already exist for the ordinary case of the crawler process itself dying, and apply equally to + an OOM-triggered container restart. Combined with `restart: unless-stopped`, the practical + effect of hitting a limit is a few seconds of downtime, not lost work. +- **On a small host (e.g. 4-5GB total), add a swapfile** as a cheap way to reduce how often a + transient spike escalates to an OOM-kill at all — the kernel pages out cold memory first + instead of immediately SIGKILLing on hitting a hard limit: + ```bash + sudo fallocate -l 2G /swapfile + sudo chmod 600 /swapfile + sudo mkswap /swapfile + sudo swapon /swapfile + echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab + ``` +- **Day/night scheduling is handled by stopping and starting the crawler service itself**, not + by anything inside the process. As part of one Compose project, check whether Dokploy's + Scheduled Tasks can target a single service within the project — if it can only toggle the + whole project, that's coarser (it would also stop api/ws/frontend), and a plain scheduled + `docker compose stop crawler` / `start crawler` on the host is the fallback. An earlier + in-app pause-only approach (an active-hours env var) was tried and reverted: pausing the + worker loops cut CPU/network but never freed the ~1.25GB RAM the crawler holds once its + language-detection models are loaded, since that memory stays resident for the life of the + process whether it's fetching or idle — only actually stopping the container frees it. +- **`LOG_LEVEL` cuts log volume at the source, on top of the disk-side cap below.** Three levels: + `info` (unset, the default) is today's full output; `warn` drops the highest-volume routine + narration (one line per page parsed, per quote saved, per URL discovered) but keeps every + skip/retry/fallback line — a rejected quote, a low-yield category abandoned, a rate-limit + backoff, a retry after a failed fetch — alongside real failures; `error` drops warnings too, + down to genuine failures only (`Could not ...`, a fetch/DB error, giving up on a URL) — never + suppressed at any level. **`error` is the recommended production setting** for a low-noise log + — `warn` is there for when you actually want to see _why_ the crawler's doing what it's doing + (rate-limited? skipping bad content? just quiet right now?) without the full per-item flood. + `docker-compose.yml` already sets `LOG_LEVEL: error` for the `crawler` service. Leave it unset + in dev if you want the full picture of what the crawler's doing. +- **Cap container log size, or the crawler's own logs can fill the disk.** Docker's default + `json-file` log driver has no size limit — the crawler logs roughly one line per URL + discovered/fetched (`internal/crawler/crawler.go`), so left running for weeks that adds up. + `docker-compose.yml` already caps every service at 10MB × 3 files via its `x-logging` anchor, + and this takes effect automatically since Dokploy's Compose deployment runs the file directly. + A host-wide default is still worth setting for anything outside this project (Dokploy's own + Traefik/dashboard containers, any other project on the same host) via `/etc/docker/daemon.json`: + ```json + { + "log-driver": "json-file", + "log-opts": { "max-size": "10m", "max-file": "3" } + } + ``` + ```bash + sudo systemctl restart docker + ``` + This only applies to _new_ containers — existing ones keep their old log config until + recreated. + +## 5. Authentication (OIDC via Keycloak) + +Login is handled by a self-hosted **Keycloak** instance, shared across projects — not something +`docker-compose.yml` deploys itself; it's a separate standing service, currently at +`https://auth.justwickedcode.dev`, realm **`apps`** (a deliberately generic realm name, not +`vroomy` — the plan is every future project registers its own client in this same realm, sharing +one user base/login session across all of them, per the original "unified login" goal). + +`frontend/src/lib/auth/oidc.ts` is provider-agnostic on purpose — it reads every endpoint +(authorize/token/jwks/userinfo) from the provider's own `/.well-known/openid-configuration` +discovery document rather than hardcoding them, and uses only standard OIDC claim names +(`sub`/`name`/`email`/`picture`). This code was built against Casdoor, then Zitadel, then +Keycloak over one long session — each switch needed zero code changes beyond env vars, which is +exactly what this design is for. If you ever switch providers again, same expectation applies. + +### Adding a new project's client to this same Keycloak realm + +1. `https://auth.justwickedcode.dev` → log in as admin → confirm you're in the **`apps`** realm + (top-left dropdown) — never configure real users in the `master` realm, that's Keycloak's own + admin account only +2. **Clients → Create client** → Client ID = your new project's name → **Client authentication: + On** (confidential client, gets a real secret — not a public/PKCE-only client) → Next → + **Valid redirect URIs**: `https:///auth/callback` → Save +3. **Credentials** tab → copy the Client Secret +4. Set that project's own `OIDC_ISSUER=https://auth.justwickedcode.dev/realms/apps`, + `OIDC_CLIENT_ID`, `OIDC_CLIENT_SECRET` env vars +5. GitHub/Google providers are already configured at the realm level (**Identity providers** in + the console) — every client in this realm gets them automatically, no per-client setup needed + +### Known gotchas, found live + +- **`OIDC_ISSUER` must include the realm path, not just the bare domain** + (`.../realms/apps`, not just `https://auth.justwickedcode.dev`) — `oidc.ts`'s discovery fetch + was originally written as `new URL('/.well-known/...', issuerBase())`, which silently drops + any path component already on the issuer URL (a leading `/` in the relative argument resets + to domain root instead of appending). Invisible with Casdoor/Zitadel (bare-domain issuers, + nothing to lose) until Keycloak's realm-scoped issuer hit it directly — fixed with plain string + concatenation instead of relative `URL` resolution. If a future provider's issuer also has a + path component, this is already handled; nothing to redo. +- **Keycloak's `firstName`/`lastName` are "root attributes" and cannot be deleted**, only + reconfigured — Realm settings → User profile → click the attribute → turn off "Required field" + and uncheck "User" under Permissions (both view and edit) to fully hide it from every + user-facing flow, including the GitHub/Google first-login screen. Vroomy never reads these + fields anyway (only the combined `name`/`email` claims), so hiding them costs nothing. +- **Cloudflare bot-protection rules must explicitly exempt the VPS's IP for every new auth + subdomain** — a rule written against one subdomain doesn't automatically cover another, even + with a broad `http.host contains "yourdomain.dev"` match, if the rule was created before that + subdomain existed. `api`/`frontend`'s own server-to-server calls to the identity provider + (token exchange, JWKS fetch, discovery) run from the VPS and will get Cloudflare-challenged + exactly like a browser-less `curl` would if this isn't set up — this is a real production + blocker, not just a local-testing inconvenience, since Vroomy's actual login flow depends on + these same calls succeeding. +- **Dokploy's "Create Environment File" toggle must be on**, and env var changes need an actual + **redeploy** (not just saving settings) to reach the running container — bit us more than once + across this whole stack, not just auth. + +## What's already handled + +- **Rate limiting** — api's one endpoint (`/api/quotes/random`) is limited to 30 requests/minute + per IP (burst of 10), tracked via `X-Forwarded-For`'s first entry when present or the direct + connection otherwise. Now that every request to api arrives from the frontend's own proxy, not + a browser, the frontend forwards the `X-Forwarded-For` header its own incoming request arrived + with (see `fetchRandomQuote` in `frontend/src/lib/typing/sentences.ts`) — without this, every + visitor's request would show up as coming from the frontend container's own address, collapsing + the per-visitor limit into one shared bucket for the whole site. `/health` is exempt — an + orchestrator's own healthcheck polling it shouldn't be able to trip a limit meant for abuse, not + routine monitoring. `ws`'s equivalent abuse control is a per-IP cap on concurrently-open + connections (see `backend/ws/README.md`'s "Concurrent-connection limiting") rather than a + request-rate limiter — a WebSocket connection is long-lived, not a discrete request a token + bucket makes sense against; this one's unaffected since browsers still connect to `ws` + directly for now. +- **Graceful shutdown** — the crawler, the API, and `ws` all stop cleanly on `SIGTERM` (what + `docker stop` and Compose both send), finishing in-flight work rather than dying mid-request. +- **Health checks** — `api`/`ws`/`frontend` all have Docker healthchecks; `crawler` doesn't + expose one (it's a background worker, not a request-serving process — its own logs are the + signal to watch, e.g. via `docker compose logs -f crawler`). + +## Findings from live testing (read before assuming this "just works") + +Real, non-obvious bugs/gotchas caught by actually running this stack end-to-end in a real +Dokploy deployment, not by reasoning about the code: + +1. **`REDIS_ADDR=redis:6379` silently connected to the wrong host.** `redis.ParseURL("redis:6379")` + doesn't error — it parses `"redis"` as a URL _scheme_ and produces `Addr="localhost:6379"`, + discarding the real host entirely. This was invisible in every local dev run only because + `REDIS_ADDR` has always literally _been_ `localhost:6379` there — the same wrong answer, by + coincidence. Fixed in `internal/db/redis.go` (`ConnectRedis` now only attempts URL parsing + for input that actually contains `://`). +2. **The crawler's real memory need is ~1.25GB, not the ~250MB first assumed.** The language + detector added for language verification (see `backend/scraper/README.md`) loads statistical + models into memory; even restricted to Latin-script languages only (~330MB, down from ~1GB + for the full 75-language set), combined with normal crawling overhead the container's real, + stable working set plateaus at ~1.25GB. `docker-compose.yml`'s `crawler` service is sized at + a 2G hard limit with `GOMEMLIMIT=1536MiB` — both **measured empirically** (memory limit + removed entirely, real usage observed over a sustained run) rather than guessed. An earlier + attempt set both values far below this real number, which didn't prevent OOM kills — it just + made the GC fight a losing battle to stay under an impossibly small target, burning 400-600% + CPU in the process before still eventually hitting the wall. `GOMEMLIMIT` only helps once + it's set _above_ actual need, giving the GC real slack to collect proactively. +3. **A plain Compose deploy lands on its own isolated network by default, not Dokploy's shared + one.** Dokploy Database resources (Postgres/Redis) run as Swarm services on `dokploy-network`; + a Compose project with no `networks:` override gets its own fresh `_default` bridge + network instead, which has no route to them at all — `DATABASE_URL`/`REDIS_ADDR` hostnames + then fail DNS resolution outright. Fixed by pinning `networks.default` to the external + `dokploy-network` in `docker-compose.yml`. `dokploy-network` is created `Attachable: true`, + so a plain (non-Swarm-service) Compose container can join it. +4. **`wget --spider http://localhost:PORT/` inside the `frontend`/`ws` containers failed with + `Connection refused`, even though the exact same request from a different container on the + same network succeeded.** Alpine/musl's `wget` resolving `localhost` prefers the IPv6 loopback + (`::1`) first; Bun/Nitro (and `ws`) only bind IPv4 `0.0.0.0`, so nothing listens on `::1` and + the healthcheck gets a real connection-refused. An unhealthy `frontend`/`ws` then gets skipped + by Dokploy's routing entirely, surfacing as a confusing plain-text `404 page not found` at the + real domain (that exact wording is Go's `net/http` default 404 — a strong tell the request + actually landed on `api`'s router instead, worth checking the Domain's Service Name field + first if you see it). Fixed by pointing both healthchecks at `127.0.0.1` instead of + `localhost`, forcing IPv4 and sidestepping the ambiguity. +5. **Dokploy's dashboard itself listens on host port 3000** — the same default port + `frontend` uses. Publishing `frontend`'s port to the host (`ports: - '127.0.0.1:3000:3000'`, + needed for the plain-`docker compose`+Caddy path) collides with it under Dokploy. Dokploy's + own Traefik doesn't need a host-published port at all — it routes to containers directly over + `dokploy-network` based on the Domain config — so `docker-compose.yml` omits `ports:` on + `frontend`/`ws` entirely now, and a would-be bind conflict never comes up under Dokploy. Only + add a host-port mapping back if deploying with plain `docker compose` + your own reverse proxy + (step 2 above), on a host that isn't also running Dokploy's dashboard on the same port. + +If you change what the crawler does (add a source, change filters) in a way that could shift its +memory profile, re-verify rather than assuming these numbers still hold — the method that found +them (`docker stats` against an unconstrained container over a sustained run) is quick to repeat. diff --git a/README.md b/README.md index 18242d0..30e8420 100644 --- a/README.md +++ b/README.md @@ -1423,6 +1423,18 @@ one to already exist for what it delivers to matter. - [ ] §6.4: `/leaderboard` endpoint + `leaderboard` view. - [ ] §11.1: `/leaderboard`, `/login`, `/register` routes. +### Phase 6 — CI and frontend test coverage + +- [ ] Add CI (e.g. GitHub Actions) running `go build`/`go test` for all + three backends (`api`, `ws`, `scraper`) and `bun run build`/`bun run + lint` for the frontend on every push/PR — currently these only run + locally, by hand, so nothing stops a broken commit from landing on + `main`. +- [ ] Add frontend test coverage. The three Go backends have real test + suites (18 `_test.go` files across crawler/dedup/fetcher/parser/db/ + api/ws); the frontend currently has none — "tested" there means + build+lint clean, not behavior-verified. + --- ## 19. Open questions & risks diff --git a/backend/api/.env.example b/backend/api/.env.example new file mode 100644 index 0000000..8c8522a --- /dev/null +++ b/backend/api/.env.example @@ -0,0 +1,14 @@ +# Copy this file to .env and fill in real values. .env is gitignored — never commit real +# credentials. + +# This API reads the same "quotes" Postgres database backend/scraper writes to — see +# backend/scraper/docker-compose.yml for that database's own container. This API does not run +# its own database and does not run migrations (backend/scraper owns the schema); run the +# scraper against a fresh database at least once before starting this API. +DATABASE_URL=postgres://postgres:postgres@localhost:55432/quotes?sslmode=disable + +# Which port this API listens on, and which frontend origin it accepts cross-origin requests +# from. In production, set this to the real deployed frontend URL (e.g. +# https://vroomy.example.com), not localhost. +API_PORT=8080 +CORS_ALLOWED_ORIGIN=http://localhost:3000 diff --git a/backend/api/Caddyfile.example b/backend/api/Caddyfile.example new file mode 100644 index 0000000..4077909 --- /dev/null +++ b/backend/api/Caddyfile.example @@ -0,0 +1,19 @@ +# Example Caddy config for terminating real HTTPS in front of the quotes API — Caddy gets you +# automatic Let's Encrypt certificates (issued and renewed with zero manual cert management) +# for the cost of one file like this. Not included as a service in docker-compose.prod.yml +# by default since it needs a real domain name pointed at this host to issue a certificate for +# — meaningless in a local/test environment, and different for every real deployment. +# +# Usage once you have a domain pointed at this host: +# 1. Copy this file to Caddyfile, replace api.example.com with your real domain. +# 2. Install Caddy on the host (not in Docker, for simplicity — apt/dnf/brew all package it, +# or see https://caddyserver.com/docs/install), or add it as its own compose service if +# you'd rather keep everything containerized. +# 3. sudo caddy run --config Caddyfile +# +# This assumes the api service is reachable at 127.0.0.1:8080 — which is exactly how +# docker-compose.prod.yml publishes it (bound to localhost only, not the public interface). + +api.example.com { + reverse_proxy 127.0.0.1:8080 +} diff --git a/backend/api/Dockerfile b/backend/api/Dockerfile new file mode 100644 index 0000000..b74ef22 --- /dev/null +++ b/backend/api/Dockerfile @@ -0,0 +1,16 @@ +# Multi-stage build: compile with the full Go toolchain, ship only the static binary in a +# minimal runtime image. +FROM golang:1.26-alpine AS build +WORKDIR /src + +COPY go.mod go.sum ./ +RUN go mod download + +COPY . . +RUN CGO_ENABLED=0 GOOS=linux go build -o /api . + +FROM alpine:3.20 +WORKDIR /app +COPY --from=build /api . +EXPOSE 8080 +ENTRYPOINT ["./api"] diff --git a/backend/api/README.md b/backend/api/README.md index 6d66218..2297426 100644 --- a/backend/api/README.md +++ b/backend/api/README.md @@ -1,13 +1,65 @@ -# Rest API +# quotes-api -## For Max: +A small read-only Go HTTP API serving quotes from the `backend/scraper` crawler's Postgres database, currently used for one thing: feeding the frontend's typing-race game real passages instead of its static placeholder pool (see `frontend/src/lib/typing/sentences.ts`). -- I was thinking about using the one of the next frameworks: - - [Elysia](https://elysiajs.com/) - - [Hono](https://hono.dev/) +## Why Go, not Elysia/Hono -You can also look into others that you like. I am personally familiar with [NestJS](https://nestjs.com/), but I think it's better to get something new for everyone. +The original plan here (see git history) was a TypeScript API using Elysia or Hono. That changed once the actual use case became concrete: this API only needs to read quotes the scraper already collected, using a schema and word-count filtering logic the scraper's own Go code already defines. Rebuilding that in a second language/runtime — a new Postgres client, a new understanding of what makes a quote "typing-race-eligible" — would have been pure duplication for zero benefit, since this service does nothing beyond query and reshape data the scraper already owns. -- They are the new "hot stuff in the town" and I taught it would be fun to try them. I personally prefer Elysia +If a broader REST API (auth, other resources, a real API for the whole app) becomes needed later, that's a fair reason to revisit the stack — Elysia/Hono are still perfectly reasonable choices for that. This one just didn't need it. -- To spin up a db, you can use the `docker-compose` file I gave you. Just run `docker-compose up -d`. To turn it off, `docker-compose down -v`. (-v = delete volume / delete the data you stored in the db already) +## Relationship to backend/scraper + +This is a **separate Go module** (its own `go.mod`), not a package inside `backend/scraper` — Go's `internal/` visibility rules mean it couldn't import `backend/scraper`'s internal packages even if it wanted to, and it doesn't need to: it only needs a Postgres connection and a couple of read queries, both trivial to have directly. + +- **Database**: points at the _same_ Postgres instance `backend/scraper` writes to (see `.env` — `DATABASE_URL`), not a separate one. Run `backend/scraper`'s own `docker-compose.yml` to get that database running. +- **Schema/migrations**: owned entirely by `backend/scraper`. This API never runs migrations and never writes — run the scraper at least once against a fresh database before starting this API, so the `quotes` table (and its `word_count` column, which this API's filtering depends on) actually exists. +- **`word_count`**: a column `backend/scraper`'s `db.SaveQuote` computes and stores on every quote at write time (indexed alongside `language`) — this API filters on it directly rather than re-splitting every candidate row's text on every request. + +## Running it + +```bash +go run . +``` + +Reads `.env` for `DATABASE_URL`, `API_PORT` (default `8080`), and `CORS_ALLOWED_ORIGIN` (default `http://localhost:3000`, i.e. the frontend's dev server). `.env` is optional — if it's just not present (e.g. running in a container, where config comes from the environment directly), that's not an error; only a genuinely malformed `.env` file that does exist is fatal. + +For a production deployment (Docker, a reverse proxy for real HTTPS, backups of the shared database) see `../../PRODUCTION.md` at the repo root. + +## Rate limiting + +`GET /api/quotes/random` is limited to 30 requests/minute per IP (burst of 10) — see `ratelimit.go`. This API has no accounts or API keys (it's called directly from any visitor's browser), so per-IP limiting is the only practical abuse control available; it's sized generously for real gameplay (one new sentence per race, roughly every 10-60+ seconds) and exists to stop a script hammering the endpoint, not to throttle real players. `/health` is exempt, since an orchestrator's own healthcheck polling it every few seconds shouldn't be able to trip a limit meant for abuse. + +Prefers `X-Forwarded-For`'s first entry when present (set this correctly in your reverse proxy — see `PRODUCTION.md`), falling back to the direct connection's address for local dev where there's no proxy in front. + +## Endpoints + +### `GET /api/quotes/random` + +Returns one random quote suitable for a typing race. + +Query params (all optional): + +| Param | Default | Notes | +| ---------- | ------- | ------------------------------------------------------------------------------------------------------------------------- | +| `language` | `en` | Matches `quotes.language` (`en` or `de` currently) | +| `minWords` | `25` | Matches the word-count floor the frontend's placeholder pool was tuned for — short passages make WPM numerically unstable | +| `maxWords` | `60` | Upper bound, same reasoning | +| `exclude` | — | Exact text to exclude, so the next race doesn't repeat the passage just shown | + +```json +{ + "text": "Racing against the clock is the only way to know how fast your fingers really are...", + "author": "Mark Twain", + "source": "wikiquote-en", + "language": "en" +} +``` + +`404` (`{"error": "no quote matches the requested filters"}`) if nothing matches — a real, expected outcome (an unusual word-count range, a language with a small corpus), not a server error. The frontend should fall back to its static placeholder pool in this case rather than treating it as broken. + +Text returned here has already been run through `MakeTypingSafe` — em/en dashes, curly quotes, and ellipsis characters are normalized to plain-ASCII equivalents (checked live: roughly a quarter of typing-eligible quotes contain one of these) so the game's exact character-match input handling doesn't punish a player for punctuation they have no reasonable way to type. Accented letters (names, borrowed words) are left as-is — typeable, just with an extra keystroke on some layouts, and not worth stripping legitimate content over. + +### `GET /health` + +`200 {"status": "ok"}` if Postgres is reachable, `503` otherwise. diff --git a/backend/api/docker-compose.yml b/backend/api/docker-compose.yml deleted file mode 100644 index 286a8a8..0000000 --- a/backend/api/docker-compose.yml +++ /dev/null @@ -1,15 +0,0 @@ -services: - db: - image: postgres:16 - restart: unless-stopped - environment: - POSTGRES_USER: ${POSTGRES_USERNAME} - POSTGRES_PASSWORD: ${POSTGRES_PASSWORD} - POSTGRES_DB: ${POSTGRES_DB} - ports: - - '5432:5432' - volumes: - - postgres_data:/var/lib/postgresql/data - -volumes: - postgres_data: diff --git a/backend/api/go.mod b/backend/api/go.mod new file mode 100644 index 0000000..8bef465 --- /dev/null +++ b/backend/api/go.mod @@ -0,0 +1,21 @@ +module quotes-api + +go 1.26.8 + +require ( + github.com/jackc/pgx/v5 v5.11.0 + github.com/joho/godotenv v1.5.1 + github.com/pressly/goose/v3 v3.27.0 + golang.org/x/time v0.16.0 +) + +require ( + github.com/jackc/pgpassfile v1.0.0 // indirect + github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect + github.com/jackc/puddle/v2 v2.2.2 // indirect + github.com/mfridman/interpolate v0.0.2 // indirect + github.com/sethvargo/go-retry v0.3.0 // indirect + go.uber.org/multierr v1.11.0 // indirect + golang.org/x/sync v0.19.0 // indirect + golang.org/x/text v0.34.0 // indirect +) diff --git a/backend/api/go.sum b/backend/api/go.sum new file mode 100644 index 0000000..1bc5519 --- /dev/null +++ b/backend/api/go.sum @@ -0,0 +1,60 @@ +github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= +github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c= +github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= +github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY= +github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto= +github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0= +github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo= +github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM= +github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg= +github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo= +github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM= +github.com/jackc/pgx/v5 v5.11.0 h1:IzBBtyK9AHqf98cctWFifYSci2hgQR/cd56wB4p+ogg= +github.com/jackc/pgx/v5 v5.11.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4= +github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo= +github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4= +github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0= +github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4= +github.com/mattn/go-isatty v0.0.20 h1:xfD0iDuEKnDkl03q4limB+vH+GxLEtL/jb4xVJSWWEY= +github.com/mattn/go-isatty v0.0.20/go.mod h1:W+V8PltTTMOvKvAeJH7IuucS94S2C6jfK/D7dTCTo3Y= +github.com/mfridman/interpolate v0.0.2 h1:pnuTK7MQIxxFz1Gr+rjSIx9u7qVjf5VOoM/u6BbAxPY= +github.com/mfridman/interpolate v0.0.2/go.mod h1:p+7uk6oE07mpE/Ik1b8EckO0O4ZXiGAfshKBWLUM9Xg= +github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w= +github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls= +github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM= +github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= +github.com/pressly/goose/v3 v3.27.0 h1:/D30gVTuQhu0WsNZYbJi4DMOsx1lNq+6SkLe+Wp59BM= +github.com/pressly/goose/v3 v3.27.0/go.mod h1:3ZBeCXqzkgIRvrEMDkYh1guvtoJTU5oMMuDdkutoM78= +github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE= +github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo= +github.com/sethvargo/go-retry v0.3.0 h1:EEt31A35QhrcRZtrYFDTBg91cqZVnFL2navjDrah2SE= +github.com/sethvargo/go-retry v0.3.0/go.mod h1:mNX17F0C/HguQMyMyJxcnU471gOZGxCLyYaFyAZraas= +github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= +github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI= +github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg= +github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U= +github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U= +go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= +go.uber.org/multierr v1.11.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y= +golang.org/x/exp v0.0.0-20260218203240-3dfff04db8fa h1:Zt3DZoOFFYkKhDT3v7Lm9FDMEV06GpzjG2jrqW+QTE0= +golang.org/x/exp v0.0.0-20260218203240-3dfff04db8fa/go.mod h1:K79w1Vqn7PoiZn+TkNpx3BUWUQksGO3JcVX6qIjytmA= +golang.org/x/sync v0.19.0 h1:vV+1eWNmZ5geRlYjzm2adRgW2/mcpevXNg50YZtPCE4= +golang.org/x/sync v0.19.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI= +golang.org/x/sys v0.41.0 h1:Ivj+2Cp/ylzLiEU89QhWblYnOE9zerudt9Ftecq2C6k= +golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks= +golang.org/x/text v0.34.0 h1:oL/Qq0Kdaqxa1KbNeMKwQq0reLCCaFtqu2eNuSeNHbk= +golang.org/x/text v0.34.0/go.mod h1:homfLqTYRFyVYemLBFl5GgL/DWEiH5wcsQ5gSh1yziA= +golang.org/x/time v0.16.0 h1:vMb6ptszcQMkcwiRTAuNNU50gom6++Q/6gY2hDM6VDE= +golang.org/x/time v0.16.0/go.mod h1:rVKOqvZeKvrDKTQiAHJ7wmwP0RzleSphoEA9RcdLA0s= +gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= +gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= +gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= +gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= +modernc.org/libc v1.68.0 h1:PJ5ikFOV5pwpW+VqCK1hKJuEWsonkIJhhIXyuF/91pQ= +modernc.org/libc v1.68.0/go.mod h1:NnKCYeoYgsEqnY3PgvNgAeaJnso968ygU8Z0DxjoEc0= +modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU= +modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg= +modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI= +modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw= +modernc.org/sqlite v1.46.1 h1:eFJ2ShBLIEnUWlLy12raN0Z1plqmFX9Qe3rjQTKt6sU= +modernc.org/sqlite v1.46.1/go.mod h1:CzbrU2lSB1DKUusvwGz7rqEKIq+NUd8GWuBBZDs9/nA= diff --git a/backend/api/internal/db/migrate.go b/backend/api/internal/db/migrate.go new file mode 100644 index 0000000..590b323 --- /dev/null +++ b/backend/api/internal/db/migrate.go @@ -0,0 +1,37 @@ +// Package db owns api's own tables (users, race_results) — distinct from backend/scraper's +// migrations, which own quotes/url_frontier in the same physical Postgres database. goose's +// migration-version tracking defaults to one shared "goose_db_version" table regardless of +// which Go binary calls it, so without SetTableName below, api's and scraper's independent +// migration histories would land in the same tracking table and get mixed together (not +// corrupted — the version numbers wouldn't collide since the two services' migration +// timestamps differ — just confusing bookkeeping). SetTableName gives api's own migrations a +// separate "api_goose_db_version" table, keeping each service's history clean. +package db + +import ( + "embed" + + "github.com/jackc/pgx/v5/pgxpool" + "github.com/jackc/pgx/v5/stdlib" + "github.com/pressly/goose/v3" +) + +//go:embed migrations/*.sql +var migrations embed.FS + +func Migrate(pool *pgxpool.Pool) error { + sqlDB := stdlib.OpenDBFromPool(pool) + + goose.SetBaseFS(migrations) + goose.SetTableName("api_goose_db_version") + + if err := goose.SetDialect("postgres"); err != nil { + return err + } + + if err := goose.Up(sqlDB, "migrations"); err != nil { + return err + } + + return nil +} diff --git a/backend/api/internal/db/migrations/20261003000000_create_users.sql b/backend/api/internal/db/migrations/20261003000000_create_users.sql new file mode 100644 index 0000000..dec8040 --- /dev/null +++ b/backend/api/internal/db/migrations/20261003000000_create_users.sql @@ -0,0 +1,19 @@ +-- +goose Up +-- id is Casdoor's own "sub" claim — a stable, globally unique string identity — not a +-- locally-generated serial. api never issues its own user ids; Casdoor already owns identity. +CREATE TABLE users ( + id TEXT PRIMARY KEY, + display_name TEXT NOT NULL, + email TEXT, + avatar_url TEXT, + car_model VARCHAR(32) NOT NULL DEFAULT 'sport', + underglow BOOLEAN NOT NULL DEFAULT false, + underglow_color VARCHAR(16) NOT NULL DEFAULT '#0284c7', + trail VARCHAR(32) NOT NULL DEFAULT 'nitro', + equipped_powerup VARCHAR(32) NOT NULL DEFAULT 'boost', + created_at TIMESTAMP NOT NULL DEFAULT NOW(), + updated_at TIMESTAMP NOT NULL DEFAULT NOW() +); + +-- +goose Down +DROP TABLE users; diff --git a/backend/api/internal/db/migrations/20261003000001_create_race_results.sql b/backend/api/internal/db/migrations/20261003000001_create_race_results.sql new file mode 100644 index 0000000..5e68e5e --- /dev/null +++ b/backend/api/internal/db/migrations/20261003000001_create_race_results.sql @@ -0,0 +1,19 @@ +-- +goose Up +CREATE TABLE race_results ( + id BIGSERIAL PRIMARY KEY, + user_id TEXT NOT NULL REFERENCES users(id) ON DELETE CASCADE, + wpm SMALLINT NOT NULL, + accuracy SMALLINT NOT NULL, + placement SMALLINT NOT NULL, + racer_count SMALLINT NOT NULL, + created_at TIMESTAMP NOT NULL DEFAULT NOW() +); + +-- Every read pattern api needs: "this user's races, newest first" (profile history) and "this +-- user's aggregate stats" (best/avg wpm, avg accuracy, wins) — both scoped by user_id, the +-- latter also scanning created_at/wpm/placement, so one composite index covers both without a +-- second index to maintain. +CREATE INDEX idx_race_results_user_id_created_at ON race_results(user_id, created_at DESC); + +-- +goose Down +DROP TABLE race_results; diff --git a/backend/api/main.go b/backend/api/main.go new file mode 100644 index 0000000..176192c --- /dev/null +++ b/backend/api/main.go @@ -0,0 +1,81 @@ +package main + +import ( + "context" + "log" + "net/http" + "os" + "os/signal" + "syscall" + "time" + + "github.com/jackc/pgx/v5/pgxpool" + "github.com/joho/godotenv" + + "quotes-api/internal/db" +) + +func main() { + ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM) + defer stop() + + // .env is a convenience for local dev — in a container, config comes from the environment + // Docker/the orchestrator already injected, and there's no .env file present at all. Only + // a genuine parse error (a malformed file that does exist) should be fatal; "the file just + // isn't there" is the expected, normal case in production and must not crash startup. + if err := godotenv.Load(); err != nil && !os.IsNotExist(err) { + log.Fatal("Error loading .env file: ", err) + } + + // This API only ever reads the quotes table — backend/scraper owns that schema (its own + // migrations create and evolve the quotes/url_frontier tables), and api never writes to or + // migrates it. api does own and migrate its own separate tables (users, race_results) — + // see internal/db/migrate.go for why that's safe to run from a second independent binary + // against the same database without colliding with the scraper's own migration history. + pool, err := pgxpool.New(ctx, os.Getenv("DATABASE_URL")) + if err != nil { + log.Fatal("Could not connect to Postgres: ", err) + } + defer pool.Close() + if err := pool.Ping(ctx); err != nil { + log.Fatal("Could not reach Postgres: ", err) + } + log.Println("Connected to Postgres!") + + if err := db.Migrate(pool); err != nil { + log.Fatal("Could not migrate api's own tables: ", err) + } + log.Println("api's own tables (users, race_results) are up to date") + + allowedOrigin := os.Getenv("CORS_ALLOWED_ORIGIN") + if allowedOrigin == "" { + allowedOrigin = "http://localhost:3000" + } + port := os.Getenv("API_PORT") + if port == "" { + port = "8080" + } + + server := &http.Server{ + Addr: ":" + port, + Handler: newServer(pool, allowedOrigin), + ReadTimeout: 10 * time.Second, + WriteTimeout: 10 * time.Second, + } + + go func() { + <-ctx.Done() + log.Println("Shutting down API server...") + shutdownCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + defer cancel() + if err := server.Shutdown(shutdownCtx); err != nil { + log.Printf("Error during server shutdown: %s\n", err) + } + }() + + log.Printf("Quotes API listening on :%s (allowing requests from %s)\n", port, allowedOrigin) + if err := server.ListenAndServe(); err != nil && err != http.ErrServerClosed { + log.Fatal("Server error: ", err) + } + log.Println("API server stopped") +} diff --git a/backend/api/quotes.go b/backend/api/quotes.go new file mode 100644 index 0000000..32b34f5 --- /dev/null +++ b/backend/api/quotes.go @@ -0,0 +1,93 @@ +package main + +import ( + "context" + "errors" + "math/rand" + + "github.com/jackc/pgx/v5" + "github.com/jackc/pgx/v5/pgxpool" +) + +// ErrNoEligibleQuote is returned when no quote matches the requested filters — a real, +// expected outcome (an unusual word-count range, a language with a small corpus), not a +// server error, so callers can map it to a 404 rather than a 500. +var ErrNoEligibleQuote = errors.New("no eligible quote found") + +// TypingQuote is the shape returned to a typing-race client — only what a typing game +// actually needs (no tags, no hashes, no internal IDs from the quotes table this reads from). +type TypingQuote struct { + Text string `json:"text"` + Author string `json:"author"` + Source string `json:"source"` + Language string `json:"language"` +} + +// RandomTypingQuote returns one random quote matching language and a word-count range +// [minWords, maxWords], excluding exclude (typically the sentence the client was just shown, +// so consecutive races don't repeat the same passage back to back) if it's non-empty. +// word_count and game_unsuitable are precomputed, indexed columns the scraper (backend/scraper) +// sets at write time — filtering on them directly is far cheaper than re-deriving them from text +// on every request. game_unsuitable flags real, correctly-sourced quotes that just aren't a good +// fit for a typing race (too short, unwritable characters, leaked citation/markup — see +// dedup.GameSuitability) — excluded here, not deleted from the corpus, since some other future +// consumer might still want them. This API only ever reads the quotes table; it never writes to +// it and never migrates it — that's the scraper's job as that schema's owner (see README). api +// does own and migrate its own separate tables (users, race_results — see internal/db). +// +// Deliberately not "ORDER BY random() LIMIT 1": that forces Postgres to evaluate random() for +// every row in the filtered candidate set and sort all of them just to keep the top 1 — an +// O(N log N) full-set sort on this API's one public endpoint, on every single request. Fine at +// today's corpus size, but the crawler this API reads from is explicitly designed to grow into +// the hundreds of thousands of rows (Goodreads' own sitemap alone lists ~5.5M quote URLs — see +// backend/scraper/README.md), at which point that sort becomes real, measurable cost paid by +// every typing-race request. Instead: pick a random starting id (from the table's live min/max, +// so it always tracks the crawler's growing corpus with no cache to invalidate) and scan forward +// from there for the first row matching the filters — an index range scan on the primary key, +// O(log N) to find the start plus a short forward scan, not a sort of the whole candidate set. +// Wraps around to the very beginning if nothing matches going forward (e.g. a narrow word-count +// range whose matches all happen to sit before the random starting point). +// +// Trade-off, deliberately accepted: not perfectly uniform — a row right after an id gap (a +// skipped ON CONFLICT insert) is marginally more likely to be picked than one in a dense run of +// consecutive ids, since it "absorbs" every random starting point landing in that gap. No player +// in a typing race could ever notice that skew; it's the standard accepted cost of avoiding a +// full-table sort for this exact problem. +func RandomTypingQuote(ctx context.Context, pool *pgxpool.Pool, language string, minWords, maxWords int, exclude string) (TypingQuote, error) { + var minID, maxID int64 + if err := pool.QueryRow(ctx, `SELECT COALESCE(MIN(id), 0), COALESCE(MAX(id), 0) FROM quotes`).Scan(&minID, &maxID); err != nil { + return TypingQuote{}, err + } + if maxID == 0 { + return TypingQuote{}, ErrNoEligibleQuote + } + randomID := minID + rand.Int63n(maxID-minID+1) + + const selectQuery = `SELECT text, author, source, language + FROM quotes + WHERE id >= $1 AND language = $2 AND word_count BETWEEN $3 AND $4 AND text != $5 + AND game_unsuitable = false + ORDER BY id + LIMIT 1` + + var q TypingQuote + err := pool.QueryRow(ctx, selectQuery, randomID, language, minWords, maxWords, exclude). + Scan(&q.Text, &q.Author, &q.Source, &q.Language) + + if errors.Is(err, pgx.ErrNoRows) { + // Nothing at or after randomID matched — wrap around to the start of the id range + // instead of giving up. + err = pool.QueryRow(ctx, selectQuery, minID, language, minWords, maxWords, exclude). + Scan(&q.Text, &q.Author, &q.Source, &q.Language) + } + + if errors.Is(err, pgx.ErrNoRows) { + return TypingQuote{}, ErrNoEligibleQuote + } + if err != nil { + return TypingQuote{}, err + } + + q.Text = MakeTypingSafe(q.Text) + return q, nil +} diff --git a/backend/api/ratelimit.go b/backend/api/ratelimit.go new file mode 100644 index 0000000..b4de480 --- /dev/null +++ b/backend/api/ratelimit.go @@ -0,0 +1,106 @@ +package main + +import ( + "net" + "net/http" + "strings" + "sync" + "time" + + "golang.org/x/time/rate" +) + +const ( + // requestsPerMinute/burstSize: this API has no accounts or API keys — it's called directly + // from any visitor's browser — so per-IP rate limiting is the only practical abuse control + // available. Sized generously for real usage: the typing race fetches one new sentence per + // race (roughly every 10-60+ seconds for an actual player), so a legitimate client should + // never come close to this; it exists to stop a script hammering the endpoint, not to + // throttle real players. + requestsPerMinute = 30 + burstSize = 10 + + // visitorTTL: how long a per-IP limiter is kept around after its last request before being + // cleaned up. Without this, every distinct IP that ever hits the API leaves a limiter in + // memory forever — a slow, unbounded leak for a public endpoint. + visitorTTL = 10 * time.Minute +) + +type visitor struct { + limiter *rate.Limiter + lastSeen time.Time +} + +// rateLimiter tracks one token-bucket limiter per client IP. Safe for concurrent use — every +// method takes the mutex; the janitor goroutine (started by withRateLimit) is the only other +// thing that touches the map. +type rateLimiter struct { + mu sync.Mutex + visitors map[string]*visitor +} + +func newRateLimiter() *rateLimiter { + return &rateLimiter{visitors: make(map[string]*visitor)} +} + +func (rl *rateLimiter) allow(ip string) bool { + rl.mu.Lock() + v, ok := rl.visitors[ip] + if !ok { + v = &visitor{limiter: rate.NewLimiter(rate.Every(time.Minute/requestsPerMinute), burstSize)} + rl.visitors[ip] = v + } + v.lastSeen = time.Now() + rl.mu.Unlock() + + return v.limiter.Allow() +} + +func (rl *rateLimiter) cleanupStale() { + rl.mu.Lock() + defer rl.mu.Unlock() + for ip, v := range rl.visitors { + if time.Since(v.lastSeen) > visitorTTL { + delete(rl.visitors, ip) + } + } +} + +// clientIP prefers X-Forwarded-For's first (original client) entry — this API is meant to run +// behind a reverse proxy in production (see README), which is what actually terminates client +// connections and sets that header; without it, every request would appear to come from the +// proxy's own IP, and rate limiting would apply to all visitors collectively instead of +// individually. Falls back to the direct connection's address, which is correct for local dev +// (no proxy in front) and is the best available answer if a proxy is present but misconfigured. +func clientIP(r *http.Request) string { + if forwarded := r.Header.Get("X-Forwarded-For"); forwarded != "" { + first, _, _ := strings.Cut(forwarded, ",") + return strings.TrimSpace(first) + } + host, _, err := net.SplitHostPort(r.RemoteAddr) + if err != nil { + return r.RemoteAddr + } + return host +} + +// withRateLimit wraps next with per-IP rate limiting and starts a background janitor that +// evicts limiters for IPs that haven't been seen in a while. The janitor stops when ctx done +// isn't wired up here deliberately — this runs for the lifetime of the process, same as the +// server itself; there's nothing else for it to be scoped to. +func withRateLimit(next http.Handler) http.Handler { + rl := newRateLimiter() + go func() { + for range time.Tick(visitorTTL) { + rl.cleanupStale() + } + }() + + return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if !rl.allow(clientIP(r)) { + writeError(w, http.StatusTooManyRequests, "rate limit exceeded") + return + } + next.ServeHTTP(w, r) + }) +} diff --git a/backend/api/ratelimit_test.go b/backend/api/ratelimit_test.go new file mode 100644 index 0000000..0ced821 --- /dev/null +++ b/backend/api/ratelimit_test.go @@ -0,0 +1,51 @@ +package main + +import ( + "net/http" + "net/http/httptest" + "testing" +) + +func TestClientIP(t *testing.T) { + cases := []struct { + name string + remote string + forward string + expected string + }{ + {"no proxy", "203.0.113.5:54321", "", "203.0.113.5"}, + {"single forwarded IP", "10.0.0.1:1234", "203.0.113.9", "203.0.113.9"}, + {"forwarded chain uses the first (original client)", "10.0.0.1:1234", "203.0.113.9, 10.0.0.1", "203.0.113.9"}, + } + + for _, c := range cases { + req := httptest.NewRequest(http.MethodGet, "/", nil) + req.RemoteAddr = c.remote + if c.forward != "" { + req.Header.Set("X-Forwarded-For", c.forward) + } + if got := clientIP(req); got != c.expected { + t.Errorf("%s: clientIP() = %q, want %q", c.name, got, c.expected) + } + } +} + +func TestRateLimiterBlocksAfterBurst(t *testing.T) { + rl := newRateLimiter() + ip := "203.0.113.1" + + allowed := 0 + for i := 0; i < burstSize+5; i++ { + if rl.allow(ip) { + allowed++ + } + } + if allowed != burstSize { + t.Errorf("expected exactly %d requests allowed before throttling, got %d", burstSize, allowed) + } + + // A different IP has its own independent budget, unaffected by the first one being spent. + if !rl.allow("203.0.113.2") { + t.Error("a different IP should not be throttled by another IP's usage") + } +} diff --git a/backend/api/server.go b/backend/api/server.go new file mode 100644 index 0000000..96f330b --- /dev/null +++ b/backend/api/server.go @@ -0,0 +1,137 @@ +package main + +import ( + "encoding/json" + "errors" + "log" + "net/http" + "strconv" + + "github.com/jackc/pgx/v5/pgxpool" +) + +const ( + // defaultMinWords/defaultMaxWords match the frontend's own placeholder-sentence design + // (frontend/src/lib/typing/sentences.ts): short passages finish in a couple of seconds + // even for an average typist, and extrapolating a multi-second burst to a per-minute rate + // is numerically unstable, so the game needs genuinely long passages, not just "a quote." + defaultMinWords = 25 + defaultMaxWords = 60 + + // maxWordsCeiling bounds what a client can ask for — not a real limitation (nothing needs + // more than a couple hundred words for a typing race), just a sanity cap so a malformed or + // hostile request can't ask for something absurd. + maxWordsCeiling = 200 +) + +// newServer builds the API's HTTP handler: CORS-wrapped routes over pool. allowedOrigin is the +// single origin allowed to call this API cross-origin (the frontend's dev or prod URL) — this +// API has no auth of its own, so it's only meant to be reachable from that one known frontend, +// not the open internet generally. +func newServer(pool *pgxpool.Pool, allowedOrigin string) http.Handler { + mux := http.NewServeMux() + // Rate limiting is scoped to the actual public endpoint only, not /health — a Docker/ + // orchestrator healthcheck polls that every few seconds, and sharing a limiter with it + // would risk the healthcheck itself tripping the limit and reporting a false outage. + mux.Handle("GET /api/quotes/random", withRateLimit(handleRandomQuote(pool))) + mux.HandleFunc("GET /health", handleHealth(pool)) + // No rate limit on these: unlike /api/quotes/random, every caller here already has to carry + // a verified identity (see withIdentity in users.go) — there's no anonymous-abuse surface + // the way there is on the one endpoint a logged-out visitor can hit freely. + mux.HandleFunc("GET /api/users/me", handleGetMe(pool)) + mux.HandleFunc("PUT /api/users/me/cosmetics", handlePutCosmetics(pool)) + mux.HandleFunc("POST /api/users/me/races", handlePostRace(pool)) + mux.HandleFunc("POST /api/users/me/races/import", handlePostRacesImport(pool)) + // CORS first: its OPTIONS preflight short-circuit must never consume a rate-limit token — + // a browser sends one before every real cross-origin request, so counting it would halve + // the effective limit for no reason. + return withCORS(allowedOrigin, mux) +} + +func withCORS(allowedOrigin string, next http.Handler) http.Handler { + return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Access-Control-Allow-Origin", allowedOrigin) + w.Header().Set("Access-Control-Allow-Methods", "GET, POST, PUT, OPTIONS") + w.Header().Set("Access-Control-Allow-Headers", "Content-Type") + if r.Method == http.MethodOptions { + w.WriteHeader(http.StatusNoContent) + return + } + next.ServeHTTP(w, r) + }) +} + +func writeJSON(w http.ResponseWriter, status int, body any) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(status) + if err := json.NewEncoder(w).Encode(body); err != nil { + log.Printf("Could not write JSON response: %s\n", err) + } +} + +func writeError(w http.ResponseWriter, status int, message string) { + writeJSON(w, status, map[string]string{"error": message}) +} + +// handleRandomQuote serves GET /api/quotes/random?language=en&minWords=25&maxWords=60&exclude=... +// All query params are optional; language/minWords/maxWords default to the same values the +// frontend's own static placeholder pool was tuned for (see defaultMinWords/defaultMaxWords). +func handleRandomQuote(pool *pgxpool.Pool) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + query := r.URL.Query() + + language := query.Get("language") + if language == "" { + language = "en" + } + + minWords, err := parseWordCountParam(query.Get("minWords"), defaultMinWords) + if err != nil { + writeError(w, http.StatusBadRequest, "invalid minWords") + return + } + maxWords, err := parseWordCountParam(query.Get("maxWords"), defaultMaxWords) + if err != nil { + writeError(w, http.StatusBadRequest, "invalid maxWords") + return + } + if minWords > maxWords { + writeError(w, http.StatusBadRequest, "minWords must not exceed maxWords") + return + } + + quote, err := RandomTypingQuote(r.Context(), pool, language, minWords, maxWords, query.Get("exclude")) + if errors.Is(err, ErrNoEligibleQuote) { + writeError(w, http.StatusNotFound, "no quote matches the requested filters") + return + } + if err != nil { + log.Printf("RandomTypingQuote failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + + writeJSON(w, http.StatusOK, quote) + } +} + +func parseWordCountParam(raw string, fallback int) (int, error) { + if raw == "" { + return fallback, nil + } + n, err := strconv.Atoi(raw) + if err != nil || n <= 0 || n > maxWordsCeiling { + return 0, errors.New("out of range") + } + return n, nil +} + +func handleHealth(pool *pgxpool.Pool) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + if err := pool.Ping(r.Context()); err != nil { + writeError(w, http.StatusServiceUnavailable, "database unreachable") + return + } + writeJSON(w, http.StatusOK, map[string]string{"status": "ok"}) + } +} diff --git a/backend/api/typingsafe.go b/backend/api/typingsafe.go new file mode 100644 index 0000000..e35f08a --- /dev/null +++ b/backend/api/typingsafe.go @@ -0,0 +1,34 @@ +package main + +import "strings" + +// typingSafeReplacer normalizes punctuation that's common in real scraped quotes but awkward +// or impossible to type on a standard keyboard, so a typing-race game can require an exact +// character-for-character match without punishing the player for the source material's own +// typography. Checked live against the corpus before deciding this was worth doing: roughly a +// quarter of quotes in the typing-game-eligible length range contain an em dash, en dash, or +// curly quote. Deliberately narrow — only swaps punctuation for a plain-ASCII equivalent, +// preserving spacing exactly (an em dash inside a word like "wisdom—knowledge" stays exactly +// one character, not " - ", so word boundaries/positions aren't shifted) — not a general +// unicode-to-ASCII transliteration, which would also strip legitimate accented letters in +// names and borrowed words that are perfectly typeable, just with an extra keystroke or two. +var typingSafeReplacer = strings.NewReplacer( + "—", "-", // em dash — + "–", "-", // en dash – + "‘", "'", // left single quote ' + "’", "'", // right single quote ' + "“", "\"", // left double quote " + "”", "\"", // right double quote " + "…", "...", // ellipsis … + " ", " ", // non-breaking space +) + +// MakeTypingSafe returns text with keyboard-unfriendly punctuation normalized to plain-ASCII +// equivalents, and any run of whitespace collapsed to a single space (defensive — the +// scraper's write path already produces single-spaced text, but this is what the typing race's +// exact-match input handling depends on, so it's worth guaranteeing here too rather than +// trusting an upstream service). +func MakeTypingSafe(text string) string { + text = typingSafeReplacer.Replace(text) + return strings.Join(strings.Fields(text), " ") +} diff --git a/backend/api/typingsafe_test.go b/backend/api/typingsafe_test.go new file mode 100644 index 0000000..f92d630 --- /dev/null +++ b/backend/api/typingsafe_test.go @@ -0,0 +1,41 @@ +package main + +import ( + "strings" + "testing" +) + +func TestMakeTypingSafe(t *testing.T) { + cases := []struct { + input string + expected string + }{ + {"wisdom—knowledge", "wisdom-knowledge"}, + {"between good and evil–ish", "between good and evil-ish"}, + {"it's a ‘quote’ within a “quote”", `it's a 'quote' within a "quote"`}, + {"wait for it…", "wait for it..."}, + {"already fine.", "already fine."}, + {"double spaced text", "double spaced text"}, + {"", ""}, + } + + for _, c := range cases { + if got := MakeTypingSafe(c.input); got != c.expected { + t.Errorf("MakeTypingSafe(%q) = %q, want %q", c.input, got, c.expected) + } + } +} + +func TestMakeTypingSafePreservesWordBoundaries(t *testing.T) { + // An em dash used mid-word must stay a single character joining the two halves, not gain + // surrounding spaces — that would split one word into two, shifting every later word's + // position, which would break the typing race's exact-match, position-based input handling. + got := MakeTypingSafe("wisdom—knowledge is power") + want := "wisdom-knowledge is power" + if got != want { + t.Errorf("MakeTypingSafe(...) = %q, want %q", got, want) + } + if fields := strings.Fields(got); len(fields) != 3 { + t.Errorf("expected 3 words after normalization, got %d: %v", len(fields), fields) + } +} diff --git a/backend/api/users.go b/backend/api/users.go new file mode 100644 index 0000000..58eb158 --- /dev/null +++ b/backend/api/users.go @@ -0,0 +1,431 @@ +package main + +import ( + "context" + "encoding/json" + "log" + "net/http" + "time" + + "github.com/jackc/pgx/v5/pgxpool" +) + +const ( + headerUserID = "X-Vroomy-User-Id" + headerUserName = "X-Vroomy-User-Name" + headerUserEmail = "X-Vroomy-User-Email" + headerUserAvatar = "X-Vroomy-User-Avatar" + + maxCosmeticFieldLen = 64 + maxDisplayNameLen = 255 + + // Sanity ceilings, not real game rules — same spirit as quotes.go's maxWordsCeiling, just + // guarding against a malformed or hostile payload rather than modeling an actual limit. + maxWPM = 500 + maxRacerCount = 64 + maxImportBatch = 50 // mirrors the old client-side MAX_RACE_HISTORY cap (frontend/useProfile.ts) +) + +// identity is the caller's verified Casdoor identity, forwarded by frontend's server as plain +// headers — see withIdentity below for why these are safe to trust without their own signature. +type identity struct { + id string + displayName string + email string + avatarURL string +} + +// withIdentity rejects any request with no X-Vroomy-User-Id header, then hands the handler a +// parsed identity. api has no public domain (see docker-compose.yml) and is reachable only from +// frontend's own internal proxy — these headers are only ever set there, after frontend has +// independently verified the caller's Casdoor session, never by a browser directly. Same trust +// model as the already-established X-Forwarded-For passthrough in handleRandomQuote: nothing +// outside this closed compose network can ever set them. +func withIdentity(next func(http.ResponseWriter, *http.Request, identity)) http.HandlerFunc { + return func(w http.ResponseWriter, r *http.Request) { + id := r.Header.Get(headerUserID) + if id == "" { + writeError(w, http.StatusUnauthorized, "missing identity") + return + } + next(w, r, identity{ + id: id, + displayName: truncate(r.Header.Get(headerUserName), maxDisplayNameLen), + email: r.Header.Get(headerUserEmail), + avatarURL: r.Header.Get(headerUserAvatar), + }) + } +} + +func truncate(s string, max int) string { + if len(s) <= max { + return s + } + return s[:max] +} + +// RaceResult is one completed race, as stored and as returned to the client. +type RaceResult struct { + ID int64 `json:"id"` + Date string `json:"date"` + WPM int `json:"wpm"` + Accuracy int `json:"accuracy"` + Placement int `json:"placement"` + RacerCount int `json:"racerCount"` +} + +// ProfileStats mirrors frontend/src/lib/profile/useProfile.ts's ProfileStats shape exactly, so +// the frontend can drop its own client-side computeStats once it reads this directly — computed +// here over the user's full race history, not a capped recent window like the old localStorage +// version needed (Postgres has no reason to cap stored rows the way local storage size did). +type ProfileStats struct { + RacesPlayed int `json:"racesPlayed"` + BestWPM int `json:"bestWpm"` + AvgWPM int `json:"avgWpm"` + AvgAccuracy int `json:"avgAccuracy"` + Wins int `json:"wins"` +} + +// UserProfile is the full GET /api/users/me response: identity, cosmetics, a bounded recent-race +// window for display, and lifetime stats computed independently of that window's size. +type UserProfile struct { + ID string `json:"id"` + DisplayName string `json:"displayName"` + Email string `json:"email,omitempty"` + AvatarURL string `json:"avatarUrl,omitempty"` + CarModel string `json:"carModel"` + Underglow bool `json:"underglow"` + UnderglowColor string `json:"underglowColor"` + Trail string `json:"trail"` + EquippedPowerup string `json:"equippedPowerup"` + Races []RaceResult `json:"races"` + Stats ProfileStats `json:"stats"` +} + +// upsertIdentity guarantees a users row exists and its identity-owned columns (name/email/ +// avatar — whatever Casdoor itself owns) stay fresh, regardless of call order between the +// endpoints below. Cosmetic columns are left at their DEFAULT on first insert and untouched on +// conflict here — only updateCosmetics (below) ever writes them. +// +// COALESCE(NULLIF(excluded.x, ”), users.x) rather than a bare "excluded.x": not every endpoint +// that calls this forwards the full identity (e.g. handlePostRace only needs X-Vroomy-User-Id to +// do its job) — an unconditional overwrite would blank a previously-stored display name/email/ +// avatar back to empty on any call that only sent the id header. Caught live: a PUT cosmetics +// call sending only the id header wiped a display name GET /me had just set moments before. +func upsertIdentity(ctx context.Context, pool *pgxpool.Pool, id identity) error { + _, err := pool.Exec(ctx, ` + INSERT INTO users (id, display_name, email, avatar_url) + VALUES ($1, $2, $3, $4) + ON CONFLICT (id) DO UPDATE SET + display_name = COALESCE(NULLIF(excluded.display_name, ''), users.display_name), + email = COALESCE(NULLIF(excluded.email, ''), users.email), + avatar_url = COALESCE(NULLIF(excluded.avatar_url, ''), users.avatar_url), + updated_at = NOW() + `, id.id, id.displayName, id.email, id.avatarURL) + return err +} + +// userRow is every column the frontend needs back on every response — read fresh from the +// database rather than echoed from the current request's headers, so a call that only forwarded +// X-Vroomy-User-Id (e.g. handlePostRace) still returns the real stored name/email/avatar instead +// of blanks. +type userRow struct { + displayName, email, avatarURL string + carModel, underglowColor, trail, equippedPowerup string + underglow bool +} + +func loadUserRow(ctx context.Context, pool *pgxpool.Pool, userID string) (userRow, error) { + var row userRow + err := pool.QueryRow(ctx, ` + SELECT display_name, email, avatar_url, car_model, underglow, underglow_color, trail, equipped_powerup + FROM users WHERE id = $1 + `, userID).Scan(&row.displayName, &row.email, &row.avatarURL, &row.carModel, &row.underglow, &row.underglowColor, &row.trail, &row.equippedPowerup) + return row, err +} + +func computeStats(ctx context.Context, pool *pgxpool.Pool, userID string) (ProfileStats, error) { + var stats ProfileStats + err := pool.QueryRow(ctx, ` + SELECT + COUNT(*)::int, + COALESCE(MAX(wpm), 0)::int, + COALESCE(ROUND(AVG(wpm))::int, 0), + COALESCE(ROUND(AVG(accuracy))::int, 0), + COUNT(*) FILTER (WHERE placement = 1)::int + FROM race_results + WHERE user_id = $1 + `, userID).Scan(&stats.RacesPlayed, &stats.BestWPM, &stats.AvgWPM, &stats.AvgAccuracy, &stats.Wins) + return stats, err +} + +// listRecentRaces returns at most limit races, newest first — a display window, not the +// lifetime record computeStats reads from (see UserProfile's own doc comment). +func listRecentRaces(ctx context.Context, pool *pgxpool.Pool, userID string, limit int) ([]RaceResult, error) { + rows, err := pool.Query(ctx, ` + SELECT id, wpm, accuracy, placement, racer_count, created_at + FROM race_results + WHERE user_id = $1 + ORDER BY created_at DESC + LIMIT $2 + `, userID, limit) + if err != nil { + return nil, err + } + defer rows.Close() + + races := []RaceResult{} + for rows.Next() { + var r RaceResult + var createdAt time.Time + if err := rows.Scan(&r.ID, &r.WPM, &r.Accuracy, &r.Placement, &r.RacerCount, &createdAt); err != nil { + return nil, err + } + r.Date = createdAt.Format(time.RFC3339) + races = append(races, r) + } + return races, rows.Err() +} + +func loadFullProfile(ctx context.Context, pool *pgxpool.Pool, userID string) (UserProfile, error) { + row, err := loadUserRow(ctx, pool, userID) + if err != nil { + return UserProfile{}, err + } + races, err := listRecentRaces(ctx, pool, userID, 50) + if err != nil { + return UserProfile{}, err + } + stats, err := computeStats(ctx, pool, userID) + if err != nil { + return UserProfile{}, err + } + return UserProfile{ + ID: userID, + DisplayName: row.displayName, + Email: row.email, + AvatarURL: row.avatarURL, + CarModel: row.carModel, + Underglow: row.underglow, + UnderglowColor: row.underglowColor, + Trail: row.trail, + EquippedPowerup: row.equippedPowerup, + Races: races, + Stats: stats, + }, nil +} + +// handleGetMe serves GET /api/users/me: ensures the row exists (first login or returning +// visitor, same codepath either way) and returns the full profile. +func handleGetMe(pool *pgxpool.Pool) http.HandlerFunc { + return withIdentity(func(w http.ResponseWriter, r *http.Request, id identity) { + if err := upsertIdentity(r.Context(), pool, id); err != nil { + log.Printf("upsertIdentity failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + profile, err := loadFullProfile(r.Context(), pool, id.id) + if err != nil { + log.Printf("loadFullProfile failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + writeJSON(w, http.StatusOK, profile) + }) +} + +type cosmeticsRequest struct { + CarModel string `json:"carModel"` + Underglow bool `json:"underglow"` + UnderglowColor string `json:"underglowColor"` + Trail string `json:"trail"` + EquippedPowerup string `json:"equippedPowerup"` +} + +// handlePutCosmetics serves PUT /api/users/me/cosmetics. Field values aren't checked against +// frontend's own allowed sets (frontend/src/lib/trails.ts, powerups.ts) — the frontend's own UI +// already only ever offers valid choices, and duplicating that allowlist here would just be two +// places to keep in sync. Only length-capped, not value-validated, same spirit as truncate above. +func handlePutCosmetics(pool *pgxpool.Pool) http.HandlerFunc { + return withIdentity(func(w http.ResponseWriter, r *http.Request, id identity) { + var req cosmeticsRequest + if err := json.NewDecoder(r.Body).Decode(&req); err != nil { + writeError(w, http.StatusBadRequest, "invalid body") + return + } + if req.CarModel == "" || req.Trail == "" || req.EquippedPowerup == "" || req.UnderglowColor == "" { + writeError(w, http.StatusBadRequest, "carModel, trail, equippedPowerup, and underglowColor are required") + return + } + + ctx := r.Context() + if err := upsertIdentity(ctx, pool, id); err != nil { + log.Printf("upsertIdentity failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + _, err := pool.Exec(ctx, ` + UPDATE users SET + car_model = $2, + underglow = $3, + underglow_color = $4, + trail = $5, + equipped_powerup = $6, + updated_at = NOW() + WHERE id = $1 + `, id.id, truncate(req.CarModel, maxCosmeticFieldLen), req.Underglow, + truncate(req.UnderglowColor, maxCosmeticFieldLen), + truncate(req.Trail, maxCosmeticFieldLen), + truncate(req.EquippedPowerup, maxCosmeticFieldLen)) + if err != nil { + log.Printf("update cosmetics failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + + profile, err := loadFullProfile(ctx, pool, id.id) + if err != nil { + log.Printf("loadFullProfile failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + writeJSON(w, http.StatusOK, profile) + }) +} + +type raceRequest struct { + WPM int `json:"wpm"` + Accuracy int `json:"accuracy"` + Placement int `json:"placement"` + RacerCount int `json:"racerCount"` +} + +func (req raceRequest) validate() string { + if req.WPM < 0 || req.WPM > maxWPM { + return "wpm out of range" + } + if req.Accuracy < 0 || req.Accuracy > 100 { + return "accuracy out of range" + } + if req.RacerCount < 1 || req.RacerCount > maxRacerCount { + return "racerCount out of range" + } + if req.Placement < 1 || req.Placement > req.RacerCount { + return "placement out of range" + } + return "" +} + +// handlePostRace serves POST /api/users/me/races — records one just-finished race. +func handlePostRace(pool *pgxpool.Pool) http.HandlerFunc { + return withIdentity(func(w http.ResponseWriter, r *http.Request, id identity) { + var req raceRequest + if err := json.NewDecoder(r.Body).Decode(&req); err != nil { + writeError(w, http.StatusBadRequest, "invalid body") + return + } + if msg := req.validate(); msg != "" { + writeError(w, http.StatusBadRequest, msg) + return + } + + ctx := r.Context() + if err := upsertIdentity(ctx, pool, id); err != nil { + log.Printf("upsertIdentity failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + + var result RaceResult + var createdAt time.Time + err := pool.QueryRow(ctx, ` + INSERT INTO race_results (user_id, wpm, accuracy, placement, racer_count) + VALUES ($1, $2, $3, $4, $5) + RETURNING id, wpm, accuracy, placement, racer_count, created_at + `, id.id, req.WPM, req.Accuracy, req.Placement, req.RacerCount). + Scan(&result.ID, &result.WPM, &result.Accuracy, &result.Placement, &result.RacerCount, &createdAt) + if err != nil { + log.Printf("insert race_results failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + result.Date = createdAt.Format(time.RFC3339) + writeJSON(w, http.StatusCreated, result) + }) +} + +type importRaceEntry struct { + raceRequest + Date string `json:"date"` +} + +// handlePostRacesImport serves POST /api/users/me/races/import — a one-shot bulk carry-over of +// whatever race history a browser already had in localStorage (see useProfile.ts) from before +// this user ever logged in, so logging in for the first time doesn't silently discard it. +// Deliberately not idempotent: calling this twice duplicates rows. That's an acceptable trade-off +// for a one-time, non-adversarial UX nicety — the frontend only ever calls it once, tracked by a +// local "already imported" flag — not something worth a dedup mechanism for. +func handlePostRacesImport(pool *pgxpool.Pool) http.HandlerFunc { + return withIdentity(func(w http.ResponseWriter, r *http.Request, id identity) { + var entries []importRaceEntry + if err := json.NewDecoder(r.Body).Decode(&entries); err != nil { + writeError(w, http.StatusBadRequest, "invalid body") + return + } + if len(entries) > maxImportBatch { + writeError(w, http.StatusBadRequest, "too many races in one import") + return + } + for _, e := range entries { + if msg := e.validate(); msg != "" { + writeError(w, http.StatusBadRequest, "race entry: "+msg) + return + } + } + + ctx := r.Context() + if err := upsertIdentity(ctx, pool, id); err != nil { + log.Printf("upsertIdentity failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + + // One multi-row INSERT rather than N round trips — same batching principle as the + // crawler's own "one INSERT per page, not per quote" (see backend/scraper). + tx, err := pool.Begin(ctx) + if err != nil { + log.Printf("begin import tx failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + defer tx.Rollback(ctx) //nolint:errcheck // no-op once Commit succeeds + + for _, e := range entries { + createdAt := time.Now() + if parsed, err := time.Parse(time.RFC3339, e.Date); err == nil { + createdAt = parsed + } + if _, err := tx.Exec(ctx, ` + INSERT INTO race_results (user_id, wpm, accuracy, placement, racer_count, created_at) + VALUES ($1, $2, $3, $4, $5, $6) + `, id.id, e.WPM, e.Accuracy, e.Placement, e.RacerCount, createdAt); err != nil { + log.Printf("insert imported race failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + } + if err := tx.Commit(ctx); err != nil { + log.Printf("commit import tx failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + + profile, err := loadFullProfile(ctx, pool, id.id) + if err != nil { + log.Printf("loadFullProfile failed: %s\n", err) + writeError(w, http.StatusInternalServerError, "internal error") + return + } + writeJSON(w, http.StatusOK, profile) + }) +} diff --git a/backend/scraper/.env.example b/backend/scraper/.env.example new file mode 100644 index 0000000..2adbc84 --- /dev/null +++ b/backend/scraper/.env.example @@ -0,0 +1,26 @@ +# Copy this file to .env and fill in real values. .env is gitignored — never commit real +# credentials. The placeholder values below (postgres/postgres, redis) are fine for local dev +# against docker-compose.yml; replace them with strong, randomly-generated secrets before any +# real deployment. + +GOOSE_DRIVER=postgres +GOOSE_DBSTRING=postgres://postgres:postgres@localhost:55432/quotes?sslmode=disable +GOOSE_MIGRATION_DIR=./internal/db/migrations + +# Postgres DB +POSTGRES_USER=postgres +POSTGRES_PASSWORD=postgres +POSTGRES_DB=quotes +DATABASE_URL=postgres://postgres:postgres@localhost:55432/quotes?sslmode=disable + +# Redis DB +REDIS_ADDR=localhost:6379 +REDIS_PASSWORD=redis +REDIS_DB=0 + +# Optional: LOG_LEVEL controls verbosity. Unset/"info" (the default) logs everything, including +# routine per-page/per-quote/per-discovery narration. "warn" drops that routine narration but +# keeps every skip/retry/fallback/pause-resume line (a rejected quote, a low-yield category +# abandoned, a rate-limit backoff) alongside real failures — the recommended production setting. +# "error" drops warnings too, leaving only genuine failures (DB/Redis errors, giving up on a URL). +# LOG_LEVEL=warn diff --git a/backend/scraper/Dockerfile b/backend/scraper/Dockerfile new file mode 100644 index 0000000..695b3fa --- /dev/null +++ b/backend/scraper/Dockerfile @@ -0,0 +1,29 @@ +# Multi-stage build: compile with the full Go toolchain, ship only the static binary in a +# minimal runtime image — the build stage (~1GB+ with the Go toolchain) never ends up in the +# final image, which stays a few MB on top of the base. +FROM golang:1.25-alpine AS build +WORKDIR /src + +# Dependencies cached in their own layer — only re-downloaded when go.mod/go.sum actually change, +# not on every source edit. +COPY go.mod go.sum ./ +RUN go mod download + +COPY . . +# CGO_ENABLED=0 for a fully static binary — runs on the minimal base image below with no libc +# dependency to worry about matching. +RUN CGO_ENABLED=0 GOOS=linux go build -o /crawler ./cmd/crawler + +FROM alpine:3.20 +# ca-certificates: required for the crawler's own outbound HTTPS fetches (Goodreads, Wikiquote) +# to verify TLS certificates — omitting this is a common "works in dev, fails in prod" trap for +# a from-scratch/alpine image, since Go's HTTP client relies on the OS certificate store. +# +# git: the crawler's build-info log line (cmd/crawler/main.go's logBuildInfo) shells out to git +# rev-parse/status. That's meaningless inside a container built from a snapshot with no .git +# directory anyway, so it'll just fail soft and log "could not determine build commit" — fine, +# not worth adding git+the full repo history to the image just to make that one line work. +RUN apk add --no-cache ca-certificates +WORKDIR /app +COPY --from=build /crawler . +ENTRYPOINT ["./crawler"] diff --git a/backend/scraper/README.md b/backend/scraper/README.md index 61207d4..fc83fe6 100644 --- a/backend/scraper/README.md +++ b/backend/scraper/README.md @@ -4,14 +4,16 @@ A scalable web crawler that collects quotes from multiple sources and stores the ## Sources -| Source | Method | Status | Notes | -| ------------------------------------------------------------------------- | ----------------------- | ------------------------- | -------------------------------------------- | -| [quotes.toscrape.com](https://quotes.toscrape.com) | Crawler | ✅ Done | Sandbox site, 100 quotes | -| [Quotable API](https://api.quotable.io) | API Fetcher | ❓ Api at the moment down | 5000+ curated quotes, no auth needed | -| [BrainyQuote](https://brainyquote.com) | Crawler | ⬜ Planned | 100k+ quotes, clean HTML | -| [Wikiquote](https://en.wikiquote.org) | Crawler (MediaWiki API) | ⬜ Planned | Millions of quotes, needs wikitext parser | -| [GoodReads](https://goodreads.com/quotes) | Crawler | ⬜ Planned | Millions of quotes, aggressive bot detection | -| [Kaggle Dataset](https://www.kaggle.com/datasets/akmittal/quotes-dataset) | CSV Import | ⬜ Planned | 500k+ quotes, bulk seed | +This project runs on Goodreads plus fourteen Wikiquote language editions — rather than spreading across every quote site that exists. All are plain Go (`net/http` + goquery), no JS rendering, no bypass of anything. Everything else considered was either cut for cause (see below) or just isn't a priority next to Goodreads' scale. + +| Source | Method | Status | Notes | +| ------------------------------------------------------------------------------------------------------------------------------------ | ---------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [GoodReads](https://goodreads.com/quotes) | Crawler | ✅ Done | Official sitemap ≈ 5.5M quote URLs. No hard block, but soft-throttles sustained sequential crawling — see note below. Primary volume source. | +| [Wikiquote](https://en.wikiquote.org) | Crawler | ✅ Done | No anti-bot wall (verified live); parses rendered author-page HTML (not wikitext). Self-sustaining primarily via a real MediaWiki title-discovery side channel, supplemented by citation-link `NextURLs` and a random-page fallback — see note below. | +| [Wikiquote (German)](https://de.wikiquote.org) | Crawler | ✅ Done | Same MediaWiki infrastructure as English Wikiquote, reused via the generalized `wikiquoteSite` discovery mechanism (`internal/crawler/wikiquote_discovery.go`), but a structurally different page layout needed its own parser (`internal/parser/germanwikiquote.go`) — see note below. Tags saved quotes `language='de'`. | +| Wikiquote (French / Spanish / Italian / Portuguese / Polish / Swedish / Romanian / Czech / Hungarian / Danish / Norwegian / Finnish) | Crawler | ✅ Done | Same `wikiquoteSite` discovery mechanism, one shared parser (`internal/parser/wikiquote_i18n.go`, `LocalizedWikiquoteParser`) — see "Localized Wikiquote editions" below for why these share one implementation instead of one file each, and its known per-article noise tradeoff. Dutch was checked and deliberately skipped — see that section. | +| [quotes.toscrape.com](https://quotes.toscrape.com) | Test infra | ✅ Done | Not a corpus source — a public scraping sandbox kept solely as the live-fetch integration test target. | +| [Kaggle Dataset](https://www.kaggle.com/datasets/akmittal/quotes-dataset) | CSV Import | ⬜ Planned | Dataset page confirmed live; download needs a Kaggle account/API token (auth, not scraping). 500k+ claimed, unverified until downloaded. | ## Stack @@ -21,6 +23,13 @@ A scalable web crawler that collects quotes from multiple sources and stores the - **HTML Parsing:** goquery - **Queue:** Redis + Asynq (planned) +## Deployment + +`Dockerfile` builds a static binary in a multi-stage build (no CGO, migrations embedded via `//go:embed` — no separate files to copy into the runtime image). For a real production deployment (docker-compose stack alongside `backend/api`, secrets, backups, reverse proxy for the API's HTTPS) see `../../PRODUCTION.md` at the repo root — it documents two real bugs a from-scratch containerized run caught live, worth reading before assuming this "just works" in a container: + +- **`ConnectRedis` used to silently connect to the wrong host** for a plain `host:port` address like a Docker service name (`redis:6379`) — invisible in local dev only because `REDIS_ADDR` there has always literally _been_ `localhost:6379`. Fixed in `internal/db/redis.go`. +- **Real memory need is ~1.25GB**, not a naive guess — the language detector (see "Real language verification" below) loads statistical models into memory, and the actual stable working set was measured live (container memory limit removed, usage observed over a sustained run), not assumed from first principles. + ## Architecture ``` @@ -105,6 +114,18 @@ go test -v ./... # verbose output go test -cover ./... # with coverage ``` +### Live integration tests + +`internal/crawler/integration_test.go` (build tag `integration`) runs against the project's own `docker-compose.yml` Postgres + Redis instead of testcontainers, and includes three tests that do a real HTTP fetch against the live internet (`goodreads.com`, `quotes.toscrape.com`, `en.wikiquote.org`) — no fixtures, no mocking — parsing real pages, saving real quotes, and (for Goodreads and toscrape) pushing a real discovered URL into the Redis `frontier` ZSET. + +```bash +docker compose up -d +go test -tags=integration -count=1 ./internal/crawler/... -v +docker compose down +``` + +State persists in the compose volumes across runs by design (dev env) — the tests are idempotent (dedup on quotes, `ON CONFLICT DO NOTHING` on URLs). + ## URL Frontier The crawler maintains a **URL frontier** — a persistent priority queue of URLs to crawl. PostgreSQL is the source of truth; Redis is the working queue. @@ -134,14 +155,15 @@ CREATE INDEX idx_url_frontier_status_priority ON url_frontier(status, priority); ``` Seed URLs → INSERT into url_frontier (status=pending) ↓ -On startup → load all pending rows → ZADD into Redis Sorted Set (score = priority) +On startup → load all pending rows → ZADD into their per-source Redis Sorted Set ↓ -Fetcher workers → ZPOPMIN from Redis → fetch page → mark status=in_progress in DB +Crawl loop → round-robin across sources (see below) → ZPOPMIN from whichever source's turn + it is → fetch page → mark status=in_progress in DB ↓ -Parser workers → extract quotes + discover next-page URLs +Parser → extract quotes + discover next-page URLs ↓ Quotes → quotes table -New URLs → INSERT INTO url_frontier ON CONFLICT DO NOTHING + ZADD Redis +New URLs → INSERT INTO url_frontier ON CONFLICT DO NOTHING + ZADD into that source's Redis set ↓ Mark URL → status=done or status=failed (increment error_count) ``` @@ -150,19 +172,29 @@ On Redis restart → reload all `status=pending` rows from DB back into Redis (s ### Storage functions -| Function | Status | Notes | -| ---------------------- | ------- | ---------------------------------------------- | -| `db.SaveURL` | ✅ Done | Insert with `ON CONFLICT (url) DO NOTHING` | -| `db.MarkURLDone` | ✅ Done | Sets `status=done`, `last_crawled_at=NOW()` | -| `db.MarkURLFailed` | ✅ Done | Sets `status=failed`, increments `error_count` | -| `db.GetPendingURLs` | ✅ Done | Returns all pending rows ordered by priority | -| `db.WarmFrontierCache` | ✅ Done | Loads pending URLs into Redis on startup | -| `db.PushURL` | ✅ Done | `ZADD frontier ` | -| `db.PopURL` | ✅ Done | `ZPOPMIN frontier` — returns next URL to crawl | +| Function | Status | Notes | +| ---------------------------------------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------- | +| `db.SaveURL` | ✅ Done | Insert with `ON CONFLICT (url) DO NOTHING` | +| `db.MarkURLDone` | ✅ Done | Sets `status=done`, `last_crawled_at=NOW()` | +| `db.MarkURLFailed` | ✅ Done | Sets `status=failed`, increments `error_count` | +| `db.GetPendingURLs` | ✅ Done | Returns all pending rows ordered by priority | +| `db.WarmFrontierCache` | ✅ Done | Loads pending URLs into their per-source Redis queue on startup | +| `db.PushURL(ctx, source, url, priority)` | ✅ Done | `ZADD frontier: ` — per-source key, not one shared queue (see below) | +| `db.PopURL(ctx, source)` | ✅ Done | `ZPOPMIN frontier:` — returns the next URL for that specific source | +| `db.HasAnyURLs` | ✅ Done | Any row, any status — used to gate seeding so it only ever runs once (see below) | +| `db.RequeueStuckInProgress` | ✅ Done | Resets `in_progress` → `pending` on startup — recovers work orphaned by a killed/crashed process | +| `db.CountPendingBySource` | ✅ Done | Used to detect when a source has run dry and needs more work discovered | +| `db.GetDiscoveryCursor` / `SetDiscoveryCursor` | ✅ Done | Redis-backed continuation token per source, so title discovery (e.g. Wikiquote's `allpages`) resumes instead of restarting from the beginning | + +### Per-source queues, not one shared priority queue — a real bug, found live + +Frontier state used to be one shared Redis sorted set (`frontier`) across all sources, with `scoring.CalculatePriority` giving each source a different base score so `ZPOPMIN` would naturally prefer the "healthier" source. **This works only as long as no source builds up a real backlog.** Confirmed live, and it's the reason the design changed: once Wikiquote's title discovery (below) gave it hundreds of pending pages, Goodreads' single pending page sat **completely unserved** — 476 Wikiquote pending vs. Goodreads' 1, for as long as the process ran — because Wikiquote's base score was always numerically lower, so `ZPOPMIN` picked a Wikiquote URL first every single time, regardless of how many were queued. + +Fixed by giving each source its own Redis key (`frontier:`) and its own dedicated worker goroutine (`Crawler.runWorker` in `crawler.go`, one per source, launched from `Run()`) instead of one shared loop picking whichever source's turn it was. Priority (`scoring.CalculatePriority`) only orders URLs _within_ a single source's own queue (e.g. by depth) — it has no say in cross-source scheduling at all, which is what actually guarantees every source keeps progressing regardless of how deep another's backlog gets. Verified live after the original round-robin fix: Goodreads advanced two full pages on its own ~20s cadence, cleanly interleaved with Wikiquote fetches every ~6-7s. German Wikiquote (`wikiquote-de`) joined the same setup once added — a third source is just another goroutine launched from `Run()` with its own cooldown and top-up function, nothing else about the mechanism changes. (This started as single-threaded round-robin across a shared loop, then became one goroutine per source once single-threading itself turned out to be the next real bottleneck — see "Concurrent per-source workers" below.) ### Priority Scoring (`internal/scoring`) -Lower score = crawled sooner: +Lower score = crawled sooner **within that source's own queue** — it no longer has any cross-source effect (see above). ``` score = source_base + (depth × DepthPenalty) + (error_count × ErrorPenalty) @@ -176,9 +208,9 @@ score = source_base + (depth × DepthPenalty) + (error_count × ErrorPenalty) | Source | Base Score | | --------------------------- | ---------- | -| `scoring.SourceQuotable` | 1.0 | -| `scoring.SourceBrainyQuote` | 5.0 | -| `scoring.SourceGoodreads` | 20.0 | +| `scoring.SourceWikiquote` | 2.0 | +| `scoring.SourceWikiquoteDE` | 2.0 | +| `scoring.SourceGoodreads` | 10.0 | ### Redis primitives used @@ -189,22 +221,102 @@ score = source_base + (depth × DepthPenalty) + (error_count × ErrorPenalty) > **Note:** Visited URL dedup is handled by the `url_frontier` table itself via `UNIQUE` on `url` + `ON CONFLICT DO NOTHING`. No separate visited set needed. -### Asynq weighted queues (planned) +### Politeness / per-source rate limiting (`crawler.go`) -``` -critical (weight 6) → high-yield, reliable sources (Quotable API) -default (weight 3) → mid-tier sources (BrainyQuote) -low (weight 1) → slow or unreliable sources (Goodreads) -``` +Each source's worker goroutine enforces its own minimum delay between two of its own fetches, **per-source, not one flat number** (`minSourceDelayBySource`) — Goodreads gets 20s, both Wikiquote editions get 6s, because they've earned different levels of caution: a 20-request no-delay burst against Wikiquote (during source research) got a flat `429` after ~11 requests but has otherwise been fine at any pace tested, while Goodreads soft-throttles regardless of pacing. Pooling every source under one delay would either be too lax for Goodreads or needlessly slow for Wikiquote. German Wikiquote reuses English Wikiquote's 6s cooldown outright rather than being separately tuned — same MediaWiki software, no throttling behavior observed yet that would justify treating it differently. + +This is a plain local `lastFetch time.Time` inside `Crawler.runWorker`'s loop, not a shared map — since each source now has its own dedicated goroutine, there's nothing else touching that variable to synchronize against. A source waiting out its cooldown just sleeps; it doesn't block any other source's worker, which is exactly what fixed the "waiting too long between fetches" feeling that motivated this (see "Concurrent per-source workers" below) — previously the entire process was one shared loop, so even with per-source cooldowns and round-robin, a source on cooldown still had to wait for whichever source got served that turn to finish its fetch first. + +**Adaptive backoff — a source backs itself off when it seems rate-limited.** A slow fetch (`slowFetchWarn`, >10s — Goodreads' tarpit signature) or an outright fetch error (including a Wikiquote `429`, which surfaces as `fetcher.Fetch` returning a `*fetcher.StatusError{StatusCode: 429}`) makes `processURL` return `stalled=true`; the calling worker then advances its own `lastFetch` so the next cooldown check yields `stallPenalty` (60s, matching Goodreads' own observed stall duration) instead of the normal per-source delay. Purely local to that source's own goroutine — it no longer needs to "switch to" another source's queue, since that other source was never blocked in the first place. + +**Bounded pacing experiment for Wikiquote (`wikiquoteExperimentalDelay`, currently 4s).** Requested explicitly, with the tradeoff understood: both Wikiquote editions now start at a narrower delay than the proven-safe 6s entry in `minSourceDelayBySource`, chosen from the two data points actually in evidence (6s: proven safe over hundreds of fetches; ~2-3s: confirmed unsafe, the exact pace that produced a live 429 from Wikiquote's discovery API) rather than a third proven-safe number — 4s is a considered guess between them, not a new fact. `fetcher.StatusError` and `fetcher.IsRateLimited(err)` (checking specifically for a 429, not any failure) let `processURL`/`topUpWikiquoteSiteIfEmpty` report a `rateLimited` signal up to `Crawler.runWorker`; the moment either Wikiquote worker actually observes one at this pace, `Crawler.abandonExperimentalPace` permanently widens that worker's `minDelay` back to the safe 6s for the rest of the run — automatically, not something a human needs to be watching logs to catch. Goodreads is unaffected; it never runs at an experimental pace. + +### Concurrent per-source workers + +Originally the crawl loop was one shared loop, round-robining across sources (`Crawler.popNextReady`) so no single source could starve another — a real fix for a real bug (see "Per-source queues" above), but it left a different problem: the loop was still single-threaded, so a slow or stalled fetch on one source (Goodreads' ~60s tarpit is the concrete example) blocked _every_ source's progress until it returned, even a different source's URL that was already off cooldown and ready to go. Round-robin, per-source cooldowns, and adaptive backoff all only act _between_ fetches — none of them can help once a fetch is already in flight. + +Fixed by giving each source in `sourceOrder`'s old place — Goodreads, English Wikiquote, German Wikiquote — its own dedicated goroutine (`Crawler.runWorker`, launched once per source from `Run()`), each running its own independent loop: wait out its own cooldown, pop from its own Redis queue, fetch, parse, save, mark done, repeat. `db.Store`'s underlying `pgxpool.Pool` and `redis.Client` are both already safe for concurrent use by design, so the three goroutines need no additional locking between them — they only ever share the store, and nothing else. Cooldown (`lastFetch`) and, for the two Wikiquote editions, the dead-streak circuit breaker are both plain local variables inside each worker's closure now, not shared maps — since exactly one goroutine ever touches either, there's nothing to synchronize. + +Verified with `go test -race ./...` (clean) plus a full rebuild; not yet measured against a live run for actual throughput gain — see handoff.md. + +**Discovery calls need the same cooldown as page fetches — found live.** `Crawler.runWorker`'s per-source cooldown only updated around an actual page fetch; the `topUp` branch (Wikiquote's MediaWiki API discovery calls) never touched it at all, so consecutive discovery calls (e.g. walking through several already-fully-known categories in a row) were only ~`idlePollInterval` (2s) apart instead of the real per-source delay (6s) — and hit a genuine `429` from Wikiquote's API as a direct result. Fixed by having each top-up function return a `topUpResult{added, calledNetwork, networkFailed}` instead of a bare `bool`: `runWorker` now updates its cooldown after any top-up call that made a real network request (`calledNetwork`), applying the same `stallPenalty` on failure as a page fetch would. Goodreads' top-up never sets `calledNetwork` — it only seeds a tag URL into Postgres/Redis, no request to goodreads.com happens during top-up itself — so it's unaffected. + +### Asynq weighted queues (superseded) + +An Asynq-based weighted-queue design was considered here as the fix for "a stalled fetch blocking other sources," before concurrent per-source goroutines (above) solved the same problem more directly, using only what the codebase already had (no new dependency, no separate worker/queue abstraction layer). Not revisited unless a reason to introduce a real task queue shows up independently (e.g. wanting workers on separate processes/machines, not just separate goroutines in one process). + +### Recursive subcategory discovery + +The curated category list (23 for English, 6 for German) is finite, and confirmed live to actually run out: English Wikiquote's worker went fully idle (0 pending, every curated category's direct members already known) after ~1,572 pages, while Goodreads and German Wikiquote kept progressing fine on their own goroutines — the concurrent-worker fix meant this wasn't a whole-crawler stall, but it was still one source sitting completely idle with real undiscovered content still out there. + +Checked live before building anything: `Category:Writers`, despite being fully mined at the top level, has real, on-topic subcategories the flat list never visited — `Category:Writers by language`, `Category:Gay writers`, `Category:Legal writers`, etc. — and those subcategories have further subcategories of their own (`Writers by language` → `Bengali writers`, `Japanese-language writers`, `Urdu-language writers`, ...). This is a genuinely large additional space, not a marginal one. + +`Crawler.topUpWikiquoteSiteIfEmpty` now walks it: `fetchWikiquoteCategoryMembers` requests `cmtype=page|subcat&cmprop=title|type` — confirmed live that MediaWiki returns a category's own member pages _and_ its direct subcategories together in a single call, tagged by an explicit `type` field, rather than needing two separate requests. Newly found subcategories are queued (breadth-first, so one path doesn't spiral arbitrarily deep before siblings get a turn) and tracked in a visited set scoped to the current top-level branch, guarding against a cyclic or diamond-shaped category graph. Once a category's own pagination (`cmContinue`) is exhausted, the cursor descends into the next queued subcategory; only once a whole branch — the top-level category and everything under it — has nothing left queued does it advance to the next curated category. + +State (`CategoryIndex`, `Current`, `CMContinue`, `Queue`, `Visited`) is packed into the same `wikiquoteDiscoveryCursor` JSON string already persisted per-edition in Redis — no schema change. Each top-up call still makes exactly one HTTP request and is reported back to `runWorker` as one `topUpResult` — deliberately kept 1:1 so the per-source cooldown fix above (each top-up call consumes exactly one cooldown slot) still holds. Combining pages+subcats into one call is a pure efficiency win on top of that, not a relaxation of it: it means a category needs _half_ as many round-trips to fully resolve (one call instead of two), at the same pace as before, rather than resolving at a faster pace. + +Also removed a redundant `idlePollInterval` sleep that used to run _in addition to_ the per-source cooldown after every discovery call — found live to be pure waste (the cooldown wait alone already paces the next attempt correctly), and it was making a chain of already-known categories feel needlessly slow to page through. Deliberately did **not** shorten the per-source discovery cooldown itself below the proven-safe 6s: the live 429 that motivated the original cooldown fix happened at almost exactly a 2-3s pace, so that's the one interval directly confirmed unsafe, not merely untested. + +Verified live (isolated probes against the real API, not assumed): `Category:Writers` → 68 pages + 6 real subcategories in one call; one of those (`Writers by language`, a pure container with 0 direct pages) → 5 further real subcategories (`Bengali writers`, `Japanese-language writers`, etc.) — confirming both that the combined-call approach works end-to-end through the actual Go code, and that the recursion reaches genuine additional content multiple levels deep. + +### Random-page fallback + +Even the recursive category walk can genuinely dry up _for a while_ — confirmed live, not hypothetical: German Wikiquote's queue emptied out (0 pending) while its cursor was deep inside a large `Wissenschaftler` (scientist) subcategory branch, most of whose ~50 sub-subcategories (physicist, chemist, biologist, ...) were already fully known. Every top-level curated category eventually funnels into this same situation once its subtree is mostly mined. + +`Crawler.wikiquoteTopUpFunc` wraps the normal category-walk top-up with a per-edition `emptyStreak` counter (local to the closure set up once per Wikiquote goroutine in `Run()`, not shared state): after `wikiquoteEmptyStreakBeforeRandom` (10) consecutive top-up calls in a row genuinely add nothing new — a network error doesn't count, only a _successful_ call that still found nothing — it falls back to `Crawler.topUpWikiquoteRandomPages`, which calls MediaWiki's `action=query&list=random&rnnamespace=0` for a batch of genuinely random article titles instead. `rnnamespace=0` keeps it to the main article namespace (excludes `Category:`/`Talk:`/`User:`/etc.), but confirmed live it still returns plenty of non-biographical content (topic pages, dates, TV show pages, even the occasional user-babel subpage) since it draws from the _entire_ wiki rather than curated occupation categories — an accepted tradeoff specific to this being a dead-end escape hatch, not the primary mechanism: the existing per-quote filters and the parser itself (no "Quotes" heading → 0 quotes) already handle a low-yield random page harmlessly, and "sometimes fetches a page with nothing on it" beats "sits idle with nothing to try at all." Doesn't touch the category-walk cursor — it's a one-off injection of extra URLs on top, not a replacement; normal category discovery resumes exactly where it left off on the next call. + +Checked whether Goodreads needed the same treatment before building anything for it: live pending counts showed Goodreads sitting at a steady 1 pending (its normal sequential-pagination steady state, not distress) while German Wikiquote was genuinely at 0 — no evidence Goodreads' own tag-cycling fallback is anywhere near its 19-tag limit, so nothing was added there. Goodreads has no MediaWiki-style random-page endpoint to leverage anyway (it's not a wiki); if its tag list ever does prove too small, the natural extension would be more curated tags or the sitemap-based discovery already noted as a future option, not a random-page equivalent. + +### Citation-link discovery (`internal/parser/wikilink.go`) + +Investigated live after a report that discovery within a parsed page "felt broken": on `de.wikiquote.org/wiki/Führer` (a thematic, not biographical, page), the visible on-page links — "Existenz," "Arbeit," "Demokratie" — are topical cross-references embedded in the quote's own body text, not people. Following those blindly would reintroduce exactly the problem an earlier discovery approach at this project already hit once and abandoned (blind `allpages` enumeration pulling in massive non-biographical junk at whole-site scale) — confirmed this is still true today, not just historical. + +The real, precise signal on that same page is different: each quote's _citation_ (e.g. "... - Volker Rühe, Konkret, Heft 2/1998") links the actual person being quoted. `resolveWikiquoteLink` (`internal/parser/wikilink.go`) implements exactly that narrower rule — only links found within a quote's citation are queued, never links found in the quote body itself: + +- Rejects anything not starting with `/wiki/` (an interwiki citation link out to `en.wikipedia.org`, seen live with `class="extiw"`, is already a full absolute URL and gets excluded by this alone). +- Rejects known non-content namespaces (`Category:`/`Kategorie:`, `Special:`/`Spezial:`, `Talk:`/`Diskussion:`, `User:`/`Benutzer:`, `File:`/`Datei:`, `Template:`/`Vorlage:`, `Help:`/`Hilfe:`, `Wikiquote:`, `Portal:`) — both languages' prefixes checked regardless of edition, since a wrong-language namespace name never legitimately collides with a real article title. +- Strips a URL fragment (`/wiki/Aristotle#Politics` → `/wiki/Aristotle`) rather than rejecting it — same page, already worth queueing. + +Wired into both Wikiquote parsers (`wikiquotes.go`, `germanwikiquote.go`), each extracting citation links from their own structurally different citation shape (English: a nested `