EtymEx is a beautiful, interactive Etymology Explorer that helps you discover the origins and roots of words. Perfect for vocabulary enthusiasts who want to understand words deeply through their etymological roots.
Try it out at etymex.com
- Public Mode: No client-side API keys needed — uses a server-side LLM with rate limiting, caching, and budget controls
- Grounded Etymology: Source-backed confidence scoring (high/medium/low) for each ancestral stage
- Romance-language Beta: Explicit Italian, Spanish, French, and neutral Portuguese lookup with paired local/English prose and no spelling-based language inference
- Etymology Lookup: Search any word to discover its linguistic origins, root morphemes, and historical evolution
- Part of Speech Tags: See grammatical categories (noun/verb/adjective) with alternate pronunciations for words like "record"
- Memorable Lore: Each word comes with a 4-6 sentence narrative that makes the etymology stick
- Related Words: Discover words that share the same roots
- Word Suggestions: Explore synonyms, antonyms, homophones, easily-confused words, and see-also links with color-coded clickable chips
- Modern Usage: Slang context gated by source significance from Urban Dictionary and supplemental Incel Wiki extracts
- Pronunciation Audio: Listen to word pronunciations powered by ElevenLabs
- Search History: Track your vocabulary exploration with a persistent sidebar
- Surprise Me: Discover random words to expand your vocabulary
- Structured Outputs: Guaranteed valid JSON via OpenRouter
json_schemastrict mode - Streaming UI: Optional
?stream=trueserver-sent events for source progress, per-section synthesis events, cached hits, and early error responses - Smart Caching: Redis-backed caching reduces costs and improves speed (30d etymology, 1yr audio)
- Shareable Word Pages:
/word/{word}is the primary search URL — cached words are server-rendered straight from the cache (crawlers never trigger LLM spend), uncached words host the live streaming trace with progressive section-by-section rendering - Rate Limiting: Per-IP protection via Upstash Redis with automatic budget enforcement
- Node.js 18+
- For self-hosted deployment:
- OpenRouter API key (required)
- Upstash Redis (optional, for rate limiting and caching)
- ElevenLabs (optional, for pronunciation audio)
# Clone the repository
git clone https://github.com/thepushkarp/etymology-explorer.git
cd etymology-explorer
# Install dependencies
bun install
# Set up environment variables (see Environment Configuration section)
cp .env.example .env.local
# Start the development server
bun devOpen http://localhost:3000 in your browser.
The app runs in public mode using a server-side OpenRouter API key and the
openai/gpt-6-luna model on OpenRouter's Responses API. All searches are
rate-limited and cost-budgeted with a monthly spend cap. Set the
OPENROUTER_API_KEY environment variable to enable it.
For self-hosted deployments, create a .env.local file:
# Required for public mode
OPENROUTER_API_KEY=your_openrouter_key_here
ADMIN_SECRET=your_admin_secret_here
# Optional: Upstash Redis (rate limiting + caching)
ETYMOLOGY_KV_REST_API_URL=your_upstash_url_here
ETYMOLOGY_KV_REST_API_TOKEN=your_upstash_token_here
# Optional: ElevenLabs (pronunciation audio)
ELEVENLABS_API_KEY=your_elevenlabs_key_here
ELEVENLABS_VOICE_ID=your_voice_id_here
# Optional: Feature flags
PUBLIC_SEARCH_ENABLED=true
PRONUNCIATION_ENABLED=true
FORCE_CACHE_ONLY=false
RATE_LIMIT_ENABLED=trueELEVENLABS_VOICE_ID must be a voice available to your account in My Voices.
Free-tier accounts cannot use Voice Library/community voices through the API.
Pronunciation requests pass the selected ISO 639-1 language_code to eleven_v3 for
language selection and text normalization. Accent still depends on the configured voice;
pt does not select a Brazilian or European accent.
See .env.example for full documentation.
For local load testing, set RATE_LIMIT_ENABLED=false in .env.local and restart bun dev.
- Framework: Next.js 16.1 with App Router
- UI: React 19.2 + Tailwind CSS v4
- LLM: OpenRouter Responses API
using
openai/gpt-6-lunawith structured outputs - Validation: Zod 4.x for schema validation
- Caching/Rate Limiting: @upstash/redis + @upstash/ratelimit
- Analytics: @vercel/analytics
- Data Sources:
- Etymonline - Historical etymology
- Wiktionary - Definitions and linguistic data
- Native Italian, Spanish, French, and Portuguese Wiktionary editions - selected-language etymology sections
- Wikidata Lexemes - language-tagged lemmas, forms, senses, and lexical claims
- FreeDictionaryAPI - multilingual senses, forms, and IPA (not etymology)
- Dicionário Aberto - older Portuguese dictionary evidence
- Free Dictionary API - Definitions, pronunciation hints, and origin data
- Wikipedia - Encyclopedic context
- Urban Dictionary - Modern slang (quality-filtered)
- Incel Wiki - Supplemental community slang context
- Audio: ElevenLabs - Text-to-speech pronunciation
- Typography: Libre Baskerville (display serif) + Alegreya Sans (body), self-hosted woff2 subsets
etymology-explorer/
├── app/
│ ├── api/
│ │ ├── admin/stats/ # Budget/usage statistics (admin-only)
│ │ ├── etymology/ # Main etymology synthesis endpoint (GET)
│ │ ├── pronunciation/ # TTS audio endpoint (ElevenLabs)
│ │ ├── random-word/ # Random word selection
│ │ └── suggestions/ # Autocomplete + typo suggestions
│ ├── faq/ # FAQ page with structured data
│ ├── learn/ # Educational content pages
│ │ └── what-is-etymology/
│ ├── og/ # Dynamic OG image generation (brand + per-word cards)
│ ├── word/[word]/ # Primary word pages: cached → SSR, uncached → live trace UI
│ ├── sitemap.ts # Dynamic sitemap (static pages + cached /word/ entries)
│ ├── robots.ts # Robots.txt configuration
│ ├── layout.tsx # Root layout with fonts
│ └── page.tsx # Landing/search page (/?q= redirects to /word/{word})
├── proxy.ts # Rate limiting + CSP headers
├── components/
│ ├── AncestryTree.tsx # Visual etymology graph
│ ├── ErrorState.tsx # Error display with retry
│ ├── EtymologyCard.tsx # Main result display
│ ├── FaqAccordion.tsx # Accessible FAQ accordion
│ ├── FaqSchema.tsx # FAQPage JSON-LD schema
│ ├── HistorySidebar.tsx # Search history panel
│ ├── JsonLd.tsx # WebApplication schema
│ ├── PronunciationButton.tsx # Audio playback
│ ├── RelatedWordsList.tsx # Related words chips
│ ├── RootChip.tsx # Expandable root morpheme
│ ├── SearchBar.tsx # Word input
│ └── SurpriseButton.tsx # Random word button
├── lib/
│ ├── research.ts # Agentic multi-source research pipeline
│ ├── llm.ts # OpenRouter-backed LLM synthesis with structured outputs
│ ├── etymologyParser.ts # CPU-only source text parser
│ ├── etymologyEnricher.ts # Post-LLM confidence enricher
│ ├── etymonline.ts # Etymonline HTML scraper
│ ├── wiktionary.ts # Wiktionary MediaWiki API client
│ ├── freeDictionary.ts # Free Dictionary API client
│ ├── wikipedia.ts # Wikipedia REST API client
│ ├── urbanDictionary.ts # Urban Dictionary API with quality scoring/filtering
│ ├── incelsWiki.ts # Incel Wiki MediaWiki API client (supplemental)
│ ├── elevenlabs.ts # ElevenLabs TTS for pronunciation audio
│ ├── spellcheck.ts # Typo detection and suggestions
│ ├── prompts.ts # System prompts and schemas
│ ├── types.ts # TypeScript interfaces
│ ├── config.ts # Centralized configuration
│ ├── env.ts # Environment variable validation
│ ├── costGuard.ts # Budget enforcement
│ ├── singleflight.ts # Request deduplication
│ ├── redis.ts # Redis client factory
│ ├── cache.ts # Caching layer
│ ├── errorUtils.ts # Secret redaction
│ ├── fetchUtils.ts # Timeout wrapper
│ ├── validation.ts # Input validation
│ ├── wordlist.ts # GRE word utilities
│ ├── hooks/ # React hooks (localStorage, history, search)
│ │ └── useStreamingEtymology.ts # SSE transport over the pure stream reducer
│ └── schemas/
│ ├── etymology.ts # Zod schema for cache validation
│ └── llm-schema.ts # Strict-mode JSON Schema for LLM structured outputs
│ # (mechanically sync-checked against the Zod schema)
├── data/
│ ├── faq.ts # FAQ content with FaqItem interface
│ └── gre-words.json # Vocabulary word list
└── .env.example # Environment variable template
| Endpoint | Method | Description | Auth Required |
|---|---|---|---|
/api/etymology |
GET | Synthesize etymology (?word=X, optional language and stream) |
No (rate-limited) |
/api/pronunciation |
GET | Get pronunciation audio (?word=X&language=it) |
No |
/api/suggestions |
GET | Get language-aware autocomplete suggestions | No |
/api/random-word |
GET | Get a random word | No |
/api/ngram |
GET | Get language-aware Google Books usage data (?word=X) |
No |
/api/health |
GET | Liveness check | No |
/api/admin/stats |
GET | Get budget/usage statistics and counters | Admin secret |
- Request Deduplication: Singleflight owner-token locks (90s TTL, heartbeat-extended) ensure one pipeline run per word; waiters poll the cache for the holder's result, and streaming waiters can take over if the holder crashes
- Rate Limiting: Per-IP rate limiting (20 req/min + 200 req/day) via Upstash Redis
- Cache Check: Redis cache lookup with versioned keys (
etymology:v2.2:), schema validation on read, and negative cache (30m) for known no-source/invalid words. Without Redis, uncached searches return 503 (fail closed) - Grounded Etymology Pipeline:
- Parser (CPU-only): Extracts "from X, from Y" chains from raw source text
- Agentic Research: Multi-phase research pipeline (aborts mid-flight if the client disconnects):
- Phase 1: Fire all 6 source fetches at once (Etymonline, Wiktionary, Free Dictionary, Wikipedia, Urban Dictionary, Incel Wiki); Etymonline + Wiktionary gate root expansion, while a Free Dictionary or Wikipedia hit can still admit synthesis when both primary sources miss. Raw Etymonline/Wiktionary pages are served from a 7-day Redis source cache when available
- Phase 2: Root morphemes extracted on-CPU from derivation formulas ("From X + Y", "equivalent to X + Y") and parsed-chain affixes (e.g., "telephone" → ["tele", "phone"]); a quick LLM call (15s timeout, truncated input) runs only when the CPU pass finds nothing
- Phase 3: One parallel wave fetches root pages (up to 4 roots, etymonline + wiktionary) and main-word related-term pages (etymonline only) within the 16-fetch budget
- LLM Synthesis: Aggregated research context sent to LLM with structured output schema
- Enricher (CPU): Post-processes LLM output, assigns confidence scores (high/medium/low) based on source evidence match
- Guaranteed JSON: Using constrained decoding, the LLM produces valid JSON matching the exact schema
- Budget Enforcement: Cost guard tracks monthly spend (OpenRouter-reported cost, with a conservative fallback-chain pricing ceiling when cost is omitted) against a $10/month cap and switches from normal to cache_only mode at 90% of budget; both the root-extraction and synthesis LLM calls are counted
- Rich Display: Etymology rendered with expandable roots, ancestry graph with confidence badges, POS tags, modern usage, related words, and source attribution (supplemental sources are only surfaced when significance checks pass)
Production synthesis and LLM root extraction use openai/gpt-6-luna through
OpenRouter's Responses API. Both paths explicitly set low reasoning and exclude
the reasoning trace from the response. Reasoning remains enabled internally but
is never returned to the browser.
OpenRouter tries these model fallbacks in order when Luna is unavailable, rate-limited, or rejected by provider routing:
openai/gpt-5.4-mini— closest proven fallback; passed the project's 15-word strict-schema bakeoff.google/gemini-3.5-flash— cross-provider redundancy; also passed the strict-schema bakeoff.
Explicit --model benchmark overrides do not inherit the production fallback
chain. openai/gpt-5.6-terra is a compatible future candidate, but remains
disabled until it completes the same etymology-quality and latency bakeoff.
┌──────────────────────────────┐
│ BROWSER UI │
│ SearchBar / History / Result │
└──────────────┬───────────────┘
│ GET /api/etymology?word=X[&stream=true]
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ Middleware (Node) + Route Entry │
│ proxy.ts: rate limit + CSP → app/api/etymology/route.ts: validate input │
└──────────────────────────────┬───────────────────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ Control Plane │
│ cache hit? → return cached result │
│ cost guard → normal | cache_only (at 90% of monthly budget) │
│ singleflight lock → dedupe concurrent lookups (Redis down → 503 uncached) │
└──────────────────────────────┬───────────────────────────────────────────────┘
│ cache miss
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ Grounded Etymology Pipeline (client disconnect aborts in-flight work) │
│ 1) Fire 6 source fetches at once (etymonline/wiktionary via 7d source cache) │
│ 2) Parse "from X, from Y" chains once etymonline+wiktionary land (CPU-only) │
│ 3) CPU root extraction from derivation formulas; LLM fallback if none found │
│ 4) ONE parallel wave: root pages + related-term pages (16-fetch budget) │
│ 5) OpenRouter synthesis, json_schema strict (sections stream when stream=true)│
│ 6) Enrich ancestry graph + confidence/evidence │
│ 7) Cache result in Redis │
└──────────────────────────────┬───────────────────────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ Response Paths │
│ stream=true → SSE events: source_* → parsing_complete → synthesis_started │
│ → synthesis_section (per top-level field, render order) │
│ → enrichment_done → result / error │
│ default → JSON response with final EtymologyResult │
└──────────────────────────────────────────────────────────────────────────────┘
# Run development server
bun dev
# Lint code
bun run lint
# Format code
bun run format
# Build for production
bun run buildDeploy easily on Vercel:
Important: Set environment variables in your Vercel project settings (see Environment Configuration section above).
MIT
- Etymology data from Etymonline and Wiktionary
- Definitions and pronunciation hints from Free Dictionary API
- Encyclopedic context from Wikipedia
- Modern slang definitions from Urban Dictionary
- Supplemental community slang context from Incel Wiki
- Pronunciation audio from ElevenLabs
- Powered by OpenRouter
openai/gpt-6-luna - Rate limiting and caching by Upstash