An AI-powered assistant for exploring U.S. National Park information using authoritative National Park Service sources. The knowledge layer uses retrieval-augmented generation to ground answers in indexed NPS content and return traceable citations.
For the end-to-end execution path, component responsibilities, algorithms, provider switching, and rebuild rules, see PROJECT_FLOW.md.
This repository implements Phase 1: stable park knowledge. It deliberately does not claim to verify live closures, weather, road status, reservations, or permit availability.
- Five parks: Yosemite, Yellowstone, Zion, Grand Canyon, and Denali
- Factual, trail, activity, comparison, safety, logistics, and simple planning questions
- Offline NPS ingestion with raw HTML snapshots and failure reports
- Heading-aware chunks with deterministic IDs and rich source metadata
- Semantic retrieval with lightweight park detection
- Grounded generation, validated citations, and explicit no-answer behavior
- Debug evidence, request IDs, latency logging, and Recall@K evaluation
OFFLINE: parks.yaml → fetch → raw HTML → parse → chunk → batch embed → vector store
ONLINE: question → park detection → embed → retrieve → context budget → LLM → validate citations
The application owns each step. Provider boundaries are explicit (EmbeddingClient, VectorStore, and LLMClient) rather than hidden inside a framework. The default stack uses Ollama with Qwen3 models, so embeddings and grounded answer generation run locally without API credentials.
Requires Python 3.12+.
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev,ui]'
cp .env.example .env
python scripts/ingest.py
uvicorn national_parks.app.main:app --reloadIn another terminal:
python scripts/test_query.py "What hiking options are available in Yosemite?" --evidence
streamlit run ui.pyAPI documentation is available at http://localhost:8000/docs.
Docker users can run cp .env.example .env && docker compose up --build, then execute ingestion in the container with docker compose exec api python scripts/ingest.py.
Install Ollama, start it, and download the two models:
ollama pull qwen3-embedding:0.6b
ollama pull qwen3:4bThe embedding model is about 639 MB and the chat model is about 2.5 GB. Configure:
LLM_PROVIDER=ollama
LLM_MODEL=qwen3:4b
EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=qwen3-embedding:0.6b
OLLAMA_URL=http://localhost:11434Changing embedding providers requires rebuilding the index (python scripts/ingest.py --force). Ollama calls use its local /api/embed and /api/chat endpoints; no document or question is sent to OpenAI.
OpenAI adapters remain available as optional alternatives, but no API key is required by the default configuration.
# All configured sources, or one park
python scripts/ingest.py
python scripts/ingest.py --park yose --force
# Remove all chunks, or one park
python scripts/clear_index.py
python scripts/clear_index.py --park zion
# Retrieval evaluation
python scripts/evaluate.py
# Tests
pytestIngestion writes raw pages under data/raw/<park-code>/ and a machine-readable summary to data/processed/ingestion_report.json. Evaluation writes timestamped experiment results under data/evaluation/results/.
POST /api/v1/chat
{
"query": "What are some easy hikes in Yosemite?",
"top_k": 8,
"include_evidence": true
}POST /api/v1/debug/retrieve returns ranked chunks and similarity scores. GET /api/v1/parks lists parks actually present in the index. GET /health reports dependency configuration.
The prompt requires claims to use supplied evidence, citations are accepted only when they map to retrieved metadata, and URLs never come from model output. Questions containing live-status language are intercepted before retrieval and receive a clear freshness warning. Each question is independent; there is no chat memory or user profile.
Qwen3 Embedding provides the semantic baseline through Ollama. Retrieval blends semantic similarity with lexical title/body matching, applies a relevance floor, and limits repeated chunks from one page. Qwen3 4B synthesizes the selected evidence into the cited answer. The JSON vector store is designed for a modest Phase 1 corpus; the VectorStore boundary allows a dedicated store to replace it as scale requires. PDF extraction, live NPS data, learned reranking, structured trail constraints, and itinerary optimization belong to later phases.
The initial dataset contains 50 questions across five parks and ten categories, including unanswerable cases. Retrieval questions identify acceptable source-title fragments. The evaluator reports Recall@1, @3, @5, and @10 and records the chunk and embedding settings with each run. Unanswerable questions remain in the dataset for generation evaluation but are excluded from Recall@K denominators.
- Phase 2: metadata constraints, hybrid BM25/vector retrieval, reranking, and query rewriting
- Phase 3: freshness classification, incremental indexing, and live conditions with source precedence
- Phase 4: preferences, query decomposition, geographic grouping, and validated itinerary planning