A retrieval-augmented chat assistant over tech documentation and source repositories. Documents and code are chunked, embedded locally, and stored in a persistent ChromaDB collection; a Chainlit web UI routes questions to one of three Gemini-backed agents, each with its own prompt, model temperature, and retrieval strategy.
src/scripts/scrape_doc.pyconverts ReadTheDocs HTML pages to markdown indata/documents/.src/scripts/ingest_repositories.pyclones the repositories listed indata/repos.txtintodata/repositories/, selects documentation, code, and configuration files, chunks them (WDL tasks, Java/Python definitions, and markdown headers are chunked structurally rather than by fixed length), and writes them to the vector store.src/scripts/setup_chromadb.pyingests the markdown underdata/documents/and runs a sample query to confirm the collection works.src/rag_chatbot/main.pyis the Chainlit entry point. Each user message goes to the currently selected agent, which searches the collection, re-ranks the hits with its own strategy, and asks Gemini to answer with that context.
Embeddings are computed locally with sentence-transformers/all-MiniLM-L6-v2
on CPU, so ingestion and search need no API key. Only answer generation calls a
hosted model.
| Type | Name | Focus | Retrieval strategy |
|---|---|---|---|
qa_agent |
Documentation Q&A Agent | Documentation and general questions | balanced: weights hits by file category (docs 0.7, code 0.3) |
code_assistant |
Code Assistant Agent | Code analysis, debugging, development | code_focused: mostly code chunks, keeps a small conversation state |
workflow_agent |
Workflow Agent | Bioinformatics pipelines and WDL | workflow_focused: prioritizes WDL, workflow, and pipeline files |
Chat commands: /agents lists the agents, /switch <agent_type> changes the
active one, /current shows its model and temperature, /help prints the list.
- Python 3.11 or 3.12
- Poetry
- git (repository ingestion shells out through GitPython)
- A Google AI API key for the Gemini models
- Docker, only if you want the Postgres container for Chainlit chat persistence
poetry install
poetry shellCreate .env.dev in the project root:
google_api_key=<your Gemini API key>
openai_api_key=<unused by the agents, but Settings requires it>
default_agent=qa_agent
debug=false
host=0.0.0.0
port=8000
qa_agent_model=gemini-2.5-flash
qa_agent_temperature=0.7
code_assistant_model=gemini-2.5-flash
code_assistant_temperature=0.2
workflow_agent_model=gemini-2.5-flash
workflow_agent_temperature=0.3
chromadb_path=./chroma_db
max_search_results=10
enable_code_execution=false
enable_state_management=true
conversation_memory_size=10
log_level=INFOEvery field in src/rag_chatbot/config.py can be set here; the ones above are
the ones that matter in practice. openai_api_key has no default, so startup
fails without it even though answer generation goes to Gemini.
Scrape the documentation pages (edit the URL list at the bottom of the script first, including the hardcoded output directory):
python src/scripts/scrape_doc.pyIndex that markdown:
python src/scripts/setup_chromadb.pyClone and index the repositories in data/repos.txt:
python src/scripts/ingest_repositories.pyBoth ingestion paths are idempotent. Markdown chunks are keyed by file and chunk
index; repository chunks are keyed by an MD5 of repo, file, chunk index, and
content, so re-running only adds what changed. Re-running the repository script
also does a git pull on each existing clone first.
chainlit run src/rag_chatbot/main.py -wThe UI is served on the configured host and port (default
http://localhost:8000). -w reloads on source changes.
For persisted chat history, start Postgres and create the Chainlit tables:
docker compose -f src/scripts/docker-compose.yaml up -d
python src/scripts/create_db_schema.pyNote that create_db_schema.py drops the Element, Step, Thread, and
User tables before recreating them, and it hardcodes the local development
connection string.
src/rag_chatbot/
main.py Chainlit entry point and message loop
config.py pydantic-settings Settings, read from .env.dev
agent/
agent_config.py AgentType enum and per-agent configuration
agent_factory.py AgentFactory (creates/caches agents) and AgentManager (commands)
agents.py BaseAgent plus the three concrete agents
prompt_templates.py Prompt text per agent
services/
vector_service.py ChromaDB collection, chunking, embedding, search
llm_service.py empty placeholder
src/scripts/
scrape_doc.py ReadTheDocs to markdown
ingest_repositories.py Repository cloning, file selection, structural chunking
setup_chromadb.py Index data/documents and smoke-test a search
create_db_schema.py Chainlit persistence tables
docker-compose.yaml Postgres for the above
data/
documents/ Scraped and hand-added source documents (gitignored)
repositories/ Repository clones (gitignored)
repos.txt Repository URLs to ingest
chroma_db/ Persistent vector store (gitignored)
scrape_doc.pyimportsbeautifulsoup4, which is not declared inpyproject.toml; install it before scraping.- The collection name is hardcoded to
documentsinVectorService, so thevector_collection_namesetting has no effect. - The PDFs under
data/documents/are not ingested; only*.mdis indexed. services/llm_service.pyis empty. The LLM client lives inBaseAgent.