Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

11 Commits

Folders and files

Repository files navigation

rag-chatbot

A retrieval-augmented chat assistant over tech documentation and source repositories. Documents and code are chunked, embedded locally, and stored in a persistent ChromaDB collection; a Chainlit web UI routes questions to one of three Gemini-backed agents, each with its own prompt, model temperature, and retrieval strategy.

How it works

  1. src/scripts/scrape_doc.py converts ReadTheDocs HTML pages to markdown in data/documents/.
  2. src/scripts/ingest_repositories.py clones the repositories listed in data/repos.txt into data/repositories/, selects documentation, code, and configuration files, chunks them (WDL tasks, Java/Python definitions, and markdown headers are chunked structurally rather than by fixed length), and writes them to the vector store.
  3. src/scripts/setup_chromadb.py ingests the markdown under data/documents/ and runs a sample query to confirm the collection works.
  4. src/rag_chatbot/main.py is the Chainlit entry point. Each user message goes to the currently selected agent, which searches the collection, re-ranks the hits with its own strategy, and asks Gemini to answer with that context.

Embeddings are computed locally with sentence-transformers/all-MiniLM-L6-v2 on CPU, so ingestion and search need no API key. Only answer generation calls a hosted model.

Agents

Type Name Focus Retrieval strategy
qa_agent Documentation Q&A Agent Documentation and general questions balanced: weights hits by file category (docs 0.7, code 0.3)
code_assistant Code Assistant Agent Code analysis, debugging, development code_focused: mostly code chunks, keeps a small conversation state
workflow_agent Workflow Agent Bioinformatics pipelines and WDL workflow_focused: prioritizes WDL, workflow, and pipeline files

Chat commands: /agents lists the agents, /switch <agent_type> changes the active one, /current shows its model and temperature, /help prints the list.

Requirements

  • Python 3.11 or 3.12
  • Poetry
  • git (repository ingestion shells out through GitPython)
  • A Google AI API key for the Gemini models
  • Docker, only if you want the Postgres container for Chainlit chat persistence

Setup

poetry install
poetry shell

Create .env.dev in the project root:

google_api_key=<your Gemini API key>
openai_api_key=<unused by the agents, but Settings requires it>

default_agent=qa_agent
debug=false
host=0.0.0.0
port=8000

qa_agent_model=gemini-2.5-flash
qa_agent_temperature=0.7
code_assistant_model=gemini-2.5-flash
code_assistant_temperature=0.2
workflow_agent_model=gemini-2.5-flash
workflow_agent_temperature=0.3

chromadb_path=./chroma_db
max_search_results=10
enable_code_execution=false
enable_state_management=true
conversation_memory_size=10
log_level=INFO

Every field in src/rag_chatbot/config.py can be set here; the ones above are the ones that matter in practice. openai_api_key has no default, so startup fails without it even though answer generation goes to Gemini.

Building the knowledge base

Scrape the documentation pages (edit the URL list at the bottom of the script first, including the hardcoded output directory):

python src/scripts/scrape_doc.py

Index that markdown:

python src/scripts/setup_chromadb.py

Clone and index the repositories in data/repos.txt:

python src/scripts/ingest_repositories.py

Both ingestion paths are idempotent. Markdown chunks are keyed by file and chunk index; repository chunks are keyed by an MD5 of repo, file, chunk index, and content, so re-running only adds what changed. Re-running the repository script also does a git pull on each existing clone first.

Running the chat app

chainlit run src/rag_chatbot/main.py -w

The UI is served on the configured host and port (default http://localhost:8000). -w reloads on source changes.

For persisted chat history, start Postgres and create the Chainlit tables:

docker compose -f src/scripts/docker-compose.yaml up -d
python src/scripts/create_db_schema.py

Note that create_db_schema.py drops the Element, Step, Thread, and User tables before recreating them, and it hardcodes the local development connection string.

Layout

src/rag_chatbot/
  main.py                    Chainlit entry point and message loop
  config.py                  pydantic-settings Settings, read from .env.dev
  agent/
    agent_config.py          AgentType enum and per-agent configuration
    agent_factory.py         AgentFactory (creates/caches agents) and AgentManager (commands)
    agents.py                BaseAgent plus the three concrete agents
    prompt_templates.py      Prompt text per agent
  services/
    vector_service.py        ChromaDB collection, chunking, embedding, search
    llm_service.py           empty placeholder
src/scripts/
  scrape_doc.py              ReadTheDocs to markdown
  ingest_repositories.py     Repository cloning, file selection, structural chunking
  setup_chromadb.py          Index data/documents and smoke-test a search
  create_db_schema.py        Chainlit persistence tables
  docker-compose.yaml        Postgres for the above
data/
  documents/                 Scraped and hand-added source documents (gitignored)
  repositories/              Repository clones (gitignored)
  repos.txt                  Repository URLs to ingest
chroma_db/                   Persistent vector store (gitignored)

Known gaps

  • scrape_doc.py imports beautifulsoup4, which is not declared in pyproject.toml; install it before scraping.
  • The collection name is hardcoded to documents in VectorService, so the vector_collection_name setting has no effect.
  • The PDFs under data/documents/ are not ingested; only *.md is indexed.
  • services/llm_service.py is empty. The LLM client lives in BaseAgent.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages