Browser (React) ──HTTP──▶ FastAPI (Python) ──▶ Groq LLM
│
┌──────┴──────┐
DuckDuckGo local uploads/
(web search) (files & PDFs)
The HTTP server, built with FastAPI.
- Loads
GROQ_API_KEYfrom.envat startup viapython-dotenv - Registers CORS middleware so the browser (port 5173) can call the backend (port 8000) — browsers block cross-origin requests by default
- Defines 5 routes:
| Route | Method | Purpose |
|---|---|---|
/health |
GET | Liveness check |
/chat |
POST | Main chat endpoint — streams AI response as SSE |
/upload |
POST | Saves a file to backend/uploads/ |
/files |
GET | Lists uploaded files |
/files/{name} |
DELETE | Deletes an uploaded file |
The /chat route receives:
{ "messages": [{ "role": "user", "content": "What is Python?" }] }…and returns a streaming SSE response rather than waiting for the full answer.
The brain of the system. Called by the /chat route.
[SYSTEM_PROMPT] + [all prior messages from the client]
The system prompt instructs the LLM to always search the web for factual/current questions and to use uploaded files when referenced.
The LLM (Llama 3.3 70B via Groq) receives the conversation plus a tool menu it can invoke:
| Tool | What it does |
|---|---|
search_web |
Queries DuckDuckGo, returns titles/URLs/snippets |
fetch_page |
Downloads and cleans a webpage as plain text |
read_file |
Reads an uploaded .txt / .md / .csv |
extract_pdf_text |
Extracts text from an uploaded PDF |
list_uploaded_files |
Lists what's been uploaded |
The LLM decides: answer directly, or call a tool first?
stream=True is set so tokens arrive incrementally. Each token is immediately forwarded to the browser as an SSE event:
data: {"type": "text", "content": "Python"}
data: {"type": "text", "content": " is"}
If the LLM decides it needs a tool, it stops generating text and outputs a structured tool call instead (e.g. {"name": "search_web", "arguments": {"query": "Python language"}}).
The agent then:
- Sends a
tool_callevent to the browser (shows the ⚙ pill in the UI) - Runs the tool in a background thread via
asyncio.to_thread()— prevents the async event loop from freezing on blocking I/O - Appends the tool result to the conversation
- Calls Groq again with the enriched conversation
- Repeats — up to 8 iterations (safety cap)
This loop is what makes it "agentic": it chains multiple tool calls autonomously before producing a final answer.
parallel_tool_calls=False— prevents the model generating malformed JSON when calling multiple tools at oncemax_iterations = 8— stops infinite loops if the model keeps calling tools- Stream wrapped in
try/except— errors return a friendly message to the client instead of crashing the server
search_web(query, max_results=5)— usesduckduckgo-search(no API key needed). Retries once on rate limit (flat 2s delay).DDGS(timeout=8)prevents indefinite hangs. Falls back gracefully on any other error.fetch_page(url, max_chars=8000)— downloads URL withrequests(6s timeout), strips<script>,<style>,<nav>,<footer>with BeautifulSoup, returns clean plain text capped at 8000 chars.
read_file(filename)— reads frombackend/uploads/. Usesos.path.basename()to prevent path traversal attacks (e.g.../../etc/passwd).list_uploaded_files()— returns names of all files in the uploads directory.
extract_pdf_text(filename)— usespdfplumberto extract text page by page, stopping once 12,000 chars are collected.
A React + TypeScript app (Vite, port 5173).
On load:
- Calls
GET /filesto populate the sidebar file list.
On message send:
- POSTs
{ messages: [...] }to/chatwith the full conversation history. - Reads the response as a raw fetch stream (not EventSource API):
const reader = res.body.getReader()
- Manually parses SSE lines (
data: {...}) and for each event:type: "text"→ appends to the assistant message (live streaming cursor effect)type: "tool_call"→ adds a tool pill above the message bubbletype: "done"→ marks the message finished, hides the blinking cursor
File upload: multipart/form-data POST to /upload, then refreshes the file list.
File delete: DELETE /files/{name}, then refreshes the file list.
EventSourceResponse (from sse_starlette) wraps each yielded string in the SSE protocol automatically:
data: {"type": "text", "content": "Hi"}\n\n
The generator in agent.py yields raw JSON strings only — no data: prefix, no \n\n. Adding those manually would cause double-wrapping (data: data: {...}).
User types: "What happened at Google I/O 2026?"
│
▼
Frontend POSTs {"messages": [...]} to /chat
│
▼
agent.py builds conversation + tool list, calls Groq (stream=True)
│
▼
Groq: "I need to search" → tool_call: search_web("Google I/O 2026")
│
▼
agent sends {"type": "tool_call"} SSE event → browser shows ⚙ pill
agent runs search_web() in background thread (asyncio.to_thread)
DuckDuckGo returns 5 results
│
▼
agent appends results to conversation, calls Groq again
│
▼
Groq: calls fetch_page() on a promising URL
agent fetches & cleans the page, appends to conversation
│
▼
Groq has enough context → streams final answer token by token
│
▼
Each token: agent yields JSON → SSE to browser
Frontend appends each chunk to message in real time
│
▼
Groq sends done signal → agent yields {"type": "done"}
Frontend hides cursor, marks message complete
| Package | Used for |
|---|---|
fastapi |
HTTP server & routing |
uvicorn |
ASGI server that runs FastAPI |
groq |
Groq API client (LLM + tool calling) |
sse-starlette |
SSE streaming response wrapper |
duckduckgo-search |
Free web search (no API key) |
beautifulsoup4 |
HTML parsing / page cleaning |
requests |
HTTP client for fetching pages |
pdfplumber |
PDF text extraction |
python-dotenv |
Loads .env for API keys |
python-multipart |
Multipart file upload parsing |
react + vite |
Frontend UI |
react-markdown + remark-gfm |
Renders markdown in chat bubbles |