Skip to content

Latest commit

 

History

History
192 lines (151 loc) · 7.06 KB

File metadata and controls

192 lines (151 loc) · 7.06 KB

Research Agent — Architecture & Pattern Reference

Overview

Browser (React) ──HTTP──▶ FastAPI (Python) ──▶ Groq LLM
                                │
                         ┌──────┴──────┐
                    DuckDuckGo   local uploads/
                    (web search)  (files & PDFs)

1. Entry point — backend/main.py

The HTTP server, built with FastAPI.

  • Loads GROQ_API_KEY from .env at startup via python-dotenv
  • Registers CORS middleware so the browser (port 5173) can call the backend (port 8000) — browsers block cross-origin requests by default
  • Defines 5 routes:
Route Method Purpose
/health GET Liveness check
/chat POST Main chat endpoint — streams AI response as SSE
/upload POST Saves a file to backend/uploads/
/files GET Lists uploaded files
/files/{name} DELETE Deletes an uploaded file

The /chat route receives:

{ "messages": [{ "role": "user", "content": "What is Python?" }] }

…and returns a streaming SSE response rather than waiting for the full answer.


2. The agent loop — backend/agent.py

The brain of the system. Called by the /chat route.

Step A — Build the conversation

[SYSTEM_PROMPT] + [all prior messages from the client]

The system prompt instructs the LLM to always search the web for factual/current questions and to use uploaded files when referenced.

Step B — Call Groq with tool definitions

The LLM (Llama 3.3 70B via Groq) receives the conversation plus a tool menu it can invoke:

Tool What it does
search_web Queries DuckDuckGo, returns titles/URLs/snippets
fetch_page Downloads and cleans a webpage as plain text
read_file Reads an uploaded .txt / .md / .csv
extract_pdf_text Extracts text from an uploaded PDF
list_uploaded_files Lists what's been uploaded

The LLM decides: answer directly, or call a tool first?

Step C — Stream tokens to the client

stream=True is set so tokens arrive incrementally. Each token is immediately forwarded to the browser as an SSE event:

data: {"type": "text", "content": "Python"}
data: {"type": "text", "content": " is"}

Step D — Tool call loop (agentic behaviour)

If the LLM decides it needs a tool, it stops generating text and outputs a structured tool call instead (e.g. {"name": "search_web", "arguments": {"query": "Python language"}}).

The agent then:

  1. Sends a tool_call event to the browser (shows the ⚙ pill in the UI)
  2. Runs the tool in a background thread via asyncio.to_thread() — prevents the async event loop from freezing on blocking I/O
  3. Appends the tool result to the conversation
  4. Calls Groq again with the enriched conversation
  5. Repeats — up to 8 iterations (safety cap)

This loop is what makes it "agentic": it chains multiple tool calls autonomously before producing a final answer.

Key safety constraints

  • parallel_tool_calls=False — prevents the model generating malformed JSON when calling multiple tools at once
  • max_iterations = 8 — stops infinite loops if the model keeps calling tools
  • Stream wrapped in try/except — errors return a friendly message to the client instead of crashing the server

3. The tools — backend/tools/

web_search.py

  • search_web(query, max_results=5) — uses duckduckgo-search (no API key needed). Retries once on rate limit (flat 2s delay). DDGS(timeout=8) prevents indefinite hangs. Falls back gracefully on any other error.
  • fetch_page(url, max_chars=8000) — downloads URL with requests (6s timeout), strips <script>, <style>, <nav>, <footer> with BeautifulSoup, returns clean plain text capped at 8000 chars.

file_reader.py

  • read_file(filename) — reads from backend/uploads/. Uses os.path.basename() to prevent path traversal attacks (e.g. ../../etc/passwd).
  • list_uploaded_files() — returns names of all files in the uploads directory.

pdf_parser.py

  • extract_pdf_text(filename) — uses pdfplumber to extract text page by page, stopping once 12,000 chars are collected.

4. The frontend — frontend/src/App.tsx

A React + TypeScript app (Vite, port 5173).

On load:

  • Calls GET /files to populate the sidebar file list.

On message send:

  • POSTs { messages: [...] } to /chat with the full conversation history.
  • Reads the response as a raw fetch stream (not EventSource API):
    const reader = res.body.getReader()
  • Manually parses SSE lines (data: {...}) and for each event:
    • type: "text" → appends to the assistant message (live streaming cursor effect)
    • type: "tool_call" → adds a tool pill above the message bubble
    • type: "done" → marks the message finished, hides the blinking cursor

File upload: multipart/form-data POST to /upload, then refreshes the file list.

File delete: DELETE /files/{name}, then refreshes the file list.


5. SSE format

EventSourceResponse (from sse_starlette) wraps each yielded string in the SSE protocol automatically:

data: {"type": "text", "content": "Hi"}\n\n

The generator in agent.py yields raw JSON strings only — no data: prefix, no \n\n. Adding those manually would cause double-wrapping (data: data: {...}).


6. Full data flow — web search query

User types: "What happened at Google I/O 2026?"
          │
          ▼
Frontend POSTs {"messages": [...]} to /chat
          │
          ▼
agent.py builds conversation + tool list, calls Groq (stream=True)
          │
          ▼
Groq: "I need to search" → tool_call: search_web("Google I/O 2026")
          │
          ▼
agent sends {"type": "tool_call"} SSE event → browser shows ⚙ pill
agent runs search_web() in background thread (asyncio.to_thread)
DuckDuckGo returns 5 results
          │
          ▼
agent appends results to conversation, calls Groq again
          │
          ▼
Groq: calls fetch_page() on a promising URL
agent fetches & cleans the page, appends to conversation
          │
          ▼
Groq has enough context → streams final answer token by token
          │
          ▼
Each token: agent yields JSON → SSE to browser
Frontend appends each chunk to message in real time
          │
          ▼
Groq sends done signal → agent yields {"type": "done"}
Frontend hides cursor, marks message complete

7. Dependency map

Package Used for
fastapi HTTP server & routing
uvicorn ASGI server that runs FastAPI
groq Groq API client (LLM + tool calling)
sse-starlette SSE streaming response wrapper
duckduckgo-search Free web search (no API key)
beautifulsoup4 HTML parsing / page cleaning
requests HTTP client for fetching pages
pdfplumber PDF text extraction
python-dotenv Loads .env for API keys
python-multipart Multipart file upload parsing
react + vite Frontend UI
react-markdown + remark-gfm Renders markdown in chat bubbles