A full-stack AI research agent built with FastAPI, Groq (Llama 3.3-70b), and React + Vite.
- Web search — DuckDuckGo search + full page fetching via the agent's tool loop
- Document analysis — Upload PDFs, text, CSV, and Markdown; the agent reads them directly
- Persistent conversations — Research sessions are saved per user and can be resumed
- Per-user file isolation — Uploaded files are stored in private per-user directories
- Auth — Email/password registration & login (bcrypt + JWT); Google OAuth sign-in (creates account on first use, returns JWT); auto-login after signup
- Streaming — Responses stream token-by-token via Server-Sent Events
ResearchAgent/
├── backend/
│ ├── main.py # FastAPI app, all HTTP routes
│ ├── agent.py # Groq tool-calling loop (streaming)
│ ├── requirements.txt
│ ├── research_agent.db # SQLite database (auto-created on first run)
│ ├── uploads/
│ │ └── {user_id}/ # Per-user uploaded files
│ ├── api/
│ │ ├── middleware/
│ │ │ └── auth.py # JWT dependency (get_current_user)
│ │ ├── model/
│ │ │ └── mydb.py # SQLite connection + init_db()
│ │ ├── repository/
│ │ │ ├── userRepository.py
│ │ │ ├── conversationRepository.py
│ │ │ └── uploadRepository.py
│ │ ├── services/
│ │ │ └── useService.py # login_user, register_user
│ │ └── controllers/
│ │ ├── authController.py
│ │ └── conversationController.py
│ └── tools/
│ ├── web_search.py
│ ├── file_reader.py
│ └── pdf_parser.py
└── frontend/
└── src/
├── App.tsx # Main UI — chat, file sidebar, conversation list
├── apiServices/api.ts # Typed fetch wrapper
├── contexts/AuthContext.tsx
└── modals/
├── Login.tsx
└── Signup.tsx
users (id, email, password_hash)
conversations (id, user_id, title, created_at)
messages (id, conversation_id, role, content, created_at)
uploads (id, user_id, filename, stored_path, created_at)All tables are created automatically on server startup via init_db().
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txtCreate a .env file in backend/:
GROQ_API_KEY=your_groq_api_key
JWT_SECRET_KEY=a_long_random_secret
GOOGLE_CLIENT_ID=optional_for_google_oauthStart the server:
uvicorn main:app --reloadAPI runs at http://localhost:8000. Interactive docs at http://localhost:8000/docs.
cd frontend
npm install
npm run devApp runs at http://localhost:5173.
Optionally create frontend/.env:
VITE_API_URL=http://localhost:8000
VITE_GOOGLE_CLIENT_ID=your_google_client_id.apps.googleusercontent.com| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /auth/signup |
— | Register with email + password; returns { token } (auto-login) |
| POST | /auth/login |
— | Login, returns { token } |
| POST | /auth/google-oauth |
— | Verify Google ID token, upsert user, return { token } |
| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /conversations |
Required | Create a new research session |
| GET | /conversations |
Required | List the user's sessions |
| GET | /conversations/{id}/messages |
Required | Load messages for a session |
| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /upload |
Required | Upload a file (PDF/TXT/MD/CSV) |
| GET | /files |
Required | List the user's uploaded files |
| DELETE | /files/{filename} |
Required | Delete an owned file |
| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /chat |
Optional | Stream agent response (SSE). Pass conversation_id to persist messages. |
{
"messages": [{ "role": "user", "content": "Summarise recent AI news" }],
"conversation_id": 42
}type |
Payload | Description |
|---|---|---|
text |
{ content: string } |
Text delta from the model |
tool_call |
{ tool, args } |
Agent is calling a tool |
done |
— | Stream complete |
The agent runs a tool-calling loop powered by Groq (llama-3.3-70b-versatile). It can call any combination of the following tools before producing its final answer. Each tool call is streamed to the frontend as a tool_call SSE event so the UI can show live progress pills.
| Tool | Source file | Description |
|---|---|---|
search_web |
tools/web_search.py |
Queries DuckDuckGo and returns up to 5 results {title, url, snippet}. Uses ddgs + curl_cffi for browser-like TLS fingerprinting. Retries up to 4× with exponential back-off on rate limits. |
fetch_page |
tools/web_search.py |
Downloads a URL, strips scripts/nav/footer with BeautifulSoup, and returns up to 8 000 chars of clean plain text. Used after search_web to get full article content. |
list_uploaded_files |
tools/file_reader.py |
Returns the list of filenames in the current user's upload directory. |
read_file |
tools/file_reader.py |
Reads up to 12 000 chars from an uploaded plain-text file (.txt, .md, .csv). Prevents path traversal via os.path.basename. |
extract_pdf_text |
tools/pdf_parser.py |
Extracts text from an uploaded PDF using pdfplumber, returning up to 12 000 chars. |
- Max 8 iterations before the loop is forced to stop.
- Tool results are executed in a thread (
asyncio.to_thread) so they don't block the async event loop. - Upload paths are per-user — each user's tools only see their own files.
- On
failed_generationfrom Groq the loop retries once automatically before giving up.