Skip to content

Repository files navigation

Local RAG API

A local Retrieval-Augmented Generation (RAG) application for querying documents through a conversational interface.

The project uses Ollama for local embeddings and text generation, PostgreSQL with pgvector for semantic search, MinIO for document storage, and a React frontend.

The goal was to build a complete RAG application, from document ingestion and vector storage to semantic retrieval, local generation and persistent conversations.

Features

  • Upload and process PDF, Markdown and plain-text documents.
  • Automatic text extraction and chunking.
  • Embeddings generated locally with Ollama.
  • Semantic search with PostgreSQL and pgvector.
  • Local text generation with Ollama.
  • Persistent conversations and messages.
  • Original documents stored in MinIO.
  • React + TypeScript frontend.
  • FastAPI backend.
  • Docker-based environment.
  • Automated backend tests.

Stack

  • Python 3.12
  • FastAPI
  • PostgreSQL
  • pgvector
  • SQLAlchemy
  • Alembic
  • MinIO
  • Ollama
  • React
  • TypeScript
  • Vite
  • Docker
  • Docker Compose
  • pytest

How it works

The application follows a RAG pipeline:

Document
   |
   v
Text extraction
   |
   v
Chunking
   |
   v
Embedding generation
   |
   v
PostgreSQL + pgvector

When the user asks a question, the question follows a similar embedding process:

User question
   |
   v
Question embedding
   |
   v
Semantic search
   |
   v
Relevant document chunks
   |
   v
Context + question
   |
   v
Ollama
   |
   v
Answer

Both document chunks and user questions are represented as vectors. The application uses cosine similarity to retrieve the document chunks that are most relevant to the question.

The current chunk configuration is:

CHUNK_SIZE = 1000
CHUNK_OVERLAP = 200

Running locally

Requirements

  • Docker
  • Docker Compose

Configure the environment:

cp .env.example .env

Then start the application:

make

That's it.

The project starts the required services and applies the database migrations.

The application is available at:

http://localhost:5173

The API is available at:

http://localhost:8000

Swagger:

http://localhost:8000/docs

API

The main question-answering endpoint is:

POST /ask

The complete API documentation is available through Swagger at /docs.

Testing

Run the backend tests with:

make test

The frontend can also be checked with:

cd frontend
npm run build
npm run lint

Project structure

Local-RAG-API/
├── backend/
│   ├── alembic/
│   ├── app/
│   └── tests/
├── frontend/
│   ├── public/
│   └── src/
├── docker-compose.yml
├── Makefile
├── .env.example
└── README.md

Why I built it

I wanted to build a project that went beyond simply calling an external AI API.

The idea was to understand and implement the infrastructure around a real AI application: document processing, embeddings, vector search, retrieval, local generation, persistence and a user interface.

The project gave me practical experience with:

  • Python and FastAPI
  • PostgreSQL and pgvector
  • RAG architecture
  • Embeddings and semantic search
  • Local LLMs with Ollama
  • Document processing
  • Object storage with MinIO
  • Docker
  • React and TypeScript
  • Automated testing

The result is a small but complete application that can ingest documents and use them as a knowledge base for a local AI assistant.

About

Retrieval-Augmented Generation with Python, PostgreSQL/pgvector and Ollama.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages