Turn your company documents into an AI assistant that answers from your data β on your own servers, with your own models.
Crafted by hand. Extended by agents.
Live demo Β· Documentation Β· Quickstart Β· API reference Β· TypeScript SDK Β· Security review Β· Self-hosting Β· Open models
Built and maintained by Web Amigos.
Most "chat with your documents" tools land in one of two places. Either they are a demo β naive top-k cosine search over a pile of PDFs, confidently wrong the moment a question needs more than one document β or they are a SaaS product you hand your contracts, personnel files and client data to, on someone else's infrastructure, under someone else's retention policy.
Ragen is neither. It is an open-source RAG platform for companies that you run yourself: retrieval that has been measured rather than assumed, access control enforced where it actually matters, and a model layer you can point at your own hardware.
Permissions are enforced at retrieval, not in the UI. A document a user
cannot open cannot appear in an answer, or in a citation, or in the context sent
to the model. Every chunk carries an accessible_by filter applied at query
time. Most tools filter the file list and pass everything to the prompt.
Retrieval that was measured, not assumed. Hybrid dense + BM25 sparse search
with server-side RRF, multi-query expansion, document summaries generated at
ingest, and cross-encoder reranking over the result set. Four composed
techniques β the first three on by default, reranking once you give it a
provider β each with an ADR recording what it changed and what it cost:
ADR-14,
ADR-15,
ADR-16,
ADR-12. Retrieval quality,
rephrasing and red-team behaviour all have eval datasets you can run yourself:
apps/web/evals.
Your infrastructure, your keys. Documents, database, vector index and encryption keys stay where you put them. No component reports back to us, and there is no hosted tier we would rather sell you. docs/security-and-privacy.md sets out exactly what leaves your network under which configuration β including the two cases where the honest answer is "it depends how you set it up".
No lock-in on the model. Ragen calls providers itself, or goes through a gateway you run β LiteLLM, Portkey, vLLM, Ollama, anything speaking OpenAI's API. Scaleway, Azure OpenAI, AWS Bedrock, Google Vertex, OpenRouter, or a model on your own GPU β changing provider is a route, not a migration.
Multi-tenancy is structural. One Qdrant collection per organization, a tenant-scope guard in the Prisma layer that flags any query missing its org filter, and two deliberately separate role hierarchies. It was designed in, not retrofitted.
Drop-in API. An OpenAI-compatible REST API and an official TypeScript SDK
(@webamigos/ragen-sdk-ts) with streaming, typed responses and upload
helpers. Most existing clients work by changing the base URL.
Ragen was not generated. Two years and 3,200 commits of hand-written architecture came first β the tenant-scope guard, the retrieval permission model, the CQRS feature modules, the shared contracts package. Thirty-four ADRs record what was rejected and why.
Only then did the agent harness go on top: AGENTS.md as the canonical brief,
five repo-specific skills for the operations that actually go wrong here
(tenant-scope audits, RAG changes with measurement, e2e triage, dependency
upgrades), and architecture tests that fail the build when a rule is broken
rather than trusting a comment to hold it.
That order is the point. An agent working in this repository is not inventing
conventions β it is following ones that survived production. The entries in
docs/lessons/ are the incidents that taught us which
guardrails had to be executable.
What that buys you: your team's AI coding agent is useful here on day one, because the invariants it has to respect are written down, linked, and enforced by CI instead of living in someone's head.
- Internal knowledge base β onboarding, procedures, contracts and project history, answerable in chat, with each answer citing the document it came from
- Customer-facing chatbot β an embeddable widget grounded in your public documentation, with a separate assistant configuration and its own limits
- RAG behind your own product β the OpenAI-compatible API and SDK as the retrieval layer of an application you are already building
- Regulated document sets β where the answer to "who could have read this?" has to be provable, and the documents cannot leave your infrastructure
- An assistant wired into your tools β Google Workspace, Slack, HubSpot, ClickUp and more over MCP, so the model can read live systems mid-conversation rather than only what was indexed last night
The app itself β an assistant answering from a knowledge base:
The knowledge base, where documents are uploaded, versioned and shared:
Settings, which is where most of the per-user and per-organization behaviour is decided:
Ragen Brain turns the documents into knowledge pages people review and own, and draws how they relate. The graph is where a curator sees what the corpus actually says, what depends on what, and where the gaps are:
A knowledge page: the claims, each with the quoted passage it came from, its owner, who may read it, and whether it is in the knowledge base yet:
The platform admin panel β one installation, every organization in it. This is the operator's surface and a separate app; the per-organization settings a customer's own owners and admins use live in the main app, and neither is a bigger version of the other. The two roles are unrelated fields on unrelated tables: Two kinds of administrator.
Connector health, showing which MCP integrations are failing and why:
Every page of the panel is documented, with screenshots regenerated from a
scripted demo state rather than captured by hand β unless one is marked
manual in scripts/screenshots/capture.mts, which is how a hand-placed
image survives the next run:
Admin panel.
Three screenshots of the app itself are currently such exceptions. They are the
design-system v2 targets from apps/web/design_handoff_ragen_panel/, so they
show where the interface is going rather than where it is; each goes back to
being generated as its phase lands. The knowledge base is the first to make
that trip β phase 7 has shipped, so its image is a capture again.
Or skip the screenshots and use it: the app is live at demo.ragen.ai, running against a seeded showcase organization rather than real customer data. The admin panel is the operator's surface and has no public instance β the page above is what it looks like.
| Hosted "chat with your docs" | Your own LangChain stack | Ragen | |
|---|---|---|---|
| Where your documents live | Vendor's cloud | Yours | Yours |
| Choice of model | Vendor's shortlist | Anything | Any provider, direct or through your own gateway |
| Access control | Usually per workspace | Whatever you build | Per file and folder, enforced at retrieval |
| Multi-tenant | Per seat, per workspace | Whatever you build | Built in β org-scoped data and index |
| Retrieval quality | Opaque | Yours to tune, and to debug | Hybrid + rerank + multi-query, ADR per decision |
| Time to a working answer | Minutes | Weeks | Minutes β one command, plus your own model keys |
| Cost shape | Per seat, forever | Your engineers' time | Your infrastructure + model spend |
| When it breaks | Support ticket | You | You, with the source and the ADRs |
Fair warning on the middle column: if your requirements are genuinely unusual, building it yourself is a legitimate answer. Ragen is the better trade when you want those decisions already made β and documented β rather than made by you.
Don't have the repo yet:
npx create-ragen-app my-ragen-appcreate-ragen-app clones the
repo, generates every secret it safely can, lets you paste a plain OpenAI or
Anthropic key instead of configuring an enterprise LLM provider, starts the
backing services in Docker, and runs the first-time Prisma setup β ending at
cd my-ragen-app && npm run web:dev. Requires Node.js ^24.15.0 || >=26.0.0 β
a range rather than a minimum, and the wizard refuses outside it before it
clones anything; --skip-docker,
--skip-install and --yes are available for a more manual or scripted run.
Already have the repo cloned:
npm run ragen:up:everythingBuilds and starts every Ragen application β web, API, ingest worker, admin β alongside Postgres, Qdrant, Redis and Docling. No Node toolchain on the host, which makes it the fastest way to evaluate a self-hosted install.
Deploying rather than evaluating? Every release publishes the application
images, so a deployment that is not tracking main does not have to compile
them:
docker pull ghcr.io/webamigos/ragen-web:latest # also: -api, -worker, -admin, -mcpEach carries four tags β the full version, the minor series, the exact commit
(sha-β¦) and latest. Pin production to the sha- tag; the other three move.
Either way, open http://localhost:3000, upload a document, and ask it something.
| Evaluating | Small production install | |
|---|---|---|
| CPU | 4 cores | 4+ cores |
| RAM | 8 GB available to Docker | 16 GB |
| Disk | 25 GB | 100 GB SSD, growing with your documents |
| Docker | >= 24.0, Compose >= v2.26 | same |
| GPU | not needed | not needed |
No GPU, unless you want one. Chat, embeddings and reranking all leave over the network, so the machine running Ragen does no model inference of its own. Point it at a hosted provider and a laptop is enough. A GPU only enters the picture if you decide to serve models yourself, which is supported and is a separate box.
Where the memory actually goes. Measured on an idle stack, backing services only:
| Service | Idle memory | Needed for |
|---|---|---|
| Presidio analyzer | 959 MB | PII masking (optional) |
| Docling | 721 MB | local document parsing |
| Postgres | 93 MB | everything |
| Presidio anonymizer | 55 MB | PII masking (optional) |
| Qdrant | 43 MB | retrieval β grows with your index |
| Redis | 11 MB | the ingest queue |
| Total | ~1.9 GB |
Three of those are optional, and together they are most of the total β the
table is the stack with every optional part switched on. Presidio is not in a
default docker compose up at all: both of its services sit behind the pii
profile, so unless you asked the installer for PII masking the idle stack is
about 900 MB rather than 1.9 GB. Drop
Presidio if you are not masking PII, DOCUMENT_PARSER=legacy skips Docling, and
Ragen calls model providers itself, so there is no proxy in this table any
more β the litellm services went with the path that used them. Qdrant is the line that moves as you add documents; the figure
above is a near-empty index, so size that one against your own corpus rather
than against this table.
Redis is not optional under the default runtime β it holds the ingest queue,
and rate limiting rides along on it. The four applications run on top of all
this and are not in the table.
npm run ragen:up:app is not a smaller way to run Ragen any more: it
brings up Postgres and Qdrant without Redis, which was a working chat-only
stack while Redis was merely a cache. Under BullMQ (ADR-44) it is not β
apps/web and apps/api validate REDIS_URL at boot whether or not you ever
upload a document, because a producer that cannot reach the queue cannot
enqueue at all. Use ragen:up:full. Sizing guidance for larger installs is in
Self-hosting.
For development, npm run ragen:up:full runs the dependencies in containers and
leaves the apps running from source with hot reload. To contribute without
installing any of it, the repository ships a dev container β Code β Codespaces
on GitHub, described in
.devcontainer/README.md. Full instructions:
Self-hosting Β·
Local development.
Retrieval Hybrid dense + BM25 sparse search over Qdrant Β· multi-query expansion Β· ingest-time document summaries Β· cross-encoder reranking (opt-in, needs a provider) Β· citations on every answer Β· per-organization vector collections
Knowledge base Nested folders Β· per-user file ownership Β· sharing with users and teams Β· document versions with diff and rollback Β· re-indexing on content change Β· RAG optimization suggestions you review before accepting
Documents PDF, DOCX, XLSX, CSV, EPUB, SRT, Markdown, plain text, images and URLs Β· local parsing with Docling by default Β· async ingest on a queue, so a large upload does not block anything
Integrations (MCP)
Client and server both Β· Google Workspace, Gmail, Slack, HubSpot, ClickUp,
Fireflies, WooCommerce Β· four auth styles including OAuth with PKCE Β· OAuth
tokens held in a separate vault service, never in the application database Β·
a Ragen assistant is also itself callable as an MCP tool (apps/mcp) by
external clients like Claude Desktop or Cursor, authenticated with the same
API key as the REST API
API and SDK OpenAI-compatible REST API Β· official TypeScript SDK Β· opaque API keys Β· streaming over SSE Β· embeddable chatbot widget
Security Opt-in AES-256-GCM envelope encryption, one key per conversation Β· optional PII masking via Presidio Β· audit log with before-and-after state Β· tenant-scope guard over ~20 models Β· no training on your documents, ever
Operations Platform admin app Β· per-organization model allowlists and usage limits Β· OpenTelemetry traces, metrics and logs Β· UI in 15 languages β English, Polish, Spanish, German, French, Portuguese, Italian, Hungarian, Bulgarian, Ukrainian, Danish, Swedish, Finnish, Czech and Slovak
Ingest β a file lands in storage, and a background job takes over: parse, chunk with a splitter chosen for the file type, generate a summary of the whole document, prepend that summary as its own chunk, embed densely and sparsely, upsert into the organization's Qdrant collection. It runs asynchronously, so a 400-page PDF does not block anything.
upload β parse β chunk β summarize β hybrid embed β Qdrant
Retrieval β a question is first rewritten into a standalone one using the conversation so far, then expanded into an alternative phrasing. Both run as hybrid searches in parallel; Qdrant fuses dense and sparse results server-side with RRF; the union is deduplicated and, when a rerank provider is configured, passed to a cross-encoder before the model sees it, with citation prompting on top.
question β standalone β +1 variant β 2Γ hybrid search β RRF β dedupe β rerank β answer
Hybrid search is not a toggle β it is the schema collections are created with,
so it is always on. Reranking is: it needs FEATURE_FLAG_RERANKING=1 and a
provider's credentials, and is skipped without them. Every stage degrades
rather than fails β a reranker error falls back to the raw vector order, an
expansion error falls back to a single query. Tuning constants and flag names:
docs/rag-pipeline.md.
Ragen is self-hosted. Documents, database, index and encryption keys stay on infrastructure you control, and nothing reports back to the vendor. Whether document content leaves your network during processing depends on how you configure the model backend β which is a real decision, not a detail, and the document below treats it as one.
docs/security-and-privacy.md answers the questions that come up in a security review: where data lives, what leaves the network, encryption, access control, audit logging, and whether documents are used for training. Every claim points at the code or the ADR behind it, and says plainly where something is configuration-dependent or not yet built.
Open models on your own hardware is the other half of that answer: how to point Ragen at a vLLM or Ollama server you run β directly or through a proxy β the four model settings that keep a cloud default until you change them, and the calls that still reach outward once you have.
Two things worth knowing before you deploy:
- Parsing is local by default, but it falls back.
DOCUMENT_PARSER=doclingparses on your own hardware; if Docling fails, the worker falls back to loaders that send PDFs to an external model. SetDOCLING_STRICT=1to fail instead of falling back. - Encryption at rest is opt-in. With no key provider configured, Ragen
starts normally and stores message content unencrypted β convenient locally,
wrong in production. Set
ENCRYPTION_PROVIDER.
We would rather tell you this here than have you find it during an audit.
Ragen is licensed under Apache 2.0. The rule for what that covers is
deliberately mechanical: if a directory contains its own LICENSE file, that
file governs everything under it. Everything else is Apache 2.0. No per-file
headers, no allowlists, no exceptions you have to go looking for. Commercial
paths today: none β the entire repository is Apache 2.0.
Three commitments constrain what may ever change:
- Security is not an upsell. Tenant isolation, encryption at rest, PII masking and access control are core and stay core. Charging extra for the mechanisms that give you control over your own data would undercut the point of the product.
- The core never degrades because the commercial layer is absent. Ragen runs without Stripe, without Presidio, without a telemetry backend and without AWS KMS. An install with no commercial licence is a complete Ragen, not a crippled one.
- Multi-tenancy is core. Organization scoping runs through the data model, the vector store and the access-control layer. It could not be withheld without dismantling the architecture, and it will not become a paid tier.
Full text, including how this affects contributions: docs/open-core-boundary.md.
No dates. The order below is what we are working on, and it changes when a real install needs something we did not expect. Open work lives in GitHub Issues; this section is the shape of it rather than a substitute for it.
Next
-
Pluggable infrastructure, and a smaller install as the payoff. Three parts of the stack are things you should be able to choose rather than inherit: the job runtime, the vector store and the document parser. Each has a spec that puts a seam in front of it β BullMQ as the worker runtime, pgvector alongside Qdrant, Mistral Document AI alongside Docling. Two of the three only add a choice: Qdrant stays the default vector store and Docling stays the default parser. The worker is the exception, and it has landed β BullMQ is the runtime (ADR-44), Temporal is an adapter behind the same seam, and what is still ahead is moving that adapter to a separate repository, where it stays supported for installs that want durable execution. A fourth spec, the in-process model gateway, replaced the LiteLLM proxy.
The payoff is that an install can be assembled to fit: select every alternative and Ragen runs on a small VM, or on managed Postgres with no container orchestration at all β without that becoming the only shape on offer. The trade-offs are real and each spec states them: a hosted parser means your documents leave your deployment, pgvector shares a failure domain with your database, and BullMQ re-runs a crashed job instead of resuming it.
-
A Python client. The TypeScript SDK is official and published. Python is the language most people integrating the API are actually writing in.
-
Slack as a place to ask. Not the existing Slack connector, which reads Slack as a source, but an assistant you can talk to in a channel, carrying the same retrieval-time access filter as the application.
-
A verified air-gapped configuration. A local model, local embeddings,
DOCLING_STRICT=1, and a script that demonstrates no outbound traffic. The pieces are all there today. -
Microsoft Entra ID sign-in. SSO and MFA are not built yet. Entra ID over OAuth comes first; SAML and SCIM directory sync come after it.
-
docker compose upto a working demo, with sample documents and sample questions, so evaluating Ragen does not begin with an empty knowledge base. -
Microsoft 365 connectors: SharePoint, OneDrive, Outlook, Teams.
-
Feedback collection in the application, which is the signal the retrieval work has been missing.
Later
- Permissions inherited from the source system. Today
accessible_byis set in Ragen. Reading permissions out of Google Drive and SharePoint and keeping them in sync is the harder and more useful version, including revocation taking effect without waiting for a reindex. - Confluence and Jira connectors
- Web search and multi-step research as tools, opt-in per organization, because some installs deliberately have no outbound path and that has to stay true
- Audit log and security-event export to a SIEM, with retention and automatic deletion
- Organization export, so leaving Ragen is a documented procedure rather than a support conversation
- Fail-closed encryption: refuse to start in production without a key provider, instead of starting and storing message content unencrypted
- MFA and passkeys
- Per-user permissions for individual MCP tools
- Usage drill-down from organization to team to user to request, with budget alerts
- A sizing guide, a versioning policy and a CHANGELOG
Not planned
Worth saying plainly, because it saves you evaluating us for something we are not building:
- Image generation, advanced voice mode, a code interpreter. Ragen answers from your documents. A general-purpose assistant is a different product and there are good ones.
- A knowledge graph layer. We would rather improve retrieval we can measure than add a stage we cannot.
- A visual workflow builder. Tools reach Ragen over MCP. The work we would rather do in that layer is approval and audit, not a node canvas.
Influencing this. Open an issue describing the install you are trying to do. A deployment blocked on a missing connector moves faster than a feature request in the abstract.
An npm-workspaces monorepo on Turborepo. Six applications and eight packages share one Prisma schema.
| Application | What it is |
|---|---|
apps/web |
The Next.js app β chat, knowledge base, projects, settings |
apps/api |
NestJS public API, the OpenAI-compatible surface |
apps/worker |
Job worker: ingest, embedding, re-indexing |
apps/admin |
Platform admin β organizations, models, limits, usage |
apps/mcp |
MCP server exposing Ragen's own chat to external MCP clients (Claude Desktop, Cursor) |
| Package | Shared by |
|---|---|
rag-core |
vector and embedding contracts β web, api, worker |
platform-contracts |
values every app must resolve identically: model catalogue, feature flags, connector metadata, tenant-scope model map |
storage |
file storage providers β local by default, S3-compatible opt-in |
vault-client, observability, db, eslint-config |
the remaining cross-app wiring |
Supporting services β Docling, Presidio, the OTel collector β live in
infra/, each with its own deployment config.
The full picture β routing, auth, the settings and admin surfaces, the feature modules, the connector registry β is in docs/architecture.md. What each companion service is and how to run the set locally: docs/companion-services.md.
Thirty-four ADRs record the decisions and what was rejected. Start with ADR-21 for the monorepo shape and ADR-33 for why shared values live in one package.
| Live demo | The app, seeded with sample data |
| Quickstart | First install, first document, first question |
| Self-hosting | Deployment, sizing, configuration |
| Open models | Running vLLM or Ollama, and what else it takes to keep traffic inside your network |
| Concepts | Assistants, knowledge bases, projects, organizations |
| API reference | The public API, endpoint by endpoint |
| Security | The document to hand a security reviewer |
| ADRs | Why the architecture is the way it is |
| AGENTS.md | The contributor and coding-agent brief |
Deeper reference, in docs/:
architecture Β·
RAG pipeline Β·
knowledge base Β·
document versioning Β·
document processing Β·
vector store Β·
MCP integrations Β·
token vault Β·
file storage Β·
companion services Β·
model routing Β·
event bus Β·
thread encryption Β·
tenant-scope guard Β·
testing conventions Β·
lessons
Pull requests are welcome on everything under Apache 2.0. Start with CONTRIBUTING.md, and read AGENTS.md before your first change β it is the canonical brief for both humans and coding agents, and it will save you a review round. Security issues go through SECURITY.md, not the public tracker.
Questions, ideas and "has anyone deployed this against X" belong in Discussions. Issues are for work with a definition of done; a discussion does not need one.
Apache 2.0. See Open core for the boundary rule.
Ragen is built and maintained by Web Amigos, an IT company in Poland. We build it because our own clients needed it and would not put their documents in someone else's cloud.
Commercial support, deployment help and enterprise terms: ragen@webamigos.pl.







