Skip to content

Latest commit

Β 

History

3,770 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Ragen AI

Ragen AI - now open-sourced πŸŽ‰

Turn your company documents into an AI assistant that answers from your data β€” on your own servers, with your own models.

Crafted by hand. Extended by agents.

Live demo Β· Documentation Β· Quickstart Β· API reference Β· TypeScript SDK Β· Security review Β· Self-hosting Β· Open models

Built and maintained by Web Amigos.


🧩 The problem

Most "chat with your documents" tools land in one of two places. Either they are a demo β€” naive top-k cosine search over a pile of PDFs, confidently wrong the moment a question needs more than one document β€” or they are a SaaS product you hand your contracts, personnel files and client data to, on someone else's infrastructure, under someone else's retention policy.

Ragen is neither. It is an open-source RAG platform for companies that you run yourself: retrieval that has been measured rather than assumed, access control enforced where it actually matters, and a model layer you can point at your own hardware.

πŸ’‘ Why Ragen

Permissions are enforced at retrieval, not in the UI. A document a user cannot open cannot appear in an answer, or in a citation, or in the context sent to the model. Every chunk carries an accessible_by filter applied at query time. Most tools filter the file list and pass everything to the prompt.

Retrieval that was measured, not assumed. Hybrid dense + BM25 sparse search with server-side RRF, multi-query expansion, document summaries generated at ingest, and cross-encoder reranking over the result set. Four composed techniques β€” the first three on by default, reranking once you give it a provider β€” each with an ADR recording what it changed and what it cost: ADR-14, ADR-15, ADR-16, ADR-12. Retrieval quality, rephrasing and red-team behaviour all have eval datasets you can run yourself: apps/web/evals.

Your infrastructure, your keys. Documents, database, vector index and encryption keys stay where you put them. No component reports back to us, and there is no hosted tier we would rather sell you. docs/security-and-privacy.md sets out exactly what leaves your network under which configuration β€” including the two cases where the honest answer is "it depends how you set it up".

No lock-in on the model. Ragen calls providers itself, or goes through a gateway you run β€” LiteLLM, Portkey, vLLM, Ollama, anything speaking OpenAI's API. Scaleway, Azure OpenAI, AWS Bedrock, Google Vertex, OpenRouter, or a model on your own GPU β€” changing provider is a route, not a migration.

Multi-tenancy is structural. One Qdrant collection per organization, a tenant-scope guard in the Prisma layer that flags any query missing its org filter, and two deliberately separate role hierarchies. It was designed in, not retrofitted.

Drop-in API. An OpenAI-compatible REST API and an official TypeScript SDK (@webamigos/ragen-sdk-ts) with streaming, typed responses and upload helpers. Most existing clients work by changing the base URL.

πŸ› οΈ Crafted by hand. Extended by agents.

Ragen was not generated. Two years and 3,200 commits of hand-written architecture came first β€” the tenant-scope guard, the retrieval permission model, the CQRS feature modules, the shared contracts package. Thirty-four ADRs record what was rejected and why.

Only then did the agent harness go on top: AGENTS.md as the canonical brief, five repo-specific skills for the operations that actually go wrong here (tenant-scope audits, RAG changes with measurement, e2e triage, dependency upgrades), and architecture tests that fail the build when a rule is broken rather than trusting a comment to hold it.

That order is the point. An agent working in this repository is not inventing conventions β€” it is following ones that survived production. The entries in docs/lessons/ are the incidents that taught us which guardrails had to be executable.

What that buys you: your team's AI coding agent is useful here on day one, because the invariants it has to respect are written down, linked, and enforced by CI instead of living in someone's head.

πŸ—οΈ What people build with it

  • Internal knowledge base β€” onboarding, procedures, contracts and project history, answerable in chat, with each answer citing the document it came from
  • Customer-facing chatbot β€” an embeddable widget grounded in your public documentation, with a separate assistant configuration and its own limits
  • RAG behind your own product β€” the OpenAI-compatible API and SDK as the retrieval layer of an application you are already building
  • Regulated document sets β€” where the answer to "who could have read this?" has to be provable, and the documents cannot leave your infrastructure
  • An assistant wired into your tools β€” Google Workspace, Slack, HubSpot, ClickUp and more over MCP, so the model can read live systems mid-conversation rather than only what was indexed last night

πŸ“Έ Screenshots

The app itself β€” an assistant answering from a knowledge base:

Ragen chat

The knowledge base, where documents are uploaded, versioned and shared:

Knowledge base

Settings, which is where most of the per-user and per-organization behaviour is decided:

Application settings

Ragen Brain turns the documents into knowledge pages people review and own, and draws how they relate. The graph is where a curator sees what the corpus actually says, what depends on what, and where the gaps are:

Ragen Brain graph

A knowledge page: the claims, each with the quoted passage it came from, its owner, who may read it, and whether it is in the knowledge base yet:

Ragen Brain knowledge page

The platform admin panel β€” one installation, every organization in it. This is the operator's surface and a separate app; the per-organization settings a customer's own owners and admins use live in the main app, and neither is a bigger version of the other. The two roles are unrelated fields on unrelated tables: Two kinds of administrator.

Ragen admin dashboard

Connector health, showing which MCP integrations are failing and why:

Connector health

Every page of the panel is documented, with screenshots regenerated from a scripted demo state rather than captured by hand β€” unless one is marked manual in scripts/screenshots/capture.mts, which is how a hand-placed image survives the next run: Admin panel.

Three screenshots of the app itself are currently such exceptions. They are the design-system v2 targets from apps/web/design_handoff_ragen_panel/, so they show where the interface is going rather than where it is; each goes back to being generated as its phase lands. The knowledge base is the first to make that trip β€” phase 7 has shipped, so its image is a capture again.

Or skip the screenshots and use it: the app is live at demo.ragen.ai, running against a seeded showcase organization rather than real customer data. The admin panel is the operator's surface and has no public instance β€” the page above is what it looks like.

βš–οΈ How it compares

Hosted "chat with your docs" Your own LangChain stack Ragen
Where your documents live Vendor's cloud Yours Yours
Choice of model Vendor's shortlist Anything Any provider, direct or through your own gateway
Access control Usually per workspace Whatever you build Per file and folder, enforced at retrieval
Multi-tenant Per seat, per workspace Whatever you build Built in β€” org-scoped data and index
Retrieval quality Opaque Yours to tune, and to debug Hybrid + rerank + multi-query, ADR per decision
Time to a working answer Minutes Weeks Minutes β€” one command, plus your own model keys
Cost shape Per seat, forever Your engineers' time Your infrastructure + model spend
When it breaks Support ticket You You, with the source and the ADRs

Fair warning on the middle column: if your requirements are genuinely unusual, building it yourself is a legitimate answer. Ragen is the better trade when you want those decisions already made β€” and documented β€” rather than made by you.

πŸš€ Quick start

Don't have the repo yet:

npx create-ragen-app my-ragen-app

create-ragen-app clones the repo, generates every secret it safely can, lets you paste a plain OpenAI or Anthropic key instead of configuring an enterprise LLM provider, starts the backing services in Docker, and runs the first-time Prisma setup β€” ending at cd my-ragen-app && npm run web:dev. Requires Node.js ^24.15.0 || >=26.0.0 β€” a range rather than a minimum, and the wizard refuses outside it before it clones anything; --skip-docker, --skip-install and --yes are available for a more manual or scripted run.

Already have the repo cloned:

npm run ragen:up:everything

Builds and starts every Ragen application β€” web, API, ingest worker, admin β€” alongside Postgres, Qdrant, Redis and Docling. No Node toolchain on the host, which makes it the fastest way to evaluate a self-hosted install.

Deploying rather than evaluating? Every release publishes the application images, so a deployment that is not tracking main does not have to compile them:

docker pull ghcr.io/webamigos/ragen-web:latest    # also: -api, -worker, -admin, -mcp

Each carries four tags β€” the full version, the minor series, the exact commit (sha-…) and latest. Pin production to the sha- tag; the other three move.

Either way, open http://localhost:3000, upload a document, and ask it something.

Requirements

Evaluating Small production install
CPU 4 cores 4+ cores
RAM 8 GB available to Docker 16 GB
Disk 25 GB 100 GB SSD, growing with your documents
Docker >= 24.0, Compose >= v2.26 same
GPU not needed not needed

No GPU, unless you want one. Chat, embeddings and reranking all leave over the network, so the machine running Ragen does no model inference of its own. Point it at a hosted provider and a laptop is enough. A GPU only enters the picture if you decide to serve models yourself, which is supported and is a separate box.

Where the memory actually goes. Measured on an idle stack, backing services only:

Service Idle memory Needed for
Presidio analyzer 959 MB PII masking (optional)
Docling 721 MB local document parsing
Postgres 93 MB everything
Presidio anonymizer 55 MB PII masking (optional)
Qdrant 43 MB retrieval β€” grows with your index
Redis 11 MB the ingest queue
Total ~1.9 GB

Three of those are optional, and together they are most of the total β€” the table is the stack with every optional part switched on. Presidio is not in a default docker compose up at all: both of its services sit behind the pii profile, so unless you asked the installer for PII masking the idle stack is about 900 MB rather than 1.9 GB. Drop Presidio if you are not masking PII, DOCUMENT_PARSER=legacy skips Docling, and Ragen calls model providers itself, so there is no proxy in this table any more β€” the litellm services went with the path that used them. Qdrant is the line that moves as you add documents; the figure above is a near-empty index, so size that one against your own corpus rather than against this table. Redis is not optional under the default runtime β€” it holds the ingest queue, and rate limiting rides along on it. The four applications run on top of all this and are not in the table.

npm run ragen:up:app is not a smaller way to run Ragen any more: it brings up Postgres and Qdrant without Redis, which was a working chat-only stack while Redis was merely a cache. Under BullMQ (ADR-44) it is not β€” apps/web and apps/api validate REDIS_URL at boot whether or not you ever upload a document, because a producer that cannot reach the queue cannot enqueue at all. Use ragen:up:full. Sizing guidance for larger installs is in Self-hosting.

For development, npm run ragen:up:full runs the dependencies in containers and leaves the apps running from source with hot reload. To contribute without installing any of it, the repository ships a dev container β€” Code β†’ Codespaces on GitHub, described in .devcontainer/README.md. Full instructions: Self-hosting Β· Local development.

✨ Features

Retrieval Hybrid dense + BM25 sparse search over Qdrant Β· multi-query expansion Β· ingest-time document summaries Β· cross-encoder reranking (opt-in, needs a provider) Β· citations on every answer Β· per-organization vector collections

Knowledge base Nested folders Β· per-user file ownership Β· sharing with users and teams Β· document versions with diff and rollback Β· re-indexing on content change Β· RAG optimization suggestions you review before accepting

Documents PDF, DOCX, XLSX, CSV, EPUB, SRT, Markdown, plain text, images and URLs Β· local parsing with Docling by default Β· async ingest on a queue, so a large upload does not block anything

Integrations (MCP) Client and server both Β· Google Workspace, Gmail, Slack, HubSpot, ClickUp, Fireflies, WooCommerce Β· four auth styles including OAuth with PKCE Β· OAuth tokens held in a separate vault service, never in the application database Β· a Ragen assistant is also itself callable as an MCP tool (apps/mcp) by external clients like Claude Desktop or Cursor, authenticated with the same API key as the REST API

API and SDK OpenAI-compatible REST API Β· official TypeScript SDK Β· opaque API keys Β· streaming over SSE Β· embeddable chatbot widget

Security Opt-in AES-256-GCM envelope encryption, one key per conversation Β· optional PII masking via Presidio Β· audit log with before-and-after state Β· tenant-scope guard over ~20 models Β· no training on your documents, ever

Operations Platform admin app Β· per-organization model allowlists and usage limits Β· OpenTelemetry traces, metrics and logs Β· UI in 15 languages β€” English, Polish, Spanish, German, French, Portuguese, Italian, Hungarian, Bulgarian, Ukrainian, Danish, Swedish, Finnish, Czech and Slovak

βš™οΈ How it works

Ingest β€” a file lands in storage, and a background job takes over: parse, chunk with a splitter chosen for the file type, generate a summary of the whole document, prepend that summary as its own chunk, embed densely and sparsely, upsert into the organization's Qdrant collection. It runs asynchronously, so a 400-page PDF does not block anything.

upload β†’ parse β†’ chunk β†’ summarize β†’ hybrid embed β†’ Qdrant

Retrieval β€” a question is first rewritten into a standalone one using the conversation so far, then expanded into an alternative phrasing. Both run as hybrid searches in parallel; Qdrant fuses dense and sparse results server-side with RRF; the union is deduplicated and, when a rerank provider is configured, passed to a cross-encoder before the model sees it, with citation prompting on top.

question β†’ standalone β†’ +1 variant β†’ 2Γ— hybrid search β†’ RRF β†’ dedupe β†’ rerank β†’ answer

Hybrid search is not a toggle β€” it is the schema collections are created with, so it is always on. Reranking is: it needs FEATURE_FLAG_RERANKING=1 and a provider's credentials, and is skipped without them. Every stage degrades rather than fails β€” a reranker error falls back to the raw vector order, an expansion error falls back to a single query. Tuning constants and flag names: docs/rag-pipeline.md.

πŸ”’ Security and privacy

Ragen is self-hosted. Documents, database, index and encryption keys stay on infrastructure you control, and nothing reports back to the vendor. Whether document content leaves your network during processing depends on how you configure the model backend β€” which is a real decision, not a detail, and the document below treats it as one.

docs/security-and-privacy.md answers the questions that come up in a security review: where data lives, what leaves the network, encryption, access control, audit logging, and whether documents are used for training. Every claim points at the code or the ADR behind it, and says plainly where something is configuration-dependent or not yet built.

Open models on your own hardware is the other half of that answer: how to point Ragen at a vLLM or Ollama server you run β€” directly or through a proxy β€” the four model settings that keep a cloud default until you change them, and the calls that still reach outward once you have.

Two things worth knowing before you deploy:

  • Parsing is local by default, but it falls back. DOCUMENT_PARSER=docling parses on your own hardware; if Docling fails, the worker falls back to loaders that send PDFs to an external model. Set DOCLING_STRICT=1 to fail instead of falling back.
  • Encryption at rest is opt-in. With no key provider configured, Ragen starts normally and stores message content unencrypted β€” convenient locally, wrong in production. Set ENCRYPTION_PROVIDER.

We would rather tell you this here than have you find it during an audit.

πŸ”“ Open core

Ragen is licensed under Apache 2.0. The rule for what that covers is deliberately mechanical: if a directory contains its own LICENSE file, that file governs everything under it. Everything else is Apache 2.0. No per-file headers, no allowlists, no exceptions you have to go looking for. Commercial paths today: none β€” the entire repository is Apache 2.0.

Three commitments constrain what may ever change:

  • Security is not an upsell. Tenant isolation, encryption at rest, PII masking and access control are core and stay core. Charging extra for the mechanisms that give you control over your own data would undercut the point of the product.
  • The core never degrades because the commercial layer is absent. Ragen runs without Stripe, without Presidio, without a telemetry backend and without AWS KMS. An install with no commercial licence is a complete Ragen, not a crippled one.
  • Multi-tenancy is core. Organization scoping runs through the data model, the vector store and the access-control layer. It could not be withheld without dismantling the architecture, and it will not become a paid tier.

Full text, including how this affects contributions: docs/open-core-boundary.md.

πŸ—ΊοΈ Roadmap

No dates. The order below is what we are working on, and it changes when a real install needs something we did not expect. Open work lives in GitHub Issues; this section is the shape of it rather than a substitute for it.

Next

  • Pluggable infrastructure, and a smaller install as the payoff. Three parts of the stack are things you should be able to choose rather than inherit: the job runtime, the vector store and the document parser. Each has a spec that puts a seam in front of it β€” BullMQ as the worker runtime, pgvector alongside Qdrant, Mistral Document AI alongside Docling. Two of the three only add a choice: Qdrant stays the default vector store and Docling stays the default parser. The worker is the exception, and it has landed β€” BullMQ is the runtime (ADR-44), Temporal is an adapter behind the same seam, and what is still ahead is moving that adapter to a separate repository, where it stays supported for installs that want durable execution. A fourth spec, the in-process model gateway, replaced the LiteLLM proxy.

    The payoff is that an install can be assembled to fit: select every alternative and Ragen runs on a small VM, or on managed Postgres with no container orchestration at all β€” without that becoming the only shape on offer. The trade-offs are real and each spec states them: a hosted parser means your documents leave your deployment, pgvector shares a failure domain with your database, and BullMQ re-runs a crashed job instead of resuming it.

  • A Python client. The TypeScript SDK is official and published. Python is the language most people integrating the API are actually writing in.

  • Slack as a place to ask. Not the existing Slack connector, which reads Slack as a source, but an assistant you can talk to in a channel, carrying the same retrieval-time access filter as the application.

  • A verified air-gapped configuration. A local model, local embeddings, DOCLING_STRICT=1, and a script that demonstrates no outbound traffic. The pieces are all there today.

  • Microsoft Entra ID sign-in. SSO and MFA are not built yet. Entra ID over OAuth comes first; SAML and SCIM directory sync come after it.

  • docker compose up to a working demo, with sample documents and sample questions, so evaluating Ragen does not begin with an empty knowledge base.

  • Microsoft 365 connectors: SharePoint, OneDrive, Outlook, Teams.

  • Feedback collection in the application, which is the signal the retrieval work has been missing.

Later

  • Permissions inherited from the source system. Today accessible_by is set in Ragen. Reading permissions out of Google Drive and SharePoint and keeping them in sync is the harder and more useful version, including revocation taking effect without waiting for a reindex.
  • Confluence and Jira connectors
  • Web search and multi-step research as tools, opt-in per organization, because some installs deliberately have no outbound path and that has to stay true
  • Audit log and security-event export to a SIEM, with retention and automatic deletion
  • Organization export, so leaving Ragen is a documented procedure rather than a support conversation
  • Fail-closed encryption: refuse to start in production without a key provider, instead of starting and storing message content unencrypted
  • MFA and passkeys
  • Per-user permissions for individual MCP tools
  • Usage drill-down from organization to team to user to request, with budget alerts
  • A sizing guide, a versioning policy and a CHANGELOG

Not planned

Worth saying plainly, because it saves you evaluating us for something we are not building:

  • Image generation, advanced voice mode, a code interpreter. Ragen answers from your documents. A general-purpose assistant is a different product and there are good ones.
  • A knowledge graph layer. We would rather improve retrieval we can measure than add a stage we cannot.
  • A visual workflow builder. Tools reach Ragen over MCP. The work we would rather do in that layer is approval and audit, not a node canvas.

Influencing this. Open an issue describing the install you are trying to do. A deployment blocked on a missing connector moves faster than a feature request in the abstract.

πŸ›οΈ Architecture at a glance

An npm-workspaces monorepo on Turborepo. Six applications and eight packages share one Prisma schema.

Application What it is
apps/web The Next.js app β€” chat, knowledge base, projects, settings
apps/api NestJS public API, the OpenAI-compatible surface
apps/worker Job worker: ingest, embedding, re-indexing
apps/admin Platform admin β€” organizations, models, limits, usage
apps/mcp MCP server exposing Ragen's own chat to external MCP clients (Claude Desktop, Cursor)
Package Shared by
rag-core vector and embedding contracts β€” web, api, worker
platform-contracts values every app must resolve identically: model catalogue, feature flags, connector metadata, tenant-scope model map
storage file storage providers β€” local by default, S3-compatible opt-in
vault-client, observability, db, eslint-config the remaining cross-app wiring

Supporting services β€” Docling, Presidio, the OTel collector β€” live in infra/, each with its own deployment config.

The full picture β€” routing, auth, the settings and admin surfaces, the feature modules, the connector registry β€” is in docs/architecture.md. What each companion service is and how to run the set locally: docs/companion-services.md.

Thirty-four ADRs record the decisions and what was rejected. Start with ADR-21 for the monorepo shape and ADR-33 for why shared values live in one package.

πŸ“š Documentation

Live demo The app, seeded with sample data
Quickstart First install, first document, first question
Self-hosting Deployment, sizing, configuration
Open models Running vLLM or Ollama, and what else it takes to keep traffic inside your network
Concepts Assistants, knowledge bases, projects, organizations
API reference The public API, endpoint by endpoint
Security The document to hand a security reviewer
ADRs Why the architecture is the way it is
AGENTS.md The contributor and coding-agent brief

Deeper reference, in docs/: architecture Β· RAG pipeline Β· knowledge base Β· document versioning Β· document processing Β· vector store Β· MCP integrations Β· token vault Β· file storage Β· companion services Β· model routing Β· event bus Β· thread encryption Β· tenant-scope guard Β· testing conventions Β· lessons

🀝 Contributing

Pull requests are welcome on everything under Apache 2.0. Start with CONTRIBUTING.md, and read AGENTS.md before your first change β€” it is the canonical brief for both humans and coding agents, and it will save you a review round. Security issues go through SECURITY.md, not the public tracker.

Questions, ideas and "has anyone deployed this against X" belong in Discussions. Issues are for work with a definition of done; a discussion does not need one.

πŸ“„ License

Apache 2.0. See Open core for the boundary rule.

ℹ️ About

Ragen is built and maintained by Web Amigos, an IT company in Poland. We build it because our own clients needed it and would not put their documents in someone else's cloud.

Commercial support, deployment help and enterprise terms: ragen@webamigos.pl.

About

Open-Source RAG Platform for Companies. Knowledge bases with multilingual retrieval and reranking, answers traceable to the source document, MCP server and a public API. No lock-in on the answering model, from AWS Bedrock, Azure, Ollama, Portkey, LiteLLM to a model running on your own infrastructure

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages