A self-hosted archive for the paper and PDFs of a household: letters, invoices, contracts, statements. Scan them, forward them or drop them in a folder; Heftig keeps the original, reads the text, sorts it, and finds it again in milliseconds. It also remembers where the paper is.
Portable. The archive is a folder of ordinary files: every original byte for byte, next to it a JSON file with everything Heftig knows about it and the recognised text as Markdown. The SQLite database is only an index – delete it and Heftig rebuilds it from the files. Copy the folder to another machine and you have moved your archive; any backup tool can back it up; you can read it without Heftig.
Small on purpose. One person's archive, one inbox, one way to file paper. Heftig collects, reads, sorts, finds and remembers where the paper is – and leaves out what a household doesn't need: roles, workflows, sharing, plugins. Where a decision was possible, it was made, so there is little to configure.
Easy to change. About 24,000 lines of plain Python, server-rendered HTML and a little JavaScript, no build step, no front-end framework, and about 500 tests that run in a minute and a half without network access. That makes it a good fit for coding agents such as Claude Code or Codex: describe what you want different, let the agent change it and run the tests. AGENTS.md gives it the map and the rules that are easy to break. Heftig itself was built this way.
On Linux, with Podman or Docker (the installer offers to install Podman if neither is there), or on macOS with Docker Desktop, OrbStack, Colima or Podman running:
curl -fsSL https://raw.githubusercontent.com/jo-tud/heftig/main/install.sh | shIt downloads the container image, creates ~/heftig, starts Heftig (also after a reboot; on
macOS only if Docker Desktop, OrbStack, Colima or the Podman machine starts at login) and prints
the address of the setup page. There you create your (local) account and, if you like, connect
an AI, a mailbox and your scanner – no configuration files. Running the installer again updates
Heftig. To remove it: podman rm -f heftig (or docker rm -f heftig) and delete ~/heftig.
Just looking? Build a demo archive with fictional documents:
uv run python scripts/demo_archive.py /tmp/heftig-demo, then
HEFTIG_ARCHIVE_DIR=/tmp/heftig-demo uv run heftig run and sign in as demo /
demo-password-1234.
Other ways to run it (Docker Compose, without containers, behind a reverse proxy): docs/operations.md.
- Getting documents in: upload in the browser, the phone's document camera, a scanner folder (network share), a mailbox for forwarded mail and scan-to-email, a REST API. Everything lands in one inbox. Pages scanned separately can be combined into one document; an e-mail can be archived itself, as one document with its attachments.
- Reading and sorting: embedded PDF text first, OCR for scans (Tesseract on your computer, or an AI model). Title, date, sender, document type, tags, amounts and contract numbers are suggested; dates and numbers are only accepted if they really appear in the text. Anything you correct is locked and never overwritten.
- Finding: full-text search built for German paperwork: word forms (Kindern → Kind), compounds in both directions (Steuerbescheid → Einkommensteuerbescheid, Stromrechnung → Strom + Rechnung), other words for the same thing (Nebenkosten → Betriebskosten, your own groups too), OCR errors in scans, typos, umlauts and spacing in numbers; it understands phrases like “March 2025” or “last year”, shows filters with counts and a timeline, suggests while you type, and marks the hits on the page. Its quality is measured (docs/search.md). With the search by meaning switched on, a built-in language model also finds documents that say it in other words (“securities” finds the ETF statement) - on your computer, no server. Optionally, an AI turns a question like “phone bills over $50 last year” into filters.
- Paper without archive numbers: file letters in binders with a section per month, newest on top. Heftig records the binder and the position when you mark a letter as filed and later shows exactly where it is, and which letters lie above and below it.
- Keeping it clean: duplicates are recognised (the identical file, and the same letter as scan and PDF); deleted documents stay in the trash for 30 days; combining and bulk deletion can be undone, and the documents that were combined are kept for good.
- Asking questions: connect Claude via MCP and ask “how much did I pay for electricity in 2025?” – with links to the pages the answer comes from (docs/mcp.md).
The interface is in English and German.
Heftig works completely offline: Tesseract reads scans, simple rules sort documents (the search by meaning, if you switch it on, downloads its model once - about 330 MB - and then runs offline too). An AI model makes titles, dates, senders and types much better. You choose it on the setup page:
| What it costs | What leaves your computer | |
|---|---|---|
| None | nothing | nothing |
| Anthropic (Claude) – what Heftig was developed with | sorting about 2 cents per document (Claude Sonnet), reading scans with AI about 2 cents per page | the document text; page images only if AI reading is switched on |
| OpenAI | depends on the model | the same |
| Your own model (Ollama, LM Studio, vLLM, llama.cpp …) | your electricity | nothing |
Cloud use is off until you switch it on and confirm what is sent, and the inbox shows what the AI has cost so far. Details, and what is sent exactly: docs/providers.md.
Heftig is young (version 0.x) but in daily use, with about 1,700 documents so far. Things to know:
- It is not a legally certified archive (no WORM storage, no signatures). Keep originals where the law asks for them.
- One user per archive.
- Tesseract is clearly weaker than AI models on handwriting and poor scans.
- The rule-based sorting without AI knows German and English letters best.
| docs/guide.md | Using Heftig: inbox, filing paper, scan sessions, duplicates, e-mails, search tips, installing as an app |
| docs/operations.md | Installation variants, HTTPS and phones, backup and restore, updates, logs |
| docs/scanner-imap.md | Scanners (network folder, scan-to-email), SMB shares, mail import |
| docs/providers.md | AI and OCR providers, local models, privacy |
| docs/search.md | How search works: ranking, word forms, synonyms, typos, syntax, date phrases, search by meaning, measured quality |
| docs/mcp.md | Ask Claude about your archive (read-only) |
| docs/data-format.md | The archive on disk, export and import |
| docs/architecture.md | Components, pipeline, security model – the place to start before changing things |
| docs/testing.md | What the tests cover, a manual smoke test, known gaps |
| docs/paperless.md | Moving to Paperless-ngx: field mapping and a migration script |
git clone https://github.com/jo-tud/heftig.git && cd heftig
uv sync # Python 3.11+, dependencies incl. dev tools
make test # the test suite, offline
uv run python scripts/demo_archive.py /tmp/demo && HEFTIG_ARCHIVE_DIR=/tmp/demo uv run heftig runStart with docs/architecture.md. Interface texts are English in the code
and translated in src/heftig/locale/; uv run python scripts/i18n.py check de shows what is
missing. Contributions are welcome – see CONTRIBUTING.md.
MIT, see LICENSE.






