Skip to content

Repository files navigation

SystemBridge AI logo

SystemBridge AI

Turn invoices, receipts and scanned documents into a balanced, audit-ready ledger — automatically.

Stop retyping paperwork. SystemBridge AI captures business documents from uploads, folders and webhooks, reads them with deterministic rules and an optional local LLM, cross-checks every total, routes anything doubtful to a human, and posts a clean ledger you can export to your accounting tool.

Python FastAPI uv Tests Lint Docker

Business value · Capabilities · Screenshots · Quick start · Architecture

SystemBridge AI — turn invoices into balanced ledgers, automatically

Why it matters to a business

An SME receives invoices by email, portal download, scan and messenger. Someone then reads each one and types it into the accounting system. That work is slow, error-prone, and impossible to audit.

SystemBridge AI replaces that manual step with a pipeline that is fast, private and provably correct:

Everyday reality With SystemBridge AI
~3 minutes of manual typing per invoice < 2 seconds per document
10–15 operational hours a week lost to data entry 10–15 hours saved per week
A percentage of entries with typos or wrong totals 100% of totals cross-checked before posting
The same invoice re-sent creates a second entry 0 duplicate ledger entries — files are fingerprinted
Doubtful reads are posted silently Anything uncertain goes to a human review queue

Money is handled as Decimal, never a float. Every document is hashed (SHA-256) for idempotency. The LLM never gets the final word — deterministic rules corroborate every field and the schema validates every total.

What it does

  • Capture from anywhere — HTTP upload (POST /api/v1/ingest/file), a watched inbox folder, and ready-to-plug channels for email/webhooks (see roadmap).
  • Read any format — digital PDF text layers via pdfplumber, with automatic OCR fallback (Tesseract, multi-language) for scans and photos.
  • Extract with intent — deterministic rules find VAT ids, dates, amounts, document numbers and currency; a local or hosted LLM (Ollama · OpenAI · Groq) repairs and completes the draft.
  • Validate like an accountant — strict schema with subtotal + tax == total and Σ line items == subtotal, within a configurable tolerance.
  • Explain its confidence — every field is scored: a rule that finds the same value in the source corroborates it (1.0); invented values keep the lower base score. Low mean → human review.
  • Keep a human in the loop — a review queue shows the source, the extracted fields and the audit trail side by side, so an operator can approve or correct in seconds.
  • Deliver the result — one-click CSV / XLSX ledger export (validated rows only) plus desktop and webhook notifications.
  • Tune it live — change the LLM, OCR, confidence threshold, tolerance, worker interval and notifications from the UI. Changes apply instantly, no restart, and are persisted.
  • Stay observable — structured logs with a per-document correlation id, plus /health, /metrics (Prometheus) and /stats (JSON).

Screenshots

Real screenshots of the running app (headless Chrome against a seeded instance).

Overview — the pipeline at a glance

Dashboard overview

Documents registry Human review queue
Documents list Document review
Ledger & exports Live settings & theming
Ledger Settings

Built for the desk — and for the phone

Mobile landing Mobile dashboard Mobile review

More mobile captures (documents, ledger, settings, login) live in docs/assets/.

How it works

flowchart LR
    A["Upload<br/>API / folder / email"] --> B["Ingest<br/>SHA-256 idempotency"]
    B --> C["Read<br/>PDF text + OCR fallback"]
    C --> D["Extract<br/>rules + optional LLM"]
    D --> E{"Validate<br/>totals & schema"}
    E -- balanced --> F["Ledger<br/>CSV / XLSX + notify"]
    E -- doubtful --> G["Human review<br/>approve / correct"]
    G --> F
Loading
  1. Drop it in — upload a file, drop it in the inbox folder, or POST it to the API. Ingestion returns immediately (202).
  2. We read it — text is pulled from the PDF or OCR'd from the scan; the structuring engine fills the gaps; deterministic rules and the schema check every total.
  3. You approve — clean documents post themselves; anything doubtful waits in the review queue for one click, then joins the ledger.

Why it is trustworthy

Principle What it means in practice
Explainable confidence Rules corroborate values, never invent them. Invented fields keep a lower score and never post silently.
Idempotent ingestion Every payload is hashed while it streams to disk; a resend can never create a second ledger entry.
Balanced by construction Decimal money and enforced subtotal + tax == total before anything is posted.
Human in the loop Low-confidence or unbalanced documents go to a review queue with source and draft side by side.
Customizable live LLM, OCR, thresholds, tolerance and notifications are editable at runtime — no restart.
Measured Structured logs and Prometheus metrics for every stage, so health and improvement are visible.

Quick start

Requires uv and Python 3.12.

uv sync
uv run uvicorn app.main:app --reload

Or the whole stack (API + Ollama + SQLite) with Docker:

docker compose up --build -d

Important

The seed script and the server must use the same SYSTEMBRIDGE_DATABASE_URL. The default is ./data/systembridge.db; the examples below use ./data/app.db, so pass that env var to both.

uv sync
SYSTEMBRIDGE_DATABASE_URL="sqlite+aiosqlite:///./data/app.db" \
  uv run python scripts/seed_demo.py          # demo data
SYSTEMBRIDGE_DATABASE_URL="sqlite+aiosqlite:///./data/app.db" \
  uv run uvicorn app.main:app --port 8010
# open http://127.0.0.1:8010/

Simulated use case (offline)

A full "day at an SME" runs deterministically and without network access: four invoices arrive through three channels (portal, email, scan), plus a duplicate resend.

uv run python scripts/run_scenario.py
Document Channel Result Confidence
acme_invoice_001.pdf portal (digital) VALIDATED 0.99
beta_invoice.txt email VALIDATED 0.99
gamma_scanned.png scan (OCR) NEEDS_REVIEW 0.73
delta_broken_totals.pdf portal NEEDS_REVIEW —
acme_invoice_001_RESEND.pdf email DUPLICATE (skipped) —

Outputs: a ledger CSV/XLSX (validated rows only, total 726.00 EUR) and four notifications. The scanned invoice is caught by rules that detect ungrounded amounts; the broken-total invoice fails strict validation and is queued for a human.

Configuration

All settings are environment variables prefixed with SYSTEMBRIDGE_ (see .env.example). A validated subset is also editable live from the Settings page.

Variable Default Purpose
SYSTEMBRIDGE_DATABASE_URL sqlite+aiosqlite:///./data/systembridge.db Storage engine
SYSTEMBRIDGE_LLM_PROVIDER ollama ollama · openai · groq
SYSTEMBRIDGE_CONFIDENCE_THRESHOLD 0.80 Below this → human review
SYSTEMBRIDGE_AMOUNT_TOLERANCE 0.01 Allowed rounding for total checks
SYSTEMBRIDGE_OCR_ENABLED true OCR fallback for scans
SYSTEMBRIDGE_WATCHER_ENABLED false Watch the inbox folder for new files
SYSTEMBRIDGE_NOTIFY_WEBHOOK_URL — Post notifications to a webhook

Tech stack

FastAPI · Jinja2 · SQLAlchemy (async) · Pydantic v2 · pdfplumber · pypdfium2 · pytesseract · instructor · structlog · uv · ruff · pytest

Runs fully on SQLite, points at Postgres via one URL, and ships a Dockerfile + docker-compose.yml.

Project layout

app/
├── main.py            # FastAPI app (JSON API + UI) + /health /metrics /stats
├── config.py          # pydantic-settings configuration (live-mutable subset)
├── core/              # ingestion, processor, worker, watcher, extraction/
├── models/            # Pydantic domain schemas + SQLAlchemy ORM
├── services/          # persistence, ledger export, notifications
├── scenarios/         # deterministic simulated SME day
├── web/               # Jinja UI (routes, templates, static)
└── utils/             # logging, metrics, telemetry
docs/                  # architecture, UI spec, observability, visual contract
scripts/               # run_scenario.py, seed_demo.py, metrics_report.py
tests/                 # 71 tests (asyncio auto mode)
spec/                  # original technical spec + Stitch design reference
prototype/             # throwaway reference UI (not wired to the API)

See docs/ARCHITECTURE.md for the pipeline, state machine, module map and design decisions.

Development & quality

uv run ruff check .        # lint
uv run ruff format .       # format
uv run pytest              # 71 tests, ~3s
uv run pytest tests/test_processor.py::test_name   # single test

Order when changing code: ruff check → pytest. There is no in-repo CI config.

Roadmap

  • F0 Scaffold, config, Docker, quality gates
  • F1 Domain schema (Decimal, cross-field validation)
  • F2 Ingestion: upload endpoint, folder watcher, SHA-256 idempotency
  • F3 Extraction: PDF → OCR → LLM fallback with per-field confidence
  • F4 Persistence, ledger export (CSV/XLSX), pluggable notifications
  • F5 Background worker, crash recovery, retries-safe batch processing
  • F6 Web UI: dashboard, documents, review queue, ledger, settings
  • F7 IMAP listener (aioimaplib), async queue (Arq + Redis), multi-tenant auth & billing

Documentation

Document Contents
docs/ARCHITECTURE.md Pipeline, state machine, module map, design decisions, confidence model, glossary
docs/UI.md Full UI specification: screens, components, data bindings, API surface
docs/OBSERVABILITY.md Logs, metrics (/metrics, /stats), test telemetry
docs/SPEC-ALIGNMENT.md How the architecture maps onto the code (concordance, gaps, dependency delta)
docs/VISUALS.md Visual-asset contract and inventory

License

Released under the MIT License.

About

Turn invoices, receipts and scanned documents into a balanced, audit-ready ledger — automatically.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages