Skip to content

Send pipeline events to maexbert-analytics for observability #75

Description

@maxi07

Goal

Send structured events from ScanSync to the existing self-hosted analytics service (analytics.maexbert.de, repo maxi07/maexbert-analytics) so that pipeline health, throughput and failures become visible in Grafana instead of only in container logs.

Today the only observability is logger output to stdout plus logfile.log (WARNING and above). To answer "how many documents failed OCR last week?" or "is OneDrive upload getting slower?" someone has to grep logs across seven containers.

Why this is worth doing

ScanSync is a multi-stage pipeline across separate services:

smb_service → detection_service → ocr_service → file_naming_service → upload_service

Each stage can fail independently, and a document that silently stalls between two stages is currently invisible. Event-based logging makes the funnel measurable: how many documents entered, how many completed, where they dropped out.

Privacy constraints (important)

ScanSync processes private documents — scanned invoices, contracts, medical letters. The analytics backend is a separate service and its database is not intended to hold document content.

Never send:

  • file names or SMB paths
  • OCR text or any document content
  • OneDrive target paths or folder names
  • user names, mail addresses, tenant IDs

Safe to send: durations, page counts, file sizes, booleans, enum-like states, exception types, service names, versions.

Where a correlation ID is useful, send a hash (e.g. first 12 chars of a SHA-256 over the internal job ID) rather than anything derived from the file name.

Proposed events

Dot notation (object.action), consistent with the other projects on the same backend.

Event When Suggested metadata
document.received new file detected on SMB share size_bytes, extension, sync_target_id
ocr.started OCR job picked up pages, size_bytes
ocr.completed OCR finished duration_ms, pages, language, skipped_text_layer (bool)
ocr.failed OCR raised error_type, duration_ms, pages
naming.completed AI renaming returned provider (openai/ollama), model, duration_ms, fallback_used (bool)
naming.failed renaming failed provider, error_type
upload.completed OneDrive upload done duration_ms, size_bytes, retries
upload.failed upload failed error_type, status_code, retries
document.completed whole pipeline done total_duration_ms, pages
queue.depth periodic sample queue_name, depth

document.received plus document.completed gives the completion rate for free; the *.failed events show where the rest is lost.

Lessons from the MatchMyTime integration

The same backend is already used by MatchMyTime, and its crash telemetry turned out to be nearly useless in practice (see maxi07/MatchMyTime issues 70 and 73). Three mistakes worth avoiding here:

  1. error_type was hardcoded to the literal string ERROR for every single event. Grouping failures then requires string matching on free-form messages. Use type(exc).__name__.
  2. No status code on failures, so it was impossible to tell whether a failure was user-visible. Include it where an HTTP call is involved (upload.failed).
  3. Expected failures and real bugs shared one event name. Invalid user input got logged as a crash and buried the genuine incidents. Keep ocr.failed for actual faults; if a file type is simply unsupported, that is its own event or a metadata flag — not a failure.

Implementation sketch

The natural place is scansynclib, next to the existing logging.py, so every service gets it from one import.

  • Add an analytics.py module wrapping the existing AnalyticsClient (client/analytics_client.py in the analytics repo — requests + daemon threads, fire-and-forget).
  • Read ANALYTICS_URL and ANALYTICS_API_KEY from the environment. If either is unset, the module becomes a no-op. ScanSync is self-hosted by other people; telemetry must be opt-in and must never be a hard dependency.
  • Failures to reach the analytics endpoint must never affect document processing — swallow and log at DEBUG.
  • Pass app_version from the existing version resolution so regressions can be attributed to a release.
  • Seed the project first: make seed NAME=scansync in the analytics repo produces the API key.

Backend API reference

POST /event
X-API-KEY: <key>

{
  "event_name": "ocr.completed",
  "user_id": null,
  "app_version": "1.4.2",
  "metadata": { "duration_ms": 8300, "pages": 4 }
}

Responses: 200 {"status":"ok"}, 400 on missing event_name, 401 on bad key.

Definition of done

  • scansynclib/analytics.py exists, is a no-op without configuration, and never raises into the caller
  • Events above are emitted from the respective services
  • No file names, paths or document content appear in any payload
  • error_type carries the real exception class name
  • Grafana dashboard scansync shows: documents per day, completion rate, failures by stage, OCR duration percentiles, queue depth
  • README documents the two environment variables and states clearly that telemetry is off by default

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions