Goal
Send structured events from ScanSync to the existing self-hosted analytics service (analytics.maexbert.de, repo maxi07/maexbert-analytics) so that pipeline health, throughput and failures become visible in Grafana instead of only in container logs.
Today the only observability is logger output to stdout plus logfile.log (WARNING and above). To answer "how many documents failed OCR last week?" or "is OneDrive upload getting slower?" someone has to grep logs across seven containers.
Why this is worth doing
ScanSync is a multi-stage pipeline across separate services:
smb_service → detection_service → ocr_service → file_naming_service → upload_service
Each stage can fail independently, and a document that silently stalls between two stages is currently invisible. Event-based logging makes the funnel measurable: how many documents entered, how many completed, where they dropped out.
Privacy constraints (important)
ScanSync processes private documents — scanned invoices, contracts, medical letters. The analytics backend is a separate service and its database is not intended to hold document content.
Never send:
- file names or SMB paths
- OCR text or any document content
- OneDrive target paths or folder names
- user names, mail addresses, tenant IDs
Safe to send: durations, page counts, file sizes, booleans, enum-like states, exception types, service names, versions.
Where a correlation ID is useful, send a hash (e.g. first 12 chars of a SHA-256 over the internal job ID) rather than anything derived from the file name.
Proposed events
Dot notation (object.action), consistent with the other projects on the same backend.
| Event |
When |
Suggested metadata |
document.received |
new file detected on SMB share |
size_bytes, extension, sync_target_id |
ocr.started |
OCR job picked up |
pages, size_bytes |
ocr.completed |
OCR finished |
duration_ms, pages, language, skipped_text_layer (bool) |
ocr.failed |
OCR raised |
error_type, duration_ms, pages |
naming.completed |
AI renaming returned |
provider (openai/ollama), model, duration_ms, fallback_used (bool) |
naming.failed |
renaming failed |
provider, error_type |
upload.completed |
OneDrive upload done |
duration_ms, size_bytes, retries |
upload.failed |
upload failed |
error_type, status_code, retries |
document.completed |
whole pipeline done |
total_duration_ms, pages |
queue.depth |
periodic sample |
queue_name, depth |
document.received plus document.completed gives the completion rate for free; the *.failed events show where the rest is lost.
Lessons from the MatchMyTime integration
The same backend is already used by MatchMyTime, and its crash telemetry turned out to be nearly useless in practice (see maxi07/MatchMyTime issues 70 and 73). Three mistakes worth avoiding here:
error_type was hardcoded to the literal string ERROR for every single event. Grouping failures then requires string matching on free-form messages. Use type(exc).__name__.
- No status code on failures, so it was impossible to tell whether a failure was user-visible. Include it where an HTTP call is involved (
upload.failed).
- Expected failures and real bugs shared one event name. Invalid user input got logged as a crash and buried the genuine incidents. Keep
ocr.failed for actual faults; if a file type is simply unsupported, that is its own event or a metadata flag — not a failure.
Implementation sketch
The natural place is scansynclib, next to the existing logging.py, so every service gets it from one import.
- Add an
analytics.py module wrapping the existing AnalyticsClient (client/analytics_client.py in the analytics repo — requests + daemon threads, fire-and-forget).
- Read
ANALYTICS_URL and ANALYTICS_API_KEY from the environment. If either is unset, the module becomes a no-op. ScanSync is self-hosted by other people; telemetry must be opt-in and must never be a hard dependency.
- Failures to reach the analytics endpoint must never affect document processing — swallow and log at DEBUG.
- Pass
app_version from the existing version resolution so regressions can be attributed to a release.
- Seed the project first:
make seed NAME=scansync in the analytics repo produces the API key.
Backend API reference
POST /event
X-API-KEY: <key>
{
"event_name": "ocr.completed",
"user_id": null,
"app_version": "1.4.2",
"metadata": { "duration_ms": 8300, "pages": 4 }
}
Responses: 200 {"status":"ok"}, 400 on missing event_name, 401 on bad key.
Definition of done
Goal
Send structured events from ScanSync to the existing self-hosted analytics service (
analytics.maexbert.de, repomaxi07/maexbert-analytics) so that pipeline health, throughput and failures become visible in Grafana instead of only in container logs.Today the only observability is
loggeroutput to stdout pluslogfile.log(WARNING and above). To answer "how many documents failed OCR last week?" or "is OneDrive upload getting slower?" someone has to grep logs across seven containers.Why this is worth doing
ScanSync is a multi-stage pipeline across separate services:
smb_service→detection_service→ocr_service→file_naming_service→upload_serviceEach stage can fail independently, and a document that silently stalls between two stages is currently invisible. Event-based logging makes the funnel measurable: how many documents entered, how many completed, where they dropped out.
Privacy constraints (important)
ScanSync processes private documents — scanned invoices, contracts, medical letters. The analytics backend is a separate service and its database is not intended to hold document content.
Never send:
Safe to send: durations, page counts, file sizes, booleans, enum-like states, exception types, service names, versions.
Where a correlation ID is useful, send a hash (e.g. first 12 chars of a SHA-256 over the internal job ID) rather than anything derived from the file name.
Proposed events
Dot notation (
object.action), consistent with the other projects on the same backend.document.receivedsize_bytes,extension,sync_target_idocr.startedpages,size_bytesocr.completedduration_ms,pages,language,skipped_text_layer(bool)ocr.failederror_type,duration_ms,pagesnaming.completedprovider(openai/ollama),model,duration_ms,fallback_used(bool)naming.failedprovider,error_typeupload.completedduration_ms,size_bytes,retriesupload.failederror_type,status_code,retriesdocument.completedtotal_duration_ms,pagesqueue.depthqueue_name,depthdocument.receivedplusdocument.completedgives the completion rate for free; the*.failedevents show where the rest is lost.Lessons from the MatchMyTime integration
The same backend is already used by MatchMyTime, and its crash telemetry turned out to be nearly useless in practice (see maxi07/MatchMyTime issues 70 and 73). Three mistakes worth avoiding here:
error_typewas hardcoded to the literal stringERRORfor every single event. Grouping failures then requires string matching on free-form messages. Usetype(exc).__name__.upload.failed).ocr.failedfor actual faults; if a file type is simply unsupported, that is its own event or a metadata flag — not a failure.Implementation sketch
The natural place is
scansynclib, next to the existinglogging.py, so every service gets it from one import.analytics.pymodule wrapping the existingAnalyticsClient(client/analytics_client.pyin the analytics repo —requests+ daemon threads, fire-and-forget).ANALYTICS_URLandANALYTICS_API_KEYfrom the environment. If either is unset, the module becomes a no-op. ScanSync is self-hosted by other people; telemetry must be opt-in and must never be a hard dependency.app_versionfrom the existing version resolution so regressions can be attributed to a release.make seed NAME=scansyncin the analytics repo produces the API key.Backend API reference
Responses:
200 {"status":"ok"},400on missingevent_name,401on bad key.Definition of done
scansynclib/analytics.pyexists, is a no-op without configuration, and never raises into the callererror_typecarries the real exception class namescansyncshows: documents per day, completion rate, failures by stage, OCR duration percentiles, queue depth