Skip to content

feat(storage): implement file-backed DLR storage #340

Description

@lykakis

Parent epic: #337
Delivery step 6; follows filesystem pending admission #339 and selected-router/routed-work persistence #345. Precedes bounded in-memory DLR backend #334.

Status and goal

Deliver a database-free, single-process profile by implementing file-backed DLR storage behind the existing DlrStorage contract. Filesystem SMS processing from #339/#345 initially continues to use PostgreSQL DLR storage; this issue adds the local durable option. Provider correlations and pending downstream HTTP/SMPP deliveries must survive application/container restart while the configured persistent volume remains available.

The physical representation is intentionally open until this issue starts. An indexed spool, H2 MVStore (including lessons from the historical MvStoreDlrStorage removed in 054a92b), or another local file-backed approach may be selected after examining indexed lookups, atomic updates, native-image support and operational cost. Choosing an engine is work within #340, not a prerequisite decision in this epic.

Required behavior

  • Persist provider acceptance and rejection, indexing provider name/message ID to the gateway ID and supporting multiple provider IDs for multipart submissions.
  • Resolve terminal receipts idempotently while preserving configured intermediate/terminal DLR behavior.
  • Track pending HTTP callbacks and downstream SMPP receipts, including receipts awaiting a client bind.
  • Select/claim due deliveries without losing them, persist attempts, next-attempt time, claim token and last result, and fence stale completions/retries/failures.
  • Recover pending correlations/deliveries after restart, apply configurable retention and cleanup without silently evicting active work, and coordinate idempotently with SMS pending completion across feat(storage): implement filesystem pending-message store and ingress #339/feat(storage): persist selected-router and routed-work state on filesystem #345.

File-backed design decisions to make when starting

  • Compare at least an indexed local file engine (MVStore is one candidate) and a purpose-built spool/index design against provider-ID lookup, due-delivery scans, atomic multi-record/index updates, file count/space, restart recovery, and the expected workload. Document the choice and rejected alternatives before implementing it; do not assume the SMS file-per-message layout fits DLR data.
  • Define a versioned persisted representation and consistent publication/commit boundary: report a successful provider outcome or downstream delivery transition only after its state is safely recorded for the promised process/container restart guarantee. Incomplete writes and corrupt committed state must fail closed rather than silently disappear.
  • Use exclusive single-process ownership of the selected DLR storage path. Its namespace, configuration and lifecycle are independent of the SMS spool/stage data, even if operators use the same persistent volume.
  • Decide schema compatibility/upgrade, retention and cleanup, stopped-instance backup/restore, optional historical MVStore-file migration or documented incompatibility, and JVM/native-image support for the chosen approach.

Configuration and operations

Use a startup-only, backend-neutral selector and configured persistent path, equivalent to:

sendium.dlr.storage.backend=file
sendium.dlr.storage.file.path=/work/data/sendium-dlr

The public backend is file; MVStore, if chosen, is an implementation detail. Unsupported SMS/DLR profiles fail startup; file storage failures never fall back to memory or disabled DLR handling. Startup/readiness must fail on open, ownership, schema, write, commit/publication, capacity, permission or corruption errors.

Log the selected backend/path, version, startup recovery and cleanup summaries, and actionable underlying storage errors without message payloads or credentials. Expose the existing sendium-dlr-storage readiness check with sanitized health data (backend=file, reason=unavailable on failure); storage failures affect readiness rather than liveness. Reuse backend-selection and bounded-operation latency/error metrics without IDs or paths in metric labels. Document persistent-volume mounts and backup/restore for JVM and native containers.

Verification and acceptance criteria

  • file is selectable for DLRs without a PostgreSQL datasource; forced process/container restart restores provider correlations, pending HTTP callbacks, downstream SMPP receipts, retry schedules and claim fencing on a retained volume.
  • Multipart provider IDs resolve correctly; intermediate receipts do not prematurely consume correlation; duplicate terminal receipts and stale claim tokens are handled idempotently/correctly under the shared DlrStorage semantics.
  • Restart/fault-injection tests prove primary records and necessary lookup/due indexes remain consistent at every committed transition; uncommitted outcomes are not reported as successful.
  • Retention keeps active correlations and deliveries; failed publication/commit, full/corrupt/locked/unwritable storage, and second-process ownership fail closed with no memory fallback.
  • HTTP/SMPP ingress and SMS-to-DLR handoff tests verify no pending SMS record is removed before the required DLR transition succeeds.
  • The chosen approach passes JVM/native restart tests; documentation covers schema upgrades, retention, repair/backup, persistent volume, and any incompatibility with historical MVStore data.

Open decisions

  • Physical representation (indexed spool, MVStore, or another file-backed approach) and associated indexes/commit boundary.
  • Retention defaults, cleanup/compaction strategy, historical data handling, backup/upgrade strategy, and exact handoff recovery rules with the SMS store.

Related issues and non-goals

Related: architecture/memory baseline #338; filesystem SMS pending #339; filesystem SMS stages #345; bounded in-memory DLR #334; PostgreSQL outbound #333.

No multi-process access to the same local DLR store, exactly-once HTTP/SMPP downstream delivery, replacement of the existing PostgreSQL DLR backend for deployments selecting it, SMS body/outbound stage storage, or automatic fallback to memory.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions