Skip to content

epic: introduce pluggable storage backends for SMS durability and DLR handling #337

Description

@lykakis

Problem

Sendium can acknowledge an outbound SMS while the work needed to deliver it exists only in memory. A restart may lose that message. Storing the accepted message solves that loss, but it does not preserve an expensive selection from a large backlog or the destinations already chosen by routing.

Delivery-report (DLR) processing has a separate problem: provider correlations and pending HTTP/SMPP receipts also need a storage choice. SMS durability alone does not make DLR delivery durable.

Goal

Let operators choose which parts of outbound processing survive a restart: accepted messages, messages selected for routing, and work assigned to destinations. Provide memory, file-backed, and eventually PostgreSQL implementations where supported, with a clear recovery guarantee for each configuration.

Keep DLR storage independently configurable. Start with the existing PostgreSQL DLR backend, then add file-backed and bounded in-memory options. Unsupported configurations must fail startup rather than silently use a less durable backend. Provider submission remains at least once.

Storage model and configuration

The three globally selected, startup-only outbound storage stages are configured as follows:

sendium.sms.pending.backend=memory|file|postgresql
sendium.sms.router-queue.backend=memory|file|postgresql
sendium.sms.routed-work.backend=memory|file|postgresql

pending holds the accepted source message/part until all required processing is terminal. A bounded select-and-stage operation records the chosen router batch; a routed-work transition records every selected destination (including copied routes) and completion of each destination. Router and worker in-memory queues are bounded scheduling projections, not the only source of work promised by the configured durability level. A router thread's in-memory dequeue does not require its own persisted stage in the single-process design.

A durable pending store admits messages or multipart parts before HTTP 202/SMPP submit_sm_resp STATUS_OK; admission failure rejects/backpressures, never silently falls back to memory. Pending records are removed only after all necessary destinations/provider parts reach terminal processing and any required DLR handoff succeeds. Processing remains at least once: a crash after provider submission but before completion can cause duplicate provider delivery.

Stage guarantees and supported profiles

Pending / selected router / routed work Process/container restart behavior
memory/memory/memory Explicit non-durable compatibility mode
file/memory/memory Recover accepted work, rerun selection and routing
file/file/memory Recover the selected batch without rerunning its expensive selection; reroute
file/memory/file Reselect not-yet-routed messages; restore recorded worker destinations
file/file/file Restore the selected batch and recorded destinations

#338 first delivers only memory/memory/memory; requesting any unimplemented profile (for example file/file/memory) fails startup with the supported choices rather than falling back to memory. The file backend then delivers these profiles in stages; memory/memory/memory remains the default for all three outbound selectors when configuration is omitted, including after file-backed stage state is implemented. Durable file/PostgreSQL profiles require explicit configuration; the default is non-durable across restarts. PostgreSQL will be added by #333 with its own explicitly documented supported combinations. This is an eventual backend vocabulary, not a promise that all 27 combinations are implemented. Startup must reject unsupported or unsafe mixes. Backend changes with outstanding durable records require a defined drain/migration policy; no silent data abandonment or guarantee downgrade.

For any combination, keep the accepted source recoverable until downstream work is safe. If selected work is durable but routed work is memory-backed, retain sufficient selected state to reroute after restart. If selection is memory-backed but routed work is durable, recovery avoids reselecting already-routed messages. Interruptions between backends require reconciliation by stable ID; do not assume cross-backend atomic transactions. Selected router batches are loaded into memory incrementally; a large pending backlog is never loaded at once. Stage semantics and eligibility/order rules are fixed in #338/#344.

These restart guarantees require surviving storage (a persistent spool volume for files). They do not promise host/power-loss recovery, exact retry timing/order, or exactly-once provider submission.

DLR storage and SMS-to-DLR handoff

Outbound SMS storage and DLR storage are independent boundaries. Outbound stages retain accepted SMS work through provider processing; the selected DLR backend retains provider-message correlations and pending downstream HTTP callbacks/SMPP receipts after the SMS has completed. Durable SMS storage alone does not make subsequent DLR delivery durable.

The DLR selector is separate from the three outbound selectors. Its planned options, delivered by their respective issues, are postgresql (existing), file (#340), memory (#334), and disabled where Sendium-owned DLR handling is not requested/supported. Only documented SMS/DLR combinations are valid; unsupported choices fail startup instead of silently weakening guarantees.

DLR profile Responsibility and restart behavior
Existing PostgreSQL DLR Initial backend for provider correlation and pending downstream HTTP/SMPP DLR delivery; available with the first filesystem SMS milestone.
File-backed DLR on a persistent local volume (#340) Completes a database-free, single-process filesystem deployment: provider correlations and pending downstream deliveries survive restart. Uses its own persistent namespace and lifecycle, separate from SMS pending/stage state; its physical representation is selected while working on #340.
Bounded in-memory DLR (#334) Correlates and delivers DLRs while running, but correlations, retries and undelivered receipts disappear on restart. DLR-requesting admissions must respect its capacity/reservation policy.
Disabled DLR handling No Sendium-owned DLR correlation or downstream delivery; documented separately from an active memory backend.
  • Keep the outbound pending source until a terminal provider outcome and any required DlrStorage acceptance/rejection handoff succeeds. A retryable DLR storage failure retains the source and backpressures/retries the handoff; it must not silently complete the SMS record.
  • Provider outcome is not separately checkpointed in the first outbound stage contract. If Sendium restarts between provider submission and a successful DLR handoff/source completion, it may submit the original SMS again; DLR handoff must handle repeated outcomes idempotently where possible.
  • Once an SMS source is completed, the DLR backend alone owns outstanding provider correlation and pending downstream delivery. With memory DLR, a restart after that point can lose the DLR even though the SMS itself was durable.
  • feat(storage): implement file-backed DLR storage #340 implements indexed provider lookup, multipart provider IDs, pending delivery scans, retry/claim state, retention and restart recovery using a file-backed approach selected during that issue (for example, an indexed spool or MVStore), rather than assuming the simple SMS file-per-message layout fits DLRs. feat(storage): implement bounded in-memory DLR storage #334 implements the same active DLR behavior within explicit non-durable capacity/retention limits. Neither is part of the initial feat(storage): implement filesystem pending-message store and ingress #339/feat(storage): persist selected-router and routed-work state on filesystem #345 SMS backend work.

Delivery sequence

Complete and review each issue before starting the next; prerequisites and the first supported profile are explicit:

  1. feat(storage): define outbound stage abstractions and memory baseline #338 - Define the stage abstractions and wire a working memory/memory/memory baseline; reject unavailable file/PostgreSQL profiles at startup.
  2. feat(storage): define the outbound pending-message contract and CBOR envelope #343 - Refine the pending-message contract and versioned CBOR envelope after feat(storage): define outbound stage abstractions and memory baseline #338.
  3. feat(storage): define selected-router and routed-work contracts #344 - Refine selected-router and routed-work contracts after feat(storage): define the outbound pending-message contract and CBOR envelope #343.
  4. feat(storage): implement filesystem pending-message store and ingress #339 - Implement filesystem pending admission and file/memory/memory after feat(storage): define selected-router and routed-work contracts #344; retain PostgreSQL DLR handling.
  5. feat(storage): persist selected-router and routed-work state on filesystem #345 - Implement opt-in filesystem router/routed state after feat(storage): implement filesystem pending-message store and ingress #339; retain the memory/memory/memory default; retain PostgreSQL DLR handling.
  6. feat(storage): implement file-backed DLR storage #340 - Implement file-backed DLR storage and the database-free SMS + DLR profile after feat(storage): persist selected-router and routed-work state on filesystem #345; choose spool, MVStore or another physical approach during feat(storage): implement file-backed DLR storage #340.
  7. feat(storage): implement bounded in-memory DLR storage #334 - Implement bounded in-memory DLR handling after feat(storage): implement file-backed DLR storage #340, distinct from disabling DLRs.
  8. feat(storage): implement PostgreSQL outbound stage backends #333 - Implement supported PostgreSQL outbound-stage profiles after feat(storage): implement bounded in-memory DLR storage #334 without replacing the selected DLR backend.
  9. Close this epic after supported-profile integration, restart tests and documentation are complete.

Delivery notes

Each implementation issue owns focused failure-boundary tests, readiness/logging/operation metrics, configuration validation, deployment changes where applicable, and documentation. The existing PostgreSQL DlrStorage remains the initial provider-correlation and downstream HTTP/SMPP DLR backend; outbound stage selection does not automatically change DLR storage.

Invariants to verify

  • Accepted payload, routing inputs and DLR return metadata are recoverable in durable pending profiles; every acknowledged multipart part survives and groups reconstruct or expire by saved absolute deadlines.
  • File-backed selected batches survive restart without repeating selection; file-backed routed destinations survive without rerouting completed branches. Memory-backed stages explicitly repeat the corresponding work.
  • Partial batch publication, router dequeue, copied-route fan-out, provider multipart, worker processing and interrupted stage handoff never make accepted work disappear. A source is completed only when every required branch and DLR handoff is terminal.
  • Forced restarts at provider outcome and SMS-to-DLR handoff preserve accepted SMS work or the committed DLR state as the selected profiles promise; repeated provider outcomes do not silently discard correlation.
  • File-backed DLR recovery restores correlations and pending HTTP/SMPP deliveries; memory DLR restart loss and capacity rejection are explicit, while a disabled DLR backend provides no Sendium-owned DLR handling.
  • Storage corruption/capacity/lock failures fail closed; readiness reflects selected SMS and DLR backends; no silent fallback or unsupported selection.
  • Forced process/container restart tests cover every supported profile and provider-send ambiguity. Logs and metrics identify selected backends without exposing SMS bodies or high-cardinality IDs.

Kannel comparison

Kannel's UUID file spool recovers accepted messages into runtime queues but does not preserve router/worker selection or processing stage. Sendium's file/memory/memory profile has comparable reroute-on-restart semantics with admission and multipart improvements. Its file/file/file profile additionally preserves selected batches and recorded destinations. Both remain at least once; durable stage state does not eliminate duplicate provider submission.

References: Kannel spool save/load, startup reconstruction.

Deferred / non-goals

  • PostgreSQL profiles beyond those explicitly implemented in feat(storage): implement PostgreSQL outbound stage backends #333; arbitrary file/PostgreSQL stage mixing or online cross-backend migration, including DLR backend migrations.
  • Durable outbound retry schedules/counters, provider-outcome checkpoints, deterministic global queue order, durable failed-work UI, online backup, and multi-process local-spool/active-active ownership.
  • StandardMessage Java serialization, exactly-once provider submission or downstream DLR delivery, replacing PostgreSQL DLR storage in the first SMS milestone, or recovery after lost/corrupt storage.

Superseded investigation: #341 (an embedded transactional engine is not a prerequisite for the initial pending-file spool; #345 chooses its durable stage representation separately).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions