You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Sendium can acknowledge an outbound SMS while the work needed to deliver it exists only in memory. A restart may lose that message. Storing the accepted message solves that loss, but it does not preserve an expensive selection from a large backlog or the destinations already chosen by routing.
Delivery-report (DLR) processing has a separate problem: provider correlations and pending HTTP/SMPP receipts also need a storage choice. SMS durability alone does not make DLR delivery durable.
Goal
Let operators choose which parts of outbound processing survive a restart: accepted messages, messages selected for routing, and work assigned to destinations. Provide memory, file-backed, and eventually PostgreSQL implementations where supported, with a clear recovery guarantee for each configuration.
Keep DLR storage independently configurable. Start with the existing PostgreSQL DLR backend, then add file-backed and bounded in-memory options. Unsupported configurations must fail startup rather than silently use a less durable backend. Provider submission remains at least once.
Storage model and configuration
The three globally selected, startup-only outbound storage stages are configured as follows:
pending holds the accepted source message/part until all required processing is terminal. A bounded select-and-stage operation records the chosen router batch; a routed-work transition records every selected destination (including copied routes) and completion of each destination. Router and worker in-memory queues are bounded scheduling projections, not the only source of work promised by the configured durability level. A router thread's in-memory dequeue does not require its own persisted stage in the single-process design.
A durable pending store admits messages or multipart parts before HTTP 202/SMPP submit_sm_resp STATUS_OK; admission failure rejects/backpressures, never silently falls back to memory. Pending records are removed only after all necessary destinations/provider parts reach terminal processing and any required DLR handoff succeeds. Processing remains at least once: a crash after provider submission but before completion can cause duplicate provider delivery.
Stage guarantees and supported profiles
Pending / selected router / routed work
Process/container restart behavior
memory/memory/memory
Explicit non-durable compatibility mode
file/memory/memory
Recover accepted work, rerun selection and routing
file/file/memory
Recover the selected batch without rerunning its expensive selection; reroute
file/memory/file
Reselect not-yet-routed messages; restore recorded worker destinations
file/file/file
Restore the selected batch and recorded destinations
#338 first delivers only memory/memory/memory; requesting any unimplemented profile (for example file/file/memory) fails startup with the supported choices rather than falling back to memory. The file backend then delivers these profiles in stages; memory/memory/memory remains the default for all three outbound selectors when configuration is omitted, including after file-backed stage state is implemented. Durable file/PostgreSQL profiles require explicit configuration; the default is non-durable across restarts. PostgreSQL will be added by #333 with its own explicitly documented supported combinations. This is an eventual backend vocabulary, not a promise that all 27 combinations are implemented. Startup must reject unsupported or unsafe mixes. Backend changes with outstanding durable records require a defined drain/migration policy; no silent data abandonment or guarantee downgrade.
For any combination, keep the accepted source recoverable until downstream work is safe. If selected work is durable but routed work is memory-backed, retain sufficient selected state to reroute after restart. If selection is memory-backed but routed work is durable, recovery avoids reselecting already-routed messages. Interruptions between backends require reconciliation by stable ID; do not assume cross-backend atomic transactions. Selected router batches are loaded into memory incrementally; a large pending backlog is never loaded at once. Stage semantics and eligibility/order rules are fixed in #338/#344.
These restart guarantees require surviving storage (a persistent spool volume for files). They do not promise host/power-loss recovery, exact retry timing/order, or exactly-once provider submission.
DLR storage and SMS-to-DLR handoff
Outbound SMS storage and DLR storage are independent boundaries. Outbound stages retain accepted SMS work through provider processing; the selected DLR backend retains provider-message correlations and pending downstream HTTP callbacks/SMPP receipts after the SMS has completed. Durable SMS storage alone does not make subsequent DLR delivery durable.
The DLR selector is separate from the three outbound selectors. Its planned options, delivered by their respective issues, are postgresql (existing), file (#340), memory (#334), and disabled where Sendium-owned DLR handling is not requested/supported. Only documented SMS/DLR combinations are valid; unsupported choices fail startup instead of silently weakening guarantees.
DLR profile
Responsibility and restart behavior
Existing PostgreSQL DLR
Initial backend for provider correlation and pending downstream HTTP/SMPP DLR delivery; available with the first filesystem SMS milestone.
File-backed DLR on a persistent local volume (#340)
Completes a database-free, single-process filesystem deployment: provider correlations and pending downstream deliveries survive restart. Uses its own persistent namespace and lifecycle, separate from SMS pending/stage state; its physical representation is selected while working on #340.
Correlates and delivers DLRs while running, but correlations, retries and undelivered receipts disappear on restart. DLR-requesting admissions must respect its capacity/reservation policy.
Disabled DLR handling
No Sendium-owned DLR correlation or downstream delivery; documented separately from an active memory backend.
Keep the outbound pending source until a terminal provider outcome and any required DlrStorage acceptance/rejection handoff succeeds. A retryable DLR storage failure retains the source and backpressures/retries the handoff; it must not silently complete the SMS record.
Provider outcome is not separately checkpointed in the first outbound stage contract. If Sendium restarts between provider submission and a successful DLR handoff/source completion, it may submit the original SMS again; DLR handoff must handle repeated outcomes idempotently where possible.
Once an SMS source is completed, the DLR backend alone owns outstanding provider correlation and pending downstream delivery. With memory DLR, a restart after that point can lose the DLR even though the SMS itself was durable.
Close this epic after supported-profile integration, restart tests and documentation are complete.
Delivery notes
Each implementation issue owns focused failure-boundary tests, readiness/logging/operation metrics, configuration validation, deployment changes where applicable, and documentation. The existing PostgreSQL DlrStorage remains the initial provider-correlation and downstream HTTP/SMPP DLR backend; outbound stage selection does not automatically change DLR storage.
Invariants to verify
Accepted payload, routing inputs and DLR return metadata are recoverable in durable pending profiles; every acknowledged multipart part survives and groups reconstruct or expire by saved absolute deadlines.
File-backed selected batches survive restart without repeating selection; file-backed routed destinations survive without rerouting completed branches. Memory-backed stages explicitly repeat the corresponding work.
Partial batch publication, router dequeue, copied-route fan-out, provider multipart, worker processing and interrupted stage handoff never make accepted work disappear. A source is completed only when every required branch and DLR handoff is terminal.
Forced restarts at provider outcome and SMS-to-DLR handoff preserve accepted SMS work or the committed DLR state as the selected profiles promise; repeated provider outcomes do not silently discard correlation.
File-backed DLR recovery restores correlations and pending HTTP/SMPP deliveries; memory DLR restart loss and capacity rejection are explicit, while a disabled DLR backend provides no Sendium-owned DLR handling.
Storage corruption/capacity/lock failures fail closed; readiness reflects selected SMS and DLR backends; no silent fallback or unsupported selection.
Forced process/container restart tests cover every supported profile and provider-send ambiguity. Logs and metrics identify selected backends without exposing SMS bodies or high-cardinality IDs.
Kannel comparison
Kannel's UUID file spool recovers accepted messages into runtime queues but does not preserve router/worker selection or processing stage. Sendium's file/memory/memory profile has comparable reroute-on-restart semantics with admission and multipart improvements. Its file/file/file profile additionally preserves selected batches and recorded destinations. Both remain at least once; durable stage state does not eliminate duplicate provider submission.
Durable outbound retry schedules/counters, provider-outcome checkpoints, deterministic global queue order, durable failed-work UI, online backup, and multi-process local-spool/active-active ownership.
StandardMessage Java serialization, exactly-once provider submission or downstream DLR delivery, replacing PostgreSQL DLR storage in the first SMS milestone, or recovery after lost/corrupt storage.
Superseded investigation: #341 (an embedded transactional engine is not a prerequisite for the initial pending-file spool; #345 chooses its durable stage representation separately).
Problem
Sendium can acknowledge an outbound SMS while the work needed to deliver it exists only in memory. A restart may lose that message. Storing the accepted message solves that loss, but it does not preserve an expensive selection from a large backlog or the destinations already chosen by routing.
Delivery-report (DLR) processing has a separate problem: provider correlations and pending HTTP/SMPP receipts also need a storage choice. SMS durability alone does not make DLR delivery durable.
Goal
Let operators choose which parts of outbound processing survive a restart: accepted messages, messages selected for routing, and work assigned to destinations. Provide memory, file-backed, and eventually PostgreSQL implementations where supported, with a clear recovery guarantee for each configuration.
Keep DLR storage independently configurable. Start with the existing PostgreSQL DLR backend, then add file-backed and bounded in-memory options. Unsupported configurations must fail startup rather than silently use a less durable backend. Provider submission remains at least once.
Storage model and configuration
The three globally selected, startup-only outbound storage stages are configured as follows:
pendingholds the accepted source message/part until all required processing is terminal. A bounded select-and-stage operation records the chosen router batch; a routed-work transition records every selected destination (including copied routes) and completion of each destination. Router and worker in-memory queues are bounded scheduling projections, not the only source of work promised by the configured durability level. A router thread's in-memory dequeue does not require its own persisted stage in the single-process design.A durable pending store admits messages or multipart parts before HTTP
202/SMPPsubmit_sm_resp STATUS_OK; admission failure rejects/backpressures, never silently falls back to memory. Pending records are removed only after all necessary destinations/provider parts reach terminal processing and any required DLR handoff succeeds. Processing remains at least once: a crash after provider submission but before completion can cause duplicate provider delivery.Stage guarantees and supported profiles
memory/memory/memoryfile/memory/memoryfile/file/memoryfile/memory/filefile/file/file#338 first delivers only
memory/memory/memory; requesting any unimplemented profile (for examplefile/file/memory) fails startup with the supported choices rather than falling back to memory. The file backend then delivers these profiles in stages;memory/memory/memoryremains the default for all three outbound selectors when configuration is omitted, including after file-backed stage state is implemented. Durable file/PostgreSQL profiles require explicit configuration; the default is non-durable across restarts. PostgreSQL will be added by #333 with its own explicitly documented supported combinations. This is an eventual backend vocabulary, not a promise that all 27 combinations are implemented. Startup must reject unsupported or unsafe mixes. Backend changes with outstanding durable records require a defined drain/migration policy; no silent data abandonment or guarantee downgrade.For any combination, keep the accepted source recoverable until downstream work is safe. If selected work is durable but routed work is memory-backed, retain sufficient selected state to reroute after restart. If selection is memory-backed but routed work is durable, recovery avoids reselecting already-routed messages. Interruptions between backends require reconciliation by stable ID; do not assume cross-backend atomic transactions. Selected router batches are loaded into memory incrementally; a large pending backlog is never loaded at once. Stage semantics and eligibility/order rules are fixed in #338/#344.
These restart guarantees require surviving storage (a persistent spool volume for files). They do not promise host/power-loss recovery, exact retry timing/order, or exactly-once provider submission.
DLR storage and SMS-to-DLR handoff
Outbound SMS storage and DLR storage are independent boundaries. Outbound stages retain accepted SMS work through provider processing; the selected DLR backend retains provider-message correlations and pending downstream HTTP callbacks/SMPP receipts after the SMS has completed. Durable SMS storage alone does not make subsequent DLR delivery durable.
The DLR selector is separate from the three outbound selectors. Its planned options, delivered by their respective issues, are
postgresql(existing),file(#340),memory(#334), anddisabledwhere Sendium-owned DLR handling is not requested/supported. Only documented SMS/DLR combinations are valid; unsupported choices fail startup instead of silently weakening guarantees.DlrStorageacceptance/rejection handoff succeeds. A retryable DLR storage failure retains the source and backpressures/retries the handoff; it must not silently complete the SMS record.Delivery sequence
Complete and review each issue before starting the next; prerequisites and the first supported profile are explicit:
memory/memory/memorybaseline; reject unavailable file/PostgreSQL profiles at startup.file/memory/memoryafter feat(storage): define selected-router and routed-work contracts #344; retain PostgreSQL DLR handling.memory/memory/memorydefault; retain PostgreSQL DLR handling.Delivery notes
Each implementation issue owns focused failure-boundary tests, readiness/logging/operation metrics, configuration validation, deployment changes where applicable, and documentation. The existing PostgreSQL
DlrStorageremains the initial provider-correlation and downstream HTTP/SMPP DLR backend; outbound stage selection does not automatically change DLR storage.Invariants to verify
Kannel comparison
Kannel's UUID file spool recovers accepted messages into runtime queues but does not preserve router/worker selection or processing stage. Sendium's
file/memory/memoryprofile has comparable reroute-on-restart semantics with admission and multipart improvements. Itsfile/file/fileprofile additionally preserves selected batches and recorded destinations. Both remain at least once; durable stage state does not eliminate duplicate provider submission.References: Kannel spool save/load, startup reconstruction.
Deferred / non-goals
StandardMessageJava serialization, exactly-once provider submission or downstream DLR delivery, replacing PostgreSQL DLR storage in the first SMS milestone, or recovery after lost/corrupt storage.Superseded investigation: #341 (an embedded transactional engine is not a prerequisite for the initial pending-file spool; #345 chooses its durable stage representation separately).