Skip to content

feat(storage): implement filesystem pending-message store and ingress #339

Description

@lykakis

Parent epic: #337
Delivery step 4; depends on the architecture, interfaces and working memory baseline #338, refined pending contract/CBOR #343, and stage contract #344. Followed by file-backed stages #345.

Goal

Implement the filesystem pending-message backend and wire durable admission into HTTP/SMPP ingress. Deliver the incremental file/memory/memory profile: acknowledged messages survive process/container restart, but selected work is reselected and routed work is rerouted. #345 adds opt-in file-backed selection/routing stages. memory/memory/memory remains the non-durable default; durable profiles require explicit configuration.

Existing PostgreSQL storage continues to handle provider correlations and downstream DLR delivery. This issue does not implement file-backed router/routed state.

File layout and admission

  • Use stable UUIDs and bounded hash buckets for complete messages and multipart parts; keep committed files distinguishable from temporary publications. Final layout is an implementation choice.
  • Validate and encode the versioned CBOR envelope from feat(storage): define the outbound pending-message contract and CBOR envelope #343, write to a unique temporary file on the same filesystem, close it, and atomically publish without replacing an existing committed record. Acknowledge only after publication succeeds; report capacity/write/codec/lock failures rather than falling back to memory.
  • Enforce single-process ownership of the spool root; require a surviving persistent volume for container replacement. Host power-loss/fsync-level guarantees are not claimed.
  • Recover committed work before opening protocol admission. Ignore/clean incomplete temporary files under documented rules; fail readiness on unknown-version or corrupt committed records, identifying the record in logs.

Ingress and lifecycle

  • HTTP 202 and SMPP submit_sm_resp STATUS_OK follow successful store admission; storage failure returns a suitable retryable response. Preserve the existing semantic of duplicate multipart ordinal acknowledgements.
  • Persist every SMPP multipart part, the group identity and absolute deadline, reconstruct incomplete/complete groups on restart, and retain source records until aggregate processing is terminal. Expired incomplete groups release accepted parts for independent routing.
  • Keep the pending source during routing, retries, asynchronous provider work, and existing PostgreSQL DLR handoff. Terminal drops/rejections and successes complete it only after all copied branches/provider parts and required DLR work are done.
  • With memory-backed selected/routed stages, recovery selects and routes surviving pending work again. A crash after provider submission can duplicate delivery; source completion must not falsely claim an uncommitted DLR handoff.
  • Existing runtime queues are bounded projections; do not materialize an entire large spool in memory. Coordinate with feat(storage): persist selected-router and routed-work state on filesystem #345's future stage metadata so recovery does not double-schedule items already represented by durable later stages.

Configuration, deployment, observability

Verification and acceptance criteria

  • Contract/codec tests from feat(storage): define the outbound pending-message contract and CBOR envelope #343 plus HTTP/SMPP commit-before-ack and storage failure tests.
  • Forced process/container restart tests for admission, router/worker processing and provider-submission ambiguity under file/memory/memory.
  • Multipart duplicate, timeout, reconstruction, copied-route/provider-multipart completion, and failed DLR handoff tests.
  • Locked/full/unwritable spool, failed move, temporary/truncated/corrupt/unknown-version files, JVM and native smoke tests.
  • Documentation clearly states accepted-message durability, reselection/rerouting on restart, volume limitations, and at-least-once provider behavior.

Out of scope

File-backed router/routed state (#345), PostgreSQL outbound backend (#333), DLR backend replacement (#340/#334), durable retry scheduling/provider outcome, multi-process spool sharing, automatic memory fallback, and exactly-once provider delivery.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions