Skip to content

feat(health): add /health/live and /health/ready probes with dependency checks - #362

Merged
Cjay-Cyber-2 merged 1 commit into
ASTROIDX556:mainfrom
AdaBliss:feat/351-health-live-ready-probes
Sep 29, 2026
Merged

Cjay-Cyber-2 merged 1 commit into
ASTROIDX556:mainfrom
AdaBliss:feat/351-health-live-ready-probes

Conversation

@AdaBliss

Copy link
Copy Markdown
Contributor

closes #351

Summary

Adds orchestrator-grade liveness and readiness probes and makes the health controller usable by load balancers.

  • GET /health/live returns 200 whenever the process is running. It performs no dependency checks, so a database or cache outage never causes an otherwise healthy process to be restarted.
  • GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with per-dependency status, latency and error detail:
{
  "status": "down",
  "timestamp": "2026-09-28T10:00:00.000Z",
  "services": {
    "database": { "status": "down", "latencyMs": 2001, "timestamp": "...", "error": "Database health check timed out after 2000ms" },
    "cache": { "status": "up", "latencyMs": 1, "timestamp": "..." }
  }
}
  • Both probes are served outside the global API prefix (like /metrics), so probe paths stay stable across API versions. The existing diagnostic routes (/{API_PREFIX}/health, /readiness, /liveness, /database) are unchanged.

Fixes found along the way

  • Health routes required authentication. HealthController was not @Public(), so every probe got a 401 from the global JwtAuthGuard. It is now public.
  • Probes could be rate limited. They are now exempt from both named throttler tiers (@SkipThrottle({ api: true, auth: true }); a bare @SkipThrottle() only skips the default throttler).
  • Probes wrote audit rows. The global AuditInterceptor persists every request, including GETs, so each probe wrote an audit row and tried to during a database outage. The controller is now @SkipAudit().
  • The Redis probe checked the wrong server. RedisHealthIndicator read an undefined REDIS_URL and always probed localhost:6379. It also held its own never-closed connection and could hang while ioredis queued the PING. It now probes the shared REDIS_CLIENT built from the validated REDIS_* config, with a timeout.

Testing

  • Unit tests for /live and /ready: all up (200), simulated database outage (503), simulated cache outage (503), empty indicator result (503). Also checked: an external Stellar outage cannot fail readiness, and liveness stays 200 during a database outage.
  • health.http.spec.ts boots the controller behind the real JwtAuthGuard with the same prefix exclusions as main.ts, and calls it over HTTP. It verifies the probes are public, unprefixed, and return 503 with structured JSON during a simulated outage. Removing @Public() makes it fail (checked).
  • Redis indicator tests: PONG, rejection, unexpected reply, closed client, and a timeout while the command is queued.
  • npm run typecheck, npm run lint, npm test (1095 passing), npm run build.

Notes for reviewers

Open PRs #354 and #358 also edit health.controller.ts. The changes are additive, but whichever merges second may need a small rebase.

…cy checks

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes ASTROIDX556#351
@drips-wave

drips-wave Bot commented Sep 28, 2026

Copy link
Copy Markdown

@AdaBliss Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits.

You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀

Learn more about application limits

@Cjay-Cyber-2
Cjay-Cyber-2 merged commit a171dc9 into ASTROIDX556:main Sep 29, 2026
1 check passed
dakwa001 pushed a commit to dakwa001/astroid-api that referenced this pull request Oct 2, 2026
)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes ASTROIDX556#351
tecch-wiz pushed a commit to tecch-wiz/astroid-api that referenced this pull request Oct 2, 2026
)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes ASTROIDX556#351
Cjay-Cyber-2 added a commit that referenced this pull request Oct 7, 2026
* feat: address requested issues - closes #264, closes #263, closes #262, closes #261

* feat: Add comprehensive observability, security, and simulation improvements

Implements four major feature requests for enhanced API observability,
security, and transaction simulation capabilities:

#247 - Prometheus Metrics Interceptor
- Create MetricsInterceptor for HTTP request metrics collection
- Add request duration histograms, active request gauges, and request counters
- Categorize metrics by route, method, and status code
- Exclude /metrics endpoint from self-instrumentation
- Add comprehensive unit tests (9 tests)

#246 - Stellar Transaction Simulation Service
- Create StellarSimulationService for transaction simulation
- Integrate with Soroban RPC for XDR validation and simulation
- Include risk assessment and fee estimation
- Add circuit breaker protection for RPC failures
- Add XDR validation helper method
- Add comprehensive unit tests with mocked Stellar RPC (24 tests)

#245 - Cryptographic API Key Hashing Upgrade
- Upgrade from SHA-256 to Argon2id for enhanced security
- Implement memory-hard algorithm resistant to GPU/ASIC attacks
- Add timing-attack resistant comparison via Argon2 verification
- Maintain SHA-256 fallback for backward compatibility
- Update ApiKeyService with dual-algorithm verification
- Add comprehensive unit tests (32 crypto tests, 15 API key tests)

#244 - Webhook Retry and Dead Letter Queue
- Verify existing implementation meets all requirements
- Confirm exponential backoff with jitter (2000ms base, 20% jitter)
- Confirm 5 max attempts and non-transient error detection
- Verify dead-letter handler via DeadLetterService
- All existing tests passing

Closes #247, #246, #245, #244

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(workers): centralize job error handling with scrubbed structured logging

Add runWorkerJob, a wrapper every background worker now routes its handler
through. It times the job via WorkerMetricsService when available, logs a
structured completion record, and classifies failures before rethrowing:
transient failures log a job.retrying warning, while a failure on the final
attempt or an UnrecoverableError logs a job.dead-lettered error with the
scrubbed payload and stack. The original error is always rethrown untouched
so BullMQ retry semantics are preserved, and logging can never mask it.

Add scrubForLog/scrubString, which redact sensitive keys (sharing the audit
sanitizer's key list) and secret-shaped substrings such as Stellar seeds,
bearer tokens and URL credentials, and coerce cycles, bigints and errors
into JSON-safe values.

QueueFailureListener now scrubs its log line as well; previously the raw
webhook payload, including its signing secret, was written to the error
log. The dead-letter copy keeps the raw payload so re-drive still works.

* feat(pagination): add offset/limit pagination to list endpoints

Extend the shared list query contract used by every list endpoint with
an offset parameter alongside the existing page parameter, and bound
the page size.

- limit defaults to 50 and is capped at 200
- offset defaults to 0; page remains supported as an alternative, and
  supplying both is rejected
- negative, non-integer, non-numeric or out-of-range values are
  rejected by the Zod pipe with 400 Bad Request
- Prisma queries use bound skip/take values and an allow-listed sort
  column, so no pagination input reaches SQL as text
- meta now includes offset, and hasNext/hasPrev are computed from the
  offset so unaligned slices report correctly
- paginated responses set an X-Total-Count header, exposed via CORS
- replace duplicated Swagger page/limit docs with ApiPaginationQuery
- add unit tests and an HTTP-level integration test

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping

* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)

* perf(analytics): batch dashboard overview queries into one transaction

Closes #338

AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.

* test(audit): cover pagination and sorting edge cases for activity log

Closes #339

The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.

* fix(migrations): resolve colliding timestamp between two merged migrations

Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.

* fix(ci): document missing env vars and fix flaky retry.util test

Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:

- docs/configuration.md was missing 7 env vars added by recent merges
  (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
  DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
  env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
  fake timers with vi.runAllTimersAsync(), then attach the rejection
  assertion afterward — a race that surfaces as an unhandled rejection
  under full-suite load (deterministic once >100 files run together).
  Attaching a no-op .catch() immediately after creating the promise
  prevents the unhandled state without changing what each test asserts.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)

* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)

* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)

* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)

* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)

* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)

* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)

* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)

* Add startup migration check and DB pool connection metrics

Wires the existing migration status checker into app bootstrap so the
process halts before accepting traffic when prisma/migrations has
pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt,
the default; 'warn' logs and continues). Gated behind
DATABASE_MIGRATION_CHECK_ENABLED.

Adds a db_pool_connections Prometheus gauge (active/idle/waiting)
sourced from pg_stat_activity, since Prisma's Rust query engine
doesn't expose pool internals through the Node client.

Closes #335
Closes #332

* Fix CI: document missing env vars, fix flaky retry backoff test

- docs/configuration.md was missing entries for DATABASE_SLOW_QUERY_THRESHOLD_MS,
  DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, and the
  PUBLIC_RATE_LIMIT_* vars, failing the configuration-documentation test.
- retry.util.spec.ts left three rejected promises unhandled between
  `runAllTimersAsync()` and the `expect(...).rejects` assertion that
  attaches the handler; attach a no-op .catch() immediately after creating
  each promise so fake-timer-driven rejections don't fire as unhandled
  rejections mid-test-run.

* Fix pre-existing typecheck/lint/test failures blocking CI

main's build/typecheck/lint/test were already broken before this branch
touched anything (confirmed by checking out upstream/main directly).
CI enforces these repo-wide, so they block this PR too. Fixed each:

- event-names.ts: duplicate object key (TransactionRiskScoringRequested)
- throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't
  exist on that type (JWT payload field name leaked into the wrong type)
- sliding-window-throttler.guard.ts: removed a user-tier rate-limit
  multiplier keyed on AuthenticatedUser.tier, a field never present
  anywhere in the auth/user model — dead, unbacked logic. Removed its
  now-orphaned tests too.
- agent.controller.ts: AstroidThrottlerGuard used but never imported;
  dropped an unused SlidingWindowThrottlerGuard import block
- risk.service.ts: event-driven risk scoring built a RiskFactorsInput
  with fields (destination/velocityCount/isNewRecipient) that don't
  exist on the current type; mapped to the real shape instead
- risk.service.spec.ts: removed orphaned unused fixtures
- stellar.service.ts (src/modules/stellar/services, unused elsewhere
  in the app but still typechecked/tested): getTransactionInfo called
  a client method that doesn't exist (real method is getTransaction);
  simulateTransaction passed a bare string where the client expects
  an options object
- stellar.service.spec.ts: rewritten against the real
  SorobanSimulationResult shape; fixed mockResolvedValueOnce/
  mockRejectedValueOnce being consumed by the test's own first
  assertion, leaving the second call unmocked
- transaction.service.spec.ts: rewritten against TransactionService's
  actual create() contract (it doesn't call Soroban simulation at all;
  the previous spec tested a flow that was never implemented) and a
  real Ed25519 checksum address
- sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist
  on this Express-platform app; switched to app.listen + fetch,
  matching the sibling public-rate-limit.integration.spec.ts pattern,
  and named the test throttler 'api' so AstroidThrottlerGuard's
  tier-matching actually engages it

* Fix remaining pre-existing lint errors and undocumented env vars

CI runs lint and test repo-wide, so these also blocked the PR:

- 4 pre-existing no-explicit-any lint errors in throttler guard code
  and specs, typed properly instead of suppressed
- 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST,
  WEBHOOK_BURST) were validated by the env schema but missing from
  docs/configuration.md, failing the docs-sync test

* perf(auth): cache session revocation answers during token verification

Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.

- Add CacheService: a small get/set/delete cache over the shared
  REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
  prefix invalidation. Every operation degrades to a no-op/miss on Redis
  failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
  answers for TOKEN_CACHE_TTL seconds (default 30, well below the
  15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
  on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
  drops the cached answer after each blacklist write (including on
  Redis-outage fallback), AuthService invalidates on logout and on
  refresh rotation, so revocations are observed immediately instead of
  after the TTL window.

Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.

Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.

Closes #341

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat: correlate requests, paginate audit, sign webhooks

* feat(rate-limit): configurable limits and client identifiers for public routes

Public endpoints are the first thing abusive traffic hits, but their
limits were a single fixed IP budget: every public route shared
PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one
shared address (NAT, office egress, CI runners) exhausted one bucket
together. This makes the public rate limiter configurable per route and
per client identifier, on the existing Redis sliding-window counter.

- Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or
  per-controller) budget overrides resolved by the guard through
  Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default.
- Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated,
  currently 'apiKey'): when enabled, a presented x-api-key or
  ApiKey/Bearer ak_... Authorization header is folded into the bucket
  key so distinct key-holding clients behind one IP get their own
  budgets. The IP always participates; keyless callers share the
  plain-IP bucket as before.
- Extract a testable PublicRateLimitGuard.check() returning the full
  decision (allowed, limit, windowSeconds, count, resetAt); canActivate
  keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset +
  Retry-After contract and in-memory fallback on Redis outage.

Tests: unit suites for per-route rule resolution, identifier bucketing
(with/without identifiers configured), header correctness on allowed
and limited requests, and @SkipPublicRateLimit() interaction with rules;
an HTTP-level integration suite simulating bursts that proves the
429-with-headers behaviour at the global limit, per-route overrides
(next to unaffected sibling routes), and per-key budget isolation.
Existing public-rate-limit suites pass unchanged (bucket keys keep the
ip: prefix, so stored counters stay compatible).

Closes #342

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat: stream audit exports and harden payment safety (#63, #64, #283)

* fix: validate nested transaction metadata as JSON (#283)

* feat(transactions): implement agent spending limit evaluation guard

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(config): deduplicate parseClientIdentifiers and pass the raw variable

The duplicated helper also ignored its parameter while the config factory
read the raw variable at the call site with none, so the module never
compiled. Collapse to a single parser that takes the raw value.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat(audit): implement structured audit logging interceptor for sensitive operations

Closes #255

Adds the @AuditLog() decorator (with optional action/entity metadata) and reworks AuditLogInterceptor so only decorated routes are persisted, keeping read-only traffic free of audit writes.

Each record captures the actor (human user id or acting agent), client IP, HTTP method, request path, a SHA-256 fingerprint of the sanitized payload, the sanitized body itself, the final response status and handler duration. Secrets, keys, tokens, signatures and mnemonics are redacted recursively without mutating the original request body, and @SkipAudit() always wins.

Applied to the sensitive policy, budget and API-key endpoints; persistence stays fire-and-forget so a failed audit write never breaks the client request.

* refactor(policies): add Prisma repository abstraction for agent spending policies

Closes #254

Introduces SpendingPolicyRepository (src/modules/policies/spending-policy.repository.ts) as the single owner of every Prisma call for policies: create/find/update/soft-delete, the paginated read, the rolling spend aggregation used by velocity checks, the POLICY_EVALUATED audit row, and an interactive withTransaction() helper. Every method funnels through one error-handling wrapper that logs the operation (plus the Prisma error code) and rethrows the original error, so error handling and transaction safety are uniform across the repository.

Adds SpendingPolicyService, which injects the repository and now owns spending-policy validation and enforcement (staleness checks, daily velocity limit, evaluation audit) with no direct Prisma access. PolicyService is reduced to orchestration - domain events plus the pure PolicyEngine - and delegates all persistence to SpendingPolicyService, which removes the direct prisma.transaction/prisma.auditLog calls from the service layer. PolicyRepository is replaced by the new, better-named repository.

Unit tests cover the repository against a mocked Prisma client (including transaction usage, Decimal summing and error propagation) and the service against a mocked repository (validation, pagination, not-found, velocity limits and audit failure swallowing).

* feat(webhooks): add BullMQ retry policy and dead-letter audit for webhook deliveries

Closes #253

Centralises the webhook queue configuration in src/queues/webhook.queue.ts: 5 attempts, exponential backoff with a 2000ms base, the jittered custom backoff strategy, 24h retention of failed jobs and the dead-letter destination. Both the queue registration and the enqueue path consume the same constants, so the API and the worker can no longer drift apart.

Adds WebhookAuditService, which appends a WEBHOOK_DELIVERY_FAILED record (subscriber, event, attempts, HTTP status, failure reason) to the audit trail for deliveries that will never be retried - either an unrecoverable 4xx or the final attempt after retries are exhausted. Both the webhook processor and the worker call it through a fire-and-forget hook that swallows and logs its own failures, so audit problems can never mask the delivery error, crash the NestJS process or interfere with BullMQ retry/backoff.

Tests cover the retry/backoff/DLQ configuration and jitter envelope, the audit service (including persistence failures) and the processor's terminal-failure behaviour - retry without auditing while attempts remain, exactly one audit entry when retries are exhausted or a 4xx is returned, and error preservation when auditing fails.

* feat(rate-limit): add Redis-backed agent throttler guard for high-frequency agent endpoints

Closes #251

Adds AgentThrottlerGuard in src/common/guards/agent-throttler.guard.ts, extending @nestjs/throttler's ThrottlerGuard so counters keep running through the shared RedisThrottlerStorage while the guard adds agent-aware behaviour: the bucket key is derived from the acting agent (x-agent-id, a route/body/query agentId, or an agent-bound API-key principal) instead of the organization or IP, falling back to the organization, then a hashed API key, then the client IP (honouring x-forwarded-for) for unauthenticated public routes.

Introduces a third 'agent' tier (THROTTLE_AGENT_LIMIT, default 300/window) alongside api/auth in config/throttler.config.ts and env.validation.ts, and routes each request to exactly one tier - explicit @ThrottleTierDecorator() metadata wins, otherwise agent-identified traffic uses the agent tier and everything else the api tier. Rejections are standard 429s that also carry plain Retry-After, X-RateLimit-Limit and X-RateLimit-Remaining headers, and the tier limit is advertised on allowed responses too.

Applied selectively to the agent-facing controllers (agents, transactions, wallets) via @UseGuards so the existing global AstroidThrottlerGuard keeps enforcing api/auth untouched. Vitest coverage asserts tier routing, tracker resolution for agents/orgs/API keys/anonymous callers, header emission and 429 enforcement when a single agent's burst exhausts its budget.

* Fix CI: pre-existing typecheck/lint/test breakage inherited from main

main's HEAD was red (27 typecheck errors) from unrelated merged PRs
(#357, #374, #381-#388). Fixed to get this branch's CI green:

- throttler.guard.ts / sliding-window-throttler.guard.ts: AuthenticatedUser
  has no `sub` or `tier` field; use `.id` and treat `tier` as an optional
  extension until the type actually carries it.
- agent.controller.ts: imported SlidingWindowThrottlerGuard/SlidingWindowLimit
  (unused) instead of the AstroidThrottlerGuard actually referenced by
  @UseGuards.
- event-names.ts: dropped a duplicate TransactionRiskScoringRequested key.
- risk.service.ts: the TransactionCreated handler built a RiskFactorsInput
  with fields (`destination`, `velocityCount`, `isNewRecipient`) that don't
  exist on the type; mapped to the real shape instead. Removed dead
  lowRisk/createEventBus fixtures left over in risk.service.spec.ts.
- stellar.service.ts (Soroban variant): fixed getTransactionInfo calling a
  nonexistent client method (getTransaction), simulateTransaction being
  called with a bare string instead of {transactionXdr}, and an error
  message interpolating the whole error object instead of `.message`.
  Rewrote stellar.service.spec.ts's mocks/assertions to match the real
  SorobanSimulationResult shape, and fixed two tests reusing a `mockOnce`
  across two separate calls to the service (second call fell through to
  the unmocked default and threw on undefined).
- transaction.service.spec.ts targeted a pre-broadcast Soroban simulation
  step that was never wired into TransactionService.create, against a
  StellarService overload TransactionService doesn't even inject; skipped
  with an explanation rather than fabricating the feature.
- sensitive-rate-limit.integration.spec.ts: replaced Fastify-only
  `app.inject()` (this app runs on platform-express) with a small
  http.request helper; the throttler in the test module was unnamed
  ('default'), so AstroidThrottlerGuard's per-tier name match against the
  route's default 'api' tier always skipped it — named it 'api' to match.
- Removed remaining `any` usages in the touched guard files/specs.
- docs/configuration.md: documented THROTTLE_WEBHOOK_LIMIT,
  THROTTLE_API_BURST, THROTTLE_AUTH_BURST, THROTTLE_WEBHOOK_BURST (missing
  from #386), failing the configuration-documentation test.

* Fix CI: repair botched merge of main into this branch

A merge of upstream main (PR #372's overlapping CI fixes) into this
branch left several files with duplicated/glued content instead of
resolved conflicts:

- sliding-window-throttler.guard.spec.ts: duplicate makeContext()
  declaration (one-line + multi-line versions concatenated).
- sliding-window-throttler.guard.ts: `limit` reverted to `const` while
  the tier-adjustment branches below still reassign it.
- sensitive-rate-limit.integration.spec.ts: my raw http.request-based
  version and another (cleaner, fetch()-based) version from main were
  concatenated rather than merged — duplicate imports, an unclosed `it`
  block. Kept the fetch()-based version.
- stellar.service.spec.ts: duplicate `error` key in a SorobanSimulationResult
  object literal.
- transaction.service.spec.ts: my placeholder describe.skip stub was
  prepended to a real, correctly implemented version of the same suite
  that showed up on main independently. Kept the real implementation,
  dropped the stub.

* Remove conflict markers from worker merge

* feat: implement risk scoring persistence and complete API documentation

Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.

Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository

Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios

Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules

Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages

* fix: add missing relation field in RiskAssessment model

Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.

* fix: resolve TypeScript errors in test files

Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests

* fix: add null checks for toPromise() results in response interceptor tests

Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.

* fix: add eslint-disable comments for temporary any types

Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.

* fix: update test expectation for date comparison in statistics test

Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.

* fix: remove duplicate Zod validation pipe integration test

* fix ci

* fix ci

* fix failing ci

* fix ci

* fix ci

* fix ci

* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)

* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files

Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* fix: use ZodValidationException in integration spec

The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: address requested issues - Done with all issues (#357)

* feat(health): dedicated GET /health/redis probe (#358)

Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.

* test(auth): lock in auth throttle-tier wiring on public routes (#359)

Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(health): add /health/live and /health/ready probes (#362)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes #351

* feat(config): validate full environment at startup (#363)

Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.

- On failure the process prints every failing variable in one message and
  exits with code 1, instead of surfacing only the first failing slice from
  inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
  default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
  set explicitly or implied by omission) and a JWT refresh secret that
  reuses the access secret.

Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.

Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.

Closes #350

* fix(tests): add SpendingLimitService mock to TransactionService spec

* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)

Closes #331
Closes #325
Closes #334
Closes #328

- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
  existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
  with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
  (WebhookWorker), keeping WebhooksProcessor which also records metrics.
  Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
  new public ingress route for Redis-backed sliding-window rate limiting
  with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
  per stream, exposed through the existing Prometheus registry.

* feat: add production rollback protection to database migration CLI (#371)

Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.

down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.

The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.

* test: add full branch coverage tests for role, permission and scope guards (#373)

Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.

The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.

Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
  bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
  leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.

* perf: add concurrent composite indexes for notification, approval and memory (#374)

Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:

- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
  mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline

Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.

* feat: add IP-based rate limiting to public endpoints (#377)

Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.

- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
  on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
  window script, so all replicas enforce one budget per IP; rejected
  requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
  per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
  PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
  PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
  network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test

* feat: return RFC 9457 problem details for all error responses (#378)

Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:

  { type, title, status, detail, instance, code, requestId, details? }

- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
  errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
  system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
  switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
  field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
  without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
  validation, authentication, domain, not-found and server errors

* fix: worker error handling + log scrubbing (#370)

* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* test: add job-worker spec and queue-failure-listener scrubbing test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)

* perf(analytics): batch dashboard overview queries into one transaction

Closes #338

AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.

* test(audit): cover pagination and sorting edge cases for activity log

Closes #339

The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.

* fix(migrations): resolve colliding timestamp between two merged migrations

Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.

* fix(ci): document missing env vars and fix flaky retry.util test

Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:

- docs/configuration.md was missing 7 env vars added by recent merges
  (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
  DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
  env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
  fake timers with vi.runAllTimersAsync(), then attach the rejection
  assertion afterward — a race that surfaces as an unhandled rejection
  under full-suite load (deterministic once >100 files run together).
  Attaching a no-op .catch() immediately after creating the promise
  prevents the unhandled state without changing what each test asserts.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)

* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)

* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)

* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)

* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)

* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)

* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)

* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)

* fix: resolve rebase conflicts and restore CI

* feat: implement risk scoring persistence and complete API documentation

Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.

Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository

Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios

Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules

Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages

* fix: add missing relation field in RiskAssessment model

Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.

* fix: resolve TypeScript errors in test files

Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests

* fix: add null checks for toPromise() results in response interceptor tests

Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.

* fix: add eslint-disable comments for temporary any types

Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.

* fix: update test expectation for date comparison in statistics test

Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.

* fix: remove duplicate Zod validation pipe integration test

* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)

* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files

Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* fix: use ZodValidationException in integration spec

The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: address requested issues - Done with all issues (#357)

* feat(health): dedicated GET /health/redis probe (#358)

Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.

* test(auth): lock in auth throttle-tier wiring on public routes (#359)

Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(health): add /health/live and /health/ready probes (#362)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes #351

* feat(config): validate full environment at startup (#363)

Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.

- On failure the process prints every failing variable in one message and
  exits with code 1, instead of surfacing only the first failing slice from
  inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
  default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
  set explicitly or implied by omission) and a JWT refresh secret that
  reuses the access secret.

Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.

Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.

Closes #350

* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)

Closes #331
Closes #325
Closes #334
Closes #328

- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
  existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
  with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
  (WebhookWorker), keeping WebhooksProcessor which also records metrics.
  Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
  new public ingress route for Redis-backed sliding-window rate limiting
  with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
  per stream, exposed through the existing Prometheus registry.

* feat: add production rollback protection to database migration CLI (#371)

Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.

down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.

The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.

* test: add full branch coverage tests for role, permission and scope guards (#373)

Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.

The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.

Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
  bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
  leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.

* perf: add concurrent composite indexes for notification, approval and memory (#374)

Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:

- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
  mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline

Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.

* feat: add IP-based rate limiting to public endpoints (#377)

Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.

- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
  on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
  window script, so all replicas enforce one budget per IP; rejected
  requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
  per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
  PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
  PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
  network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test

* feat: return RFC 9457 problem details for all error responses (#378)

Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:

  { type, title, status, detail, instance, code, requestId, details? }

- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
  errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
  system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
  switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
  field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
  without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
  validation, authentication, domain, not-found and server errors

* fix: worker error handling + log scrubbing (#370)

* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* test: add job-worker spec and queue-failure-listener scrubbing test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)

* perf(analytics): batch dashboard overview queries into one transaction

Closes #338

AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.

* test(audit): cover pagination and sorting edge cases for activity log

Closes #339

The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.

* fix(migrations): resolve colliding timestamp between two merged migrations

Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.

* fix(ci): document missing env vars and fix flaky retry.util test

Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:

- docs/configuration.md was missing 7 env vars added by recent merges
  (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
  DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
  env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
  fake timers with vi.runAllTimersAsync(), then attach the rejection
  assertion afterward — a race that surfaces as an unhandled rejection
  under full-suite load (deterministic once >100 files run together).
  Attaching a no-op .catch() immediately after creating the promise
  prevents the unhandled state without changing what each test asserts.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)

* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)

* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)

* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)

* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)

* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)

* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)

* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)

* Add startup migration check and DB pool connection metrics

Wires the existing migration status checker into app bootstrap so the
process halts before accepting traffic when prisma/migrations has
pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt,
the default; 'warn' logs and continues). Gated behind
DATABASE_MIGRATION_CHECK_ENABLED.

Adds a db_pool_connections Prometheus gauge (active/idle/waiting)
sourced from pg_stat_activity, since Prisma's Rust query engine
doesn't expose pool internals through the Node client.

Closes #335
Closes #332

* Fix pre-existing typecheck/lint/test failures blocking CI

main's build/typecheck/lint/test were already broken before this branch
touched anything (confirmed by checking out upstream/main directly).
CI enforces these repo-wide, so they block this PR too. Fixed each:

- event-names.ts: duplicate object key (TransactionRiskScoringRequested)
- throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't
  exist on that type (JWT payload field name leaked into the wrong type)
- sliding-window-throttler.guard.ts: removed a user-tier rate-limit
  multiplier keyed on AuthenticatedUser.tier, a field never present
  anywhere in the auth/user model — dead, unbacked logic. Removed its
  now-orphaned tests too.
- agent.controller.ts: AstroidThrottlerGuard used but never imported;
  dropped an unused SlidingWindowThrottlerGuard import block
- risk.service.ts: event-driven risk scoring built a RiskFactorsInput
  with fields (destination/velocityCount/isNewRecipient) that don't
  exist on the current type; mapped to the real shape instead
- risk.service.spec.ts: removed orphaned unused fixtures
- stellar.service.ts (src/modules/stellar/services, unused elsewhere
  in the app but still typechecked/tested): getTransactionInfo called
  a client method that doesn't exist (real method is getTransaction);
  simulateTransaction passed a bare string where the client expects
  an options object
- stellar.service.spec.ts: rewritten against the real
  SorobanSimulationResult shape; fixed mockResolvedValueOnce/
  mockRejectedValueOnce being consumed by the test's own first
  assertion, leaving the second call unmocked
- transaction.service.spec.ts: rewritten against TransactionService's
  actual create() contract (it doesn't call Soroban simulation at all;
  the previous spec tested a flow that was never implemented) and a
  real Ed25519 checksum address
- sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist
  on this Express-platform app; switched to app.listen + fetch,
  matching the sibling public-rate-limit.integration.spec.ts pattern,
  and named the test throttler 'api' so AstroidThrottlerGuard's
  tier-matching actually engages it

* Fix remaining pre-existing lint errors and undocumented env vars

CI runs lint and test repo-wide, so these also blocked the PR:

- 4 pre-existing no-explicit-any lint errors in throttler guard code
  and specs, typed properly instead of suppressed
- 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST,
  WEBHOOK_BURST) were validated by the env schema but missing from
  docs/configuration.md, failing the docs-sync test

* perf(auth): cache session revocation answers during token verification

Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.

- Add CacheService: a small get/set/delete cache over the shared
  REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
  prefix invalidation. Every operation degrades to a no-op/miss on Redis
  failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
  answers for TOKEN_CACHE_TTL seconds (default 30, well below the
  15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
  on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
  drops the cached answer after each blacklist write (including on
  Redis-outage fallback), AuthService invalidates on logout and on
  refresh rotation, so revocations are observed immediately instead of
  after the TTL window.

Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.

Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.

Closes #341

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(ci): resolv…
Cjay-Cyber-2 added a commit that referenced this pull request Oct 7, 2026
…ed CI job (#364)

* feat: address requested issues - closes #264, closes #263, closes #262, closes #261

* ci(migrations): harden migration verification and run it as a dedicated CI job

Expand scripts/verify-migrations.sh into three tiers of checks that report
every problem in one run, as GitHub annotations in CI:

1. Structure and SQL (always): prisma validate; migration_lock.toml present
   and matching the schema provider; folder names follow Prisma's
   <timestamp>_<name> convention (plus the 0_ baseline) with unique
   timestamps; each folder has a migration.sql with at least one statement;
   no stray files. A comment-, string- and dollar-quote-aware SQL lexer
   (POSIX awk, runs under mawk) rejects merge-conflict markers, unbalanced
   parentheses, unterminated strings/identifiers/comments and a missing
   final semicolon, with file and line.
2. History (MIGRATION_BASE_REF): migrations already on the base branch must
   not be modified or deleted (deployed databases store their checksums),
   new migrations must sort after the newest base migration, and
   destructive statements in new migrations are reported as warnings.
3. Replay and drift (SHADOW_DATABASE_URL): replay the full history on an
   empty database with prisma migrate diff, catching invalid SQL and
   forward dependencies between migrations, and fail on schema drift.

Run it in a dedicated verify-migrations job with its own disposable
Postgres, full git history, and a base ref derived from the PR target or
the previous push tip. This replaces the previous step, whose shadow
database URL the old script never used.

Make `npm run db:verify` invoke bash explicitly so it works on Windows,
mark the script executable, and document the requirements in
CONTRIBUTING.md and docs/database.md.

Closes #256

* feat: Add comprehensive observability, security, and simulation improvements

Implements four major feature requests for enhanced API observability,
security, and transaction simulation capabilities:

#247 - Prometheus Metrics Interceptor
- Create MetricsInterceptor for HTTP request metrics collection
- Add request duration histograms, active request gauges, and request counters
- Categorize metrics by route, method, and status code
- Exclude /metrics endpoint from self-instrumentation
- Add comprehensive unit tests (9 tests)

#246 - Stellar Transaction Simulation Service
- Create StellarSimulationService for transaction simulation
- Integrate with Soroban RPC for XDR validation and simulation
- Include risk assessment and fee estimation
- Add circuit breaker protection for RPC failures
- Add XDR validation helper method
- Add comprehensive unit tests with mocked Stellar RPC (24 tests)

#245 - Cryptographic API Key Hashing Upgrade
- Upgrade from SHA-256 to Argon2id for enhanced security
- Implement memory-hard algorithm resistant to GPU/ASIC attacks
- Add timing-attack resistant comparison via Argon2 verification
- Maintain SHA-256 fallback for backward compatibility
- Update ApiKeyService with dual-algorithm verification
- Add comprehensive unit tests (32 crypto tests, 15 API key tests)

#244 - Webhook Retry and Dead Letter Queue
- Verify existing implementation meets all requirements
- Confirm exponential backoff with jitter (2000ms base, 20% jitter)
- Confirm 5 max attempts and non-transient error detection
- Verify dead-letter handler via DeadLetterService
- All existing tests passing

Closes #247, #246, #245, #244

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(workers): centralize job error handling with scrubbed structured logging

Add runWorkerJob, a wrapper every background worker now routes its handler
through. It times the job via WorkerMetricsService when available, logs a
structured completion record, and classifies failures before rethrowing:
transient failures log a job.retrying warning, while a failure on the final
attempt or an UnrecoverableError logs a job.dead-lettered error with the
scrubbed payload and stack. The original error is always rethrown untouched
so BullMQ retry semantics are preserved, and logging can never mask it.

Add scrubForLog/scrubString, which redact sensitive keys (sharing the audit
sanitizer's key list) and secret-shaped substrings such as Stellar seeds,
bearer tokens and URL credentials, and coerce cycles, bigints and errors
into JSON-safe values.

QueueFailureListener now scrubs its log line as well; previously the raw
webhook payload, including its signing secret, was written to the error
log. The dead-letter copy keeps the raw payload so re-drive still works.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping

* Fix CI: document missing env vars, fix flaky retry backoff test

- docs/configuration.md was missing entries for DATABASE_SLOW_QUERY_THRESHOLD_MS,
  DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, and the
  PUBLIC_RATE_LIMIT_* vars, failing the configuration-documentation test.
- retry.util.spec.ts left three rejected promises unhandled between
  `runAllTimersAsync()` and the `expect(...).rejects` assertion that
  attaches the handler; attach a no-op .catch() immediately after creating
  each promise so fake-timer-driven rejections don't fire as unhandled
  rejections mid-test-run.

* perf(auth): cache session revocation answers during token verification

Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.

- Add CacheService: a small get/set/delete cache over the shared
  REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
  prefix invalidation. Every operation degrades to a no-op/miss on Redis
  failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
  answers for TOKEN_CACHE_TTL seconds (default 30, well below the
  15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
  on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
  drops the cached answer after each blacklist write (including on
  Redis-outage fallback), AuthService invalidates on logout and on
  refresh rotation, so revocations are observed immediately instead of
  after the TTL window.

Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.

Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.

Closes #341

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat: correlate requests, paginate audit, sign webhooks

* feat(rate-limit): configurable limits and client identifiers for public routes

Public endpoints are the first thing abusive traffic hits, but their
limits were a single fixed IP budget: every public route shared
PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one
shared address (NAT, office egress, CI runners) exhausted one bucket
together. This makes the public rate limiter configurable per route and
per client identifier, on the existing Redis sliding-window counter.

- Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or
  per-controller) budget overrides resolved by the guard through
  Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default.
- Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated,
  currently 'apiKey'): when enabled, a presented x-api-key or
  ApiKey/Bearer ak_... Authorization header is folded into the bucket
  key so distinct key-holding clients behind one IP get their own
  budgets. The IP always participates; keyless callers share the
  plain-IP bucket as before.
- Extract a testable PublicRateLimitGuard.check() returning the full
  decision (allowed, limit, windowSeconds, count, resetAt); canActivate
  keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset +
  Retry-After contract and in-memory fallback on Redis outage.

Tests: unit suites for per-route rule resolution, identifier bucketing
(with/without identifiers configured), header correctness on allowed
and limited requests, and @SkipPublicRateLimit() interaction with rules;
an HTTP-level integration suite simulating bursts that proves the
429-with-headers behaviour at the global limit, per-route overrides
(next to unaffected sibling routes), and per-key budget isolation.
Existing public-rate-limit suites pass unchanged (bucket keys keep the
ip: prefix, so stored counters stay compatible).

Closes #342

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat: stream audit exports and harden payment safety (#63, #64, #283)

* fix: validate nested transaction metadata as JSON (#283)

* feat(transactions): implement agent spending limit evaluation guard

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(config): deduplicate parseClientIdentifiers and pass the raw variable

The duplicated helper also ignored its parameter while the config factory
read the raw variable at the call site with none, so the module never
compiled. Collapse to a single parser that takes the raw value.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat(audit): implement structured audit logging interceptor for sensitive operations

Closes #255

Adds the @AuditLog() decorator (with optional action/entity metadata) and reworks AuditLogInterceptor so only decorated routes are persisted, keeping read-only traffic free of audit writes.

Each record captures the actor (human user id or acting agent), client IP, HTTP method, request path, a SHA-256 fingerprint of the sanitized payload, the sanitized body itself, the final response status and handler duration. Secrets, keys, tokens, signatures and mnemonics are redacted recursively without mutating the original request body, and @SkipAudit() always wins.

Applied to the sensitive policy, budget and API-key endpoints; persistence stays fire-and-forget so a failed audit write never breaks the client request.

* refactor(policies): add Prisma repository abstraction for agent spending policies

Closes #254

Introduces SpendingPolicyRepository (src/modules/policies/spending-policy.repository.ts) as the single owner of every Prisma call for policies: create/find/update/soft-delete, the paginated read, the rolling spend aggregation used by velocity checks, the POLICY_EVALUATED audit row, and an interactive withTransaction() helper. Every method funnels through one error-handling wrapper that logs the operation (plus the Prisma error code) and rethrows the original error, so error handling and transaction safety are uniform across the repository.

Adds SpendingPolicyService, which injects the repository and now owns spending-policy validation and enforcement (staleness checks, daily velocity limit, evaluation audit) with no direct Prisma access. PolicyService is reduced to orchestration - domain events plus the pure PolicyEngine - and delegates all persistence to SpendingPolicyService, which removes the direct prisma.transaction/prisma.auditLog calls from the service layer. PolicyRepository is replaced by the new, better-named repository.

Unit tests cover the repository against a mocked Prisma client (including transaction usage, Decimal summing and error propagation) and the service against a mocked repository (validation, pagination, not-found, velocity limits and audit failure swallowing).

* feat(webhooks): add BullMQ retry policy and dead-letter audit for webhook deliveries

Closes #253

Centralises the webhook queue configuration in src/queues/webhook.queue.ts: 5 attempts, exponential backoff with a 2000ms base, the jittered custom backoff strategy, 24h retention of failed jobs and the dead-letter destination. Both the queue registration and the enqueue path consume the same constants, so the API and the worker can no longer drift apart.

Adds WebhookAuditService, which appends a WEBHOOK_DELIVERY_FAILED record (subscriber, event, attempts, HTTP status, failure reason) to the audit trail for deliveries that will never be retried - either an unrecoverable 4xx or the final attempt after retries are exhausted. Both the webhook processor and the worker call it through a fire-and-forget hook that swallows and logs its own failures, so audit problems can never mask the delivery error, crash the NestJS process or interfere with BullMQ retry/backoff.

Tests cover the retry/backoff/DLQ configuration and jitter envelope, the audit service (including persistence failures) and the processor's terminal-failure behaviour - retry without auditing while attempts remain, exactly one audit entry when retries are exhausted or a 4xx is returned, and error preservation when auditing fails.

* feat(rate-limit): add Redis-backed agent throttler guard for high-frequency agent endpoints

Closes #251

Adds AgentThrottlerGuard in src/common/guards/agent-throttler.guard.ts, extending @nestjs/throttler's ThrottlerGuard so counters keep running through the shared RedisThrottlerStorage while the guard adds agent-aware behaviour: the bucket key is derived from the acting agent (x-agent-id, a route/body/query agentId, or an agent-bound API-key principal) instead of the organization or IP, falling back to the organization, then a hashed API key, then the client IP (honouring x-forwarded-for) for unauthenticated public routes.

Introduces a third 'agent' tier (THROTTLE_AGENT_LIMIT, default 300/window) alongside api/auth in config/throttler.config.ts and env.validation.ts, and routes each request to exactly one tier - explicit @ThrottleTierDecorator() metadata wins, otherwise agent-identified traffic uses the agent tier and everything else the api tier. Rejections are standard 429s that also carry plain Retry-After, X-RateLimit-Limit and X-RateLimit-Remaining headers, and the tier limit is advertised on allowed responses too.

Applied selectively to the agent-facing controllers (agents, transactions, wallets) via @UseGuards so the existing global AstroidThrottlerGuard keeps enforcing api/auth untouched. Vitest coverage asserts tier routing, tracker resolution for agents/orgs/API keys/anonymous callers, header emission and 429 enforcement when a single agent's burst exhausts its budget.

* Fix CI: pre-existing typecheck/lint/test breakage inherited from main

main's HEAD was red (27 typecheck errors) from unrelated merged PRs
(#357, #374, #381-#388). Fixed to get this branch's CI green:

- throttler.guard.ts / sliding-window-throttler.guard.ts: AuthenticatedUser
  has no `sub` or `tier` field; use `.id` and treat `tier` as an optional
  extension until the type actually carries it.
- agent.controller.ts: imported SlidingWindowThrottlerGuard/SlidingWindowLimit
  (unused) instead of the AstroidThrottlerGuard actually referenced by
  @UseGuards.
- event-names.ts: dropped a duplicate TransactionRiskScoringRequested key.
- risk.service.ts: the TransactionCreated handler built a RiskFactorsInput
  with fields (`destination`, `velocityCount`, `isNewRecipient`) that don't
  exist on the type; mapped to the real shape instead. Removed dead
  lowRisk/createEventBus fixtures left over in risk.service.spec.ts.
- stellar.service.ts (Soroban variant): fixed getTransactionInfo calling a
  nonexistent client method (getTransaction), simulateTransaction being
  called with a bare string instead of {transactionXdr}, and an error
  message interpolating the whole error object instead of `.message`.
  Rewrote stellar.service.spec.ts's mocks/assertions to match the real
  SorobanSimulationResult shape, and fixed two tests reusing a `mockOnce`
  across two separate calls to the service (second call fell through to
  the unmocked default and threw on undefined).
- transaction.service.spec.ts targeted a pre-broadcast Soroban simulation
  step that was never wired into TransactionService.create, against a
  StellarService overload TransactionService doesn't even inject; skipped
  with an explanation rather than fabricating the feature.
- sensitive-rate-limit.integration.spec.ts: replaced Fastify-only
  `app.inject()` (this app runs on platform-express) with a small
  http.request helper; the throttler in the test module was unnamed
  ('default'), so AstroidThrottlerGuard's per-tier name match against the
  route's default 'api' tier always skipped it — named it 'api' to match.
- Removed remaining `any` usages in the touched guard files/specs.
- docs/configuration.md: documented THROTTLE_WEBHOOK_LIMIT,
  THROTTLE_API_BURST, THROTTLE_AUTH_BURST, THROTTLE_WEBHOOK_BURST (missing
  from #386), failing the configuration-documentation test.

* Fix CI: repair botched merge of main into this branch

A merge of upstream main (PR #372's overlapping CI fixes) into this
branch left several files with duplicated/glued content instead of
resolved conflicts:

- sliding-window-throttler.guard.spec.ts: duplicate makeContext()
  declaration (one-line + multi-line versions concatenated).
- sliding-window-throttler.guard.ts: `limit` reverted to `const` while
  the tier-adjustment branches below still reassign it.
- sensitive-rate-limit.integration.spec.ts: my raw http.request-based
  version and another (cleaner, fetch()-based) version from main were
  concatenated rather than merged — duplicate imports, an unclosed `it`
  block. Kept the fetch()-based version.
- stellar.service.spec.ts: duplicate `error` key in a SorobanSimulationResult
  object literal.
- transaction.service.spec.ts: my placeholder describe.skip stub was
  prepended to a real, correctly implemented version of the same suite
  that showed up on main independently. Kept the real implementation,
  dropped the stub.

* Remove conflict markers from worker merge

* feat: implement risk scoring persistence and complete API documentation

Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.

Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository

Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios

Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules

Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages

* fix: add missing relation field in RiskAssessment model

Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.

* fix: resolve TypeScript errors in test files

Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests

* fix: add null checks for toPromise() results in response interceptor tests

Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.

* fix: add eslint-disable comments for temporary any types

Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.

* fix: update test expectation for date comparison in statistics test

Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.

* fix: remove duplicate Zod validation pipe integration test

* fix ci

* fix ci

* fix failing ci

* fix ci

* fix ci

* fix ci

* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)

* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files

Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* fix: use ZodValidationException in integration spec

The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: address requested issues - Done with all issues (#357)

* feat(health): dedicated GET /health/redis probe (#358)

Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.

* test(auth): lock in auth throttle-tier wiring on public routes (#359)

Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(health): add /health/live and /health/ready probes (#362)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes #351

* feat(config): validate full environment at startup (#363)

Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.

- On failure the process prints every failing variable in one message and
  exits with code 1, instead of surfacing only the first failing slice from
  inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
  default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
  set explicitly or implied by omission) and a JWT refresh secret that
  reuses the access secret.

Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.

Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.

Closes #350

* fix(tests): add SpendingLimitService mock to TransactionService spec

* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)

Closes #331
Closes #325
Closes #334
Closes #328

- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
  existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
  with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
  (WebhookWorker), keeping WebhooksProcessor which also records metrics.
  Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
  new public ingress route for Redis-backed sliding-window rate limiting
  with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
  per stream, exposed through the existing Prometheus registry.

* feat: add production rollback protection to database migration CLI (#371)

Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.

down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.

The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.

* test: add full branch coverage tests for role, permission and scope guards (#373)

Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.

The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.

Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
  bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
  leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.

* perf: add concurrent composite indexes for notification, approval and memory (#374)

Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:

- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
  mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline

Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.

* feat: add IP-based rate limiting to public endpoints (#377)

Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.

- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
  on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
  window script, so all replicas enforce one budget per IP; rejected
  requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
  per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
  PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
  PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
  network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test

* feat: return RFC 9457 problem details for all error responses (#378)

Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:

  { type, title, status, detail, instance, code, requestId, details? }

- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
  errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
  system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
  switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
  field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
  without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
  validation, authentication, domain, not-found and server errors

* fix: worker error handling + log scrubbing (#370)

* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* test: add job-worker spec and queue-failure-listener scrubbing test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)

* perf(analytics): batch dashboard overview queries into one transaction

Closes #338

AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.

* test(audit): cover pagination and sorting edge cases for activity log

Closes #339

The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.

* fix(migrations): resolve colliding timestamp between two merged migrations

Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.

* fix(ci): document missing env vars and fix flaky retry.util test

Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:

- docs/configuration.md was missing 7 env vars added by recent merges
  (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
  DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
  env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
  fake timers with vi.runAllTimersAsync(), then attach the rejection
  assertion afterward — a race that surfaces as an unhandled rejection
  under full-suite load (deterministic once >100 files run together).
  Attaching a no-op .catch() immediately after creating the promise
  prevents the unhandled state without changing what each test asserts.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)

* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)

* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)

* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)

* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)

* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)

* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)

* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)

* fix: resolve rebase conflicts and restore CI

* feat: implement risk scoring persistence and complete API documentation

Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.

Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository

Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios

Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules

Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages

* fix: add missing relation field in RiskAssessment model

Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.

* fix: resolve TypeScript errors in test files

Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests

* fix: add null checks for toPromise() results in response interceptor tests

Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.

* fix: add eslint-disable comments for temporary any types

Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.

* fix: update test expectation for date comparison in statistics test

Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.

* fix: remove duplicate Zod validation pipe integration test

* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)

* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files

Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* fix: use ZodValidationException in integration spec

The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: address requested issues - Done with all issues (#357)

* feat(health): dedicated GET /health/redis probe (#358)

Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.

* test(auth): lock in auth throttle-tier wiring on public routes (#359)

Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(health): add /health/live and /health/ready probes (#362)

Add orchestrator-grade liveness and readiness probes:

- GET /health/live returns 200 whenever the process is running and performs
  no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
  in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
  per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
  probe paths are stable across API versions.

Make the health controller usable by load balancers:

- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
  request (or attempt to while the database is down).

Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.

Closes #351

* feat(config): validate full environment at startup (#363)

Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.

- On failure the process prints every failing variable in one message and
  exits with code 1, instead of surfacing only the first failing slice from
  inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
  default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
  set explicitly or implied by omission) and a JWT refresh secret that
  reuses the access secret.

Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.

Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.

Closes #350

* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)

Closes #331
Closes #325
Closes #334
Closes #328

- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
  existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
  with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
  (WebhookWorker), keeping WebhooksProcessor which also records metrics.
  Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
  new public ingress route for Redis-backed sliding-window rate limiting
  with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
  per stream, exposed through the existing Prometheus registry.

* feat: add production rollback protection to database migration CLI (#371)

Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.

down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.

The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.

* test: add full branch coverage tests for role, permission and scope guards (#373)

Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.

The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.

Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
  bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
  leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.

* perf: add concurrent composite indexes for notification, approval and memory (#374)

Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:

- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
  mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline

Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.

* feat: add IP-based rate limiting to public endpoints (#377)

Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.

- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
  on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
  window script, so all replicas enforce one budget per IP; rejected
  requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
  per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
  PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
  PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
  network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test

* feat: return RFC 9457 problem details for all error responses (#378)

Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:

  { type, title, status, detail, instance, code, requestId, details? }

- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
  errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
  system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
  switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
  field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
  without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
  validation, authentication, domain, not-found and server errors

* fix: worker error handling + log scrubbing (#370)

* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* test: add job-worker spec and queue-failure-listener scrubbing test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8

* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)

* perf(analytics): batch dashboard overview queries into one transaction

Closes #338

AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.

* test(audit): cover pagination and sorting edge cases for activity log

Closes #339

The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.

* fix(migrations): resolve colliding timestamp between two merged migrations

Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.

* fix(ci): document missing env vars and fix flaky retry.util test

Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:

- docs/configuration.md was missing 7 env vars added by recent merges
  (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
  DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
  env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
  fake timers with vi.runAllTimersAsync(), then attach the rejection
  assertion afterward — a race that surfaces as an unhandled rejection
  under full-suite load (deterministic once >100 files run together).
  Attaching a no-op .catch() immediately after creating the promise
  prevents the unhandled state without changing what each test asserts.

* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)

* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)

* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)

* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)

* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)

* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)

* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)

* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)

* Add startup migration check and DB pool connection metrics

Wires the existing migration status checker into app bootstrap so the
process halts before accepting traffic when prisma/migrations has
pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt,
the default; 'warn' logs and continues). Gated behind
DATABASE_MIGRATION_CHECK_ENABLED.

Adds a db_pool_connections Prometheus gauge (active/idle/waiting)
sourced from pg_stat_activity, since Prisma's Rust query engine
doesn't expose pool internals through the Node client.

Closes #335
Closes #332

* Fix pre-existing typecheck/lint/test failures blocking CI

main's build/typecheck/lint/test were already broken before this branch
touched anything (confirmed by checking out upstream/main directly).
CI enforces these repo-wide, so they block this PR too. Fixed each:

- event-names.ts: duplicate object key (TransactionRiskScoringRequested)
- throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't
  exist on that type (JWT payload field name leaked into the wrong type)
- sliding-window-throttler.guard.ts: removed a user-tier rate-limit
  multiplier keyed on AuthenticatedUser.tier, a field never present
  anywhere in the auth/user model — dead, unbacked logic. Removed its
  now-orphaned tests too.
- agent.controller.ts: AstroidThrottlerGuard used but never imported;
  dropped an unused SlidingWindowThrottlerGuard import block
- risk.service.ts: event-driven risk scoring built a RiskFactorsInput
  with fields (destination/velocityCount/isNewRecipient) that don't
  exist on the current type; mapped to the real shape instead
- risk.service.spec.ts: removed orphaned unused fixtures
- stellar.service.ts (src/modules/stellar/services, unused elsewhere
  in the app but still typechecked/tested): getTransactionInfo called
  a client method that doesn't exist (real method is getTransaction);
  simulateTransaction passed a bare string where the client expects
  an options object
- stellar.service.spec.ts: rewritten against the real
  SorobanSimulationResult shape; fixed mockResolvedValueOnce/
  mockRejectedValueOnce being consumed by the test's own first
  assertion, leaving the second call unmocked
- transaction.service.spec.ts: rewritten against TransactionService's
  actual create() contract (it doesn't call Soroban simulation at all;
  the previous spec tested a flow that was never implemented) and a
  real Ed25519 checksum address
- sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist
  on this Express-platform app; switched to app.listen + fetch,
  matching the sibling public-rate-limit.integration.spec.ts pattern,
  and named the test throttler 'api' so AstroidThrottlerGuard's
  tier-matching actually engages it

* Fix remaining pre-existing lint errors and undocumented env vars

CI runs lint and test repo-wide, so these also blocked the PR:

- 4 pre-existing no-explicit-any lint errors in throttler guard code
  and specs, typed properly instead of suppressed
- 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST,
  WEBHOOK_BURST) were validated by the env schema but missing from
  docs/configuration.md, failing the docs-sync test

* perf(auth): cache session revocation answers during token verification

Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.

- Add CacheService: a small get/set/delete cache over the shared
  REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
  prefix invalidation. Every operation degrades to a no-op/miss on Redis
  failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
  answers for TOKEN_CACHE_TTL seconds (default 30, well below the
  15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
  on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
  drops the cached answer after each blacklist write (including on
  Redis-outage fallback), AuthService invalidates on logout and on
  refresh rotation, so revocations are observed immediately instead of
  after the TTL window.

Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.

Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.

Closes #341

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* feat(rate-limit): configurable limits and client identifiers for public routes

Public endpoints are the first thing abusive traffic hits, but their
limits were a single fixed IP budget: every public route shared
PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one
shared address (NAT, office egress, CI runners) exhausted one bucket
together. This makes the public rate limiter configurable per route and
per client identifier, on the existing Redis sliding-window counter.

- Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or
  per-controller) budget overrides resolved by the guard through
  Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default.
- Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated,
  currently 'apiKey'): when enabled, a presented x-api-key or
  ApiKey/Bearer ak_... Authorization header is folded into the bucket
  key so distinct key-holding clients behind one IP get their own
  budgets. The IP always participates; keyless callers share the
  plain-IP bucket as before.
- Extract a testable PublicRateLimitGuard.check() returning the full
  decision (allowed, limit, windowSeconds, count, resetAt); canActivate
  keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset +
  Retry-After contract and in-memory fallback on Redis outage.

Tests: unit suites for per-route rule resolution, identifier bucketing
(with/without identifiers configured), header correctness on allowed
and limited requests, and @SkipPublicRateLimit() interaction with rules;
an HTTP-level integration suite simulating bursts that proves the
429-with-headers behaviour at the global limit, per-route overrides
(next to unaffected sibling routes), and per-key budget isolation.
Existing public-rate-limit suites pass unchanged (bucket keys keep the
ip: prefix, so stored counters stay compatible).

Closes #342

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(ci): resolve typecheck and lint failures across specs, guards, and docs

Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* fix(config): deduplicate parseClientIdentifiers and pass the raw variable

The duplicated helper also ignored its parameter while the config factory
read the raw variable at the call site with none, so the module never
compiled. Collapse to a single parser that takes the raw value.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

* Provide spending limit dependency in transaction tests

* Fix merged audit repository test structure

* Fix merged API audit and webhook tests

* Align policy velocity test with service architecture

* Preserve webhook failure reason in audit

* Fix duplicate event emitter declarations

* Fix: Remove duplicate db:verify key breaking package.json JSON parsing

A duplicate db:verify script entry (missing comma) made package.json
invalid JSON, causing npm ci to fail in CI with EJSONPARSE.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fix: Remove redundant transactionId index causing migration drift

RiskAssessment.transactionId is already @unique, which Prisma covers
with its own index; the extra @@index([transactionId]) had no
corresponding migration and tripped the CI drift check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: aetheron06 <asynchronousawait@gmail.com>
Co-authored-by: tecch-wiz <zynstro@gmail.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: AdaBebe0 <adabebe4040@gmail.com>
Co-authored-by: Isaac Gellu <113101848+gelluisaac@users.noreply.github.com>
Co-authored-by: IyanuOluwaJesuloba <jesulobaowoseni1@gmail.com>
Co-authored-by: Depo-dev <ed4557316@gmail.com>
Co-authored-by: S13 <61961655+samad13@users.noreply.github.com>
Co-authored-by: Codebuff <noreply@codebuff.com>
Co-authored-by: IyanuOluwa Owoseni <141356521+IyanuOluwaJesuloba@users.noreply.github.com>
Co-authored-by: Depo.dev <depolonedev@outlook.com>
Co-authored-by: Astroid Dev <astroid@dev.local>
Co-authored-by: Chijioke Joseph <chijiokejoseph2022@gmail.com>
Co-authored-by: xeladev4 <xeladev4@gmail.com>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-authored-by: gelluisaac <isaacgellu6@gmail.com>
Co-authored-by: eliasaph01 <eliasaphmarkus405@gmail.com>
Co-authored-by: DarcKnight000 <ahmadry2508@gmail.com>
Co-authored-by: Oladayo Oladipupo <oladayo225@gmail.com>
Co-authored-by: Deon <110722148+0xDeon@users.noreply.github.com>
Co…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add automated health check endpoint with dependency probing

2 participants