Repository navigation
feat(health): add /health/live and /health/ready probes with dependency checks - #362
Merged
Cjay-Cyber-2 merged 1 commit intoSep 29, 2026
Conversation
…cy checks Add orchestrator-grade liveness and readiness probes: - GET /health/live returns 200 whenever the process is running and performs no dependency checks, so a downstream outage never triggers a restart. - GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a per-dependency status, latency and error report under `services`. - Both probes are served outside the global API prefix, like /metrics, so probe paths are stable across API versions. Make the health controller usable by load balancers: - Mark it @public(); previously every health route required a JWT. - Exempt it from both named throttler tiers so probes cannot receive 429s. - Exclude it from the audit trail so probes do not write an audit row per request (or attempt to while the database is down). Fix the Redis indicator to probe the shared REDIS_CLIENT built from the validated REDIS_* config. It previously read an undefined REDIS_URL, always probed localhost:6379, kept its own never-closed connection, and could hang while ioredis queued the PING during an outage. Closes ASTROIDX556#351
|
@AdaBliss Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits. You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀 |
dakwa001
pushed a commit
to dakwa001/astroid-api
that referenced
this pull request
Oct 2, 2026
) Add orchestrator-grade liveness and readiness probes: - GET /health/live returns 200 whenever the process is running and performs no dependency checks, so a downstream outage never triggers a restart. - GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a per-dependency status, latency and error report under `services`. - Both probes are served outside the global API prefix, like /metrics, so probe paths are stable across API versions. Make the health controller usable by load balancers: - Mark it @public(); previously every health route required a JWT. - Exempt it from both named throttler tiers so probes cannot receive 429s. - Exclude it from the audit trail so probes do not write an audit row per request (or attempt to while the database is down). Fix the Redis indicator to probe the shared REDIS_CLIENT built from the validated REDIS_* config. It previously read an undefined REDIS_URL, always probed localhost:6379, kept its own never-closed connection, and could hang while ioredis queued the PING during an outage. Closes ASTROIDX556#351
tecch-wiz
pushed a commit
to tecch-wiz/astroid-api
that referenced
this pull request
Oct 2, 2026
) Add orchestrator-grade liveness and readiness probes: - GET /health/live returns 200 whenever the process is running and performs no dependency checks, so a downstream outage never triggers a restart. - GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a per-dependency status, latency and error report under `services`. - Both probes are served outside the global API prefix, like /metrics, so probe paths are stable across API versions. Make the health controller usable by load balancers: - Mark it @public(); previously every health route required a JWT. - Exempt it from both named throttler tiers so probes cannot receive 429s. - Exclude it from the audit trail so probes do not write an audit row per request (or attempt to while the database is down). Fix the Redis indicator to probe the shared REDIS_CLIENT built from the validated REDIS_* config. It previously read an undefined REDIS_URL, always probed localhost:6379, kept its own never-closed connection, and could hang while ioredis queued the PING during an outage. Closes ASTROIDX556#351
Cjay-Cyber-2
added a commit
that referenced
this pull request
Oct 7, 2026
* feat: address requested issues - closes #264, closes #263, closes #262, closes #261
* feat: Add comprehensive observability, security, and simulation improvements
Implements four major feature requests for enhanced API observability,
security, and transaction simulation capabilities:
#247 - Prometheus Metrics Interceptor
- Create MetricsInterceptor for HTTP request metrics collection
- Add request duration histograms, active request gauges, and request counters
- Categorize metrics by route, method, and status code
- Exclude /metrics endpoint from self-instrumentation
- Add comprehensive unit tests (9 tests)
#246 - Stellar Transaction Simulation Service
- Create StellarSimulationService for transaction simulation
- Integrate with Soroban RPC for XDR validation and simulation
- Include risk assessment and fee estimation
- Add circuit breaker protection for RPC failures
- Add XDR validation helper method
- Add comprehensive unit tests with mocked Stellar RPC (24 tests)
#245 - Cryptographic API Key Hashing Upgrade
- Upgrade from SHA-256 to Argon2id for enhanced security
- Implement memory-hard algorithm resistant to GPU/ASIC attacks
- Add timing-attack resistant comparison via Argon2 verification
- Maintain SHA-256 fallback for backward compatibility
- Update ApiKeyService with dual-algorithm verification
- Add comprehensive unit tests (32 crypto tests, 15 API key tests)
#244 - Webhook Retry and Dead Letter Queue
- Verify existing implementation meets all requirements
- Confirm exponential backoff with jitter (2000ms base, 20% jitter)
- Confirm 5 max attempts and non-transient error detection
- Verify dead-letter handler via DeadLetterService
- All existing tests passing
Closes #247, #246, #245, #244
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(workers): centralize job error handling with scrubbed structured logging
Add runWorkerJob, a wrapper every background worker now routes its handler
through. It times the job via WorkerMetricsService when available, logs a
structured completion record, and classifies failures before rethrowing:
transient failures log a job.retrying warning, while a failure on the final
attempt or an UnrecoverableError logs a job.dead-lettered error with the
scrubbed payload and stack. The original error is always rethrown untouched
so BullMQ retry semantics are preserved, and logging can never mask it.
Add scrubForLog/scrubString, which redact sensitive keys (sharing the audit
sanitizer's key list) and secret-shaped substrings such as Stellar seeds,
bearer tokens and URL credentials, and coerce cycles, bigints and errors
into JSON-safe values.
QueueFailureListener now scrubs its log line as well; previously the raw
webhook payload, including its signing secret, was written to the error
log. The dead-letter copy keeps the raw payload so re-drive still works.
* feat(pagination): add offset/limit pagination to list endpoints
Extend the shared list query contract used by every list endpoint with
an offset parameter alongside the existing page parameter, and bound
the page size.
- limit defaults to 50 and is capped at 200
- offset defaults to 0; page remains supported as an alternative, and
supplying both is rejected
- negative, non-integer, non-numeric or out-of-range values are
rejected by the Zod pipe with 400 Bad Request
- Prisma queries use bound skip/take values and an allow-listed sort
column, so no pagination input reaches SQL as text
- meta now includes offset, and hasNext/hasPrev are computed from the
offset so unaligned slices report correctly
- paginated responses set an X-Total-Count header, exposed via CORS
- replace duplicated Swagger page/limit docs with ApiPaginationQuery
- add unit tests and an HTTP-level integration test
* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping
* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)
* perf(analytics): batch dashboard overview queries into one transaction
Closes #338
AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.
* test(audit): cover pagination and sorting edge cases for activity log
Closes #339
The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.
* fix(migrations): resolve colliding timestamp between two merged migrations
Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.
* fix(ci): document missing env vars and fix flaky retry.util test
Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:
- docs/configuration.md was missing 7 env vars added by recent merges
(DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
fake timers with vi.runAllTimersAsync(), then attach the rejection
assertion afterward — a race that surfaces as an unhandled rejection
under full-suite load (deterministic once >100 files run together).
Attaching a no-op .catch() immediately after creating the promise
prevents the unhandled state without changing what each test asserts.
* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)
* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)
* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)
* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)
* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)
* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)
* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)
* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)
* Add startup migration check and DB pool connection metrics
Wires the existing migration status checker into app bootstrap so the
process halts before accepting traffic when prisma/migrations has
pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt,
the default; 'warn' logs and continues). Gated behind
DATABASE_MIGRATION_CHECK_ENABLED.
Adds a db_pool_connections Prometheus gauge (active/idle/waiting)
sourced from pg_stat_activity, since Prisma's Rust query engine
doesn't expose pool internals through the Node client.
Closes #335
Closes #332
* Fix CI: document missing env vars, fix flaky retry backoff test
- docs/configuration.md was missing entries for DATABASE_SLOW_QUERY_THRESHOLD_MS,
DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, and the
PUBLIC_RATE_LIMIT_* vars, failing the configuration-documentation test.
- retry.util.spec.ts left three rejected promises unhandled between
`runAllTimersAsync()` and the `expect(...).rejects` assertion that
attaches the handler; attach a no-op .catch() immediately after creating
each promise so fake-timer-driven rejections don't fire as unhandled
rejections mid-test-run.
* Fix pre-existing typecheck/lint/test failures blocking CI
main's build/typecheck/lint/test were already broken before this branch
touched anything (confirmed by checking out upstream/main directly).
CI enforces these repo-wide, so they block this PR too. Fixed each:
- event-names.ts: duplicate object key (TransactionRiskScoringRequested)
- throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't
exist on that type (JWT payload field name leaked into the wrong type)
- sliding-window-throttler.guard.ts: removed a user-tier rate-limit
multiplier keyed on AuthenticatedUser.tier, a field never present
anywhere in the auth/user model — dead, unbacked logic. Removed its
now-orphaned tests too.
- agent.controller.ts: AstroidThrottlerGuard used but never imported;
dropped an unused SlidingWindowThrottlerGuard import block
- risk.service.ts: event-driven risk scoring built a RiskFactorsInput
with fields (destination/velocityCount/isNewRecipient) that don't
exist on the current type; mapped to the real shape instead
- risk.service.spec.ts: removed orphaned unused fixtures
- stellar.service.ts (src/modules/stellar/services, unused elsewhere
in the app but still typechecked/tested): getTransactionInfo called
a client method that doesn't exist (real method is getTransaction);
simulateTransaction passed a bare string where the client expects
an options object
- stellar.service.spec.ts: rewritten against the real
SorobanSimulationResult shape; fixed mockResolvedValueOnce/
mockRejectedValueOnce being consumed by the test's own first
assertion, leaving the second call unmocked
- transaction.service.spec.ts: rewritten against TransactionService's
actual create() contract (it doesn't call Soroban simulation at all;
the previous spec tested a flow that was never implemented) and a
real Ed25519 checksum address
- sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist
on this Express-platform app; switched to app.listen + fetch,
matching the sibling public-rate-limit.integration.spec.ts pattern,
and named the test throttler 'api' so AstroidThrottlerGuard's
tier-matching actually engages it
* Fix remaining pre-existing lint errors and undocumented env vars
CI runs lint and test repo-wide, so these also blocked the PR:
- 4 pre-existing no-explicit-any lint errors in throttler guard code
and specs, typed properly instead of suppressed
- 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST,
WEBHOOK_BURST) were validated by the env schema but missing from
docs/configuration.md, failing the docs-sync test
* perf(auth): cache session revocation answers during token verification
Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.
- Add CacheService: a small get/set/delete cache over the shared
REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
prefix invalidation. Every operation degrades to a no-op/miss on Redis
failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
answers for TOKEN_CACHE_TTL seconds (default 30, well below the
15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
drops the cached answer after each blacklist write (including on
Redis-outage fallback), AuthService invalidates on logout and on
refresh rotation, so revocations are observed immediately instead of
after the TTL window.
Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.
Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.
Closes #341
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* feat: correlate requests, paginate audit, sign webhooks
* feat(rate-limit): configurable limits and client identifiers for public routes
Public endpoints are the first thing abusive traffic hits, but their
limits were a single fixed IP budget: every public route shared
PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one
shared address (NAT, office egress, CI runners) exhausted one bucket
together. This makes the public rate limiter configurable per route and
per client identifier, on the existing Redis sliding-window counter.
- Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or
per-controller) budget overrides resolved by the guard through
Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default.
- Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated,
currently 'apiKey'): when enabled, a presented x-api-key or
ApiKey/Bearer ak_... Authorization header is folded into the bucket
key so distinct key-holding clients behind one IP get their own
budgets. The IP always participates; keyless callers share the
plain-IP bucket as before.
- Extract a testable PublicRateLimitGuard.check() returning the full
decision (allowed, limit, windowSeconds, count, resetAt); canActivate
keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset +
Retry-After contract and in-memory fallback on Redis outage.
Tests: unit suites for per-route rule resolution, identifier bucketing
(with/without identifiers configured), header correctness on allowed
and limited requests, and @SkipPublicRateLimit() interaction with rules;
an HTTP-level integration suite simulating bursts that proves the
429-with-headers behaviour at the global limit, per-route overrides
(next to unaffected sibling routes), and per-key budget isolation.
Existing public-rate-limit suites pass unchanged (bucket keys keep the
ip: prefix, so stored counters stay compatible).
Closes #342
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* feat: stream audit exports and harden payment safety (#63, #64, #283)
* fix: validate nested transaction metadata as JSON (#283)
* feat(transactions): implement agent spending limit evaluation guard
* fix(ci): resolve typecheck and lint failures across specs, guards, and docs
Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* fix(ci): resolve typecheck and lint failures across specs, guards, and docs
Get CI green for the token-verification cache PR by fixing pre-existing
main-branch type errors alongside PR-specific ones: Express specs no longer
use Fastify-only app.inject, Stellar mocks match the real Soroban result
interface, the transaction spec exercises the actual create pipeline,
TokenBlacklistService resolves the global REDIS_CLIENT token explicitly,
and the configuration docs cover every THROTTLE_* env var the docs test
asserts.
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* fix(config): deduplicate parseClientIdentifiers and pass the raw variable
The duplicated helper also ignored its parameter while the config factory
read the raw variable at the call site with none, so the module never
compiled. Collapse to a single parser that takes the raw value.
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* feat(audit): implement structured audit logging interceptor for sensitive operations
Closes #255
Adds the @AuditLog() decorator (with optional action/entity metadata) and reworks AuditLogInterceptor so only decorated routes are persisted, keeping read-only traffic free of audit writes.
Each record captures the actor (human user id or acting agent), client IP, HTTP method, request path, a SHA-256 fingerprint of the sanitized payload, the sanitized body itself, the final response status and handler duration. Secrets, keys, tokens, signatures and mnemonics are redacted recursively without mutating the original request body, and @SkipAudit() always wins.
Applied to the sensitive policy, budget and API-key endpoints; persistence stays fire-and-forget so a failed audit write never breaks the client request.
* refactor(policies): add Prisma repository abstraction for agent spending policies
Closes #254
Introduces SpendingPolicyRepository (src/modules/policies/spending-policy.repository.ts) as the single owner of every Prisma call for policies: create/find/update/soft-delete, the paginated read, the rolling spend aggregation used by velocity checks, the POLICY_EVALUATED audit row, and an interactive withTransaction() helper. Every method funnels through one error-handling wrapper that logs the operation (plus the Prisma error code) and rethrows the original error, so error handling and transaction safety are uniform across the repository.
Adds SpendingPolicyService, which injects the repository and now owns spending-policy validation and enforcement (staleness checks, daily velocity limit, evaluation audit) with no direct Prisma access. PolicyService is reduced to orchestration - domain events plus the pure PolicyEngine - and delegates all persistence to SpendingPolicyService, which removes the direct prisma.transaction/prisma.auditLog calls from the service layer. PolicyRepository is replaced by the new, better-named repository.
Unit tests cover the repository against a mocked Prisma client (including transaction usage, Decimal summing and error propagation) and the service against a mocked repository (validation, pagination, not-found, velocity limits and audit failure swallowing).
* feat(webhooks): add BullMQ retry policy and dead-letter audit for webhook deliveries
Closes #253
Centralises the webhook queue configuration in src/queues/webhook.queue.ts: 5 attempts, exponential backoff with a 2000ms base, the jittered custom backoff strategy, 24h retention of failed jobs and the dead-letter destination. Both the queue registration and the enqueue path consume the same constants, so the API and the worker can no longer drift apart.
Adds WebhookAuditService, which appends a WEBHOOK_DELIVERY_FAILED record (subscriber, event, attempts, HTTP status, failure reason) to the audit trail for deliveries that will never be retried - either an unrecoverable 4xx or the final attempt after retries are exhausted. Both the webhook processor and the worker call it through a fire-and-forget hook that swallows and logs its own failures, so audit problems can never mask the delivery error, crash the NestJS process or interfere with BullMQ retry/backoff.
Tests cover the retry/backoff/DLQ configuration and jitter envelope, the audit service (including persistence failures) and the processor's terminal-failure behaviour - retry without auditing while attempts remain, exactly one audit entry when retries are exhausted or a 4xx is returned, and error preservation when auditing fails.
* feat(rate-limit): add Redis-backed agent throttler guard for high-frequency agent endpoints
Closes #251
Adds AgentThrottlerGuard in src/common/guards/agent-throttler.guard.ts, extending @nestjs/throttler's ThrottlerGuard so counters keep running through the shared RedisThrottlerStorage while the guard adds agent-aware behaviour: the bucket key is derived from the acting agent (x-agent-id, a route/body/query agentId, or an agent-bound API-key principal) instead of the organization or IP, falling back to the organization, then a hashed API key, then the client IP (honouring x-forwarded-for) for unauthenticated public routes.
Introduces a third 'agent' tier (THROTTLE_AGENT_LIMIT, default 300/window) alongside api/auth in config/throttler.config.ts and env.validation.ts, and routes each request to exactly one tier - explicit @ThrottleTierDecorator() metadata wins, otherwise agent-identified traffic uses the agent tier and everything else the api tier. Rejections are standard 429s that also carry plain Retry-After, X-RateLimit-Limit and X-RateLimit-Remaining headers, and the tier limit is advertised on allowed responses too.
Applied selectively to the agent-facing controllers (agents, transactions, wallets) via @UseGuards so the existing global AstroidThrottlerGuard keeps enforcing api/auth untouched. Vitest coverage asserts tier routing, tracker resolution for agents/orgs/API keys/anonymous callers, header emission and 429 enforcement when a single agent's burst exhausts its budget.
* Fix CI: pre-existing typecheck/lint/test breakage inherited from main
main's HEAD was red (27 typecheck errors) from unrelated merged PRs
(#357, #374, #381-#388). Fixed to get this branch's CI green:
- throttler.guard.ts / sliding-window-throttler.guard.ts: AuthenticatedUser
has no `sub` or `tier` field; use `.id` and treat `tier` as an optional
extension until the type actually carries it.
- agent.controller.ts: imported SlidingWindowThrottlerGuard/SlidingWindowLimit
(unused) instead of the AstroidThrottlerGuard actually referenced by
@UseGuards.
- event-names.ts: dropped a duplicate TransactionRiskScoringRequested key.
- risk.service.ts: the TransactionCreated handler built a RiskFactorsInput
with fields (`destination`, `velocityCount`, `isNewRecipient`) that don't
exist on the type; mapped to the real shape instead. Removed dead
lowRisk/createEventBus fixtures left over in risk.service.spec.ts.
- stellar.service.ts (Soroban variant): fixed getTransactionInfo calling a
nonexistent client method (getTransaction), simulateTransaction being
called with a bare string instead of {transactionXdr}, and an error
message interpolating the whole error object instead of `.message`.
Rewrote stellar.service.spec.ts's mocks/assertions to match the real
SorobanSimulationResult shape, and fixed two tests reusing a `mockOnce`
across two separate calls to the service (second call fell through to
the unmocked default and threw on undefined).
- transaction.service.spec.ts targeted a pre-broadcast Soroban simulation
step that was never wired into TransactionService.create, against a
StellarService overload TransactionService doesn't even inject; skipped
with an explanation rather than fabricating the feature.
- sensitive-rate-limit.integration.spec.ts: replaced Fastify-only
`app.inject()` (this app runs on platform-express) with a small
http.request helper; the throttler in the test module was unnamed
('default'), so AstroidThrottlerGuard's per-tier name match against the
route's default 'api' tier always skipped it — named it 'api' to match.
- Removed remaining `any` usages in the touched guard files/specs.
- docs/configuration.md: documented THROTTLE_WEBHOOK_LIMIT,
THROTTLE_API_BURST, THROTTLE_AUTH_BURST, THROTTLE_WEBHOOK_BURST (missing
from #386), failing the configuration-documentation test.
* Fix CI: repair botched merge of main into this branch
A merge of upstream main (PR #372's overlapping CI fixes) into this
branch left several files with duplicated/glued content instead of
resolved conflicts:
- sliding-window-throttler.guard.spec.ts: duplicate makeContext()
declaration (one-line + multi-line versions concatenated).
- sliding-window-throttler.guard.ts: `limit` reverted to `const` while
the tier-adjustment branches below still reassign it.
- sensitive-rate-limit.integration.spec.ts: my raw http.request-based
version and another (cleaner, fetch()-based) version from main were
concatenated rather than merged — duplicate imports, an unclosed `it`
block. Kept the fetch()-based version.
- stellar.service.spec.ts: duplicate `error` key in a SorobanSimulationResult
object literal.
- transaction.service.spec.ts: my placeholder describe.skip stub was
prepended to a real, correctly implemented version of the same suite
that showed up on main independently. Kept the real implementation,
dropped the stub.
* Remove conflict markers from worker merge
* feat: implement risk scoring persistence and complete API documentation
Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.
Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository
Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios
Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules
Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages
* fix: add missing relation field in RiskAssessment model
Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.
* fix: resolve TypeScript errors in test files
Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests
* fix: add null checks for toPromise() results in response interceptor tests
Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.
* fix: add eslint-disable comments for temporary any types
Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.
* fix: update test expectation for date comparison in statistics test
Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.
* fix: remove duplicate Zod validation pipe integration test
* fix ci
* fix ci
* fix failing ci
* fix ci
* fix ci
* fix ci
* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)
* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files
Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* fix: use ZodValidationException in integration spec
The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: address requested issues - Done with all issues (#357)
* feat(health): dedicated GET /health/redis probe (#358)
Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.
* test(auth): lock in auth throttle-tier wiring on public routes (#359)
Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(health): add /health/live and /health/ready probes (#362)
Add orchestrator-grade liveness and readiness probes:
- GET /health/live returns 200 whenever the process is running and performs
no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
probe paths are stable across API versions.
Make the health controller usable by load balancers:
- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
request (or attempt to while the database is down).
Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.
Closes #351
* feat(config): validate full environment at startup (#363)
Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.
- On failure the process prints every failing variable in one message and
exits with code 1, instead of surfacing only the first failing slice from
inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
set explicitly or implied by omission) and a JWT refresh secret that
reuses the access secret.
Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.
Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.
Closes #350
* fix(tests): add SpendingLimitService mock to TransactionService spec
* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)
Closes #331
Closes #325
Closes #334
Closes #328
- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
(WebhookWorker), keeping WebhooksProcessor which also records metrics.
Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
new public ingress route for Redis-backed sliding-window rate limiting
with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
per stream, exposed through the existing Prometheus registry.
* feat: add production rollback protection to database migration CLI (#371)
Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.
down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.
The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.
* test: add full branch coverage tests for role, permission and scope guards (#373)
Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.
The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.
Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.
* perf: add concurrent composite indexes for notification, approval and memory (#374)
Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:
- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline
Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.
* feat: add IP-based rate limiting to public endpoints (#377)
Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.
- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
window script, so all replicas enforce one budget per IP; rejected
requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test
* feat: return RFC 9457 problem details for all error responses (#378)
Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:
{ type, title, status, detail, instance, code, requestId, details? }
- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
validation, authentication, domain, not-found and server errors
* fix: worker error handling + log scrubbing (#370)
* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* test: add job-worker spec and queue-failure-listener scrubbing test
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)
* perf(analytics): batch dashboard overview queries into one transaction
Closes #338
AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.
* test(audit): cover pagination and sorting edge cases for activity log
Closes #339
The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.
* fix(migrations): resolve colliding timestamp between two merged migrations
Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.
* fix(ci): document missing env vars and fix flaky retry.util test
Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:
- docs/configuration.md was missing 7 env vars added by recent merges
(DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
fake timers with vi.runAllTimersAsync(), then attach the rejection
assertion afterward — a race that surfaces as an unhandled rejection
under full-suite load (deterministic once >100 files run together).
Attaching a no-op .catch() immediately after creating the promise
prevents the unhandled state without changing what each test asserts.
* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)
* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)
* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)
* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)
* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)
* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)
* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)
* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)
* fix: resolve rebase conflicts and restore CI
* feat: implement risk scoring persistence and complete API documentation
Implement comprehensive risk scoring service with data persistence for
compliance tracking, complete OpenAPI documentation coverage, and add
integration tests for core infrastructure components.
Risk Scoring Service (#213):
- Add RiskRepository for assessment record persistence and historical analysis
- Update RiskService to persist assessment records via repository
- Add getHistory() and getStatistics() methods for risk analytics
- Add RiskAssessment model to Prisma schema with proper indexes
- Create database migration for risk assessments table
- Add comprehensive integration tests for RiskRepository
Response Envelope Interceptor (#216):
- Verify ResponseInterceptor is registered globally in app.module.ts
- Add integration tests for response wrapping and pagination handling
- Test requestId header handling and null data scenarios
Swagger Documentation (#218):
- Add @ApiProperty decorators to admin DLQ and queue management DTOs
- Add @ApiProperty decorators to audit export DTOs
- Add @ApiProperty decorators to risk assessment DTOs
- Update risk controller to use DTO for request body documentation
- Complete OpenAPI documentation coverage across all domain modules
Zod Validation Pipe (#220):
- Verify ZodValidationPipe supports custom error formatting
- Add integration tests for validation scenarios and error handling
- Test optional fields, nested objects, arrays, and custom messages
* fix: add missing relation field in RiskAssessment model
Add organization relation field to RiskAssessment model to fix
Prisma schema validation error. Update migration to add foreign
key constraint separately for proper schema validation.
* fix: resolve TypeScript errors in test files
Fix TypeScript errors in test files by:
- Converting async tests to use toPromise() instead of callbacks
- Adding proper type assertions for exception details
- Adding RiskRepository to RiskService constructor
- Using type assertions for Prisma client methods until migration runs
- Adding missing PaginationMeta properties in tests
* fix: add null checks for toPromise() results in response interceptor tests
Add null checks after toPromise() calls to resolve TypeScript
'possibly undefined' errors in response interceptor tests.
* fix: add eslint-disable comments for temporary any types
Add eslint-disable comments for @typescript-eslint/no-explicit-any
where we use type assertions for Prisma client methods until
the migration runs and generates the proper types.
* fix: update test expectation for date comparison in statistics test
Change the test expectation to use expect.objectContaining for the
nested createdAt object instead of expect.any(Date) for the entire
field, as the Date object is being compared with its actual value.
* fix: remove duplicate Zod validation pipe integration test
* feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356)
* feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files
Resolves the merge conflict from PR #284 by applying all genuinely new
additions (RiskRepository, schema migration, test specs, service updates)
while keeping main's already-improved Swagger DTO implementations.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* fix: use ZodValidationException in integration spec
The pipe throws ZodValidationException (a BadRequestException subclass
defined in zod-validation.pipe.ts), not the domain-layer ValidationException.
Update all assertions to match the actual thrown type.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: address requested issues - Done with all issues (#357)
* feat(health): dedicated GET /health/redis probe (#358)
Adds GET /health/redis, mirroring GET /health/database, so container
orchestration and uptime monitors can probe Redis alone instead of only
seeing it as one service inside /health/readiness. Reports status,
latency and the ping error, with 200/503 semantics, plus unit tests for
the up, down and isolation paths.
* test(auth): lock in auth throttle-tier wiring on public routes (#359)
Guard and config behavior for rate limiting is already covered in
isolation, but nothing asserted that register/login/refresh actually
declare the auth tier. Adds a regression test against the decorator
metadata so removing @ThrottleTierDecorator('auth') from a handler
fails CI instead of silently dropping back to the looser api limit.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(health): add /health/live and /health/ready probes (#362)
Add orchestrator-grade liveness and readiness probes:
- GET /health/live returns 200 whenever the process is running and performs
no dependency checks, so a downstream outage never triggers a restart.
- GET /health/ready probes the database (SELECT 1) and cache (Redis PING)
in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a
per-dependency status, latency and error report under `services`.
- Both probes are served outside the global API prefix, like /metrics, so
probe paths are stable across API versions.
Make the health controller usable by load balancers:
- Mark it @Public(); previously every health route required a JWT.
- Exempt it from both named throttler tiers so probes cannot receive 429s.
- Exclude it from the audit trail so probes do not write an audit row per
request (or attempt to while the database is down).
Fix the Redis indicator to probe the shared REDIS_CLIENT built from the
validated REDIS_* config. It previously read an undefined REDIS_URL, always
probed localhost:6379, kept its own never-closed connection, and could hang
while ioredis queued the PING during an outage.
Closes #351
* feat(config): validate full environment at startup (#363)
Compose the per-slice Zod schemas into a single `environmentSchema` and
validate `process.env` against it at the start of bootstrap(), before any
Nest module is constructed or any connection is opened.
- On failure the process prints every failing variable in one message and
exits with code 1, instead of surfacing only the first failing slice from
inside Nest's module initialization with a stack trace.
- Messages are value-free (e.g. "must be one of: ..." rather than Zod's
default "received '<value>'") so secrets never reach logs.
- In production, reject the publicly known default ENCRYPTION_KEY (whether
set explicitly or implied by omission) and a JWT refresh secret that
reuses the access secret.
Per-slice validation in each registerAs factory is unchanged, so the typed
ConfigService namespaces keep their guarantees outside main.ts.
Document every variable, its type, default and production rules in
docs/configuration.md, and correct the README's list of required
variables. Tests keep the docs and .env.example in sync with the schema,
and exercise the real main.ts entrypoint to prove missing or malformed
variables halt startup before NestFactory.create is called.
Closes #350
* feat: webhook ingress validation, retry queue cleanup, rate limiting (#367)
Closes #331
Closes #325
Closes #334
Closes #328
- Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the
existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard,
with a Zod schema validating the payload shape.
- Removed duplicate BullMQ processor consuming the webhooks queue
(WebhookWorker), keeping WebhooksProcessor which also records metrics.
Deleted dead WebhookDeliveryWorker, never registered as a real processor.
- Wired the existing, previously unused SlidingWindowThrottlerGuard onto the
new public ingress route for Redis-backed sliding-window rate limiting
with standard rate-limit headers.
- Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation
per stream, exposed through the existing Prometheus registry.
* feat: add production rollback protection to database migration CLI (#371)
Add a migration CLI (npm run db:migrate -- <command>) with a production
safety guard. The destructive commands, down and reset, are rejected when
NODE_ENV=production unless --force is supplied. A blocked run prints a
warning to stderr, exits with code 1 and never contacts the database. A
forced run is also announced on stderr.
down <migration> runs the migration's hand-written down.sql and removes
its _prisma_migrations row in a single script, so the history row is only
dropped when every rollback statement succeeded. Migration names are
validated against the folder on disk before being used in SQL.
The guard logic lives in src/database/migration-guard.ts with injected
process dependencies so it is fully unit tested. The entry point
src/database/migrate.cli.ts is compiled into dist and can run in images
without ts-node. Usage is documented in docs/database.md.
* test: add full branch coverage tests for role, permission and scope guards (#373)
Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard,
including matchScope, and extend the RbacGuard spec. The guards now have
100% statement, branch, function and line coverage.
The tests attach real @Roles, @RequirePermissions, @RequireScopes and
@Public metadata to fixture controllers and run them through a real
Reflector, so handler-over-class inheritance is exercised rather than
mocked. Assertion matrices cover every UserRole, every wildcard shape in
matchScope, including nested scopes, and the AND semantics of multi-
permission routes.
Behaviour pinned by these tests:
- OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN
bypasses ScopesGuard; API-key principals with those roles do not.
- PermissionsGuard grants no role-based override and expands no wildcards.
- Guests and expired or revoked principals, which the auth strategies
leave without request.user, get a 401 on every restricted route.
- Stale or differently cased role names are rejected.
* perf: add concurrent composite indexes for notification, approval and memory (#374)
Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but
the tables behind the busiest ones only had single-column indexes. Postgres
therefore had to fetch all of a tenant's or user's rows, or walk the global
createdAt index and filter, before it could return a page. These composite
indexes match the actual query shapes:
- notifications (userId, createdAt): inbox list
- notifications (organizationId, userId, read, createdAt): unread badge,
mark-all-read, unread filter (index-only count)
- proposals (organizationId, createdAt): approval queue
- proposals (organizationId, status, createdAt): pending count, status filter
- memory_records (organizationId, createdAt): memory browser
- memory_records (agentId, createdAt): per-agent memory timeline
Each index is built with CREATE INDEX CONCURRENTLY, so writes are never
blocked. Each one lives in its own single-statement migration, because
Prisma runs a multi-statement migration as one implicit transaction, where
Postgres rejects CONCURRENTLY. The matching @@index entries are added to
schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md
records the convention and how to recover from a failed concurrent build.
* feat: add IP-based rate limiting to public endpoints (#377)
Add PublicRateLimitGuard, a global guard that applies a per-IP sliding
window limit to every unauthenticated route: handlers marked @Public()
and any route under /<API_PREFIX>/public/. Requests beyond the limit
are rejected with 429 Too Many Requests and a Retry-After header.
- X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set
on every limited response
- counters live in the shared REDIS_CLIENT via an atomic Lua sliding
window script, so all replicas enforce one budget per IP; rejected
requests are not recorded
- if Redis is unavailable the guard falls back to an in-memory window
per instance instead of failing open
- thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED,
PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and
PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default)
- @SkipPublicRateLimit() exempts routes; applied to the
network-restricted /metrics scrape endpoint
- add store and guard unit tests plus an HTTP burst integration test
* feat: return RFC 9457 problem details for all error responses (#378)
Replace the { success, error: { code, message }, requestId } error
envelope with a uniform problem details body served as
application/problem+json:
{ type, title, status, detail, instance, code, requestId, details? }
- type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP
errors without a dedicated code use about:blank with the reason phrase
- title comes from a new ERROR_TITLE map kept exhaustive by the type
system; instance is the request path without its query string
- code and requestId are kept as extension members so clients can keep
switching on the machine-readable code
- ZodValidationException now keeps its VALIDATION_ERROR code and
field-level details instead of collapsing to BAD_REQUEST
- unhandled exceptions still map to a generic 500 INTERNAL_ERROR
without leaking internals
- update the shared response types and API documentation
- rewrite the filter spec and add an HTTP integration test covering
validation, authentication, domain, not-found and server errors
* fix: worker error handling + log scrubbing (#370)
* fix: worker error handling + log scrubbing (resolve PR #370 conflicts)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* test: add job-worker spec and queue-failure-listener scrubbing test
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8
* Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369)
* perf(analytics): batch dashboard overview queries into one transaction
Closes #338
AnalyticsService.overview() issued 7 independent queries (counts,
spend aggregates, status/risk group-bys) via Promise.all, each its own
roundtrip. Batches the 3 counts and 2 aggregates into a single
$transaction([...]) call; the 2 groupBy calls stay outside the batch
since Prisma's groupBy return type doesn't infer correctly inside a
$transaction array. Response shape is unchanged. Existing indexes on
organizationId/status/createdAt already cover these queries.
* test(audit): cover pagination and sorting edge cases for activity log
Closes #339
The audit log (this repo's activity log) already supported
page/limit/sort/order/filter query params and returned pagination
metadata (total, totalPages, hasNext, hasPrev), with limit capped at
100. Adds unit test coverage for the previously-untested list()
method: normal pagination, empty results, out-of-bounds pages,
invalid sort field fallback, ascending order, and entity filtering.
* fix(migrations): resolve colliding timestamp between two merged migrations
Migrations 20260928120000_add_agent_contribution_stats_index and
20260928120000_add_notifications_user_created_at_index landed with the
same 14-digit timestamp prefix from two separately merged PRs (#374,
#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps
the notifications index migration to 20260928120001; both migrations
are independent, additive CREATE INDEX statements with no ordering
dependency between them, so the rename is safe.
* fix(ci): document missing env vars and fix flaky retry.util test
Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking
this branch:
- docs/configuration.md was missing 7 env vars added by recent merges
(DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS,
DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which
env.validation.spec.ts asserts against. Documented all 7.
- retry.util.spec.ts had 3 tests that create a rejecting promise, advance
fake timers with vi.runAllTimersAsync(), then attach the rejection
assertion afterward — a race that surfaces as an unhandled rejection
under full-suite load (deterministic once >100 files run together).
Attaching a no-op .catch() immediately after creating the promise
prevents the unhandled state without changing what each test asserts.
* feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388)
* Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381)
* Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384)
* Fix #223: Implement Stellar Transaction Simulation Service Integration (#385)
* Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382)
* Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383)
* feat(interceptors): add global RequestIdInterceptor for correlation logging (#387)
* feat(throttler): add Redis-backed rate limiting guard and throttler config (#386)
* Add startup migration check and DB pool connection metrics
Wires the existing migration status checker into app bootstrap so the
process halts before accepting traffic when prisma/migrations has
pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt,
the default; 'warn' logs and continues). Gated behind
DATABASE_MIGRATION_CHECK_ENABLED.
Adds a db_pool_connections Prometheus gauge (active/idle/waiting)
sourced from pg_stat_activity, since Prisma's Rust query engine
doesn't expose pool internals through the Node client.
Closes #335
Closes #332
* Fix pre-existing typecheck/lint/test failures blocking CI
main's build/typecheck/lint/test were already broken before this branch
touched anything (confirmed by checking out upstream/main directly).
CI enforces these repo-wide, so they block this PR too. Fixed each:
- event-names.ts: duplicate object key (TransactionRiskScoringRequested)
- throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't
exist on that type (JWT payload field name leaked into the wrong type)
- sliding-window-throttler.guard.ts: removed a user-tier rate-limit
multiplier keyed on AuthenticatedUser.tier, a field never present
anywhere in the auth/user model — dead, unbacked logic. Removed its
now-orphaned tests too.
- agent.controller.ts: AstroidThrottlerGuard used but never imported;
dropped an unused SlidingWindowThrottlerGuard import block
- risk.service.ts: event-driven risk scoring built a RiskFactorsInput
with fields (destination/velocityCount/isNewRecipient) that don't
exist on the current type; mapped to the real shape instead
- risk.service.spec.ts: removed orphaned unused fixtures
- stellar.service.ts (src/modules/stellar/services, unused elsewhere
in the app but still typechecked/tested): getTransactionInfo called
a client method that doesn't exist (real method is getTransaction);
simulateTransaction passed a bare string where the client expects
an options object
- stellar.service.spec.ts: rewritten against the real
SorobanSimulationResult shape; fixed mockResolvedValueOnce/
mockRejectedValueOnce being consumed by the test's own first
assertion, leaving the second call unmocked
- transaction.service.spec.ts: rewritten against TransactionService's
actual create() contract (it doesn't call Soroban simulation at all;
the previous spec tested a flow that was never implemented) and a
real Ed25519 checksum address
- sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist
on this Express-platform app; switched to app.listen + fetch,
matching the sibling public-rate-limit.integration.spec.ts pattern,
and named the test throttler 'api' so AstroidThrottlerGuard's
tier-matching actually engages it
* Fix remaining pre-existing lint errors and undocumented env vars
CI runs lint and test repo-wide, so these also blocked the PR:
- 4 pre-existing no-explicit-any lint errors in throttler guard code
and specs, typed properly instead of suppressed
- 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST,
WEBHOOK_BURST) were validated by the env schema but missing from
docs/configuration.md, failing the docs-sync test
* perf(auth): cache session revocation answers during token verification
Every authenticated request performed a Redis round trip against the
token blacklist to answer "is this session still revoked?". This adds a
short-TTL caching layer in front of the blacklist so repeated
verifications within one window skip the Redis query entirely.
- Add CacheService: a small get/set/delete cache over the shared
REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based
prefix invalidation. Every operation degrades to a no-op/miss on Redis
failure so caching can never break the request path.
- Add TokenVerificationCacheService: caches per-session revocation
answers for TOKEN_CACHE_TTL seconds (default 30, well below the
15-minute access-token lifetime) and exposes invalidation hooks.
- Wire the cache into JwtStrategy.validate (cache-first, source of truth
on miss, fail-open unchanged on Redis outages).
- Hook invalidation into every revocation path: TokenBlacklistService
drops the cached answer after each blacklist write (including on
Redis-outage fallback), AuthService invalidates on logout and on
refresh rotation, so revocations are observed immediately instead of
after the TTL window.
Revocation reliability is preserved because no cached answer outlives
its TTL, and explicit logout/rotation clears the entry at once.
Tests: unit suites for CacheService and TokenVerificationCacheService
(hits, misses, resolver fallback, invalidation hooks), an integration
suite proving repeated authentications trigger a single blacklist
lookup and that logout flips a cached-valid session to 401 immediately,
plus updated JwtStrategy/TokenBlacklistService/api-key integration
suites for the new wiring.
Closes #341
🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
* fix(ci): resolv…
Cjay-Cyber-2
added a commit
that referenced
this pull request
Oct 7, 2026
…ed CI job (#364) * feat: address requested issues - closes #264, closes #263, closes #262, closes #261 * ci(migrations): harden migration verification and run it as a dedicated CI job Expand scripts/verify-migrations.sh into three tiers of checks that report every problem in one run, as GitHub annotations in CI: 1. Structure and SQL (always): prisma validate; migration_lock.toml present and matching the schema provider; folder names follow Prisma's <timestamp>_<name> convention (plus the 0_ baseline) with unique timestamps; each folder has a migration.sql with at least one statement; no stray files. A comment-, string- and dollar-quote-aware SQL lexer (POSIX awk, runs under mawk) rejects merge-conflict markers, unbalanced parentheses, unterminated strings/identifiers/comments and a missing final semicolon, with file and line. 2. History (MIGRATION_BASE_REF): migrations already on the base branch must not be modified or deleted (deployed databases store their checksums), new migrations must sort after the newest base migration, and destructive statements in new migrations are reported as warnings. 3. Replay and drift (SHADOW_DATABASE_URL): replay the full history on an empty database with prisma migrate diff, catching invalid SQL and forward dependencies between migrations, and fail on schema drift. Run it in a dedicated verify-migrations job with its own disposable Postgres, full git history, and a base ref derived from the PR target or the previous push tip. This replaces the previous step, whose shadow database URL the old script never used. Make `npm run db:verify` invoke bash explicitly so it works on Windows, mark the script executable, and document the requirements in CONTRIBUTING.md and docs/database.md. Closes #256 * feat: Add comprehensive observability, security, and simulation improvements Implements four major feature requests for enhanced API observability, security, and transaction simulation capabilities: #247 - Prometheus Metrics Interceptor - Create MetricsInterceptor for HTTP request metrics collection - Add request duration histograms, active request gauges, and request counters - Categorize metrics by route, method, and status code - Exclude /metrics endpoint from self-instrumentation - Add comprehensive unit tests (9 tests) #246 - Stellar Transaction Simulation Service - Create StellarSimulationService for transaction simulation - Integrate with Soroban RPC for XDR validation and simulation - Include risk assessment and fee estimation - Add circuit breaker protection for RPC failures - Add XDR validation helper method - Add comprehensive unit tests with mocked Stellar RPC (24 tests) #245 - Cryptographic API Key Hashing Upgrade - Upgrade from SHA-256 to Argon2id for enhanced security - Implement memory-hard algorithm resistant to GPU/ASIC attacks - Add timing-attack resistant comparison via Argon2 verification - Maintain SHA-256 fallback for backward compatibility - Update ApiKeyService with dual-algorithm verification - Add comprehensive unit tests (32 crypto tests, 15 API key tests) #244 - Webhook Retry and Dead Letter Queue - Verify existing implementation meets all requirements - Confirm exponential backoff with jitter (2000ms base, 20% jitter) - Confirm 5 max attempts and non-transient error detection - Verify dead-letter handler via DeadLetterService - All existing tests passing Closes #247, #246, #245, #244 Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(workers): centralize job error handling with scrubbed structured logging Add runWorkerJob, a wrapper every background worker now routes its handler through. It times the job via WorkerMetricsService when available, logs a structured completion record, and classifies failures before rethrowing: transient failures log a job.retrying warning, while a failure on the final attempt or an UnrecoverableError logs a job.dead-lettered error with the scrubbed payload and stack. The original error is always rethrown untouched so BullMQ retry semantics are preserved, and logging can never mask it. Add scrubForLog/scrubString, which redact sensitive keys (sharing the audit sanitizer's key list) and secret-shaped substrings such as Stellar seeds, bearer tokens and URL credentials, and coerce cycles, bigints and errors into JSON-safe values. QueueFailureListener now scrubs its log line as well; previously the raw webhook payload, including its signing secret, was written to the error log. The dead-letter copy keeps the raw payload so re-drive still works. * feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping * Fix CI: document missing env vars, fix flaky retry backoff test - docs/configuration.md was missing entries for DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, and the PUBLIC_RATE_LIMIT_* vars, failing the configuration-documentation test. - retry.util.spec.ts left three rejected promises unhandled between `runAllTimersAsync()` and the `expect(...).rejects` assertion that attaches the handler; attach a no-op .catch() immediately after creating each promise so fake-timer-driven rejections don't fire as unhandled rejections mid-test-run. * perf(auth): cache session revocation answers during token verification Every authenticated request performed a Redis round trip against the token blacklist to answer "is this session still revoked?". This adds a short-TTL caching layer in front of the blacklist so repeated verifications within one window skip the Redis query entirely. - Add CacheService: a small get/set/delete cache over the shared REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based prefix invalidation. Every operation degrades to a no-op/miss on Redis failure so caching can never break the request path. - Add TokenVerificationCacheService: caches per-session revocation answers for TOKEN_CACHE_TTL seconds (default 30, well below the 15-minute access-token lifetime) and exposes invalidation hooks. - Wire the cache into JwtStrategy.validate (cache-first, source of truth on miss, fail-open unchanged on Redis outages). - Hook invalidation into every revocation path: TokenBlacklistService drops the cached answer after each blacklist write (including on Redis-outage fallback), AuthService invalidates on logout and on refresh rotation, so revocations are observed immediately instead of after the TTL window. Revocation reliability is preserved because no cached answer outlives its TTL, and explicit logout/rotation clears the entry at once. Tests: unit suites for CacheService and TokenVerificationCacheService (hits, misses, resolver fallback, invalidation hooks), an integration suite proving repeated authentications trigger a single blacklist lookup and that logout flips a cached-valid session to 401 immediately, plus updated JwtStrategy/TokenBlacklistService/api-key integration suites for the new wiring. Closes #341 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * feat: correlate requests, paginate audit, sign webhooks * feat(rate-limit): configurable limits and client identifiers for public routes Public endpoints are the first thing abusive traffic hits, but their limits were a single fixed IP budget: every public route shared PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one shared address (NAT, office egress, CI runners) exhausted one bucket together. This makes the public rate limiter configurable per route and per client identifier, on the existing Redis sliding-window counter. - Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or per-controller) budget overrides resolved by the guard through Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default. - Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated, currently 'apiKey'): when enabled, a presented x-api-key or ApiKey/Bearer ak_... Authorization header is folded into the bucket key so distinct key-holding clients behind one IP get their own budgets. The IP always participates; keyless callers share the plain-IP bucket as before. - Extract a testable PublicRateLimitGuard.check() returning the full decision (allowed, limit, windowSeconds, count, resetAt); canActivate keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset + Retry-After contract and in-memory fallback on Redis outage. Tests: unit suites for per-route rule resolution, identifier bucketing (with/without identifiers configured), header correctness on allowed and limited requests, and @SkipPublicRateLimit() interaction with rules; an HTTP-level integration suite simulating bursts that proves the 429-with-headers behaviour at the global limit, per-route overrides (next to unaffected sibling routes), and per-key budget isolation. Existing public-rate-limit suites pass unchanged (bucket keys keep the ip: prefix, so stored counters stay compatible). Closes #342 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * feat: stream audit exports and harden payment safety (#63, #64, #283) * fix: validate nested transaction metadata as JSON (#283) * feat(transactions): implement agent spending limit evaluation guard * fix(ci): resolve typecheck and lint failures across specs, guards, and docs Get CI green for the token-verification cache PR by fixing pre-existing main-branch type errors alongside PR-specific ones: Express specs no longer use Fastify-only app.inject, Stellar mocks match the real Soroban result interface, the transaction spec exercises the actual create pipeline, TokenBlacklistService resolves the global REDIS_CLIENT token explicitly, and the configuration docs cover every THROTTLE_* env var the docs test asserts. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * fix(ci): resolve typecheck and lint failures across specs, guards, and docs Get CI green for the token-verification cache PR by fixing pre-existing main-branch type errors alongside PR-specific ones: Express specs no longer use Fastify-only app.inject, Stellar mocks match the real Soroban result interface, the transaction spec exercises the actual create pipeline, TokenBlacklistService resolves the global REDIS_CLIENT token explicitly, and the configuration docs cover every THROTTLE_* env var the docs test asserts. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * fix(config): deduplicate parseClientIdentifiers and pass the raw variable The duplicated helper also ignored its parameter while the config factory read the raw variable at the call site with none, so the module never compiled. Collapse to a single parser that takes the raw value. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * feat(audit): implement structured audit logging interceptor for sensitive operations Closes #255 Adds the @AuditLog() decorator (with optional action/entity metadata) and reworks AuditLogInterceptor so only decorated routes are persisted, keeping read-only traffic free of audit writes. Each record captures the actor (human user id or acting agent), client IP, HTTP method, request path, a SHA-256 fingerprint of the sanitized payload, the sanitized body itself, the final response status and handler duration. Secrets, keys, tokens, signatures and mnemonics are redacted recursively without mutating the original request body, and @SkipAudit() always wins. Applied to the sensitive policy, budget and API-key endpoints; persistence stays fire-and-forget so a failed audit write never breaks the client request. * refactor(policies): add Prisma repository abstraction for agent spending policies Closes #254 Introduces SpendingPolicyRepository (src/modules/policies/spending-policy.repository.ts) as the single owner of every Prisma call for policies: create/find/update/soft-delete, the paginated read, the rolling spend aggregation used by velocity checks, the POLICY_EVALUATED audit row, and an interactive withTransaction() helper. Every method funnels through one error-handling wrapper that logs the operation (plus the Prisma error code) and rethrows the original error, so error handling and transaction safety are uniform across the repository. Adds SpendingPolicyService, which injects the repository and now owns spending-policy validation and enforcement (staleness checks, daily velocity limit, evaluation audit) with no direct Prisma access. PolicyService is reduced to orchestration - domain events plus the pure PolicyEngine - and delegates all persistence to SpendingPolicyService, which removes the direct prisma.transaction/prisma.auditLog calls from the service layer. PolicyRepository is replaced by the new, better-named repository. Unit tests cover the repository against a mocked Prisma client (including transaction usage, Decimal summing and error propagation) and the service against a mocked repository (validation, pagination, not-found, velocity limits and audit failure swallowing). * feat(webhooks): add BullMQ retry policy and dead-letter audit for webhook deliveries Closes #253 Centralises the webhook queue configuration in src/queues/webhook.queue.ts: 5 attempts, exponential backoff with a 2000ms base, the jittered custom backoff strategy, 24h retention of failed jobs and the dead-letter destination. Both the queue registration and the enqueue path consume the same constants, so the API and the worker can no longer drift apart. Adds WebhookAuditService, which appends a WEBHOOK_DELIVERY_FAILED record (subscriber, event, attempts, HTTP status, failure reason) to the audit trail for deliveries that will never be retried - either an unrecoverable 4xx or the final attempt after retries are exhausted. Both the webhook processor and the worker call it through a fire-and-forget hook that swallows and logs its own failures, so audit problems can never mask the delivery error, crash the NestJS process or interfere with BullMQ retry/backoff. Tests cover the retry/backoff/DLQ configuration and jitter envelope, the audit service (including persistence failures) and the processor's terminal-failure behaviour - retry without auditing while attempts remain, exactly one audit entry when retries are exhausted or a 4xx is returned, and error preservation when auditing fails. * feat(rate-limit): add Redis-backed agent throttler guard for high-frequency agent endpoints Closes #251 Adds AgentThrottlerGuard in src/common/guards/agent-throttler.guard.ts, extending @nestjs/throttler's ThrottlerGuard so counters keep running through the shared RedisThrottlerStorage while the guard adds agent-aware behaviour: the bucket key is derived from the acting agent (x-agent-id, a route/body/query agentId, or an agent-bound API-key principal) instead of the organization or IP, falling back to the organization, then a hashed API key, then the client IP (honouring x-forwarded-for) for unauthenticated public routes. Introduces a third 'agent' tier (THROTTLE_AGENT_LIMIT, default 300/window) alongside api/auth in config/throttler.config.ts and env.validation.ts, and routes each request to exactly one tier - explicit @ThrottleTierDecorator() metadata wins, otherwise agent-identified traffic uses the agent tier and everything else the api tier. Rejections are standard 429s that also carry plain Retry-After, X-RateLimit-Limit and X-RateLimit-Remaining headers, and the tier limit is advertised on allowed responses too. Applied selectively to the agent-facing controllers (agents, transactions, wallets) via @UseGuards so the existing global AstroidThrottlerGuard keeps enforcing api/auth untouched. Vitest coverage asserts tier routing, tracker resolution for agents/orgs/API keys/anonymous callers, header emission and 429 enforcement when a single agent's burst exhausts its budget. * Fix CI: pre-existing typecheck/lint/test breakage inherited from main main's HEAD was red (27 typecheck errors) from unrelated merged PRs (#357, #374, #381-#388). Fixed to get this branch's CI green: - throttler.guard.ts / sliding-window-throttler.guard.ts: AuthenticatedUser has no `sub` or `tier` field; use `.id` and treat `tier` as an optional extension until the type actually carries it. - agent.controller.ts: imported SlidingWindowThrottlerGuard/SlidingWindowLimit (unused) instead of the AstroidThrottlerGuard actually referenced by @UseGuards. - event-names.ts: dropped a duplicate TransactionRiskScoringRequested key. - risk.service.ts: the TransactionCreated handler built a RiskFactorsInput with fields (`destination`, `velocityCount`, `isNewRecipient`) that don't exist on the type; mapped to the real shape instead. Removed dead lowRisk/createEventBus fixtures left over in risk.service.spec.ts. - stellar.service.ts (Soroban variant): fixed getTransactionInfo calling a nonexistent client method (getTransaction), simulateTransaction being called with a bare string instead of {transactionXdr}, and an error message interpolating the whole error object instead of `.message`. Rewrote stellar.service.spec.ts's mocks/assertions to match the real SorobanSimulationResult shape, and fixed two tests reusing a `mockOnce` across two separate calls to the service (second call fell through to the unmocked default and threw on undefined). - transaction.service.spec.ts targeted a pre-broadcast Soroban simulation step that was never wired into TransactionService.create, against a StellarService overload TransactionService doesn't even inject; skipped with an explanation rather than fabricating the feature. - sensitive-rate-limit.integration.spec.ts: replaced Fastify-only `app.inject()` (this app runs on platform-express) with a small http.request helper; the throttler in the test module was unnamed ('default'), so AstroidThrottlerGuard's per-tier name match against the route's default 'api' tier always skipped it — named it 'api' to match. - Removed remaining `any` usages in the touched guard files/specs. - docs/configuration.md: documented THROTTLE_WEBHOOK_LIMIT, THROTTLE_API_BURST, THROTTLE_AUTH_BURST, THROTTLE_WEBHOOK_BURST (missing from #386), failing the configuration-documentation test. * Fix CI: repair botched merge of main into this branch A merge of upstream main (PR #372's overlapping CI fixes) into this branch left several files with duplicated/glued content instead of resolved conflicts: - sliding-window-throttler.guard.spec.ts: duplicate makeContext() declaration (one-line + multi-line versions concatenated). - sliding-window-throttler.guard.ts: `limit` reverted to `const` while the tier-adjustment branches below still reassign it. - sensitive-rate-limit.integration.spec.ts: my raw http.request-based version and another (cleaner, fetch()-based) version from main were concatenated rather than merged — duplicate imports, an unclosed `it` block. Kept the fetch()-based version. - stellar.service.spec.ts: duplicate `error` key in a SorobanSimulationResult object literal. - transaction.service.spec.ts: my placeholder describe.skip stub was prepended to a real, correctly implemented version of the same suite that showed up on main independently. Kept the real implementation, dropped the stub. * Remove conflict markers from worker merge * feat: implement risk scoring persistence and complete API documentation Implement comprehensive risk scoring service with data persistence for compliance tracking, complete OpenAPI documentation coverage, and add integration tests for core infrastructure components. Risk Scoring Service (#213): - Add RiskRepository for assessment record persistence and historical analysis - Update RiskService to persist assessment records via repository - Add getHistory() and getStatistics() methods for risk analytics - Add RiskAssessment model to Prisma schema with proper indexes - Create database migration for risk assessments table - Add comprehensive integration tests for RiskRepository Response Envelope Interceptor (#216): - Verify ResponseInterceptor is registered globally in app.module.ts - Add integration tests for response wrapping and pagination handling - Test requestId header handling and null data scenarios Swagger Documentation (#218): - Add @ApiProperty decorators to admin DLQ and queue management DTOs - Add @ApiProperty decorators to audit export DTOs - Add @ApiProperty decorators to risk assessment DTOs - Update risk controller to use DTO for request body documentation - Complete OpenAPI documentation coverage across all domain modules Zod Validation Pipe (#220): - Verify ZodValidationPipe supports custom error formatting - Add integration tests for validation scenarios and error handling - Test optional fields, nested objects, arrays, and custom messages * fix: add missing relation field in RiskAssessment model Add organization relation field to RiskAssessment model to fix Prisma schema validation error. Update migration to add foreign key constraint separately for proper schema validation. * fix: resolve TypeScript errors in test files Fix TypeScript errors in test files by: - Converting async tests to use toPromise() instead of callbacks - Adding proper type assertions for exception details - Adding RiskRepository to RiskService constructor - Using type assertions for Prisma client methods until migration runs - Adding missing PaginationMeta properties in tests * fix: add null checks for toPromise() results in response interceptor tests Add null checks after toPromise() calls to resolve TypeScript 'possibly undefined' errors in response interceptor tests. * fix: add eslint-disable comments for temporary any types Add eslint-disable comments for @typescript-eslint/no-explicit-any where we use type assertions for Prisma client methods until the migration runs and generates the proper types. * fix: update test expectation for date comparison in statistics test Change the test expectation to use expect.objectContaining for the nested createdAt object instead of expect.any(Date) for the entire field, as the Date object is being compared with its actual value. * fix: remove duplicate Zod validation pipe integration test * fix ci * fix ci * fix failing ci * fix ci * fix ci * fix ci * feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356) * feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files Resolves the merge conflict from PR #284 by applying all genuinely new additions (RiskRepository, schema migration, test specs, service updates) while keeping main's already-improved Swagger DTO implementations. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * fix: use ZodValidationException in integration spec The pipe throws ZodValidationException (a BadRequestException subclass defined in zod-validation.pipe.ts), not the domain-layer ValidationException. Update all assertions to match the actual thrown type. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: address requested issues - Done with all issues (#357) * feat(health): dedicated GET /health/redis probe (#358) Adds GET /health/redis, mirroring GET /health/database, so container orchestration and uptime monitors can probe Redis alone instead of only seeing it as one service inside /health/readiness. Reports status, latency and the ping error, with 200/503 semantics, plus unit tests for the up, down and isolation paths. * test(auth): lock in auth throttle-tier wiring on public routes (#359) Guard and config behavior for rate limiting is already covered in isolation, but nothing asserted that register/login/refresh actually declare the auth tier. Adds a regression test against the decorator metadata so removing @ThrottleTierDecorator('auth') from a handler fails CI instead of silently dropping back to the looser api limit. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(health): add /health/live and /health/ready probes (#362) Add orchestrator-grade liveness and readiness probes: - GET /health/live returns 200 whenever the process is running and performs no dependency checks, so a downstream outage never triggers a restart. - GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a per-dependency status, latency and error report under `services`. - Both probes are served outside the global API prefix, like /metrics, so probe paths are stable across API versions. Make the health controller usable by load balancers: - Mark it @Public(); previously every health route required a JWT. - Exempt it from both named throttler tiers so probes cannot receive 429s. - Exclude it from the audit trail so probes do not write an audit row per request (or attempt to while the database is down). Fix the Redis indicator to probe the shared REDIS_CLIENT built from the validated REDIS_* config. It previously read an undefined REDIS_URL, always probed localhost:6379, kept its own never-closed connection, and could hang while ioredis queued the PING during an outage. Closes #351 * feat(config): validate full environment at startup (#363) Compose the per-slice Zod schemas into a single `environmentSchema` and validate `process.env` against it at the start of bootstrap(), before any Nest module is constructed or any connection is opened. - On failure the process prints every failing variable in one message and exits with code 1, instead of surfacing only the first failing slice from inside Nest's module initialization with a stack trace. - Messages are value-free (e.g. "must be one of: ..." rather than Zod's default "received '<value>'") so secrets never reach logs. - In production, reject the publicly known default ENCRYPTION_KEY (whether set explicitly or implied by omission) and a JWT refresh secret that reuses the access secret. Per-slice validation in each registerAs factory is unchanged, so the typed ConfigService namespaces keep their guarantees outside main.ts. Document every variable, its type, default and production rules in docs/configuration.md, and correct the README's list of required variables. Tests keep the docs and .env.example in sync with the schema, and exercise the real main.ts entrypoint to prove missing or malformed variables halt startup before NestFactory.create is called. Closes #350 * fix(tests): add SpendingLimitService mock to TransactionService spec * feat: webhook ingress validation, retry queue cleanup, rate limiting (#367) Closes #331 Closes #325 Closes #334 Closes #328 - Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard, with a Zod schema validating the payload shape. - Removed duplicate BullMQ processor consuming the webhooks queue (WebhookWorker), keeping WebhooksProcessor which also records metrics. Deleted dead WebhookDeliveryWorker, never registered as a real processor. - Wired the existing, previously unused SlidingWindowThrottlerGuard onto the new public ingress route for Redis-backed sliding-window rate limiting with standard rate-limit headers. - Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation per stream, exposed through the existing Prometheus registry. * feat: add production rollback protection to database migration CLI (#371) Add a migration CLI (npm run db:migrate -- <command>) with a production safety guard. The destructive commands, down and reset, are rejected when NODE_ENV=production unless --force is supplied. A blocked run prints a warning to stderr, exits with code 1 and never contacts the database. A forced run is also announced on stderr. down <migration> runs the migration's hand-written down.sql and removes its _prisma_migrations row in a single script, so the history row is only dropped when every rollback statement succeeded. Migration names are validated against the folder on disk before being used in SQL. The guard logic lives in src/database/migration-guard.ts with injected process dependencies so it is fully unit tested. The entry point src/database/migrate.cli.ts is compiled into dist and can run in images without ts-node. Usage is documented in docs/database.md. * test: add full branch coverage tests for role, permission and scope guards (#373) Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard, including matchScope, and extend the RbacGuard spec. The guards now have 100% statement, branch, function and line coverage. The tests attach real @Roles, @RequirePermissions, @RequireScopes and @Public metadata to fixture controllers and run them through a real Reflector, so handler-over-class inheritance is exercised rather than mocked. Assertion matrices cover every UserRole, every wildcard shape in matchScope, including nested scopes, and the AND semantics of multi- permission routes. Behaviour pinned by these tests: - OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN bypasses ScopesGuard; API-key principals with those roles do not. - PermissionsGuard grants no role-based override and expands no wildcards. - Guests and expired or revoked principals, which the auth strategies leave without request.user, get a 401 on every restricted route. - Stale or differently cased role names are rejected. * perf: add concurrent composite indexes for notification, approval and memory (#374) Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but the tables behind the busiest ones only had single-column indexes. Postgres therefore had to fetch all of a tenant's or user's rows, or walk the global createdAt index and filter, before it could return a page. These composite indexes match the actual query shapes: - notifications (userId, createdAt): inbox list - notifications (organizationId, userId, read, createdAt): unread badge, mark-all-read, unread filter (index-only count) - proposals (organizationId, createdAt): approval queue - proposals (organizationId, status, createdAt): pending count, status filter - memory_records (organizationId, createdAt): memory browser - memory_records (agentId, createdAt): per-agent memory timeline Each index is built with CREATE INDEX CONCURRENTLY, so writes are never blocked. Each one lives in its own single-statement migration, because Prisma runs a multi-statement migration as one implicit transaction, where Postgres rejects CONCURRENTLY. The matching @@index entries are added to schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md records the convention and how to recover from a failed concurrent build. * feat: add IP-based rate limiting to public endpoints (#377) Add PublicRateLimitGuard, a global guard that applies a per-IP sliding window limit to every unauthenticated route: handlers marked @Public() and any route under /<API_PREFIX>/public/. Requests beyond the limit are rejected with 429 Too Many Requests and a Retry-After header. - X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set on every limited response - counters live in the shared REDIS_CLIENT via an atomic Lua sliding window script, so all replicas enforce one budget per IP; rejected requests are not recorded - if Redis is unavailable the guard falls back to an in-memory window per instance instead of failing open - thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED, PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default) - @SkipPublicRateLimit() exempts routes; applied to the network-restricted /metrics scrape endpoint - add store and guard unit tests plus an HTTP burst integration test * feat: return RFC 9457 problem details for all error responses (#378) Replace the { success, error: { code, message }, requestId } error envelope with a uniform problem details body served as application/problem+json: { type, title, status, detail, instance, code, requestId, details? } - type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP errors without a dedicated code use about:blank with the reason phrase - title comes from a new ERROR_TITLE map kept exhaustive by the type system; instance is the request path without its query string - code and requestId are kept as extension members so clients can keep switching on the machine-readable code - ZodValidationException now keeps its VALIDATION_ERROR code and field-level details instead of collapsing to BAD_REQUEST - unhandled exceptions still map to a generic 500 INTERNAL_ERROR without leaking internals - update the shared response types and API documentation - rewrite the filter spec and add an HTTP integration test covering validation, authentication, domain, not-found and server errors * fix: worker error handling + log scrubbing (#370) * fix: worker error handling + log scrubbing (resolve PR #370 conflicts) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * test: add job-worker spec and queue-failure-listener scrubbing test Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369) * perf(analytics): batch dashboard overview queries into one transaction Closes #338 AnalyticsService.overview() issued 7 independent queries (counts, spend aggregates, status/risk group-bys) via Promise.all, each its own roundtrip. Batches the 3 counts and 2 aggregates into a single $transaction([...]) call; the 2 groupBy calls stay outside the batch since Prisma's groupBy return type doesn't infer correctly inside a $transaction array. Response shape is unchanged. Existing indexes on organizationId/status/createdAt already cover these queries. * test(audit): cover pagination and sorting edge cases for activity log Closes #339 The audit log (this repo's activity log) already supported page/limit/sort/order/filter query params and returned pagination metadata (total, totalPages, hasNext, hasPrev), with limit capped at 100. Adds unit test coverage for the previously-untested list() method: normal pagination, empty results, out-of-bounds pages, invalid sort field fallback, ascending order, and entity filtering. * fix(migrations): resolve colliding timestamp between two merged migrations Migrations 20260928120000_add_agent_contribution_stats_index and 20260928120000_add_notifications_user_created_at_index landed with the same 14-digit timestamp prefix from two separately merged PRs (#374, #357), which scripts/verify-migrations.sh rejects as a conflict. Bumps the notifications index migration to 20260928120001; both migrations are independent, additive CREATE INDEX statements with no ordering dependency between them, so the rename is safe. * fix(ci): document missing env vars and fix flaky retry.util test Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking this branch: - docs/configuration.md was missing 7 env vars added by recent merges (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which env.validation.spec.ts asserts against. Documented all 7. - retry.util.spec.ts had 3 tests that create a rejecting promise, advance fake timers with vi.runAllTimersAsync(), then attach the rejection assertion afterward — a race that surfaces as an unhandled rejection under full-suite load (deterministic once >100 files run together). Attaching a no-op .catch() immediately after creating the promise prevents the unhandled state without changing what each test asserts. * feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388) * Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381) * Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384) * Fix #223: Implement Stellar Transaction Simulation Service Integration (#385) * Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382) * Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383) * feat(interceptors): add global RequestIdInterceptor for correlation logging (#387) * feat(throttler): add Redis-backed rate limiting guard and throttler config (#386) * fix: resolve rebase conflicts and restore CI * feat: implement risk scoring persistence and complete API documentation Implement comprehensive risk scoring service with data persistence for compliance tracking, complete OpenAPI documentation coverage, and add integration tests for core infrastructure components. Risk Scoring Service (#213): - Add RiskRepository for assessment record persistence and historical analysis - Update RiskService to persist assessment records via repository - Add getHistory() and getStatistics() methods for risk analytics - Add RiskAssessment model to Prisma schema with proper indexes - Create database migration for risk assessments table - Add comprehensive integration tests for RiskRepository Response Envelope Interceptor (#216): - Verify ResponseInterceptor is registered globally in app.module.ts - Add integration tests for response wrapping and pagination handling - Test requestId header handling and null data scenarios Swagger Documentation (#218): - Add @ApiProperty decorators to admin DLQ and queue management DTOs - Add @ApiProperty decorators to audit export DTOs - Add @ApiProperty decorators to risk assessment DTOs - Update risk controller to use DTO for request body documentation - Complete OpenAPI documentation coverage across all domain modules Zod Validation Pipe (#220): - Verify ZodValidationPipe supports custom error formatting - Add integration tests for validation scenarios and error handling - Test optional fields, nested objects, arrays, and custom messages * fix: add missing relation field in RiskAssessment model Add organization relation field to RiskAssessment model to fix Prisma schema validation error. Update migration to add foreign key constraint separately for proper schema validation. * fix: resolve TypeScript errors in test files Fix TypeScript errors in test files by: - Converting async tests to use toPromise() instead of callbacks - Adding proper type assertions for exception details - Adding RiskRepository to RiskService constructor - Using type assertions for Prisma client methods until migration runs - Adding missing PaginationMeta properties in tests * fix: add null checks for toPromise() results in response interceptor tests Add null checks after toPromise() calls to resolve TypeScript 'possibly undefined' errors in response interceptor tests. * fix: add eslint-disable comments for temporary any types Add eslint-disable comments for @typescript-eslint/no-explicit-any where we use type assertions for Prisma client methods until the migration runs and generates the proper types. * fix: update test expectation for date comparison in statistics test Change the test expectation to use expect.objectContaining for the nested createdAt object instead of expect.any(Date) for the entire field, as the Date object is being compared with its actual value. * fix: remove duplicate Zod validation pipe integration test * feat: add RiskRepository, RiskAssessment schema, and Swagger decorators (#356) * feat: merge PR #284 — add RiskRepository, RiskAssessment schema, and test files Resolves the merge conflict from PR #284 by applying all genuinely new additions (RiskRepository, schema migration, test specs, service updates) while keeping main's already-improved Swagger DTO implementations. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * fix: use ZodValidationException in integration spec The pipe throws ZodValidationException (a BadRequestException subclass defined in zod-validation.pipe.ts), not the domain-layer ValidationException. Update all assertions to match the actual thrown type. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: address requested issues - Done with all issues (#357) * feat(health): dedicated GET /health/redis probe (#358) Adds GET /health/redis, mirroring GET /health/database, so container orchestration and uptime monitors can probe Redis alone instead of only seeing it as one service inside /health/readiness. Reports status, latency and the ping error, with 200/503 semantics, plus unit tests for the up, down and isolation paths. * test(auth): lock in auth throttle-tier wiring on public routes (#359) Guard and config behavior for rate limiting is already covered in isolation, but nothing asserted that register/login/refresh actually declare the auth tier. Adds a regression test against the decorator metadata so removing @ThrottleTierDecorator('auth') from a handler fails CI instead of silently dropping back to the looser api limit. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(health): add /health/live and /health/ready probes (#362) Add orchestrator-grade liveness and readiness probes: - GET /health/live returns 200 whenever the process is running and performs no dependency checks, so a downstream outage never triggers a restart. - GET /health/ready probes the database (SELECT 1) and cache (Redis PING) in parallel, each bounded by a 2s timeout, and returns 200 or 503 with a per-dependency status, latency and error report under `services`. - Both probes are served outside the global API prefix, like /metrics, so probe paths are stable across API versions. Make the health controller usable by load balancers: - Mark it @Public(); previously every health route required a JWT. - Exempt it from both named throttler tiers so probes cannot receive 429s. - Exclude it from the audit trail so probes do not write an audit row per request (or attempt to while the database is down). Fix the Redis indicator to probe the shared REDIS_CLIENT built from the validated REDIS_* config. It previously read an undefined REDIS_URL, always probed localhost:6379, kept its own never-closed connection, and could hang while ioredis queued the PING during an outage. Closes #351 * feat(config): validate full environment at startup (#363) Compose the per-slice Zod schemas into a single `environmentSchema` and validate `process.env` against it at the start of bootstrap(), before any Nest module is constructed or any connection is opened. - On failure the process prints every failing variable in one message and exits with code 1, instead of surfacing only the first failing slice from inside Nest's module initialization with a stack trace. - Messages are value-free (e.g. "must be one of: ..." rather than Zod's default "received '<value>'") so secrets never reach logs. - In production, reject the publicly known default ENCRYPTION_KEY (whether set explicitly or implied by omission) and a JWT refresh secret that reuses the access secret. Per-slice validation in each registerAs factory is unchanged, so the typed ConfigService namespaces keep their guarantees outside main.ts. Document every variable, its type, default and production rules in docs/configuration.md, and correct the README's list of required variables. Tests keep the docs and .env.example in sync with the schema, and exercise the real main.ts entrypoint to prove missing or malformed variables halt startup before NestFactory.create is called. Closes #350 * feat: webhook ingress validation, retry queue cleanup, rate limiting (#367) Closes #331 Closes #325 Closes #334 Closes #328 - Inbound webhook receiving endpoint (POST /webhooks/receive) wired to the existing but previously unwired RawBodyMiddleware + WebhookSignatureGuard, with a Zod schema validating the payload shape. - Removed duplicate BullMQ processor consuming the webhooks queue (WebhookWorker), keeping WebhooksProcessor which also records metrics. Deleted dead WebhookDeliveryWorker, never registered as a real processor. - Wired the existing, previously unused SlidingWindowThrottlerGuard onto the new public ingress route for Redis-backed sliding-window rate limiting with standard rate-limit headers. - Added StreamMetricsService: ring-buffer based p95/p99 latency aggregation per stream, exposed through the existing Prometheus registry. * feat: add production rollback protection to database migration CLI (#371) Add a migration CLI (npm run db:migrate -- <command>) with a production safety guard. The destructive commands, down and reset, are rejected when NODE_ENV=production unless --force is supplied. A blocked run prints a warning to stderr, exits with code 1 and never contacts the database. A forced run is also announced on stderr. down <migration> runs the migration's hand-written down.sql and removes its _prisma_migrations row in a single script, so the history row is only dropped when every rollback statement succeeded. Migration names are validated against the folder on disk before being used in SQL. The guard logic lives in src/database/migration-guard.ts with injected process dependencies so it is fully unit tested. The entry point src/database/migrate.cli.ts is compiled into dist and can run in images without ts-node. Usage is documented in docs/database.md. * test: add full branch coverage tests for role, permission and scope guards (#373) Add dedicated specs for RolesGuard, PermissionsGuard and ScopesGuard, including matchScope, and extend the RbacGuard spec. The guards now have 100% statement, branch, function and line coverage. The tests attach real @Roles, @RequirePermissions, @RequireScopes and @Public metadata to fixture controllers and run them through a real Reflector, so handler-over-class inheritance is exercised rather than mocked. Assertion matrices cover every UserRole, every wildcard shape in matchScope, including nested scopes, and the AND semantics of multi- permission routes. Behaviour pinned by these tests: - OWNER bypasses RolesGuard, and a JWT-authenticated OWNER or ADMIN bypasses ScopesGuard; API-key principals with those roles do not. - PermissionsGuard grants no role-based override and expands no wildcards. - Guests and expired or revoked principals, which the auth strategies leave without request.user, get a 401 on every restricted route. - Stale or differently cased role names are rejected. * perf: add concurrent composite indexes for notification, approval and memory (#374) Every paginated list endpoint defaults to ORDER BY "createdAt" DESC, but the tables behind the busiest ones only had single-column indexes. Postgres therefore had to fetch all of a tenant's or user's rows, or walk the global createdAt index and filter, before it could return a page. These composite indexes match the actual query shapes: - notifications (userId, createdAt): inbox list - notifications (organizationId, userId, read, createdAt): unread badge, mark-all-read, unread filter (index-only count) - proposals (organizationId, createdAt): approval queue - proposals (organizationId, status, createdAt): pending count, status filter - memory_records (organizationId, createdAt): memory browser - memory_records (agentId, createdAt): per-agent memory timeline Each index is built with CREATE INDEX CONCURRENTLY, so writes are never blocked. Each one lives in its own single-statement migration, because Prisma runs a multi-statement migration as one implicit transaction, where Postgres rejects CONCURRENTLY. The matching @@index entries are added to schema.prisma so migrate dev reports no drift. docs/concurrent-indexes.md records the convention and how to recover from a failed concurrent build. * feat: add IP-based rate limiting to public endpoints (#377) Add PublicRateLimitGuard, a global guard that applies a per-IP sliding window limit to every unauthenticated route: handlers marked @Public() and any route under /<API_PREFIX>/public/. Requests beyond the limit are rejected with 429 Too Many Requests and a Retry-After header. - X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset are set on every limited response - counters live in the shared REDIS_CLIENT via an atomic Lua sliding window script, so all replicas enforce one budget per IP; rejected requests are not recorded - if Redis is unavailable the guard falls back to an in-memory window per instance instead of failing open - thresholds are configurable with PUBLIC_RATE_LIMIT_ENABLED, PUBLIC_RATE_LIMIT_MAX_REQUESTS, PUBLIC_RATE_LIMIT_WINDOW_SECONDS and PUBLIC_RATE_LIMIT_TRUST_PROXY (X-Forwarded-For is ignored by default) - @SkipPublicRateLimit() exempts routes; applied to the network-restricted /metrics scrape endpoint - add store and guard unit tests plus an HTTP burst integration test * feat: return RFC 9457 problem details for all error responses (#378) Replace the { success, error: { code, message }, requestId } error envelope with a uniform problem details body served as application/problem+json: { type, title, status, detail, instance, code, requestId, details? } - type is a stable URN per ErrorCode (urn:astroid:problem:<code>); HTTP errors without a dedicated code use about:blank with the reason phrase - title comes from a new ERROR_TITLE map kept exhaustive by the type system; instance is the request path without its query string - code and requestId are kept as extension members so clients can keep switching on the machine-readable code - ZodValidationException now keeps its VALIDATION_ERROR code and field-level details instead of collapsing to BAD_REQUEST - unhandled exceptions still map to a generic 500 INTERNAL_ERROR without leaking internals - update the shared response types and API documentation - rewrite the filter spec and add an HTTP integration test covering validation, authentication, domain, not-found and server errors * fix: worker error handling + log scrubbing (#370) * fix: worker error handling + log scrubbing (resolve PR #370 conflicts) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * test: add job-worker spec and queue-failure-listener scrubbing test Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: query metrics, retry utility, and input sanitization (resolve PR #368 conflicts) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012xRrTYwEy2y8yow4o8iUQ8 * Batch dashboard queries and add activity log pagination test coverage (#338, #339) (#369) * perf(analytics): batch dashboard overview queries into one transaction Closes #338 AnalyticsService.overview() issued 7 independent queries (counts, spend aggregates, status/risk group-bys) via Promise.all, each its own roundtrip. Batches the 3 counts and 2 aggregates into a single $transaction([...]) call; the 2 groupBy calls stay outside the batch since Prisma's groupBy return type doesn't infer correctly inside a $transaction array. Response shape is unchanged. Existing indexes on organizationId/status/createdAt already cover these queries. * test(audit): cover pagination and sorting edge cases for activity log Closes #339 The audit log (this repo's activity log) already supported page/limit/sort/order/filter query params and returned pagination metadata (total, totalPages, hasNext, hasPrev), with limit capped at 100. Adds unit test coverage for the previously-untested list() method: normal pagination, empty results, out-of-bounds pages, invalid sort field fallback, ascending order, and entity filtering. * fix(migrations): resolve colliding timestamp between two merged migrations Migrations 20260928120000_add_agent_contribution_stats_index and 20260928120000_add_notifications_user_created_at_index landed with the same 14-digit timestamp prefix from two separately merged PRs (#374, #357), which scripts/verify-migrations.sh rejects as a conflict. Bumps the notifications index migration to 20260928120001; both migrations are independent, additive CREATE INDEX statements with no ordering dependency between them, so the rename is safe. * fix(ci): document missing env vars and fix flaky retry.util test Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking this branch: - docs/configuration.md was missing 7 env vars added by recent merges (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which env.validation.spec.ts asserts against. Documented all 7. - retry.util.spec.ts had 3 tests that create a rejecting promise, advance fake timers with vi.runAllTimersAsync(), then attach the rejection assertion afterward — a race that surfaces as an unhandled rejection under full-suite load (deterministic once >100 files run together). Attaching a no-op .catch() immediately after creating the promise prevents the unhandled state without changing what each test asserts. * feat(filters): add GlobalExceptionFilter with Prisma and validation error mapping (#388) * Fix #234: Implement Event Emitter Domain Event Handlers for Transaction Risk Scoring (#381) * Fix #225: Implement Structured Audit Log Interceptor for Mutating Operations (#384) * Fix #223: Implement Stellar Transaction Simulation Service Integration (#385) * Fix #226: Add Redis-Backed Rate Limiting Guard with Dynamic Tier Support (#382) * Fix #231: Implement Redis-backed Rate Limiting Guard for Sensitive API Endpoints (#383) * feat(interceptors): add global RequestIdInterceptor for correlation logging (#387) * feat(throttler): add Redis-backed rate limiting guard and throttler config (#386) * Add startup migration check and DB pool connection metrics Wires the existing migration status checker into app bootstrap so the process halts before accepting traffic when prisma/migrations has pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt, the default; 'warn' logs and continues). Gated behind DATABASE_MIGRATION_CHECK_ENABLED. Adds a db_pool_connections Prometheus gauge (active/idle/waiting) sourced from pg_stat_activity, since Prisma's Rust query engine doesn't expose pool internals through the Node client. Closes #335 Closes #332 * Fix pre-existing typecheck/lint/test failures blocking CI main's build/typecheck/lint/test were already broken before this branch touched anything (confirmed by checking out upstream/main directly). CI enforces these repo-wide, so they block this PR too. Fixed each: - event-names.ts: duplicate object key (TransactionRiskScoringRequested) - throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't exist on that type (JWT payload field name leaked into the wrong type) - sliding-window-throttler.guard.ts: removed a user-tier rate-limit multiplier keyed on AuthenticatedUser.tier, a field never present anywhere in the auth/user model — dead, unbacked logic. Removed its now-orphaned tests too. - agent.controller.ts: AstroidThrottlerGuard used but never imported; dropped an unused SlidingWindowThrottlerGuard import block - risk.service.ts: event-driven risk scoring built a RiskFactorsInput with fields (destination/velocityCount/isNewRecipient) that don't exist on the current type; mapped to the real shape instead - risk.service.spec.ts: removed orphaned unused fixtures - stellar.service.ts (src/modules/stellar/services, unused elsewhere in the app but still typechecked/tested): getTransactionInfo called a client method that doesn't exist (real method is getTransaction); simulateTransaction passed a bare string where the client expects an options object - stellar.service.spec.ts: rewritten against the real SorobanSimulationResult shape; fixed mockResolvedValueOnce/ mockRejectedValueOnce being consumed by the test's own first assertion, leaving the second call unmocked - transaction.service.spec.ts: rewritten against TransactionService's actual create() contract (it doesn't call Soroban simulation at all; the previous spec tested a flow that was never implemented) and a real Ed25519 checksum address - sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist on this Express-platform app; switched to app.listen + fetch, matching the sibling public-rate-limit.integration.spec.ts pattern, and named the test throttler 'api' so AstroidThrottlerGuard's tier-matching actually engages it * Fix remaining pre-existing lint errors and undocumented env vars CI runs lint and test repo-wide, so these also blocked the PR: - 4 pre-existing no-explicit-any lint errors in throttler guard code and specs, typed properly instead of suppressed - 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST, WEBHOOK_BURST) were validated by the env schema but missing from docs/configuration.md, failing the docs-sync test * perf(auth): cache session revocation answers during token verification Every authenticated request performed a Redis round trip against the token blacklist to answer "is this session still revoked?". This adds a short-TTL caching layer in front of the blacklist so repeated verifications within one window skip the Redis query entirely. - Add CacheService: a small get/set/delete cache over the shared REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based prefix invalidation. Every operation degrades to a no-op/miss on Redis failure so caching can never break the request path. - Add TokenVerificationCacheService: caches per-session revocation answers for TOKEN_CACHE_TTL seconds (default 30, well below the 15-minute access-token lifetime) and exposes invalidation hooks. - Wire the cache into JwtStrategy.validate (cache-first, source of truth on miss, fail-open unchanged on Redis outages). - Hook invalidation into every revocation path: TokenBlacklistService drops the cached answer after each blacklist write (including on Redis-outage fallback), AuthService invalidates on logout and on refresh rotation, so revocations are observed immediately instead of after the TTL window. Revocation reliability is preserved because no cached answer outlives its TTL, and explicit logout/rotation clears the entry at once. Tests: unit suites for CacheService and TokenVerificationCacheService (hits, misses, resolver fallback, invalidation hooks), an integration suite proving repeated authentications trigger a single blacklist lookup and that logout flips a cached-valid session to 401 immediately, plus updated JwtStrategy/TokenBlacklistService/api-key integration suites for the new wiring. Closes #341 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * fix(ci): resolve typecheck and lint failures across specs, guards, and docs Get CI green for the token-verification cache PR by fixing pre-existing main-branch type errors alongside PR-specific ones: Express specs no longer use Fastify-only app.inject, Stellar mocks match the real Soroban result interface, the transaction spec exercises the actual create pipeline, TokenBlacklistService resolves the global REDIS_CLIENT token explicitly, and the configuration docs cover every THROTTLE_* env var the docs test asserts. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * feat(rate-limit): configurable limits and client identifiers for public routes Public endpoints are the first thing abusive traffic hits, but their limits were a single fixed IP budget: every public route shared PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one shared address (NAT, office egress, CI runners) exhausted one bucket together. This makes the public rate limiter configurable per route and per client identifier, on the existing Redis sliding-window counter. - Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or per-controller) budget overrides resolved by the guard through Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default. - Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated, currently 'apiKey'): when enabled, a presented x-api-key or ApiKey/Bearer ak_... Authorization header is folded into the bucket key so distinct key-holding clients behind one IP get their own budgets. The IP always participates; keyless callers share the plain-IP bucket as before. - Extract a testable PublicRateLimitGuard.check() returning the full decision (allowed, limit, windowSeconds, count, resetAt); canActivate keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset + Retry-After contract and in-memory fallback on Redis outage. Tests: unit suites for per-route rule resolution, identifier bucketing (with/without identifiers configured), header correctness on allowed and limited requests, and @SkipPublicRateLimit() interaction with rules; an HTTP-level integration suite simulating bursts that proves the 429-with-headers behaviour at the global limit, per-route overrides (next to unaffected sibling routes), and per-key budget isolation. Existing public-rate-limit suites pass unchanged (bucket keys keep the ip: prefix, so stored counters stay compatible). Closes #342 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * fix(ci): resolve typecheck and lint failures across specs, guards, and docs Get CI green for the token-verification cache PR by fixing pre-existing main-branch type errors alongside PR-specific ones: Express specs no longer use Fastify-only app.inject, Stellar mocks match the real Soroban result interface, the transaction spec exercises the actual create pipeline, TokenBlacklistService resolves the global REDIS_CLIENT token explicitly, and the configuration docs cover every THROTTLE_* env var the docs test asserts. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * fix(config): deduplicate parseClientIdentifiers and pass the raw variable The duplicated helper also ignored its parameter while the config factory read the raw variable at the call site with none, so the module never compiled. Collapse to a single parser that takes the raw value. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com> * Provide spending limit dependency in transaction tests * Fix merged audit repository test structure * Fix merged API audit and webhook tests * Align policy velocity test with service architecture * Preserve webhook failure reason in audit * Fix duplicate event emitter declarations * Fix: Remove duplicate db:verify key breaking package.json JSON parsing A duplicate db:verify script entry (missing comma) made package.json invalid JSON, causing npm ci to fail in CI with EJSONPARSE. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Fix: Remove redundant transactionId index causing migration drift RiskAssessment.transactionId is already @unique, which Prisma covers with its own index; the extra @@index([transactionId]) had no corresponding migration and tripped the CI drift check. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: aetheron06 <asynchronousawait@gmail.com> Co-authored-by: tecch-wiz <zynstro@gmail.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: AdaBebe0 <adabebe4040@gmail.com> Co-authored-by: Isaac Gellu <113101848+gelluisaac@users.noreply.github.com> Co-authored-by: IyanuOluwaJesuloba <jesulobaowoseni1@gmail.com> Co-authored-by: Depo-dev <ed4557316@gmail.com> Co-authored-by: S13 <61961655+samad13@users.noreply.github.com> Co-authored-by: Codebuff <noreply@codebuff.com> Co-authored-by: IyanuOluwa Owoseni <141356521+IyanuOluwaJesuloba@users.noreply.github.com> Co-authored-by: Depo.dev <depolonedev@outlook.com> Co-authored-by: Astroid Dev <astroid@dev.local> Co-authored-by: Chijioke Joseph <chijiokejoseph2022@gmail.com> Co-authored-by: xeladev4 <xeladev4@gmail.com> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com> Co-authored-by: gelluisaac <isaacgellu6@gmail.com> Co-authored-by: eliasaph01 <eliasaphmarkus405@gmail.com> Co-authored-by: DarcKnight000 <ahmadry2508@gmail.com> Co-authored-by: Oladayo Oladipupo <oladayo225@gmail.com> Co-authored-by: Deon <110722148+0xDeon@users.noreply.github.com> Co…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
closes #351
Summary
Adds orchestrator-grade liveness and readiness probes and makes the health controller usable by load balancers.
GET /health/livereturns200whenever the process is running. It performs no dependency checks, so a database or cache outage never causes an otherwise healthy process to be restarted.GET /health/readyprobes the database (SELECT 1) and cache (RedisPING) in parallel, each bounded by a 2s timeout, and returns200or503with per-dependency status, latency and error detail:{ "status": "down", "timestamp": "2026-09-28T10:00:00.000Z", "services": { "database": { "status": "down", "latencyMs": 2001, "timestamp": "...", "error": "Database health check timed out after 2000ms" }, "cache": { "status": "up", "latencyMs": 1, "timestamp": "..." } } }/metrics), so probe paths stay stable across API versions. The existing diagnostic routes (/{API_PREFIX}/health,/readiness,/liveness,/database) are unchanged.Fixes found along the way
HealthControllerwas not@Public(), so every probe got a 401 from the globalJwtAuthGuard. It is now public.@SkipThrottle({ api: true, auth: true }); a bare@SkipThrottle()only skips thedefaultthrottler).AuditInterceptorpersists every request, including GETs, so each probe wrote an audit row and tried to during a database outage. The controller is now@SkipAudit().RedisHealthIndicatorread an undefinedREDIS_URLand always probedlocalhost:6379. It also held its own never-closed connection and could hang while ioredis queued the PING. It now probes the sharedREDIS_CLIENTbuilt from the validatedREDIS_*config, with a timeout.Testing
/liveand/ready: all up (200), simulated database outage (503), simulated cache outage (503), empty indicator result (503). Also checked: an external Stellar outage cannot fail readiness, and liveness stays 200 during a database outage.health.http.spec.tsboots the controller behind the realJwtAuthGuardwith the same prefix exclusions asmain.ts, and calls it over HTTP. It verifies the probes are public, unprefixed, and return 503 with structured JSON during a simulated outage. Removing@Public()makes it fail (checked).npm run typecheck,npm run lint,npm test(1095 passing),npm run build.Notes for reviewers
Open PRs #354 and #358 also edit
health.controller.ts. The changes are additive, but whichever merges second may need a small rebase.