From 1700586b799162172ec00d3446bc6848dc1cb33d Mon Sep 17 00:00:00 2001 From: Manas Srivastava Date: Sat, 30 May 2026 17:05:25 +0530 Subject: [PATCH] feat(obs): AUTH-004 synthetic prober alert + dashboard + catalog MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds the rule-25 observability surface for the worker AUTH-004 synthetic prober (worker repo PR #68): - NR alert (P0): newrelic/alerts/auth-probe-fail.json fires CRITICAL on any instant_auth_probe_outcome_total{result="fail"} over a 10m window (two consecutive 5-min ticks of fail). - Prom rule: k8s/prometheus-rules.yaml::AuthProbeFail (same shape). - Dashboard tiles: 3 new tiles on instanode-reliability.json — per-leg outcomes line, fails billboard (CRITICAL threshold ≥1), P95 latency per leg. - Catalog rows: instant_auth_probe_outcome_total + instant_auth_probe_latency_seconds in observability/METRICS-CATALOG.md. Lands in lockstep with worker PR #68 per CLAUDE.md rule 25 ("every new metric ships with its alert + dashboard tile in the same PR"). The two repos auto-deploy on separate workflows so they ship as two PRs against their respective masters. Co-Authored-By: Claude Opus 4.7 (1M context) --- k8s/prometheus-rules.yaml | 28 ++++++++++++ newrelic/alerts/auth-probe-fail.json | 31 +++++++++++++ .../dashboards/instanode-reliability.json | 45 +++++++++++++++++++ observability/METRICS-CATALOG.md | 2 + 4 files changed, 106 insertions(+) create mode 100644 newrelic/alerts/auth-probe-fail.json diff --git a/k8s/prometheus-rules.yaml b/k8s/prometheus-rules.yaml index 3065bbd..3e56266 100644 --- a/k8s/prometheus-rules.yaml +++ b/k8s/prometheus-rules.yaml @@ -614,3 +614,31 @@ spec: upgrade. Find the team in audit_log where audit_kind = 'billing.charge_undeliverable', cross-check against the Razorpay dashboard, refund or hand-grant the tier. + + # AUTH-004 synthetic prober. Every 5 minutes the worker drives the full + # browser-shaped login loop against prod (/auth/email/start + + # /auth/exchange CORS contract + /auth/me bearer). Any fail outcome + # over a 10-minute window pages — two consecutive ticks of fail is + # unambiguous regression, not flake. Mirrors auth-probe-fail.json in + # newrelic/alerts/. + - name: instant-worker-auth-probe + rules: + - alert: AuthProbeFail + expr: | + sum(increase(instant_auth_probe_outcome_total{result="fail"}[10m])) > 0 + for: 10m + labels: + severity: critical + service: worker + annotations: + summary: "AUTH-004 synthetic prober fail (login loop broken in prod)" + description: | + instant_auth_probe_outcome_total{result="fail"} > 0 in 10m. + The synthetic prober drives the full browser-shaped login + loop against prod every 5 minutes. A fail outcome means one + of: /auth/email/start non-202, /auth/exchange missing + Access-Control-Allow-Origin / Access-Control-Allow-Credentials + (the AUTH-004 regression class), or /auth/me bearer rejected. + Cross-correlate against audit_log kind=auth_probe_failed + + the worker structured slog ERROR `auth_probe_failed leg=... + reason=...`. Source: worker/internal/jobs/auth_probe.go. diff --git a/newrelic/alerts/auth-probe-fail.json b/newrelic/alerts/auth-probe-fail.json new file mode 100644 index 0000000..a5c935c --- /dev/null +++ b/newrelic/alerts/auth-probe-fail.json @@ -0,0 +1,31 @@ +{ + "name": "instant-worker — auth_probe_failed [synthetic login loop broken in prod]", + "type": "NRQL", + "description": "P0 page on ANY occurrence of instant_auth_probe_outcome_total{result=\"fail\"}. The AUTH-004 synthetic prober drives the full browser-shaped login loop against prod every 5 minutes (POST /auth/email/start + OPTIONS/POST /auth/exchange CORS contract + GET /auth/me bearer). A fail outcome means one of: (a) /auth/email/start is returning non-202, (b) /auth/exchange is missing the Access-Control-Allow-Origin or Access-Control-Allow-Credentials response header (the exact regression class that hid prod login broken for ~24h on 2026-05-29), or (c) the probe bearer token rejected at /auth/me. Cross-correlate against the audit_log row written by worker auth_probe_failed (kind=auth_probe_failed, actor='system:auth_probe') AND the structured slog ERROR line auth_probe_failed leg=... reason=... — same content on both surfaces. Source: worker/internal/jobs/auth_probe.go (AuthProbeWorker), metric registered in worker/internal/metrics/metrics.go (AuthProbeOutcomeTotal). Threshold ABOVE 0 with 10m window: two consecutive ticks (one tick = 5min) of fail is unambiguous regression, not flake.", + "enabled": true, + "nrql": { + "query": "SELECT sum(instant_auth_probe_outcome_total) FROM Metric WHERE metricName = 'instant_auth_probe_outcome_total' AND result = 'fail'" + }, + "terms": [ + { + "priority": "CRITICAL", + "operator": "ABOVE", + "threshold": 0, + "thresholdDuration": 600, + "thresholdOccurrences": "AT_LEAST_ONCE" + } + ], + "signal": { + "aggregationWindow": 60, + "aggregationMethod": "EVENT_FLOW", + "aggregationDelay": 120, + "fillOption": "STATIC", + "fillValue": 0 + }, + "expiration": { + "expirationDuration": 3600, + "openViolationOnExpiration": false, + "closeViolationsOnExpiration": true + }, + "violationTimeLimitSeconds": 86400 +} diff --git a/newrelic/dashboards/instanode-reliability.json b/newrelic/dashboards/instanode-reliability.json index 69ca320..3082c16 100644 --- a/newrelic/dashboards/instanode-reliability.json +++ b/newrelic/dashboards/instanode-reliability.json @@ -258,6 +258,51 @@ ], "platformOptions": { "ignoreTimeRange": false } } + }, + { + "title": "AUTH-004 synthetic prober — outcomes per leg (1h)", + "layout": { "column": 1, "row": 25, "width": 6, "height": 3 }, + "visualization": { "id": "viz.line" }, + "rawConfiguration": { + "nrqlQueries": [ + { + "accountIds": [0], + "query": "SELECT sum(instant_auth_probe_outcome_total) FROM Metric WHERE metricName = 'instant_auth_probe_outcome_total' FACET leg, result TIMESERIES SINCE 1 hour ago" + } + ], + "platformOptions": { "ignoreTimeRange": false } + } + }, + { + "title": "AUTH-004 synthetic prober — fails (last 1h, must be 0)", + "layout": { "column": 7, "row": 25, "width": 3, "height": 3 }, + "visualization": { "id": "viz.billboard" }, + "rawConfiguration": { + "nrqlQueries": [ + { + "accountIds": [0], + "query": "SELECT sum(instant_auth_probe_outcome_total) AS 'fails' FROM Metric WHERE metricName = 'instant_auth_probe_outcome_total' AND result = 'fail' SINCE 1 hour ago" + } + ], + "platformOptions": { "ignoreTimeRange": false }, + "thresholds": [ + { "alertSeverity": "CRITICAL", "value": 1 } + ] + } + }, + { + "title": "AUTH-004 synthetic prober — P95 latency per leg (1h)", + "layout": { "column": 10, "row": 25, "width": 3, "height": 3 }, + "visualization": { "id": "viz.line" }, + "rawConfiguration": { + "nrqlQueries": [ + { + "accountIds": [0], + "query": "SELECT percentile(instant_auth_probe_latency_seconds, 95) AS 'p95' FROM Metric WHERE metricName = 'instant_auth_probe_latency_seconds' FACET leg TIMESERIES SINCE 1 hour ago" + } + ], + "platformOptions": { "ignoreTimeRange": false } + } } ] } diff --git a/observability/METRICS-CATALOG.md b/observability/METRICS-CATALOG.md index 059bf9a..ebcce2b 100644 --- a/observability/METRICS-CATALOG.md +++ b/observability/METRICS-CATALOG.md @@ -39,6 +39,8 @@ fires. Operators need this so they don't panic when a fresh deploy looks | `email_missing_renderer_total` | worker | `kind` | lazy (CounterVec — any tick is a bug, label series only appears on the broken kind) | `email-missing-renderer.json` | `EmailMissingRenderer` | "Email missing-renderer ticks (any > 0 == P0)" | | `migration_version`, `migration_count`, `migration_status` (worker `/healthz` JSON fields, NOT Prometheus metrics) | worker | n/a (log-based) | **eager** (read live from `schema_migrations` table by `migrations.Reader`, cached 60s) | `worker-migration-mismatch.json` (log-based) | n/a (log-based) | "Worker /healthz migration_count drift" | | `instant_idempotency_replay_refunded_total` | api | `route` | lazy (CounterVec — first cache HIT on each route materialises the label series; a fresh deploy with no retries reports nothing until the first agent retries with the same `Idempotency-Key`) | `idempotency-replay-refund-spike.json` | `IdempotencyReplayRefundSpike` | "Idempotency replay refunds by route (1h) — FINDING API-1" | +| `instant_auth_probe_outcome_total` | worker | `leg,result` | lazy (CounterVec — `pass`/`degraded` materialise on the first happy tick; `fail` only appears after a real regression. AUTH-004 synthetic prober: every 5 min the worker drives /auth/email/start + /auth/exchange CORS contract + /auth/me bearer against prod) | `auth-probe-fail.json` | `AuthProbeFail` | "AUTH-004 synthetic prober — outcomes per leg (1h)", "AUTH-004 synthetic prober — fails (last 1h, must be 0)" | +| `instant_auth_probe_latency_seconds` | worker | `leg` | lazy (HistogramVec — observation only on a real HTTP response; DNS/TCP errors omit the observation so the histogram isn't polluted with 0s timeouts) | (covered by `auth-probe-fail.json`) | (covered by `AuthProbeFail`) | "AUTH-004 synthetic prober — P95 latency per leg (1h)" | ## Lazy-emit gotcha — what operators should expect