Skip to content

fix(deploy_probe): self-heal on first-tick redeploy 404 (DEPLOY-002) - #71

Merged
mastermanas805 merged 4 commits into
masterfrom
fix/deploy-probe-bootstrap-on-404
May 30, 2026
Merged

mastermanas805 merged 4 commits into
masterfrom
fix/deploy-probe-bootstrap-on-404

Conversation

@mastermanas805

Copy link
Copy Markdown
Member

Summary

The hourly synthetic deploy prober (worker#69) wedged on its first tick in prod with 404 no_existing_deployment_to_redeploy — it POSTs redeploy=true from tick 1, but the persistent deploy-probe-hourly app row doesn't exist yet, so the api's typed-error path correctly refused. The probe was DOA until manual bootstrap.

Fix (Option A from the brief): legSubmit detects the canonical error code on a 404 response, logs jobs.deploy_probe.bootstrap_retry, and transparently retries ONCE without redeploy=true. The retry's outcome is reported as result=bootstrap on instant_deploy_probe_outcome_total{leg=submit} — distinct from pass. Subsequent ticks find the row, get a 2xx, and report pass.

  • ONLY the canonical no_existing_deployment_to_redeploy error code triggers the retry. A non-canonical 404 (auth misroute, future api-side regression that drops error) still fails the leg with audit_log + ERROR slog — never mask a real outage as a self-heal.
  • Bootstrap counts as a successful leg-1 for downstream leg dispatch, so status + serve run on the freshly-bootstrapped row.
  • bootstrap is NOT degraded/fail; recordLeg logs at INFO, skips the audit_log INSERT.

Coverage block (rule 17)

Symptom:        404 no_existing_deployment_to_redeploy on first probe tick
Enumeration:    rg -F 'no_existing_deployment_to_redeploy' (api + worker)
Sites found:    2 in api/internal/handlers/deploy.go (sql.ErrNoRows
                and wrong-team defence-in-depth), 0 in worker pre-fix
Sites touched:  1 in worker (new deployProbeRedeployMissingCode const
                consumed by legSubmit retry guard); api sites untouched
                (already correct typed-error contract per api#206)
Coverage test:  TestDeployProbe_Bootstrap_FirstTick404RetriesAsCreate
                + TestDeployProbe_Bootstrap_NonCanonical404StillFails
                + TestBuildDeployProbeMultipart_BootstrapShape
Live verified:  pending merge + deploy; gated on rule 14 SHA check post-rollout

Test plan

  • go test ./internal/jobs/ -run 'TestDeployProbe' -count=1 -short → ok
  • Full make gate equivalent (go build ./... && go vet ./... && go test ./... -short -count=1) → all packages ok
  • New TestDeployProbe_Bootstrap_FirstTick404RetriesAsCreate asserts: exactly 1 POST with redeploy=true, exactly 1 follow-up POST without redeploy, submit=bootstrap, status=pass, serve=pass, no audit_log INSERT
  • New TestDeployProbe_Bootstrap_NonCanonical404StillFails asserts: exactly 1 POST (no retry on error="route_not_found"), submit=fail, audit_log INSERT fires
  • New TestBuildDeployProbeMultipart_BootstrapShape asserts the body OMITS the redeploy field when redeploy=false (so the api takes create semantics)
  • Post-merge: verify /healthz SHA matches HEAD (rule 14), watch first probe tick in NR for result=bootstrap event then steady-state result=pass

Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com

mastermanas805 and others added 2 commits May 31, 2026 00:41
The hourly synthetic deploy prober (worker#69) wedged on its first tick
in prod with `404 no_existing_deployment_to_redeploy` — it POSTs with
`redeploy=true` from tick 1, but the persistent `deploy-probe-hourly`
app row doesn't exist yet, so the api's typed-error path correctly
refused. The probe was DOA until the operator manually bootstrapped.

Fix: legSubmit detects the canonical error code on a 404 response,
logs `jobs.deploy_probe.bootstrap_retry`, and transparently retries
ONCE without `redeploy=true` (create semantics). The retry's outcome
is reported as `result=bootstrap` on
`instant_deploy_probe_outcome_total{leg=submit}` — distinct from
`pass` so the dashboard sees the one-time self-heal as its own event.
Subsequent ticks find the row, get a 2xx, and report `pass`.

Anti-design guardrails:
- ONLY the canonical `no_existing_deployment_to_redeploy` error code
  triggers the retry. A non-canonical 404 (auth misroute, future
  api-side regression that drops `error`) still fails the leg with
  audit_log + ERROR slog — never mask a real outage as a self-heal.
- Bootstrap counts as a successful leg-1 for downstream leg dispatch
  so status + serve run on the freshly-bootstrapped row.
- `bootstrap` is NOT a degraded/fail; recordLeg logs at INFO, skips
  the audit_log INSERT, and the alert NRQL keys only on `result=fail`.

Coverage block (rule 17):
Symptom:        404 no_existing_deployment_to_redeploy on first probe tick
Enumeration:    rg -F 'no_existing_deployment_to_redeploy' (api + worker)
Sites found:    2 in api/internal/handlers/deploy.go (sql.ErrNoRows
                and wrong-team defence-in-depth), 0 in worker pre-fix
Sites touched:  1 in worker (new deployProbeRedeployMissingCode const
                consumed by legSubmit retry guard); api sites untouched
                (already correct typed-error contract per api#206)
Coverage test:  TestDeployProbe_Bootstrap_FirstTick404RetriesAsCreate
                + TestDeployProbe_Bootstrap_NonCanonical404StillFails
                + TestBuildDeployProbeMultipart_BootstrapShape
Live verified:  pending merge + deploy; gated on rule 14 SHA check post-rollout

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@mastermanas805

Copy link
Copy Markdown
Member Author

re-triggering CI (workflow didn't fire on initial push)

@mastermanas805
mastermanas805 merged commit 3b39263 into master May 30, 2026
11 checks passed
@mastermanas805
mastermanas805 deleted the fix/deploy-probe-bootstrap-on-404 branch May 30, 2026 20:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant