fix(deploy_probe): self-heal on first-tick redeploy 404 (DEPLOY-002) - #71
Merged
Merged
Conversation
The hourly synthetic deploy prober (worker#69) wedged on its first tick
in prod with `404 no_existing_deployment_to_redeploy` — it POSTs with
`redeploy=true` from tick 1, but the persistent `deploy-probe-hourly`
app row doesn't exist yet, so the api's typed-error path correctly
refused. The probe was DOA until the operator manually bootstrapped.
Fix: legSubmit detects the canonical error code on a 404 response,
logs `jobs.deploy_probe.bootstrap_retry`, and transparently retries
ONCE without `redeploy=true` (create semantics). The retry's outcome
is reported as `result=bootstrap` on
`instant_deploy_probe_outcome_total{leg=submit}` — distinct from
`pass` so the dashboard sees the one-time self-heal as its own event.
Subsequent ticks find the row, get a 2xx, and report `pass`.
Anti-design guardrails:
- ONLY the canonical `no_existing_deployment_to_redeploy` error code
triggers the retry. A non-canonical 404 (auth misroute, future
api-side regression that drops `error`) still fails the leg with
audit_log + ERROR slog — never mask a real outage as a self-heal.
- Bootstrap counts as a successful leg-1 for downstream leg dispatch
so status + serve run on the freshly-bootstrapped row.
- `bootstrap` is NOT a degraded/fail; recordLeg logs at INFO, skips
the audit_log INSERT, and the alert NRQL keys only on `result=fail`.
Coverage block (rule 17):
Symptom: 404 no_existing_deployment_to_redeploy on first probe tick
Enumeration: rg -F 'no_existing_deployment_to_redeploy' (api + worker)
Sites found: 2 in api/internal/handlers/deploy.go (sql.ErrNoRows
and wrong-team defence-in-depth), 0 in worker pre-fix
Sites touched: 1 in worker (new deployProbeRedeployMissingCode const
consumed by legSubmit retry guard); api sites untouched
(already correct typed-error contract per api#206)
Coverage test: TestDeployProbe_Bootstrap_FirstTick404RetriesAsCreate
+ TestDeployProbe_Bootstrap_NonCanonical404StillFails
+ TestBuildDeployProbeMultipart_BootstrapShape
Live verified: pending merge + deploy; gated on rule 14 SHA check post-rollout
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Member
Author
|
re-triggering CI (workflow didn't fire on initial push) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The hourly synthetic deploy prober (worker#69) wedged on its first tick in prod with
404 no_existing_deployment_to_redeploy— it POSTsredeploy=truefrom tick 1, but the persistentdeploy-probe-hourlyapp row doesn't exist yet, so the api's typed-error path correctly refused. The probe was DOA until manual bootstrap.Fix (Option A from the brief):
legSubmitdetects the canonical error code on a 404 response, logsjobs.deploy_probe.bootstrap_retry, and transparently retries ONCE withoutredeploy=true. The retry's outcome is reported asresult=bootstraponinstant_deploy_probe_outcome_total{leg=submit}— distinct frompass. Subsequent ticks find the row, get a 2xx, and reportpass.no_existing_deployment_to_redeployerror code triggers the retry. A non-canonical 404 (auth misroute, future api-side regression that dropserror) still fails the leg with audit_log + ERROR slog — never mask a real outage as a self-heal.bootstrapis NOT degraded/fail;recordLeglogs at INFO, skips the audit_log INSERT.Coverage block (rule 17)
Test plan
go test ./internal/jobs/ -run 'TestDeployProbe' -count=1 -short→okmake gateequivalent (go build ./... && go vet ./... && go test ./... -short -count=1) → all packagesokTestDeployProbe_Bootstrap_FirstTick404RetriesAsCreateasserts: exactly 1 POST withredeploy=true, exactly 1 follow-up POST withoutredeploy,submit=bootstrap,status=pass,serve=pass, noaudit_logINSERTTestDeployProbe_Bootstrap_NonCanonical404StillFailsasserts: exactly 1 POST (no retry onerror="route_not_found"),submit=fail,audit_logINSERT firesTestBuildDeployProbeMultipart_BootstrapShapeasserts the body OMITS theredeployfield whenredeploy=false(so the api takes create semantics)/healthzSHA matches HEAD (rule 14), watch first probe tick in NR forresult=bootstrapevent then steady-stateresult=passCo-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com