Skip to content

Make a failure and the build that produced it visible in the logs - #1399

Merged
ppXD merged 1 commit into
mainfrom
fix/failures-are-always-logged
Aug 14, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/failures-are-always-logged

Conversation

@ppXD

@ppXD ppXD commented Aug 14, 2026

Copy link
Copy Markdown
Owner

Summary

A pod answered /api/auth/sign-in with 42703: column t.owner_user_id does not exist — a column #1379 renamed, which only pre-#1379 code asks for — while the latest build had been pushed. Diagnosing it took three rounds because of three separate gaps, each fixed here.

Nothing recorded which build a pod was running. Every event now carries Build (from AssemblyInformationalVersion, which the SDK stamps with the commit sha), and the first line of each boot names it, the environment, and the log destination.

The searchable copy went nowhere. Serilog:Seq:ServerUrl shipped as http://localhost:5341; every deployment inherits it, and in a container that resolves to itself. The batched sink retries against nothing and an empty Seq reads as silence. The base file now ships blank — console-only, and it says so at startup — with the zero-config local Seq moved to appsettings.Development.json. Also adds ReadFrom.Configuration + Serilog:MinimumLevel, so levels are tunable without a rebuild.

Failures thrown outside the MediatR pipeline reached no sink. GlobalExceptionFilter logged nothing, reasoning that RequestFailureObserver already had — but that only sees exceptions escaping a handler. Anything from a controller, model binding or an auth filter was a masked 500 with no line on any sink. The observer now marks what it records; the filter records what nothing else did, at the same severity.

DbUp no longer "succeeds" having found nothing. Scripts travel as content files beside the assembly; if that copy is ever missing, PerformUpgrade returns success having applied nothing and the app serves an unmigrated database — which presents exactly as this incident did.

Note on embedding the migrations

Adding <EmbeddedResource> alongside the existing content copy — the obvious hardening — would re-run all 124 migrations on any live database. DbUp journals by name, and the two providers name the same file differently:

0001_initial.sql                                       ← file system, what every DB has journalled
CodeSpace.Core.Persistence.DbUpFiles.0001_initial.sql  ← embedded, unapplied under this name

MigrationDiscoveryTests now fails if anyone adds a second script source, and the dead embedded provider (which matched nothing, and is what made the change look free) is removed.

Test plan

  • FailuresAreAlwaysLoggedTests (6) — mutation-verified: filter stops logging → 3 red; observer stops marking → 1 red
  • MigrationDiscoveryTests (3) — reded on the real duplicate-name hazard before it was reverted
  • SeqSinkSettingsTests updated to pin both halves: base ships blank, Development ships localhost
  • Unit 6474/6474 · Integration 2745/2745 · E2E 180/180

A pod answered sign-in with a Postgres 42703 for a column a migration had
already renamed -- which only the code from before that migration asks
for -- while the operator had built and pushed the version after it.
Three things had to be true for that to take three rounds to diagnose.

Nothing recorded the build. Every log event now carries Build, and the
first line of every boot names it alongside the environment and where
searchable logs are going, so "is this pod running the code I think it
is" is answerable from one line.

The searchable copy went nowhere. Serilog:Seq:ServerUrl shipped as
http://localhost:5341, which every deployment inherits and which in a
container resolves to itself. The batched sink retries against nothing
and the operator reads an empty Seq as silence. The base file now ships
blank -- console-only, and it says so -- and the local convenience moves
to appsettings.Development.json.

Failures thrown outside the MediatR pipeline reached no sink at all.
GlobalExceptionFilter logged nothing on the grounds that
RequestFailureObserver had, but that only sees exceptions escaping a
handler; anything from a controller, model binding or an auth filter was
a masked 500 with no line anywhere. The filter now records what the
observer marked as unrecorded, at the same severity.

Also refuses to start when DbUp discovers no migration scripts. They
travel as content files beside the assembly, and if that copy is ever
missing PerformUpgrade reports success having applied nothing, which
presents exactly as this incident did.
@ppXD
ppXD merged commit 39b59e2 into main Aug 14, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant