Resolve the storage driver in a scope of its own - #2031
Merged
Merged
Conversation
A healthy agent run - 22 minutes of review work on codespace-test - landed Failed with "A second operation was started on this context instance before a previous operation completed". The executor, the log capture bridge and everything the bridge reaches are resolved from one lifetime scope, the Hangfire job's, and the capture loop runs beside the durable runner's drain tick. The CAS coordinator runs its own statements on contexts of its own, but the storage driver broker it was handed carried the scope's CodeSpaceDbContext in its profile and credential readers. So a segment append resolving its driver while the tick flushed events or wrote the spool offset put two statements on one context. EF refused whichever started second: on the capture side that read as a transient storage stall, on the tick's side it escaped the observer and failed the run (observer-failed-before-terminal, the attempt closed Lost). The coordinator now resolves the broker in a child scope per open, so nothing it reaches shares its caller's context. Every coordinator entry point resolves the profile revision on its own context first, so the broker never needed the caller's transaction; the consumers that do (default materialization, route probing) call the broker directly and are unchanged. Every earlier capture test resolved the bridge from a different scope than the executor, which is why none saw this. The new tests resolve it the way the job scope does and hold the broker's read in flight.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
DbContext. On codespace-test, a 22-minute PR review agent run landed Failed with "A second operation was started on this context instance before a previous operation completed". Its log streams closedobserver-failed-before-terminaland the attempt closed Lost.AgentRunLogCaptureBridge.CaptureLoopAsync) runs beside the durable runner's drain tick.ArtifactCasRuntimeCoordinatoralready runs its own statements on contexts of its own. The storage driver broker it was constructed with, however, carried the scope'sCodeSpaceDbContextin its profile and credential readers (StorageProfileSnapshotResolver,StorageCredentialSecretResolver).ArtifactCasRuntimeCoordinatornow resolves the broker in a child lifetime scope per open (ArtifactCasRuntimeCoordinator.cs,OpenDriverInOwnScopeAsync). It follows the same child-of-the-owning-scope patternNodeInvocationExecutorandWorkflowEngineuse.Test plan
AgentRunExecutorSharedScopeTests(integration, high fidelity). The executor is resolved the way the job scope resolves it, so bridge, log service, coordinator, broker and resolvers are the production graph. The runner is the realLocalProcessRunnerand the database is real Postgres. Every earlier capture test resolved the bridge from a different scope, which is why none caught this.ACCESS EXCLUSIVElock onstorage_credential, held on a separate connection, keeps the broker's credential read in flight while the agent keeps printing. On the append path nothing else reads that table.pg_stat_activitywitness asserts the read really was waiting (fixture check).