Environment
- Orca:
0.4.32
- Source:
66a68035c6ee5851fa519aae1c6a131c9699f2ac
- Host: Linux x86_64
- Rust:
1.98.1
- Runner:
cargo nextest, profile ci, framework retries disabled
Summary
The following test can fail during a parallel workspace run because workflow replay attempts to acquire the typed surface owner lease before the previous owner has been observably released:
orca-runtime::runtime_surface_host
workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing
Failure
During the parallel run:
take over workflow thread: ThreadStartFailed {
message: "failed to acquire typed surface owner lease: AlreadyOwned"
}
Location:
crates/orca-runtime/tests/runtime_surface_host.rs:1692
Reproduction evidence
A workspace run excluding the sandbox-dependent suites executed 3169 tests:
The only failing test was:
workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing
The same test was then run independently and serially three times:
for i in 1 2 3; do
cargo test \
-p orca-runtime \
--test runtime_surface_host \
workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing \
--locked \
-- \
--nocapture \
--test-threads=1
done
Result:
This points to a parallel scheduling/lifecycle race rather than a deterministic functional failure.
Expected behavior
Once the first runtime owner has been shut down or relinquished, replay should wait for or observe the committed lease-release boundary before attempting takeover.
The test should not depend on scheduler timing or filesystem visibility.
Suggested fix
- Introduce an observable synchronization boundary proving that the previous owner lease has been released before takeover begins.
- If lease release is durable, wait for the committed revision or ownership transition rather than process/thread completion alone.
- Ensure all owner exit paths release or fence the lease symmetrically.
- Add a parallel stress regression that repeatedly exercises:
- initial owner acquisition;
- activation-store removal;
- owner shutdown/release;
- replay takeover.
Please avoid fixing this with a larger sleep, blind retry, or relaxed assertion; those would only lower the failure probability.
Environment
0.4.3266a68035c6ee5851fa519aae1c6a131c9699f2ac1.98.1cargo nextest, profileci, framework retries disabledSummary
The following test can fail during a parallel workspace run because workflow replay attempts to acquire the typed surface owner lease before the previous owner has been observably released:
Failure
During the parallel run:
Location:
Reproduction evidence
A workspace run excluding the sandbox-dependent suites executed 3169 tests:
The only failing test was:
The same test was then run independently and serially three times:
Result:
This points to a parallel scheduling/lifecycle race rather than a deterministic functional failure.
Expected behavior
Once the first runtime owner has been shut down or relinquished, replay should wait for or observe the committed lease-release boundary before attempting takeover.
The test should not depend on scheduler timing or filesystem visibility.
Suggested fix
Please avoid fixing this with a larger sleep, blind retry, or relaxed assertion; those would only lower the failure probability.