Skip to content

[Flaky test]: workflow launch replay can race durable surface owner lease release under parallel Linux tests #119

Description

@echoVic

Environment

  • Orca: 0.4.32
  • Source: 66a68035c6ee5851fa519aae1c6a131c9699f2ac
  • Host: Linux x86_64
  • Rust: 1.98.1
  • Runner: cargo nextest, profile ci, framework retries disabled

Summary

The following test can fail during a parallel workspace run because workflow replay attempts to acquire the typed surface owner lease before the previous owner has been observably released:

orca-runtime::runtime_surface_host
workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing

Failure

During the parallel run:

take over workflow thread: ThreadStartFailed {
    message: "failed to acquire typed surface owner lease: AlreadyOwned"
}

Location:

crates/orca-runtime/tests/runtime_surface_host.rs:1692

Reproduction evidence

A workspace run excluding the sandbox-dependent suites executed 3169 tests:

3168 passed
1 failed

The only failing test was:

workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing

The same test was then run independently and serially three times:

for i in 1 2 3; do
  cargo test \
    -p orca-runtime \
    --test runtime_surface_host \
    workflow_launch_replay_uses_surface_identity_when_activation_store_is_missing \
    --locked \
    -- \
    --nocapture \
    --test-threads=1
done

Result:

3/3 passed

This points to a parallel scheduling/lifecycle race rather than a deterministic functional failure.

Expected behavior

Once the first runtime owner has been shut down or relinquished, replay should wait for or observe the committed lease-release boundary before attempting takeover.

The test should not depend on scheduler timing or filesystem visibility.

Suggested fix

  • Introduce an observable synchronization boundary proving that the previous owner lease has been released before takeover begins.
  • If lease release is durable, wait for the committed revision or ownership transition rather than process/thread completion alone.
  • Ensure all owner exit paths release or fence the lease symmetrically.
  • Add a parallel stress regression that repeatedly exercises:
    1. initial owner acquisition;
    2. activation-store removal;
    3. owner shutdown/release;
    4. replay takeover.

Please avoid fixing this with a larger sleep, blind retry, or relaxed assertion; those would only lower the failure probability.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions