Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 16 additions & 12 deletions docs/backends/wslc/wslc-state-aware.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,10 +40,12 @@ the live SDK handles. Each phase process is a thin client that contacts the daem
pipe; the daemon performs the actual SDK calls and streams stdio back. This mirrors the Windows
Sandbox daemon pattern.

The daemon runs all SDK calls on a **single apartment-affine worker thread** (the WSLc SDK handles
are not thread-agnostic). Today `exec` blocks that worker for the duration of the run, so commands
against different sandboxes are serialized β€” correct, just not concurrent. See
[Known limitations](#known-limitations).
The daemon owns the live SDK handles on a **single apartment-affine worker thread**, which services
every lifecycle command. Any thread that has joined the MTA may use those handles, so an image
pull and an `exec` each run on an MTA thread of their own and post their outcome back to the
worker, leaving it free to serve other sandboxes for the duration of a run. A command naming a
container with a run in flight waits for that run, because deleting the container would free a
handle the run is using. See [Known limitations](#known-limitations).

## Components

Expand Down Expand Up @@ -141,8 +143,7 @@ the streaming path reports it through its wait result.

### exec admission, cancellation, and failure containment

The daemon runs one exec at a time because all WSLc SDK operations are confined
to one apartment-affine worker. A concurrent exec is rejected with
The daemon admits one exec at a time. A concurrent exec is rejected with
`backend_error` rather than queued behind an unknown-duration workload. Up to
eight additional control connections can be serviced while an exec owns the
stream slot; connections beyond the daemon's bounded client capacity are
Expand Down Expand Up @@ -301,12 +302,15 @@ fixtures **through the harness**, not by pointing `wxc-exec --config` at them di

## Known limitations

- **Serialized exec (deferred).** Because the daemon's single worker thread blocks on
`WaitForSingleObject` for the whole `exec`, no other sandbox can provision or exec while one
command runs, and a per-container single-flight `Busy` guard is not yet meaningful. The intended
fix splits `exec` into an on-worker `ExecStart` (extract the thread-agnostic Win32 exit-event
handle) + an off-thread wait + an on-worker `ExecReap`, with a per-container `in_flight` slot. This
is tracked as follow-up work.
- **Multiple exec streams are deferred.** The daemon admits one exec stream at a time, so a
client's concurrent exec is refused rather than run alongside the first, and the per-container
single-flight slot is not reported as `Busy`. Raising that bound is tracked as follow-up work.
Lifecycle calls on another sandbox are a separate matter: they are admitted through the control
client reserve and proceed while a run is in flight.

- **Ordering is per-container, not global.** Commands naming a container with a run in flight wait
for that run; commands for other sandboxes proceed independently. A caller cannot infer that
work on one sandbox completed because work on another did.

- **No typed SDK can set port mappings yet.** The Rust, Node, and .NET v1 SDKs
all pin the published stable contract `1.0.0`, which does not declare the
Expand Down
70 changes: 63 additions & 7 deletions src/mxc-sdk/src/bin/wslc_daemon/control_server.rs
Original file line number Diff line number Diff line change
Expand Up @@ -48,8 +48,8 @@ use mxc_sdk::wslc_common::daemon_protocol::{

use crate::session_manager::{ExecStream, SessionHandle, WorkerError};

/// The WSLc SDK worker is apartment-affine and executes one command at a time.
/// Refuse additional execs instead of admitting a queue that cannot run.
/// Exec streams admitted at once. A further request is refused rather than
/// queued behind a workload of unknown duration.
const MAX_CONCURRENT_EXECS: usize = 1;

/// Capacity reserved for cancellation and lifecycle requests while all exec
Expand Down Expand Up @@ -510,10 +510,11 @@ where
/// [`StreamFrame`]s, followed by a terminal frame.
///
/// The sandbox is validated (exists + started) *before* the `Ok` admission is
/// written, and β€” critically β€” admission is **atomic** with the start of the
/// run on the worker thread (see [`SessionHandle::exec`]): the worker validates
/// and begins running within one command handler, so no `Stop`/`Deprovision`
/// can invalidate the checked state between the admission and the run. An
/// written, and β€” critically β€” admission is **atomic** with the claim the
/// worker takes on the container (see [`SessionHandle::exec`]): the worker
/// validates, claims the container and hands the run to a thread of its own
/// without yielding, and every later command naming that container parks behind
/// the claim, so no `Stop`/`Deprovision` can invalidate the checked state. An
/// unknown/not-started sandbox therefore comes back as a pre-admission typed
/// [`DaemonResponse::Err`] rather than a post-admission stream `Error` frame.
///
Expand All @@ -533,9 +534,21 @@ async fn handle_exec<S>(
where
S: AsyncWrite + Unpin,
{
let exec_id = config.exec_id.clone();
let run_token = config.run_token.clone();

// Await the worker's admission decision before writing anything: a rejected
// exec is a pre-admission typed error, never a post-admission stream frame.
write_exec_result(&mut pipe, session.exec(config).await).await
let delivered = write_exec_result(&mut pipe, session.exec(config).await).await;

if delivered.is_err() {
Comment on lines +542 to +544
// The run outlives this handler on a thread of its own, and the permit
// that bounds exec capacity is released as this returns. Without a kill
// a client could disconnect in a loop and leave runs going unbounded.
session.cancel_exec(&exec_id, &run_token);
}

delivered
}

/// Turn an exec **admission** outcome into the client's frame sequence, generic
Expand Down Expand Up @@ -739,6 +752,49 @@ mod tests {
session.shutdown().await.unwrap();
}

#[tokio::test]
async fn an_undelivered_exec_is_cancelled_so_its_run_cannot_outlive_the_permit() {
use crate::session_manager::register_exec;

let session = crate::session_manager::spawn().unwrap();
let cancellation = Arc::new(AtomicBool::new(false));
let registration = register_exec(
session.active_execs(),
"orphan-1",
"orphan-run-1",
&cancellation,
)
.unwrap();

let limiter = Arc::new(Semaphore::new(MAX_CONCURRENT_EXECS));
let permit = limiter.clone().try_acquire_owned().unwrap();
let delivered = handle_exec(
BrokenPipe,
session.clone(),
mxc_sdk::wslc_common::daemon_protocol::ExecConfig {
exec_id: "orphan-1".to_string(),
run_token: "orphan-run-1".to_string(),
sandbox_id: "wslc:does-not-exist".to_string(),
script_code: "sleep 600".to_string(),
working_directory: String::new(),
env: Vec::new(),
env_scope: mxc_sdk::wslc_common::process_env::EnvScope::Merge,
timeout_ms: 0,
},
permit,
)
.await;

assert!(delivered.is_err(), "the broken pipe must fail delivery");
assert!(
cancellation.load(Ordering::Acquire),
"a run the client can no longer read must be cancelled, or it keeps \
going after its capacity permit is released"
);
drop(registration);
session.shutdown().await.unwrap();
}

/// A writer that always fails, standing in for a client that has gone.
struct BrokenPipe;

Expand Down
Loading
Loading