Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 51 additions & 17 deletions docs/backends/wslc/wslc-state-aware.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,9 +43,10 @@ Sandbox daemon pattern.
The daemon owns the live SDK handles on a **single apartment-affine worker thread**, which services
every lifecycle command. Any thread that has joined the MTA may use those handles, so an image
pull and an `exec` each run on an MTA thread of their own and post their outcome back to the
worker, leaving it free to serve other sandboxes for the duration of a run. A command naming a
container with a run in flight waits for that run, because deleting the container would free a
handle the run is using. See [Known limitations](#known-limitations).
worker, leaving it free to serve other sandboxes for the duration of a run. A second `exec` on a
container with a run in flight is refused as `busy`; a lifecycle command naming that container
waits for the run, because deleting the container would free a handle the run is using. See
[Known limitations](#known-limitations).

## Components

Expand Down Expand Up @@ -143,11 +144,39 @@ the streaming path reports it through its wait result.

### exec admission, cancellation, and failure containment

The daemon admits one exec at a time. A concurrent exec is rejected with
`backend_error` rather than queued behind an unknown-duration workload. Up to
eight additional control connections can be serviced while an exec owns the
stream slot; connections beyond the daemon's bounded client capacity are
refused.
The daemon admits up to eight exec streams at once, one per container. A second
exec naming a container that already has one in flight is refused with `busy`
before any admission reaches the client; the outward SDK error is
`backend_error`. The bound of eight comes from memory: a streaming exec's output
lives only in its bounded live-output queue, so the persistent per-user daemon
stays near 128 MB of live output even against clients that never drain. An
exec's slot is held until the run has reported back and its client has been
written to, so a client that disconnects mid-run keeps counting against that
bound while its process is still going, and one that drains slowly keeps
counting while its output is still queued. A run whose termination could not be
confirmed leaves its sandbox quarantined and keeps the slot until that sandbox
is deprovisioned, because the process may still be alive.
Comment on lines +156 to +158

Three conditions surface as `busy`, and all reach an SDK caller as
`backend_error`:

| Condition | Message | Retry |
| --- | --- | --- |
| The daemon's eight exec slots are all occupied | `WSLc daemon exec capacity is exhausted` | Succeeds once any exec finishes |
| The named container already has an exec in flight | `sandbox <id> already has an exec in flight` | Succeeds once that container's run finishes |
| Every client slot is occupied and the request is not a cancellation | `WSLc daemon client capacity is exhausted` | Succeeds once any client disconnects |

Up to eight additional control connections can be serviced while every exec slot
is occupied. Beyond that, a connection is admitted only to cancel: cancellation
is the one request that never waits on the worker, and the only way to end a run
with no timeout, so it keeps capacity of its own that lifecycle work cannot
consume. A connection on that lane must send its request within two seconds and
in under 1 KB, so one that connects and stalls cannot hold a cancellation slot
for the general deadline.
Comment on lines +173 to +175

`start` / `stop` / `deprovision` naming a container with an exec in flight
**wait** for that run, because deleting the container would free a handle the
run is still using.

Each exec carries an internal ID and per-run token. A duplicate live ID is
rejected, and cancellation must match both values so a delayed cancellation
Expand Down Expand Up @@ -277,9 +306,14 @@ it can observe idle-teardown within seconds.
host that can reach a registry or already has the image cached, and `wxc-wslc-daemon.exe` staged next to `wxc-exec.exe`). It exercises
core lifecycle, warm-reuse (a marker written by one `exec` is read back by a separate `exec`
process β€” only possible if the container stayed warm), filesystem volumes, bridged networking +
proxy, validation rejections, and idle teardown. Fixtures live in
proxy, validation rejections, exec concurrency, and idle teardown. Fixtures live in
`tests/configs/wslc_state_aware_*.json`.

The concurrency section launches a phase without waiting for it (`Start-StateAware` /
`Wait-StateAware`), which is what lets it observe two sandboxes running at once, a refused
same-container second exec, and a lifecycle command issued while a run is in flight. Every other
section drives one phase process at a time.

### Running the fixtures (ordering + id substitution)

The `wslc_state_aware_*.json` fixtures are **stateful** β€” unlike the one-shot configs, they cannot be
Expand All @@ -302,15 +336,15 @@ fixtures **through the harness**, not by pointing `wxc-exec --config` at them di

## Known limitations

- **Multiple exec streams are deferred.** The daemon admits one exec stream at a time, so a
client's concurrent exec is refused rather than run alongside the first, and the per-container
single-flight slot is not reported as `Busy`. Raising that bound is tracked as follow-up work.
Lifecycle calls on another sandbox are a separate matter: they are admitted through the control
client reserve and proceed while a run is in flight.
- **Ordering is per-container, not global.** A lifecycle command naming a container with a run in
flight waits for that run; commands for other sandboxes proceed independently. A caller cannot
infer that work on one sandbox completed because work on another did.

- **Ordering is per-container, not global.** Commands naming a container with a run in flight wait
for that run; commands for other sandboxes proceed independently. A caller cannot infer that
work on one sandbox completed because work on another did.
- **`busy` collapses to `backend_error` (deferred).** A refused exec reaches an SDK caller as a
generic `backend_error` with no indication that retrying would succeed. A retryable wire code
needs a new `MxcErrorCode` variant, which is a closed set matching the SDK `ErrorCode` union
one-for-one, so it spans Rust, Node, .NET, the versioned references, and schema regeneration.
This is tracked as follow-up work.

- **No typed SDK can set port mappings yet.** The Rust, Node, and .NET v1 SDKs
all pin the published stable contract `1.0.0`, which does not declare the
Expand Down
77 changes: 55 additions & 22 deletions src/mxc-sdk/src/backends/wslc/common/container_steps.rs
Original file line number Diff line number Diff line change
Expand Up @@ -118,11 +118,8 @@ pub enum OutStream {
Stderr,
}

/// Optional live-output sink invoked from the SDK's stdout/stderr callbacks in
/// addition to the capped capture buffers. The daemon supplies one to stream a
/// container's output to the client as bytes arrive; paths that only need the
/// final captured blob (one-shot, detached init) leave it unset. It receives
/// the full callback bytes, independent of the capped buffers' truncation.
/// Optional live-output sink that replaces the capped capture buffers, so a
/// caller that sets one gets empty captured stdout/stderr.
///
/// **Two distinct live-output architectures β€” why they don't share plumbing.**
/// The one-shot runner streams via `OutputMode::Stream` (in `wsl_container_runner`)
Expand All @@ -137,7 +134,7 @@ pub enum OutStream {
/// block** β€” the same SDK thread also delivers the process-exit callback, so the
/// daemon path uses a non-blocking `try_send` that drops on overflow rather than
/// stalling teardown. Keep the two paths separate for those reasons; share only
/// the leaf primitives ([`OutStream`], the capped capture buffers).
/// the leaf primitives ([`OutStream`], [`IoContext`]).
pub type OutputSink = Box<dyn Fn(OutStream, &[u8]) + Send + Sync>;

/// Shared buffer for capturing process I/O via SDK callbacks. Fields are
Expand All @@ -147,10 +144,11 @@ pub struct IoContext {
pub(crate) stdout: Arc<Mutex<Vec<u8>>>,
pub(crate) stderr: Arc<Mutex<Vec<u8>>>,
pub(crate) exited: Arc<(Mutex<bool>, Condvar)>,
/// Live sink for streaming output alongside the capped buffers; `None` when
/// only the final captured blob is needed. Its owning `Arc<IoContext>` is
/// released on the same schedule as the capture buffers, so on the
/// deliberate kill-path leak the sink (and its sender) is leaked too.
/// Live sink for streaming output; `None` when the caller wants the final
/// captured blob instead. A streaming caller reads its bytes from the sink
/// as they arrive, so the capture buffers stay empty. Its owning
/// `Arc<IoContext>` is released on the same schedule as those buffers, so
/// on the deliberate kill-path leak the sink (and its sender) is leaked too.
sink: Option<OutputSink>,
}

Expand Down Expand Up @@ -201,24 +199,20 @@ unsafe extern "C" fn io_callback(
let ctx = &*(context as *const IoContext);
let bytes = std::slice::from_raw_parts(data, data_size as usize);
match io_handle {
WslcProcessIOHandle::WSLC_PROCESS_IO_HANDLE_STDOUT => {
{
WslcProcessIOHandle::WSLC_PROCESS_IO_HANDLE_STDOUT => match ctx.sink.as_ref() {
Comment thread
SohamDas2021 marked this conversation as resolved.
Some(sink) => sink(OutStream::Stdout, bytes),
None => {
let mut buf = ctx.stdout.lock().unwrap_or_else(|e| e.into_inner());
append_capped(&mut buf, bytes);
}
if let Some(sink) = ctx.sink.as_ref() {
sink(OutStream::Stdout, bytes);
}
}
WslcProcessIOHandle::WSLC_PROCESS_IO_HANDLE_STDERR => {
{
},
WslcProcessIOHandle::WSLC_PROCESS_IO_HANDLE_STDERR => match ctx.sink.as_ref() {
Some(sink) => sink(OutStream::Stderr, bytes),
None => {
let mut buf = ctx.stderr.lock().unwrap_or_else(|e| e.into_inner());
append_capped(&mut buf, bytes);
}
if let Some(sink) = ctx.sink.as_ref() {
sink(OutStream::Stderr, bytes);
}
}
},
_ => {}
}
}
Expand Down Expand Up @@ -1351,6 +1345,45 @@ mod tests {
);
}

/// A streaming caller reads its bytes from the sink, so a second capped
/// copy would double the daemon's per-exec output memory.
#[test]
fn a_streaming_callback_captures_nothing() {
let streamed = Arc::new(Mutex::new(Vec::new()));
let seen = Arc::clone(&streamed);
let io_ctx = Arc::new(IoContext {
stdout: Arc::new(Mutex::new(Vec::new())),
stderr: Arc::new(Mutex::new(Vec::new())),
exited: Arc::new((Mutex::new(false), Condvar::new())),
sink: Some(Box::new(move |kind, bytes| {
seen.lock().unwrap().push((kind, bytes.to_vec()));
})),
});

let payload = b"streamed-not-captured";
let raw = Arc::into_raw(Arc::clone(&io_ctx));
// SAFETY: `raw` is a live `Arc<IoContext>` pointer and `payload` is
// valid for its full length, matching what the SDK passes.
unsafe {
io_callback(
WslcProcessIOHandle::WSLC_PROCESS_IO_HANDLE_STDOUT,
payload.as_ptr(),
payload.len() as u32,
raw as *mut c_void,
);
drop(Arc::from_raw(raw));
}

assert_eq!(
streamed.lock().unwrap().as_slice(),
&[(OutStream::Stdout, payload.to_vec())]
);
assert!(
io_ctx.stdout.lock().unwrap().is_empty(),
"a streamed chunk must not also be captured"
);
}

#[test]
fn null_exit_event_still_observes_cancellation() {
let io_ctx = IoContext {
Expand Down
Loading
Loading