Skip to content

Enforce workspace storage bounds - #665

Open
SaladDay wants to merge 3 commits into
mainfrom
enforce-workspace-storage-bounds
Open

SaladDay wants to merge 3 commits into
mainfrom
enforce-workspace-storage-bounds

Conversation

@SaladDay

@SaladDay SaladDay commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

Workspace capacity was not enforced by the filesystem, and a tenant could reserve unbounded aggregate storage across workspace configurations. Persist each workspace's admitted capacity and enforce one tenant capacity policy across all non-deleted reservations. Existing bindings retain their configuration and capacity; a retained unbounded workspace blocks further admission for that tenant while the policy is enabled.

The filesystem protocol now explicitly separates credential normalization, credentialed Core control and public resolution. Private workspace credentials are encrypted in PostgreSQL and omitted from public configuration, bindings and audit output. The NFS adapter implements positive capacity through a pinned-host-key, restricted OpenSSH command backed by XFS byte/inode project quotas and a durable SQLite allocation journal. Full-quota deletion keeps lifecycle metadata outside the charged data project. The operator documentation, Chinese translations, generated contracts/client and distribution helper are included.

Preparation failures now release their exact unallocated node placement under the Session/deployment locks and rejoin demand behind waiting requests. This prevents a quota-refused Session from consuming compute capacity indefinitely. Existing filesystem identity and uncertain compute ownership remain intact; stale node calls and incompatible retained-workspace generations are rejected.

Validation:

  • Existing full Core package/DB tests and fresh independent review passed. Generated OpenAPI/SQLc, client TypeScript, translations, distribution-manifest, helper tests and affected Darwin/Windows compilation checks passed.

  • Real kernel NFSv4.2 + restricted OpenSSH + Core/store qualification on isolated mx containers passed: 64 MiB byte limit, 256-inode limit (252 files plus envelope), sibling/tenant isolation, full-quota deletion, four 64 MiB admissions under a 256 MiB tenant cap, retained-unbounded rejection, configuration-switch accounting, credential redaction, host-key rejection and uncertain-outcome retry.

  • Real XFS crash/replay and refusal cases passed. A frozen block-image backup with matching PostgreSQL restore preserved backed-up inode identities, quotas and retained content; deleted objects remained deleted, post-backup objects disappeared, and new writes still hit their quota. All temporary remote resources were removed.

  • Real native Codex microsandbox qualification on mx1/mx2 passed through the pinned public SDK, integrated Core/node, guest virtiofs and quota NFS: two 1024 MiB workspaces; guest ENOSPC near the hard limit; a public 16 MiB Environment-file copy returned 503 execution_unavailable with a settled rejected receipt and no destination; sibling writes succeeded under pressure; normal next input continued with the same native Session identity; deletion physically removed A while B stayed usable, then B was also deleted. Original Claude Runtime/storage/resources were restored as generation 4 with both nodes ready; isolated quota resources were removed.

  • All 25 PR CI checks passed on the original quota implementation. The final combined head 00643759 includes the focused scheduling correction and admission documentation; all 25 applicable CI jobs and the aggregate check passed.

  • Rebased without conflicts onto b69ff07e; Clarify Session workspace ownership #661 Session workspace ownership documentation and tests are preserved. Focused hosted-environment/API and translation checks rerun after rebase.

  • A real PostgreSQL regression first reproduced one active/reserved slot with zero running compute after quota refusal. The corrected shared flow passed first and retained demand fairness, capacity recovery, exact workspace-reference retry after unknown FS creation, preserved unknown compute ownership, stale-node/generation protection, and concurrent allocation/archive/deletion checks. Execution, deployment, placement and both database packages passed; targeted race checks passed five repetitions, build/vet and a fresh independent review passed.

Qualification scope: the native test used the integrated quota/recovery/performance Core and node with the existing qualified Codex Runtime. The subsequent scheduling-only correction was validated with real PostgreSQL lifecycle tests; it does not change the native filesystem or checkpoint adapters and their successful live gates were not repeated. Local Darwin/Windows quota-package checks were cross-compilation; the earlier standard Linux, macOS and Windows native CI jobs passed, and the final head passed all three platforms again. Session creation can precede asynchronous workspace admission; exhausted aggregate capacity follows the existing input reservation deadline rather than introducing a direct public creation error.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant