Probe bubblewrap with the argv a launch runs - #2039
Merged
Merged
Conversation
The availability probe ran a hand-picked subset of flags with no --proc, --dev or --unshare-net. On a host that allows user namespaces but masks /proc (Docker or a Kubernetes pod with seccomp Unconfined and the default masked paths or procMount) that subset passed, so the worker reported confinement while every real launch died in bwrap with "Can't mount proc". RequireConfinement could not refuse such a host, and the run record claimed isolation the launch never got. Under Docker's full defaults nothing changes: the default seccomp profile already blocks the user namespace, so both probes report no-userns. The probe now runs BuildArgs of the default tier's own plan (network severed, running true), so "available" means the launch argv runs here. Only when bwrap starts and refuses that argv does it re-run the old minimal user-namespace argv, to tell no-userns apart from the new mounts-denied reason. The two are told apart by which argv ran, not by parsing bwrap's stderr, because the fixes differ: allow user namespaces versus unmask /proc. A masked-/proc host goes from "every launch fails" to an honest Unconfined record, or a refusal under RequireConfinement. One host loses confinement it had: one that allows user namespaces but denies network namespaces (user.max_net_namespaces=0, or a profile that denies only CLONE_NEWNET). Its network-off launches already died in bwrap; now the probe fails as well, so its Trusted runs, which share the network and did confine, run unconfined, and the record names the wall mounts-denied. That is accepted, because "available" answers for the default tier and no probed deployment posture denies only network namespaces. The refusal and the RequireConfinement docs name the network namespace, so an operator on such a host is not sent to /proc. The reason column has no constraint, so the new value needs no migration. It reaches user-visible text only through the network-posture sentence, which the Room shows as its network stat line and the shared fixture pins.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
BuildArgs(ProbePlan): the default tier's own launch argv, including--proc,--devand--unshare-net, confiningtrue(BubblewrapSandbox.Classify). The old subset passed on a host that allows user namespaces but masks /proc, for example Docker or a Kubernetes pod with seccomp Unconfined and the default masked paths or procMount. Every real launch then died withbwrap: Can't mount proc on /newroot/proc,RequireConfinementcould not refuse the host, and the run record claimed isolation it never got.no-usernsapart from the newSandboxConfinement.ReasonMountsDenied = "mounts-denied"by which argv ran, never by parsing stderr. Under Docker's full defaults nothing changes: the default seccomp profile already blocks the user namespace, so both probes reportno-userns.user.max_net_namespaces=0, or a profile that denies onlyCLONE_NEWNET) now reportsmounts-denied.RequireConfinement.BubblewrapSandbox.cs:69) and theRequireConfinementdocs (appsettings.jsonSandbox._doc,RuntimeSettings.RequireSandboxConfinement) now name the network namespace, so an operator on such a host is not sent to /proc. The trade-off is documented onProbePlan.:stat:networkStatBlock detail). The frontend renders that detail raw (SessionRoomView.tsx:1601), so no frontend change is needed. The shared wording fixturefrontend/src/lib/networkPosture.fixture.jsongains amounts-deniedcase.Test plan
BuildArgs(ProbePlan)and holds--proc,--devand--unshare-net.mounts-denied.RoomProjectorFlowTests, including a newmounts-deniedRoom row, andSandboxConfigurationStartupTests: 142/142.BubblewrapProbeE2ETests: /dev/null over /proc/kcore gives unavailable withmounts-denied; without it, available.Category=Sandbox: 69/69. Host left clean (ip rule, ip netns, nft tables, kcore overmounts, staged dirs).no-userns; seccomp Unconfined givesmounts-denied; seccomp and systempaths Unconfined gives available.max_net_namespaces=0in a child user namespace):mounts-denied, as accepted above.sandbox-isolation.yml: at least 69 tests executed and both probe-arm markers present.