Skip to content

filesystem: map host paths from any volume, each with its own access - #459

Merged
Pedro Henrique Penna (ppenna) merged 4 commits into
devfrom
agents/fix-issue-281-filesystem-mappings
Oct 10, 2026
Merged

Pedro Henrique Penna (ppenna) merged 4 commits into
devfrom
agents/fix-issue-281-filesystem-mappings

Conversation

@ppenna

@ppenna Pedro Henrique Penna (ppenna) commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #281.

This PR lets an NVX workload map any number of host directories and files, on any volume, each read-only or read-write and enforced by the host. It uses the aggregate virtio-fs device from nanvix/openvmm#127, which serves several host directories through the microVM's single virtio-fs slot.

Why

MXC drives the aci_edge_sandboxes OpenVMM backend. That backend exported one virtio-fs share: the deepest directory that contained every mapped path. As a result:

  • Support multiple filesystem mappings with separate access modes #281's layout was refused: C:\projects\app and C:\build-output read-write beside C:\tools read-only. Paths on different volumes were refused too.
  • The export was writable whenever any mapping was, so read-only mappings and unmapped siblings were protected only by the guest kernel.
  • The 1024-byte kernel command line held only about a dozen mappings.

The microVM also had no free interrupt for more virtio-fs devices.

Changes

Commit Change
9f3fb167f Pins OpenVMM 7bf0ee28b, the head of nanvix/openvmm#127. Changes only the gitlink.
55c31c4a5 aci_edge_sandboxes always exports one aggregate. Its children are the outermost mapped directories and the parents of mapped files; a file's parent has a hidden root. Each child is read-only unless it holds a read-write mapping, and --mount-write limits writes to the read-write mappings. The mapping table travels over the control channel in APP_MAPS requests, with only nvx_maps=COUNT on the kernel command line. The guest agent refuses to run anything until every entry is mounted. It retires HOST_MAPPINGS (bit 1) and advertises HOST_MAPPING_TABLE (bit 6), which the backend now requires. Sandbox records move to STATE_FORMAT 3. Mapping a whole volume, or a file directly in a volume's root, is refused.
74bb05291 In run and sandbox, one --mount stays a plain share. Two or more attach as children of --mount-aggregate /run/nvx/shares, and the guest binds each at its target from nvx_share= tokens. Each child's name is its index and a digest of its guest target, so a snapshot pins every share's target, and a restore with a different target fails before the guest runs. With one share or several, a share's root may be its own --mount-deny path when --mount-allow exposes paths inside it. The two-share cap is replaced by the kernel command-line budget.
d5888a607 test-microvm's filesystem-shares scenario shares three directories through the aggregate, including a hidden root and a name with spaces, and covers snapshot capture and restore. It checks that a hard link from one child into a non-writable directory fails with EROFS, that one into another child's writable directory fails with EXDEV, and that a restore that renames a child fails before the guest runs. denied-filesystem-paths now expects OpenVMM's new message when it refuses to deny the share's root without --mount-allow paths.

Updated docs: aci_edge_sandboxes/README.md, doc/run.md, doc/usage.md, doc/ci.md, doc/design/machine-and-device-abi.md, doc/design/sandbox-filesystem-and-agent-architecture.md, and doc/design/cold-boot.md.

Limits and behavior changes

  • The host enforces every mapping's access except a read-only path nested inside a read-write mapping, which the guest kernel still enforces.
  • Limits:
    • At most 256 shared directories.
    • At most 128 hidden, exposed, or writable paths in each, so more than 128 files from one directory must be mapped as that directory.
    • At most 4096 mappings in total.
    • The OpenVMM command line must fit in a Windows process command line.
  • This change is not backward compatible:
    • Guest images without HOST_MAPPING_TABLE fail start with backend_unavailable, and older crates refuse the new image.
    • Format-2 sandbox records are rejected.
    • A repeated OpenVMM --mount and two-slot snapshots no longer work.
  • Bundled builds that download the release pinned in aci_edge_sandboxes/artifacts.json can't start sandboxes until it names a release with the new guest agent.

OpenVMM promotion

Validation

Run locally on Windows/WHP at this head:

  • python scripts\nvx.py verify.
  • aci_edge_sandboxes: fmt, clippy (all features, and --lib --no-default-features), tests with both feature sets, doc, the bundling test, and the MSRV 1.89 check.
  • ruff, pyright, shellcheck, and shfmt pass.
  • validate-nvx unit tests:
    • Linux (WSL): 872/872, including the agent's C tests.
    • Windows: 871/872. test_materialize_kernel_provenance_inputs_bypasses_mutable_index fails with or without this change because of this machine's git signing configuration.
  • On WHP, with OpenVMM, the guest kernel, and the Alpine initramfs rebuilt from this head:
    • The full test-microvm suite: 39 scenarios, including IPv6 egress.
    • test-aci-edge-sandboxes: 14/14.
    • Two-share nvx.py sandbox runs with a hidden root, one-shot and managed.
    • A manual restore through nvx.py run: the same targets resume the guest, and a changed target fails with OpenVMM's snapshot-contract error before the guest runs.
    • Before the review fixes, a manual run of Support multiple filesystem mappings with separate access modes #281's layout directly under C:\: app and build-output read-write, tools read-only, and the unmapped private sibling unreachable.

Still pending in CI: KVM and MSHV, the Ubuntu and Azure Linux initramfs builds, and the second-volume test on runners that have a second volume.

NVX head submitted to CI: d5888a6072595ec6372674826c3c08c877a66e78.

Follow-ups

  • Repin aci_edge_sandboxes/artifacts.json once a dev release ships the new guest agent.
  • Update the MXC integration doc (feat: NVX integration design doc mxc#1448).
  • Bring the separately supplied native aci_edge_agent backend to the same rules.
  • Remove the retired slot's MMIO base from kernel/patches/0003-microvm-virtio-mmio-shared-status.patch.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Aggregate restores ignore requested guest targets, and multi-share validation rejects the documented hidden-root policy.

2 open findings
What changed in this PR

Adds aggregate virtio-fs support so workloads can map multiple host paths with independently enforced access modes.

Changes:

  • Pins OpenVMM’s aggregate-share implementation.
  • Moves host mapping tables from the kernel command line to the control protocol.
  • Expands multi-share, snapshot, policy, and cross-volume coverage.
File Description
openvmm Pins aggregate virtio-fs support.
aci_edge_sandboxes/​README.md Documents mapping behavior and limits.
aci_edge_sandboxes/​src/​bin/​aci-edge-sandboxes-fake-openvmm.rs Simulates mapping-table delivery.
aci_edge_sandboxes/​src/​openvmm/​filesystem.rs Plans aggregate exports and policies.
aci_edge_sandboxes/​src/​openvmm/​launch.rs Builds and bounds launch arguments.
aci_edge_sandboxes/​src/​openvmm/​mod.rs Sends mappings during startup.
aci_edge_sandboxes/​src/​openvmm/​protocol.rs Defines mapping-table protocol encoding.
aci_edge_sandboxes/​src/​openvmm/​session.rs Transfers mapping records.
aci_edge_sandboxes/​src/​openvmm/​state.rs Advances persisted-state format.
aci_edge_sandboxes/​tests/​openvmm_e2e.rs Tests mappings on real guests.
aci_edge_sandboxes/​tests/​openvmm_fake.rs Tests backend behavior with the fake VMM.
doc/​ci.md Updates filesystem CI coverage.
doc/​design/​cold-boot.md Updates boot-token ordering.
doc/​design/​machine-and-device-abi.md Defines the aggregate device ABI.
doc/​design/​sandbox-filesystem-and-agent-architecture.md Documents aggregate architecture.
doc/​run.md Documents runtime mounting behavior.
doc/​usage.md Updates CLI option reference.
guest/​common/​nvx-hostmount Mounts and unmounts aggregate children.
guest/​common/​nvx-init-agent Binds aggregate children into sandboxes.
guest/​common/​nvx-managed-agent.c Receives and applies mapping tables.
guest/​common/​nvx-virtio-restore-probe Handles aggregate restore probing.
scripts/​nvx.py Emits aggregate launch arguments.
scripts/​nvx_tools/​microvm_test_scripts/​filesystem-shares-snapshot.sh Exercises aggregate snapshot restore.
scripts/​nvx_tools/​microvm_test_scripts/​filesystem-shares.sh Exercises aggregate access controls.
scripts/​nvx_tools/​microvm_tests.py Expands filesystem-share scenarios.
scripts/​nvx_tools/​sandbox.py Models aggregate sandbox shares.
scripts/​test_managed_agent.py Tests mapping-table handling.
scripts/​test_microvm_tests.py Tests aggregate scenario orchestration.
scripts/​test_nvx_tools.py Updates CLI and guest-agent tests.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/nvx.py Outdated
Comment thread scripts/nvx_tools/sandbox.py
Copilot AI balanced review requested due to automatic review settings October 10, 2026 01:11
@ppenna
Pedro Henrique Penna (ppenna) force-pushed the agents/fix-issue-281-filesystem-mappings branch from 7467936 to 13bf70e Compare October 10, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes a security-sensitive cross-language ABI while its required OpenVMM PR remains unmerged and platform CI is still running.

2 open findings

🧠 Review effort: Balanced

7bf0ee28b (fix/aggregate-microvm-mounts-281 in nanvix/openvmm) stacks on
e5bf79776, the IPv6 egress of nanvix/openvmm#126, which dev already pins
and nanvix/openvmm main has merged. It gives a microVM a single
virtio-fs slot, fs:microvm0, and retires the second one, fs:microvm1
with tag microvm1 at 0xd0008000 and IRQ 13. --mount may appear once.
--mount-aggregate GUEST_TARGET with a repeatable --mount-child
NAME,HOST_PATH[,ro|rw] instead shares several host directories through
that slot, as named children of a synthetic read-only root. Each child
has its own mode and its own --mount-deny, --mount-allow, and
--mount-write paths, which are absolute and belong to the child that
contains them. A rename between children fails with EXDEV, and so does a
hard link, unless its destination is read-only, which fails with EROFS
first, as on Linux. An aggregate share is snapshot-capable, a restore
requires the same children, names, and order, and a snapshot that
recorded a second slot no longer restores.

--mount-deny may also name the root of a share or child when
--mount-allow names paths inside it, which leaves that root
traverse-only, and policy path names may contain inner spaces.

run and sandbox still pass a second share as a repeated --mount; the
following commits move them and aci_edge_sandboxes to the aggregate
form.

Part of #281.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…own access

The OpenVMM backend exported one virtio-fs share, the deepest directory
that contained every mapped path, and its guest agent bind-mounted each
mapped path from nvx_map= kernel command-line tokens. Paths on different
volumes, or under nothing narrower than a volume's root, were refused,
as is the layout of #281: C:\projects\app and C:\build-output read-write
beside C:\tools read-only. The export was writable whenever any mapping
was, so read-only mappings and unmapped siblings were protected only by
the guest kernel, and the 1024-byte kernel command line held about a
dozen mappings.

The pinned OpenVMM shares several host directories through its one
virtio-fs device as the children of an aggregate. The backend now always
exports one: its children are the outermost mapped directories and the
parents of the outermost mapped files, named by their index. A child is
read-write only if it holds a read-write mapping, and --mount-write then
limits writes to its read-write mappings, so the host enforces each
mapping's access, except that of a read-only path inside a read-write
mapping. A child that is a file's parent hides its root behind its
mapped paths, so unmapped siblings are not exported. Denied paths are
attributed to the child that contains them, and those already hidden are
dropped. Mapping a whole volume, or a file directly in a volume's root,
whose parent would export that root, is refused. Provision checks
OpenVMM's limits: 256 children; 128 hidden, exposed, and writable paths
in each, so more than 128 files of one directory must be mapped as that
directory; their path syntax, which now admits spaces inside names; and a
command line that fits a Windows process with 1,024 characters to spare.

The mapping table no longer travels on the kernel command line, which
carries only nvx_maps=COUNT. Start sends it in APP_MAPS requests that
each fit one 64 KiB control record, after FEATURES and before the
sandbox counts as started, and aborts the start if the guest cannot
mount an entry. The guest agent accepts the entries in order and only up
to the announced count, checks a whole request before it mounts
anything, and refuses EXEC as mappings-incomplete until every entry is
mounted, so a workload never runs without its host paths. It ignores
nvx_map= tokens, stops advertising HOST_MAPPINGS (bit 1), and advertises
HOST_MAPPING_TABLE (bit 6), which the backend now requires instead, so
images and crates of either generation refuse each other. Sandbox
records move to STATE_FORMAT 3, and records of format 2 fail with the
unsupported-format error.

Unit tests cover the planner, the batching of APP_MAPS, and the size
check; the fake OpenVMM applies the table and enforces the new contract;
the agent's C tests decode tables, refuse workloads before a table is
complete, refuse extra and malformed entries, and ignore stale tokens;
and new end-to-end tests map the layout of #281, files beside hidden
siblings, names with spaces, 50 mappings, and a second volume where one
exists.

Part of #281.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The pinned OpenVMM gives a microVM one virtio-fs slot and refuses a
repeated --mount, so run and sandbox could no longer attach the second
share that they placed in the retired slot, and they never could attach
a third.

One --mount still attaches the slot as before. Two or more now attach
it as an aggregate: --mount-aggregate /run/nvx/shares, with one
--mount-child NAME,HOST_PATH,MODE per share in order, and absolute
policy paths that OpenVMM attributes to the child that contains them.
A share's name is its index, a hyphen, and the first 32 hexadecimal
digits of the SHA-256 digest of its guest target. NVX adds an
nvx_share=NAME,TARGET,MODE token per share to the kernel command line,
and the two-share cap gives way to that command line's 1024-byte
budget, which the shares' tokens and OpenVMM's bootstrap tokens share.
run rejects targets that overlap each other or /run/nvx/shares, and
relative policy paths among several shares. A restore passes the same
aggregate without the tokens, because the restored guest keeps its
binds. OpenVMM does not record the guest targets, but its snapshot
contract records the children's names, so a restore that requests
another target fails before the guest runs instead of resuming binds at
the snapshot's targets. With several shares, sandbox accepts a
--mount-deny of a share's own directory, which OpenVMM accepts when
--mount-allow paths inside it remain, and still requires every other
policy path to lie inside its share. The managed configuration still
lists the shares, so its format is unchanged.

In the guest, nvx-hostmount mounts an aggregate at /run/nvx/shares,
whose root only root can enter, and binds each child at its target with
its mode and nosuid,nodev; with --unmount, it unmounts the binds in
reverse order and then the share. nvx-init-agent checks every child and
creates every target before it mounts anything, mounts the aggregate
outside the container root, binds each child into it, and unmounts the
binds before the aggregate on teardown. Both accept a child name of
ASCII letters, digits, '.', '_', and '-', other than '.' and '..'.
nvx-virtio-restore-probe reads and writes through an aggregate's first
child and unmounts with nvx-hostmount --unmount.

Tests cover the argument translation, the children's names, the tokens
and their budget, a hidden share root, a restore's names, the managed
lifecycle's replay, and the guest's child names, binds, refusals, and
teardown order. The run and usage guides, the machine and device ABI,
which frees IRQ 13 and the 0xd0008000 window, and the sandbox design
describe one slot, the aggregate, and its children's names. A rename
between children fails with EXDEV, and so does a hard link, unless its
destination is read-only, which fails with EROFS first, as on Linux.

Part of #281.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… device

The filesystem-shares scenario attached a read-write workspace and a
read-only tool cache to the two virtio-fs slots, with a repeated --mount
and the microvm1 tag, which the pinned OpenVMM no longer offers.

The scenario now attaches the workspace, the tool cache, and a data
directory as the children of one aggregate, --mount-aggregate
/run/nvx/shares with three --mount-child options, and passes the
nvx_share= tokens with which the guest binds each child at its target.
The data directory exercises a hidden root: --mount-deny names the
directory itself, --mount-allow exposes its public directory and a file
whose name holds a space, and --mount-write makes only that file
writable.

The guest checks that only the 0xd0001000 slot exists, that the
aggregate is mounted at /run/nvx/shares with a root-only root that lists
the three children, and that each bind carries its share's mode. It
checks each share's denied paths and the data directory's hidden root,
writes where the policy allows it, and through /run/nvx/shares, where
guest root bypasses the binds' flags, OpenVMM still rejects every write
to the read-only child with EROFS. A hard link between children fails
with EROFS into the data directory's read-only public directory, which
Linux's linkat reports before a link across filesystems, and with EXDEV
into the writable workspace, and neither link appears. The host
requires the guest's writes and nothing else.

OpenVMM must reject a repeated --mount, --mount beside
--mount-aggregate, overlapping children, and a relative policy path
before boot. A capture with every child resumes the guest's open
journal on restore, and restores that drop a child, swap two, rename
one, or offer a plain --mount fail before the guest runs. Renaming a
child, with the same directory, mode, and order, is enough, which the
names that run and sandbox derive from the guest targets rely on.

The denied-filesystem-paths scenario still requires OpenVMM to refuse,
before boot, a --mount-deny of the share's root without --mount-allow
paths, now with the message that a root can be hidden only to expose
allowed paths inside it.

Part of #281.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The documented OpenVMM pin is no longer the upstream PR head, which also retains unresolved restore-path findings.

2 open findings
1 resolved since last review

🧠 Review effort: Balanced

Comment thread doc/design/machine-and-device-abi.md
@ppenna
Pedro Henrique Penna (ppenna) merged commit 009f05b into dev Oct 10, 2026
32 checks passed
@ppenna
Pedro Henrique Penna (ppenna) deleted the agents/fix-issue-281-filesystem-mappings branch October 10, 2026 04:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support multiple filesystem mappings with separate access modes

2 participants