Refresh CMDQV series and SMMUv3 Resolve AUTO property series - #29
Refresh CMDQV series and SMMUv3 Resolve AUTO property series#29kobak2026 wants to merge 83 commits into
Conversation
Use smmuv3_ats_enabled() to determine whether ATS is enabled for the guest when handling an ATC invalidation command, as setting the ATS property value to 'auto' will resolve to ATS being detected as enabled in the ATC invalidation check otherwise. Fixes: f7f5013 ("hw/arm/smmuv3-accel: Add support for ATS") Reported-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-2-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 8addbdb793b7 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Add a statement to clarify that the host SMMUv3 must support HW-accelerated vfio-pci device assignment when setting accel=on. Reported-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Message-id: 20260608174900.2227340-3-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 6328edbe5d72 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Introduce smmuv3_accel_auto_finalise() to resolve properties that are set to 'auto' for accelerated SMMUv3. This helper function allows properties such as ats, ril, ssidsize, and oas support to be resolved from host IOMMU capabilities via IOMMU_GET_HW_INFO. The later commits in this series set the auto_mode flag to true when an accel SMMUv3 property value is explicitly set to 'auto', or if the property value is not set and defaults to auto mode. Setting these property values to 'auto' requires at least one cold-plugged device to retrieve and finalise these properties. If the auto_mode flag is true, register a machine_init_done notifier to verify this requirement and fail boot if it is not met. Hot-plugged devices into an accel SMMUv3-associated bus will re-use the resolved host values from the initial cold-plug. Subsequent patches will make use of this helper to resolve 'auto' to what is reported by host IOMMU capabilities. Suggested-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-4-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 16c547e9cb91 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Allow accelerated SMMUv3 Address Translation Services support property to be derived from host IOMMU capabilities. Derive host values using IOMMU_GET_HW_INFO, retrieving ATS capability from IDR0. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-5-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 7e88c1a9d0f9 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Allow accelerated SMMUv3 Range Invalidation support property to be derived from host IOMMU capabilities. Derive host values using IOMMU_GET_HW_INFO, retrieving RIL capability from IDR3. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-6-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 8e0a2aa777f6 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Allow accelerated SMMUv3 SSID size property to be derived from host IOMMU capabilities. Derive host values using IOMMU_GET_HW_INFO, retrieving SSID size from IDR1. When the auto SSID size is resolved to a non-zero value, PASID capability is advertised to the vIOMMU and accelerated use cases such as Shared Virtual Addressing (SVA) are supported. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-7-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 13b941aaff48 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Allow accelerated SMMUv3 OAS property to be derived from host IOMMU capabilities. Derive host values using IOMMU_GET_HW_INFO, retrieving OAS from IDR5. This keeps the OAS value advertised by the virtual SMMU compatible with the capabilities of the host SMMUv3, so that the intermediate physical addresses (IPA) consumed by host SMMU for stage-2 translation do not exceed the host's max supported IPA size. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-8-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 86ec6639b18c https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
|
@kobak2026 @JiandiAnNVIDIA
|
|
@kobak2026 Given this is QEMU 11.0, it should be added to the hw_compat_10_2[] array. Other than that. Acked-by: Jiandi An jan@nvidia.com |
Set the default value of ATS, RIL, SSIDSIZE, and OAS to auto, in order to match the host IOMMU properties when accel=on. If accel=off and these property values are set to auto, the default property values defined in smmuv3_init_id_regs() for OAS and RIL will remain unchanged, while SSIDSIZE and ATS values will remain initialized at 0. Introduce a new compat for the changed defaults. Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-9-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> [ koba: on nvidia_unstable-11.0 put SMMUv3 default-change compat in hw_compat_10_2[] (active for older machine types); do not introduce hw_compat_11_0[] (post-11.0 only). Per Jiandi PR NVIDIA#29 review. ] (backported from e61f8d1138aa https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Update documentation now that "auto" is supported for accelerated SMMUv3 properties. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-10-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 1b8c0edfee08 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
By turning the inline functions into stubs we can avoid the use of target-specific CONFIG_DEVICES include in a hw/ header, allowing to build the source files including it as common objects. Signed-off-by: Philippe Mathieu-Daudé <philmd@linaro.org> Reviewed-by: Thomas Huth <thuth@redhat.com> Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Message-Id: <20260410200031.18572-3-philmd@linaro.org> (cherry picked from commit 12ce4d9630ae09edec706f6254417e3ec96e08e5) Signed-off-by: Jiandi An <jan@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Acked-by: Nathan Chen <nathanc@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
We used idev and idevc naming for HostIOMMUDeviceIOMMUFD and corresponding class variables which followed the iommufd_device naming in linux kernel. This is mixed with the hiod naming for base type HostIOMMUDevice. Rename HostIOMMUDeviceIOMMUFD* to hiodi* for consistency in QEMU. No functional change intended. Suggested-by: Cédric Le Goater <clg@redhat.com> Signed-off-by: Zhenzhong Duan <zhenzhong.duan@intel.com> Reviewed-by: Philippe Mathieu-Daudé <philmd@linaro.org> Reviewed-by: Cédric Le Goater <clg@redhat.com> Link: https://lore.kernel.org/qemu-devel/20260401080354.1347212-1-zhenzhong.duan@intel.com Signed-off-by: Cédric Le Goater <clg@redhat.com> (cherry picked from commit e2b7310276087351fadd6afbe9e5a3fff9c25f8d) Signed-off-by: Jiandi An <jan@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Acked-by: Nathan Chen <nathanc@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
The updated IOMMUFD uAPI introduces the ability for userspace to request a specific hardware info data type via IOMMU_GET_HW_INFO. Update iommufd_backend_get_device_info() to set IOMMU_HW_INFO_FLAG_INPUT_TYPE when a non-zero type is supplied, and adjust all callers to pass a type value explicitly initialised to zero (IOMMU_HW_INFO_TYPE_DEFAULT) when no specific type is requested. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-2-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 3e79b45e0478 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
The updated IOMMUFD VIOMMU_ALLOC uAPI allows userspace to provide a data buffer when creating a vIOMMU (e.g. for Tegra241 CMDQV). Extend iommufd_backend_alloc_viommu() to pass a user pointer and size to the kernel. Update the caller accordingly. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-3-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from d2f3b3156dbd https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Add a helper to allocate an iommufd backed HW queue for a vIOMMU. While at it, define a struct IOMMUFDHWqueue for use by vendor implementations. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-4-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 41dae463f487 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Add a backend helper to mmap hardware MMIO regions exposed via iommufd for a vIOMMU instance. This allows user space to access HW-accelerated MMIO pages provided by the vIOMMU. The caller is responsible for unmapping the returned region. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-5-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from b391e41e53b5 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
The viommu field is assigned but never used. Callers freeing the veventq already have access to the IOMMUFDViommu object through other references, so this field is redundant. Removing it also simplifies upcoming changes where veventq is allocated based on the viommu id before the IOMMUFDViommu object is created (e.g. vendor CMDQV-based veventq allocation). No functional change. Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-6-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 9e67a50ea91d https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Command Queue Virtualization (CMDQV) is a hardware extension available on certain platforms that allows the SMMUv3 command queue to be virtualized and passed through to a VM, improving performance. For example, NVIDIA Tegra241 implements CMDQV to support virtualization of multiple command queues (VCMDQs). The term CMDQV is used here generically to refer to any platform that provides hardware support to virtualize the SMMUv3 command queue. CMDQV support is a specialization of the IOMMUFD-backed accelerated SMMUv3 path. Introduce an ops interface to factor out CMDQV-specific probe, initialization, and vIOMMU allocation logic from the base implementation. The ops pointer and associated state are stored in the accelerated SMMUv3 state. This provides an extensible design to support future vendor-specific CMDQV implementations. No functional change. Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-7-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 4912111c682f https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Introduce a Tegra241 CMDQV backend that plugs into the SMMUv3 accelerated CMDQV ops interface. This patch wires up the Tegra241 CMDQV backend and provides a stub implementation for CMDQV probe, initialization, vIOMMU allocation and reset handling. Functional CMDQV support is added in follow-up patches. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-8-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> [ koba: drop unrelated FSL_IMX8MM/FSL_IMX8MM_EVK Kconfig context; keep TEGRA241_CMDQV only ] (backported from bc2fd5ce6ca0 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Add support for selecting and initializing a CMDQV backend based on the cmdqv OnOffAuto property. If set to OFF, CMDQV is not used and the default IOMMUFD-backed allocation path is taken. If set to AUTO, QEMU attempts to probe a CMDQV backend during device setup. If probing succeeds, the selected ops are stored in the accelerated SMMUv3 state and used. If probing fails, QEMU silently falls back to the default path. If set to ON, QEMU requires CMDQV support. Probing is performed during setup and failure results in an error. When a CMDQV backend is active, its callbacks are used for vIOMMU allocation, free, and reset handling. Otherwise, the base implementation is used. The current implementation wires up the Tegra241 CMDQV backend through the generic ops interface. Functional CMDQV behaviour is added in subsequent patches. No functional change. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-9-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from f4ec4767ce0c https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Introduce a GPtrArray in VirtMachineState to track all SMMUv3 devices created on the virt machine, and use it when building the IORT table instead of relying on object_child_foreach_recursive() walks of the object tree. This avoids recursive object traversal and provides a foundation for subsequent patches that need direct access to SMMUv3 instances for CMDQV-related handling. No functional change. No bios-tables qtest failures observed. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-10-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 92e340997f86 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Use IOMMU_GET_HW_INFO to query host support for Tegra241 CMDQV. Validate the returned data type, version, and minimum number of vCMDQs and SIDs per Tegra241 CMDQ Virtual Interface(VI). Fail the probe if the host does not meet these requirements. The QEMU model supports one Virtual Interface(VI) per VM with 2 vCMDQs and 16 SIDs per VI, so the probe ensures the host implementation is compatible with these limits. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-11-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 61e7de842085 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Tegra241 CMDQV extends SMMUv3 with support for virtual command queues (VCMDQs) exposed via a CMDQV MMIO region. The CMDQV MMIO space is split into 64KB pages: 0x00000 (CMDQ-V Config page) 0x10000 (CMDQ-V CMDQ Page0) 0x20000 (CMDQ-V CMDQ Page1) 0x30000 (Virtual Interface Page0) 0x40000 (Virtual Interface Page1) This patch wires up the Tegra241 CMDQV init callback and allocates vendor-specific CMDQV state. The state pointer is stored in SMMUv3AccelState for use by subsequent CMDQV operations. The CMDQV MMIO region and a dedicated IRQ line are registered with the SMMUv3 device. The MMIO read/write handlers are currently stubs and will be implemented in later patches. The CMDQV interrupt is edge-triggered and indicates VCMDQ or VINTF error conditions. This patch only registers the IRQ line. Interrupt generation and propagation to the guest will be added in a subsequent patch. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Message-id: 20260609112552.378999-12-skolothumtho@nvidia.com Co-developed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from b110918bce7c https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
SMMUv3 devices with acceleration may enable CMDQV extensions after device realize. In that case, additional MMIO regions and IRQ lines may be registered but not yet mapped to the platform bus. Ensure SMMUv3 device resources are linked to the platform bus during machine_done(). This is safe to do unconditionally since the platform bus helpers skip resources that are already mapped. Reviewed-by: Eric Auger <eric.auger@redhat.com> Reviewed-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-13-skolothumtho@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from dc5726da05a4 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Replace the stub implementation with real vIOMMU allocation for Tegra241 CMDQV. Allocate a matching vEVENTQ together with the vIOMMU, since it is specific to the Tegra241 CMDQV vIOMMU and used to receive CMDQV events. Free both objects on teardown. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Message-id: 20260609112552.378999-14-skolothumtho@nvidia.com Co-developed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from 43763a1ab76f https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
The kernel currently exposes a single VINTF per emulated SMMUv3 instance. IOMMU_VIOMMU_ALLOC returns an mmap offset for the host VINTF Page0 allocated for this SMMU. However, VCMDQs only become bound to that VINTF after IOMMU_HW_QUEUE_ALLOC, so until then the mapped Page0 does not back any real VCMDQ state. mmap the host VINTF Page0 right after IOMMU_VIOMMU_ALLOC, as the host VINTF is already enabled at that point, and unmap it when the vIOMMU is freed. The mapping shares the vIOMMU's lifetime. This prepares the VINTF mapping in advance of subsequent patches that add VCMDQ allocation. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260609112552.378999-15-skolothumtho@nvidia.com Co-developed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from ab15bda3ac13 https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Tegra241 CMDQV exposes control and status registers in the CMDQ-V Config page (offset [0x0, 0x10000)) used to configure virtual command queue allocation and interrupt behavior. Add read/write emulation for the CMDQ-V Config region ([CMDQV_BASE, CMDQV_CMDQ_BASE]), backed by a simple register cache. This includes CONFIG, PARAM, STATUS, VI error and interrupt maps, CMDQ allocation map and the VINTF0 related registers defined in the CMDQ-V Config space. Only VINTF0 is supported; VINTF1-63 are not. Dispatch writes on access size: Introduced writel_mmio for 4-byte and writell_mmio for 8-byte. Reads need no split as the MMIO framework masks the returned value to the access size. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Tested-by: Nicolin Chen <nicolinc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Message-id: 20260609112552.378999-16-skolothumtho@nvidia.com Co-developed-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Peter Maydell <peter.maydell@linaro.org> (cherry-picked from ab9ead11f03d https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com>
Tegra241 CMDQV exposes per-VCMDQ register windows through two MMIO
apertures:
Direct VCMDQ aperture (0x10000/0x20000): VCMDQ Page0/Page1
VINTF logical aperture (0x30000/0x40000): VINTF0 LVCMDQ Page0/Page1
Both apertures are hardware aliases of the same underlying registers:
Page 0 (control/status): CONS_INDX, PROD_INDX, CONFIG, STATUS,
GERROR, GERRORN
Page 1 (base/DRAM): BASE_L/H, CONS_INDX_BASE_DRAM_L/H
The direct aperture Page 0 is programmable at any time so long as
CMDQV_EN is enabled. The VINTF (logical) aperture Page 0 is
programmable only once SW has mapped a VCMDQ to a VINTF; the
"logical" view is local to that VINTF.
Add read emulation for both apertures, backed by a single per-VCMDQ
register cache. VINTF aperture reads are translated to their
equivalent direct-aperture offset and served from the same cached
state.
Per the CMDQV architecture, a VCMDQ must be allocated to a Virtual
Interface before it is used to send commands to the SMMU. Until that
allocation happens, reads return cached register state with no HW
interaction. Subsequent patches wire up IOMMU_HW_QUEUE_ALLOC, mmap
the host VINTF Page 0, and install it into guest MMIO; after that,
Page 0 reads from either aperture are served from the hardware-backed
mmap'd page instead of the cache. Page 1 is also a hardware alias,
but the kernel only exposes mmap for Page 0, so Page 1 reads always
trap to QEMU and are served from cache.
Signed-off-by: Nicolin Chen <nicolinc@nvidia.com>
Reviewed-by: Eric Auger <eric.auger@redhat.com>
Tested-by: Nicolin Chen <nicolinc@nvidia.com>
Tested-by: Eric Auger <eric.auger@redhat.com>
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
Message-id: 20260609112552.378999-17-skolothumtho@nvidia.com
Co-developed-by: Shameer Kolothum <skolothumtho@nvidia.com>
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
Signed-off-by: Peter Maydell <peter.maydell@linaro.org>
(cherry-picked from 0d06b1cc2050 https://gitlab.com/qemu-project/qemu.git)
Signed-off-by: Koba Ko <kobak@nvidia.com>
Filter -Werror=implicit-function-declaration out of the extra CFLAGS passed through QEMU configure. QEMU forwards --extra-cflags to C++ builds as well, and GCC 15 rejects that C-only warning option for C++ when global -Werror is enabled. Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
QEMU 11.0 creates an offline Python venv during configure. Noble does not install python3-wheel through python3-setuptools, so mkvenv cannot satisfy wheel>=0.34.2 unless python3-wheel is an explicit build dependency. Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
The dma_map_file pathway doesn't fit into the kernel WAR that bypasses PFNMAP, resulting dma_map failures. Only the dma_map pathway could work. So retry with that upon a failure. Keep this WAR until the kernel WAR for PFNMAP is lifted. Signed-off-by: Nicolin Chen <nicolinc@nvidia.com> (backported from commit faff92d https://github.com/nvmochs/QEMU/tree/stable101_smmuv3-accel-07212025_egm) [nathanc: Moved changes from hw/vfio/container-base.c to hw/vfio/container.c for 11.0 base] Signed-off-by: Nathan Chen <nathanc@nvidia.com> Acked-by: Ian May <ianm@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
CXL VFIO passthrough needs a stable guest physical address range for device memory (DPA) that falls inside a CFMWS entry the guest discovers from ACPI CEDT. Without a dedicated range in the address map, the HDM decoder has nowhere to point. Add VIRT_HIGH_CXL_MMIO immediately after the second PCIe MMIO window. It gets its own highmem_cxl_mmio flag in VirtMachineState rather than sharing highmem_cxl, so the two slots are independently controllable even though both are currently tied to CXL bridge presence. The base and size flow through GPEXConfig.cxl_mmio to acpi_dsdt_add_gpex(), which carves out a QWord memory descriptor in the first CXL root bridge's _CRS. The CFMWS window is system-wide, so only the first CXL bridge gets the descriptor - subsequent ones would produce duplicate resource claims for the same range. build_crs() already emits the bridge's own 64-bit ranges into crs. The CFMWS window is a separate system-wide range, so only that window is appended as a new QWord descriptor; the bridge ranges are not re-emitted. A warn_report() fires if the CFMWS window overlaps any existing bridge 64-bit range, since that would indicate an address layout conflict. Signed-off-by: Zhi Wang <zhiw@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) [jan: Resolve #include conflict with hw/acpi/acpi_egm_memory.h added by nvidia_unstable-10.1 base; keep both includes] Signed-off-by: Jiandi An <jan@nvidia.com> (backported from 7831a4a https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Resolved the 11.0 include layout by keeping the existing ACPI EGM include and adding the CXL host MMIO plumbing.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Before this patch, pxb-cxl bridges had no _DSM method at all. When the OS called _DSM on a CXL host bridge, ACPI returned an error and the OS defaulted to reassigning resources across suspend/resume. On machines where firmware pre-commits the HDM decoder, that reassignment breaks the DPA mapping. Wire preserve_config through GPEXConfig into build_cxl_osc_method() so pxb-cxl host bridges get a _DSM method that signals the OS to keep resource assignments stable when needed. The _DSM function 5 (preserve firmware PCI configuration) is the mechanism used to convey this. build_pci_host_bridge_dsm_method() is promoted from static to exported so cxl.c can call it without duplicating the AML. The x86 build_cxl_osc_method() call site passes false since x86 does not use firmware-committed HDM decoders. build_cxl_osc_method is renamed to acpi_dsdt_add_cxl_host_bridge_methods The function now appends both the CXL _OSC method and the _DSM method, so its old name is misleading. Renamed it to match the pxb-pcie analogue acpi_dsdt_add_host_bridge_methods(), making the two root bridge code paths symmetric. No AML change. Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) Signed-off-by: Jiandi An <jan@nvidia.com> (cherry-picked from a8cf3b4 https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
… passthrough Sync the VFIO UAPI additions from the kernel CXL Type-2 passthrough series. VFIO_DEVICE_FLAGS_CXL (bit 9) marks a device as CXL Type-2 and guarantees the capability chain includes a vfio_device_info_cap_cxl entry (cap id 6). That capability carries the BAR index holding the CXL component registers, flags for firmware-committed and cache-capable devices, the byte offset to the HDM Decoder Capability block within that BAR, and region indices for both the DPA memory region and the Component Register shadow. Two new region subtypes: VFIO_REGION_SUBTYPE_CXL (1): mmappable DPA memory VFIO_REGION_SUBTYPE_CXL_COMP_REGS (2): HDM decoder shadow, r/w only Note: UAPI headers are normally kept in sync via scripts/update-linux-headers.sh once upstream kernel changes merge. This patch manually adds the CXL Type-2 additions as a temporary measure to unblock QEMU development. It should be dropped and replaced with a proper header sync once the kernel series is accepted. Signed-off-by: Zhi Wang <zhiw@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) Signed-off-by: Jiandi An <jan@nvidia.com> (cherry-picked from 076f74c https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…or custom region ops vfio_region_setup() always initializes the region MemoryRegion with vfio_region_ops. CXL needs custom pread/pwrite ops for the Component Register shadow region. Add vfio_region_setup_with_ops() which accepts a const MemoryRegionOps * parameter. When non-NULL it is passed to memory_region_init_io(); when NULL the existing vfio_region_ops is used. vfio_region_setup() is retained unchanged as a thin wrapper for all existing callers. Signed-off-by: Zhi Wang <zhiw@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) Signed-off-by: Jiandi An <jan@nvidia.com> (backported from 37cad16 https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Kept the custom region-ops helper, addressed the review finding by retaining sparse-mmap unwind, and preserved vfio_device_get_region_info() borrowed ownership instead of adding unsafe g_autofree cleanup.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…egion setup When VFIO_DEVICE_FLAGS_CXL is set, the kernel has identified a CXL Type-2 device and populated the capability chain with a vfio_device_info_cap_cxl entry. Read that entry to locate the DPA and CXL Component Register shadow regions, then call vfio_region_setup() for each. DPA covers the device's host-managed memory and is faulted in lazily by the VMM. The CXL Component Register shadow gives the VMM access to the HDM Decoder Capability block so it can intercept decoder commits without touching the hardware register page directly. vfio_cxl_derive_hdm_info() walks the CXL Capability Array inside the Component Register shadow to find the HDM Decoder capability (ID 0x5) and extracts hdm_decoder_offset and hdm_count. All reads use le32_to_cpu() since the capability array is little-endian per the CXL spec. Dword 0 is the array header; capability entries start at dword 1, which is why the loop begins at i = 1. CXL register constants are defined here using names that mirror <linux/cxl.h> to make cross-referencing straightforward. Add the VFIOCXL struct embedded in VFIOPCIDevice. Signed-off-by: Zhi Wang <zhiw@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) [jan: Resolve include path conflict (nvidia_unstable-10.1 uses hw/hw.h and hw/iommu.h not yet renamed to hw/core/); add VMChangeStateEntry vmstate field from base; Reject non-firmware-committed CXL devices early in vfio_cxl_setup(); add machine_done notifier cleanup in vfio_pci_put_device() to prevent use-after-free on teardown; Fix double-errp in CXL DPA and COMP_REGS setup; use error_prepend instead of error_setg since vfio_region_setup already sets errp] Signed-off-by: Jiandi An <jan@nvidia.com> (backported from e312bc7 https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Resolved 11.0 include and VFIOPCIDevice state layout, kept only VFIOCXL cxl, preserved firmware-committed/error propagation fixes, and kept DPA/COMP_REGS region-info ownership borrowed.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…COMP_REGS overlay
The CXL Component Register BAR contains two types of ranges that need
different handling:
- Accelerator register windows: passed through as direct hardware
mmaps for performance. The kernel reports the real BAR size and
lists mmappable windows via VFIO_REGION_INFO_CAP_SPARSE_MMAP,
excluding the HDM Decoder Capability block. vfio_region_mmap()
creates hardware-backed sub-regions for each sparse area.
- HDM Decoder Capability block: guest accesses must go through
emulated ops so QEMU can observe and program decoder state. The
kernel blocks direct mmap of this range.
vfio_bar_register(): after the normal mmap path, overlay the COMP_REGS
emulation region at hdm_regs_offset with priority 1. In QEMU's
MemoryRegion model, overlapping subregions are resolved by priority;
the default is 0. Priority 1 ensures guest accesses to the HDM range
always dispatch through the emulated COMP_REGS ops regardless of any
hardware-backed sub-region at a neighbouring offset.
vfio_pci_bars_exit(): remove the COMP_REGS overlay before the normal
BAR teardown path.
Signed-off-by: Zhi Wang <zhiw@nvidia.com>
Signed-off-by: Manish Honap <mhonap@nvidia.com>
(backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/)
[jan: Resolve pci_register_bar conflict; nvidia_unstable-10.1 uses &vdev->pdev not local pdev variable]
Signed-off-by: Jiandi An <jan@nvidia.com>
(backported from ae09b90 https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21)
[kobak: Resolved the 11.0 BAR registration shape while preserving the COMP_REGS overlay.]
Signed-off-by: Koba Ko <kobak@nvidia.com>
Acked-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
… for firmware-committed devices setup_locked_hdm() runs as a machine_done notifier after all devices have been realized. It programs HDM decoder 0 with the CFMWS base address so the guest can fault into device memory from the first instruction. The notifier is only registered when the kernel reports the device as firmware-committed (VFIO_CXL_CAP_FIRMWARE_COMMITTED). The host is responsible for HDM decoder programming; the guest has no mechanism to remap host physical address mappings. The function uses cxl_fmws_base captured by cxl_fmws_set_memmap() during machine memory-map init. If no CFMWS base is available, it warns and returns without programming anything. If COMMIT_LOCK is set in decoder 0 CTRL at machine_done time (left-over from a prior FLR?), it is cleared before writing BASE so the subsequent write is not blocked. COMMIT_LOCK is re-set after programming so the hardware enforces the committed base. read_region() return is checked; failure aborts programming rather than leaving ctrl uninitialized. All write_region() failures are propagated. The function exits cleanly rather than leaving the decoder half-programmed. Add cxl_fmws_base as a hwaddr global in cxl-host.c (and a stub in cxl-host-stubs.c). It is set once by cxl_fmws_set_memmap() and read later at machine_done time. Signed-off-by: Zhi Wang <zhiw@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) [jan: Set machine_done_registered flag after motifier registration for safe cleanup] Signed-off-by: Jiandi An <jan@nvidia.com> (backported from c091724 https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Preserved safe machine_done cleanup and the explicit VFIO-CXL DPA overlap over CFMWS. CFMWS count/size validation is kept in the later FMWS placement fix so this commit remains bisectable.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…y bus The SMMUv3 primary bus check only accepted pxb-pcie as a valid root. pxb-cxl uses the same PCIe-compatible bus implementation; reject it and CXL devices behind it cannot reach the IOMMU. Extend the check to also accept CXL buses so SMMUv3 translation applies to passthrough CXL devices. Update the comment above the check to mention pxb-cxl alongside pxb-pcie. Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com> Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) Signed-off-by: Jiandi An <jan@nvidia.com> (backported from 1e4510b https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Adapted the pxb-cxl primary-bus check to the 11.0 SMMUv3 bus matching code.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…-device regions vfio_container_region_add() attempts an IOMMU DMA mapping for every RAM section that enters the guest address space. For VFIO mmap-backed regions (PCI BAR windows, CXL.mem regions), this mapping always fails: the backing VMAs carry VM_IO | VM_PFNMAP flags and pin_user_pages() refuses to pin VM_IO pages, so IOMMU_IOAS_MAP returns -EFAULT. CPU access to these regions goes through KVM Stage-2 page faults independently of the SMMU/IOMMU, so no IOMMU entry is required for correct operation. Add an early return for RAM-device sections owned by a VFIO device. vfio_get_vfio_device(memory_region_owner(section->mr)) returns non-NULL for any mmap subregion created by vfio_region_mmap(), since memory_region_init_ram_device_ptr() propagates the VFIOPCIDevice owner from the containing region. Matching on ownership covers both normal PCI BAR windows and CXL.mem regions uniformly; non-VFIO RAM-device regions such as NVDIMMs are unaffected and continue through the normal mapping path. Signed-off-by: Manish Honap <mhonap@nvidia.com> (backported from https://lore.kernel.org/all/20260427181235.3003865-1-mhonap@nvidia.com/) Signed-off-by: Jiandi An <jan@nvidia.com> (backported from 695286c https://github.com/JiandiAnNVIDIA/QEMU/tree/vfio-cxl-2026-05-21) [kobak: Adapted the VFIO-owned RAM-device skip to the 11.0 listener ownership path.] Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Validation found that the ARM CXL highmem state was set after the compact highmem layout decision. That left the CXL MMIO window and ACPI exposure vulnerable to disagreeing with the actual memory map. The same check found cxl_fmws_base, cxl_fmws_size, and the placed-window count could be treated as valid before a CFMWS had actually been placed. Set the ARM CXL highmem flags before the compact highmem layout is computed, and only advertise the CXL MMIO window through ACPI when that window remains enabled. Only publish cxl_fmws_base after a CFMWS is actually placed, record the first placed CFMWS size and placed-window count, skip unplaced CFMWS windows when building CEDT/MMIO mappings, and reject multiple VFIO-CXL DPA mappings while the implementation still has a single global fallback window. Fixes: 5b3ffc5 ("NVIDIA: VR: SAUCE: hw/arm/virt: Add CXL FMWS PA window for device memory") Fixes: 1e536ed23d ("NVIDIA: VR: SAUCE: hw/vfio+cxl: Program HDM decoder 0 at machine_done for firmware-committed devices") Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
The VFIO-CXL fallback path provides one global CFMWS window. A
second firmware-committed VFIO-CXL device could otherwise proceed far
enough for the later machine_done path to discover the conflict, which
is too late for clean realize-time error propagation.
Reject a second firmware-committed VFIO-CXL device during setup while
only one global CFMWS fallback window is available. Reserve the window
in the error-propagating setup path and release it during device teardown,
instead of letting machine_done warn after realization. Keep
the CFMWS placed-window count and first-window size populated so
machine_done can validate the DPA mapping against the CFMWS window.
Fixes: 1e536ed23d ("NVIDIA: VR: SAUCE: hw/vfio+cxl: Program HDM decoder 0 at machine_done for firmware-committed devices")
Signed-off-by: Koba Ko <kobak@nvidia.com>
Acked-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
The duplicate-FMWS check still ran after BAR and CXL region population.
That meant duplicate users of the single global CFMWS fallback could fail
through a deeper setup path after state had already been populated.
Reserve the single global CFMWS fallback immediately after VFIO attach,
before BAR and CXL regions are populated. This makes duplicate VFIO-CXL
users fail through the shallow realize error path without leaking setup
state. Preserve the placed-window count and first-window size so
machine_done can validate the DPA mapping against the selected CFMWS.
Fixes: 1e536ed23d ("NVIDIA: VR: SAUCE: hw/vfio+cxl: Program HDM decoder 0 at machine_done for firmware-committed devices")
Signed-off-by: Koba Ko <kobak@nvidia.com>
Acked-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Guest log before this fix: pci 0000:34:00.0: bridge window [mem size 0xc000000000 64bit pref]: can't assign; no space pci 0000:35:00.0: BAR 2 [mem size 0x8000000000 64bit pref]: failed to assign pci 0000:35:00.0: BAR 0 [mem size 0x10000000 64bit pref]: failed to assign pci 0000:35:00.0: BAR 4 [mem size 0x02000000 64bit pref]: failed to assign The virt highmem-mmio-size property only resized the PCIe high MMIO window. Keep the CXL high MMIO window in sync too, so large VFIO-CXL Type-2 BARs have enough guest bridge aperture when the machine is booted with a larger highmem-mmio-size. After rebuilding QEMU and booting with highmem-mmio-size=1T, the guest assigned the endpoint memory BARs: pci 0000:35:00.0: BAR 2 [mem 0x20000000000-0x27fffffffff 64bit pref]: assigned pci 0000:35:00.0: BAR 0 [mem 0x28000000000-0x2800fffffff 64bit pref]: assigned pci 0000:35:00.0: BAR 4 [mem 0x28010000000-0x28011ffffff 64bit pref]: assigned Fixes: 5b3ffc5 ("NVIDIA: VR: SAUCE: hw/arm/virt: Add CXL FMWS PA window for device memory") Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Current workflow actions are deprecated, this updates them to newer versions. Signed-off-by: Morgan Hunter <morgan.hunter@canonical.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
…wned RAM-device regions" This reverts commit be3d015. The commit added an early return in vfio_container_region_add() for any RAM-device section owned by a VFIO device (vfio_get_vfio_device(memory_region_owner(section->mr)) != NULL), skipping vfio_container_dma_map() for it. In practice that excludes every VFIO mmap subregion from the IOMMU IOAS (the SMMU Stage-2 page tables): PCI BAR windows, and the CXL.mem coherent device memory of a CXL Type-2 device. The commit was added becuase for earlier testing during device boot, I was seeing these errors in the QEMU log: qemu-system-aarch64: warning: IOMMU_IOAS_MAP failed: Bad address, PCI BAR? While testing with kernel 6.17.9, I only had tested with the CXL mode and only CUDA tests. This was a testing miss on my side. During testing, I had missed to test UVM ATS tests and verification of PCI-E mode with my QEMU patches. Earlier during debugging I misdiagnosed the issue to be the mapping was always failing as the backing VMAs are VM_IO | VM_PFNMAP, pin_user_pages() refused VM_IO pages, and IOMMU_IOAS_MAP therefore returns -EFAULT resulting in failure. For overcoming this failure, I had added the mapping skip part. This was not correct and the skip seems to break CUDA UVM ATS on CXL Type-2 passthrough. The bug caused by mapping skip ------------------------------ In accelerated nested SMMUv3 mode the GPU translates shared virtual addresses through the hardware SMMU (Stage-1 is the guest page tables, Stage-2 is the host iommufd tables). When CUDA UVM migrates a managed buffer into the device's coherent memory, that page's guest-physical address lies in the CXL HDM window. The GPU reaches it with an ATS request, and to answer that request the SMMU must complete the Stage-1 and Stage-2 walk. With the HDM region skipped there is no Stage-2 entry, so the translation faults. The GPU posts a replayable fault, UVM services it and replays, the access faults again, and the GPU spins in a fault/service/replay livelock. The guest test hangs, and on cancel it reports "Xid 31 ... FAULT_PTE ACCESS_TYPE_VIRT_WRITE". Non-ATS workloads and PCI-E-mode (nvgrace-gpu) passthrough are unaffected, because they never reach the unmapped path. IOMMU_IOAS_MAP is now succeeding for VM_PFNMAP GPU device memory on kernel 6.17.13 because of the host-side workaround 68a70c30f8ce ("NVIDIA: SAUCE: WAR: iommufd/pages: Bypass PFNMAP") That patch teaches iommufd's pfn_reader to handle VM_PFNMAP VMAs. Instead of pin_user_pages() (which does refuse VM_IO), it calls follow_fault_pfn(), which uses follow_pfnmap_start(), faults the lazily inserted PFN in through fixup_user_fault(), and takes the raw struct-page-less PFN. d814's premise was therefore only true on a kernel without this workaround. With it present, IOMMU_IOAS_MAP for the HDM region returns success (0), as the QEMU trace shows: iommufd_backend_map_dma ... iova=0x80000000000 size=0x2330000000 ... readonly=0 (0) I will check in next series of vfio-cxl QEMU support series if any additional fix is required for this. Signed-off-by: Manish Honap <mhonap@nvidia.com> Acked-by: Shameer Kolothum <skolothumtho@nvidia.com> Acked-by: Matthew R. Ochs <mochs@nvidia.com> Signed-off-by: Matthew R. Ochs <mochs@nvidia.com> (cherry picked from commit c6d9738 nvidia_unstable-10.1) [mochs: Updated reverted commit SHA to match commit on this branch] Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
Signed-off-by: Matthew R. Ochs <mochs@nvidia.com>
e2fe530 to
a34bfda
Compare
|
@JiandiAnNVIDIA @NathanChenNVIDIA Addressed the Fix
Why configure/ninja still passed before the fix Verify Merge note Thanks for the review / Ack. |
|
LGTM, thanks Koba!
|
Set the default value of ATS, RIL, SSIDSIZE, and OAS to auto, in order to match the host IOMMU properties when accel=on. If accel=off and these property values are set to auto, the default property values defined in smmuv3_init_id_regs() for OAS and RIL will remain unchanged, while SSIDSIZE and ATS values will remain initialized at 0. Introduce a new compat for the changed defaults. Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-9-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> [ koba: on nvidia_unstable-11.0 put SMMUv3 default-change compat in hw_compat_10_2[] (active for older machine types); do not introduce hw_compat_11_0[] (post-11.0 only). Per Jiandi PR #29 review. ] (backported from e61f8d1138aa https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Jiandi An <jan@nvidia.com> Acked-by: Nathan Chen <nathanc@nvidia.com> Signed-off-by: Jiandi An <jan@nvidia.com>
54a7b96 to
5752c6c
Compare
|
Merged |
Set the default value of ATS, RIL, SSIDSIZE, and OAS to auto, in order to match the host IOMMU properties when accel=on. If accel=off and these property values are set to auto, the default property values defined in smmuv3_init_id_regs() for OAS and RIL will remain unchanged, while SSIDSIZE and ATS values will remain initialized at 0. Introduce a new compat for the changed defaults. Reviewed-by: Shameer Kolothum <skolothumtho@nvidia.com> Reviewed-by: Eric Auger <eric.auger@redhat.com> Signed-off-by: Nathan Chen <nathanc@nvidia.com> Tested-by: Eric Auger <eric.auger@redhat.com> Message-id: 20260608174900.2227340-9-nathanc@nvidia.com Signed-off-by: Peter Maydell <peter.maydell@linaro.org> [ koba: on nvidia_unstable-11.0 put SMMUv3 default-change compat in hw_compat_10_2[] (active for older machine types); do not introduce hw_compat_11_0[] (post-11.0 only). Per Jiandi PR #29 review. ] (backported from e61f8d1138aa https://gitlab.com/qemu-project/qemu.git) Signed-off-by: Koba Ko <kobak@nvidia.com> Acked-by: Jiandi An <jan@nvidia.com> Acked-by: Nathan Chen <nathanc@nvidia.com> Signed-off-by: Jiandi An <jan@nvidia.com>
Summary
This PR replaces two in-tree NVIDIA SAUCE patch series on
nvidia_unstable-11.0with their merged upstreamqemu-project/qemuequivalents, preserving every other commit unchanged (EGM, packaging, VFIO-CXL, WARs, reverts).Methodology
b358d8a11b(smmuv3: Have smmuv3_accel_init() take an Error parameter*)qemu-project/qemu2e2530239c05)Commit message policy
(cherry-picked from …)or(backported from …)provenance URL +Signed-off-by: Koba Ko <kobak@nvidia.com>last (38 cherry-picked, 2 backported)Verification
=(patch + message identical vs backup)Config confirmed:
CONFIG_ARM_SMMUV3=y,CONFIG_ARM_SMMUV3_ACCEL=y,CONFIG_TEGRA241_CMDQV=y.Note on commit SHAs
Unrelated downstream commits (EGM, packaging, VFIO-CXL, WARs, reverts) receive new SHAs because parent hashes change when AUTO/CMDQV are swapped — but patches and messages remain byte-identical (
range-diff =,patch-idmatch). This is expected, not a defect.Merge plan
Jiandi requested a single PR to
nvidia_unstable-11.0. Manish's PR #28 will be merged on top separately after this lands.Refreshed series
https://gitlab.com/qemu-project/qemu.git/qemu-upstream/masterhttps://gitlab.com/qemu-project/qemu.git/qemu-upstream/masterTest plan
git range-diff nvidia_unstable-11.0-backup...HEADlego-cg1-qs-207.ipp2a2.colossus.nvidia.com)