· hardware
A Hardware Model of Device Memory Access
DMA is often described as allowing a device to access memory “without going through the CPU.” That description identifies an important benefit, but obscures where the transfer actually goes. A DMA-capable device initiates its own memory transactions, relieving a processor core of the instructions needed to move the payload. Those transactions still pass through the platform’s I/O and memory infrastructure, where they encounter routing, translation, permission checks, and competition from other requesters.
Following that path provides a useful way to understand several mechanisms that often appear together in a virtualization configuration: the IOMMU, Intel VT-d, PCIe Access Control Services, and Linux IOMMU groups. Each addresses a different part of the problem. Their relationship becomes clearer once the device is treated as an independent requester rather than an extension of the CPU core.
What DMA changes
Without DMA, a processor core could transfer data by repeatedly reading a device register and writing the result into RAM. DMA gives another hardware agent responsibility for the transfer. An NVMe controller, NIC, HBA, or GPU can act as a bus master, originating transactions instead of merely responding to requests from the processor.
The driver still sets up the operation. A simple DMA engine might receive a source, destination, length, and start command. Modern devices generally use descriptors and queues in memory: the driver allocates buffers, describes the work, and tells the device where to find it. The device can then read those descriptors and transfer the payload without having a CPU core execute a load and store for each portion.
For a transfer into host DRAM, the conceptual path is:
Driver configures descriptors and queues
|
v
PCIe device
|
DMA transactions
v
PCIe fabric / root complex
|
IOMMU translation and checks
|
v
System interconnect / caches
|
v
Memory controller / DRAM
This diagram describes responsibilities rather than a motherboard layout. The root complex and IOMMU may be integrated, and a coherent cache can service traffic before it reaches DRAM. Some PCIe root ports reside in the processor or SoC; others can reside in a chipset connected to it. In each case, the transfer uses a hardware path distinct from the CPU cores.
The term bus master comes from older shared-bus systems, where arbitration could literally give the CPU or DMA controller control of the same electrical bus. PCIe uses switched, point-to-point links, but the distinction between an initiator and a responder remains useful.
The memory traffic remains
DMA removes CPU instruction execution from the bulk transfer. It does not remove the payload’s memory traffic. Once a device request enters the host memory system, it can compete with requests from processor cores and other devices for interconnect, cache, and DRAM resources.
A NIC sustaining tens of gigabytes per second into host buffers consumes bandwidth even if the processor spends little time copying the packets. Under sufficient load, that traffic can increase the latency observed by a CPU workload. How much reaches DRAM depends on the destination, cache behavior, coherency, and the platform’s handling of I/O traffic; a transfer between devices need not have the same path as a transfer into host RAM.
The memory controller also does not simply serve one requester at a time in arrival order. It can reorder accepted requests to exploit DRAM banks, row locality, and read/write batching. A DMA operation therefore obtains access through an arbitrated system rather than acquiring exclusive use of RAM.
Coherency introduces a separate requirement. A device may modify memory while a CPU has the same lines cached. Coherent platforms provide machinery for maintaining consistent copies; non-coherent platforms require appropriate cache maintenance. In either case, the driver must use the platform’s DMA mapping API and the ordering barriers required by the device protocol. Even coherent memory does not ensure that a device observes descriptor fields in the order software intends. Linux DMA mapping guide
A device needs its own address space
A processor’s virtual address is not ordinarily an address that a PCIe device can use. CPU page tables translate addresses generated by CPU instructions. Device transactions use DMA addresses supplied through the platform’s DMA interfaces, which may differ from the host physical addresses backing the buffers.
An I/O Memory Management Unit, or IOMMU, translates and checks those device accesses. Software can give a device an I/O virtual address (IOVA), while the IOMMU maps that address to a selected host page. For example, three consecutive pages in the device’s address space could have the following mappings:
| Device IOVA | Host physical page |
|---|---|
0x10000000 |
0x824510000 |
0x10001000 |
0x391A20000 |
0x10002000 |
0xF00230000 |
The example shows both purposes of translation. Physically scattered pages can appear contiguous to the device, and software can restrict the device to the pages included in its mappings. The latter matters because an unconstrained bus master could otherwise read or overwrite sensitive host state.
An IOMMU’s presence alone does not establish that restriction. The device needs an appropriate translation domain and permission policy. Broad identity mappings or a bypass configuration can leave it with broad access despite the hardware’s availability.
The IOMMU selects a translation context using requester information. PCIe requests normally identify their origin through a Requester ID associated with a bus, device, and function. Additional mechanisms such as a Process Address Space ID (PASID) can distinguish contexts within a requester, while bridges and aliases can reduce the identity granularity the platform actually sees. The lookup therefore involves both the requester and its IOVA; two devices can use the same numerical address and reach different host pages. Translation caches, commonly called IOTLBs, avoid a full page-table walk for every transaction. Intel VT-d architecture specification
Intel calls its architecture for DMA remapping and related I/O virtualization functions VT-d. AMD’s corresponding architecture is commonly called AMD-Vi or AMD IOMMU. IOMMU is the generic term, rather than another name for Intel’s implementation.
Why virtualization needs DMA remapping
A guest CPU normally translates a guest virtual address to a guest physical address, after which Intel EPT or AMD NPT translates it to the host memory backing the guest. A device assigned directly to the guest cannot use that CPU translation path: its transactions originate at the device.
The hypervisor must arrange a separate IOMMU mapping so that the device’s DMA reaches the host pages assigned to the guest. The guest can then program the device without intentionally granting it access to the host’s entire physical address space. This is one of the mechanisms that makes direct device assignment practical.
Interrupt delivery needs protection as well. A passed-through device produces interrupts, including MSI/MSI-X messages, in addition to reading and writing buffers. VT-d includes interrupt remapping to constrain how those device-generated interrupts are delivered. DMA remapping protects the memory-access path; interrupt remapping addresses interrupt routing. Intel’s description of VT-d functions
Neither feature removes the need to examine the device’s surroundings. The IOMMU can enforce permissions only on transactions that traverse the relevant protection path.
PCIe routing can create another path
Consider two endpoints below the same PCIe switch:
Root complex / host enforcement
|
Root port
|
PCIe switch
/ \
Device A Device B
\_______/
possible local P2P path
The switch may route a peer-to-peer transaction directly between the endpoints. If that transaction never reaches the host’s IOMMU path, the IOMMU cannot reject it. Knowing that the IOMMU distinguishes Device A from Device B is therefore insufficient to prove that the devices are isolated.
PCIe Access Control Services (ACS) supplies controls for source validation and transaction routing at appropriate ports and functions. Among these are controls that redirect peer requests and completions upstream, allowing the hierarchy to prevent an unsafe local shortcut. The available capabilities and their configuration determine which boundaries the topology can establish.
Upstream redirection should not be equated with a guarantee that every peer-to-peer request undergoes IOMMU translation. Routing beyond a root port or host bridge depends on the platform. Linux’s peer-to-peer DMA documentation distinguishes routes contained within a PCIe hierarchy from those that depend on host-bridge behavior. The security argument must establish both the route and the enforcement it encounters. Linux PCI peer-to-peer DMA documentation
The IOMMU and ACS consequently address complementary problems. The IOMMU translates and restricts requests presented to it. ACS helps constrain routes and source behavior that could otherwise undermine the intended isolation boundary.
What an IOMMU group represents
Linux uses IOMMU groups to expose the smallest collection of devices it can establish as isolated from the rest of the system. That decision considers requester identities, topology, and the isolation capabilities the kernel recognizes. A group is an ownership boundary, rather than a requirement that every group use a unique set of translation tables; different groups can deliberately share a translation domain.
For example, a GPU, NIC, and HBA below a non-isolating bridge may remain in one group even though their PCIe functions have different Requester IDs. The hardware can distinguish their upstream requests while still permitting communication that bypasses the relevant checks. Linux cannot infer safe independent assignment from those IDs alone.
Multifunction devices present a similar problem. A GPU could expose graphics at 41:00.0, audio at 41:00.1, a USB controller at 41:00.2, and a USB-C controller at 41:00.3. Those are separate PCI functions, but internal communication paths can exist within the physical device. The functions’ addresses do not establish that upstream controls can isolate them from one another. Linux VFIO documentation
Moving a card between slots can change its group because the new slot places it behind a different root port, bridge, or switch. The card and IOMMU may be unchanged while the path between them crosses a different set of isolation boundaries. Grouping is therefore evidence about the topology recognized by the running kernel, not simply a property printed on the card’s specification sheet.
What ACS override actually changes
The commonly used pcie_acs_override patch changes the kernel’s assumptions about those boundaries. It can allow Linux to separate devices that would otherwise remain grouped because sufficient isolation has not been demonstrated. It does not add a hardware check or close an existing peer-to-peer route. The original patch discussion explicitly distinguishes assuming isolation from establishing it. ACS override patch discussion
This is an out-of-tree mechanism: the running kernel must contain an implementation of the patch. Merely adding the parameter to a command line does not establish that anything implements it. Nor should this software override be confused with a mechanism that actually configures ACS control registers in hardware. The latter can change routing; an override that changes capability decisions changes what Linux is willing to treat as isolated.
An administrator can accept that weaker assurance when the devices and guests are trusted. The resulting separate groups nevertheless cannot support the same claim of separation in an adversarial environment. Once software uses the artificial split to assign devices independently, shared reset behavior, peer traffic, or other hardware dependencies can also affect correctness.
This distinction limits what can be inferred about stability. Changing a grouping decision does not directly alter electrical link training or lane negotiation. That makes a direct electrical explanation unlikely, but does not exclude the patched kernel or the resulting assignment configuration from a hang or reset. Diagnosing those failures requires evidence from the particular system.
The complete device-access model is thus a chain of responsibilities: software describes the transfer, the device originates transactions, PCIe routes them, the IOMMU checks those that reach its translation path, and the memory system services accepted traffic. Linux groups describe the ownership boundary it can establish from that chain. A successful device assignment shows that the configuration operates; establishing isolation additionally requires that no usable path evades the intended checks.