diff --git a/src/runtime-rs/crates/hypervisor/src/qemu/ARCHITECTURE.md b/src/runtime-rs/crates/hypervisor/src/qemu/ARCHITECTURE.md new file mode 100644 index 0000000000..d1e4fda2e4 --- /dev/null +++ b/src/runtime-rs/crates/hypervisor/src/qemu/ARCHITECTURE.md @@ -0,0 +1,651 @@ +# QEMU Command-Line Architecture + +> **Status:** Migration in progress — see the [phase tracker](#migration-phases) below. +> **Issues:** [#12187](https://github.com/kata-containers/kata-containers/issues/12187) (refactor), +> [#12125](https://github.com/kata-containers/kata-containers/issues/12125) (NUMA / hugepages) + +> **Living document:** This file is updated with every PR that lands a migration +> phase. It intentionally contains open questions and decision notes while the +> migration is in progress. Once Phase 6 completes, all such notes will be +> removed and the document will serve as the definitive, stable reference for the +> QEMU command-line architecture — no TODOs, no decision stubs, just the design +> as implemented. + +## Overview + +This document describes the target machine-centric architecture for the QEMU +command-line generator in `runtime-rs`, why the current design needs changing, +and how each migration phase moves the code toward the target. + +The [Grace Platform Configurations](#grace-platform-configurations) section +enumerates 7 concrete command-line topologies derived from tested production +deployments. Each becomes a golden test fixture; the architecture must be +able to generate every one of them exactly from typed `Platform` inputs. + +--- + +## Current State (the Problem) + +All QEMU command-line construction lives in `cmdline_generator.rs` (~3 700 lines). +The design is inspired by `govmm` and works well for simple cases, but has +accumulated three structural problems: + +### 1. Architecture decisions scattered across devices + +Whether a virtio device needs `bus=pcie.0` depends on the machine layout +(Q35 with `pxb-pcie`, `virt` with extra root ports, NUMA enabled, …), yet every +device struct encodes that decision itself via `#[cfg(target_arch = ...)]`: + +```rust +// Only valid for Q35 or VIRT machine types aka x86 or aarch64 +#[cfg(any(target_arch = "x86_64", target_arch = "aarch64"))] +params.push("bus=pcie.0".to_owned()); +``` + +This pattern is duplicated across `DeviceVhostUserFs`, `DeviceVirtioBlk`, +`VhostVsock`, `DeviceVirtioNet`, `DeviceVirtioSerial`, `DeviceVirtconsole`, +`DeviceRng`, `DeviceIntelIommu`, and `DeviceVirtioScsi`. + +### 2. Stringly-typed flat `Machine` struct + +```rust +struct Machine { + r#type: String, // "q35", "virt", "s390-ccw-virtio", … + accel: String, + options: String, // raw accelerator string from config + nvdimm: bool, + kernel_irqchip: Option, + confidential_guest_support: String, + memory_backend: Option, +} +``` + +All machine types share one struct. `kernel_irqchip` is meaningless on `virt` +or `s390x`; `gic-version` is meaningful only on `virt`. There is no type-level +enforcement. + +### 3. No hugepages / NUMA-aware memory backends + +`runtime-rs` does not support `memory-backend-file` with `hugepages` today +(tracked in [#12125](https://github.com/kata-containers/kata-containers/issues/12125)). +`MemoryBackendFile` exists in `cmdline_generator.rs` but is only wired for +`/dev/shm` (virtiofs shared memory) and nvdimm paths — never for the +`/hugepages` mount needed to back the entire guest with huge pages. + +--- + +## Target Architecture + +The refactor replaces the flat struct with a **machine-centric** model. All +topology decisions (PCIe bus assignment, IOMMU wiring, NUMA layout, memory +backends) are made when constructing the `Machine` and `Platform`. Devices +receive a pre-resolved bus string and typed references to shared objects; +they no longer branch on architecture. + +### Module layout (target) + +``` +src/runtime-rs/crates/hypervisor/src/qemu/ +├── ARCHITECTURE.md ← this file +├── cmdline_generator.rs ← legacy; shrinks as phases complete +├── inner.rs +├── mod.rs +├── qmp.rs +└── machine/ ← new; introduced in Phase 0 + ├── mod.rs (Machine enum + Platform + PciTopology) + ├── q35.rs + ├── virt.rs + ├── pseries.rs + ├── s390x.rs + └── probe.rs (PlatformProbe trait + HostTopology) +``` + +### Core types + +#### `Machine` — per-machine-type structs + +```rust +pub enum Machine { + Q35(machine::Q35), + Virt(machine::Virt), + Pseries(machine::Pseries), + S390xCcwVirtio(machine::S390xCcwVirtio), +} + +// Q35: intel-iommu is a global singleton on the machine, not per-bus. +pub struct Q35 { + pub base: BaseMachine, + pub kernel_irqchip: Option, + /// Global IOMMU device for Q35. Emitted as a top-level -device intel-iommu, + /// not attached to any pxb-pcie. Contrast with SmmuV3 on PciRootComplex. + pub intel_iommu: Option, + pub runtime: RuntimeFeatures, +} + +pub struct IntelIommuConfig { + pub intremap: bool, + pub caching_mode: bool, +} + +// Virt: arm-smmuv3 is per-bus; it lives on PciRootComplex, not here. +pub struct Virt { + pub base: BaseMachine, + pub gic_version: Option, + pub ras: bool, + /// Required for Grace GPU passthrough; must be a power of 2. + /// 4T for GH200/GB200 (≤4 GPUs), 8T for GB300 NVL72 (4 GPUs). + pub highmem_mmio_size: Option, + pub runtime: RuntimeFeatures, +} +``` + +`kernel_irqchip` and `intel_iommu` live exclusively on `Q35`; `gic_version` and +`highmem_mmio_size` live exclusively on `Virt`. The compiler enforces this — +no runtime guards needed. + +#### `PciTopology` — bus resolution moved here + +```rust +pub struct PciTopology { + pub default_bus: Option, // "pcie.0" when NUMA / multi-RC is active + pub roots: Vec, +} + +pub struct PciRootComplex { + pub id: String, // "pcie.N" + pub bus_nr: u8, + /// Maps to pxb-pcie `numa_node=N`. Required on Grace; omitting it causes + /// "Unknown NUMA node; performance will be reduced" in the guest kernel. + pub numa_node: Option, + /// Bus-attached IOMMU (arm-smmuv3 on aarch64). Intel IOMMU is a global + /// device on Q35 and lives on Machine::Q35, not here. + pub iommu: Option, + /// One entry per passthrough device on this SMMU. + /// 1 root port = 1 GPU (1:1 SMMU mapping). + /// N root ports = N GPUs sharing the same physical SMMU. + pub root_ports: Vec, +} + +pub struct PciRootPort { + pub id: String, // "pcie.portN" + pub chassis: u8, + pub device: Option, +} + +/// IOMMU that attaches to a specific PCIe expander bus (pxb-pcie). +/// Intel IOMMU is a Q35-global device and is NOT represented here — +/// see Machine::Q35::intel_iommu. +pub enum BusIommu { + SmmuV3 { + accel: bool, + ats: bool, + pasid: bool, + oas: u8, + ril: bool, + /// Enable SMMU command-queue virtualisation (vCMDQ). Requires + /// physically contiguous guest memory (hugepages or EGM). + cmdqv: bool, + }, +} +``` + +**SMMU grouping rule:** GPUs that share a physical SMMU on the host **must** be +placed on the same `PciRootComplex` in the guest (they share the same +`arm-smmuv3` device). The IOMMU group boundaries in host sysfs determine the +grouping. See [Config 3](#config-3--4-gpus-2-gpus-per-smmu-33-numa-nodes) for +the 2-GPUs-per-SMMU topology. + +#### `Objects` — shared QEMU `-object` backends + +```rust +pub struct Objects { + pub iommufd: Option, + pub memory_backends: Vec, + pub thread_contexts: Vec, + pub acpi_links: Vec, + pub rng: Option, +} + +pub enum MemoryBackend { + Ram { id: String, size: u64 }, + /// File-backed memory. Two uses: + /// path = "/dev/hugepages/" → hugepages-backed guest RAM (vCMDQ) + /// path = "/dev/egmN" → per-socket EGM region (vEGM) + File { id: String, size: u64, path: String, prealloc: bool, share: bool }, +} + +pub enum AcpiPciNodeLink { + /// Emitted 8× per passthrough GPU. The GPU driver uses these nodes to + /// online GPU memory to the guest kernel (required for MIG regardless of + /// whether MIG is actually enabled). + GenericInitiator { id: String, pci_dev: String, node: u32 }, + /// Emitted 1× per passthrough GPU. Links the GPU to the per-socket EGM + /// memory-backend file. `node` is the CpuMem NUMA node for the socket + /// that holds this GPU's EGM device, not a GPU initiator node. + EgmMemory { id: String, pci_dev: String, node: u32 }, +} +``` + +**EGM is per socket, not per GPU:** one `MemoryBackend::File` with `/dev/egmN` +per CPU socket. Two GPUs on the same socket share the full EGM backing; each +gets its own `acpi-egm-memory` pointing to that socket's CpuMem NUMA node. +See [Config 7](#config-7--vegm-2-gpus-per-socket-4-gpus-2-sockets). + +#### `HostTopology` — probe result driving `Platform` + +```rust +/// Read from host sysfs and IOMMU group layout before constructing Platform. +pub struct HostTopology { + pub sockets: Vec, + pub gpu_smmu_groups: Vec, + pub egm_sockets: Vec, +} + +pub struct SocketInfo { + pub id: u32, + pub cpu_range: std::ops::Range, +} + +/// All GPUs in this group share a physical SMMU and must be placed on the same +/// pxb-pcie + arm-smmuv3 in the guest. Derived from /sys/kernel/iommu_groups. +pub struct GpuSmmuGroup { + pub pci_bus_addrs: Vec, // e.g. ["0008:06:00.0", "0009:06:00.0"] + pub socket: u32, +} + +/// One entry per /dev/egmN device (created by the nvgrace-egm kernel module). +pub struct EgmSocketInfo { + pub path: String, // "/dev/egm4" + pub socket: u32, + pub total_size: u64, +} +``` + +`Platform::apply_host_defaults(topo)` consumes `HostTopology` to populate +`PciTopology::roots` (one `PciRootComplex` per `GpuSmmuGroup`) and +`Objects::memory_backends` (one `MemoryBackend::File` per `EgmSocketInfo`). +This is the **only** location that knows about DGX, GB300, or any host flavour. + +#### `Platform` — single wiring point + +```rust +pub struct Platform { + pub machine: Machine, + pub pci: PciTopology, + pub objects: Objects, +} + +impl Platform { + pub fn from_config(config: &HypervisorConfig) -> Result { … } + pub fn apply_host_defaults(&mut self, topo: &HostTopology) { … } + pub fn with_hugepages(mut self, path: &str) -> Self { … } +} +``` + +### NUMA Layout Rules + +The guest Linux kernel processes ACPI SRAT entries in a fixed order: + +1. **CPU Affinity** — nodes with a `cpus=` range (CpuMem nodes, one per socket) +2. **Generic Affinity** — initiator nodes for PCIe devices (8 per GPU for MIG) +3. **Memory-only Affinity** — nodes without CPUs (EGM backing, hotplug regions) + +The `-numa node` arguments **must appear in this order** in the QEMU command +line. Placing Generic Affinity nodes before CpuMem nodes causes the kernel to +assign wrong NUMA node IDs. + +**8 NUMA nodes per GPU (MIG):** Each passthrough GPU requires exactly 8 dedicated +generic-initiator NUMA nodes regardless of whether MIG is in use. The GPU +driver (CUDA) uses these nodes to online GPU memory to the guest kernel. Total +node count with 4 GPUs on a single-socket host: 1 CpuMem + 4 × 8 = 33 nodes. + +**Memory hotplug placeholder:** QEMU assigns the last-defined NUMA node to the +hotplug region. When DIMM/NVDIMM hotplug is enabled, always append one empty +NUMA node after all real nodes so hotplug memory does not overlap GPU nodes. + +**GPU memory spill prevention:** GPU NUMA nodes may attract page migration from +`autonuma` or systemd NUMA policies. Mitigate with explicit NUMA distances: + +```text +-numa dist,src=,dst=,val=254 +``` + +Or by disabling NUMA balancing in the guest OS. + +**`highmem-mmio-size` sizing on `-machine virt`:** +- GH200 / GB200 with ≤ 4 GPUs → `4T` +- GB300 NVL72 with 4 GPUs → `8T` +- Must be a power of 2; round up to the next power when in doubt. + +### Hugepages wiring (`with_hugepages`) + +```rust +pub fn with_hugepages(mut self, path: &str) -> Self { + let id = "m0".to_owned(); + self.objects.memory_backends.push(MemoryBackend::File { + id: id.clone(), + size: self.machine.memory_size(), + path: path.to_owned(), // "/dev/hugepages/" + prealloc: true, + share: true, + }); + self.machine.set_memory_backend(&id); // -machine memory-backend=m0 + self +} +``` + +This replaces the current pattern where `set_memory_backend_file` + +`set_memory_backend` are called ad-hoc from `QemuCmdLine` internals. The +caller checks `config.memory_info.enable_hugepages` and calls +`Platform::with_hugepages("/dev/hugepages/")` once. No device code changes. + +EGM memory is natively physically contiguous, so vEGM implicitly satisfies the +vCMDQ contiguity requirement without hugepages. `cmdqv=on` is still required +in the `arm-smmuv3` device args when vCMDQ is enabled. + +### Emission order + +`QemuCmdLine::build()` becomes a thin orchestrator: + +1. Emit `-object iommufd,id=iommufd0` first (all vfio devices and SMMU SID + tables reference it). +2. Emit remaining `Objects` — `memory_backends`, `thread_contexts`, `rng` — + all `-object` lines, IDs defined before any reference. +3. Emit `Machine` — picks up `memory-backend=` from objects, `highmem-mmio-size`. +4. Emit **CpuMem** `-numa node` entries: one per socket, with `cpus=` + `memdev=`. +5. Emit **GPU initiator** `-numa node` entries: 8 per GPU, no `cpus`/`memdev`, + ordered by GPU index. +6. Emit **EGM / hotplug** `-numa node` entries: memory-only nodes, no `cpus`. +7. Emit `PciTopology` in bus-number order: `pxb-pcie` (with `numa_node=`), + `arm-smmuv3`, root ports, vfio devices per root complex. +8. Emit `Objects.acpi_links`: `acpi-generic-initiator` (8 per GPU) then + `acpi-egm-memory` (1 per GPU). These reference device IDs emitted in step 7. + +Steps 4–6 must be in that order to match Linux ACPI SRAT processing. + +--- + +## Grace Platform Configurations + +The 7 configurations below are derived from tested production deployments of +NVIDIA Grace GPU passthrough. Each becomes a golden test fixture in **Phase 0b**. +The implementation must reproduce every one exactly from the corresponding +`Platform` + `HostTopology` input. + +All Grace configurations share these constants: +- `-device vfio-pci-nohotplug` (not `vfio-pci`) — required for C2C interconnect +- `-object iommufd,id=iommufd0` — modern IOMMU fd interface; legacy VFIO groups + not supported on Grace +- `arm-smmuv3` fixed parameters: `accel=on,ats=on,ril=off,pasid=on,oas=48` +- Host kernel driver: `nvgrace-gpu-vfio-pci` (replaces standard `vfio-pci`) +- EGM kernel module: `nvgrace-egm` (creates `/dev/egm*` character devices) + +### Config 1 — Single GPU, 1 SMMU (9 NUMA nodes) + +```text +-object iommufd,id=iommufd0 +-object memory-backend-ram,size=16G,id=m0 +-machine virt,accel=kvm,gic-version=3,ras=on,highmem-mmio-size=4T,memory-backend=m0 +-numa node,memdev=m0,cpus=0-3,nodeid=0 +-numa node,nodeid=1 +... +-numa node,nodeid=8 +-device pxb-pcie,id=pcie.1,bus_nr=1,bus=pcie.0,numa_node=0 +-device arm-smmuv3,primary-bus=pcie.1,id=smmuv3.1,accel=on,ats=on,ril=off,pasid=on,oas=48 +-device pcie-root-port,id=pcie.port1,bus=pcie.1,chassis=1,io-reserve=0 +-device vfio-pci-nohotplug,host=0008:06:00.0,bus=pcie.port1,rombar=0,id=dev0,iommufd=iommufd0 +-object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=1 +... +-object acpi-generic-initiator,id=gi7,pci-dev=dev0,node=8 +``` + +`HostTopology`: 1 socket, 1 `GpuSmmuGroup { pci_bus_addrs: ["0008:06:00.0"], socket: 0 }`. + +### Config 2 — 4 GPUs, 1 GPU per SMMU (33 NUMA nodes) + +Each GPU gets its own `PciRootComplex` (one `pxb-pcie` + one `arm-smmuv3` + one +root port). Repeat the pxb-pcie/smmuv3/root-port/vfio block 4 times: + +```text +-object iommufd,id=iommufd0 +-object memory-backend-ram,size=16G,id=m0 +-machine virt,...,highmem-mmio-size=4T,memory-backend=m0 +-numa node,memdev=m0,cpus=0-3,nodeid=0 +-numa node,nodeid=1 ... -numa node,nodeid=32 # 4×8 = 32 GPU initiator nodes + +# Per GPU (N = 1..4): +-device pxb-pcie,id=pcie.N,bus_nr=N,bus=pcie.0,numa_node=0 +-device arm-smmuv3,primary-bus=pcie.N,id=smmuv3.N,accel=on,ats=on,ril=off,pasid=on,oas=48 +-device pcie-root-port,id=pcie.portN,bus=pcie.N,chassis=N,io-reserve=0 +-device vfio-pci-nohotplug,host=,bus=pcie.portN,rombar=0,id=dev,iommufd=iommufd0 +-object acpi-generic-initiator,id=gi<8*(N-1)>,pci-dev=dev,node=<1+8*(N-1)> +... # ×8 per GPU +``` + +`HostTopology`: 1 socket, 4 `GpuSmmuGroup` each with 1 address. + +### Config 3 — 4 GPUs, 2 GPUs per SMMU (33 NUMA nodes) + +GPUs sharing a physical SMMU share one `PciRootComplex` with **2 root ports**. +2 complexes × 2 GPUs each: + +```text +-device pxb-pcie,id=pcie.1,bus_nr=1,bus=pcie.0,numa_node=0 +-device arm-smmuv3,primary-bus=pcie.1,id=smmuv3.1,accel=on,ats=on,ril=off,pasid=on,oas=48 +-device pcie-root-port,id=pcie.port1,bus=pcie.1,chassis=1,io-reserve=0 +-device vfio-pci-nohotplug,host=0008:06:00.0,bus=pcie.port1,rombar=0,id=dev0,iommufd=iommufd0 +-device pcie-root-port,id=pcie.port2,bus=pcie.1,chassis=2,io-reserve=0 +-device vfio-pci-nohotplug,host=0009:06:00.0,bus=pcie.port2,rombar=0,id=dev1,iommufd=iommufd0 + +-device pxb-pcie,id=pcie.2,bus_nr=9,bus=pcie.0,numa_node=0 +-device arm-smmuv3,primary-bus=pcie.2,id=smmuv3.2,accel=on,ats=on,ril=off,pasid=on,oas=48 +-device pcie-root-port,id=pcie.port3,bus=pcie.2,chassis=3,io-reserve=0 +-device vfio-pci-nohotplug,host=0010:06:00.0,bus=pcie.port3,rombar=0,id=dev2,iommufd=iommufd0 +-device pcie-root-port,id=pcie.port4,bus=pcie.2,chassis=4,io-reserve=0 +-device vfio-pci-nohotplug,host=0011:06:00.0,bus=pcie.port4,rombar=0,id=dev3,iommufd=iommufd0 + +-object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=1 +... +-object acpi-generic-initiator,id=gi7,pci-dev=dev0,node=8 +-object acpi-generic-initiator,id=gi8,pci-dev=dev1,node=9 +... +-object acpi-generic-initiator,id=gi15,pci-dev=dev1,node=16 +# × 2 more for dev2, dev3 +``` + +`HostTopology`: 1 socket, 2 `GpuSmmuGroup` each with 2 addresses. + +### Config 4 — GPU + NIC passthrough + +Same structure as Config 2 but one `PciRootPort` holds a NIC (`vfio-pci-nohotplug` +with the NIC's PCI address). That root port does **not** emit +`acpi-generic-initiator` links — the NIC has no GPU memory to online. + +`VfioDevice` carries a `kind: VfioDeviceKind` field (enum `Gpu` / `Nic` / …) +that gates initiator emission. The NIC shares the host SMMU with no GPU on its +bus, so it gets its own `PciRootComplex`. + +### Config 5 — vCMDQ (hugepages + SMMU command-queue virtualisation) + +Same PCIe topology as Config 1 or 2, but `MemoryBackend::Ram` is replaced with +`MemoryBackend::File` for physically contiguous memory (required by the vCMDQ +hardware for the queue base address), and `cmdqv=on` is added to `arm-smmuv3`: + +```text +-object memory-backend-file,id=m0,size=16G,mem-path=/dev/hugepages/,prealloc=on,share=on +-machine virt,...,memory-backend=m0 +-device arm-smmuv3,...,cmdqv=on +``` + +`Platform::with_hugepages("/dev/hugepages/")` + `IommuKind::SmmuV3 { cmdqv: true }`. + +### Config 6 — vEGM, 1 GPU per socket (4 GPUs, 4 sockets) + +One `memory-backend-file` per socket using the `/dev/egmN` device created by +`nvgrace-egm`. One `acpi-egm-memory` per GPU pointing to its socket's CpuMem +NUMA node: + +```text +-object memory-backend-file,id=m0,mem-path=/dev/egm4,size=56896M,share=on,prealloc=on +-object memory-backend-file,id=m1,mem-path=/dev/egm5,size=56896M,share=on,prealloc=on +-object memory-backend-file,id=m2,mem-path=/dev/egm6,size=56896M,share=on,prealloc=on +-object memory-backend-file,id=m3,mem-path=/dev/egm7,size=56896M,share=on,prealloc=on +-machine virt,... +-numa node,memdev=m0,cpus=0,nodeid=0 +-numa node,memdev=m1,cpus=1,nodeid=1 +-numa node,memdev=m2,cpus=2,nodeid=2 +-numa node,memdev=m3,cpus=3,nodeid=3 +-numa node,nodeid=4 ... -numa node,nodeid=35 # 4×8 GPU initiator nodes + +# PCI topology: 4× (pxb-pcie + smmuv3 + root-port + vfio) — same shape as Config 2 + +-object acpi-egm-memory,id=egm0,pci-dev=dev0,node=0 # GPU on socket 0 +-object acpi-egm-memory,id=egm1,pci-dev=dev1,node=1 # GPU on socket 1 +-object acpi-egm-memory,id=egm2,pci-dev=dev2,node=2 +-object acpi-egm-memory,id=egm3,pci-dev=dev3,node=3 +``` + +`HostTopology`: 4 sockets, 4 `GpuSmmuGroup` (1 GPU each), 4 `EgmSocketInfo`. + +### Config 7 — vEGM, 2 GPUs per socket (4 GPUs, 2 sockets) + +Two GPUs per socket share the socket's EGM device. The `/dev/egmN` path appears +in one `memory-backend-file` at full socket size. Both `acpi-egm-memory` entries +for that socket point to the same CpuMem NUMA node: + +```text +-object memory-backend-file,id=m0,mem-path=/dev/egm4,size=56896M,share=on,prealloc=on +-object memory-backend-file,id=m1,mem-path=/dev/egm5,size=56896M,share=on,prealloc=on +-machine virt,... +-numa node,memdev=m0,cpus=0-1,nodeid=0 +-numa node,memdev=m1,cpus=2-3,nodeid=1 +-numa node,nodeid=2 ... -numa node,nodeid=33 # 4×8 GPU initiator nodes + +# PCI topology: 2× (pxb-pcie + smmuv3 + 2 root ports + 2 vfio) — Config 3 shape + +-object acpi-egm-memory,id=egm0,pci-dev=dev0,node=0 # both GPUs on socket 0 → node=0 +-object acpi-egm-memory,id=egm1,pci-dev=dev1,node=0 +-object acpi-egm-memory,id=egm2,pci-dev=dev2,node=1 # both GPUs on socket 1 → node=1 +-object acpi-egm-memory,id=egm3,pci-dev=dev3,node=1 +``` + +`HostTopology`: 2 sockets, 2 `GpuSmmuGroup` (2 GPUs each), 2 `EgmSocketInfo`. + +--- + +## Migration Phases + +Each phase is a self-contained PR. Phases 0–1 introduce new types without +touching the hot path; Phases 2–5 strangle the old code one device at a time. + +### Phase 0 — Test harness and empty data types + +- **0a** — golden-test harness + one trivial fixture (basic `virt` machine) +- **0b** — All 7 Grace configurations as command-line fixtures + parse smoke test. + Each fixture provides the expected QEMU argument list and the `HostTopology` + input that produces it. Zero implementation; tests all fail intentionally. +- **0c** — empty `machine/` module with unit tests on pure helpers + (`format_memory`, `numa_node` string, bus-name helpers, etc.) + +No behaviour changes. CI green throughout. + +### Phase 1 — Platform probe (unused) + +Introduce `PlatformProbe` trait, `HostTopology` struct, and `Platform::from_config`. +Nothing in the hot path calls them yet. Tests assert construction succeeds for +each supported machine type and that `HostTopology` round-trips through +`apply_host_defaults` without panic. + +### Phase 2 — Strangle bus resolution (one device per PR) + +For each device listed below, pass a resolved `bus: String` instead of +computing it inside `ToQemuParams`: + +- `DeviceVhostUserFs` +- `DeviceVirtioBlk` +- `VhostVsock` +- `DeviceVirtioNet` +- `DeviceVirtioSerial` +- `DeviceVirtconsole` +- `DeviceRng` +- `DeviceIntelIommu` +- `DeviceVirtioScsi` + +Each PR removes one `#[cfg(target_arch)]` block and one duplicated comment. +The golden fixtures validate that the emitted command lines are unchanged. + +Final PR in Phase 2: remove all remaining `#[cfg(target_arch)]` blocks from +`cmdline_generator.rs`. + +### Phase 3 — Objects registry + +- Lift `MemoryBackendFile` into `Objects::memory_backends` as + `MemoryBackend::File`. +- Wire hugepages via `Platform::with_hugepages`. +- Lift `ObjectIoThread` / `ObjectRngRandom` into `Objects`. + +This is the phase that enables hugepages for `runtime-rs` (issue #12125). + +### Phase 4 — Multi-RC PCIe and NUMA layout + +- Emit `pxb-pcie` (with `numa_node=`) + per-RC `arm-smmuv3` from `PciTopology`. +- Support N root ports per `PciRootComplex` (Config 3 shape: 2 GPUs per SMMU). +- Add `VfioPciNoHotplug` with typed `IommufdRef` and `VfioDeviceKind`. +- Emit NUMA nodes in the correct order (CpuMem → GPU initiators → memory-only). +- `apply_host_defaults` wired end-to-end: Configs 1–4 golden fixtures pass. + +### Phase 5 — vCMDQ and vEGM + +- `SmmuV3 { cmdqv: true }` + `MemoryBackend::File { path: "/dev/hugepages/", … }`. +- `AcpiPciNodeLink::EgmMemory` + per-socket `MemoryBackend::File { path: "/dev/egmN", … }`. +- `EgmSocketInfo` probe wired into `apply_host_defaults`. +- Configs 5–7 golden fixtures pass. + +### Phase 6 — Cleanup + +- Delete dead code from `cmdline_generator.rs`. +- Remove feature flags / compat shims introduced during migration. +- Final golden-test sweep across all 7 Grace configs plus existing machine types. +- Remove all TODOs and decision stubs from this document. + +--- + +## Design Principles + +1. **Machine decisions stay in `Machine` and `PciTopology`.** + Devices receive resolved values; they do not branch on architecture or + machine type. + +2. **Singletons are first-class.** + `iommufd0` is emitted once and referenced via `IommufdRef`. No string + concatenation in device impls. + +3. **One location for host-specific wiring.** + `Platform::apply_host_defaults` is the only place that knows about DGX, + GB300, or any future host flavour. + +4. **IOMMU placement matches hardware placement.** + Bus-attached IOMMUs (`arm-smmuv3`) live on `PciRootComplex`. Global IOMMUs + (`intel-iommu`) live on the machine type (`Machine::Q35`). These are + different device placement models and must not share a field. + SMMU grouping: devices sharing a physical SMMU are placed on the same + `PciRootComplex`; devices on separate physical SMMUs get separate entries. + +5. **NUMA emission order is non-negotiable.** + CpuMem → GenericInitiator → MemoryOnly. The `Platform` builder enforces + this by construction; there is no API to emit them in a different order. + +6. **Incremental migration, no flag day.** + Each phase leaves CI green. Old `QemuCmdLine` and new `Platform` coexist + until the strangle is complete. + +7. **Types enforce constraints.** + `kernel_irqchip` compiles only on `Q35`. `cmdqv` compiles only on + `IommuKind::SmmuV3`. You cannot emit `gic-version` on a Q35 machine. + +--- + +## Related Documents + +- [Issue #12187](https://github.com/kata-containers/kata-containers/issues/12187) — full design spec with data-model definitions and worked examples +- [Issue #12125](https://github.com/kata-containers/kata-containers/issues/12125) — NUMA and hugepages roadmap