hypervisor/qemu: add Q35 fixture, platform parity, and planned fixtures

Q35 vanilla kata fixture (q35_vanilla_kata_x86.args) from a production
DGX x86 invocation (2026-07-07): 2-socket NUMA, 65 vCPUs, 36864M per
socket pinned to /dev/shm, 8 cold-plug pcie-root-ports on pcie.0.
Test added as #[ignore = "Phase 3"]; it drives the Phase 3 API design.

Key differences from the Grace/virt topology that inform Phase 3 scope:
- No kernel-irqchip on vanilla Q35 (only required for CoCo)
- NUMA SHM memory model: per-socket file-backed backends with host-nodes
  and policy=bind rather than a single backend on the -machine line
- Q35 GPU passthrough uses cold-plug onto pre-provisioned pcie-root-ports
  (cold_plug_vfio=root-port, pcie_root_port=8, hot_plug_vfio=no-port);
  no pxb-pcie, no arm-smmuv3

Platform Parity section: documents legacy Machine fields without typed
homes (machine_accelerators, confidential_guest_support), baseline
-machine output for each supported machine type, and Platform fields
required before the legacy struct can be deleted.

Planned Fixture Configurations section: enumerates the four fixture sets
needed before Phase 6 closes (vanilla virt, Q35 vanilla, CoCo+GPU,
8-GPU+4-NVSwitch), notes what data is captured vs. still needed, and
lists the new types required for each.

Known Issues and Follow-up Items section: QMP startup timeout (#13343),
seccomp_sandbox not yet in Platform, machine_accelerators /
confidential_guest_support cross-references.

Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Zvonko Kaiser
2026-07-13 22:20:37 +00:00
parent f8af92b8be
commit 43ac4f83c9
3 changed files with 263 additions and 0 deletions

View File

@@ -606,6 +606,192 @@ This is the phase that enables hugepages for `runtime-rs` (issue #12125).
---
## Platform Parity
The `Platform` type must generate correct QEMU command lines for every supported
machine type, not just Grace/virt. The legacy `Machine` struct carried two raw
string fields that have no typed home yet:
| Field | Legacy value | Machine types |
|------------------------------|--------------------------------------|-------------------|
| `accel` | `"kvm"`, `"tcg"`, `"kvm:tcg"` | all |
| `options` | raw KVM accelerator options | x86/arm |
| `kernel_irqchip` | `"on"`, `"split"`, `"off"` | Q35 only |
| `confidential_guest_support` | `"sev-snp0"`, `"tdx0"`, `""` | Q35 / virt |
| `memory_backend` | e.g. `"m0"` | virt (NUMA/EGM) |
These are preserved as raw strings in `BaseMachine` for now. Phase 3 will
introduce typed representations.
### Baseline `-machine` output per type
The examples below show the minimum expected output for vanilla (no GPU) configs
derived from tested deployments. They anchor the per-machine fixture set.
**virt (aarch64 — vanilla kata)**
```text
-machine virt,accel=kvm,gic-version=3,ras=on
```
**Q35 (x86_64 — vanilla kata)**
```text
-machine q35,accel=kvm
```
(`kernel-irqchip` is absent on vanilla Q35; it is only required for CoCo.)
**Q35 + TDX (x86_64 — CoCo)**
```text
-object tdx-guest,id=tdx,...
-machine q35,accel=kvm,kernel-irqchip=split,confidential-guest-support=tdx
```
**Q35 + SEV-SNP (x86_64 — CoCo)**
```text
-object sev-snp-guest,id=sev-snp,...
-machine q35,accel=kvm,kernel-irqchip=split,confidential-guest-support=sev-snp
```
**s390-ccw-virtio (s390x)**
```text
-machine s390-ccw-virtio,accel=kvm
```
Both `machine_accelerators` (the raw KVM option string) and
`confidential_guest_support` need typed representations before the legacy
`Machine` struct can be deleted. Tracked in Phase 3.
---
## Planned Fixture Configurations
The 7 Grace fixtures cover Grace GPU passthrough thoroughly. The configurations
below must also be captured as golden fixtures before Phase 6 closes. Each
entry notes the data source required: fixture content must come from actual
production QEMU invocations, not from documentation.
### Vanilla kata — virt (aarch64)
Basic `virt` machine with no GPU passthrough, no NUMA, no hugepages.
Represents the common ARM64 kata use-case.
**Data needed:** capture `qemu-system-aarch64` invocation from a running
non-GPU kata pod on an ARM64 host.
### Vanilla kata — Q35 (x86_64)
**Production data captured** (DGX x86 host, 2026-07-07).
Fixture: `q35_vanilla_kata_x86.args`. Test: `q35_vanilla_kata_x86` (ignored, Phase 3).
Key observations from the production invocation:
- `-machine q35,accel=kvm` — no `kernel-irqchip` on vanilla; only required for CoCo
- NUMA memory model differs from Grace: total memory via `-m 73728M,slots=10,maxmem=127052M`;
NUMA pinning via separate `memory-backend-file` objects with `host-nodes=N,policy=bind`
backed by `/dev/shm` (not `/dev/hugepages` or `/dev/egm*`)
- Two NUMA nodes: socket 0 cpus 0-32 / 36864M, socket 1 cpus 33-65 / 36864M;
distance 20 between them
- 8 `pcie-root-port` pre-provisioned on `pcie.0` (slots 0-7) for GPU cold-plug
(`cold_plug_vfio=root-port`, `pcie_root_port=8`, `hot_plug_vfio=no-port` in
`configuration-qemu-nvidia-gpu.toml.in`); GPU VFIO devices are added to the
static QEMU command line before the VM boots, not via QMP after boot
- No `pxb-pcie`, no `arm-smmuv3` — Q35 GPU passthrough uses cold-plug onto `pcie.0`
root ports, not the static vfio-pci-nohotplug topology used on Grace
New Platform fields required (Phase 3):
- `MemoryBackend::File { host_nodes: Option<u32>, policy: Option<String> }` for NUMA SHM
- `Objects::numa_distances: Vec<(u32, u32, u32)>` for `-numa dist` entries
- `HostTopology` additions for NUMA SHM paths and distances
### CoCo + GPU passthrough (SEV-SNP or TDX)
Confidential compute requires:
- A protection object (`sev-snp-guest` or `tdx-guest`) before the machine line
- `confidential-guest-support=<id>` on the `-machine` line
- `kernel-irqchip=split` instead of `on` (TDX/SNP requirement)
- OVMF firmware path instead of the standard BIOS (`-bios/-pflash`)
- Memory hotplug disabled (confidential memory cannot be hot-added)
CoCo + GPU passthrough combines these with the VFIO topology from the Grace
configs. The `ProtectionDeviceConfig::SevSnp` and `::Tdx` paths exist in
`cmdline_generator.rs` (`add_sev_snp_protection_device` / `add_tdx_protection_device`);
Platform needs a typed `Objects::protection` field to carry these through.
**Data needed:** capture from a CoCo-enabled kata pod with at least one GPU
passed through on both an AMD (SEV-SNP) and Intel (TDX) host.
### 8 GPUs + 4 NVSwitches (DGX/HGX topology)
NVSwitch passthrough adds a new device kind and a multi-level PCIe hierarchy
not present in the Grace configs:
- `VfioDeviceKind::NvSwitch` is not yet defined in `topology.rs` (only `Gpu`
and `Nic` exist).
- NVSwitches currently use `VfioDeviceConfig` (not `VfioDeviceGroup`) in the
legacy path (`add_gpu_nvswitch_setup` at cmdline_generator.rs:3373).
- PCIe hierarchy: root port → `x3130-upstream``xio3130-downstream`
device (three levels vs. the two levels used for GPU direct attachment).
- `add_pcie_switch_ports` (cmdline_generator.rs:3508) emits this hierarchy;
`PciTopology` has no equivalent typed representation yet.
New types needed before a fixture can be written:
- `VfioDeviceKind::NvSwitch`
- `PciSwitchPort { upstream: PcieUpstreamPort, downstream: Vec<PcieDownstreamPort> }` on `PciRootComplex`
- `HostTopology::nvswitch_addrs` or equivalent probe field
**Data needed:** capture from a DGX/HGX or GB200 NVL system with 8 GPUs and
4 NVSwitches passed through. Exact bus_nr arithmetic and PCIe address
assignments must come from a live invocation, not from inference.
---
## Known Issues and Follow-up Items
Misconfiguration and known defects in the current runtime-rs/QEMU path that
are related to this refactor. Items marked **post-refactor** require the new
`Platform` emission path to be wired end-to-end before they can be addressed
cleanly.
### QMP startup timeout is hard-coded and ignores the caller's deadline (#13343)
`QemuInner::start_vm(_timeout)` accepts the timeout argument from the
`Hypervisor` trait but silently ignores it. QMP initialisation instead uses
two independent hard-coded values:
- 5 s per-read socket timeout (`QMP_SOCKET_TIMEOUT`)
- 50 s overall connect/init deadline (`QMP_INIT_TIMEOUT`)
Under large-memory / VFIO / GPU passthrough conditions QEMU can take longer
than 50 s before QMP is fully responsive, causing a spurious timeout failure
even though the higher-level `start_vm` caller's intended budget was never
applied.
The `Hypervisor::start_vm(timeout: i32)` contract is also inconsistently
interpreted across backends: QEMU and Firecracker ignore it, Dragonball
treats it as milliseconds (despite comments saying seconds), Cloud Hypervisor
treats it as seconds, and Remote treats it as seconds for the RPC deadline.
**Minimum fix for GPU passthrough (pre-refactor):** make `QMP_INIT_TIMEOUT`
configurable via `HypervisorConfig` (e.g. `qmp_init_timeout_secs`) so
operators can raise it without patching.
**Post-refactor:** once `Platform::to_qemu_args` drives QEMU startup, wire
the QMP connect/init step to consume the caller's remaining `start_vm` budget
rather than maintaining an independent fixed deadline.
### `seccomp_sandbox` (-sandbox) not plumbed through Platform
Pre-refactor the option works via the legacy `cmdline_generator.rs` path.
Post-refactor it needs a typed `Objects::seccomp_sandbox` field so the legacy
generator can be removed. Tracked in Phase 3.
### `machine_accelerators` and `confidential_guest_support` have no typed home
Tracked in the Platform Parity section. Both are passed as raw strings
through `BaseMachine` today; they need typed representations before the
legacy `Machine` struct can be deleted.
---
## Design Principles
1. **Machine decisions stay in `Machine` and `PciTopology`.**

View File

@@ -194,3 +194,26 @@ fn grace_7_vegm_2_per_socket() {
"grace_7_vegm_2_per_socket.args",
);
}
// ---- Q35 x86_64: vanilla kata, 2-socket NUMA, 8 cold-plug root ports ----
//
// Production capture: DGX x86 host, 2026-07-07. 65 vCPUs, 73728M total,
// 36864M per socket pinned to host NUMA node via /dev/shm. 8 pcie-root-ports
// pre-provisioned on pcie.0 for GPU cold-plug (hot_plug_vfio=no-port).
//
// Blocked on Phase 3:
// - Q35 machine in Platform::to_qemu_args (no gic-version, no highmem-mmio-size)
// - MemoryBackend::File { host_nodes, policy } fields for NUMA SHM pinning
// - Objects::numa_distances for -numa dist entries
// - HostTopology fields for NUMA node + SHM path per socket
#[test]
#[ignore = "Phase 3: Q35 machine + NUMA SHM memory model not yet implemented"]
fn q35_vanilla_kata_x86() {
// HostTopology shape TBD in Phase 3.
// 2 sockets: socket 0 cpus 0-32 (host-node 0), socket 1 cpus 33-65 (host-node 1).
// No gpu_smmu_groups, no egm_sockets.
// 8 cold-plug root ports on pcie.0 (cold_plug_vfio=root-port, pcie_root_port=8);
// NUMA distance 20 between the two nodes.
todo!("Phase 3")
}

View File

@@ -0,0 +1,54 @@
# Q35 vanilla kata — x86_64, 2-socket NUMA, 65 vCPUs, no GPU assigned
# Source: production kata-containers log, July 2026 (NVIDIA DGX x86 host)
#
# Platform scope: machine line + NUMA SHM memory backends + NUMA topology +
# pre-provisioned pcie-root-ports for GPU cold-plug.
# NOT included here (legacy generator): -name, -uuid, -cpu, -m, -smp, all
# virtio devices, -netdev, -kernel, -append, -bios, -rtc, -pidfile, etc.
#
# HostTopology: 2 sockets
# socket 0: cpus 0-32 (33 vCPUs), host-node 0
# socket 1: cpus 33-65 (33 vCPUs), host-node 1
# Memory: 73728M total — 36864M per socket, pinned to host NUMA node via /dev/shm
# PCI: 8 pcie-root-ports on pcie.0 — cold-plug slots (cold_plug_vfio=root-port,
# pcie_root_port=8 in configuration-qemu-nvidia-gpu.toml.in);
# hot_plug_vfio=no-port, hotplug is disabled
# NUMA distance: 20 between the two nodes
#
# When a GPU is assigned to a container, -device vfio-pci,host=<bdf>,bus=rpN
# is added to the static QEMU command line (before VM boot) by the legacy
# generator. This fixture captures the base VM without any GPU assigned.
#
# New Platform fields required (Phase 3):
# MemoryBackend::File { host_nodes: Some(N), policy: "bind" }
# Objects::numa_distances: Vec<(u32, u32, u32)> (src, dst, val)
-machine
q35,accel=kvm
-object
memory-backend-file,id=numa-mem0,size=36864M,mem-path=/dev/shm,host-nodes=0,policy=bind,share=on
-numa
node,nodeid=0,memdev=numa-mem0,cpus=0-32
-object
memory-backend-file,id=numa-mem1,size=36864M,mem-path=/dev/shm,host-nodes=1,policy=bind,share=on
-numa
node,nodeid=1,memdev=numa-mem1,cpus=33-65
-numa
dist,src=0,dst=1,val=20
-numa
dist,src=1,dst=0,val=20
-device
pcie-root-port,id=rp0,bus=pcie.0,chassis=0,slot=0,multifunction=off
-device
pcie-root-port,id=rp1,bus=pcie.0,chassis=0,slot=1,multifunction=off
-device
pcie-root-port,id=rp2,bus=pcie.0,chassis=0,slot=2,multifunction=off
-device
pcie-root-port,id=rp3,bus=pcie.0,chassis=0,slot=3,multifunction=off
-device
pcie-root-port,id=rp4,bus=pcie.0,chassis=0,slot=4,multifunction=off
-device
pcie-root-port,id=rp5,bus=pcie.0,chassis=0,slot=5,multifunction=off
-device
pcie-root-port,id=rp6,bus=pcie.0,chassis=0,slot=6,multifunction=off
-device
pcie-root-port,id=rp7,bus=pcie.0,chassis=0,slot=7,multifunction=off