Emit the EGM memory-backend-file line as id, size, mem-path,
prealloc=on, share=on, matching the vCMDQ and Q35 SHM emission paths
and the normalized golden fixtures. QEMU treats the options as
unordered; one canonical order keeps the emitter free of per-backend
special cases.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Q35 NVIDIA GPU configs set hot_plug_vfio=no-port and cold_plug_vfio=root-port
with pcie_root_port=8. The 8 pcie-root-ports in the production invocation are
pre-provisioned for cold-plug: GPU VFIO devices are added to the static QEMU
command line before the VM boots, not via QMP after boot.
Update fixture comments, test comments, and ARCHITECTURE.md to use the correct
term and cite the config knobs.
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds fixture and ignored test (Phase 3) for the x86_64 kata use-case
based on a production QEMU invocation from a DGX x86 host (2026-07-07):
65 vCPUs across two NUMA sockets, 73728M total, 36864M per socket pinned
to host NUMA nodes via /dev/shm, 8 pcie-root-ports pre-provisioned on
pcie.0 for GPU hotplug via QMP.
Key architectural differences from the Grace/virt topology documented
in ARCHITECTURE.md:
- No kernel-irqchip on vanilla Q35 (only required for CoCo)
- NUMA memory model: separate file-backed backends with host-nodes+policy
rather than a single backend referenced on the -machine line
- GPU passthrough via QMP hotplug onto pcie-root-ports, not static
vfio-pci-nohotplug on pxb-pcie buses
Phase 3 items needed before the test can pass:
MemoryBackend::File { host_nodes, policy }, Objects::numa_distances,
HostTopology NUMA SHM fields, Q35 branch in to_qemu_args.
Also corrects the Platform Parity baseline for Q35 (no kernel-irqchip
in vanilla; add note that it is only needed for CoCo).
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
build("virt", ...) is called from both from_config (production) and
from_config_defaults (test). Hardcoding gic_version=3, ras=on and
highmem_mmio_size=4T in build would set Grace-specific values on every
virt machine regardless of the actual hardware, breaking vanilla aarch64
VMs that need none of these options.
Move the Grace defaults into from_config_defaults, which is #[cfg(test)]
and explicitly models the Grace fixture topology. The production path via
from_config will read these values from HypervisorConfig when that wiring
lands in a later phase.
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
seccompsandbox already maps to -sandbox in cmdline_generator.rs via
seccomp_sandbox: Option<String> in HypervisorConfig. When Platform
takes over argument emission in Phase 3, sandbox support needs a typed
Objects::seccomp_sandbox field so the legacy generator can be removed
cleanly.
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
from_config and from_config_defaults now build real Platform values
instead of todo!-ing. apply_host_defaults populates PciTopology from
HostTopology (one PciRootComplex + arm-smmuv3 per GpuSmmuGroup, one
PciRootPort per GPU), assigns GenericInitiator NUMA nodes (8 per GPU
after all CpuMem nodes), and adds EGM File backends plus EgmMemory
links when egm_sockets is present. with_hugepages swaps the primary
RAM backend for a File backend backed by the hugepages mount.
Six structural unit tests cover: default Virt construction, single GPU
topology, multi-GPU/multi-SMMU grouping, EGM backend and link counts,
hugepages swap, and hugepages-preserves-EGM. Golden fixture tests
remain #[ignore]d until to_qemu_args is wired in Phase 2+.
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
grace_4_gpu_and_nic: state explicitly that HostTopology cannot
represent NIC passthrough yet and that Phase 4 rewrites the body;
the fixture already defines the full expected output.
grace_5_vcmdq: expand the body to call with_hugepages() before
emission, documenting the step Config 5 requires. The method is a
stub until Phase 5, so the test stays ignored, but the body now
matches the shape Phase 5 lands.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
The vEGM fixtures used id,mem-path,size,share,prealloc while the vCMDQ
and Q35 SHM fixtures use id,size,mem-path with prealloc before share.
QEMU treats the options as unordered, but the golden tests compare
strings exactly, so a single canonical order keeps the emitter free of
per-backend special cases. Canonical: id, size, mem-path, then
host-nodes/policy where pinned, then prealloc=on, share=on.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
The highmem_mmio_size doc comment ended with a dangling "bytes." line;
fold the unit into the opening sentence, in virt.rs and the copy quoted
in ARCHITECTURE.md.
ARCHITECTURE.md claimed the Makefile already defaults CPUFEATURES and
TDXCPUFEATURES to include host-phys-bits=on; the repo defaults are
still pmu=off only. State the actual situation: the flag must be set
per configuration file until the Makefile defaults are extended.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
cspell flags domain terms and identifiers (SMMU, CMDQ, SRAT, VAES,
iommufd, pseries, prefetchable, and friends) as unknown words. Wrap
them in backticks: the spellcheck config ignores inline code spans,
and these tokens are identifiers or hardware acronyms that belong in
code style anyway.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
- q35_vanilla_kata_x86.args: header said 65 vCPUs but the NUMA ranges
(0-32, 33-65) cover 66; fix the count.
- q35_coco_snp_single_gpu.args: replace the production-captured
host-data blob with a clearly synthetic placeholder (base64 of
"KATA-SYNTHETIC-HOST-DATA-0000000"); the fixture validates arg shape
and ordering, not attestation material.
- ARCHITECTURE.md: kernel-irqchip -> kernel_irqchip in four places to
match the actual QEMU -machine option spelling used by the code and
fixtures.
Assisted-by: Claude <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
For CoCo SNP the CPU model is part of the attestation identity: a fleet
with mixed Milan/Genoa/Turin nodes produces different CPUID per node with
cpu=host, so attestation reference values diverge and rescheduled Pods
fail (#12329). Pinning to EPYC-v4 gives a stable, portable identity
across the fleet, but strips AVX-512 and VAES extensions, halving
AES-GCM throughput to ~4 GB/s (#12382).
CpuModel::EpycV4 { extra_features } resolves both: attestation identity
is fixed while SNP_CRYPTO_FEATURES re-enables the stripped extensions.
CpuModel::Host { extra_features } carries pmu=off and host-phys-bits=on
for vanilla/TDX x86_64 guests (#13270).
ARCHITECTURE.md gains a BaseMachine/CpuConfig section with the full
rationale table and links to the four issues as breadcrumbs.
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Explains the root cause (40-bit GPA without host-phys-bits=on), the
short-term fix in the Makefile, and the post-refactor path (Q35 Platform
builder should emit host-phys-bits=on unconditionally for x86_64/KVM).
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Four items promised in replies to Copilot review:
virt.rs: add "bytes." to highmem_mmio_size doc comment so the unit is
unambiguous when emitting QEMU highmem-mmio-size=.
platform.rs: clarify BaseMachine::memory_backend doc — None is correct
for multi-socket vEGM where each socket supplies its own memory-backend
via -numa node,memdev= rather than a single machine-wide backend.
ARCHITECTURE.md: update module layout diagram to reflect what Phase 0
actually introduces (platform.rs, topology.rs, tests.rs all present;
probe.rs holds HostTopology only — PlatformProbe is Phase 1). Annotate
runtime: RuntimeFeatures fields in both Machine code samples as Phase 3+
so the doc is self-consistent with the current stubs.
tests.rs: run cargo fmt (long assignment line on check() helper).
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds fixture and ignored test (Phase 3) for Q35 CoCo (SEV-SNP) with one
GPU passed through, based on a production invocation (AMD EPYC host,
2026-07-13, H100 80GB at 0000:e1:00.0).
Key differences from Grace/virt and vanilla Q35 documented in fixture
and ARCHITECTURE.md Planned Fixtures section:
- sev-snp-guest object emitted before the -machine line (QEMU requires
the protection object first); kernel_irqchip=split required for SNP
- Memory model: memory-backend-ram with host-nodes=1,policy=bind (not
file-backed; CoCo uses RAM backend for NUMA pinning on x86)
- GPU passthrough via pxb-pcie + pcie-root-port + vfio-pci (not
vfio-pci-nohotplug); no arm-smmuv3 — x86 uses global AMD/Intel IOMMU
- iommufd is per-device (id=iommufdvfio-<uuid>), not the shared
iommufd0 used on Grace
- x-pci-vendor-id/x-pci-device-id overrides required so the guest sees
correct device IDs for measured boot attestation
- AMDSEV.fd firmware (AMD-specific OVMF, not generic OVMF.fd)
- pxb-pcie bus_nr=32 (not the Grace 1-indexed cumulative formula)
Phase 3 items enumerated: Objects::protection, Q35::kernel_irqchip and
confidential_guest_support typed fields, MemoryBackend::Ram host_nodes
and policy, per-device iommufd, VfioDevice vendor/device id overrides.
TDX capture still needed.
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Q35 vanilla kata fixture (q35_vanilla_kata_x86.args) from a production
DGX x86 invocation (2026-07-07): 2-socket NUMA, 65 vCPUs, 36864M per
socket pinned to /dev/shm, 8 cold-plug pcie-root-ports on pcie.0.
Test added as #[ignore = "Phase 3"]; it drives the Phase 3 API design.
Key differences from the Grace/virt topology that inform Phase 3 scope:
- No kernel-irqchip on vanilla Q35 (only required for CoCo)
- NUMA SHM memory model: per-socket file-backed backends with host-nodes
and policy=bind rather than a single backend on the -machine line
- Q35 GPU passthrough uses cold-plug onto pre-provisioned pcie-root-ports
(cold_plug_vfio=root-port, pcie_root_port=8, hot_plug_vfio=no-port);
no pxb-pcie, no arm-smmuv3
Platform Parity section: documents legacy Machine fields without typed
homes (machine_accelerators, confidential_guest_support), baseline
-machine output for each supported machine type, and Platform fields
required before the legacy struct can be deleted.
Planned Fixture Configurations section: enumerates the four fixture sets
needed before Phase 6 closes (vanilla virt, Q35 vanilla, CoCo+GPU,
8-GPU+4-NVSwitch), notes what data is captured vs. still needed, and
lists the new types required for each.
Known Issues and Follow-up Items section: QMP startup timeout (#13343),
seccomp_sandbox not yet in Platform, machine_accelerators /
confidential_guest_support cross-references.
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Seven fixture files encode the expected QEMU argument list for each
Grace passthrough topology derived from tested production deployments:
grace_1_single_gpu -- 1 GPU, 1 SMMU, 9 NUMA nodes
grace_2_four_gpus_1_per_smmu -- 4 GPUs x 1-per-SMMU, 33 NUMA nodes
grace_3_four_gpus_2_per_smmu -- 4 GPUs x 2-per-SMMU, 33 NUMA nodes
grace_4_gpu_and_nic -- GPU + NIC, NIC emits no generic-initiators
grace_5_vcmdq -- hugepages backing + cmdqv=on on arm-smmuv3
grace_6_vegm_1_per_socket -- vEGM, 4 sockets, 1 GPU each
grace_7_vegm_2_per_socket -- vEGM, 2 sockets, 2 GPUs each
Each test is #[ignore] until the corresponding phase lands. The harness
reads fixtures as one Vec<String> element per line; blank lines and
lines starting with '#' are skipped.
The fixtures are the contract: any Platform implementation must reproduce
them exactly before a phase PR can be merged.
Refs: #12187, #12125
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
The existing cmdline_generator.rs (~3700 lines) encodes bus-assignment
logic inside each device and uses a single stringly-typed Machine struct
for all machine types. This makes adding Grace/vCMDQ/vEGM support
impractical without touching every device.
Introduce src/qemu/machine/ with the target types:
Platform -- top-level wiring point (machine + PCIe topology + objects)
Machine -- per-machine-type structs (Q35, Virt, Pseries, S390xCcwVirtio)
PciTopology -- PCIe expander buses with per-RC numa_node and BusIommu
BusIommu -- bus-attached IOMMU (SmmuV3); IntelIommu lives on Q35 directly
Objects -- typed -object registry (MemoryBackend, AcpiPciNodeLink, ...)
HostTopology -- probe result consumed by apply_host_defaults
All methods are todo!() stubs. No behaviour change; cmdline_generator.rs
is untouched. Implementations land phase by phase starting at Phase 2.
Refs: #12187, #12125
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
The current cmdline_generator.rs (~3700 lines) has three structural
problems: bus-assignment logic duplicated across every device via
#[cfg(target_arch)], a single stringly-typed Machine struct shared by
all machine types, and no NUMA-aware memory backend support.
Add ARCHITECTURE.md describing the target design before any code
changes land. The document covers:
current pain points
target module layout and core types (Platform, Machine, PciTopology,
BusIommu, Objects, HostTopology)
NUMA layout rules (ACPI SRAT ordering, 8 nodes per GPU for MIG,
highmem-mmio-size sizing, hotplug placeholder)
7 Grace platform configurations as golden test fixtures
6-phase migration plan (Phase 0 test harness through Phase 6 cleanup)
design principles
The document is a living reference updated with each phase PR.
It becomes a clean stable reference once Phase 6 completes.
Refs: #12187, #12125
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Assisted-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Signed-off-by: Zvonko Kaiser <zkaiser@nvidia.com>
Only enable discard for block-plain emptyDir volumes when the
hypervisor reports block device discard support.
The resource manager now queries hypervisor capabilities once and
passes block discard support through VolumeContext. Block emptyDir
volume setup uses that capability to decide whether to add the
discard mount option and related mount metadata, avoiding unsupported
discard mounts on hypervisors that cannot expose discard/unmap to the
guest.
It will correctly adjust the related codes with the new support.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Map BlockConfigModern.discard_unmap to Dragonball block device
sparse support when adding modern block devices.
This lets runtime-rs request guest-visible discard support through
the Dragonball block device configuration while leaving vhost-user
block devices disabled.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Add the Cloud Hypervisor disk sparse option to the local DiskConfig
model and keep its default aligned with upstream. Map
BlockConfigModern.discard_unmap to DiskConfig.sparse so block devices
that request discard/unmap explicitly expose sparse discard support
through the CH API.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Remove the remaining legacy Block device implementation now that
BlockModern covers all block paths, finishing the migration started
by the QEMU and DeviceManager legacy-path drops.
Delete virtio_blk.rs (BlockConfig, BlockDevice and its Device impl)
and stop exporting it from driver/mod.rs. Drop the BlockCfg variant
from DeviceConfig and the Block variant from DeviceType. Remove the
DeviceConfig::BlockCfg handling and the create_block_device helper
from DeviceManager, and drop DeviceType::Block from the QEMU
hotunplug unsupported list.
No new behavior; pure removal of the dead legacy path.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Remove DeviceType::Block so BlockModern is the single code path,
completing the BlockConfig -> BlockConfigModern migration at the
Device Manager layer.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
The preceding "Use BlockConfigModern for ..." series converted every
block-device producer in the Qemu path to emit BlockModern. The Qemu
inner module still held the matching legacy Block consumer code across
all three lifecycle ops (coldplug/hotplug/hotunplug), now unreachable.
Remove it so BlockModern is the single code path, completing the
BlockConfig -> BlockConfigModern migration at the Qemu layer.
The hotunplug arm for Block now returns an explicit "unsupported"
error instead of a stale QMP detach. The new added Block(_) is just a
place holder and it will be removed in a follow-up commit.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Migrate the block config from BlockConfig to BlockConfigModern,
routing it through DeviceConfig::BlockCfgModern for firecacker.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Migrate the erofs rootfs path from BlockConfig to BlockConfigModern,
routing its block-device attachment through BlockCfgModern.
This covers the rw layer, single-device raw, multi-device VMDK, and
GPT-partitioned VMDK attachment points in erofs_rootfs.rs.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Migrate the block config from BlockConfig to BlockConfigModern,
routing it through DeviceConfig::BlockCfgModern for swap task.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Migrate the block config from BlockConfig to BlockConfigModern,
routing it through DeviceConfig::BlockCfgModern for block rootfs.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Migrate the host block-device (major/minor) passthrough path in
ResourceManagerInner from BlockConfig to BlockConfigModern, routing it
through DeviceConfig::BlockCfgModern. This continues the BlockModern
migration to cover block-mode devices carried in the OCI spec.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Switch the InitData resource path from BlockConfig to BlockConfigModern
so the InitData is handled through the modern block device pipeline,
consistent with the recent BlockModern migration for block volumes.
This changes ResourceConfig::InitData to wrap BlockConfigModern,
routes its handling through DeviceConfig::BlockCfgModern, and updates
its related helper function in virt_container to build a
BlockConfigModern.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Switch the VmRootfs resource path from BlockConfig to BlockConfigModern
so the VM rootfs is handled through the modern block device pipeline,
consistent with the recent BlockModern migration for block volumes.
This changes ResourceConfig::VmRootfs to wrap BlockConfigModern, routes
its handling through DeviceConfig::BlockCfgModern, and updates
prepare_rootfs_config() in virt_container to build a BlockConfigModern.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Introduce a new field of serial_override in the BlockModern config.
Add a DeviceType::BlockModern branch to the QEMU QMP hotplug path,
mirroring the existing Block semantics: skip block devices that
represent the initrd, and dispatch to nvdimm or blk/ccw/scsi
according to driver_option. Snapshot the config parameters under
lock and run the asynchronous hotplug without holding it to avoid
awaiting across the lock.
Pass the real block format and, when enabled, the first independent
iothread (indep_iothread_0) to hotplug_block_device instead of the
default format and None, and return DeviceType::BlockModern so the
hotplugged device is tracked.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Treat SPDK and VFIO direct volumes as a two-step mount. The agent
storage entry mounts the device inside the guest with block storage
options, while the OCI mount bind-mounts that guest path into the
container.
Filter bind flags from storage options, preserve direct-volume
mountInfo options when present, and add ro when the volume is
read-only.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Build regular block volumes with BlockConfigModern and submit them
through DeviceConfig::BlockCfgModern. Preserve the existing block
device attributes while also recording the host path required by the
modern block device model.
Update block volume helper call sites with the default option slice
so this change remains buildable with the modern helper signature.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Switch the Cloud Hypervisor device handling path from the legacy
Block device variant to BlockModern. Keep the block device behind
its shared mutex while snapshotting the configuration for the
hotplug request, and persist the PCI path assigned by Cloud
Hypervisor back into the shared device state.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Introduce build_bind_mount_options() to construct OCI mount options for
container-side bind mount, preferring volume_options over OCI options
and ensuring a bind/rbind flag is present.
And addfilter_block_storage_options() to strip bind/rbind flags from
Storage.options, preventing ENOTDIR when the agent mounts a block
device to a filesystem inside the guest.
And then Extend handle_block_volume() with a volume_options parameter
and replace hardcoded ro/empty options with the new filter+ro logic,
narrowing the function to a BlockModern-only path by removing the
DeviceType::Block branch.
Fix generate_shared_path to always create under rw/ since ro/
is a read-only bind mount of rw/ and direct creation fails with ReadOnly
FS. Switch the OCI Mount type from fs_type to KATA_MOUNT_BIND_TYPE because
the mount source is a directory (guest_path), not a block device.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Move the block device type definitions shared between BlockDevice and
BlockDeviceModern — VIRTIO_BLOCK_{PCI,MMIO,CCW}/VIRTIO_PMEM constants,
KATA_*_DEV_TYPE aliases, and the BlockDeviceAio / BlockDeviceFormat
enums — from virtio_blk.rs into virtio_blk_modern.rs, so they live
alongside the BlockModern handler that will replace the legacy
BlockDevice.
Add a `format: BlockDeviceFormat` field to BlockConfigModern to match
the legacy BlockConfig. Update mod.rs re-exports accordingly.
No behavior change; pure relocation plus the new field.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Extend add_block_device with num_queues and queue_size parameters
to support multi-queue virtio-blk configuration, and add BlockModern
hotplug/hotunplug handler that extracts config from
Arc<Mutex<BlockDeviceModern>>, passes the new parameters to
add_block_device, and updates pci_path on successful hotplug.
And we will use BlockDeviceModern to handle block device instead of
Legacy BlockDevice.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Add BlockModern handling in DeviceManager so that device index is
correctly released both in get_device_info and in the error path.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Use the hotplug-returned SCSI address as Storage.source when a
container rootfs is backed by a block device.
Signed-off-by: Manuel Huber <manuelh@nvidia.com>
The network hotplug (handle_network_device) passed `queue_num`, a
queue pair count, straight as `num_queues`, while cloud-hypervisor's
virtio-net expects a raw virtqueue count.
With the default network_queues=1 this hit "Number of queues (1) to
virtio_net should be higher than 2" on the post-start netdev add,
failing sandbox start.
Convert the pair count to the raw queue count (queue_num.max(1) * 2),
matching the coldplug path which already does network_queues_pairs * 2.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
NetworkConfig.queue_num is a queue pair count, but Dragonball's
virtio/vhost-net backends expect an even raw virtqueue count. Passing
it through without doubling caused InvalidQueueNum(1) on the default
network_queues=1, failing sandbox start.
Convert pair count to raw count (queue_num.max(1) * 2) at both
Dragonball network entry points, mirroring the Cloud Hypervisor
backend.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
As part of the effort to thread network queues from the top-level
configuration all the way down to each endpoint instead of having
low-level helpers re-read the global configuration, propagate the
per-network queue_num into the QEMU cmdline generator and hotplug
path.
add_network_device() and get_network_device() now take an explicit
num_queues argument sourced from network.config.queue_num, replacing
the direct read of network_info.network_queues, so each network
endpoint honors its own configured queue count.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
As there's no field to map the configuration's network_queues item
in the `DanConfig`, this commit introduces a network_queues to do this.
And accordingly, we also make it passed down from sandbox layer to
Dan network configurations.
To make it work well, it make it more robust for queues and queue_size
settings with checking logics. And related UT is added.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>
Set queue_num (derived from net_pair.queues, one TX + one RX per
queue pair) and queue_size (256) when building the network pair for
veth, macvlan, ipvlan and vlan endpoints so the virtio-net device is
created with the configured multi-queue layout.
Meanwhile, update endpoint tests accordingly.
Signed-off-by: Alex Lyn <alex.lyn@antgroup.com>