Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Device support matrix

The backend support table answers “which dtypes and programs does each backend admit”. This document answers the different question underneath it: which physical devices actually execute the work, and what exactly stops the rest from doing so.

The distinction matters because the host backends choose devices differently. Core ML and OpenVINO are runtimes that dispatch across a machine’s inference estate, not NPU drivers, so parts of this project already run on CPUs and GPUs. XDNA and Hexagon are single-device by design. Vulkan spans vendors and enumerates every suitable physical device, but each backend instance binds to one of them. Naming those differences explicitly is more useful than an “NPU” label that is true of the intent and only partly true of the code.

How to read this

Every row carries one of four statuses. They are claims about this repository, not about the hardware.

StatusMeaning
ValidatedExecuted on that device, with an evidence pin in-repo naming the host, driver, and runtime versions.
ReachableThe selection and dispatch path drives the device today with no code change, but no in-repo evidence pins that part.
One change awayNot reachable today. Each row names the single gate — a constant, a path, or a build condition — and what it would take.
Out of scopeNo path, and none implied by the current design.

“Reachable” is deliberately weaker than “supported”. It means the code will select the device and try; it does not promise the program admits, the numerics match, or the performance is sane.

The matrix

DeviceClassBackendHost OS / archStatus
Apple Neural Engine, Apple silicon (M-series)NPUcoremlmacOS 14+Validated — Apple M4, macOS 26.5.2 (performance.md)
Apple CPU, as Core ML per-operator placementCPUcoremlmacOS 14+Reachable — see Core ML
Apple GPUGPUcoremlmacOS 14+One change away — compute units are pinned
Intel Mac (CPU, AMD/Intel GPU)CPU / GPUcoremlmacOS 14+One change away — ANE gate refuses construction
Intel NPU, Core Ultra (Meteor Lake / Lunar Lake / Arrow Lake / Panther Lake)NPUopenvinoAny host with the runtimeValidated — arch 5010 (Panther Lake), OpenVINO 2026.4: the FP8 tier’s graphs compile and FP8 movement round-trips bit-exactly, 2026-09-18 (ADR 0009). Arch 40XX (Lunar Lake) refuses FP8 outright and the withholding is pinned the same day: the suite passes with the tier absent and the refusal asserted, so the backend probes at open rather than branding by arch; the float and INT8 tiers remain first device preference on every generation
Intel GPU, integrated Xe/UHD and discrete ArcGPUopenvinoAny host with the runtimeReachable — including indexed GPU.1
x86-64 CPU, IntelCPUopenvinoAny host with the runtimeValidated — openvino-host-test CI lane, OpenVINO 2026.3.0
x86-64 CPU, AMDCPUopenvinoAny host with the runtimeReachable — misreports vendor, see OpenVINO
ARM64 CPU (Apple silicon, Ampere, Raspberry Pi)CPUopenvinoAny host with an ARM CPU-plugin buildReachable — enumerates as CPU, misreports vendor
OpenVINO virtual devices (AUTO, MULTI, HETERO, BATCH)—openvino—One change away — resolution requires enumeration
AMD XDNA2 NPU, Strix Point / Strix Halo / Krackan PointNPUxdnaLinux, amdxdna driverValidated — PCI 1022:17f0 rev 0x20, Fedora 44 (baseline)
AMD XDNA1 NPU, Phoenix / Hawk PointNPUxdnaLinux, amdxdna driverReachable — ungated but unvalidated, see XDNA
Second and later XDNA NPUs in one hostNPUxdnaLinuxOne change away — device index is fixed at 0
Qualcomm Hexagon HTP v73, Snapdragon XNPUhexagonWindows 11 ARM64Validated — Snapdragon X126100, QAIRT 2.49 (baseline)
Qualcomm Hexagon HTP v75+, newer SnapdragonNPUhexagonWindows 11 ARM64Reachable — ungated, misreports v73
Qualcomm Adreno GPU / Kryo CPU via QNNGPU / CPUhexagonWindows 11 ARM64One change away — backend library path is fixed, and deliberately so
Snapdragon on Linux or AndroidNPUhexagon—One change away — build target gate
Intel Arc 140V, Lunar LakeGPUvulkanLinux x86-64Validated — Mesa 26.0.8 ANV (baseline); full FP32 operator tier suite passed 2026-09-06 (ADR 0007), FP8 operator tier 2026-09-18 (ADR 0009)
Intel Arc B390, Panther LakeGPUvulkanLinux x86-64Validated — Mesa 26.0.8 ANV, Vulkan 1.4.335; full device suite including the FP8 operator tier passed 2026-09-18 (ADR 0009)
AMD Radeon 860M, Krackan PointGPUvulkanLinux x86-64, amdgpuValidated — RADV Mesa 26.1.8, Vulkan 1.4.354; full FP32 operator tier suite passed 2026-09-08 in every advertised memory domain, and clean under Khronos synchronization validation
Apple M3 / M4 GPU, Apple siliconGPUvulkanmacOS, MoltenVK loaderValidated — M4 on MoltenVK 1.4.2, full FP32 and FP16 corpora 2026-09-17; M3 on MoltenVK, FP8 operator tier 2026-09-18 (ADR 0009)
Other Vulkan 1.3 compute devicesGPU / virtual GPU / CPUvulkanLinux, Android, Windows, macOSReachable — enumerated and selected at run time; no other hardware evidence pin yet
lavapipe / llvmpipe software ICDCPUvulkanLinux x86-64Validated — pinned by the vulkan-lavapipe-test CI lane and exercised by the full backend suite, including the FP32 operator corpus; the FP8 corpus also passed on llvmpipe (LLVM 21.1.8) on 2026-09-18 (ADR 0009)
No device, executed in software—mockAnyDeterministic in-memory reference; outside this vocabulary
Whatever a wrapped provider drives—vaccelAny (std)Pass-through; the wrapped backend decides

Per-backend detail

Apple Core ML (virtio-accel-coreml)

What selects the device. Nothing does, explicitly. The backend hands Core ML a fixed compute budget and Core ML places each operator itself:

configuration.computeUnits = MLComputeUnitsCPUAndNeuralEngine;

— coreml_bridge.m:277

That constant is the whole device policy. It grants the ANE and the CPU, and withholds the GPU.

The CPU is already in play. Because MLComputeUnitsCPUAndNeuralEngine includes the CPU, a model containing operators the ANE declines runs those operators on the CPU — silently, inside a submission that device_info() reports as AcceleratorClass::NPU. This is not a defect; it is Core ML’s documented placement model, and the crate README states it. It does mean the project’s first non-NPU execution path already exists and already ships, and that per-submission device attribution is not observable through the Accelerator contract.

What excludes everything else. Construction refuses any host without an ANE:

#![allow(unused)]
fn main() {
if unsafe { va_coreml_has_neural_engine() } == 0 {
    return Err(InitError::NeuralEngineUnavailable);
}
}

— macos.rs:383, backed by a MLNeuralEngineComputeDevice scan at coreml_bridge.m:121

This gate, not the framework, is what excludes Intel Macs — Core ML runs there, it simply has no ANE. It also excludes ANE-less VMs, which is why the CI example on macos-latest skips rather than fails.

The two changes.

  • Apple GPU: switch the constant to MLComputeUnitsAll. One line. It widens placement to the GPU without adding a device to select, so nothing above the bridge changes.
  • Intel Macs and ANE-less hosts: soften the ANE gate to a capability probe. Larger than it looks — identity.uuid (apple-coreml-ane) and identity.class (NPU) are compile-time constants at macos.rs:393 and would have to become runtime-derived to stay truthful.

Intel OpenVINO (virtio-accel-openvino)

The genuinely heterogeneous backend. Device selection is explicit, ordered, and already covers three classes:

#![allow(unused)]
fn main() {
let device = ["NPU", "GPU", "CPU"]
    .into_iter()
    .find_map(|preferred| { /* first enumerated match */ })
}

— native.rs:842

with_device (native.rs:855) pins one device by enumerated name or class prefix, so "GPU.1" selects the second GPU and "NPU" selects whatever NPU instance exists. matches_device (native.rs:634) makes prefix matching strict: "GPU" matches GPU and GPU.1, never GPUX.

Not an Intel-only backend. The build gate is a pkg-config probe for the runtime, not a target OS or vendor check (build.rs:35). Consequently the CPU plugin admits any x86-64 host including AMD, and an ARM CPU-plugin build enumerates CPU on Apple silicon or Ampere just as readily. The openvino-host-test CI lane installs OpenVINO 2026.3.0 on an x86-64 Ubuntu runner and executes the real path against the CPU plugin. Vulkan has a separate native CI lane pinned to the lavapipe software ICD.

Two truthfulness gaps. device_info_for (native.rs:642) hardcodes vendor_id: 0x8086, so an AMD or ARM CPU host reports itself as Intel. And a CPU device falls through to AcceleratorClass::OTHER, because the class enum defines OTHER, NPU, GPU, and DSP but no CPU. A guest cannot currently distinguish “a CPU” from “a device this backend does not recognize”. Adding AcceleratorClass::CPU is additive — the type is a #[repr(transparent)] u16 designed for exactly this.

The one change. Virtual devices (AUTO, MULTI, HETERO, BATCH) are unreachable because both constructors resolve requests against the enumerated device list, and the standard plugin set does not enumerate them. with_device("AUTO") therefore returns DeviceUnavailable. A pass-through for a known set of virtual names would unlock OpenVINO’s own scheduling — worth weighing against this project’s preference for one submission mapping to one identified device.

AMD XDNA (virtio-accel-xdna)

Single device, index zero. The process-wide owner takes the first device HRX reports and never enumerates further:

#![allow(unused)]
fn main() {
check(ffi::hrx_gpu_device_get(0, &mut device))
}

— native.rs:132

Multi-NPU hosts are therefore reachable only at index 0. The fix is mechanical — plumb an index through shared_device — but the HRX fork’s model is one process-wide device, so it is a design question rather than a parameter change.

No generation gate. Nothing checks the PCI ID. An XDNA1 part (Phoenix / Hawk Point) would be opened and driven, while device_info reports it as XDNA2 regardless:

#![allow(unused)]
fn main() {
uuid: *b"amd.xdna.npu\0\0\0\0",
vendor_id: 0x1022,
device_id: 0x17f0,
}

— native.rs:691

Program admission also applies the XDNA2 shape and local-memory envelope, so an XDNA1 host would most likely fail during compilation rather than produce wrong numerics — but it would fail confusingly, and it would misidentify itself first.

Linux only, by runtime. The build script requires libhrx.so and the amdxdna-native headers (build.rs:58), and the validated stack is the in-tree amdxdna driver exposing /dev/accel/accel0. Windows XDNA uses a different driver stack entirely and is not addressed.

Qualcomm Hexagon (virtio-accel-hexagon)

HTP only, by construction. The QNN backend library is a fixed path:

#![allow(unused)]
fn main() {
let path = root.join("lib/aarch64-windows-msvc/QnnHtp.dll");
}

— native.rs:81

QAIRT also ships QnnCpu, QnnGpu, and QnnDsp libraries, and parameterizing this path would reach the Adreno GPU and Kryo CPU. This exclusion is deliberate, not an oversight: the crate documents that SDK-free builds fail explicitly rather than fall back, so that a host silently executing on the CPU can never be mistaken for NPU evidence. Treat it as a policy to revisit consciously, not a gap to close.

No SoC gate. The build pins the target, windows + aarch64 (build.rs:31), but nothing pins the SoC. A newer Snapdragon with HTP v75 or v79 would load and run, reporting itself as v73 the whole time (uuid: *b"qualcomm-htp-v73", device_id: 73 at native.rs:457). In practice such a host also needs ADSP_LIBRARY_PATH pointed at its own DSP libraries — an environment concern the crate README covers — and the FP32/FP8 rejections recorded for v73 may not describe it correctly.

Linux and Android are a build gate, not a port. QAIRT ships aarch64 Linux and Android libraries. The windows/aarch64 assertion and the hardcoded lib/aarch64-windows-msvc path are the only two things naming the OS.

Vendor-neutral Vulkan (virtio-accel-vulkan)

One physical device per instance. The run-time-loaded Vulkan 1.3 path enumerates every device with a compute queue and synchronization2, then prefers discrete GPU, integrated GPU, virtual GPU, CPU, and other devices in that order. with_device selects an exact enumerated name instead:

#![allow(unused)]
fn main() {
let physical = devices
    .into_iter()
    .min_by_key(PhysicalDeviceRecord::rank)
}

— native.rs:1205

The identity is probed, not branded. UUID, vendor ID, and device ID come from Vulkan physical device properties. GPU-like devices report AcceleratorClass::GPU; a CPU ICD such as lavapipe reports OTHER, because the protocol 1.0 class set has no CPU value. Memory domains are likewise per-device: host-coherent Host is required, Shared is advertised only for a device-local and host-visible type, and Device only for device-local memory. Every submitted buffer remains a direct storage-buffer binding; staging occurs only during explicit reads and writes of device-local memory.

Current execution boundary. The advertised tier is the static 42-operator FP32 tier with BOOL/INT32 auxiliaries, plus the FP16 tier: the same operators over binary16 tensors, advertised on every device the backend opens — the kernels’ binary16 conversions are crate-owned integer and binary32 code, so the tier needs no device feature (ADR 0008). The provisional integer target is declared but not advertised. The native path and full backend conformance suite are validated on Intel Arc 140V through Mesa ANV, on llvmpipe/lavapipe, and on Apple M4 via MoltenVK; CI pins lavapipe so the native path cannot silently turn into the portable placeholder. The FP16 corpus passes on Apple M4 via MoltenVK and on Intel Arc LNL (Mesa ANV) and AMD Radeon 860M (RADV); the lavapipe CI lane exercises the tier on every change.

What each backend reports

Useful when reading DeviceInfo in a trace. OpenVINO varies its UUID and class from the enumerated device name; Vulkan reports the physical device’s actual Vulkan identity. The other provider rows are compile-time constants, which is why a v75 Snapdragon still reports qualcomm-htp-v73.

Backenduuidclassvendor_iddevice_id
coremlapple-coreml-aneNPU0x106b (Apple)0
openvinointel-ov-<device>NPU / GPU / OTHER0x8086 (always)0
xdnaamd.xdna.npuNPU0x1022 (AMD)0x17f0 (always)
hexagonqualcomm-htp-v73NPU0x17cb (Qualcomm)73 (always)
vulkanVulkan deviceUUIDGPU / OTHER for a CPU ICDVulkan physical-device propertyVulkan physical-device property
mockvirtio-accelmockNPU00

class comes from AcceleratorClass, an extensible u16 newtype: OTHER = 0, NPU = 1, GPU = 2, DSP = 3. Unknown values stay representable across implementations, so new classes are additive.

Out of scope

  • Vendor-specific GPU APIs beyond Vulkan. There is no CUDA, TensorRT, cuDNN, ROCm, or Metal backend. A conformant NVIDIA, AMD, Intel, or Apple portability-layer device may still be reachable through the Vulkan backend.
  • Guest-side device passthrough. The project claims no Virtio device ID yet; guests reach hardware through a host backend behind the vAccel adapter.
  • Non-Apple ANE-class fixed-function blocks with no runtime this project speaks to.

Keeping this current

This document tracks device reachability, which changes for different reasons than dtype coverage. Revisit it when:

  • a device selection constant moves — the compute-unit budget, the QNN library path, the HRX device index, the OpenVINO preference order, or Vulkan’s physical-device ranking;
  • a build.rs target or runtime gate changes, which is what most often converts “one change away” into “reachable”;
  • a hardware evidence pin lands in a crate README or performance.md, which is what converts “reachable” into “validated”; or
  • a DeviceIdentity field stops being a constant and starts being probed.

Adding a dtype or operator to an existing backend does not require a change here.