Protocol 1.0 frozen · workspace 0.3.x
Give a guest an accelerator, not a driver.
virtio-accel is a transport-neutral virtual accelerator device: a frozen wire
contract plus executable no_std guest, device, transport and TOSA layers in
Rust. Native backends now reach NPUs, GPUs and CPUs while the core model remains neutral
enough for DSPs and other program-driven accelerators.
// device-owned state, bounded generational IDs
let ctx = device.create_context(&desc)?;
let buf = ctx.import_buffer(bytes, Access::ReadWrite)?;
let prog = ctx.load_program(tosa_artifact)?;
let queue = ctx.create_queue(QueueDesc::default())?;
let token = queue.submit(prog, &[buf])?;
// success or failure, an event always arrives
match queue.wait(token)? {
Event::Complete(e) => use_outputs(e),
Event::Failed(e) => reclaim(e),
}
What it is
The portable majority of the system
virtio-accel concentrates on the part that should not have to be rewritten per platform: protocol wire structures, transport ports, device-owned state, a guest client, a mock backend and a transport-free conformance suite.
What it deliberately does not contain: Linux ioctls, macOS frameworks, Windows APIs, guest physical addresses, vendor command formats, or a claimed Virtio device ID.
Platform adapters are meant to be written against these layers rather than
inside them. Each crate is pinned to a portability tier that CI enforces on bare-metal
aarch64-unknown-none,
riscv64gc-unknown-none-elf and
wasm32-unknown-unknown, so a crate cannot quietly
acquire a host dependency.
Twelve of the eighteen crates forbid unsafe
outright. The rest quarantine it: virtio-accel-tosa to
flatc-generated schema accessors, and the five host backends to their audited native code, which
compile only when the native runtime is present — every other build of them forbids
unsafe too.
Protocol 1.0, frozen
Wire ABI →A versioned wire contract for exposing an accelerator to a guest: contexts, buffers, programs, execution queues, submissions and events.
Decoding is total
Unknown opcodes, statuses and event states stay raw integers until validated, so decoding untrusted bytes never constructs an invalid Rust enum.
Failure still returns
A successful submit yields an event, and an indeterminate failure must yield one too — the operation's resources are still owned by the device.
Tiers are enforced
Portability tiers are checked in CI on bare-metal targets rather than asserted in a
README, so a no_std crate stays no_std.
Real backends, real silicon
Device support matrix →Five backends run end to end today. Each takes the same device-neutral TOSA 1.0 artifact, owns its native lowering and execution path, and submits directly bound program-visible buffers. “Supported” means the backend admits that program and dtype and executes it through the provider; parsing a graph is not executing one.
Lowers TOSA into Core ML and predicts against buffers the caller already owns, accepting a result only when Core ML wrote into that same allocation. Core ML places each operator on the ANE or the CPU as it sees fit. Direct INT8 boundaries need macOS 26+ and cover identity plus zero-point-aware MATMUL; the INT4 that Core ML advertises is compressed weight storage, not INT4 tensor execution.
Apple Neural Engine + CPUvalidated on Apple M4 / macOS 26.5.2
Program dtypes
- FP32 supported
- FP16 supported
- FP8 not supported
- INT8 supported
- INT4 not supported
Lowers the same artifact into in-memory OpenVINO IR and compiles it per device — NPU first, then GPU, then CPU — under accuracy-preserving execution mode. It reports completion only if the runtime actually wrote into the caller's output allocation, which is a stronger claim than “the call returned”. Its INT8 tier legalizes zero points explicitly in INT32, and an FP8 tier carries both encodings at the model boundary on the same target the Vulkan backend uses — movement bit-exact, MATMUL widened to binary16, which is the accumulator type TOSA already assigns FP8 MATMUL.
Intel NPU, GPU or CPUCPU plugin proven in CI; the FP8 tier compiles and round-trips on an Intel NPU at arch 5010 (Panther Lake), and is withheld on devices whose compiler rejects FP8 — arch 40XX among them — by probing at open rather than trusting an allowlist
Program dtypes
- FP32 supported
- FP16 supported
- FP8 supported as a widened tier, on devices whose compiler accepts FP8
- INT8 supported
- INT4 not supported
BF16 TOSA over the HRX runtime with no XRT userspace dependency, compiling admitted graphs with the pinned aiecc toolchain as a bounded subprocess. There is no native FP32 or FP16 path: FP32 appears only as MATMUL accumulator output, and FP8 is an explicit E4M3/E5M2 → BF16 storage CAST. The INT8 tier adds exact RESCALE.
Ryzen AI NPU — PCI 1022:17f0Strix, Strix Halo and Krackan; validated on Fedora 44 with the in-tree amdxdna driver
Program dtypes
- FP32 accumulator outputs only
- FP16 not supported
- FP8 supported as storage CAST
- INT8 supported
- INT4 not supported
Lowers a strict FP16 and INT8 subset to QNN and runs it on Hexagon HTP with direct client buffers. FP16 covers 41 of the 42 operators Core ML and OpenVINO share — ERF has no QAIRT node. FP32 is refused on purpose: the v73 precision probe caught MATMUL rounding to FP16 even for FLOAT_32 tensors, so admitting it would be a lie about precision.
Hexagon HTP v73Snapdragon X126100 on Windows 11 ARM64
Program dtypes
- FP32 not supported
- FP16 supported
- FP8 not supported
- INT8 supported
- INT4 not supported
Executes the FP32 operator tier shared with Core ML and OpenVINO — 42 TOSA operators with BOOL/INT32 auxiliaries — on crate-authored SPIR-V kernels specialized when the TOSA program loads, one submission per graph. Host, shared and device-local storage buffers bind directly for submission; explicit staging is confined to reads and writes of device-local memory. Fence status provides nonblocking completion without a worker thread. The same 42 operators run over binary16 tensors on every device — the conversions are crate-owned, so the tier needs no device feature. A separate FP8 tier adds TOSA's (FP8, FP8) → FP16 MATMUL over both E4M3 and E5M2, CAST in both directions, MAX_POOL2D and ARGMAX — eleven operators rather than 42, because TOSA defines no FP8 elementwise operator at all — and it too needs no device feature, so FP8 numerics cannot vary by device.
Vulkan 1.3 compute devicesFP32 operator tier validated on Intel Arc 140V (Mesa ANV), AMD Radeon 860M (RADV), Apple M4 (MoltenVK), and continuously verified on Mesa lavapipe; the FP16 tier's corpus passes on Intel Arc LNL (Mesa ANV), AMD Radeon 860M (RADV), Apple M4, and the lavapipe CI lane; the FP8 tier's corpus passes on Intel Arc B390 (Panther Lake, Mesa ANV), a Lunar Lake host (Xe2, Mesa ANV), and Apple M3 (MoltenVK)
Program dtypes
- FP32 supported
- FP16 supported
- FP8 supported as an eleven-operator tier
- INT8 target declared but not advertised
- INT4 not supported
None of these silently dequantize an unsupported graph: unsupported INT8 and packed INT4 graphs are rejected at load, where you can see it. Vulkan advertises the FP32, FP16 and FP8 tiers on every device it opens; the device support matrix records which parts are validated on hardware, which are merely reachable, and which are one named change away; the Hexagon operator matrix has the per-operator detail.
Eighteen crates, one release train
Full crate reference →This project is pre-standardization and experimental. Protocol 1.0 is frozen as a versioned review input for independent implementation — stable enough to build against and to disagree with in writing, not an approved Virtio specification.
The workspace publishes together at 0.3.x. That Cargo version tracks the Rust API, which is still young and expected to change as adapter authors build against it; the protocol freeze is a separate guarantee, so a pre-1.0 crate version is not a statement about protocol stability.