Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Getting started

virtio-accel logo

CI Crates.io docs.rs GitHub last commit License: MIT OR Apache-2.0 MSRV: Rust 1.85+ no_std supported

virtio-accel defines a protocol and contains executable no_std guest, device, transport, queue, and TOSA layers for exposing an accelerator to a guest: contexts, buffers, programs, execution queues, submissions, and events.

▶ Watch the Kerr black-hole demo

Exterior view of the wormhole demo Throat view of the wormhole demo
Demo: NPU-assisted live geodesic ray tracing of a GR wormhole metric

The project claims no Virtio device ID (yet). For guest environments, use the vAccel adapter; see crates/virtio-accel-vaccel/README.md. virtio-accel is currently pre-standardization; protocol 1.0 is frozen as a versioned review input for independent implementation — it is stable enough to build against and to disagree with in writing, not an approved Virtio specification.

Backend support

“Supported” below means that the backend admits the declared program and dtype and exercises it end-to-end; support in the TOSA parser or shared numerical corpus alone does not imply hardware execution. “Not implemented” describes this repository, not necessarily the underlying hardware.

This table is organized by program and dtype. For the physical devices behind it — which parts are validated on hardware, which are merely reachable, and which are one named constant away, including the non-NPU CPU and GPU paths Core ML and OpenVINO already expose — see the device support matrix.

BackendStatusProgram admissionFP32FP16FP8 E4M3/E5M2INT8Packed INT4Program-visible buffers
Apple Core ML / ANE (virtio-accel-coreml)Implemented; macOS 14+Static TOSA 1.0 FP; INT8 tier on macOS 26+SupportedSupportedNot implementedIdentity + MATMULNot implementedDirect host/shared bindings
Intel OpenVINO (virtio-accel-openvino)Implemented; OpenVINO 2026.xStatic TOSA 1.0 FP + INT8 tier; FP8 tierSupportedSupportedWidened tier (ADR 0009)Identity + MATMULNot implementedDirect host/shared bindings
AMD XDNA (virtio-accel-xdna)Experimental; HRX on XDNA2Static BF16 TOSA + explicit FP8 storage CAST + INT8 tierAccumulator outputs onlyNot implementedE4M3/E5M2 → BF16 CASTIdentity + MATMUL + RESCALENot implementedDirect host/shared bindings
Qualcomm Hexagon (virtio-accel-hexagon)Experimental; QAIRT 2.49 on Windows ARM64Static TOSA 1.0 FP16 + BOOL/INT32 auxiliaries; INT8 tierBlocked by v73 precision evidence41/42 shared operators (ERF blocked)Blocked: ambiguous encodingIdentity + MATMULNot implementedDirect host/shared bindings
Vulkan (virtio-accel-vulkan)Experimental; Vulkan 1.3 loaderStatic TOSA 1.0 FP32/FP16 + BOOL/INT32 auxiliaries; FP8 tierSupportedSupported11-operator tier (ADR 0009)Target declared, not advertisedNot implementedDirect host/shared/device bindings

Core ML (Apple Neural Engine)

  • Execution: Core ML selects ANE or CPU placement for each operation. The support table refers to model-boundary dtypes; restricted INT32 outputs are also available.
  • INT8: Direct INT8 model boundaries require macOS 26+ and currently support identity plus zero-point-aware MATMUL.
  • Explicit limits: FP8, unsupported INT8 operators, and packed INT4 graphs are rejected rather than dequantized. Core ML’s INT4 support is compressed-weight storage, not TOSA INT4 execution.

See the virtio-accel-coreml support boundary.

OpenVINO (Intel NPU/GPU/CPU)

  • Execution: The backend compiles separately for each available device—NPU, then GPU, then CPU by default—using OpenVINO’s accuracy-preserving mode. A submission completes only after the runtime writes into the caller’s output allocation.
  • INT8: Direct INT8 model boundaries are supported; MATMUL uses explicit INT32 zero-point legalization. Restricted INT32 outputs are also available.
  • FP8 tier (ADR 0009): both encodings at the model boundary, on the same target the Vulkan tier uses, so a graph admitted by one backend is admitted by the other. FP8 data movement is bit-exact; MATMUL widens its operands to binary16, which is TOSA’s own accumulator type for FP8 MATMUL, so nothing narrows back; MAX_POOL2D widens around the window only. The widening is not a preference — the NPU compiler’s IE dialect declares MatMul operands without the FP8 types, so MLIR’s verifier rejects a raw FP8 MatMul before the hardware is consulted. What is native is the FP8 boundary: parameters stay FP8 through compilation, so nothing converts on the host. The tier is advertised per device, not per backend: FP8 support turned out to be arch-gated — Intel NPU arch 5010 (Panther Lake) compiles FP8 while arch 40XX (Lunar Lake) refuses even an FP8 IDENTITY — so each instance compiles a one-element FP8 graph at open and withholds the descriptor when the device rejects it. There is no property to read instead: OPTIMIZATION_CAPABILITIES omits FP8 on 5010, where it works.
  • Runtime: NPU and GPU require their Intel Level Zero driver or compute runtime. The CPU plugin is exercised in CI.
  • Explicit limits: unsupported INT8 operators and packed INT4 graphs are rejected rather than dequantized. FP8 arithmetic beyond MATMUL is not reachable: TOSA admits no FP8 elementwise operator at all.

See the virtio-accel-openvino support boundary.

Hexagon (Qualcomm Snapdragon X126100/QAIRT 2.49)

  • Evidence scope: The supported configuration is Snapdragon X126100 with QAIRT 2.49 on Windows ARM64.
  • FP16: 41 of the 42 operators shared by Core ML and OpenVINO work. ERF is excluded because QAIRT’s public operation definitions provide no ERF node.
  • INT8: Exact identity and zero-point-aware MATMUL are supported, with INT32 output.
  • Explicit limits: FP32 is rejected because the v73 probe observed FP16-rounded MATMUL even for FLOAT_32 tensors. FP8 is rejected because this client path has no unambiguous E4M3/E5M2 QAIRT selector. A missing complete SDK reports RuntimeUnavailable.

See the virtio-accel-hexagon support boundary and the operator matrix.

XDNA/2 (AMD XDNA2 NPU over HRX runtime)

  • Execution: The backend uses HRX-owned buffers and a serialized dispatch worker; submitted program buffers bind directly, without a submission-time bounce copy.
  • BF16: IDENTITY, BF16→FP32 MATMUL, and NHWC MAX_POOL2D are implemented within documented static shape envelopes. FP32 is exposed only where TOSA requires the MATMUL accumulator output; FP32 and FP16 arithmetic are rejected.
  • FP8 storage: Explicit TOSA CAST from E4M3 or E5M2 to BF16 runs on the NPU and is bit-exact for every finite value. NaN payload canonicalization is permitted. FP8 arithmetic is not advertised.
  • INT8: Exact identity, zero-point-aware MATMUL with INT32 output, and signed per-tensor INT32→INT8 RESCALE are supported. The CONST/IDENTITY/MATMUL baseline mirrors OpenVINO; RESCALE is an intentional released-TOSA expansion needed to turn exact accumulator output back into INT8. XDNA admission adds a one-core static memory envelope and explicit four-byte DMA slot padding where a logical tensor is not word-sized.
  • Runtime: Native execution requires the pinned amdxdna-native HRX runtime and compiler toolchain. Portable admission and offline artifact compilation remain available without a device.

Vulkan (FP32, FP16, and FP8 tiers on any Vulkan 1.3 compute device)

virtio-accel-vulkan is a vendor-neutral Vulkan 1.3 compute backend bound through the pinned ash crate with run-time loader discovery. It admits static single-block TOSA 1.0 graphs over the same 42 FP32-tier operators Core ML and OpenVINO share — elementwise, activation, comparison and logical, reduction, ARGMAX, MATMUL, MAX_POOL2D, and data-movement operators with BOOL and INT32 auxiliaries — and executes a whole graph as one submission: crate-authored SPIR-V kernels specialized at load_program, constants and intermediates in a per-program arena, compute barriers between dependent dispatches. Buffers are dedicated directly bound storage buffers in Host/Shared/Device memory domains; completion is a bounded per-context ring polled through vkGetFenceStatus.

  • Numerics: every float operation is NoContraction; SIN, COS, TANH, and ERF are crate-authored range reductions and polynomials (Payne–Hanek beyond |x| = 8192) rather than the driver’s loosely specified built-ins; NaN modes follow the TOSA pseudocode literally. MATMUL is a register-tiled shared-memory kernel, or a split-k streaming kernel for eight rows or fewer, accumulating in binary32 with fused multiply-add — within a stated bound of the exact sum and deterministic per device (ADR 0011).
  • FP16 tier (ADR 0008): the same 42 operators over binary16 tensors, advertised on every device the backend opens — the tier needs no device feature. Packed binary16 tensors are unpacked and widened to binary32 by crate-owned integer code, the float lanes evaluate in binary32 — the implementation choice TOSA 1.0 §1.10.3 names explicitly, and the correctly rounded binary16 result for ADD/SUB/MUL — and results narrow back through crate-owned round-to-nearest-even code that produces subnormals on every device. NEGATE/ABS are integer sign operations and data movement copies lanes as integers, both exact for every bit pattern; MATMUL and reductions accumulate in binary32 (the accumulator width TOSA assigns FP16). Numerics are bit-identical across devices by construction for every operator but MATMUL, whose fused multiply-add is deterministic per device.
  • FP8 tier (ADR 0009): a separate target, because TOSA gates FP8 on the FP8E4M3 and FP8E5M2 extensions rather than the base floating-point profile, carrying eleven operators: TOSA’s (FP8, FP8) -> FP16 MATMUL over either encoding, CAST in both directions, MAX_POOL2D, ARGMAX, and exact data movement. It is a subset rather than the shared 42 because TOSA admits no FP8 elementwise operator at all — no arithmetic, comparison, selection, reduction or transcendental lane takes FP8. Like the FP16 tier it needs no device feature: packed bytes widen to binary32 by crate-owned integer code — every FP8 value is representable in binary32, so the widening is exact — and nothing writes FP8 except a raw byte copy, so the tier is advertised on every device the backend opens and its numerics cannot vary by device. Float-to-FP8 overflow is this crate’s policy rather than TOSA’s, which leaves it undefined: a value too large for E4M3 becomes NaN, because saturation stays expressible as a CLAMP before the cast while a saturated 448 would be indistinguishable from a genuine one.
  • Constraints: MATMUL and NEGATE admit zero zero-points only, MUL a zero shift, and RESHAPE a constant shape (the TOSA 1.0 CONST-producer forms).
  • Evidence: the shared FP32 operator corpus, the conformance suite, and the kernel-level tests pass on Mesa lavapipe in CI; on 2026-09-06 on Intel Arc 140V (Lunar Lake, Mesa 26.0.8 ANV, Vulkan 1.4.335) and the same host’s llvmpipe (LLVM 21.1.8); on 2026-09-08 on AMD Radeon 860M (Krackan Point, RADV Mesa 26.1.8, Vulkan 1.4.354), which also runs clean under Khronos synchronization validation in every advertised memory domain; and on 2026-09-17 on Apple M4 via MoltenVK 1.4.2. sin/cos/tanh land within 1 ulp and erf within 2 ulp of binary64 on every one. The FP16 corpus — every bit-exact case, the ulp-tolerated groups, an exhaustive 65536-pattern NEGATE round trip, the higher-precision lanes within 1 ulp of the binary64 references over the whole finite binary16 domain, and the subnormal-arithmetic probe — passes on Apple M4 via MoltenVK 1.4.2 and, against the final kernels, is owed the confirmation runs on Intel Arc LNL (Mesa ANV) and AMD Radeon 860M (RADV), whose earlier runs passed the corpus identically. The FP8 corpus — all 256 patterns of both encodings through IDENTITY bit-for-bit, MATMUL against a widened host reference with input and with constant operands, CAST round-tripping every encoding in both directions and honouring the overflow policy, two FP8 matmuls chained through a CAST, and MAX_POOL2D/ARGMAX — passed on 2026-09-18 on Intel Arc B390 (Panther Lake, Mesa 26.0.8 ANV), independently on a Lunar Lake host (Xe2, Mesa ANV), and on an Apple M3 via MoltenVK, in every advertised memory domain with identical results; all 170 assembled kernel variants pass spirv-val --target-env vulkan1.3. INT8 gating remains under the Vulkan wayfinder map; design decisions are recorded in docs/adr/ (ADR 0007 covers the FP32 tier, ADR 0008 the FP16 tier, ADR 0009 the FP8 tier).

TOSA 1.0

Independently of backend execution, virtio-accel-tosa validates the TOSA 1.0 profiles and extensions for all five dtype columns, and virtio-accel-conformance ships shared fixtures and oracles for them. virtio-accel-tosa-build provides matching borrowed and incrementally owned safe authoring paths for static single-block graphs and validates every result through that ingestion boundary. The byte-oriented virtio-accel-mock backend remains test infrastructure rather than a typed hardware implementation.

Workspace

CrateTierDescription
virtio-accel-vaccelcoreAdapter seam for mapping native provider contracts (including vAccel-style backends) to virtio-accel-core
virtio-accel-coremlstdTOSA-to-Core ML lowering, direct buffers, and asynchronous ANE-capable prediction
virtio-accel-openvinostdTOSA-to-OpenVINO IR lowering, direct host-pointer tensors, and asynchronous NPU/GPU/CPU inference
virtio-accel-vulkanstdVendor-neutral Vulkan 1.3 compute backend over ash: crate-authored SPIR-V kernels, direct storage-buffer binding, and the shared FP32 plus native FP16 and FP8 operator tiers
virtio-accel-xdnastdAMD XDNA2 NPU backend over HRX with direct buffers, asynchronous dispatch, and strict BF16/FP8/INT8 TOSA tiers
virtio-accel-hexagonstd (Windows ARM64)Strict FP16/INT8 TOSA-to-QNN lowering, direct buffers, and asynchronous Hexagon HTP execution
virtio-accelcore + allocFacade re-exporting the portable layers
virtio-accel-protocorePointer-free, little-endian protocol 1.0 wire structures
virtio-accel-transportcoreDependency-free descriptor-chain, queue, reset, and notification ports
virtio-accel-corecoreBackend lifecycle, memory, program, queue, and event contracts
virtio-accel-tosacore + allocBounded zero-copy TOSA 1.0 validation, lowering analysis, specialization, and packed low-precision utilities
virtio-accel-tosa-buildcore + allocBorrowed and incrementally owned static TOSA 1.0 authoring with mandatory validation round trips
virtio-accel-split-queuecore + allocBounded in-memory split-ring reference model
virtio-accel-guestcore + allocTyped reference client with bounded request tracking
virtio-accel-devicecore + allocDevice-owned state, including bounded generational IDs
virtio-accel-mockstdIn-memory backend with deterministic test-only artifacts and scripted faults
virtio-accel-conformancestdTransport-free semantic suite and shared FP32/FP16/FP8/INT8/INT4 numerical TOSA corpus
virtio-accel-cleanroomcoreIndependent conformance codec, written without the shared protocol types

Dependency graph

virtio-accel-split-queue ---> virtio-accel-transport
                                      ^
                                      |
virtio-accel-device ----------+-------+------> virtio-accel-core
          |
          +-----> virtio-accel-proto

virtio-accel-guest -----------> virtio-accel-transport
          |
          +--------------------> virtio-accel-proto

virtio-accel-conformance --------------------> virtio-accel-core
virtio-accel-tosa ---------------------------> virtio-accel-core
virtio-accel-tosa-build ---------------------> virtio-accel-tosa
virtio-accel-xdna ---------+--------------> virtio-accel-core
                              |
                              +--------------> virtio-accel-tosa
virtio-accel-coreml ----------+--------------> virtio-accel-core
                              |
                              +--------------> virtio-accel-tosa
virtio-accel-openvino --------+--------------> virtio-accel-core
                              |
                              +--------------> virtio-accel-tosa
virtio-accel-vulkan ----------+--------------> virtio-accel-core
                              |
                              +--------------> virtio-accel-tosa
virtio-accel-hexagon ---------+--------------> virtio-accel-core
                              |
                              +--------------> virtio-accel-tosa
virtio-accel-vaccel -----------------------> virtio-accel-core
other provider adapters --------------------> virtio-accel-core

The transport crate exposes reset-scoped chain identities, flattened direction/length metadata, and owned publication/completion tokens. Neither it nor the device-state layer leaks guest addresses, ring pointers, or concrete descriptor types into the command engine or provider backend.

Install

[dependencies]
virtio-accel = "0.4"

The facade is no_std. Add the reference backend as a dev-dependency to run the example below:

[dev-dependencies]
virtio-accel-mock = "0.4"

Host backends are separate dependencies and are never re-exported by the portable facade. Add virtio-accel-coreml = "0.4" on an ANE-capable Mac, virtio-accel-openvino = "0.4" on a host with OpenVINO 2026.x, or virtio-accel-vulkan = "0.4" on a supported host with a Vulkan 1.3 loader and compute device. Each accepts the device-neutral TOSA 1.0 program format and owns its provider-specific validation, lowering, and execution path.

For portable adapter-boundary validation while the native vAccel path is wired, add virtio-accel-vaccel = "0.4". The crate exposes a vAccel seam with an in-repo representative conformance recipe and explicit copy-path diagnostics.

Adapter profiles

  • Portable-only profile: use virtio-accel (+ virtio-accel-mock) to keep all portable layers and conformance fixtures inside the workspace.
  • Adapter profile: add virtio-accel-vaccel when you need an adapter seam for native/vAccel-like implementations that still re-export the Accelerator contract from virtio-accel-core.
  • Production host profile: add virtio-accel-coreml, virtio-accel-openvino, virtio-accel-hexagon, virtio-accel-xdna, and/or virtio-accel-vulkan instead of any mock backend once provider licensing and native runtime availability are in place.

virtio-accel-hexagon = "0.4" exposes the separate Qualcomm adapter. A complete QAIRT/QNN SDK on Windows ARM64 enables its HTP backend; SDK-free builds validate its strict FP16 graph planner and constructors return RuntimeUnavailable.

virtio-accel-vulkan = "0.4" loads the platform Vulkan loader at run time. It admits the shared FP32 and FP16 operator tiers (with BOOL/INT32 auxiliaries) and a separate eleven-operator FP8 tier. It returns RuntimeUnavailable or DeviceUnavailable when no suitable Vulkan 1.3 compute path exists; it does not silently fall back to the mock backend.

Add virtio-accel-tosa = "0.4" separately to validate TOSA 1.0 artifacts, inspect safe borrowed graph and typed-attribute views, enforce complete stable-op semantics for a declared target, and construct the device-neutral TOSA artifact envelope. Model::analyze_for also produces bounded dense IDs, topological order, liveness, runtime obligations, and specialization keys for Core ML, OpenVINO, or another provider. It is intentionally not re-exported by the facade.

Add virtio-accel-tosa-build = "0.4" to produce static single-block TOSA artifacts through typed tensor and operator definitions. Borrowed definitions suit graph literals; owned definitions let compiler frontends assemble runtime-discovered metadata without a parallel owned-to-borrowed adapter, while existing constant storage can remain borrowed. Both surfaces pass the same parser and target validator providers use at admission.

Production backend examples

On macOS 14+ with an accessible Apple Neural Engine, the backend-local example sends a TOSA 1.0 IDENTITY graph through the real lowering, compilation, direct-binding, asynchronous prediction, and teardown path:

cargo run -p virtio-accel-coreml --example tosa_coreml
TOSA -> Core ML -> ANE-capable result: 3.25

On a Linux host with an OpenVINO 2026.x runtime, the equivalent backend-local example executes the same graph on the preferred available Intel inference device (NPU, then GPU, then CPU):

cargo run -p virtio-accel-openvino --example tosa_openvino
TOSA -> OpenVINO -> CPU result: 3.25

With the documented QAIRT environment, the Qualcomm adapter’s example executes FP16 identity on HTP and verifies the shared numerical oracle. SDK-free builds fail explicitly without a CPU/GPU fallback:

cargo run -p virtio-accel-hexagon --example tosa_hexagon
cargo run -p virtio-accel-hexagon --example mock_classifier

On any supported host with a Vulkan 1.3 loader and compute-capable device, the Vulkan example runs the FP32 identity artifact on the preferred device (discrete, integrated, virtual, then CPU). It reports a clean skip when no loader or device is available:

cargo run -p virtio-accel-vulkan --example tosa_vulkan

The portable facade, device engine, transport, and guest layers see only the TOSA artifact format, target identity, and opaque bytes. Provider objects, shaders, generated models, native runtimes, and FFI bridges remain owned by their adapter crates.

Portable lifecycle example

A full submission against the in-memory reference backend — allocate a buffer, load an artifact, bind it to a slot, submit, and observe the event:

#![allow(unused)]
fn main() {
use virtio_accel::core::{
    Accelerator, AccessMode, ArtifactRef, BindingRef, BufferDesc, BufferRange, BufferUsage,
    ContextDesc, EventState, MemoryDomain, QueueDesc, SubmitFailure, Timeout,
};
use virtio_accel_mock::{MockAccelerator, reference};

let backend = MockAccelerator::default();
let context = backend.create_context(ContextDesc::default())?;

// An 8-byte shared buffer the program may read and write.
let desc = BufferDesc::new(
    8,
    8,
    MemoryDomain::Shared,
    BufferUsage::TRANSFER_SOURCE
        | BufferUsage::TRANSFER_DESTINATION
        | BufferUsage::PROGRAM_INPUT
        | BufferUsage::PROGRAM_OUTPUT
        | BufferUsage::MUTABLE_STATE,
)?;
let (mut buffer, _) = backend.allocate_buffer(&context, desc)?.into_parts();
backend.write_buffer(&mut buffer, 0, &[0x00, 0x11, 0x7f, 0x80, 0xa5, 0xff, 0x3c, 0xc3])?;

// A deterministic test-only artifact: XOR every byte bound to slot 7 with 0x5a.
let artifact = reference::ReferenceArtifact::xor(7, 0x5a);
let program = backend.load_program(
    &context,
    ArtifactRef {
        format: reference::ARTIFACT_FORMAT,
        target: reference::TARGET_IDENTITY,
        payload: artifact.as_bytes(),
        resident_bytes: reference::RESIDENT_BYTES,
    },
)?;
let queue = backend.create_queue(&context, QueueDesc::default())?;

let bindings = [BindingRef {
    slot: 7,
    buffer: &buffer,
    range: BufferRange::new(0, 8)?,
    access: AccessMode::ReadWrite,
}];

// Submission is asynchronous at the ownership boundary, so it always yields an event.
let event = backend
    .submit(&queue, &program, &bindings, Timeout::Infinite)
    .map_err(|failure| match failure {
        SubmitFailure::Rejected(error) | SubmitFailure::Indeterminate { error, .. } => error,
    })?;
assert_eq!(backend.poll_event(&event)?, EventState::Pending);

// The mock backend runs under harness control, so the caller drives completion.
backend.complete(&event)?;
assert_eq!(backend.poll_event(&event)?, EventState::Complete);

let mut output = [0_u8; 8];
backend.read_buffer(&buffer, 0, &mut output)?;
assert_eq!(output, [0x5a, 0x4b, 0x25, 0xda, 0xff, 0xa5, 0x66, 0x99]);
}

Every object is released explicitly, and a release can itself fail; see examples/reference_execution.rs for the teardown path.

cargo run --example reference_execution

Protocol 1.0

The protocol defines fixed headers and payloads for device discovery, contexts, buffers, programs, execution queues, submissions, and events. Two properties shape most of the API:

  1. Unknown values stay raw. Unrecognized opcodes, statuses, and event states remain integers until validated, so decoding untrusted bytes never constructs an invalid Rust enum.
  2. Failure still returns an event. A successful submit returns an event; an indeterminate failure must also return one, because the operation’s resources are still owned by the device. Guest-visible object IDs are opaque, kind-tagged, generational, and never reused after generation exhaustion.

The primary zerocopy ABI and the manual clean-room codec both decode and re-encode every canonical frame. Their bridge test exchanges bytes only, providing an independent implementation check without making the conformance codec a production dependency.

Non-Rust device and driver implementations can include include/virtio_accel.h. The header is a packed C projection of the wire contract, not a host backend plugin ABI. CI compiles it as C11 and C++11 and derives constant, size, alignment, and offset assertions from the frozen layout manifest.

Writing a backend

Implement the Accelerator contract from virtio-accel-core, then run the standard semantic suite against it. The suite is transport-free: no wire format, virtqueue, OS, or vendor dependency.

cargo run --example backend_conformance
memory.shared: Passed
buffer.transfer-permissions: Passed
submission.context-isolation: Passed
event.cancellation-races: Passed
accounting.resource-lifecycle: Passed
...

The backend implementer guide walks through the hooks, the optional resource-accounting and progress adapters, and the fault-injection harness.

Documentation

DocumentCovers
specification.mdNormative terminology, object model, compatibility rules, mandatory baseline
wire-abi.mdExact byte layouts and the coordinated change procedure
virtio_accel.hChecked C and C++ projection of the protocol 1.0 wire contract
virtqueue.mdCommand-chain rules
architecture.mdImplementation invariants
threat-model.mdTrust boundaries and finite resource policy
portability.mdEnforced target matrix and crate tiers
device-support-matrix.mdPhysical devices each backend reaches, and the named gate holding back the rest
performance.mdv1 performance and copy budgets
public-api.mdPublic rustdoc policy
release-policy.mdRelease governance and evolution rules
backend-implementer-guide.mdRunning the semantic suite against a new backend
releases/v1.0.mdProtocol 1.0 release note
conformance/v1.0Golden artifacts, canonical frames, and the freeze audit
CONTRIBUTING.mdDevelopment gates, protocol change classification, and scope boundaries
CODE_OF_CONDUCT.mdExpected conduct in project spaces
SECURITY.mdReporting a vulnerability

Portability

Project-authored portable and reference code forbids or denies unsafe code. The TOSA crate confines official generated FlatBuffers accessors to a private module behind bounded verification, while the five host backends isolate audited native FFI behind their host/runtime gates. CI enforces each portability tier, including compile-only checks of every adapter’s unsupported-platform surface.

TierAllowed runtime surface
corecore only; no allocation
core + alloccore + alloc; no OS, filesystem, sockets, threads, or host synchronization
stdPortable std; no host-OS or vendor-specific API
host-nativeCore ML, OpenVINO, QNN, HRX, or Vulkan behind an adapter-specific runtime gate

Concrete VMM, kernel, OS, and vendor adapters do not change the portable v1 protocol and must not become default dependencies of a portable crate. Cargo features must be additive: disabling default features may remove convenience behavior, but must never select a different protocol interpretation.

Development

Minimum supported Rust version is 1.85 (edition 2024), checked in CI.

cargo fmt --all -- --check
python3 ci/check-release-policy.py
cargo clippy --workspace --all-targets --all-features -- -D warnings
cargo test --workspace --all-targets --all-features
cargo run --example backend_conformance
cargo run --example reference_execution
cargo run -p virtio-accel-coreml --example tosa_coreml # macOS 14+ with ANE
cargo run -p virtio-accel-openvino --example tosa_openvino # Linux with OpenVINO 2026.x
cargo run -p virtio-accel-hexagon --example tosa_hexagon # Windows ARM64 with the documented QAIRT setup
cargo run -p virtio-accel-hexagon --example mock_classifier # FP16 linear classifier on Hexagon HTP
cargo run -p virtio-accel-vulkan --example tosa_vulkan # Vulkan 1.3 loader and compute device
python3 ci/publish-dry-run.py

Target checks need the corresponding standard libraries:

rustup target add aarch64-unknown-none riscv64gc-unknown-none-elf wasm32-unknown-unknown

Status

Included in protocol 1.0:

  • one command virtqueue at index zero
  • device discovery and exact protocol compatibility checks
  • contexts, buffers, opaque programs, execution queues, submissions, and events
  • bounded explicit buffer transfers
  • event polling, optional cancellation, release, reset, and backend-discard recovery
  • direct-binding requirements for program-visible buffers
  • checked finite limits for untrusted byte counts, descriptor counts, object counts, and retained backend storage
  • an independent clean-room codec and a transport-free semantic conformance suite

Reserved and unadvertised — an implementation that advertises one of these is not 1.0 conformant until a future version assigns its negotiation, ownership, synchronization, and conformance rules:

  • multi-queue and event queues
  • external memory import/export
  • timeline fences
  • secure contexts
  • packed virtqueues
  • protocol-level negotiation for additional VMM, kernel, OS, and vendor integrations
  • a standardized graph IR, compiler, or executable format

Protocol 1.0 numeric opcodes, statuses, and payload layouts are frozen for the portable v1.0 baseline by the final freeze audit. Future changes must follow the coordinated change procedure in wire-abi.md and the release and evolution policy; incompatible changes require a new protocol major version.

Contributing

Contributions are welcome, including disagreement with frozen decisions — a reasoned objection is worth more than a workaround built on top of one. See CONTRIBUTING.md for the local gates, the scope boundaries, and how wire changes are classified before code is merged.

Citation

If virtio-accel supports your work, use GitHub’s Cite this repository control. The canonical citation metadata is in CITATION.cff.

License

Licensed under either of Apache License, Version 2.0 or MIT license at your option.

Contributions are dual-licensed on the same terms, with no separate CLA.

virtio-accel portable protocol foundations

Status: frozen portable protocol 1.0 candidate. The final freeze audit is recorded in conformance/v1.0/freeze-audit.md.

This document defines the portable semantic foundation for virtio-accel. It does not assign a Virtio device ID and is not an OASIS Virtio specification. The transport model is intended to remain compatible with the general facilities defined by the Virtio specification.

Sections 1 through 11, wire-abi.md, and virtqueue.md are normative for the portable protocol 1.0 candidate. Appendix A maps the current Rust surface to the terms in this document and is non-normative implementation guidance.

1. Normative language

The key words MUST, MUST NOT, REQUIRED, SHOULD, SHOULD NOT, MAY, and OPTIONAL are to be interpreted as described by RFC 2119 and RFC 8174 when, and only when, they appear in bold capitals.

Explanatory paragraphs beginning with “Rationale” are non-normative.

2. Scope

The portable v1 contract defines:

  • device discovery and protocol compatibility;
  • contexts and context-scoped resource ownership;
  • buffers and bounded byte transfers;
  • opaque program artifacts;
  • accelerator execution queues;
  • submissions, bindings, relative timeouts, and asynchronous events;
  • cancellation when supported;
  • explicit destruction, reset, and failure recovery; and
  • a single baseline command virtqueue carrying request and response frames.

The portable v1 contract deliberately does not define:

  • a standardized accelerator graph, compiler IR, or executable format;
  • an OASIS device ID or standards submission;
  • PCI, MMIO, vhost-user, QEMU, kernel, DRM, VFIO, or hypervisor integration;
  • operating-system or vendor accelerator APIs;
  • DMA-BUF, platform shared handles, zero-copy external memory, cache-coherency protocols, or timeline fences;
  • packed virtqueues; or
  • process, thread, executor, or async-runtime policy.

Adding an integration listed above MUST NOT change the portable object lifecycle by convention. If it changes driver/device-visible behavior, it requires a negotiated feature or a new protocol version.

3. Roles and boundaries

3.1 Driver

The driver is the request initiator. It discovers the device, validates protocol compatibility, negotiates Virtio features, constructs request frames, publishes descriptor chains, validates responses, and owns guest-side request tracking.

The driver MUST treat all device-written bytes as untrusted.

3.2 Device

The device owns the guest-visible protocol state. It validates descriptor direction and frame bytes, maps opaque object IDs to live resources, enforces limits and context isolation, invokes an accelerator backend, writes responses, and performs reset.

The device MUST treat all driver-written bytes, lengths, counts, object IDs, flags, and timing as untrusted.

3.3 Transport adapter

A transport adapter maps a concrete Virtio implementation to portable readable/writable regions, queue notification, used-length reporting, and reset operations.

A transport adapter MUST NOT expose guest addresses, descriptor objects, host mappings, or transport-specific synchronization types to the accelerator backend.

3.4 Command engine

The command engine is the transport-neutral device state machine. It decodes one validated request, performs object and quota checks, invokes the backend, updates ownership state, and produces one response.

The command engine MUST NOT depend on an operating system, VMM, kernel, guest-memory crate, vendor API, or compiler API.

3.5 Accelerator backend

The accelerator backend implements device-local execution semantics using provider-owned native handles. It is represented by virtio_accel_core::Accelerator.

A backend MUST NOT receive wire structures, object IDs, guest addresses, or virtqueue descriptors. It receives validated semantic values and provider-owned handles.

4. Terminology and object model

4.1 Device instance

A device instance is one initialized protocol endpoint and its backend. It owns:

  • one device identity and set of semantic capabilities;
  • advertised limits;
  • one object-ID namespace;
  • zero or more contexts; and
  • one reset epoch.

Object IDs are meaningful only within the device instance and reset epoch that created them.

4.2 Command virtqueue

The command virtqueue is the Virtio transport queue that carries request and response frames. Queue index zero is the only command virtqueue in the mandatory baseline.

The command virtqueue is not an accelerator execution queue.

4.3 Context

A context is the isolation and ownership parent for buffers, programs, accelerator execution queues, and the events created by its submissions.

A device MUST reject an operation that combines objects from different contexts before invoking the backend. A context MUST NOT be destroyed while any child object or in-flight reference remains live.

4.4 Buffer

A buffer is a bounded backend allocation with:

  • a nonzero byte length;
  • a nonzero power-of-two alignment;
  • a memory domain; and
  • declared usage flags.

The buffer’s guest-visible object ID is not an address. Every transfer and binding range MUST fit within the buffer without integer overflow.

Memory domains are strict allocation requirements:

  • Host requests provider memory optimized for host transfers;
  • Device requests the provider’s accelerator-local placement class; and
  • Shared requests one provider-owned allocation that is both host visible and directly bindable by the accelerator.

Shared does not mean guest-memory import, cross-process export, a platform shared handle, or implicit cache coherence. Those operations require the reserved external-memory feature and have no protocol 1.0 semantics.

An allocation result MUST report the requested descriptor, actual retained allocation bytes, actual guaranteed alignment, and honest backing properties. Actual bytes MUST NOT be smaller than the logical buffer and actual alignment MUST NOT be weaker than requested. Device resource accounting uses actual retained bytes rather than assuming logical bytes include provider padding. A provider MUST NOT return an allocation whose reported placement can be achieved only by copying through another full-size allocation during submission.

A buffer whose usage includes program input, program output, or mutable state MUST report direct binding. A compatible submission MUST bind the exact provider allocation without copying the bound range into or out of another allocation. If the provider cannot satisfy that invariant, it MUST reject allocation. If a particular program is incompatible with an otherwise valid buffer, submission MUST be rejected as INCOMPATIBLE; the provider MUST NOT silently stage it.

WRITE_BUFFER and READ_BUFFER are the only baseline operations that explicitly transfer buffer contents. A write requires TRANSFER_DESTINATION; a read requires TRANSFER_SOURCE. Device-local memory may use bounded provider staging during these explicit transfers, but no transfer may retain the request or response byte region after the backend call returns.

The semantic transfer API accepts bounded byte sources and sinks that may be segmented. This avoids requiring the command engine to allocate a contiguous copy of a valid transfer payload or response before calling the backend. A contiguous region remains available to providers as an optional fast path.

4.5 Program

A program is a resident backend object created from an opaque artifact format, opaque target identity, payload, and caller-authorized upper bound on retained resident bytes.

The transport MUST NOT interpret vendor artifact contents. Artifact-format adapters own their validation beyond the portable envelope fields.

The semantic artifact payload is a bounded byte source rather than a required contiguous slice. Providers may inspect an available contiguous view or read segmented payload bytes directly into final resident storage.

4.6 Accelerator execution queue

An accelerator execution queue is a context-scoped backend object used to submit programs for execution. It is created by the protocol CreateQueue command and represented by Accelerator::Queue.

It is unrelated to a Virtio queue index. Documentation and APIs SHOULD use “execution queue” whenever omitting “accelerator” would make the distinction unclear.

4.7 Binding

A binding associates one program slot with a nonempty buffer range and an intended access mode for one submission.

Binding slot numbers MUST be unique within a submission. A binding does not transfer ownership of its buffer, but the device MUST retain the buffer until the resulting event is safely reclaimed.

Read access requires a buffer declared for program input or mutable state. Write access requires a buffer declared for program output or mutable state. Read-write access requires mutable state. Binding-array order has no semantic meaning; the slot number identifies the program argument.

4.8 Submission

A submission is an attempt to admit one program, execution queue, binding list, and relative timeout to the backend.

Admission has an explicit acceptance boundary:

  • rejected means the backend guarantees that execution was not accepted and no operation resources were retained;
  • accepted returns an event; and
  • indeterminate means acceptance cannot be established and therefore also returns an event that owns the retained resources.

The command engine MUST NOT convert an indeterminate submission into an ordinary error.

4.9 Event

An event is the ownership and completion token for one accepted or indeterminate submission. Its state is pending, complete, failed, or cancelled.

An event MUST retain its execution queue, program, buffers, and per-invocation backend state until it is terminal and successfully destroyed. A pending event MUST NOT be destroyed.

Terminal event states are stable. Cancellation and completion select exactly one terminal result: if cancellation wins, the cancellation command succeeds and polling reports cancelled; if completion wins, cancellation returns BUSY and polling reports the completed or failed result.

4.10 Request and response

A request is one driver-readable command frame identified by a request ID. A response is the corresponding device-written frame with the same request ID.

Request IDs correlate command completion only. They are not object IDs or execution event IDs.

4.11 Reset

A reset stops admission, disposes or quarantines in-flight state according to its known ownership, invalidates every guest-visible object ID, advances the device epoch, resets command queues, and returns the device to its initial negotiation state.

The transport MUST stop fetching command chains and publishing completions before portable object teardown begins. Teardown MUST be bounded and child-before-parent: events precede their execution queues, programs, and buffers, and all context children precede the context.

The device MAY reuse a backend instance only when every known resource is released successfully. An unresolved pending event, a rejected reset-time release, an indeterminate release, backend device loss, or an accounting contradiction requires discarding the complete backend instance. Once discard is required, repeated reset attempts MUST NOT invoke that backend again.

Successful reinitialization MUST use a fresh nonzero object namespace. No object ID created before reset may resolve after reset.

4.12 Host resource policy and hostile input

Every guest-controlled byte length, object count, array count, descriptor count, queue occupancy, and in-flight reference count MUST be checked against an authoritative nonzero bound before it can drive allocation, iteration, or provider invocation. Checked arithmetic MUST reject a value whose derived size is not representable. Wire framing bounds belong to WireConfig, provider capability and per-object bounds belong to DeviceLimits, and aggregate provider-retained storage bounds belong to a host-supplied ResourcePolicy; an implementation MUST NOT use conflicting duplicate limits for the same boundary.

A device integration MUST provide nonzero aggregate limits for actual buffer backing bytes and declared program resident bytes. The command engine MUST charge the actual BufferInfo::allocation_bytes returned by the provider and the ArtifactRef::resident_bytes authorized before program loading. A provider MUST NOT retain program storage attributable to a returned program handle beyond that program’s resident-byte charge.

A retained-byte charge begins before the associated object ID is published and MUST remain accounted while the provider handle is live or releasing. Rejected release MUST restore the charge with the same live object. Indeterminate release MUST transfer the charge to quarantine and require discard of the complete backend instance.

If actual buffer backing would exceed aggregate policy, the device MUST NOT expose the new object ID. It MUST attempt release through the provider ownership boundary; unless successful release is proven, the device MUST stop admission and require backend discard rather than report ordinary resource exhaustion.

Portable parsing, lookup, queue, polling, cancellation, and reset work MUST terminate after a validated finite amount of local work and MUST NOT busy-wait for guest or provider progress. Malformed guest input and configured resource exhaustion MUST NOT cause a panic or expose uninitialized response bytes. Repeated-command rate limits, guest-memory pinning budgets, and provider-hang watchdogs belong to the transport or host integration because the portable layer has no clock, scheduler, tenant, or safe way to preempt arbitrary provider code.

5. Baseline capabilities and feature policy

5.1 Mandatory baseline

The portable v1 baseline MUST provide:

  • command virtqueue zero;
  • fixed-width little-endian request and response headers;
  • device information;
  • context creation and destruction;
  • buffer allocation, destruction, and bounded read/write transfer;
  • opaque program loading and destruction;
  • accelerator execution queue creation and destruction;
  • submission with bounded bindings and relative timeout;
  • event polling and destruction;
  • stale-object and cross-context rejection; and
  • whole-device reset.

The command opcode for cancellation is part of the baseline namespace. If the backend does not advertise semantic event-cancellation capability, the device MUST return UNSUPPORTED without changing the event.

5.2 Transport feature bits

Transport feature bits alter driver/device protocol behavior. The mandatory baseline has no device-specific transport features: BASELINE_FEATURES is empty.

The currently reserved bits MULTI_QUEUE, EVENT_QUEUE, EXTERNAL_MEMORY, TIMELINE_FENCES, and SECURE_CONTEXTS reserve numeric positions only. A baseline device MUST NOT advertise them and a baseline driver MUST NOT accept them. Protocol 1.0 assigns no semantics to these positions.

Unknown device-specific feature bits MUST NOT be accepted by a driver. A device MUST NOT offer a feature it cannot honor if accepted.

5.3 Semantic backend capabilities

Backend Capabilities report whether a semantic operation or resource class is supported. They do not independently change wire framing.

Protocol 1.0 assigns these semantic capability bits:

BitNameMeaning
0HOST_VISIBLE_MEMORYMemoryDomain::Host allocation is supported
1DEVICE_LOCAL_MEMORYMemoryDomain::Device allocation is supported
2EVENT_CANCELLATIONCANCEL_EVENT is implemented for pending events
5SHARED_MEMORYProvider-owned MemoryDomain::Shared allocation is supported

Semantic bits 3 (EXTERNAL_MEMORY) and 4 (SECURE_CONTEXTS) are reserved until their transport, ownership, synchronization, and isolation rules are specified. A protocol 1.0 device MUST NOT advertise either bit.

The device MUST reject an allocation for an unsupported memory domain before invoking the backend. Advertising a memory-domain capability commits the backend to the corresponding allocation properties and direct-binding rules from section 4.4; capability reporting is not permission to substitute a staged implementation.

A baseline backend MUST advertise at least one assigned memory-domain capability. Every resource-count, binding-count, and byte limit in DeviceInfo MUST be nonzero. Identity, capabilities, and limits MUST remain stable for the lifetime of one backend instance so the command engine can validate and cache them once before constructing object state.

If EVENT_CANCELLATION is absent, the command engine MUST reject cancellation before backend invocation. If it is present, the backend MUST implement cancellation and MUST NOT return UNSUPPORTED for a valid pending event.

If enabling a backend capability would require different descriptor direction, additional queues, new synchronization, or changed lifetime rules, the device MUST also negotiate an appropriate transport feature.

ContextFlags::SECURE and QueueFlags::IN_ORDER reserve semantic API positions only. Protocol 1.0 defines neither behavior and accepts only empty context and execution-queue flags. The command engine MUST reject nonempty values before backend invocation; a direct caller of the backend trait receives UNSUPPORTED without a retained resource. Enabling either reservation in a later version requires explicit ownership, ordering, synchronization, and negotiation rules.

5.4 Request and object flags

All request-header flags are zero in protocol 1.0. The former draft NO_WAIT position, bit zero, is reserved and MUST be zero.

Unknown request, context, execution-queue, submission, buffer-usage, or binding-access flag bits MUST be rejected before backend invocation. Reserved fields MUST be zero.

6. Versioning and compatibility

6.1 Candidate version

This specification assigns version 1.0 to the candidate exercised by the conformance artifacts. Candidate implementations MUST expose major 1 and minor 0 in device-specific configuration and MUST NOT claim candidate protocol 1.0 conformance unless they satisfy the normative wire, queue, lifecycle, and compatibility requirements.

Protocol 1.0 compatibility is frozen by the final audit recorded in conformance/v1.0/freeze-audit.md. A future erratum, extension, or incompatible change MUST follow the coordinated procedure in wire-abi.md and the classification rules in release-policy.md.

6.2 Stable major versions

For protocol major one and later:

  • a driver MUST reject a device with an unsupported major version;
  • a device minor version is backward compatible with all earlier minor versions of the same major;
  • a driver MUST NOT use behavior newer than the device minor version it observed; and
  • feature negotiation remains required even when both endpoints know a feature’s numeric bit.

A major version changes when compatibility cannot be preserved through a new opcode, a new negotiated feature, or a previously reserved value.

6.3 Extending existing frames

The config space does not communicate the driver’s supported minor version back to the device. Therefore a minor version alone MUST NOT append fields to an existing response that an older driver would receive.

Within a stable major version, an existing request or response payload length is immutable unless a negotiated feature explicitly selects a different layout. Extensions SHOULD use a new opcode or feature-gated payload.

Without such a feature, a receiver MUST require the exact payload length for the opcode and MUST reject trailing bytes. It MUST NOT silently reinterpret or ignore an unknown tail.

6.4 Unknown numeric values

  • An unknown opcode produces UNSUPPORTED and no semantic state change.
  • An unknown request or object flag bit produces UNSUPPORTED and no backend call.
  • An unknown object ID or zero object ID produces the specified invalid/stale-object failure.
  • An unknown response status is an opaque failure. A driver MUST NOT treat it as success or construct an invalid Rust enum value.
  • Unknown accelerator classes and provider-owned artifact formats remain representable as opaque numeric values, but using them may produce UNSUPPORTED or INCOMPATIBLE.

7. Ownership and destruction

The device owns the mapping from opaque object IDs to backend handles. The mapping MUST reject:

  • zero IDs;
  • stale generations;
  • wrong resource kinds;
  • IDs from another context; and
  • IDs invalidated by reset.

Object-ID encoding is implementation-private. Drivers MUST NOT infer slot, kind, generation, or address information from an ID.

Destructive backend calls consume handles and have an explicit release boundary:

  • a rejected release returns the still-live handle and the device MUST restore it for retry;
  • a successful release removes the object permanently; and
  • an indeterminate release invalidates the guest ID and requires recovery. The device MUST NOT reuse or free the handle based on an assumption.

Drop may provide defensive cleanup inside an implementation but MUST NOT be used as guest-visible protocol state.

A creation method that returns an ordinary error MUST NOT retain a newly created resource. A failed explicit write may have modified part of its requested buffer range; that complete range is then unspecified, the buffer remains live, and a later successful full-range write reestablishes its contents. A failed explicit read may have partially initialized its destination but does not modify the buffer; the command engine MUST NOT publish that destination as a successful payload.

8. Time and progress

Wire timeouts are relative nanosecond durations measured from backend admission. Zero means infinite. Absolute guest timestamps MUST NOT be compared with a host clock because their monotonic epochs are unrelated.

Polling an event MUST be nonblocking and bounded. The portable contract creates no background thread and selects no async executor.

Discovery, context, allocation, transfer, program, execution-queue, and release methods may perform synchronous provider work. Submission performs validation and admission only and MUST NOT wait for execution to become terminal. Cancellation MUST also be nonblocking and bounded. No portable method requires a provider handle to contain a lock or atomic value; cross-thread calls are available only when the concrete backend and handle types opt into the corresponding Rust auto traits.

A timeout before admission is rejected. A timeout or communication failure after admission that cannot prove rejection is indeterminate and MUST retain an event.

9. Errors

The portable status namespace distinguishes unsupported behavior, incompatibility, invalid arguments, out-of-bounds access, busy resources, host allocation failure, configured resource-limit exhaustion, deadline expiration, device loss, permission failure, stale objects, and internal errors.

A device MUST map errors deterministically and MUST NOT expose uninitialized response bytes. Provider-specific External domains and codes are not part of the baseline wire ABI; absent a future negotiated diagnostic extension, they map to INTERNAL_ERROR.

An error response MUST NOT imply that a backend operation was rejected when acceptance was indeterminate.

10. Portability requirements

Portable crates MUST NOT select an operating system, VMM, kernel, vendor API, or global runtime.

  • virtio-accel-proto, virtio-accel-transport, and virtio-accel-core require neither std nor alloc.
  • virtio-accel-device and the future reference guest/queue layers may require alloc but not std.
  • reference host tooling and mock backends may require std.

Platform adapters depend inward on the portable layers. A portable layer MUST NOT conditionally expose different semantics by host OS.

11. Release governance and audit trail

This document, wire-abi.md, and virtqueue.md define the frozen portable driver/device protocol 1.0 baseline. Post-freeze change classification, Cargo semver, feature, MSRV, target, unsafe-code, dependency, license, and publish-metadata policy are documented in release-policy.md.

The versioned audit result is conformance/v1.0/freeze-audit.md and the release note is docs/releases/v1.0.md. Future releases MUST preserve that audit trail by linking the versioned checklist, release note, and conformance directory for the protocol version they claim.

The portable security assumptions, finite attacker-controlled dimensions, resource-policy owners, and explicit exclusions are documented in threat-model.md.

No implementation MAY advertise a reserved optional feature merely because a Rust constant records its numeric position.

Appendix A: Rust surface mapping

This appendix is non-normative. It ensures every current public concept has an explicit place in the model or is identified as an implementation helper.

Rust conceptProtocol term or classification
BackendErrorBackend failure taxonomy translated by the command engine into portable status
AcceleratorClassExtensible device class in device identity
CapabilitiesSemantic backend capabilities; not Virtio feature bits
DeviceIdentityDevice instance identity
DeviceLimitsContext, buffer, program, execution-queue, event, binding, and byte limits enforced before backend invocation
DeviceInfoDevice identity, semantic capabilities, and limits
DeviceInfoErrorConstruction-time rejection of reserved capabilities, absent baseline memory domains, or zero limits
ContextFlags, ContextDescContext creation intent; protocol 1.0 accepts only empty context flags
MemoryDomainStrict provider-owned buffer placement requirement
BufferUsage, BufferDesc, BufferRangeBuffer allocation intent and checked byte range
BufferProperties, BufferInfo, AllocatedBufferVerified backing properties kept out of the backend hot path
AccessModeBinding read/write intent
ByteSource, ByteSinkBounded contiguous-or-segmented bulk byte ports
ArtifactFormat, TargetIdentity, ArtifactRefOpaque program artifact description over a byte source
QueueFlags, QueueDescAccelerator execution-queue creation intent, not a Virtio queue
TimeoutRelative admission timeout; zero on wire means infinite
BindingRefBorrowed program binding view. Hosts must reject AccessMode incompatible with BufferUsage before backend admission
EventStateEvent completion state
SubmitFailureRejected versus indeterminate submission acceptance boundary
ReleaseFailureRejected versus indeterminate resource-release boundary
AcceleratorAccelerator backend contract
DeviceHealthRunning, known-state reset required, or backend-discard-required processor state
ResourcePolicyMandatory host-private aggregate limits for provider-retained buffer and program storage
RetainedBytesExact live and quarantined provider-retained bulk-storage charges
ResourceCounts, ResetDisposition, ResetReport, ResetErrorReset object and retained-byte accounting, reuse decision, and namespace validation
Le16, Le32, Le64Wire implementation aliases; not semantic API
PROTOCOL_MAJOR, PROTOCOL_MINORCandidate protocol version 1.0
COMMAND_QUEUEBaseline command virtqueue index
BASELINE_COMMAND_QUEUES, HARD_MAX_*, MIN_MAX_*Candidate queue, frame, binding, and configuration bounds
KNOWN_*_BITS, RESERVED_REQUEST_FLAG_NO_WAITAssigned and reserved-zero flag namespaces
FeatureBitsDevice-specific Virtio transport features
BASELINE_FEATURESMandatory feature set, currently empty
RESERVED_FEATURESReserved namespace that must not be advertised
RequestFlagsPer-request wire flags; empty in protocol 1.0
KnownOpcode, UnknownOpcodeValidated command namespace and opaque unknown opcode
StatusCodeExtensible response status namespace
WireConfig, ConfigErrorDevice-specific configuration and executable validity rules
RequestHeader, ResponseHeaderRequest and response frame headers
WireDeviceInfoDevice-information response payload
CreateContextRequestContext-create request payload
ObjectPayloadOne opaque object ID in a request or response
AllocateBufferRequestBuffer-allocation request payload
TransferBufferRequestBounded buffer-transfer request prefix
LoadProgramRequestProgram-load request prefix
CreateQueueRequestAccelerator execution-queue-create payload
SubmitRequest, WireBindingSubmission prefix and binding array entry
SubmitResponseEvent ownership payload for accepted or indeterminate submission
WireEventState, KnownEventState, UnknownEventStateEvent-poll response payload and extensible raw state handling
DecodeError, read_exact, checked_array_bytesProtocol decoding helpers; implementation-only
QueueSize, QueueState, QueueEpoch, QueueControlValidated Virtio queue configuration and reset lifecycle
ChainId, ChainRegion, ChainLayout, ChainIo, DeviceChainReset-scoped descriptor-chain identity, flattened metadata, and mapped byte-port presentation
DriverQueue, PublishedChain, UsedChain, ReclaimedChain, PublishErrorOwnership-preserving driver publication, completion, reset reclamation, and retryable backpressure
DeviceQueue, UsedLength, NotificationHint, NotificationRecheckConsuming device completion, exact used length, and lost-wakeup-safe notification decisions
SegmentedSource, SegmentedSinkDevice test/reference byte ports over segmented storage
FrameDecoder, DecodedRequest, DecodedRequestBody, FramePreflightComplete untrusted request validation before semantic dispatch
ResponseWriter, ResponsePayloadBounded response framing and exact direct payload destination
ObjectNamespace, ObjectKind, ObjectIdDevice-private namespaced typed object identity
ObjectTable, ObjectTableErrorDevice implementation state; encoding is not wire ABI
DeviceState, typed record types, ChildCountsContext-scoped ownership graph, quotas, and in-flight reference accounting
ReleaseState, CreateError, RestoreErrorDevice-state admission and release rollback helpers
status_from_backend_error, status_from_object_table_error, status_from_device_state_errorDevice implementation mapping helpers
MockAccelerator, FaultAccelerator, and their handle typesReference/test implementation, not normative protocol
virtio-accel-conformance cases, target fixtures, hooks, and reportsReusable provider-test implementation, not normative protocol

virtio-accel protocol 1.0 wire ABI

This document defines the protocol 1.0 candidate byte contract used by the command virtqueue. It is normative together with specification.md and virtqueue.md. Structure names refer to the Rust implementation for convenience; implementations in other languages depend only on the byte layouts and rules below.

All multibyte integers are unsigned little-endian values unless a field explicitly says otherwise. Every structure has byte alignment one and contains no implicit padding. Offsets and sizes are listed in the checked-in layout manifest.

1. Global limits

ConstantValueRequirement
Protocol version1.0Configuration MUST report major 1 and minor 0
Baseline command queues1Only command virtqueue index 0 exists
Hard maximum chain descriptors256The advertised maximum MUST be 2 through 256
Hard maximum request frame16 MiBIncludes the 16-byte request header
Hard maximum response frame16 MiBIncludes the 16-byte response header
Hard maximum submission bindings4096The semantic advertised limit may be smaller

The device-specific configuration limits are additional negotiated bounds. A device MUST NOT advertise a limit above a hard maximum. A driver MUST reject invalid configuration rather than allocating from it.

Every count-to-byte conversion MUST use checked multiplication and addition. A value that overflows the implementation address space or the fixed-width containing field is invalid.

2. Device-specific configuration

WireConfig is 16 bytes:

OffsetBytesFieldProtocol 1.0 rule
02protocol_majorMUST be 1
22protocol_minorA conforming 1.0 device reports 0; a 1.0 driver MUST accept a higher minor and use only 1.0 behavior
42command_queue_countMUST be 1
62max_chain_descriptorsMUST be 2 through 256 and no greater than the configured queue size
84max_request_bytesMUST be 97 through 16 MiB
124max_response_bytesMUST be 92 through 16 MiB

The minimum request limit admits a one-byte program artifact. The minimum response limit admits the complete device-information response.

The 16 bytes above are the protocol 1.0 configuration prefix. A future minor version may append configuration only when a 1.0 driver can safely ignore it and the extension does not alter baseline behavior without feature negotiation. A 1.0 driver reads and validates the prefix and does not require the entire transport-specific configuration region to be exactly 16 bytes.

The baseline device-specific feature set is empty. Feature-bit positions 0 through 4 are reserved for multi-queue, event-queue, external-memory, timeline-fence, and secure-context proposals. A protocol 1.0 device MUST NOT advertise them and a protocol 1.0 driver MUST NOT accept them.

3. Request and response frames

3.1 Request header

Every request begins with the 16-byte RequestHeader:

OffsetBytesFieldRule
02opcodeRaw opcode from section 5
22flagsMUST be zero
44payload_bytesExact number of readable bytes after this header
88request_idNonzero and unique among requests outstanding on this device instance

Bit zero of flags, formerly drafted as NO_WAIT, is reserved and has no 1.0 semantics.

The readable byte count MUST equal 16 + payload_bytes exactly. There is no tolerated trailing extension area. An unknown opcode remains a raw integer long enough to produce UNSUPPORTED; it MUST NOT be materialized as an invalid language enum.

A request ID may be reused only after the corresponding descriptor chain has been returned used or after the driver has completed a device reset. Request IDs correlate command completion; they do not identify accelerator events.

3.2 Response header

Every written response begins with the 16-byte ResponseHeader:

OffsetBytesFieldRule
02statusRaw status from section 6
22flagsMUST be zero
44payload_bytesExact number of bytes written after this header
88request_idExact request ID from the corresponding valid request header

The driver MUST validate that the response request ID matches the request associated with the used descriptor head. An unknown status is an opaque failure, never success.

Except for an indeterminate SUBMIT, every non-OK response has an empty payload. An indeterminate SUBMIT has the original mapped failure status and an eight-byte SubmitResponse; possession of that event ID prevents premature resource release.

4. Common value namespaces

4.1 Object IDs

An object ID is an opaque nonzero u64. Zero is invalid. Its encoding is device-private and no bit has driver-visible meaning.

4.2 Memory domains

ValueMeaning
1Host-preferred memory
2Device-preferred memory
3Shared/coherent memory class

All other values produce INVALID_ARGUMENT. These values express placement intent only and do not create a host mapping or external-memory handle.

4.3 Buffer usage bits

BitValueMeaning
00x00000001Transfer source
10x00000002Transfer destination
20x00000004Program input
30x00000008Program output
40x00000010Mutable program state

At least one usage bit MUST be set. Unknown bits produce UNSUPPORTED.

4.4 Binding access

ValueMeaning
1Read
2Write
3Read and write

All other values produce INVALID_ARGUMENT.

Read requires PROGRAM_INPUT or MUTABLE_STATE, Write requires PROGRAM_OUTPUT or MUTABLE_STATE, and Read and write requires MUTABLE_STATE. A mismatch produces PERMISSION_DENIED before backend admission.

4.5 Event states

ValueStateerror field
0PendingOK
1CompleteOK
2FailedNon-OK status explaining execution failure
3CancelledOK

Unknown event states make the response invalid to a 1.0 driver. A driver MUST retain the event and request recovery rather than guessing that the event is terminal.

5. Opcodes and payloads

The request payload length is exact. A fixed prefix followed by variable bytes has no alignment padding between the prefix and tail.

OpcodeValueRequest payloadOK response payloadMaximum required writable capacity
GET_DEVICE_INFO0x0001EmptyWireDeviceInfo92
CREATE_CONTEXT0x0100CreateContextRequestObjectPayload context ID24
DESTROY_CONTEXT0x0101ObjectPayload context IDEmpty16
ALLOCATE_BUFFER0x0200AllocateBufferRequestObjectPayload buffer ID24
FREE_BUFFER0x0201ObjectPayload buffer IDEmpty16
WRITE_BUFFER0x0202TransferBufferRequest + dataEmpty16
READ_BUFFER0x0203TransferBufferRequestExactly bytes data16 + bytes
LOAD_PROGRAM0x0300LoadProgramRequest + artifactObjectPayload program ID24
UNLOAD_PROGRAM0x0301ObjectPayload program IDEmpty16
CREATE_QUEUE0x0400CreateQueueRequestObjectPayload execution-queue ID24
DESTROY_QUEUE0x0401ObjectPayload execution-queue IDEmpty16
SUBMIT0x0500SubmitRequest + WireBinding[]SubmitResponse event ID24
POLL_EVENT0x0501ObjectPayload event IDWireEventState24
CANCEL_EVENT0x0502ObjectPayload event IDEmpty16
DESTROY_EVENT0x0503ObjectPayload event IDEmpty16

The maximum required writable capacity is validated before any semantic mutation or backend call. Protocol 1.0 has no object-list payload: every destruction or event operation names exactly one object ID. The binding array is the only variable-count structured array in a baseline request.

5.1 WireDeviceInfo

WireDeviceInfo is 76 bytes:

FieldRule
uuid[16]Stable identity for this accelerator device
classExtensible raw class; 0 other, 1 NPU, 2 GPU, 3 DSP
reservedZero
vendor_id, device_idProvider identity; zero means unspecified
capabilitiesAssigned semantic capability bits only
max_contextsNonzero device-wide live-context limit
max_buffers_per_contextNonzero live-buffer limit
max_programs_per_contextNonzero live-program limit
max_queues_per_contextNonzero live execution-queue limit
max_events_per_contextNonzero live-event/in-flight-submission limit
max_bindings_per_submission1 through 4096
max_buffer_bytesNonzero maximum allocation size
max_artifact_bytesNonzero maximum artifact tail size, additionally bounded by the request-frame limit

Capability bits are semantic reports, not Virtio feature bits. A capability MUST NOT alter wire framing without a separately negotiated feature. Unknown capability bits are ignored for operation selection and preserved by diagnostic tooling.

Assigned protocol 1.0 semantic capability bits are:

BitName
0HOST_VISIBLE_MEMORY
1DEVICE_LOCAL_MEMORY
2EVENT_CANCELLATION
5SHARED_MEMORY

Bits 3 (EXTERNAL_MEMORY) and 4 (SECURE_CONTEXTS) are reserved and MUST NOT be advertised by a protocol 1.0 device.

5.2 Context

CreateContextRequest is eight bytes: flags: u32 followed by reserved: u32. Both fields MUST be zero in protocol 1.0.

Context destruction uses an eight-byte ObjectPayload.

5.3 Buffers

AllocateBufferRequest is 40 bytes:

FieldRule
context_idLive context
bytesNonzero and no greater than max_buffer_bytes
alignmentNonzero power of two
memory_domainAssigned value from section 4.2
reserved0[7]All zero
usageNonempty subset of assigned usage bits
reserved1Zero

The device MUST reject a memory domain whose corresponding semantic capability is absent before backend invocation. Host, Device, and provider-owned Shared allocations use capability bits 0, 1, and 5 respectively. Successful allocation commits the backend to the placement and direct-binding rules in specification.md; it may not silently substitute a staged submission path.

TransferBufferRequest is 24 bytes containing buffer_id, offset, and bytes. bytes MUST be nonzero. offset + bytes MUST NOT overflow and MUST fit in the buffer.

For WRITE_BUFFER, the request payload length MUST be 24 + bytes, and bytes following the prefix are copied to the buffer. For READ_BUFFER, the request payload is exactly 24 bytes and the success response payload contains exactly bytes bytes. Transfers must also fit the configured request or response frame maximum.

WRITE_BUFFER requires buffer usage TRANSFER_DESTINATION; READ_BUFFER requires TRANSFER_SOURCE. These commands are explicit copy boundaries. Their existence does not permit allocation or submission to copy program bindings through hidden bounce buffers.

5.4 Programs

LoadProgramRequest is an 80-byte prefix:

FieldRule
context_idLive context
formatNonzero provider-owned format ID
flagsZero
target[12]Opaque format-owned target words
payload_bytesNonzero, equals the exact artifact tail length
resident_bytesNonzero upper bound on provider storage retained for the returned program

80 + payload_bytes MUST fit the request payload and configured frame limit.

5.5 Execution queues

CreateQueueRequest is 16 bytes: context_id: u64, flags: u32, and reserved: u32. Both flag and reserved words MUST be zero in protocol 1.0.

5.6 Submission and events

SubmitRequest is a 32-byte prefix containing queue_id, program_id, binding_count, flags, and timeout_ns.

  • binding_count MUST be 1 through both advertised max_bindings_per_submission and 4096.
  • flags MUST be zero.
  • timeout_ns is relative to backend admission; zero means infinite.
  • The payload length MUST be 32 + binding_count * 32.

Each 32-byte WireBinding contains buffer_id, offset, bytes, slot, access, and three reserved-zero bytes. Buffer ranges are nonempty and checked for overflow. Slots are unique within the submission. Every object belongs to the same context.

SubmitResponse is an eight-byte event ID. It is returned with OK after accepted admission and with the mapped non-OK status when admission is indeterminate. A rejected submission has an empty error payload.

WireEventState is eight bytes: state: u16, error: u16, and reserved: u32. Reserved bytes are zero and the state/error combinations are exactly those in section 4.5. Polling is nonblocking. Once a terminal state is observed, later successful polls return the same terminal state.

6. Status namespace

ValueNameMeaning
0OKCommand completed successfully
1UNSUPPORTEDOpcode, flag, feature, capability, or operation is not supported
2INCOMPATIBLEKnown artifact, target, object, or capability combination is incompatible
3INVALID_ARGUMENTMalformed value, reserved field, length mismatch, zero required value, or duplicate slot
4OUT_OF_BOUNDSChecked byte range does not fit its object
5BUSYObject is live, referenced, pending, or otherwise retryable
6OUT_OF_MEMORYHost/provider allocation failed
7RESOURCE_LIMITConfigured count or byte limit would be exceeded
8DEADLINE_EXPIREDOperation expired according to the relative timeout contract
9DEVICE_LOSTBackend/device state cannot continue normally
10PERMISSION_DENIEDIsolation or provider policy denied the operation
11STALE_OBJECTNonzero object ID is stale, wrong-kind, reset-invalidated, or not valid in this context
65535INTERNAL_ERRORUnclassified implementation/provider failure

Unknown status values are opaque non-success failures. Provider-specific error domains do not cross the 1.0 wire boundary; absent a future diagnostic feature, they map to INTERNAL_ERROR.

Malformed input is classified before backend invocation:

  • unknown opcode or nonzero/unknown flags: UNSUPPORTED;
  • fixed-length mismatch, trailing bytes, reserved nonzero, invalid scalar, zero required value, or arithmetic overflow: INVALID_ARGUMENT;
  • frame, binding, object, or configured quota exceeded: RESOURCE_LIMIT;
  • valid object ID with wrong generation, kind, reset epoch, or context: STALE_OBJECT; and
  • valid byte range outside the selected object: OUT_OF_BOUNDS.

7. Response atomicity

Before backend invocation, the device MUST validate the complete readable frame and enough writable capacity for every possible response shape of that command. It MUST initialize every response byte it reports used.

Ordinary protocol errors have no semantic state change. If an unexpected transport write failure occurs after semantic mutation, or a release becomes indeterminate, the device MUST enter recovery and expose the Virtio DEVICE_NEEDS_RESET condition. It MUST NOT report an ordinary rejected response that would let the driver free resources whose ownership is uncertain.

8. Versioned compatibility artifacts

Protocol constants, layouts, and canonical bytes are checked in under conformance/v1.0. They are review inputs, not test-generated output. The final freeze audit in ../conformance/v1.0/freeze-audit.md makes these files the frozen 1.0 inputs. Post-freeze changes are classified by section 9 and release-policy.md.

9. Candidate and post-freeze change procedure

A proposed wire change MUST be classified before code is merged:

  1. An erratum that changes no accepted or emitted bytes may clarify the 1.0 documents and tests.
  2. A compatible extension uses a previously reserved number plus explicit feature or new-opcode negotiation, preserves every 1.0 frame, and receives a new minor-version conformance directory.
  3. Any changed assigned number, existing payload length, field meaning, required response, or ownership interpretation requires a new protocol major version and a new conformance directory.

The same reviewed change MUST update the normative documents, Rust constants/layout assertions, machine layout manifest, canonical vectors, and compatibility tests. Protocol version directories are never regenerated opportunistically from current Rust types.

virtio-accel protocol 1.0 command-virtqueue contract

This document defines how protocol 1.0 request and response frames inhabit Virtio descriptor chains. It is normative together with specification.md and wire-abi.md.

The portable reference path targets a split virtqueue. It relies only on ordered readable and writable byte regions plus queue publication/completion operations, so future transport adapters do not need to expose guest addresses or transport-specific descriptor types to the command engine.

1. Relationship to the base Virtio specification

The Virtio 1.3 specification continues to supply the base rules, including descriptor-loop prohibition, negotiated indirect descriptors, queue publication barriers, notification suppression, and transport reset. This document adds device-specific framing requirements.

Protocol 1.0 defines command virtqueue index zero. It defines no device-specific multi-queue feature. The reference device MUST NOT offer VIRTIO_F_RING_PACKED, VIRTIO_F_IN_ORDER, VIRTIO_F_RING_RESET, VIRTIO_F_INDIRECT_DESC, or VIRTIO_F_EVENT_IDX. It uses direct split-ring descriptors and the basic available/used notification-suppression flags. A future adapter may support one of these base Virtio features only when it preserves the device-specific rules here and adds corresponding conformance evidence. Whole-device reset remains mandatory.

2. Flattened chain model

For validation, a transport adapter presents one available descriptor head as:

  • an ordered list of device-readable byte regions;
  • followed by an ordered list of device-writable byte regions; and
  • the original queue head token used for completion.

A future adapter may flatten a valid indirect descriptor table when VIRTIO_F_INDIRECT_DESC was negotiated. The configured max_chain_descriptors then applies after flattening and includes both readable and writable descriptors. The protocol 1.0 reference profile rejects indirect descriptors because it does not negotiate that feature.

The command engine sees region lengths and byte access only. It MUST NOT receive guest physical addresses, descriptor indices, indirect-table addresses, ring pointers, or notification objects.

3. Descriptor-chain validity

A protocol 1.0 command chain is valid only when all of the following hold:

  1. The chain is acyclic and structurally valid under the base Virtio specification.
  2. Its flattened descriptor count is 2 through configured max_chain_descriptors.
  3. Every descriptor length is nonzero.
  4. At least one device-readable descriptor precedes at least one device-writable descriptor.
  5. No device-readable descriptor follows a device-writable descriptor.
  6. The readable total is representable, no greater than configured max_request_bytes, and exactly one complete request frame.
  7. The writable total is representable and large enough for the command’s maximum possible response shape.

Cross-region reads and writes concatenate regions without inserting padding. Header or payload fields may cross descriptor boundaries at any byte.

The transport adapter MUST validate the descriptor topology and map every readable and required writable byte before the command engine performs semantic validation or invokes a backend.

4. One command per chain

Exactly one request frame occupies the complete readable portion of a chain. The first readable byte is request-header byte zero. The readable total MUST equal:

16 + request_header.payload_bytes

No bytes before the header, between fixed and variable payload portions, or after the declared payload are permitted.

The response begins at writable byte zero. Writable capacity beyond the actual response remains untouched and is not included in used length.

5. Validation order and failure atomicity

The device validates a chain in this order:

  1. descriptor topology, direction, count, nonzero lengths, addressability, and total-length arithmetic;
  2. availability of a complete 16-byte request header and 16-byte writable response header;
  3. configured request limit and exact payload_bytes equality;
  4. opcode and request-header flags;
  5. command-specific fixed length, variable counts, reserved fields, scalar namespaces, object relationships, and required writable capacity;
  6. quotas and semantic preconditions; and
  7. backend invocation.

No failure in steps 1 through 5 may mutate device semantic state or invoke the backend.

If descriptor topology is invalid, the request header is truncated, or writable capacity is less than 16 bytes, the device returns the descriptor head used with used length zero and writes no bytes.

If the request header is valid enough to recover request_id, and writable capacity is at least 16, later validation failures produce a complete error response header. The used length is 16.

Required success capacity is checked before semantic mutation. A short success buffer therefore produces used length zero, no response bytes, and no semantic state change; it is not converted to a smaller protocol error because the driver failed the response-buffer contract.

6. Response writing and used length

The device MUST commit a response atomically from the protocol’s point of view:

  • every byte counted as used is initialized;
  • the response header’s payload_bytes equals the bytes written after it;
  • no byte beyond 16 + payload_bytes is written; and
  • the split-ring used element length equals 16 + payload_bytes.

Used length counts device-written bytes only. It never includes readable request bytes.

Response bytes may span writable descriptors. A device MUST write them as if to one concatenated byte sequence.

An unexpected mapping or write failure after the preflight succeeds indicates broken transport or memory state. If semantic mutation has not occurred, the chain may complete with used length zero. If mutation or uncertain backend acceptance has occurred, the device MUST set DEVICE_NEEDS_RESET, quarantine retained resources, and avoid a response that falsely claims rejection.

7. Command and execution ordering

The baseline command queue permits out-of-order command completion. The device MUST consume available heads in available-ring order, but may dispatch their semantic work concurrently and return them used in a different order as work completes. Drivers therefore track both descriptor heads and nonzero request IDs.

Request-ID uniqueness lasts until the corresponding chain is returned used. A driver MUST NOT infer ordering from numeric request IDs.

Command completion and accelerator execution completion are distinct:

  • SUBMIT command completion reports admission and returns an event ID;
  • the accelerator operation may remain pending after the command chain is used; and
  • POLL_EVENT observes execution completion independently of command ordering.

Implementations may serialize command execution, but they MUST NOT promise in-order completion as protocol behavior.

8. Backpressure and queue fullness

Queue capacity is transport state, not a protocol error response.

  • A driver with no free descriptor head or insufficient writable buffers MUST retain the command locally and report backpressure to its caller.
  • It MUST NOT publish a partial command or reuse descriptors still owned by the device.
  • A device may defer consumption or completion without busy-looping.
  • Resource limits discovered after a valid chain is consumed produce RESOURCE_LIMIT; descriptor scarcity before publication does not.

The reference guest API should represent pre-publication backpressure separately from protocol statuses so callers can retry without constructing a new semantic command.

9. Notifications

Available and used notifications follow the base Virtio split-ring suppression rules. Suppressing a notification changes neither command visibility nor completion semantics.

  • The driver publishes descriptors and the available-ring entry before deciding whether to notify.
  • The device publishes the used element and used index before deciding whether to notify.
  • Both sides must recheck queue state when required by the base specification to avoid lost wakeups.

The device-specific protocol defines no polling interval, thread, executor, or interrupt affinity.

10. Malformed chains

The following are isolated command-chain failures and do not by themselves require device reset:

  • descriptor loops or an invalid indirect table caught before byte access;
  • too many, zero-length, missing-readable, missing-writable, or interleaved-direction descriptors;
  • truncated headers or payloads;
  • oversized frames;
  • exact-length or reserved-zero violations;
  • unknown opcodes, flags, or scalar values; and
  • insufficient writable capacity detected before semantic mutation.

The device completes a structurally recoverable malformed chain according to section 5 and continues with later available chains.

Repeated malformed input may be rate-limited by the transport or device integration, but rate limiting MUST NOT change object ownership or fabricate successful responses.

11. Device reset and split-ring disposition

Protocol 1.0 uses whole-device reset. The reference device does not negotiate independent VIRTIO_F_RING_RESET.

When reset begins, the device:

  1. stops fetching new available chains;
  2. stops publishing used entries and notifications;
  3. prevents new backend admission;
  4. resolves, cancels, or quarantines already-started operations according to known ownership;
  5. invalidates all object IDs and request tracking from the old reset epoch; and
  6. resets command-queue available/used state as required by the base transport.

After steps 1 and 2 quiesce queue access, the transport gives exclusive ownership to the portable command processor for one bounded teardown pass. A reusable result permits queue and object-table reinitialization with a fresh namespace. A discard-required result forbids further backend calls; the transport discards that processor/backend instance and creates a new one before exposing new queues.

Descriptor chains that were available or in progress when reset began are not completed after reset. Once the driver has observed reset completion, it may reclaim all queue memory and descriptors under the base Virtio reset rule. It MUST NOT expect response bytes or used entries for those chains.

After reinitialization, request IDs may be reused, all object IDs from the prior epoch are stale, and the queue starts from its base initial state. No late backend completion from the old epoch may write guest memory or a new used ring.

12. DEVICE_NEEDS_RESET

Ordinary hostile input does not set DEVICE_NEEDS_RESET. The device sets it when continued protocol operation cannot preserve ownership or response truth, including:

  • indeterminate provider release that invalidates a guest object but leaves backend ownership unknown;
  • inability to report or retain an indeterminate accepted submission;
  • response-write failure after semantic mutation;
  • internal state corruption or accounting contradiction; or
  • backend device loss that prevents bounded recovery.

After observing DEVICE_NEEDS_RESET, the driver SHOULD stop submitting commands and perform whole-device reset. The device may finish responses whose ownership and output remain provably safe, but it MUST NOT accept new semantic work.

13. Protocol conformance cases

The portable conformance suite uses these stable case identifiers:

IDRequired assertion
VQ-001One readable and one writable descriptor carries a valid command
VQ-002Header and every fixed payload decode across every possible byte split
VQ-003Multiple readable and writable segments concatenate without padding
VQ-004Missing readable or writable region completes with used length zero
VQ-005Readable descriptor after a writable descriptor is rejected before backend invocation
VQ-006Zero-length, looping, excessive, or invalid-indirect chains are rejected
VQ-007Truncated request header writes nothing and reports used length zero
VQ-008Valid header plus malformed payload writes exactly a 16-byte error response
VQ-009Unknown opcode returns UNSUPPORTED without semantic mutation
VQ-010Nonzero request, command, or reserved flags are rejected
VQ-011Trailing readable bytes are rejected
VQ-012Short command-specific response capacity is detected before mutation and writes nothing
VQ-013Success and error used lengths equal exact initialized response bytes
VQ-014Distinct requests can complete out of publication order and retain correct request IDs
VQ-015SUBMIT command completion is independent of event completion
VQ-016Queue-full pre-publication behavior is retryable backpressure, not a protocol status
VQ-017Notification suppression does not lose available or used work
VQ-018Reset produces no late used entries or guest writes from the old epoch
VQ-019Reset invalidates every old object ID and permits request-ID reuse
VQ-020Post-mutation output failure or indeterminate release sets DEVICE_NEEDS_RESET

The byte vectors under conformance/v1.0 cover protocol frames. Queue-model implementations consume the case identifiers above so the same behavioral scenarios can be reused by the in-memory split ring and future platform transports. The dependency-free virtio-accel-transport crate provides the executable region, ownership, reset-epoch, backpressure, and notification port contracts. virtio-accel-split-queue exercises the ring-level portions of VQ-001, VQ-003 through VQ-006, VQ-013, VQ-014, and VQ-016 through VQ-018. The reference guest and end-to-end scenarios exercise the complete VQ-001 through VQ-020 lifecycle.

Appendix A: portable Rust queue-port mapping

This appendix is non-normative. DriverQueue::publish consumes a complete driver chain on success and returns it unchanged through PublishError on pre-publication failure. pop_used or reset returns every successfully published chain, preventing safe callers from reusing descriptors while the device owns them.

DeviceQueue::pop_available returns a non-Copy DeviceChain. Completion consumes that value and must compare its ChainId epoch with current QueueState before publishing bytes or ring state. ChainIo contains only flattened ChainRegion values and generic readable/writable byte ports; it cannot carry guest addresses or a concrete descriptor type.

All steady-state queue-port methods are specified as nonblocking and allocation-free in their Rust documentation. Publication/completion establish release ordering, their peer-side pops establish acquire ordering, and notification enablement returns WorkPending when its mandatory recheck finds work that raced with suppression.

Appendix B: in-memory split-ring model

This appendix is non-normative. SplitQueue preallocates one descriptor-ownership table, chain record table, available ring, and used ring from a validated power-of-two QueueSize. Publication, available consumption, out-of-order completion, used consumption, and notification rechecks move owned values or update fixed slots; none allocates or coalesces payload bytes.

DriverChain records each descriptor’s logical byte span during its bounded construction pass. Segmented byte access binary-searches the first intersecting span and visits only descriptors that contain requested bytes, avoiding repeated prefix scans when a command decoder performs many small reads. The span metadata remains caller-owned chain storage and is never allocated from descriptor byte lengths.

DriverChain::direct constructs the baseline direct profile. DriverChain::raw retains malformed topology for deterministic device tests, while SplitQueue::inject_available bypasses only normal driver-side profile validation. Traversal remains bounded by the supplied descriptor table and never allocates from a descriptor byte length, next value, or other guest-controlled scalar. Unknown flags and VIRTQ_DESC_F_INDIRECT are rejected because protocol 1.0 negotiates neither.

The four public RingCounters use the split ring’s wrapping u16 index arithmetic. Test hooks can place an empty queue at a chosen index and choose the next descriptor allocation slot, making ring wrap, descriptor wrap, exhaustion, and notification races reproducible without timing or threads. Reset publishes a new atomic epoch before removing ring storage; every old ChainIo byte access and completion checks that epoch before touching a payload or used entry.

Protocol 1.0 threat model and resource policy

Status: security model for the portable protocol 1.0 candidate. This document explains the threat assumptions and maps them to normative requirements, implementation owners, and executable evidence. The requirements themselves live in specification.md, wire-abi.md, and virtqueue.md.

Security goals

The portable device treats every guest-controlled byte, count, object ID, descriptor shape, timeout, command ordering, and reset timing as hostile. Its goals are to:

  • preserve memory safety without trusting wire values or guest mappings;
  • keep guest-visible objects isolated by device, reset epoch, kind, and context;
  • bound the CPU work and allocation attributable to one command;
  • bound live object counts and provider-retained bulk storage;
  • publish no uninitialized or ownership-ambiguous response;
  • recover deterministically when provider ownership is known; and
  • stop admitting work and require backend discard when ownership cannot be proven.

Availability of the host process or physical accelerator against a malicious host backend is not a portable-layer guarantee. The device still preserves guest-visible ownership truth when a trusted backend reports a hang, device loss, or indeterminate result.

Trust boundaries

untrusted guest
    |
    | descriptors, bytes, counts, IDs, timing
    v
transport adapter  -- validates/maps guest memory and queue topology
    |
    | address-free readable/writable ports and owned chain tokens
    v
command engine     -- validates frames, ownership, quotas, and resource policy
    |
    | semantic values and provider-owned handles only
    v
accelerator backend -- trusted implementation, failure-prone provider/device
    |
    v
trusted host process and platform integration

The guest is adversarial. The transport adapter, command engine, backend implementation, and host process are trusted code, but mappings, allocations, provider calls, and hardware may fail. Safe Rust prevents memory unsafety in the portable crates; it does not make a dishonest backend conformant.

The transport adapter owns guest-address validation, access permissions, pinning, and base Virtio ring synchronization. The command engine never receives an address. The backend receives neither wire structures nor guest memory and must not retain borrowed byte ports after a call returns.

In-scope threats and bounds

Attacker-controlled dimension or failureBound and enforcement ownerFailure behaviorExecutable evidence
Descriptor topology, direction, loops, and segmentationThe transport validates a flattened chain in one pass. WireConfig::max_chain_descriptors is 2 through HARD_MAX_CHAIN_DESCRIPTORS (256); the reference split queue also bounds traversal by its configured table.Invalid topology completes without backend invocation; unusable chains write no bytes.chain_layout_rejects_direction_and_length_errors, vq_004_vq_005_vq_006_malformed_topology_is_classified_boundedly, and unusable_frames_write_nothing_and_never_become_dispatchable
Request and response byte totalsConfiguration limits are capped by HARD_MAX_REQUEST_BYTES and HARD_MAX_RESPONSE_BYTES (16 MiB each). Checked addition enforces exact frame lengths and response capacity before mutation.Oversized or inconsistent frames return RESOURCE_LIMIT or INVALID_ARGUMENT; short success capacity writes nothing.edge_vectors_cover_limits_reserved_bits_and_exact_lengths, short_success_capacity_writes_nothing, and short_responses_and_backend_rejection_never_publish_state
Variable binding arraysDeviceLimits::max_bindings_per_submission is nonzero and at most HARD_MAX_BINDINGS (4096). Decode performs one fallible bounded reservation and sorts in place. Event metadata retains at most the same bounded ID count.Limit excess returns RESOURCE_LIMIT; allocation failure returns OUT_OF_MEMORY; no backend call occurs.duplicate_submission_slots_are_rejected_before_dispatch, bindings_are_nonempty_bounded_and_unique, and submission_validation_rejects_before_backend_admission
Object IDs and cross-context referencesIDs are nonzero, namespace-, kind-, slot-, and generation-tagged. Slots retire before generation wrap. Every combined operation checks one context before provider invocation.Zero or malformed IDs are invalid; stale, wrong-kind, cross-context, and reset-invalidated IDs are rejected without handle access.stale_ids_never_resolve_after_slot_reuse, exhausted_generations_retire_the_slot_before_an_id_can_revive, and wrong_kind_cross_context_and_cross_device_ids_fail_before_provider_use
Live contexts, buffers, programs, execution queues, and eventsNonzero DeviceLimits cap contexts and each per-context child class. Aggregate table capacities use checked multiplication. Event/reference maxima are checked before state exists.Exhaustion returns RESOURCE_LIMIT before provider invocation or table growth. Fallible table reservation maps to OUT_OF_MEMORY.limits_are_enforced_before_growth, invalid_limits_are_rejected_before_tables_exist, and quota_exhaustion_never_invokes_provider_callbacks
Logical buffer size and transfer rangesmax_buffer_bytes caps each logical allocation. Nonzero ranges use checked end arithmetic and must fit the logical buffer. Usage bits gate read, write, and program access.Oversize returns RESOURCE_LIMIT; arithmetic/range failure returns INVALID_ARGUMENT or OUT_OF_BOUNDS; permission mismatch returns PERMISSION_DENIED.transfer_range_overflow_is_a_protocol_error, semantic_validation_prevents_backend_calls, and mutable_state_allows_every_program_access_mode
Aggregate provider buffer backingA mandatory device-private ResourcePolicy::max_buffer_backing_bytes limits the sum of actual BufferInfo::allocation_bytes, not logical guest bytes. The state prechecks the logical lower bound before allocation and reconciles the actual charge before publishing an ID.An over-budget allocation is released before its ID is exposed. Rejected or indeterminate cleanup makes the backend discard-required and retains or quarantines the exact charge.aggregate_policy_uses_actual_backing_and_charges_until_release_commits, aggregate_retained_byte_policy_cleans_up_before_exposing_an_id, rejected_cleanup_of_an_unpublished_allocation_requires_backend_discard, and indeterminate_cleanup_of_an_unpublished_allocation_quarantines_actual_backing
Artifact input and resident program storagemax_artifact_bytes bounds the borrowed input payload. ResourcePolicy::max_program_resident_bytes bounds the aggregate declared resident charge before provider invocation. A conforming backend retains no more storage than ArtifactRef::resident_bytes.Excess returns RESOURCE_LIMIT before loading. Invalid or unsupported artifacts create no program resource.byte_limits_and_resident_charges_have_distinct_semantics, program_policy_rejects_before_provider_invocation, and malformed_reference_artifacts_do_not_create_resident_programs
TOSA structure, semantics, and graph cardinalityvirtio-accel-tosa::Limits caps input bytes, verifier depth/tables/apparent bytes, graph objects, edges, names, ranks, and constants before returning a safe view. Model::validate_for checks stable-op semantics and Model::analyze_for derives bounded dense IDs, dependency order, liveness, constant state, runtime obligations, and specialization inputs without copying graph payloads. Runtime values use exact bounded encodings; specialization keys and caches have explicit word and entry limits.Malformed tables/unions, draft schema entries, invalid references, unsupported target combinations, illegal constants/geometry, missing or malformed dynamic compile-time inputs, allocation failure, and limit excess return an error before provider compilation. Advisory REQUIRE failures are reported as unpredictable rather than misclassified as mandatory errors.virtio-accel-tosa unit/integration tests plus the tosa_parse fuzz target seeded from an upstream stable graph and exercising semantic analysis and runtime validation for two targets
Command-queue occupancy and guest request trackingVirtio QueueSize bounds descriptor/ring ownership. GuestConfig::max_inflight is nonzero and no greater than queue capacity. The split reference model preallocates ring state at configuration.Prepublication pressure returns caller ownership as retryable backpressure; it is not fabricated as a protocol response.publication_backpressure_returns_the_chain, prepublication_backpressure_returns_chain_and_operation, and queue_pressure_returns_caller_ownership_before_draining
Submission and in-flight referencesPer-context events and per-submission bindings bound retained invocation metadata. Queue, program, and buffer reference increments are checked before admission and released exactly once with the event.Referenced objects return BUSY; rejected admission rolls back references; indeterminate admission retains an event ownership token.complete_lifecycle_tracks_children_references_and_release_rollback, rejected_and_indeterminate_submissions_preserve_the_admission_boundary, and finite_timeout_rejection_retains_all_referenced_resources
Cancellation, completion, release, and reset racesOne mutable command processor serializes semantic transitions. Provider events own native completion races. Reset makes one bounded child-before-parent pass and never waits for progress.Rejected release restores the same ID; indeterminate release invalidates it and requires backend discard. A pending uncancellable event prevents backend reuse.accepted_events_retain_resources_and_resolve_cancel_completion_races, rejected_reset_release_is_quarantined_and_reset_is_idempotent, and reset_reclaims_completed_and_cancelled_events_exactly_once
Response truncation and information disclosureCommand-specific capacity is preflighted before mutation. Payload bytes are initialized before the header commits their length; used length includes only initialized bytes.Unusable capacity writes nothing. Post-mutation write failure sets recovery instead of claiming rejection.payload_guard_commits_the_preflighted_length, direct_payload_region_does_not_touch_excess_capacity, and response_failure_after_creation_requires_reset
Backend device loss or ownership uncertaintyThe acceptance and release result types distinguish rejected from indeterminate outcomes. Quarantined counts and retained-byte charges survive until the complete backend instance is discarded.New semantic work stops. Reset is sticky, makes no later provider calls, and reports represented plus quarantined resources.event_faults_and_unreportable_admission_require_recovery, indeterminate_event_release_quarantines_the_complete_object_graph, and device_loss_crosses_backend_engine_and_guest_recovery_boundaries

All named tests are located in the corresponding crate’s src unit tests, crates/virtio-accel-device/tests/command_processor.rs, crates/virtio-accel-device/tests/state_model.rs, or tests/portable_end_to_end.rs. Coverage-guided malformed input checks live under fuzz/: protocol decode is compared with the clean-room codec, descriptor segmentation is driven through the split-queue model, and bounded stateful sequences check resource accounting after every action. The TOSA target mutates a stable graph, traverses all safe views, decodes stable attributes, and runs target-semantic validation. Another target drives the reference guest client against a non-conforming device. That direction is reference-implementation robustness rather than a portable-layer security boundary, because a malicious backend or host process is excluded below; it exists so the driver-side obligations in this protocol — opaque unknown statuses, recovery on unknown event states, bounded in-flight tracking, and epoch-scoped handle staleness — hold against inputs no example test enumerates. The deterministic state-model suite runs a bounded seed set in normal CI, asserts stale-ID retirement, context isolation, reference release, quota, and retained-byte invariants after every action, and prints minimized replayable schedules for failures. An ignored deep test provides manual or nightly-style exploration with VIRTIO_ACCEL_STATE_MODEL_SEED.

Resource-accounting rules

DeviceLimits are backend capabilities and per-object/count bounds visible to the driver. WireConfig bounds command framing and descriptor presentation. ResourcePolicy is host policy: it is supplied when constructing a CommandProcessor, is never guest-controlled, and is not a promise that an allocation will succeed.

The policy has two nonzero aggregate limits:

  • max_buffer_backing_bytes: actual provider allocation bytes across all live buffer handles; and
  • max_program_resident_bytes: declared resident charges across all live program handles.

RetainedBytes uses u128 totals so the sum of a u32 object count and u64 per-object charges is representable. A charge begins before the corresponding object ID is returned. It remains while a handle is live or in the Releasing state and ends only when provider release succeeds. Rejected release restores the handle without dropping its charge. Indeterminate release transfers the charge to quarantine and requires the entire backend instance to be discarded.

Buffer padding is provider knowledge, so the command engine first checks the logical byte lower bound, then reconciles the returned BufferInfo::allocation_bytes. If the actual charge crosses the aggregate budget, the engine attempts release before exposing the new ID. Failure to prove that cleanup succeeded cannot be represented as an ordinary RESOURCE_LIMIT; it is device loss.

Program resident storage is declared before provider invocation. ArtifactRef::resident_bytes is the maximum storage the provider may retain for the returned program, including compiled code and provider metadata attributable to that program. A backend that cannot honor the charge rejects the load. Temporary provider memory used only during the synchronous call is governed by the backend and host allocator, not counted as retained program storage. The portable layer cannot observe that internal allocation, so a backend or isolated provider worker must enforce its own transient compile memory and work budget when hostile artifacts can expand beyond their bounded input.

Object-table capacity and per-event binding vectors are bounded by count limits rather than byte estimates because their Rust layout is implementation-specific. Quantitative retained metadata, allocation, and copy-path budgets are tracked in performance.md; changing their representation does not change the security bounds above.

Progress and denial of service

Every portable parsing, lookup, queue, polling, cancellation, and reset operation is bounded by a validated count and contains no retry loop waiting for guest or provider progress. Repeated hostile commands can consume an unbounded amount of CPU over unbounded time; rate limiting belongs to the transport or host integration because the portable layer has no clock, scheduler, tenant identity, or worker policy. Rate limiting must not change ownership or fabricate successful responses. Malformed guest values and configured resource exhaustion are ordinary protocol errors: they do not panic, create an unbounded allocation, or expose uninitialized response storage.

Provider lifecycle and explicit transfer calls are synchronous and may block according to the Accelerator contract. The portable layer cannot safely preempt arbitrary provider code or reclaim native handles from a hung call. Integrations that require availability against provider hangs must use an appropriate process or worker with CPU and memory limits, a watchdog, or a hardware-reset boundary and discard the complete backend instance after loss. Turning provider calls into futures would not itself enforce bounded polling or safe cancellation.

Timeouts are relative admission constraints, not host watchdogs. A zero timeout is infinite. After admission, an uncertain timeout retains an event rather than permitting a duplicate submission.

Exclusions

Protocol 1.0 does not claim protection against:

  • a malicious backend or compromised host process;
  • physical attacks, accelerator side channels, power analysis, or denial of service by hardware;
  • confidentiality between mutually untrusted workloads inside one provider context unless the provider supplies that isolation;
  • guest-memory pinning limits, IOMMU policy, or cache-coherency bugs in a future platform adapter;
  • artifact-language safety inside an opaque vendor compiler or executable; or
  • platform external-memory lifetime and synchronization, which remain an unadvertised feature.

These exclusions do not permit a portable adapter to weaken object ownership, response atomicity, or resource accounting. A future feature that crosses one of these boundaries needs its own threat model, negotiation, conformance evidence, and release audit. The non-normative external-memory handoff design enumerates those gates; it does not narrow this Protocol 1.0 exclusion or authorize advertising the reserved feature.

Review obligations

A change that adds a guest-controlled count, retained allocation, new descriptor topology, new queue, new asynchronous ownership state, or new platform handle must update this document and map the dimension to one authoritative limit. It must also add executable evidence and refresh the normative requirement ledger. Silent defaults, duplicate conflicting caps, and best-effort cleanup across an indeterminate ownership boundary are not conformant.

Protocol 1.0 release notes

Status: frozen portable protocol 1.0 candidate after the v1.0 freeze audit.

The release surface is the portable protocol and Rust foundation for transport-neutral accelerator devices. It does not assign a Virtio device ID, standardize an NPU graph/compiler IR, or include platform adapters.

Frozen artifacts

The protocol 1.0 assigned constants, exact structure layouts, canonical frame bytes, scenario traces, and hot-path budgets are frozen in conformance/v1.0. Post-freeze incompatible wire or semantic changes require a new protocol major version and a new conformance directory.

Baseline included in 1.0

  • one command virtqueue at index zero;
  • device discovery and exact protocol compatibility checks;
  • contexts, buffers, opaque programs, execution queues, submissions, and events;
  • bounded explicit buffer transfers;
  • event polling, optional cancellation, release, reset, and backend-discard recovery;
  • direct-binding requirements for program-visible buffers; and
  • checked finite limits for untrusted byte counts, descriptor counts, object counts, and retained backend storage.

Deferred optional features

The following positions remain reserved and unadvertised in protocol 1.0:

  • multi-queue;
  • event queues;
  • external memory import/export;
  • timeline fences;
  • secure contexts;
  • packed virtqueues;
  • concrete VMM, kernel, OS, and vendor SDK adapters; and
  • standardized graph IR, compiler, or executable formats.

An implementation that advertises one of those features is not protocol 1.0 conformant unless a future version assigns its negotiation, ownership, synchronization, and conformance rules.

Release evidence

The freeze audit links every reviewed area to executable or documented evidence: conformance/v1.0/freeze-audit.md.

The clean-room Rust codec remains dependency-free and decodes/encodes canonical bytes without importing the primary protocol types. The backend conformance suite exercises the semantic provider contract without wire, virtqueue, OS, or vendor dependencies.

Architecture

Scope

The repository contains both the portable protocol implementation and concrete host backend work. The portable layers define the semantics that must agree across every guest transport, virtual machine monitor, host operating system, and hardware provider. Platform integrations are adapters; they implement those semantics without owning or redefining them.

Protocol 1.0 does not assign a Virtio device ID. It also does not standardize vendor executable formats: artifact format identifiers and target words remain opaque to the transport.

The normative terminology, object model, and compatibility rules live in specification.md, with exact layouts in wire-abi.md and command queue rules in virtqueue.md. This document explains the implementation boundaries that preserve those rules. The portable trust assumptions and denial-of-service bounds are mapped in threat-model.md.

Load-bearing invariants

Wire safety

Every wire structure is fixed-width, little-endian, pointer-free, and valid at byte alignment. zerocopy derives reject layouts containing implicit padding. Raw numeric opcodes and status values are validated before conversion to semantic types, preserving forward compatibility without invalid Rust enum discriminants.

Request and response buffers are untrusted. The command-frame preflight validates descriptor direction, total byte counts, reserved-zero fields, array multiplication, configured limits, and command-specific response capacity before a decoded request can reach semantic dispatch.

Object identity and ownership

Guest-visible handles are opaque u64 values. The device table combines a slot number with a device-instance namespace, resource-kind tag, and generation. Removing an object increments its generation; a slot is permanently retired before generation overflow. Therefore stale, wrong-kind, cross-device, and pre-reset handles never alias a live object when each reset epoch receives a fresh namespace.

DeviceState composes typed context, buffer, program, execution-queue, and event tables. Context records retain live-child counts, while event records retain queue, program, and buffer references until event destruction. Destroying a context with children or a referenced buffer, program, or queue returns BUSY.

The state model has no internal locks or interior mutability. Every transition requires exclusive access, so a future concurrent command engine has one outer synchronization boundary and no nested resource-lock order. Creation checks quotas and reserves fallible table capacity before invoking a provider. Release moves a handle to an explicit Releasing state and either commits removal or restores the same live ID after a rejected provider release.

CommandProcessor preserves that ownership model rather than placing an atomic or lock in each record. One mutable processor owns one backend and one object graph. A transport may move that owner between workers or serialize admission at its queue boundary; provider-native asynchronous work continues behind borrowed handles and event objects. Atomics belong in provider completion tokens or the concrete Virtio status publication path where state is genuinely shared, not in portable object lookup or reference counting.

Provider releases have an explicit failure boundary too. A rejected release returns the still-live handle for retry. An indeterminate release invalidates the guest ID and requires device recovery; the adapter must never guess that the resource is either safe to reuse or safe to free.

Reset and quarantine

The transport stops fetching chains and publishing completions before handing exclusive ownership to CommandProcessor::reset. The processor then makes one bounded pass: events first, followed by execution queues, programs, buffers, and contexts. Pending events are cancelled only when the backend advertises cancellation; no reset path spins, waits, or creates a background executor.

A completely drained graph receives a fresh ObjectNamespace and may continue with the same backend. Any unresolved pending event, rejected reset release, indeterminate release, device loss, or accounting contradiction produces BackendDiscardRequired. That result reports both resources released during the pass and resources still represented or previously orphaned in quarantine. The result is sticky: later reset calls make no provider calls, and the complete processor/backend instance must be discarded rather than reattached to newly initialized queues.

This keeps synchronization at the existing owner boundary. Reset needs no per-record atomics or locks; provider completion tokens remain responsible for the cancellation/completion race.

Submission acceptance

A rejected submission guarantees that the backend accepted no execution and retained no resources. When acceptance cannot be determined, SubmitFailure::Indeterminate carries an event. That event is the ownership token for all referenced resources until it becomes terminal and is destroyed.

This distinction must survive every provider and transport adapter. Collapsing the two cases into a single error would permit use-after-free during device-reset and timeout races.

Time

Wire timeouts are relative nanosecond durations measured from device admission. Guest and host monotonic clocks do not share an epoch, so absolute guest timestamps are never compared with a host clock. A zero timeout means infinite.

Memory

The baseline contract uses device-owned buffers plus bounded read/write transfers. External memory, shared mappings, and fences are optional features because their ownership, cache visibility, and synchronization rules differ by transport and host OS. Adding them requires a separate invariant and threat-model pass.

Provider-owned shared memory is distinct from external memory. MemoryDomain::Shared requests one allocation that the provider can access through a host mapping and bind directly for accelerator execution. It does not expose a guest address or platform handle and does not imply cross-process sharing or implicit cache coherence.

The allocation result reports verified backing properties, actual retained bytes, and actual alignment separately from the provider-native handle. Logical buffer bytes may be smaller than a page-, section-, or device-aligned backing allocation. The command engine retains those facts in its buffer record for compatibility checks and resource accounting, while submission passes only borrowed native handles. This lets the device reject a dishonest or degraded allocation before it becomes guest-visible without adding metadata lookup or boxing to the execution hot path.

ResourcePolicy supplies mandatory host-private aggregate limits without duplicating the wire or provider capability limits. The command engine prechecks logical buffer bytes, reconciles the provider’s actual backing charge before publishing an ID, and charges program resident storage before provider invocation. Charges remain represented across rejected releases; indeterminate ownership transfers them to quarantine and makes the entire backend discard-only. ResetReport therefore reports both object counts and exact released or quarantined retained bytes.

Bulk byte payloads cross the semantic boundary through ByteSource and ByteSink. Both abstractions support checked random access over segmented storage and an optional contiguous view. A command engine can therefore expose validated descriptor-backed regions directly: a provider streams them into final storage, while an already contiguous payload remains one borrowed slice.

Queue model

Command virtqueue zero is the baseline bidirectional transport queue. One descriptor chain contains device-readable request bytes and device-writable response bytes. Completion may be out of submission order, keyed by the request ID.

virtio-accel-transport defines the queue boundary without choosing a ring implementation or guest memory library. Driver publication transfers ownership of a complete chain until used-ring consumption or reset returns it. Device pop returns a non-Copy chain consumed by completion. Queue identities include a monotonic reset epoch, so stale completion is rejected before guest bytes, a used element, or a notification can be published.

Publication and completion are release boundaries for request and response bytes; the corresponding pop operations are acquire boundaries. Notification enablement includes the required atomic recheck, represented explicitly as Idle or WorkPending. Concrete adapters may use atomics and atomic pointers for shared indices and ownership transfer, but the portable traits require no lock, thread, executor, or global runtime.

Queue configuration may reserve storage bounded by the validated queue size. Every steady-state operation is nonblocking and heap-allocation-free: publish, pop, complete, notification suppression, notification recheck, and reset. Reset may move already-owned storage into a reclamation result but does not allocate or wait for a peer.

virtio-accel-guest owns one portable driver queue without internal synchronization. It preallocates a caller-selected number of tracking slots, writes fixed prefixes directly into caller-owned chains, and retains bulk read responses in reclaimed chain storage. Prepared transfer and artifact tails are published without another payload copy. Non-Copy typed handles carry the queue reset epoch; release operations consume them and report whether a failure is retryable, invalidated, indeterminate, or an opaque unknown status.

virtio-accel-split-queue is the deterministic executable implementation of that boundary. It preallocates descriptor ownership, chain records, available entries, and used entries at configuration. Its split-ring counters use wrapping u16 arithmetic, direct chains retain their scatter/gather buffers, and profile-invalid flags or indirect descriptors are classified before byte access. Driver and device operations take &mut SplitQueue, so ordinary ring state needs no atomics, atomic pointers, locks, or compare-and-swap loops. Non-atomic Rc ownership keeps a reset reclamation token and a consumed device token tied to the same buffers; one AtomicU64 reset epoch is the only synchronization primitive, because it must invalidate byte ports already issued to a device token before driver ownership is reclaimed.

The baseline SUBMIT command returns an event object; POLL_EVENT provides portable progress without requiring unsolicited device writes. Optional multi-queue and event-queue features are reserved for later validation. Split and packed virtqueue mechanics belong to transport adapters, not the command engine.

An accelerator execution queue is a separate context-scoped backend object. It never denotes a Virtio queue index.

Performance posture

The semantic hot path uses associated handle types and borrowed binding slices, avoiding trait-object dispatch and per-binding boxing. Wire decoding will operate directly over validated descriptor-backed regions. Object lookup is constant time and bounded by advertised limits. The queue ports add no allocation or copy to the steady-state path; mapping implementations can present borrowed segmented byte ports directly to the command processor.

Accelerator deliberately places no Send or Sync bound on the backend or its associated handle types. The reference command engine specializes over one concrete backend and owns it behind one mutable admission boundary. A provider can therefore preserve thread-affine native handles without boxing, atomics, or locks. Providers that opt into cross-thread auto traits own the synchronization needed by their actual shared state; the portable object graph does not speculate by adding it to every handle.

The source-level trait may be erased only after an adapter fixes all associated handle types. Stable binary plugin loading, cross-module allocation ownership, and an erased handle ABI are deliberately outside v1. A future plugin adapter can add those policies without changing static providers or weakening the submit and release contracts.

Backend metadata is fetched and validated once before object tables are constructed. Assigned reserved capabilities, a missing usable memory domain, and zero advertised limits fail construction. Unknown capability bits remain available for diagnostics but do not select operations. The command engine then uses the cached capabilities and limits to reject unsupported work before provider invocation.

WRITE_BUFFER and READ_BUFFER are the baseline’s explicit content-copy boundaries. Device-local memory may require bounded staging during those operations. Allocation, submission, polling, and release do not receive permission to copy a bound buffer merely because a native import or binding path is inconvenient.

Every buffer declared for program input, output, or mutable state reports BufferProperties::DIRECT_BINDING. This means a compatible submission binds that exact allocation without copying the bound range to or from a different allocation. A backend that cannot honor the requested placement and direct-binding contract rejects allocation; a program-specific alignment or format mismatch rejects submission as INCOMPATIBLE. Neither path may silently degrade to a bounce buffer.

The mutable side of an explicit write receives &mut Buffer, allowing implementations to use ordinary provider handles and mappings rather than forcing interior mutability or a lock into every buffer. Submission remains borrowed and allocation-free in the semantic API.

Program artifacts use the same byte-source abstraction. Program loading is a slow lifecycle path, so an object-safe source is an acceptable dispatch cost; forcing a frame-sized allocation and copy for every segmented artifact is not. Providers can parse a contiguous borrowed artifact in place or read segmented bytes directly into final resident storage.

Zero-copy guest-memory imports are deliberately deferred rather than pretending that DMA-BUF, Windows shared handles, and other mechanisms have identical lifetime or coherency semantics. When external memory is specified, fallback staging will require explicit negotiation and copy accounting; it will not weaken the provider-owned direct-binding rule. The non-normative external-memory handoff design fixes the proposed ownership, visibility, reset, isolation, fallback, and conformance boundaries without assigning or advertising the reserved protocol feature.

Backend fast-path checklist

A provider implementation should make the native buffer handle own or reference everything needed to reuse the allocation efficiently: the final backing object, device address or import, host mapping when present, alignment facts, and synchronization state. Native mapping or import setup belongs at allocation or another amortized lifecycle boundary, not in every submission.

The intended steady-state submission path is a bounded walk over the borrowed binding slice, validation of program-specific compatibility, native handle/address binding, and queue admission. It does not allocate per binding, assemble a second binding array with owned payloads, or copy tensor contents. Small command and metadata writes are not buffer staging and remain provider-specific.

The command engine uses three bounded metadata allocations for SUBMIT: decoded bindings, retained buffer IDs, and borrowed native binding references. Duplicate-slot detection sorts the decoded allocation in place; event state takes ownership of the retained-ID allocation. No allocation owns buffer contents, boxes individual bindings, or survives event reclamation except the retained ID list required for exactly-once reference release.

performance.md owns quantitative evidence: explicit-transfer bytes, staged bytes and allocations, submission allocations, retained memory, and host preparation versus device execution time. A backend should be diagnosable when it misses the intended path rather than requiring a profiler to discover an undocumented copy.

Deterministic reference execution

virtio-accel-mock::reference defines a test-only artifact envelope for executable backend tests. Its fixed 24-byte payload carries an artifact version, a binding-ABI version, an operation, binding slots, one byte operand, and reserved-zero bytes. The mock additionally requires its provider-owned format ID, target identity, and exact resident charge before accepting a program. These values and payload bytes are implementation fixtures, not additions to the normative accelerator ABI; production command and transport layers continue to pass artifact formats, targets, and payloads through opaquely.

The reference operations are a lifecycle barrier, equal-length copy, fill, and in-place XOR. Each operation validates its exact slot and access contract before event admission. Buffers use shared atomic-byte backing so an accepted event retains only fixed operation metadata, ranges, and atomic reference-counted backing pointers. Submission does not lock, stage buffer contents, or allocate an owned binding mirror. Explicit segmented transfers use a fixed-size stack window rather than coalescing the complete transfer.

Events remain pending until the harness calls complete. A single compare-exchange chooses among execution, cancellation, and injected device loss; after execution starts, cancellation and device loss report Busy. Completion publishes buffer mutations before the terminal event state, while the harness controls latency and completion order by deciding when and in which order to complete accepted events.

Deterministic fault injection

virtio-accel-mock::fault::FaultAccelerator<A> wraps any backend with a validated explicit fault script. Each step selects one Accelerator method, its one-based call occurrence, and a compatible action: error before invocation, error after successful invocation, rejected ownership transfer, indeterminate ownership transfer, or a persistent terminal event completion. Post-create errors synchronously release the newly created provider resource before returning the injected error. Post-admission submission errors always return an event as indeterminate rather than misreporting accepted work as rejected.

The wrapper assigns harness-local IDs to contexts, buffers, programs, queues, and events. Its audit snapshot records every method call, injected action, release attempt, rollback, remaining script step, and last known provider ownership state. It also tracks context children and the resources retained by each event, rejecting double release, use after release, parent release with live children, and release of an event-retained resource before the wrapped provider sees the invalid call. Live or indeterminate resources become clean only after successful release or an explicit backend-discard acknowledgement.

Fault scripting is single-threaded test control built on Rc<RefCell<_>>. It may allocate a mapped binding vector to interpose on submission and is intentionally outside production performance claims; the wrapped backend and ordinary command path retain their synchronization and copy contracts.

Reusable backend conformance

virtio-accel-conformance depends only on virtio-accel-core and runs each semantic case against a fresh backend instance. A provider supplies one executable target fixture plus test-only progress and optional resource-accounting hooks. Stable case IDs cover metadata, reserved intent, every advertised memory domain, segmented transfers and artifacts, bounds, permissions, bindings, context isolation, admission, pending-event release, terminal stability, finite timeout, and both cancellation race outcomes.

Mandatory cases cannot skip. Memory-domain and cancellation cases skip only when the corresponding semantic capability is absent, and reports preserve the explicit reason. Accounting, when exposed, is sampled before and after every case and counts both live and indeterminate native resources. The reference backend passes the suite directly and through FaultAccelerator; intentionally broken adapters prove that each major contract area produces a named failure.

The suite is portable std test tooling rather than a production dependency. It preserves static backend and handle dispatch and introduces no wire types, virtqueues, host APIs, threads, or global synchronization. The backend implementer guide maps each case to trait obligations and separates semantic evidence from the quantitative allocation and copy budgets in performance.md.

Its numerics module complements lifecycle conformance with checked-in TOSA graphs and shared FP32, FP16, BF16, FP8E4M3, FP8E5M2, INT8, INT32-accumulator, and packed INT4 oracles. Identity edge values, non-square batched matrix multiplication, NHWC max pooling, explicit FP8-to-BF16 CAST, and signed INT32-to-INT8 RESCALE are shared cases. A backend runs every case it advertises and rejects unsupported profiles, extensions, and dtypes at program admission; provider-specific graphs cannot stand in for the shared bytes.

The exact integer oracle is implemented in portable Rust from the TOSA scaling and accumulation rules. The first production integer tiers use direct INT8 boundaries for identity and widen INT8 MATMUL operands to INT32 before explicit zero-point subtraction and INT32 accumulation. XDNA also implements the released per-tensor scale32/single-round RESCALE back to signed INT8. Core ML encodes MATMUL as an ML Program on macOS 26+; OpenVINO encodes the same MATMUL semantics as IR Convert, Subtract, and MatMul nodes; XDNA specializes the exact expressions into AIE kernels. None of these paths converts through floating point.

Production lowering boundaries

The first production artifact path is now TOSA-to-Core ML. The facade, guest, transport, and device engine pass the TOSA format ID, target words, and opaque bytes without importing Core ML. The Core ML adapter depends inward on virtio-accel-tosa, verifies and analyzes one static TOSA graph, derives ordinal input/output slots, emits a Core ML model, and confines temporary model compilation plus Foundation/Objective-C state to the host-native boundary. Program-visible buffers remain exact provider allocations from load through asynchronous prediction; lowering never authorizes staging at submission.

The provider-neutral analysis is intentionally the shared seam rather than a second repository-wide graph IR. A host backend may lower its supported subset directly, reject unsupported semantics at program load, and add capability coverage without making its native SDK or model representation a dependency of the portable stack.

Before constructing partitions, a host scheduler may query the optional virtio_accel_tosa::TosaCapabilityProvider interface. Descriptors remain outside DeviceInfo and the wire contract: they distinguish exact targets, dtype roles, operator/attribute constraints, static-shape limits, and runtime-condition policy. They are conservative preflight data, while the provider retains final authority at load_program for concrete graphs and current resources.

The Intel OpenVINO provider consumes the same verified model and analysis. The adapter lowers one static TOSA graph to an in-memory OpenVINO IR document plus weights blob, compiles it for one enumerated NPU, GPU, or CPU device with the accuracy-preserving execution hint, and binds exact provider allocations as host-pointer tensors from load through asynchronous inference; completion is accepted only when the runtime reports the caller’s own allocation as its output storage. Its validation runs the backend conformance suite and the shared numerical TOSA corpus on every enumerated device and reports the direct-binding and explicit-transfer counters through the conformance diagnostics hook. This keeps the Core ML and Intel paths on one bounded TOSA contract without requiring either provider to adopt the other’s native graph representation.

The AMD XDNA provider follows that OpenVINO boundary with one ecosystem-forced compiler divergence. Safe Rust verifies and analyzes one static TOSA graph, matches only the advertised BF16, explicit FP8-to-BF16 storage-conversion, or exact integer templates, and invokes the pinned MLIR-AIE/IRON compiler as a bounded subprocess during program load or offline catalog population. Guest TOSA bytes never enter Python: the subprocess receives a small validated specialization, and its measured toolchain identity participates in the content-addressed cache key. A serving host may load the resulting crate-local XDNP artifact without Python or a compiler. Each artifact carries an exact per-slot byte/access plan so fixed DMA extents are checked again at submission.

Native execution uses the amdxdna HRX C ABI rather than XRT. One process-wide device owner creates per-backend streams; each backend serializes accepted work through a bounded ring and worker that bridges blocking HRX synchronization to latched, nonblocking event polling. Persistent HRX mappings are the exact allocations bound to dispatch, and diagnostics prove that submission introduces no bounce allocation or transfer. HRX exposes no bounded cancellation, so finite deadlines are rejected before acceptance and a two-tier poison/watchdog model handles device loss. These choices preserve OpenVINO’s static admission, direct-binding, capability, and conformance seams; the helper process, serialized stream, and XDNA-specific local-memory envelopes are documented hardware/runtime constraints rather than portable API changes or silent fallback paths.

The Vulkan provider is the first GPU-class consumer of the seam and keeps the data plane graph-shaped (ADR 0001): one admitted TOSA graph becomes one sequence of compute pipelines created at load from crate-authored SPIR-V kernels specialized by validated shape constants, so guest bytes never reach a driver’s shader compiler (ADR 0003, ADR 0007). Every kernel addresses tensors through one descriptor — an array of storage buffers holding the submission’s bound slots and a per-program arena for constants and intermediates — so one module per kernel serves every binding layout, and a whole graph is recorded into one command buffer with compute barriers between dependent dispatches. Buffers are dedicated VkDeviceMemory allocations bound directly as storage buffers; host-visible domains stay persistently mapped, and device-local memory is reached only through bounded staging inside the explicit transfer calls. Each context owns a bounded ring of command buffers, fences, and descriptor sets; vkQueueSubmit2 success is the admission boundary and vkGetFenceStatus is the whole completion path, so no worker thread bridges the runtime. Device loss poisons the instance. The backend runs the conformance suite and the shared FP32 operator corpus on every device it enumerates; the FP32 operator tier is verified on the Mesa lavapipe CI lane, on Intel ANV (Arc 140V), and on Apple M4 via MoltenVK. The FP16 tier (ADR 0008) — the same operators over packed binary16 storage with crate-owned conversions and binary32 evaluation, advertised on every device — has executed its corpus on Apple M4 via MoltenVK, Intel Arc LNL (Mesa ANV), and AMD Radeon 860M (RADV), and the lavapipe CI lane exercises it on every change.

The Qualcomm adapter uses the same seam. Its safe planner admits 41 of the 42 floating-point operators shared by Core ML and OpenVINO, including owned constants/data movement, FP16 unary and binary computation, BOOL comparison/selection/logical tensors, and INT32 indexing results. ERF is the explicit exception because QAIRT 2.49 exposes no public QNN ERF operation. A separate exact integer target supports INT8 identity and zero-point-aware INT8 MATMUL with INT32 output. Typed tensor plans carry scalar size, owned constant bytes, and QNN scale-offset metadata through an owned ABI, so submission range checks use exact one-, two-, or four-byte element storage. QNN static parameters, descriptor arity, axis/permutation vectors, constant byte lengths, and generated-tensor counts are checked before provider calls. Unsupported native pool, reverse, and product-reduction forms decompose into HTP Gather and elementwise nodes rather than falling back to a host runtime. FP32 remains rejected because a pinned v73 probe returned FP16-rounded results, and ambiguous generic FLOAT8 cannot be advertised as either TOSA format. With a complete QAIRT/QNN C development package on Windows ARM64, the audited boundary creates and finalizes QNN HTP graphs, binds exact caller buffers, and publishes completion from a bounded worker. Driver-only and AppBuilder/Genie installations are intentionally insufficient to enable that boundary.

The portable command engine depends on virtio-accel-proto, virtio-accel-core, and the transport-neutral region metadata re-exported by virtio-accel-device. Its baseline processor:

  1. Decodes one bounded request from abstract readable/writable byte regions.
  2. Maintains typed object records and context dependency counts.
  3. Translates wire types into validated semantic values.
  4. Passes transfer and artifact regions directly to backend byte ports.
  5. Produces a response without knowing about rust-vmm or a host operating system.

Submission/event retention, deterministic reset, the bounded split-virtqueue model, the no-std reference guest, deterministic reference execution, scripted ownership-boundary faults, and the reusable backend conformance suite now complete both portable queue endpoints and the provider contract evidence. The threat model and enforceable aggregate resource policy close the security model. Coverage-guided fuzzing exercises protocol decoding, descriptor segmentation, and bounded stateful command sequences. The deterministic state-model replay suite extends that coverage with random context/resource graphs, stale-ID probes, submission/cancel/completion/reset race schedules, and minimized replay output for failures. A thin rust-vmm adapter supplying virtio-device, virtio-queue, and vm-memory integration remains a later platform layer, as do Linux vhost-user and an in-kernel guest driver.

Portability and continuous integration

The portable v1 promise is enforced by .github/workflows/ci.yml. This file is the source of truth for required checks; the table below explains why each job exists.

The minimum supported Rust version (MSRV) is 1.85.0. Rust 1.85 is the first release supporting the Rust 2024 edition selected by the workspace. Every package inherits the same rust-version.

CI matrix

JobToolchain and targetsContract enforced
style-and-apiCurrent stable on UbuntuFormatting, complete normative-requirement ledger, release-policy invariants, Clippy with warnings denied, and warning-free public docs
native-testCurrent stable on Ubuntu, macOS, and WindowsAll workspace unit, integration, target, feature, and documentation tests, runnable examples, and release-profile checking
openvino-host-testCurrent stable on Ubuntu with a pinned Intel OpenVINO runtimeThe real (probed) OpenVINO backend: lint, unit, integration, semantic-conformance, and example runs against the CPU plugin
vulkan-lavapipe-testCurrent stable on Ubuntu with the Mesa lavapipe ICD pinnedThe native Vulkan 1.3 backend: lint, unit, integration, semantic-conformance, the shared FP32 operator corpus, and example runs without allowing a placeholder skip
msrvRust 1.85.0 on UbuntuEvery workspace target and test continues to compile at the declared MSRV
portable-targetStable aarch64-unknown-none, riscv64gc-unknown-none-elf, and wasm32-unknown-unknowncleanroom, proto, transport, and core remain no_std; guest, split-queue, device, and facade layers require at most alloc; Wasm also checks the std reference crates
feature-sets-and-dependenciesStable on UbuntuEvery Cargo feature combination plus dependency and std/alloc leakage guards for the portable codecs, queue ports, and core
dependency-policyCargo-deny with Rust 1.85.0Advisories, yanked crates, duplicate versions, wildcard requirements, licenses, and dependency sources
fuzz-smokePinned nightly on UbuntuShared harness tests plus bounded smoke iterations over generated seed corpora and committed minimized regressions for every fuzz target
publish-dry-runStable on UbuntuEvery published crate packages in documented order into an isolated local registry, and each one builds, tests, and documents from its own extracted tarball rather than from the workspace

The native jobs intentionally use GitHub-hosted *-latest images so runner security and supported host versions advance without changing the project’s semantic target list. Release evidence records the concrete runner versions used for the release.

Crate portability tiers

Every crate below is published to crates.io, so its tier is a promise to downstream users rather than an internal convention: a crate may not move to a less-portable tier without the release-note entry and portability review required by release-policy.md. The std-reference tier is the ceiling for the reference and conformance crates, not a licence for host-OS or vendor-specific APIs — third parties depend on virtio-accel-conformance and virtio-accel-mock directly, and they must keep working on any platform the portable crates support.

TierCratesAllowed runtime surface
core-onlyvirtio-accel-cleanroom, virtio-accel-proto, virtio-accel-transport, virtio-accel-corecore; the clean-room codec and transport ports have no normal/build dependencies, while proc-macros used by other crates may execute with std on the build host
alloc-portablevirtio-accel-guest, virtio-accel-split-queue, virtio-accel-device, virtio-accel-tosa, virtio-accel-tosa-build, virtio-accelcore + alloc; no OS, filesystem, sockets, threads, or host synchronization
std-referencevirtio-accel-mock, virtio-accel-conformancePortable std; no host-OS or vendor-specific API
host-nativevirtio-accel-coreml, virtio-accel-openvino, virtio-accel-hexagon, virtio-accel-xdna, virtio-accel-vulkanCore ML/Foundation on macOS 14+, the OpenVINO C runtime (libopenvino_c 2026.x), the complete QAIRT/QNN C SDK on Windows ARM64, the amdxdna-native HRX runtime (libhrx), or a dynamically loaded Vulkan loader (via ash) on the build.rs-enumerated host targets; a compile-only unsupported-platform or unsupported-runtime placeholder elsewhere

No host-native crate is a dependency of the facade or any portable layer. The Core ML crate’s Objective-C bridge and TOSA-to-Core ML model compilation are built only when the Cargo target is macOS. The Linux, Windows, and Wasm workspace jobs compile the placeholder API and backend-local lowering utilities, while the macOS native job executes its real model and semantic-conformance tests. An accessible Apple Neural Engine is required to construct the real backend; macOS runners without one skip execution after checking that availability through Core ML.

The OpenVINO crate’s boundary is a build environment rather than a target operating system: its build script probes pkg-config for openvino and compiles the native FFI modules only on success. VIRTIO_ACCEL_OPENVINO=1 turns a missing runtime into a loud build failure, =0 forces the placeholder, and VIRTIO_ACCEL_OPENVINO_LIB_DIR links installations that ship no pkg-config metadata. The default native-test runners have no OpenVINO and compile the placeholder plus the portable TOSA-to-IR encoder; the dedicated openvino-host-test job installs a pinned runtime and exercises the real backend against the CPU plugin. NPU and GPU devices additionally require the Intel Level Zero NPU driver or the Intel OpenCL/Level Zero GPU runtime on the host; hosts without an inference device skip execution after enumeration.

The Hexagon crate exercises its SDK-free placeholder and strict FP16/BOOL/INT32 plus INT8 TOSA graph planner in portable CI. Its parity test compares the real Core ML/OpenVINO 42-operator surface and allowlists only ERF; portable fixtures lower every other shared operator without QAIRT. Its build script distinguishes a complete public QAIRT/QNN development installation from driver-only and AppBuilder/Genie bundles by requiring QnnInterface.h and the Windows ARM64 QnnHtp import library. VIRTIO_ACCEL_HEXAGON=0 forces the placeholder; VIRTIO_ACCEL_HEXAGON=1 makes missing requirements a build failure; VIRTIO_ACCEL_QNN_SDK_ROOT/QNN_SDK_ROOT and VIRTIO_ACCEL_QNN_LIB_DIR select the SDK. On the pinned Windows ARM64 hardware tier, backend-local tests execute numerical fixtures for all 41 advertised operators plus INT8 identity and zero-point-aware INT8 MATMUL through QNN HTP, followed by the reusable semantic suite. A public hardware CI lane remains unavailable, so the README publishes the exact manual replacement commands.

The AMD XDNA crate compiles its portable admission surface (lower, including the TOSA IDENTITY, MATMUL, MAX_POOL2D, explicit FP8-to-BF16 CAST, and exact INT8 IDENTITY/MATMUL/RESCALE admissions and its advertised Target constants), the portable precompiled-artifact codec, and a placeholder on every host; portable CI unit-tests admission. HRX exposes a plain C ABI, so its build script has no cc/CMake step; it enables the native modules (HRX FFI and the Accelerator implementation with its serialized dispatch worker) only when an amdxdna-native HRX prefix (VIRTIO_ACCEL_HRX_DIR/HRX_DIR) carries both HRX headers — the amdxdna header must declare hrx_amdxdna_executable_create, whose absence marks an older, incompatible libhrx generation — and lib/libhrx.so. VIRTIO_ACCEL_XDNA=0 forces the placeholder, VIRTIO_ACCEL_XDNA=1 makes a missing or incomplete runtime a build failure, and VIRTIO_ACCEL_HRX_LIB_DIR links a bare lib directory. No standard locations are scanned, keeping the runtime pin authoritative. On Unix, the bounded compiler-helper subprocess and compile_artifact remain available without HRX for offline catalog population; they require only the separately pinned VIRTIO_ACCEL_AMDXDNA_TOOLCHAIN. On the reference machine with that toolchain and an NPU, backend-local tests run a DMA passthrough, a compiled TOSA BF16 IDENTITY, a bit-exact BF16 → FP32 MATMUL, BF16 NHWC MAX_POOL2D against the shared bit-exact oracle, both FP8E4M3/E5M2 → BF16 CASTs against shared bit-exact oracles, and the shared exact INT8 identity, nonzero-zero-point MATMUL, and signed INT32-to-INT8 RESCALE oracles, plus the shared semantic conformance suite end to end; compilation invokes the aiecc helper as a bounded subprocess (never a Cargo dependency). The conformance run includes provider-resource accounting and direct-binding diagnostics; a feature-gated on-metal fault suite proves definite device loss, the 120-second tier-2 watchdog state machine with a shortened test deadline, rejected pending releases, stable terminal polling, and nonblocking poisoned-instance discard. MAX_POOL2D mirrors OpenVINO’s propagating-NaN and zero-padding semantics, but XDNA admission is deliberately narrower: batch 1, kernel and stride dimensions at most 8, and at most 8,192 input-plus-output elements so the full tensors fit in the AIE2P worker’s local memory. A public hardware CI lane remains unavailable. The INT8 capability preserves OpenVINO’s exact CONST/IDENTITY/MATMUL semantic baseline and adds RESCALE as a documented operator-surface expansion. The expansion uses released TOSA semantics and is needed to return an INT32 accumulator to signed INT8 without host arithmetic; XDNA admits only per-tensor scale32/SINGLE_ROUND forms proven against the shared exact oracle. An XDNA-specific one-core memory envelope still applies. AIE DMA’s four-byte transfer granularity is represented as explicit per-slot padding in the compiled artifact, never as hidden submission-time staging. Tile-compatible MATMUL shapes widen/subtract zero points on the NPU and use the native 4x4x8 INT16 matrix unit; smaller shapes and exact 64-bit RESCALE arithmetic use scalar on-NPU kernels. RESCALE clears the at-most-three-byte padded output tail on the NPU.

The Vulkan crate binds Vulkan 1.3 through the pinned ash crate and executes admitted TOSA graphs on crate-authored SPIR-V compute kernels specialized at load_program; the admitted tier is the FP32 operator set shared with Core ML and OpenVINO, executed graph-at-a-time (ADR 0007). Unlike the SDK-probing backends, there is nothing to detect at build time — ash loads the platform’s Vulkan loader dynamically at run time — so the va_vulkan cfg enumerates the host target operating systems (Linux, Android, Windows, macOS) and runtime loader or device absence surfaces as InitError. VIRTIO_ACCEL_VULKAN=0 forces the placeholder everywhere; VIRTIO_ACCEL_VULKAN=1 makes an unsupported target a loud build failure; unset is auto (ADR 0002 in docs/adr/). Buffers are dedicated allocations bound directly as storage buffers; Host and Shared domains are persistently mapped host-coherent memory, Device is device-local memory reached only through bounded staging inside the explicit transfer calls (ADR 0005). The dedicated vulkan-lavapipe-test job forces the native build, requires an enumerated device, and pins Mesa’s software ICD before running Clippy, all Vulkan tests, and the example. This keeps the run deterministic while ensuring loader discovery and the real Vulkan lifecycle execute in CI.

Concrete VMM, kernel, OS, and vendor adapters are outside the portable-v1 milestone and must not become default dependencies of a portable crate.

Cargo feature policy

CI runs cargo hack check --feature-powerset --no-dev-deps across the workspace. Features must be additive: disabling default features may remove convenience behavior but must not select a different protocol interpretation.

The portable dependency guard inspects normal and build target features for virtio-accel-cleanroom, virtio-accel-proto, virtio-accel-transport, and virtio-accel-core, virtio-accel-guest, virtio-accel-split-queue, and virtio-accel-tosa; test-only development dependencies are intentionally outside the target runtime graph. It additionally proves that the clean-room codec and transport ports have no normal or build dependencies at all. A dependency’s host-side derive macro may use std, but the target graph for these crates must not enable a dependency feature named std or alloc.

The official FlatBuffers runtime uses a pure-Rust build script with rustc_version and semver to detect the compiler. Those host-only dependencies use std; the guard checks the TOSA crate’s normal target graph separately and still forbids a std or alloc dependency feature there.

Dependency policy

deny.toml permits only crates.io dependencies and the workspace’s path dependencies. It rejects:

  • known advisories and yanked releases;
  • unmaintained direct workspace dependencies;
  • unsound dependencies;
  • multiple resolved versions of the same crate;
  • wildcard version requirements; and
  • licenses outside Apache-2.0, BSD-2-Clause, MIT, and Unicode-3.0.

GitHub Actions are pinned to immutable commit SHAs with their human-readable release versions kept in comments. Standalone CI tools are pinned too: cargo-hack 0.6.45, cargo-fuzz 0.13.2 on nightly-2026-07-13, and cargo-deny 0.20.2 (bundled by the pinned cargo-deny action).

Local verification

The host-independent checks can be reproduced with:

cargo fmt --all -- --check
python3 ci/check-normative-requirements.py --check
python3 ci/check-release-policy.py
cargo clippy --workspace --all-targets --all-features -- -D warnings
RUSTDOCFLAGS="-D warnings" cargo doc --workspace --all-features --no-deps
cargo test --workspace --all-targets --all-features
cargo test -p virtio-accel-device --test state_model
cargo test --workspace --doc --all-features
cargo run --example backend_conformance
cargo run --example reference_execution
cargo run -p virtio-accel-coreml --example tosa_coreml
cargo run -p virtio-accel-openvino --example tosa_openvino
cargo run -p virtio-accel-hexagon --example tosa_hexagon
cargo run -p virtio-accel-hexagon --example mock_classifier
cargo +1.85.0 check --workspace --all-targets --all-features
cargo hack check --workspace --feature-powerset --no-dev-deps
bash ci/check-portable-dependencies.sh
cargo deny --all-features check

The ordered publication dry run needs network access the first time, to vendor third-party dependencies into its local registry:

python3 ci/publish-dry-run.py

Add --keep to inspect the registry and the extracted per-crate sources afterwards.

Deeper deterministic state-model exploration can be run manually with:

VIRTIO_ACCEL_STATE_MODEL_SEED=9e3779b97f4a7c15 cargo test -p virtio-accel-device --test state_model deep_generated_object_graphs_match_the_reference_model -- --ignored

Fuzz smoke coverage can be reproduced with:

cargo test --manifest-path fuzz/Cargo.toml --lib --no-default-features
python3 ci/seed-fuzz-corpus.py
cargo fuzz run protocol_decode fuzz/corpus/protocol_decode fuzz/regressions/protocol_decode -- -runs=256 -max_total_time=20 -timeout=5 -rss_limit_mb=2048 -max_len=65536
cargo fuzz run descriptor_end_to_end fuzz/corpus/descriptor_end_to_end fuzz/regressions/descriptor_end_to_end -- -runs=256 -max_total_time=20 -timeout=5 -rss_limit_mb=2048 -max_len=65536
cargo fuzz run stateful_commands fuzz/corpus/stateful_commands fuzz/regressions/stateful_commands -- -runs=256 -max_total_time=20 -timeout=5 -rss_limit_mb=2048 -max_len=65536
cargo fuzz run guest_client fuzz/corpus/guest_client fuzz/regressions/guest_client -- -runs=256 -max_total_time=20 -timeout=5 -rss_limit_mb=2048 -max_len=65536
cargo fuzz run tosa_parse fuzz/corpus/tosa_parse fuzz/regressions/tosa_parse -- -runs=256 -max_total_time=20 -timeout=5 -rss_limit_mb=2048 -max_len=65536

Target checks require the corresponding Rust standard libraries:

rustup target add aarch64-unknown-none riscv64gc-unknown-none-elf wasm32-unknown-unknown
cargo check \
  -p virtio-accel-cleanroom \
  -p virtio-accel-proto \
  -p virtio-accel-transport \
  -p virtio-accel-core \
  -p virtio-accel-device \
  -p virtio-accel-guest \
  -p virtio-accel-split-queue \
  -p virtio-accel-tosa \
  -p virtio-accel \
  --target aarch64-unknown-none \
  --no-default-features

Device support matrix

The backend support table answers “which dtypes and programs does each backend admit”. This document answers the different question underneath it: which physical devices actually execute the work, and what exactly stops the rest from doing so.

The distinction matters because the host backends choose devices differently. Core ML and OpenVINO are runtimes that dispatch across a machine’s inference estate, not NPU drivers, so parts of this project already run on CPUs and GPUs. XDNA and Hexagon are single-device by design. Vulkan spans vendors and enumerates every suitable physical device, but each backend instance binds to one of them. Naming those differences explicitly is more useful than an “NPU” label that is true of the intent and only partly true of the code.

How to read this

Every row carries one of four statuses. They are claims about this repository, not about the hardware.

StatusMeaning
ValidatedExecuted on that device, with an evidence pin in-repo naming the host, driver, and runtime versions.
ReachableThe selection and dispatch path drives the device today with no code change, but no in-repo evidence pins that part.
One change awayNot reachable today. Each row names the single gate — a constant, a path, or a build condition — and what it would take.
Out of scopeNo path, and none implied by the current design.

“Reachable” is deliberately weaker than “supported”. It means the code will select the device and try; it does not promise the program admits, the numerics match, or the performance is sane.

The matrix

DeviceClassBackendHost OS / archStatus
Apple Neural Engine, Apple silicon (M-series)NPUcoremlmacOS 14+Validated — Apple M4, macOS 26.5.2 (performance.md)
Apple CPU, as Core ML per-operator placementCPUcoremlmacOS 14+Reachable — see Core ML
Apple GPUGPUcoremlmacOS 14+One change away — compute units are pinned
Intel Mac (CPU, AMD/Intel GPU)CPU / GPUcoremlmacOS 14+One change away — ANE gate refuses construction
Intel NPU, Core Ultra (Meteor Lake / Lunar Lake / Arrow Lake / Panther Lake)NPUopenvinoAny host with the runtimeValidated — arch 5010 (Panther Lake), OpenVINO 2026.4: the FP8 tier’s graphs compile and FP8 movement round-trips bit-exactly, 2026-09-18 (ADR 0009). Arch 40XX (Lunar Lake) refuses FP8 outright and the withholding is pinned the same day: the suite passes with the tier absent and the refusal asserted, so the backend probes at open rather than branding by arch; the float and INT8 tiers remain first device preference on every generation
Intel GPU, integrated Xe/UHD and discrete ArcGPUopenvinoAny host with the runtimeReachable — including indexed GPU.1
x86-64 CPU, IntelCPUopenvinoAny host with the runtimeValidated — openvino-host-test CI lane, OpenVINO 2026.3.0
x86-64 CPU, AMDCPUopenvinoAny host with the runtimeReachable — misreports vendor, see OpenVINO
ARM64 CPU (Apple silicon, Ampere, Raspberry Pi)CPUopenvinoAny host with an ARM CPU-plugin buildReachable — enumerates as CPU, misreports vendor
OpenVINO virtual devices (AUTO, MULTI, HETERO, BATCH)—openvino—One change away — resolution requires enumeration
AMD XDNA2 NPU, Strix Point / Strix Halo / Krackan PointNPUxdnaLinux, amdxdna driverValidated — PCI 1022:17f0 rev 0x20, Fedora 44 (baseline)
AMD XDNA1 NPU, Phoenix / Hawk PointNPUxdnaLinux, amdxdna driverReachable — ungated but unvalidated, see XDNA
Second and later XDNA NPUs in one hostNPUxdnaLinuxOne change away — device index is fixed at 0
Qualcomm Hexagon HTP v73, Snapdragon XNPUhexagonWindows 11 ARM64Validated — Snapdragon X126100, QAIRT 2.49 (baseline)
Qualcomm Hexagon HTP v75+, newer SnapdragonNPUhexagonWindows 11 ARM64Reachable — ungated, misreports v73
Qualcomm Adreno GPU / Kryo CPU via QNNGPU / CPUhexagonWindows 11 ARM64One change away — backend library path is fixed, and deliberately so
Snapdragon on Linux or AndroidNPUhexagon—One change away — build target gate
Intel Arc 140V, Lunar LakeGPUvulkanLinux x86-64Validated — Mesa 26.0.8 ANV (baseline); full FP32 operator tier suite passed 2026-09-06 (ADR 0007), FP8 operator tier 2026-09-18 (ADR 0009)
Intel Arc B390, Panther LakeGPUvulkanLinux x86-64Validated — Mesa 26.0.8 ANV, Vulkan 1.4.335; full device suite including the FP8 operator tier passed 2026-09-18 (ADR 0009)
AMD Radeon 860M, Krackan PointGPUvulkanLinux x86-64, amdgpuValidated — RADV Mesa 26.1.8, Vulkan 1.4.354; full FP32 operator tier suite passed 2026-09-08 in every advertised memory domain, and clean under Khronos synchronization validation
Apple M3 / M4 GPU, Apple siliconGPUvulkanmacOS, MoltenVK loaderValidated — M4 on MoltenVK 1.4.2, full FP32 and FP16 corpora 2026-09-17; M3 on MoltenVK, FP8 operator tier 2026-09-18 (ADR 0009)
Other Vulkan 1.3 compute devicesGPU / virtual GPU / CPUvulkanLinux, Android, Windows, macOSReachable — enumerated and selected at run time; no other hardware evidence pin yet
lavapipe / llvmpipe software ICDCPUvulkanLinux x86-64Validated — pinned by the vulkan-lavapipe-test CI lane and exercised by the full backend suite, including the FP32 operator corpus; the FP8 corpus also passed on llvmpipe (LLVM 21.1.8) on 2026-09-18 (ADR 0009)
No device, executed in software—mockAnyDeterministic in-memory reference; outside this vocabulary
Whatever a wrapped provider drives—vaccelAny (std)Pass-through; the wrapped backend decides

Per-backend detail

Apple Core ML (virtio-accel-coreml)

What selects the device. Nothing does, explicitly. The backend hands Core ML a fixed compute budget and Core ML places each operator itself:

configuration.computeUnits = MLComputeUnitsCPUAndNeuralEngine;

— coreml_bridge.m:277

That constant is the whole device policy. It grants the ANE and the CPU, and withholds the GPU.

The CPU is already in play. Because MLComputeUnitsCPUAndNeuralEngine includes the CPU, a model containing operators the ANE declines runs those operators on the CPU — silently, inside a submission that device_info() reports as AcceleratorClass::NPU. This is not a defect; it is Core ML’s documented placement model, and the crate README states it. It does mean the project’s first non-NPU execution path already exists and already ships, and that per-submission device attribution is not observable through the Accelerator contract.

What excludes everything else. Construction refuses any host without an ANE:

#![allow(unused)]
fn main() {
if unsafe { va_coreml_has_neural_engine() } == 0 {
    return Err(InitError::NeuralEngineUnavailable);
}
}

— macos.rs:383, backed by a MLNeuralEngineComputeDevice scan at coreml_bridge.m:121

This gate, not the framework, is what excludes Intel Macs — Core ML runs there, it simply has no ANE. It also excludes ANE-less VMs, which is why the CI example on macos-latest skips rather than fails.

The two changes.

  • Apple GPU: switch the constant to MLComputeUnitsAll. One line. It widens placement to the GPU without adding a device to select, so nothing above the bridge changes.
  • Intel Macs and ANE-less hosts: soften the ANE gate to a capability probe. Larger than it looks — identity.uuid (apple-coreml-ane) and identity.class (NPU) are compile-time constants at macos.rs:393 and would have to become runtime-derived to stay truthful.

Intel OpenVINO (virtio-accel-openvino)

The genuinely heterogeneous backend. Device selection is explicit, ordered, and already covers three classes:

#![allow(unused)]
fn main() {
let device = ["NPU", "GPU", "CPU"]
    .into_iter()
    .find_map(|preferred| { /* first enumerated match */ })
}

— native.rs:842

with_device (native.rs:855) pins one device by enumerated name or class prefix, so "GPU.1" selects the second GPU and "NPU" selects whatever NPU instance exists. matches_device (native.rs:634) makes prefix matching strict: "GPU" matches GPU and GPU.1, never GPUX.

Not an Intel-only backend. The build gate is a pkg-config probe for the runtime, not a target OS or vendor check (build.rs:35). Consequently the CPU plugin admits any x86-64 host including AMD, and an ARM CPU-plugin build enumerates CPU on Apple silicon or Ampere just as readily. The openvino-host-test CI lane installs OpenVINO 2026.3.0 on an x86-64 Ubuntu runner and executes the real path against the CPU plugin. Vulkan has a separate native CI lane pinned to the lavapipe software ICD.

Two truthfulness gaps. device_info_for (native.rs:642) hardcodes vendor_id: 0x8086, so an AMD or ARM CPU host reports itself as Intel. And a CPU device falls through to AcceleratorClass::OTHER, because the class enum defines OTHER, NPU, GPU, and DSP but no CPU. A guest cannot currently distinguish “a CPU” from “a device this backend does not recognize”. Adding AcceleratorClass::CPU is additive — the type is a #[repr(transparent)] u16 designed for exactly this.

The one change. Virtual devices (AUTO, MULTI, HETERO, BATCH) are unreachable because both constructors resolve requests against the enumerated device list, and the standard plugin set does not enumerate them. with_device("AUTO") therefore returns DeviceUnavailable. A pass-through for a known set of virtual names would unlock OpenVINO’s own scheduling — worth weighing against this project’s preference for one submission mapping to one identified device.

AMD XDNA (virtio-accel-xdna)

Single device, index zero. The process-wide owner takes the first device HRX reports and never enumerates further:

#![allow(unused)]
fn main() {
check(ffi::hrx_gpu_device_get(0, &mut device))
}

— native.rs:132

Multi-NPU hosts are therefore reachable only at index 0. The fix is mechanical — plumb an index through shared_device — but the HRX fork’s model is one process-wide device, so it is a design question rather than a parameter change.

No generation gate. Nothing checks the PCI ID. An XDNA1 part (Phoenix / Hawk Point) would be opened and driven, while device_info reports it as XDNA2 regardless:

#![allow(unused)]
fn main() {
uuid: *b"amd.xdna.npu\0\0\0\0",
vendor_id: 0x1022,
device_id: 0x17f0,
}

— native.rs:691

Program admission also applies the XDNA2 shape and local-memory envelope, so an XDNA1 host would most likely fail during compilation rather than produce wrong numerics — but it would fail confusingly, and it would misidentify itself first.

Linux only, by runtime. The build script requires libhrx.so and the amdxdna-native headers (build.rs:58), and the validated stack is the in-tree amdxdna driver exposing /dev/accel/accel0. Windows XDNA uses a different driver stack entirely and is not addressed.

Qualcomm Hexagon (virtio-accel-hexagon)

HTP only, by construction. The QNN backend library is a fixed path:

#![allow(unused)]
fn main() {
let path = root.join("lib/aarch64-windows-msvc/QnnHtp.dll");
}

— native.rs:81

QAIRT also ships QnnCpu, QnnGpu, and QnnDsp libraries, and parameterizing this path would reach the Adreno GPU and Kryo CPU. This exclusion is deliberate, not an oversight: the crate documents that SDK-free builds fail explicitly rather than fall back, so that a host silently executing on the CPU can never be mistaken for NPU evidence. Treat it as a policy to revisit consciously, not a gap to close.

No SoC gate. The build pins the target, windows + aarch64 (build.rs:31), but nothing pins the SoC. A newer Snapdragon with HTP v75 or v79 would load and run, reporting itself as v73 the whole time (uuid: *b"qualcomm-htp-v73", device_id: 73 at native.rs:457). In practice such a host also needs ADSP_LIBRARY_PATH pointed at its own DSP libraries — an environment concern the crate README covers — and the FP32/FP8 rejections recorded for v73 may not describe it correctly.

Linux and Android are a build gate, not a port. QAIRT ships aarch64 Linux and Android libraries. The windows/aarch64 assertion and the hardcoded lib/aarch64-windows-msvc path are the only two things naming the OS.

Vendor-neutral Vulkan (virtio-accel-vulkan)

One physical device per instance. The run-time-loaded Vulkan 1.3 path enumerates every device with a compute queue and synchronization2, then prefers discrete GPU, integrated GPU, virtual GPU, CPU, and other devices in that order. with_device selects an exact enumerated name instead:

#![allow(unused)]
fn main() {
let physical = devices
    .into_iter()
    .min_by_key(PhysicalDeviceRecord::rank)
}

— native.rs:1205

The identity is probed, not branded. UUID, vendor ID, and device ID come from Vulkan physical device properties. GPU-like devices report AcceleratorClass::GPU; a CPU ICD such as lavapipe reports OTHER, because the protocol 1.0 class set has no CPU value. Memory domains are likewise per-device: host-coherent Host is required, Shared is advertised only for a device-local and host-visible type, and Device only for device-local memory. Every submitted buffer remains a direct storage-buffer binding; staging occurs only during explicit reads and writes of device-local memory.

Current execution boundary. The advertised tier is the static 42-operator FP32 tier with BOOL/INT32 auxiliaries, plus the FP16 tier: the same operators over binary16 tensors, advertised on every device the backend opens — the kernels’ binary16 conversions are crate-owned integer and binary32 code, so the tier needs no device feature (ADR 0008). The provisional integer target is declared but not advertised. The native path and full backend conformance suite are validated on Intel Arc 140V through Mesa ANV, on llvmpipe/lavapipe, and on Apple M4 via MoltenVK; CI pins lavapipe so the native path cannot silently turn into the portable placeholder. The FP16 corpus passes on Apple M4 via MoltenVK and on Intel Arc LNL (Mesa ANV) and AMD Radeon 860M (RADV); the lavapipe CI lane exercises the tier on every change.

What each backend reports

Useful when reading DeviceInfo in a trace. OpenVINO varies its UUID and class from the enumerated device name; Vulkan reports the physical device’s actual Vulkan identity. The other provider rows are compile-time constants, which is why a v75 Snapdragon still reports qualcomm-htp-v73.

Backenduuidclassvendor_iddevice_id
coremlapple-coreml-aneNPU0x106b (Apple)0
openvinointel-ov-<device>NPU / GPU / OTHER0x8086 (always)0
xdnaamd.xdna.npuNPU0x1022 (AMD)0x17f0 (always)
hexagonqualcomm-htp-v73NPU0x17cb (Qualcomm)73 (always)
vulkanVulkan deviceUUIDGPU / OTHER for a CPU ICDVulkan physical-device propertyVulkan physical-device property
mockvirtio-accelmockNPU00

class comes from AcceleratorClass, an extensible u16 newtype: OTHER = 0, NPU = 1, GPU = 2, DSP = 3. Unknown values stay representable across implementations, so new classes are additive.

Out of scope

  • Vendor-specific GPU APIs beyond Vulkan. There is no CUDA, TensorRT, cuDNN, ROCm, or Metal backend. A conformant NVIDIA, AMD, Intel, or Apple portability-layer device may still be reachable through the Vulkan backend.
  • Guest-side device passthrough. The project claims no Virtio device ID yet; guests reach hardware through a host backend behind the vAccel adapter.
  • Non-Apple ANE-class fixed-function blocks with no runtime this project speaks to.

Keeping this current

This document tracks device reachability, which changes for different reasons than dtype coverage. Revisit it when:

  • a device selection constant moves — the compute-unit budget, the QNN library path, the HRX device index, the OpenVINO preference order, or Vulkan’s physical-device ranking;
  • a build.rs target or runtime gate changes, which is what most often converts “one change away” into “reachable”;
  • a hardware evidence pin lands in a crate README or performance.md, which is what converts “reachable” into “validated”; or
  • a DeviceIdentity field stops being a constant and starts being probed.

Adding a dtype or operator to an existing backend does not require a change here.

Performance and memory budgets

The portable v1 performance posture is explicit before the API freezes. The checked-in budget artifact is performance-budgets.json, and the baseline metadata is performance-baseline.json.

The default CI budget is deterministic. It checks complexity classes, allocation boundaries, copy-path counters, and representative hot-path byte reads. Wall-clock timings are useful release evidence, but they are not stable enough for ordinary pull-request gating across hosted runners.

Operation areaExpected costAllocation profileCopy boundary
Config and scalar request decodeO(1)nonefixed scalar bytes only
Non-SUBMIT request decodeO(1) plus descriptor validationnonetransfer and artifact tails stay borrowed
SUBMIT decodeO(b log b)one bounded metadata vector after binding-count validationbinding metadata only
Segmented byte-port accessportable worst case O(s + n); indexed split queue O(log s + k + n)none per accessexact caller-requested range
Object lookupO(1)nonenone
Command dispatchrequest-specificbounded object-table reservation before mutationresponse publication, except explicit transfers
Submission admissionO(b log b) plus lookups; canonical binding revalidation is O(b)bounded event dependency and binding metadatano hidden buffer staging
PollingO(1)noneevent-state response only
Resetobject graph walkreleases existing state; no new guest-count allocationnone
TOSA parse + target semanticsO(f + g log s + c)bounded borrowed-name/symbol/control-flow metadata after FlatBuffer verificationno graph, string, or constant-data copy
TOSA lowering analysisO(g log g)compact dense IDs, spans, topological/liveness metadata, and runtime obligationsborrowed graph and constant payloads remain in place
TOSA dynamic specializationO(d) plus exact-key cache lookupcaller-bounded key and LRU entriesdynamic CTC bytes only; ordinary tensor inputs are not scanned

b is a validated binding count, s is segment count (or the largest TOSA symbol table in the TOSA row), k is the number of descriptor segments touched by one logical byte access, n is explicitly requested bytes, f is verified FlatBuffer structure, g is the number of graph objects and edges, and c is compile-time-constant data inspected by the semantic pass.

Copy accounting

The baseline content-copy boundaries are Accelerator::write_buffer and Accelerator::read_buffer. They report explicit transfer bytes separately from provider staging. Submission binds the exact provider allocation. If a provider stages a direct-binding buffer through a hidden bounce allocation during submission, the conformance diagnostics case fails.

The ConformanceHooks::submission_path_diagnostics hook reports cumulative direct, shared/imported, staged-direct, staged-byte, and explicit-transfer counters. Providers that cannot report these counters skip the diagnostics case, but release evidence should include them for any hardware adapter claiming v1 performance conformance.

Budget exceptions

The portable decoder keeps one bounded DecodedBinding vector for SUBMIT duplicate-slot validation. The command engine also owns bounded event dependency and binding metadata while admitting a submission. These allocations are deliberately after guest count validation and contain metadata only, never program-buffer contents.

The decoder’s slot sort is also the canonical handoff to core admission. Core and guest validation recognize strictly increasing slot order in O(b) without allocation. Their public APIs continue to accept arbitrary binding order through an allocation-free fallback, so this optimization does not make ordering semantic.

Split-queue chain construction records bounded logical descriptor spans alongside the flattened regions. Each later byte access binary-searches the first touched span instead of rescanning from descriptor zero. This metadata is allocated only while the driver owns and constructs the bounded chain; queue publication, command decoding, completion, and reset remain allocation-free.

Device admission validates each resolved buffer descriptor in place instead of retaining a parallel descriptor vector. The provider-facing binding vector and event-owned buffer dependency vector remain necessary, but descriptor validation adds no per-submission allocation.

The current v1 budget treats those metadata allocations as acceptable. It does not permit an allocation sized by an unvalidated guest count and does not permit full-range program-buffer copies outside explicit transfer calls.

Local checks

python3 ci/check-performance-budgets.py --check
cargo test --test performance_budgets --all-features
cargo test -p virtio-accel-conformance --all-features

TOSA artifact evidence

virtio-accel-tosa first runs the official FlatBuffers verifier with explicit depth, table-count, apparent-size, and input-byte limits. Its one structural pass stores borrowed &str keys in bounded, fallibly reserved vectors, sorts them once per scope, and uses binary search for reference lookup. It never creates an owned graph and never copies names, tensor bytes, shape data, or appended constant buffers. Returned views read the already-verified buffer in place.

Model::validate_for then walks those borrowed views without constructing an owned IR. It keeps fallibly reserved symbol and control-flow bookkeeping, validates bounded compile-time constants in place, and performs rank-bounded shape arithmetic. No tensor or shape payload is copied.

Model::analyze_for amortizes provider lowering work at program load: every name lookup becomes a dense ID/span access, topological order and liveness are retained, dead/layout/constant-folding opportunities are marked conservatively, and runtime ERROR_IF work is separated from advisory per-element REQUIRE conditions. Dynamic CTC validation is allocation-free after the caller has assembled its sorted borrowed value list. Specialization keys are caller-bounded and collision-safe; the portable LRU uses exact words after its fingerprint and caps retained compiled variants.

The parser’s default graph counts and byte ceilings are finite and callers can lower every one via Limits. Tests compare the input and returned buffer pointers, exercise caller-selected ceilings, parse a flatc-encoded upstream stable graph, and traverse all public views. The tosa_parse fuzz target mutates that upstream seed, cross-checks traversal counts and constant bytes against the validation statistics, materializes every safe attribute view, and runs both a fully enabled and a minimal Level 8K semantic target to exercise rejection paths.

Core ML provider evidence

virtio-accel-coreml builds a sorted slot/access plan at model load. Warm submission reuses the queue’s native-binding array, resolves arbitrary binding order against that plan, and performs one O(b log b) retained-allocation deduplication before admission. The event keeps that one backing vector directly, avoiding the previous second vector allocation/conversion. Submission copies no tensor bytes. Read-only allocations may be shared by overlapping predictions; any output or read-write use remains exclusive.

The crate includes an ignored release-mode measurement for fixed provider overhead:

cargo test --release -p virtio-accel-coreml \
  measures_warm_submission_and_completion_latency -- --ignored --nocapture

On an Apple M4 running macOS 26.5.2, five runs of 200 measured iterations after 20 warmups reported per-run median admission between 5.00 and 5.46 microseconds and median completion between 95.42 and 103.29 microseconds for the embedded Float32[8] model. The pre-pass measurement was 5.25 microseconds admission and 98.46 microseconds completion. The optimization therefore removes submission allocation/scan work without claiming a timing improvement below the noise floor of this micro-model. This is evidence for host and Core ML fixed overhead, not representative ANE throughput, and remains non-gating wall-clock data.

OpenVINO provider evidence

virtio-accel-openvino builds the sorted slot/access/shape plan once at program load. Warm submission reuses the queue’s pointer-slot storage plus one empty high-water vector allocation for backing guards and one for tensor/check metadata. Each spare is cleared before the queue can retain it; therefore reuse removes Rust metadata allocations without retaining buffer pointers, backing guards, tensor handles, or native requests. Concurrent events remain supported: when the one spare is occupied, another event allocates independently, and completion retains at most the larger returned allocation.

The native infer request remains event-owned and is created for every submission. Pooling it would be an unsafe optimization because OpenVINO copies bound tensor objects into the request and its C API has no reset operation that detaches all input and output tensors. Submission still copies no tensor bytes.

The crate includes an ignored release-mode measurement that reports admission separately from submit-to-complete latency:

cargo test --release -p virtio-accel-openvino \
  measures_warm_submission_and_completion_latency -- --ignored --nocapture

Wall-clock results must be recorded on a pinned OpenVINO runtime and identified device before a timing claim is made; the deterministic regression tests instead pin capacity reuse, pointer scrubbing, guard release at terminal observation, and tensor-metadata release after request destruction.

AMD XDNA provider evidence

virtio-accel-xdna compiles each admitted TOSA shape once at program load and stores the resulting precompiled artifact in a content-addressed cache. Warm FP8 submission binds the caller’s FP8 input and BF16 output allocations directly. The conversion streams fixed 1,024-element tiles through one AIE2P worker; no host conversion, submission-time bounce copy, or tensor-sized Rust allocation is part of the warm path.

The crate includes an ignored release-mode scaling measurement matching the OpenVINO structure:

source ~/toolchains/amdxdna-hrx-v2026.08/env.sh
export VIRTIO_ACCEL_AMDXDNA_TOOLCHAIN=~/toolchains/amdxdna-hrx-v2026.08
cargo test --release -p virtio-accel-xdna --test hardware \
  measures_fp8_cast_scaling_on_one_aie_worker -- --ignored --nocapture --test-threads=1

On August 25, 2026, a 1022:17f0 XDNA2 NPU with the v2026.08 HRX/aiecc toolchain produced the following E4M3-to-BF16 results. Each shape was loaded once, warmed up for 20 submissions, and then measured for 200 sequential submissions. Effective I/O counts one FP8 input byte plus two BF16 output bytes per element.

ElementsAdmission median / p95Submit-to-complete median / p95Effective I/ODiagnostics
1,0240.692 / 1.513 µs0.086 / 0.113 ms0.033 GiB/s440 direct bindings; 0 explicit bytes
16,3840.631 / 1.513 µs0.486 / 0.512 ms0.094 GiB/s440 direct bindings; 0 explicit bytes
262,1441.072 / 5.080 µs6.743 / 6.790 ms0.109 GiB/s440 direct bindings; 0 explicit bytes
1,048,5762.585 / 8.526 µs26.656 / 26.816 ms0.110 GiB/s440 direct bindings; 0 explicit bytes

The benchmark validates every output against the exact FP8 oracle after timing. The near-constant large-shape rate documents the current single-worker envelope without claiming it is the final throughput configuration. Multi-worker striping is an optional optimization; deterministic CI continues to gate exact numerics, direct binding, and zero submission-time transfer bytes instead of wall-clock latency.

The exact INT8 MATMUL benchmark uses the same 20 warmups and 200 measured submissions. Shapes on the native 8x8x8 INT8 MMUL grid (M % 16, K % 8 with K >= 16, N % 16) stream through DMA-side micro-tile layout transforms into the fork’s vectorized i8/i32 mm kernel on raw INT8 values, followed by an exact zero-point correction pass (C = R - zb*rowsum(A) - za*colsum(B) + K*za*zb, every term provably inside INT32); no core cycle widens or repacks an operand. Off-grid shapes retain the scalar exact kernel. Run it with:

source ~/toolchains/amdxdna-hrx-v2026.08/env.sh
export VIRTIO_ACCEL_AMDXDNA_TOOLCHAIN=~/toolchains/amdxdna-hrx-v2026.08
cargo test --release -p virtio-accel-xdna --test hardware \
  measures_exact_int8_matmul_latency -- --ignored --nocapture --test-threads=1

On August 27, 2026, the same 1022:17f0 XDNA2 NPU and v2026.08 toolchain measured the 64x64x32 specialization at 0.851 microseconds admission median, 72.089 microseconds submit-to-complete median, and 99.427 microseconds p95, or 3.636 effective GOPS — 4.7x the August 26 widening-kernel baseline (334.065 microseconds, 0.785 GOPS), with every output still matching the shared exact oracle. All 660 bindings across 220 submissions were direct and submission reported zero explicit-transfer bytes. Disassembly of the retired kernel attributed ~99% of its time to scalar zero-point widening and lane-by-lane packing around a single vmac; the correction-term formulation removed that work entirely, and the submit-to-complete median now sits at the measured per-submission overhead floor (the 1,024-element FP8 case measures 86 microseconds), so the remaining latency is submission-path cost, not kernel cost. Issue #151 tracks the next steps (submission overlap, worker striping) without weakening exactness or direct binding.

The pipelined-throughput benchmark keeps four submissions in flight over four rotating buffer sets (measures_pipelined_int8_matmul_throughput, same shape and oracle). On August 27, 2026 it measured 73.3-74.9 microseconds amortized per inference across three 400-completion runs – statistically identical to the sequential submit-to-complete median (69.3-74.6 microseconds across three runs of the latency benchmark on the same worker). A batched-flush variant (all in-flight dispatches submitted under one hrx_stream_flush) measured 70.1 microseconds, also identical. The conclusion this evidence supports: the per-submission floor is per-command driver/firmware round-trip cost inside one hardware context, and neither deeper host-side pipelining nor flush batching moves it. Raising effective throughput therefore requires more work per dispatch (larger admitted envelopes, worker striping – issue #151 steps 5-6) or parallel hardware contexts (issue #121), not further submission-path restructuring. The ring depth of four still pays for itself in semantics: submissions overlap with host-side polling and readback, and completion waits no longer serialize against allocate_buffer.

Vulkan provider evidence status

virtio-accel-vulkan’s warm path is a descriptor update plus vkQueueSubmit on a preallocated bounded ring of (command buffer, fence, descriptor set) slots — no worker thread, no submission-time staging; completion is a nonblocking vkGetFenceStatus poll (ADR 0006). Programs compile once at load_program (checked-in SPIR-V + specialization constants) and are charged against ArtifactRef::resident_bytes.

Manual hardware commands (per the #75 precedent — no self-hosted runner on a public repo):

# Real GPU (any Vulkan 1.3 compute device; pin the ICD explicitly):
VIRTIO_ACCEL_VULKAN_REQUIRE_DEVICE=1 \
  cargo test -p virtio-accel-vulkan --test vulkan -- --nocapture

# Software ICD rehearsal (the CI lane's shape):
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/lvp_icd.x86_64.json \
VIRTIO_ACCEL_VULKAN_REQUIRE_DEVICE=1 \
  cargo test -p virtio-accel-vulkan --test vulkan -- --nocapture

Verified driver stacks for the FP32 operator tier (ADR 0007): Intel Arc 140V (Lunar Lake, Mesa 26.0.8 ANV, Vulkan 1.4.335) (full suite, 2026-09-06, alongside the same host’s llvmpipe LLVM 21.1.8), Mesa lavapipe in the vulkan-lavapipe-test CI lane, and Apple M4 via MoltenVK 1.4.2 (full suite, 2026-09-17; local validation only, not a CI lane). On ANV and llvmpipe alike the crate-authored transcendentals measured 1 ulp (sin, cos, tanh) and 2 ulp (erf) worst case against binary64 over 4096 samples spanning ±8000, ±1e6, f32::MAX, and the non-finite edges — the same numbers, which is what the NoContraction and software-reduction policy exists to guarantee. Copy-path diagnostics across all three: every submission is a direct binding; explicit_transfer_bytes stays zero for Host and Shared domains, and Device staging is confined to write_buffer/read_buffer as the memory-domain contract requires.

The FP16 tier (ADR 0008) shares the FP32 submission path exactly — the same dispatch geometry, arena, and ring; only the element storage is packed two per word — so no separate timing claims are made. Its corpus has executed end-to-end on Apple M4 via MoltenVK, Intel Arc LNL (Mesa ANV), and AMD Radeon 860M (RADV) (2026-09-17), and the lavapipe CI lane exercises the tier on every change.

The FP32 operator tier (ADR 0007) adds the structural optimizations a real graph needs before any timing is worth publishing: a whole graph is one command buffer with barriers only between dependent dispatches; constants and intermediates live in one device-local arena per program with lifetime-packed regions, RESHAPE/IDENTITY views instead of copies, and dead operators elided; MATMUL is a register-tiled shared-memory kernel (a 64 × 64 block per workgroup, ADR 0010) or, for eight rows or fewer, a barrier-free split-k streaming kernel (ADR 0011), both with fused multiply-add and a stated error bound rather than bit-identity to the sequential sum; pipelines are created against a per-instance VkPipelineCache; every 1-D kernel is a grid-stride loop so dispatch counts stay inside maxComputeWorkGroupCount at any tensor size. Known costs, recorded so they are measured rather than assumed: predicate (BOOL) elementwise outputs, strided sub-word moves, and MATMUL’s FP16 result are written with two atomics per element, and SIN/COS evaluate both range reductions and select.

Warm-latency numbers in the XDNA structure (load once, warm 20, measure 200) are not yet published; the throughput benchmark below reports the submission floor it measures alongside each case, and the broadened tier still owes a MoltenVK run (the same commands above).

Vulkan FP8 tier throughput (ADR 0010)

cargo bench -p virtio-accel-vulkan times whole TOSA graphs submit-to-fence through the public Accelerator surface, per enumerated device, after a warm-up of at least 300 ms per case so a frequency-scaling GPU is at clock. On 2026-09-18, Intel Arc (Panther Lake, Mesa 26.0.8 ANV, Vulkan 1.4.335), Device memory domain, median of 30 timed submissions, 16 Mi elements for the elementwise cases; “before” is the kernels as shipped by ADR 0009 under the same harness:

CaseBeforeAfter
IDENTITY FP84.60 ms, 7.3 GB/s0.37 ms, 92 GB/s
IDENTITY FP161.70 ms, 39 GB/s0.65 ms, 103 GB/s
IDENTITY FP321.22 ms, 110 GB/s1.23 ms, 109 GB/s
CAST FP8 → FP161.67 ms, 30 GB/s0.55 ms, 91 GB/s
CAST FP16 → FP83.27 ms, 15 GB/s0.52 ms, 97 GB/s
CAST FP32 → FP83.27 ms, 26 GB/s0.80 ms, 104 GB/s
CAST FP32 → FP161.88 ms, 54 GB/s0.95 ms, 106 GB/s
MATMUL FP8 → FP16, 1024³3.76 ms, 572 GFLOP/s1.43 ms, 1503 GFLOP/s
MATMUL FP16, 1024³2.86 ms, 751 GFLOP/s1.14 ms, 1882 GFLOP/s
MATMUL FP32, 1024³2.74 ms, 785 GFLOP/s1.11 ms, 1928 GFLOP/s
CAST FP8 → FP16 + MATMUL FP16, 1024³3.24 ms1.28 ms
GEMV FP8 → FP16, 1 × 4096 × 40960.98 ms, 17 GB/s of weights0.44 ms, 39 GB/s of weights
GEMV FP16, 1 × 4096 × 40961.06 ms, 32 GB/s of weights0.53 ms, 63 GB/s of weights
GEMV FP32, 1 × 4096 × 40961.06 ms, 63 GB/s of weights0.72 ms, 93 GB/s of weights
CAST FP8 → FP16 + GEMV FP16, 1 × 4096 × 40962.73 ms0.95 ms

GB/s counts bytes read plus written; the GEMV weight figures count the weight matrix alone. The submission floor (a four-element identity) measured 100–170 µs across runs and is included in every number. What the table says about the FP8 tier: data movement now runs at the device’s copy rate, so the quarter-width storage delivers its bandwidth, and GEMV now orders the right way (FP8 0.44 ms, FP16 0.53, FP32 0.72). Net of the floor the FP8 GEMV streams weights at roughly half the rate the FP32 kernel shows the memory system delivers; the remainder is per-step fixed cost, recorded in ADR 0010 as the next objective. The 1024³ cases vary about ±15% run to run on this device even at 30 samples.

Transfers (ADR 0012), same device, 64 MiB, median of 10: write_buffer into the Device domain took 32.6 ms (2.1 GB/s) through the staged path and 2.59 ms (25.9 GB/s) once a single-heap device maps that domain; read_buffer 25.1 → 2.47 ms. Shared measures the same 2.59 / 2.48 ms, and kernel throughput is identical in all three domains on this unified-memory device. The bench takes VIRTIO_ACCEL_VULKAN_BENCH_DOMAIN=host|shared|device and times both transfers per run.

Qualcomm Hexagon evidence status

virtio-accel-hexagon includes an ignored release-mode measurement for fixed submission overhead:

cargo test --release -p virtio-accel-hexagon --test hexagon `
  measures_warm_submission_and_completion_latency -- --ignored --nocapture --test-threads=1

On August 17, 2026, a Snapdragon X126100 Hexagon HTP v73 with NPU driver 30.0.222.0, Windows Balanced power mode, QAIRT 2.49.0.260730, provider build v2.49.0.260730134355, QNN core API 2.38.0, and HTP backend API 5.49.0 produced the following single-run results. Each graph was loaded once, warmed up for 20 submissions, and then measured for 200 sequential submissions.

GraphDtypeAdmission median / p95Submit-to-complete median / p95Diagnostics
identity, 8 elementsFP1627.6 / 61.5 µs2.8098 / 3.0682 ms440 direct bindings; 0 explicit submission bytes
identity, 8 elementsINT823.1 / 58.9 µs2.8465 / 3.0333 ms440 direct bindings; 0 explicit submission bytes

The counts cover two exact caller-owned bindings for all 20 warmups and 200 samples. The input was initialized before counters were sampled; no read or write occurred during measured submission. These are fixed-overhead micro-model results, not throughput claims or representative large-model latency. Ordinary CI gates correctness and copy diagnostics rather than wall-clock values.

Accelerator backend implementer guide

This guide is for providers implementing virtio_accel_core::Accelerator. The trait is the native, transport-independent boundary: backend crates do not decode wire frames, inspect virtqueues, map guest object IDs, or depend on a host operating-system adapter.

The reusable virtio-accel-conformance crate exercises that boundary directly. It needs three provider-owned inputs:

  1. A factory that returns a fresh backend instance for each case.
  2. One executable TargetDescription containing an artifact and its bindings — a single observable binding, or one fixture per slot for programs with disjoint input and output slots.
  3. A ConformanceHooks implementation that advances a pending event to successful completion.

The same flow is available as a runnable crate-level example: cargo run --example backend_conformance.

Both crates are published, so a provider outside this repository depends on them directly. The conformance suite is test-only and belongs in [dev-dependencies], which keeps it out of the shipped dependency graph:

[dependencies]
virtio-accel-core = "0.3"

[dev-dependencies]
virtio-accel-conformance = "0.3"

virtio-accel-mock is also published, if a reference backend is useful to compare against while bringing a provider up. Like the conformance suite it is test-only reference code, with deterministic non-secret artifacts and scripted faults; see ../SECURITY.md for what that does and does not cover.

The completion hook is test control, not a new production requirement. A backend with an external scheduler can signal that scheduler; a deterministic backend can execute the retained invocation directly. The target operation must remain pending until the hook runs, transform the fixture’s initial bytes into its expected bytes, and use the declared slot and access mode. Its observable binding is Write or ReadWrite; a read-only binding cannot carry the required output evidence.

A program whose artifact declares disjoint input and output slots — the shape every lowered TOSA graph produces — is described with TargetDescription::with_bindings, pairing BindingFixture::read_only inputs with at least one observable writable fixture. The executing cases bind every fixture in declared order and, after completion, verify every fixture’s expected bytes, so clobbered inputs fail the suite alongside wrong outputs.

use virtio_accel_conformance::{
    BindingFixture, ConformanceHooks, ProgramFixture, TargetDescription, run,
};

struct Hooks;

impl ConformanceHooks<MyBackend> for Hooks {
    fn complete_event(
        &self,
        backend: &MyBackend,
        event: &MyEvent,
    ) -> Result<(), virtio_accel_core::BackendError> {
        backend.test_control().complete(event)
    }
}

#[test]
fn backend_conformance() {
    let target = TargetDescription::new(
        ProgramFixture::new(FORMAT, TARGET, ARTIFACT, RESIDENT_BYTES).unwrap(),
        BindingFixture::new(
            SLOT,
            ACCESS,
            DOMAIN,
            ALIGNMENT,
            INITIAL_BYTES,
            EXPECTED_BYTES,
        )
        .unwrap(),
    );
    run(MyBackend::new, &target, &Hooks).assert_conformant();
}

The factory must not return a process-global singleton carrying state from another case. Fresh instances are how the suite isolates indeterminate ownership and mirror the device recovery rule: once reuse cannot be proved safe, discard the complete backend instance.

Standard cases

Case IDs are stable diagnostic names. Mandatory cases never skip. Capability cases skip only when the corresponding DeviceInfo.capabilities bit is absent, and the report records that reason. Accounting is optional and is reported separately rather than disguised as a passing mandatory case.

Case IDRequirementContract exercised
metadata.stable-validMandatorydevice_info succeeds, validates, and remains stable for one backend instance
intent.reserved-flagsMandatoryReserved context and execution-queue flags return Unsupported without resources
memory.hostHOST_VISIBLE_MEMORYHonest host allocation metadata and successful release
memory.deviceDEVICE_LOCAL_MEMORYHonest device-local allocation metadata and successful release
memory.sharedSHARED_MEMORYHost-visible, directly bindable shared allocation metadata
buffer.segmented-transfer-boundsMandatorySegmented explicit transfers, exact bytes, and out-of-bounds rejection
buffer.transfer-permissionsMandatoryTransfer-source and transfer-destination usage are enforced independently
program.segmented-artifact-boundsMandatorySegmented artifact input, target loading, advertised artifact limit, and release
submission.binding-validationMandatoryNonempty/unique slots, ranges, program access, valid execution, and acceptance truth
submission.context-isolationMandatoryA buffer from another context is rejected before admission
event.pending-release-terminal-stabilityMandatoryPending release returns the live event, completion is stable, output is visible, and retry releases
timeout.finite-admissionMandatoryFinite timeout results preserve rejected, accepted, and indeterminate ownership shapes
event.cancellation-racesEVENT_CANCELLATIONCancellation-first and completion-first select one stable terminal result
accounting.resource-lifecycleOptional hookFresh and post-case live/indeterminate provider totals are zero

The target’s memory domain must be advertised even though the three general memory-domain cases are capability-conditional. This prevents an implementer from selecting an unusable target fixture and then mistaking the resulting skips for execution evidence.

Accounting hook

Implement ConformanceHooks::resource_counts when provider diagnostics can report retained native resources. Return the sum of resources known live and resources whose release or admission is indeterminate. Do not count an unknown resource as freed merely because its guest-visible ID was invalidated.

The runner samples accounting before and after every case. A nonzero post-case count is attached to that case’s failure, so a semantic error that drops a Rust handle without crossing the provider release boundary is also reported as a leak. The reference backend runs under virtio_accel_mock::fault::FaultAccelerator to exercise this path.

Trait obligations

Discovery and capabilities

device_info is immutable for the backend instance. Every resource and byte limit is nonzero, at least one provider-owned memory domain is available, and reserved capabilities are not advertised. Advertise EVENT_CANCELLATION only when pending-event cancellation is implemented and never returns Unsupported.

Capability bits are promises, not hints. An advertised memory domain must allocate the requested backing or return a request-specific resource error; an unadvertised optional capability cannot be used to skip mandatory context, transfer, program, queue, submission, polling, or release behavior.

TOSA providers may additionally implement virtio_accel_tosa::TosaCapabilityProvider. This is a host-side artifact-planning interface, not a protocol or DeviceInfo capability. Return one CapabilityDescriptor per exact target tier and keep role-specific dtypes, operator constraints, shape limits, and runtime-condition policy conservative. An unavailable runtime returns an empty slice. A positive descriptor query permits a load attempt only; load_program remains responsible for concrete shape relationships, resources, native compilation, and device-state failures.

Handles and ownership

Handles own provider state and borrow no call argument. Context children and event dependencies are caller-managed, but accepted events must retain all provider invocation state until terminal release. Do not use Drop timing as protocol state.

Creation errors retain no resource. Submission rejection proves that no work was accepted and no event exists. If acceptance cannot be proved false, return SubmitFailure::Indeterminate with the event. A rejected release returns the same live handle for retry; an indeterminate release consumes the handle and requires complete backend recovery.

The suite retries one rejected release. It never retries an indeterminate operation on the same backend instance.

Buffers and copies

BufferInfo describes the backing that was actually allocated. Its descriptor exactly matches the request, allocation size and alignment meet the request, and properties are truthful. Every program-visible buffer reports DIRECT_BINDING; inability to bind the exact allocation is an allocation or submission error, not permission for a hidden bounce buffer.

write_buffer and read_buffer are the only baseline bulk-copy boundaries. They accept segmented ports and may use bounded staging for device-local transfers. Allocation, submission, polling, and release do not copy complete bound ranges. The semantic suite validates direct-binding metadata, observable output, and optional copy-path diagnostics; the v1 budgets in performance-budgets.json define the regression thresholds.

Programs, bindings, and admission

Artifacts are opaque to transports but not to the provider. Read segmented bytes without requiring one artifact-sized coalescing allocation, validate the format/target envelope, and retain no borrow of the source. ArtifactRef::resident_bytes is the caller-authorized upper bound for all storage retained by the returned program handle, including compiled code and provider metadata attributable to that program. Reject the load when that bound cannot be honored; do not treat it as an estimate.

Submission validates nonempty bounded bindings, unique slots, nonempty in-range regions, declared buffer usage (access must be compatible with the buffer’s usage bits), program-specific slot/access compatibility, and one context across queue, program, and buffers. Reject usage mismatches before provider admission. Validation and admission are bounded. The borrowed binding slice is not an owned per-binding mirror and must not survive the call as Rust references.

Events, cancellation, and time

Polling is bounded and nonblocking. Once a terminal state is observed, every later successful poll returns exactly that state. A pending event cannot be destroyed. Cancellation and completion race to one terminal result: cancellation success means later polls are Cancelled; if completion won, cancellation returns Busy and preserves the completed or failed result.

Timeouts are relative to backend admission. Never compare a guest duration with an absolute host timestamp. A timeout before acceptance is rejected as DeadlineExpired; uncertainty after the acceptance boundary is indeterminate and carries an event. The standard finite-timeout case accepts any of the three truthful ownership shapes because provider scheduling policy is implementation specific.

Reset and device loss

Accelerator deliberately has no reset method. The command engine owns object IDs and teardown; the provider owns native resource truth. Successful child-before-parent release permits backend reuse. Device loss, indeterminate release, unresolved pending work, or accounting contradiction requires discarding the entire instance and constructing a fresh one through the factory.

Use virtio_accel_mock::fault::FaultAccelerator or an equivalent provider-local injector to test rejected and indeterminate outcomes at each native API boundary. The standard suite remains usable without a vendor fault API, while the command-engine and full-stack fault tests prove the portable recovery policy.

For a complete portable lifecycle without the conformance harness, run cargo run --example reference_execution. It demonstrates the context, buffer, program, queue, submit, poll, transfer, and teardown sequence against the mock backend. On macOS, run cargo run -p virtio-accel-coreml --example tosa_coreml for the production artifact path from a device-neutral TOSA graph through Core ML execution and the same backend lifecycle.

cargo run -p virtio-accel-hexagon --example tosa_hexagon exercises the QNN HTP lifecycle and direct bindings on a configured Windows ARM64 QAIRT host. Without that build environment, it retains the explicit unavailable-runtime surface.

Common traps

  • Advertising a capability because a slow fallback exists, while the fallback violates direct binding or synchronization semantics.
  • Returning Rejected after a native queue accepted work or a timeout raced with admission.
  • Dropping an event on failed destruction instead of returning it in ReleaseFailure::Rejected.
  • Letting a second poll regress from complete, failed, or cancelled to pending.
  • Accepting duplicate slots or matching bindings by array position rather than slot number.
  • Retaining ArtifactRef, BindingRef, or byte-port references after their call returns.
  • Adding a lock, atomic, allocation, or owned binding mirror to every submission when native handle ownership already provides the required exclusivity.
  • Treating a Rust handle drop, ID invalidation, or reset request as proof that provider state was released.

Qualcomm Hexagon TOSA operator matrix

This matrix records the TOSA 1.0 floating-point surface validated on Snapdragon X126100, Hexagon HTP v73, driver 30.0.222.0, and QAIRT 2.49.0.260730. “HTP” means the checked-in numerical fixture passed through the QNN HTP backend; “portable” means lowering is also covered without the SDK. The floating target admits FP16, BOOL conditions/results, and required INT32 indexing results. The separate integer target remains limited to INT8 IDENTITY and zero-point-aware INT8 MATMUL with INT32 output.

TOSA operatorQNN representationValidated restrictionEvidence
ARGMAXArgMaxFP16 input, INT32 output, one static axis, keep_dims=falseportable + HTP
MAX_POOL2DGather + ElementWiseMaximumFP16 NHWC, positive static kernel/stride, zero paddingportable + HTP
MATMULMatMulFP16, or exact INT8/INT32 tier; no transposeportable + HTP
CLAMPReluMinMaxFP16 bounds, NaN propagateportable + HTP
ERFnoneblocked: QAIRT 2.49 public QnnOpDef.h defines no ERF operationparity exception
SIGMOIDSigmoidFP16portable + HTP
TANHTanhFP16portable + HTP
ADDElementWiseAddFP16 broadcastingportable + HTP
LOGICAL_ANDElementWiseAndBOOLportable + HTP
LOGICAL_ORElementWiseOrBOOLportable + HTP
LOGICAL_XORElementWiseXorBOOLportable + HTP
MAXIMUMElementWiseMaximumFP16 broadcasting, NaN propagateportable + HTP
MINIMUMElementWiseMinimumFP16 broadcasting, NaN propagateportable + HTP
MULElementWiseMultiplyFP16 broadcasting, validated zero shiftportable + HTP
POWElementWisePowerFP16 broadcasting; TOSA domain requirement remains caller-visibleportable + HTP
SUBElementWiseSubtractFP16 broadcastingportable + HTP
ABSElementWiseAbsFP16portable + HTP
CEILElementWiseCeilFP16portable + HTP
COSElementWiseCosFP16, 8-ULP oracle boundportable + HTP
EXPElementWiseExpFP16portable + HTP
FLOORElementWiseFloorFP16portable + HTP
LOGElementWiseLogpositive FP16 fixture domainportable + HTP
LOGICAL_NOTElementWiseNotBOOLportable + HTP
NEGATEElementWiseNegFP16, zero input/output zero pointsportable + HTP
RECIPROCALElementWiseUnary(RECIPROCAL)nonzero FP16 fixture domainportable + HTP
RSQRTElementWiseRsqrtpositive FP16 fixture domainportable + HTP
SINElementWiseSinFP16, 8-ULP oracle boundportable + HTP
SELECTElementWiseSelectBOOL condition, FP16 valuesportable + HTP
EQUALElementWiseEqualFP16 input, BOOL outputportable + HTP
GREATERElementWiseGreaterFP16 input, BOOL outputportable + HTP
GREATER_EQUALElementWiseGreaterEqualFP16 input, BOOL outputportable + HTP
REDUCE_MAXReduceMaxFP16, one static axis, keep_dims=trueportable + HTP
REDUCE_MINReduceMinFP16, one static axis, keep_dims=trueportable + HTP
REDUCE_PRODUCTGather + ElementWiseMultiplyFP16, one static axis; public ReduceProd rejected by HTP v73portable + HTP
REDUCE_SUMReduceSumFP16, one static axis, keep_dims=trueportable + HTP
CONCATConcatFP16, static valid axisportable + HTP
RESHAPEReshapeFP16, compile-time CONST_SHAPEportable + HTP
REVERSEdescending static indices + GatherFP16, one static valid axisportable + HTP
TRANSPOSETransposeFP16, static complete permutationportable + HTP
CONSTowned QNN static tensorinternal FP16 constants with exact byte lengthportable + HTP
CONST_SHAPEconsumed during loweringstatic valid reshape shapeportable + HTP
IDENTITYReshapeFP16 or exact INT8 tierportable + HTP

Every attribute, dtype, rank, axis, permutation, shape, and constant payload first passes the shared bounded TOSA verifier and semantic analyzer. The Hexagon planner then owns constants, uses checked shape/byte arithmetic, and rejects unsupported graphs before entering QNN. The native bridge repeats descriptor arity, pointer, element-size, constant-size, reference, and generated-tensor resource checks before calling the provider.

Reproduce the portable and hardware evidence with the commands in the Hexagon README.

API reference

The public Rust API is documented with rustdoc and built for the whole workspace with all features enabled. The generated documentation is served under /api/.

Crates

CrateTierRole
virtio-accelcore + allocFacade re-exporting the portable layers
virtio-accel-protocorePointer-free, little-endian protocol 1.0 wire structures
virtio-accel-transportcoreDescriptor-chain, queue, reset, and notification ports
virtio-accel-corecoreBackend lifecycle, memory, program, queue, and event contracts
virtio-accel-tosacore + allocBounded zero-copy TOSA 1.0 validation and lowering analysis
virtio-accel-tosa-buildcore + allocBorrowed and owned static TOSA 1.0 authoring
virtio-accel-vaccelcoreAdapter seam for native provider contracts
virtio-accel-coremlmacOS stdTOSA-to-Core ML lowering and ANE-capable prediction
virtio-accel-openvinoLinux stdTOSA-to-OpenVINO IR lowering and NPU/GPU/CPU inference
virtio-accel-hexagonWindows ARM64 stdStrict FP16/INT8 TOSA-to-QNN lowering
virtio-accel-xdnastdAMD XDNA2 NPU backend over HRX with BF16/FP8/INT8 TOSA tiers
virtio-accel-vulkanstdVendor-neutral Vulkan 1.3 compute backend with the shared FP32 tier plus the native FP16 and FP8 tiers
virtio-accel-split-queuecore + allocBounded in-memory split-ring reference model
virtio-accel-guestcore + allocTyped reference client
virtio-accel-devicecore + allocDevice-owned state with generational IDs
virtio-accel-mockstdIn-memory backend with deterministic test-only artifacts
virtio-accel-conformancestdTransport-free semantic suite and numerical corpus
virtio-accel-cleanroomcoreIndependent conformance codec

The rustdoc index at /api/ lists every crate and its feature-gated items. See the public API policy for how the documentation layers are split and what is guaranteed stable.

Public API documentation

The public Rust API is split into three documentation layers:

  1. Human-facing contracts in virtio-accel-core, virtio-accel-guest, virtio-accel-device, virtio-accel-transport, virtio-accel-tosa, virtio-accel-tosa-build, and virtio-accel-conformance.
  2. Pointer-free wire mirrors in virtio-accel-proto and the checked C projection in include/virtio_accel.h.
  3. Independent conformance artifacts and the clean-room codec under conformance/.

The first layer documents ownership, blocking behavior, allocation, copy boundaries, and recovery semantics in rustdoc. The normative protocol text remains in specification.md, wire-abi.md, and virtqueue.md; rustdoc links back to those documents rather than restating every wire rule in multiple places.

Raw wire and clean-room exceptions

virtio-accel-proto deliberately exposes repr(C) structures whose public fields match the normative wire names exactly. Field-level rustdoc would mostly duplicate wire-abi.md and increase drift risk. The authoritative field semantics are the normative document, layout.json, and vectors.json.

include/virtio_accel.h exposes the same constants and packed layouts to C11 and C++11 consumers. It deliberately contains no functions, provider handles, allocator hooks, or callbacks: it is a wire header, not a stable backend plugin ABI. ci/check-c-header.py compiles manifest-derived assertions so adding or changing a recorded namespace or layout cannot silently drift the header.

virtio-accel-cleanroom is also intentionally raw. It is a dependency-free independent decoder for golden-vector validation, not the ergonomic production API. Its public names mirror protocol terms so another implementer can compare behavior without importing the primary wire crate.

These are documented exceptions, not permission for platform adapters or future public crates to skip API docs. New ergonomic APIs should document ownership, lifetime, error, blocking, allocation, copy, and portability behavior at the item where consumers call it.

virtio-accel-tosa is an ergonomic exception around private raw bindings: its public API exposes only verified borrowed views, typed stable-op attributes, and raw forward-compatible enum numbers. Model::validate_for applies the complete stable TOSA 1.0 target semantic pass. Model::analyze_for additionally returns the compact dense-ID execution/liveness/constant/runtime plan intended for provider lowering, and retains verified Operator/Tensor/Shape views rather than copying an owned graph. Dynamic providers can validate host-readable CTC values and use the bounded exact-key specialization cache. Provider-specific capability utilities extend the layer through ModelValidator. The optional TosaCapabilityProvider trait exposes conservative, target-specific host descriptors for schedulers without adding TOSA concepts to virtio-accel-core; successful discovery never replaces authoritative program admission. Generated FlatBuffers tables and unchecked roots remain private.

virtio-accel-tosa-build is the authoring companion, not a second ingestion path. It exposes borrowed definitions for statically declared graphs and owned definitions for incremental compiler frontends, keeps raw table slots and union tags private, and invokes the same virtio-accel-tosa structural and target validation before either build surface returns owned output bytes.

Runnable entry points

Six examples are part of the default CI workflow:

  • cargo run --example backend_conformance
  • cargo run --example reference_execution
  • cargo run -p virtio-accel-coreml --example tosa_coreml
  • cargo run -p virtio-accel-openvino --example tosa_openvino
  • cargo run -p virtio-accel-hexagon --example tosa_hexagon
  • cargo run -p virtio-accel-hexagon --example mock_classifier

backend_conformance shows how a backend author wires a provider to the reusable conformance suite. reference_execution runs a complete context/buffer/program/queue/submit/poll/read/release lifecycle against the portable mock backend. tosa_coreml proves the production path from a device-neutral TOSA artifact through backend-local Core ML lowering and direct-bound asynchronous execution; non-macOS hosts compile the placeholder, and macOS hosts without an ANE skip execution. tosa_openvino proves the same production path through backend-local OpenVINO IR lowering on the preferred available Intel inference device; hosts without an OpenVINO runtime compile the placeholder, and hosts without an inference device skip execution. tosa_hexagon executes the shared FP16 identity graph through QNN HTP when the complete QAIRT SDK is selected on Windows ARM64. The backend advertises 41 of the 42 Core ML/OpenVINO TOSA operators; the exact restrictions are in the Hexagon operator matrix. SDK-free hosts retain the compile-only RuntimeUnavailable surface without falling back to CPU or GPU. mock_classifier uses the same native lifecycle to compute two sets of class logits from three FP16 features and a direct-bound 3x2 weight matrix.

Baseline, reserved, and post-v1 work

Baseline v1 is the mandatory command, queue, object, reset, error, and conformance behavior in the normative documents. Reserved feature bits, opcodes, flags, and fields are not optional features; they are invalid until a later policy assigns semantics. Platform integrations such as KVM, vhost-user, VFIO, Windows, macOS, or vendor SDK adapters do not change protocol 1.0 and must not leak into portable default dependencies. virtio-accel-coreml and virtio-accel-openvino are the first concrete host backends. virtio-accel-hexagon is a separately packaged, pinned experimental QNN HTP adapter. Each depends inward on virtio-accel-core and virtio-accel-tosa, while the facade and portable runtime crates depend on none of them.

The compatibility and release classification rules are in release-policy.md. The protocol 1.0 frozen surface is summarized in releases/v1.0.md and audited in ../conformance/v1.0/freeze-audit.md.

Release and evolution policy

This policy applies after the protocol 1.0 freeze audit. It keeps Cargo package evolution, wire compatibility, feature selection, unsafe code, dependency selection, and target support aligned.

Version dimensions

There are two separate version axes:

  • Cargo crate versions describe the Rust API and package graph.
  • Protocol versions describe driver/device wire compatibility and conformance artifacts.

The two axes advance independently. A Cargo patch, minor, or major change does not by itself change the wire protocol, and a wire protocol change must follow the protocol classification below even when the Rust crate version is still pre-1.0. A public protocol 1.0 release needs a matching release note and a frozen conformance/v1.0 directory; it does not require a 1.0.0 Cargo version.

Cargo version posture

The workspace publishes at 0.1.x while carrying the frozen protocol 1.0 baseline. This is deliberate rather than an unreconciled gap:

  • The Cargo version tracks the Rust API and package graph, which is young and expected to change as backend, guest, device, and transport adapter authors build against it.
  • The protocol version tracks wire compatibility, which is frozen by the v1.0 freeze audit and governed by the classification table below.

A pre-1.0 Cargo version is therefore the accurate signal on both axes, and is not a statement about protocol stability. Consumers who need the stable artifact should depend on protocol 1.0 and its conformance directory, not on a Cargo version number.

Moving the workspace to 1.0.0 is a separate, later decision. It requires the public Rust API to have real downstream users and an explicit semver promise recorded in a release note. Until then, breaking Rust API changes ship as 0.x minor bumps under the classification table below.

Change classification examples

ChangeCargo classificationProtocol classificationRequired evidence
Fix rustdoc, examples, comments, non-normative rationale, or tests without changing accepted/emitted bytesPatchErratum or no protocol changeCI plus updated docs when relevant
Add a new helper type or trait method with a default implementation that preserves existing behaviorMinor while pre-1.0/public policy permits it; otherwise semver-compatible minorNo protocol changeAPI review and downstream compile coverage
Remove, rename, or change the meaning of a public Rust item used by backend, guest, device, or transport authorsCargo majorNo protocol change unless wire behavior also changesMigration note and affected-crate review
Raise MSRV or remove a supported portable targetCargo minor only if release notes document it and no public semver promise forbids it; otherwise Cargo majorNo protocol changeTarget/MSRV rationale and CI matrix update
Add a platform adapter crate that depends on portable crates but is not a default dependencyAdditive Cargo minorNo protocol changeDependency-policy and portability review
Add a default feature that selects Linux, macOS, Windows, VMM, kernel, vendor SDK, filesystem, socket, thread, or runtime behavior in a portable crateForbiddenForbidden unless it is a negotiated protocol feature and isolated from portable defaultsMust be redesigned
Assign a reserved feature bit, opcode, status, flag, field, or capability with negotiated behavior and unchanged 1.0 framesCargo minor or major depending on Rust API impactProtocol minor with a new conformance directoryNormative docs, feature negotiation tests, vectors, scenarios, and clean-room coverage
Append fields to an existing response a 1.0 driver can receive without a negotiated featureForbidden in protocol 1.xProtocol major if requiredNew major-version directory
Change an assigned opcode value, structure size, field meaning, status success/failure interpretation, ownership rule, reset rule, or existing payload lengthCargo major if Rust API also changesProtocol majorNew normative documents and conformance directory

Wire evolution

Protocol 1.0 freezes the assigned values, exact payload lengths, ownership rules, and golden bytes in conformance/v1.0. Unknown fields are not a baseline extension mechanism: protocol 1.0 receivers validate exact payload lengths and reject trailing bytes unless a negotiated feature explicitly selects a different layout. Unknown opcodes remain unsupported without side effects, unknown request or object flags are rejected before backend invocation, unknown response statuses are opaque failures, and unknown event states require recovery rather than being guessed terminal.

New behavior should prefer one of these forms, in order:

  1. A new opcode with exact request and response payloads.
  2. A previously reserved feature bit that gates all changed behavior.
  3. A previously reserved value whose semantics are specified in full before it is advertised.
  4. A new protocol major version when compatibility cannot be preserved.

Reserved values are invalid until assigned by a later policy. A constant that records a reserved number is not permission for a device to advertise it or a driver to accept it.

Cargo feature policy

Cargo features are additive. Disabling a default feature may remove convenience code but must not select a different protocol interpretation. Enabling a feature must not make a portable crate depend on an operating system, VMM, kernel, guest-memory library, vendor SDK, filesystem, socket, thread, global runtime, or platform synchronization primitive.

Platform integrations must live in adapter crates that depend inward on the portable layers. They must not become default dependencies of virtio-accel-core, virtio-accel-proto, virtio-accel-transport, virtio-accel-device, virtio-accel-guest, virtio-accel-split-queue, or the facade crate.

MSRV and supported targets

The minimum supported Rust version is the workspace rust-version. A change to MSRV requires a release-note entry, a CI matrix update, and an explanation of why the old compiler cannot preserve the current API or implementation invariants.

The supported portable target set is the one documented in docs/portability.md and enforced by CI. Removing a target or moving a crate to a less-portable runtime tier requires a release-note entry and an explicit portability review. Adding a platform adapter cannot reduce the portability tier of an existing crate.

Unsafe-code policy

Project-authored code in portable crates, reference crates, and fuzz harness support code forbids or denies unsafe code at the crate root. This is an intentional v1 invariant, not incidental linting. There are six reviewed, confined exceptions. The host-native virtio-accel-coreml adapter’s Rust FFI and aligned-allocation code is documented in crates/virtio-accel-coreml/SAFETY.md; non-macOS builds still forbid unsafe code. The host-native virtio-accel-openvino adapter’s OpenVINO C API FFI and aligned-allocation code is documented in crates/virtio-accel-openvino/SAFETY.md; builds without a detected OpenVINO runtime still forbid unsafe code. virtio-accel-tosa denies unsafe code globally but locally permits its private, checked-in official FlatBuffers bindings after bounded verification; the boundary and regeneration procedure are documented in crates/virtio-accel-tosa/SAFETY.md. The host-native virtio-accel-hexagon adapter’s QNN C API boundary is documented in crates/virtio-accel-hexagon/SAFETY.md; builds without a detected QNN runtime still forbid unsafe code. The host-native virtio-accel-xdna adapter’s HRX C ABI, persistent mappings, and dispatch ownership are documented in crates/virtio-accel-xdna/SAFETY.md; builds without a detected HRX runtime still forbid unsafe code. The host-native virtio-accel-vulkan adapter’s ash entry-point inventory, persistent mappings, submission ring, and device-loss handling are documented in crates/virtio-accel-vulkan/SAFETY.md; builds forced to the placeholder (or on a host outside the Vulkan loader target set) still forbid unsafe code.

A future unsafe exception requires all of the following in one reviewed change:

  • the crate-level forbid(unsafe_code) removal or replacement is explicit;
  • every unsafe block has a local safety comment naming the invariant it relies on;
  • the release review records why a safe abstraction, zerocopy validation, ownership token, or adapter boundary could not preserve the invariant;
  • tests or conformance evidence exercise the unsafe boundary; and
  • the public API does not transfer unsafe obligations to downstream users unless those obligations are documented on the item that requires them.

Dependency and license policy

Workspace dependencies must be centralized in [workspace.dependencies] unless there is a narrow crate-local reason to diverge. Normal dependencies for portable crates should use minimal features and default-features = false when the dependency supports it.

Dependency review must check:

  • cargo-deny advisories, yanked crates, duplicate versions, wildcard requirements, unknown sources, and licenses;
  • whether a build dependency leaks std or alloc into a target graph;
  • whether a proc macro or helper crate is build-host-only or runtime-visible;
  • whether a dependency introduces platform defaults; and
  • whether its license remains inside the workspace allowlist.

The workspace license is MIT OR Apache-2.0. Every published manifest inherits license, rust-version, repository, homepage, keywords, and categories from [workspace.package], declares its own description and readme, and carries its own byte-identical copies of LICENSE-MIT and LICENSE-APACHE. Cargo only packages files inside a package directory, so the root license files do not reach the sub-crate tarballs; the copies exist for that reason and are copies rather than symlinks because CI runs windows-latest.

ci/check-release-policy.py enforces all of this against an explicit eighteen-crate allowlist. A new package fails that check until it is added to the allowlist, which forces a decision about whether it is public rather than letting it default either way. The check also validates the crates.io keyword and category limits, which neither cargo package nor cargo publish --dry-run catches before an upload is attempted.

The fuzz/ harness is a separate workspace at version 0.0.0 and stays publish = false.

Publication, yank, and rollback

Eighteen packages are published to crates.io. Publication is ordered: a crate cannot be published before every crate it depends on, and that includes development dependencies, because a published crate’s versioned dev-dependencies must resolve from the registry for cargo test to run on the packaged source.

#CrateNormal dependenciesDevelopment dependencies
1virtio-accel-transport——
2virtio-accel-cleanroom——
3virtio-accel-proto—cleanroom
4virtio-accel-coretransport—
5virtio-accel-tosacore, FlatBuffers—
6virtio-accel-tosa-buildtosa, FlatBuffers—
7virtio-accel-split-queuetransport—
8virtio-accel-guestproto, transportsplit-queue
9virtio-accel-mockcore—
10virtio-accel-devicecore, proto, transportmock
11virtio-accel-conformancecoremock, tosa
12virtio-accel-vaccelcoreconformance
13virtio-accel-xdnacore, tosaconformance, tosa-build
14virtio-accel-coremlcore, tosaconformance
15virtio-accel-openvinocore, tosaconformance
16virtio-accel-hexagoncore, tosaconformance
17virtio-accel-vulkancore, tosaconformance
18virtio-accelthe six runtime cratesconformance, mock, cleanroom

This order is executable, not just documentary: ci/publish-dry-run.py walks it against an isolated local registry, adding each crate only after it has been built, tested, and documented from its own extracted tarball. A crate can therefore only ever resolve its predecessors, so a wrong order fails with an unresolvable dependency instead of passing quietly. The same script is a required CI job.

All eighteen packages share the workspace version and are published together. A GitHub release tag must be exactly v<workspace-version> and must point at the commit being released. Publishing a GitHub release triggers .github/workflows/publish.yml, which checks out that tag, runs the release policy checks and publication-driver tests, repeats the full ordered local-registry dry run on the tagged source, and only then gives ci/publish.py the crates.io token from the CRATES_IO_KEY repository secret. The job has read-only GitHub permissions and serializes releases so two publication sequences cannot overlap. The tagged commit must be reachable from the repository’s default branch; a release cannot use an unmerged tag to run token-bearing repository code.

The production publisher uses the same order from ci/publication.py. Before each upload, it builds the actual .crate archive and queries crates.io. An existing version is skipped only when the registry checksum exactly matches the local archive. This makes a rerun safe after a partial publish or an upload whose result was ambiguous, while refusing to bless a tag whose immutable crates.io version contains different bytes. The upload itself names the crates-io registry explicitly and uses Cargo’s --no-verify mode: all compilation and tests have already completed in the prior step, so package build scripts never execute in the step that holds the crates.io token.

cargo package’s own verify step is not sufficient and must not be treated as sufficient: it builds only the library target. That is how four cross-package include_str! sites reached outside their package directories unnoticed, leaving assertions that could never have compiled from a published tarball. Any check on packaged output must run the tests inside the packaged source.

When a mid-order publish fails

crates.io publication is not transactional across crates. If crate N fails after 1..N-1 succeeded, those earlier versions are live and permanent.

  1. Stop. Do not publish the remaining crates by hand.
  2. Inspect crates.io before rerunning. The automated job may be rerun unchanged only when every version that appeared has the exact checksum of the tagged source; it will enforce this check.
  3. Diagnose against the local registry, not against crates.io. Reproduce with ci/publish-dry-run.py.
  4. If any source or packaging metadata must change, fix forward: bump the lockstep workspace patch version, create the matching tag, and publish a new GitHub release. Versions already accepted by crates.io are immutable and cannot be replaced.
  5. Yank only if a published version is actively harmful — see below. A version that is merely stranded, because its dependents were never published, is not harmful; it is unreachable.

Yank versus patch

A crates.io version is immutable. It cannot be edited, replaced, or deleted, and its contents remain downloadable even after a yank. Publishing is therefore a one-way action, and a mistaken publish is corrected by publishing again, never by trying to undo.

Yanking only stops new resolution: existing Cargo.lock files continue to resolve a yanked version, so a yank is not a security control and never a substitute for an advisory.

Publish a patch, and do not yank, when:

  • the defect is a bug, a missing file, or wrong metadata that a newer version supersedes;
  • the version is stranded but harmless; or
  • downstream users are better served by upgrading than by a broken resolution.

Yank, in addition to publishing a patch, when:

  • the version is a security risk to anyone who resolves it — coordinate with SECURITY.md and publish an advisory, since the yank alone protects nobody;
  • it claims a protocol conformance it does not have, so a driver or device could interoperate incorrectly on the wire; or
  • it was published in error and has no valid use, such as a wrong version number or a crate published out of order with an unsatisfiable dependency.

Never un-yank to “restore” a version that was yanked for a wire-compatibility or security reason. Publish a new version instead.

Rollback

There is no rollback. The recovery path for every publication mistake is a new version, in the same documented order, with a release-note entry recording what happened and why. If a protocol-affecting defect ships, the classification table above governs whether the fix is an erratum, a protocol minor extension, or a new protocol major version with its own conformance directory — a security fix is not exempt from that classification.

Release review checklist

Every release or compatibility-affecting PR should answer:

  • Does this change alter accepted or emitted protocol bytes, exact payload lengths, ownership, reset, error, timeout, or feature-negotiation behavior?
  • If yes, is it a protocol erratum, minor extension, or major-version change under this policy?
  • Are layout.json, vectors.json, scenarios.json, requirements.json, and performance budgets still authoritative inputs rather than regenerated by accident?
  • Did any public Rust API change affect backend implementers, guest/device authors, or transport adapters?
  • Did any default dependency, Cargo feature, or target move platform behavior into a portable crate?
  • Did any crate add or permit unsafe code, and is the audit trail complete?
  • Did dependency, license, advisory, and MSRV checks pass?
  • Are deferred optional features still unadvertised and documented as out of scope?