Skip to main content

Crate virtio_accel_xdna

Crate virtio_accel_xdna 

Source
Expand description

AMD XDNA (Ryzen AI NPU) host backend for virtio-accel.

This crate executes device-neutral TOSA 1.0 programs on an AMD XDNA2 NPU through the HRX runtime (libhrx), compiling admitted graphs with the pinned aiecc toolchain as a bounded subprocess. The design is recorded across the AMD XDNA wayfinder map (issue #78) and its decision tickets (#82 numerical tier, #83 crate layout, #84 compiler helper, #85 execution model), whose resolution records live on their respective ticket branches.

The native modules (ffi, native) compile only when the build script finds a complete amdxdna-native HRX prefix (VIRTIO_ACCEL_HRX_DIR/HRX_DIR, or the VIRTIO_ACCEL_HRX_LIB_DIR escape hatch) and sets the va_xdna cfg; VIRTIO_ACCEL_XDNA forces the probe on (1, failing loudly) or off (0). Hosts without HRX build the portable admission surface (lower), the artifact codec, the offline compiler driver (compile_artifact, unix), and a placeholder — and compile no unsafe at all.

Scope today: the full Accelerator lifecycle — the HRX device/stream owner, hrx_buffer primitives (persistent mapping, range flush/invalidate, release), and the serialized dispatch worker bridging hrx_stream_dispatch/synchronize to a latched nonblocking poll_event (execution-model spec, issue #85). load_program accepts the crate-local precompiled format (artifact) directly, and a TOSA artifact by admitting it and compiling it with the bounded aiecc helper subprocess (issue #84). The compilable TOSA subsets today are BF16 IDENTITY, BF16 → FP32 MATMUL, BF16 MAX_POOL2D, explicit FP8 → BF16 storage conversion, the fused FP8 → FP32 MATMUL that keeps that promotion on the compute core, exact INT8 IDENTITY, zero-point-aware INT8 → INT32 MATMUL, and exact INT32 → INT8 RESCALE. Admission (lower) unit-tests on every host.

Re-exports§

pub use artifact::PrecompiledArtifact;
pub use artifact::XDNA_PRECOMPILED_FORMAT;

Modules§

artifact
The crate-local precompiled XDNA artifact format.
bfp_experiment
AMD bfp16ebs8 vendor experiment (block-8): the XBFP artifact container.

Structs§

XdnaAccelerator
Placeholder that keeps workspace consumers portable when no HRX runtime was detected.

Enums§

AdmitError
Why a TOSA artifact is not admissible to the compilable subset.
CompilerSpec
A validated operator specialization ready for the compiler helper. Each variant names its input and output dtypes; the closed shape is integers only, so no guest bytes cross the boundary.
Fp8Format
The two TOSA/OCP FP8 storage encodings accepted by the conversion template.
InitError
Failure to initialize an XDNA backend instance.

Constants§

REQUIRED_RESIDENT_BYTES
The HRX runtime publishes no finite upper bound for a loaded program’s device residency.
XDNA_ERROR_DOMAIN
BackendError::External domain tag for this backend’s failures (“XDNA” in ASCII), covering both HRX runtime errors and compiler-helper failures.
XDNA_TOSA_CAPABILITY
Conservative capability boundary for the implemented XDNA BF16 execution tier.
XDNA_TOSA_FP8_CAPABILITY
Conservative capability boundary for explicit FP8 storage conversion.
XDNA_TOSA_FP8_TARGET
The FP8 storage tier: graph-visible E4M3/E5M2 inputs explicitly cast to BF16 on the NPU.
XDNA_TOSA_INTEGER_CAPABILITY
Conservative exact integer-profile capability boundary for XDNA lowering.
XDNA_TOSA_INTEGER_TARGET
The integer tier: TOSA 1.0, integer profile, level 8K, no extensions.
XDNA_TOSA_TARGET
The BF16 floating-point tier: TOSA 1.0, floating-point profile, level 8K, BF16 extension.

Functions§

admit
Admit a TOSA artifact for target, returning the specialization the helper compiles.
compile_artifact
Admit a TOSA artifact and compile it to a precompiled-artifact container (artifact) with the pinned toolchain, without touching the device.