Skip to main content

Crate virtio_accel_vulkan

Crate virtio_accel_vulkan 

Source
Expand description

Vendor-neutral Vulkan compute host backend for virtio-accel.

The native path binds Vulkan 1.3 through the pinned [ash] crate, loading the platform’s Vulkan loader at run time (ADR 0002). It executes device-neutral TOSA 1.0 programs admitted by lower — the 42 FP32-tier operators shared with the Core ML and OpenVINO backends over FP32 and FP16 tensors, with BOOL/INT32 auxiliaries — on crate-authored SPIR-V compute kernels specialized at load_program (ADR 0003, ADR 0007). The FP16 tier needs no device feature: its binary16 conversions are crate-owned integer and binary32 code (ADR 0008), so every device that hosts the FP32 tier hosts the FP16 tier with identical numerics. A whole graph is one submission: constants and intermediates live in a per-program arena and dependent dispatches are separated by compute barriers. Buffers are dedicated VkDeviceMemory allocations bound directly as storage buffers, persistently mapped wherever their memory type allows (ADR 0012); completion is a nonblocking vkGetFenceStatus read over a bounded per-context ring of command buffers, fences, and descriptor sets (ADR 0006); no worker thread exists.

The native module compiles on the host operating systems enumerated by build.rs (va_vulkan). Loader absence is a run-time fact reported as InitError::RuntimeUnavailable, never a build probe. VIRTIO_ACCEL_VULKAN=0 forces the placeholder, =1 makes an unsupported target a loud build failure. The design decisions live in docs/adr/ (ADRs 0001–0013); ADR 0010 records the kernel geometries and the benchmark (cargo bench -p virtio-accel-vulkan) that measures them, ADR 0011 the MATMUL numerics (fused multiply-add, split-k), and ADR 0012 the mapped Device domain on unified-memory devices and the boundary views.

Modules§

nvfp4
Provider-native NVFP4 matrix products.
shader
The crate-authored SPIR-V compute kernels (ADR 0003, ADR 0007).

Structs§

LiveResources
Live provider resource totals for accounting hooks.
VulkanAccelerator
VulkanBuffer
One dedicated VkBuffer + VkDeviceMemory, persistently mapped unless device-local.
VulkanContext
Vulkan context handle: pools plus the bounded submission ring.
VulkanEvent
One submission: a claimed ring slot, its fence, and the guards it holds until terminal.
VulkanGateSignal
The raising half of a VulkanHostGate, for any thread.
VulkanHostGate
Vulkan backend instance bound to one physical device. A timeline semaphore the host raises, which VulkanAccelerator::submit_after submissions wait on (ADR 0013): work is queued before its inputs exist, and whoever produces them (a thread finishing a storage read, say) releases it with a VulkanGateSignal, without a round trip through the thread that owns the backend.
VulkanOptions
Backend options a host may set when opening a device. The defaults are what VulkanAccelerator::new and VulkanAccelerator::with_device use.
VulkanProgram
Resident compute pipelines specialized for one admitted TOSA graph.
VulkanQueue
Vulkan execution queue handle. Every queue of a context feeds the device’s one compute queue.

Enums§

InitError
Failure to initialize a Vulkan backend instance.
LoweringError
Why an artifact was not admitted.

Constants§

REQUIRED_RESIDENT_BYTES
Vulkan publishes no bound on what a driver retains for a compiled pipeline, so the provider promises the maximal charge; anything less would under-report ArtifactRef::resident_bytes.
VULKAN_TOSA_CAPABILITY
The FP32 tier’s admitted boundary: exactly what the generated kernels execute.
VULKAN_TOSA_FP8_CAPABILITY
The FP8 tier’s admitted boundary (ADR 0009): (FP8, FP8) -> FP16 MATMUL and exact FP8 data movement. Like the FP16 tier it needs no device feature — the widening is crate-owned integer and binary32 code and nothing writes FP8 except a raw byte copy — so it is advertised on every device the backend opens, with numerics identical everywhere.
VULKAN_TOSA_FP8_TARGET
The FP8 tier’s target: TOSA 1.0, floating-point profile, level 8K, both FP8 extensions.
VULKAN_TOSA_FP16_CAPABILITY
The FP16 tier’s admitted boundary (ADR 0008): the same 42 operators and graph envelope with binary16 tensors in every role. The tier needs no device feature — the kernels’ binary16 conversions are crate-owned integer and binary32 code, the implementation choice TOSA 1.0 §1.10.3 names explicitly — so it is advertised on every device the backend opens, with numerics identical to the FP32 tier’s everywhere.
VULKAN_TOSA_INTEGER_TARGET
The provisional integer tier: TOSA 1.0, integer profile, level 8K, no extensions.
VULKAN_TOSA_TARGET
The FP32 base tier: TOSA 1.0, floating-point profile, level 8K, no extensions.

Functions§

supports_tosa_dtype
Whether the advertised FP16 tier exposes dtype at a program boundary.
supports_tosa_operator
Whether the FP32 tier admits op.