Expand description
Vendor-neutral Vulkan compute host backend for virtio-accel.
The native path binds Vulkan 1.3 through the pinned [ash] crate, loading the platform’s
Vulkan loader at run time (ADR 0002). It executes device-neutral TOSA 1.0 programs admitted by
lower — the 42 FP32-tier operators shared with the
Core ML and OpenVINO backends over FP32 and FP16 tensors, with BOOL/INT32 auxiliaries —
on crate-authored SPIR-V compute kernels specialized at load_program (ADR 0003, ADR 0007).
The FP16 tier needs no device feature: its binary16 conversions are crate-owned integer and
binary32 code (ADR 0008), so every device that hosts the FP32 tier hosts the FP16 tier with
identical numerics. A whole graph is one submission: constants and intermediates live in a
per-program arena and dependent dispatches are separated by compute barriers. Buffers are
dedicated VkDeviceMemory allocations bound directly as storage buffers, persistently mapped
wherever their memory type allows (ADR 0012); completion is a
nonblocking vkGetFenceStatus read over a bounded per-context ring of command buffers,
fences, and descriptor sets (ADR 0006); no worker thread exists.
The native module compiles on the host operating systems enumerated by build.rs (va_vulkan).
Loader absence is a run-time fact reported as InitError::RuntimeUnavailable, never a build
probe. VIRTIO_ACCEL_VULKAN=0 forces the placeholder, =1 makes an unsupported target a loud
build failure. The design decisions live in docs/adr/ (ADRs 0001–0013);
ADR 0010 records the kernel geometries and the benchmark (cargo bench -p virtio-accel-vulkan)
that measures them, ADR 0011 the MATMUL numerics (fused multiply-add, split-k), and
ADR 0012 the mapped Device domain on unified-memory devices and the boundary views.
Modules§
- nvfp4
- Provider-native NVFP4 matrix products.
- shader
- The crate-authored SPIR-V compute kernels (ADR 0003, ADR 0007).
Structs§
- Live
Resources - Live provider resource totals for accounting hooks.
- Vulkan
Accelerator - Vulkan
Buffer - One dedicated
VkBuffer+VkDeviceMemory, persistently mapped unless device-local. - Vulkan
Context - Vulkan context handle: pools plus the bounded submission ring.
- Vulkan
Event - One submission: a claimed ring slot, its fence, and the guards it holds until terminal.
- Vulkan
Gate Signal - The raising half of a
VulkanHostGate, for any thread. - Vulkan
Host Gate - Vulkan backend instance bound to one physical device.
A timeline semaphore the host raises, which
VulkanAccelerator::submit_aftersubmissions wait on (ADR 0013): work is queued before its inputs exist, and whoever produces them (a thread finishing a storage read, say) releases it with aVulkanGateSignal, without a round trip through the thread that owns the backend. - Vulkan
Options - Backend options a host may set when opening a device. The defaults are what
VulkanAccelerator::newandVulkanAccelerator::with_deviceuse. - Vulkan
Program - Resident compute pipelines specialized for one admitted TOSA graph.
- Vulkan
Queue - Vulkan execution queue handle. Every queue of a context feeds the device’s one compute queue.
Enums§
- Init
Error - Failure to initialize a Vulkan backend instance.
- Lowering
Error - Why an artifact was not admitted.
Constants§
- REQUIRED_
RESIDENT_ BYTES - Vulkan publishes no bound on what a driver retains for a compiled pipeline, so the provider
promises the maximal charge; anything less would under-report
ArtifactRef::resident_bytes. - VULKAN_
TOSA_ CAPABILITY - The FP32 tier’s admitted boundary: exactly what the generated kernels execute.
- VULKAN_
TOSA_ FP8_ CAPABILITY - The FP8 tier’s admitted boundary (ADR 0009):
(FP8, FP8) -> FP16MATMUL and exact FP8 data movement. Like the FP16 tier it needs no device feature — the widening is crate-owned integer and binary32 code and nothing writes FP8 except a raw byte copy — so it is advertised on every device the backend opens, with numerics identical everywhere. - VULKAN_
TOSA_ FP8_ TARGET - The FP8 tier’s target: TOSA 1.0, floating-point profile, level 8K, both FP8 extensions.
- VULKAN_
TOSA_ FP16_ CAPABILITY - The FP16 tier’s admitted boundary (ADR 0008): the same 42 operators and graph envelope with binary16 tensors in every role. The tier needs no device feature — the kernels’ binary16 conversions are crate-owned integer and binary32 code, the implementation choice TOSA 1.0 §1.10.3 names explicitly — so it is advertised on every device the backend opens, with numerics identical to the FP32 tier’s everywhere.
- VULKAN_
TOSA_ INTEGER_ TARGET - The provisional integer tier: TOSA 1.0, integer profile, level 8K, no extensions.
- VULKAN_
TOSA_ TARGET - The FP32 base tier: TOSA 1.0, floating-point profile, level 8K, no extensions.
Functions§
- supports_
tosa_ dtype - Whether the advertised FP16 tier exposes
dtypeat a program boundary. - supports_
tosa_ operator - Whether the FP32 tier admits
op.