Skip to main content

KernelKey

Enum KernelKey 

Source
pub enum KernelKey {
    Nvfp4Matmul {
        buffers: u32,
        cooperative: bool,
        subgroup: bool,
    },
    Elementwise {
        op: ElementwiseOp,
        float: Storage,
        broadcast: bool,
        workgroup: u32,
        buffers: u32,
    },
    Reduce {
        op: ReduceOp,
        float: Storage,
        workgroup: u32,
        buffers: u32,
    },
    Matmul {
        input: Storage,
        output: Storage,
        tile: u32,
        buffers: u32,
    },
    MatmulStream {
        rhs: Storage,
        output: Storage,
        buffers: u32,
    },
    MaxPool {
        nan_mode: NanMode,
        float: Storage,
        workgroup: u32,
        buffers: u32,
    },
    Cast {
        input: Storage,
        output: Storage,
        workgroup: u32,
        buffers: u32,
    },
    Move {
        storage: Storage,
        contiguous: bool,
        workgroup: u32,
        buffers: u32,
    },
}
Expand description

One assembled kernel variant. Everything that changes instructions is in the key; everything that changes only numbers is a specialization constant.

Variants§

§

Nvfp4Matmul

F32 activations times row-major packed E2M1 weights and E4M3 block scales.

Fields

§buffers: u32
§cooperative: bool
§subgroup: bool
§

Elementwise

Elementwise lanes over count output elements; broadcast selects the strided multi-index addressing, otherwise every operand shares the output’s linear index. float is the storage of the operator’s floating-point tensors (Word for FP32, Half for FP16); BOOL lanes are byte storage in either variant. Binary16 lanes evaluate in binary32 and narrow once, except the integer NEGATE/ABS sign lanes (ADR 0008).

Fields

§float: Storage
§broadcast: bool
§workgroup: u32
§buffers: u32
§

Reduce

Axis reduction: one invocation per output element, sequential ascending-axis fold. The input is read at float storage; sums and products fold in binary32 (the TOSA accumulator width), and the output is stored back at float storage.

Fields

§float: Storage
§workgroup: u32
§buffers: u32
§

Matmul

Batched, register-tiled matrix multiplication: a tile × tile workgroup computes a matmul_block-sided output square from workgroup-shared binary32 slabs. Both operands are read at input storage and accumulate in binary32 — the accumulator width TOSA assigns FP16 and FP8 MATMUL — and the result is stored at output storage. The two differ only for the FP8 tier, where TOSA defines MATMUL as (FP8, FP8) -> FP16.

Fields

§input: Storage
§output: Storage
§tile: u32
§buffers: u32
§

MatmulStream

Split-k streaming MATMUL for m ≤ STREAM_ROWS rows: a 1-D workgroup of STREAM_WORKGROUP invocations over STREAM_COLUMNS columns of rhs (read at rhs storage, a word per invocation), the lhs read as binary32 words — lowering widens a narrower lhs beforehand — and one fixed-order reduction of the k slices at the end. Same accumulator width as Self::Matmul.

Fields

§output: Storage
§buffers: u32
§

MaxPool

NHWC max pooling with padding excluded from the window, at float storage.

Fields

§nan_mode: NanMode
§float: Storage
§workgroup: u32
§buffers: u32
§

Cast

Elementwise float conversion: read at input storage, write at output. One dispatch per CAST between float dtypes, including both FP8 directions (ADR 0009).

Fields

§input: Storage
§output: Storage
§workgroup: u32
§buffers: u32
§

Move

Strided copy over a rank-MAX_RANK iteration space (TRANSPOSE, REVERSE, CONCAT segments); contiguous collapses to a linear copy.

Fields

§storage: Storage
§contiguous: bool
§workgroup: u32
§buffers: u32

Implementations§

Source§

impl KernelKey

Source

pub fn assemble(self) -> Vec<u32>

Assemble the module for this variant.

Source

pub fn every_variant() -> Vec<KernelKey>

Every kernel variant this backend can assemble, at one representative tuning. The SPIR-V validation sweep and the specialization-count test both walk this list, so a new variant is validated the moment it is added here.

Source

pub const fn spec_constant_count(self) -> u32

Number of specialization constants the module declares, ids 0..count.

Source

pub const fn local_size(self) -> [u32; 3]

OpExecutionMode LocalSize of the module.

Trait Implementations§

Source§

impl Clone for KernelKey

Source§

fn clone(&self) -> KernelKey

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Copy for KernelKey

Source§

impl Debug for KernelKey

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl Eq for KernelKey

Source§

impl Hash for KernelKey

Source§

fn hash<__H: Hasher>(&self, state: &mut __H)

Feeds this value into the given Hasher. Read more
1.3.0 · Source§

fn hash_slice<H>(data: &[Self], state: &mut H)
where H: Hasher, Self: Sized,

Feeds a slice of this type into the given Hasher. Read more
Source§

impl PartialEq for KernelKey

Source§

fn eq(&self, other: &KernelKey) -> bool

Tests for self and other values to be equal, and is used by ==.
1.0.0 (const: unstable) · Source§

fn ne(&self, other: &Rhs) -> bool

Tests for !=. The default implementation is almost always sufficient, and should not be overridden without very good reason.
Source§

impl StructuralPartialEq for KernelKey

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.