pub enum CompilerSpec {
Identity {
elements: usize,
},
Int8Identity {
elements: usize,
line_size: usize,
},
Fp8ToBf16 {
format: Fp8Format,
elements: usize,
},
Matmul {
m: usize,
k: usize,
n: usize,
},
Fp8Matmul {
format: Fp8Format,
m: usize,
k: usize,
n: usize,
},
Int8Matmul {
m: usize,
k: usize,
n: usize,
left_zero_point: i8,
right_zero_point: i8,
},
Int32ToInt8Rescale {
elements: usize,
multiplier: i32,
shift: i8,
output_zero_point: i8,
},
MaxPool2d {
input_h: usize,
input_w: usize,
channels: usize,
output_h: usize,
output_w: usize,
kernel_h: usize,
kernel_w: usize,
stride_h: usize,
stride_w: usize,
},
}Expand description
A validated operator specialization ready for the compiler helper. Each variant names its input and output dtypes; the closed shape is integers only, so no guest bytes cross the boundary.
Variants§
Identity
BF16 → BF16 elementwise copy of elements values (a positive multiple of
1,024 values).
Int8Identity
Exact INT8 → INT8 elementwise copy. line_size is the complete DMA line selected by
admission; no tail staging or dtype conversion is permitted.
Fp8ToBf16
Explicit FP8 storage conversion to BF16. Every finite source value is exactly representable, so the output is bit-exact except for permitted NaN canonicalization.
Matmul
BF16 × BF16 → FP32 matrix multiply C[M, N] = A[M, K] · B[K, N] (batch 1). Each of m,
k, n is a positive multiple of the corresponding MATMUL tile dimension and at most
512. The FP32 output is the TOSA-mandated accumulator (issue #82).
Fp8Matmul
Fused FP8 × FP8 → FP32 matrix multiply (batch 1): the graph’s explicit BF16 promotion is performed on the compute core, per L1 tile, instead of through DDR.
Numerically identical to CompilerSpec::Fp8ToBf16 followed by CompilerSpec::Matmul —
FP8 → BF16 is exact for every encoding and the multiply is the same BF16 → FP32 kernel — so
fusing is a placement choice, not a change of numerical contract. It removes the
caller-visible BF16 tensors and their DDR round trip.
Int8Matmul
Exact zero-point-aware INT8 × INT8 → INT32 matrix multiply (batch 1).
The serialized TOSA zero points are part of the specialization and therefore also part of
the compiler cache key. Arithmetic is (a - a_zp) * (b - b_zp) accumulated in INT32.
Int32ToInt8Rescale
Exact signed INT32 → INT8 scale32 RESCALE with one shared multiplier and shift.
MaxPool2d
Batch-1 BF16 NHWC MAX_POOL2D with zero padding and propagating NaNs. The complete static specialization is carried to the helper so no TOSA bytes cross the subprocess boundary.
Trait Implementations§
Source§impl Clone for CompilerSpec
impl Clone for CompilerSpec
Source§fn clone(&self) -> CompilerSpec
fn clone(&self) -> CompilerSpec
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreimpl Copy for CompilerSpec
Source§impl Debug for CompilerSpec
impl Debug for CompilerSpec
impl Eq for CompilerSpec
Source§impl Hash for CompilerSpec
impl Hash for CompilerSpec
Source§impl PartialEq for CompilerSpec
impl PartialEq for CompilerSpec
Source§fn eq(&self, other: &CompilerSpec) -> bool
fn eq(&self, other: &CompilerSpec) -> bool
self and other values to be equal, and is used by ==.