Skip to main content

Module numerics

Module numerics 

Source
Expand description

Device-neutral numerical acceptance cases for hardware backends.

Each case couples a stable TOSA artifact with exact input shapes and a numerical oracle. Host backends consume the same bytes and values, so a provider cannot quietly substitute a backend-specific graph while claiming cross-device equivalence.

Structs§

Bfloat16Tensor
One immutable bfloat16 tensor represented by exact storage bits.
Float16Tensor
One immutable IEEE-754 binary16 tensor in a numerical acceptance case.
Float32Tensor
One immutable FP32 tensor in a numerical acceptance case.
Int32Tensor
One immutable INT32 tensor produced by an integer-profile acceptance case.
PackedTensor
One immutable packed low-precision tensor.
TosaBfloat16Case
A stable TOSA EXT-BF16 graph and bit-exact bfloat16 numerical oracle.
TosaFloat16Case
A stable TOSA graph and binary16 numerical oracle shared by host backends.
TosaFloat32Case
A stable TOSA graph and FP32 numerical oracle shared by host backends.
TosaFp8ToBfloat16Case
A stable explicit FP8 → BF16 CAST with a bit-exact output oracle.
TosaFp32OperatorCase
One FP32-tier operator case shared by host backends: a stable TOSA 1.0 graph over FP32 tensors with BOOL and INT32 auxiliaries, and a numerical oracle for its single output.
TosaInt8MatmulCase
A stable TOSA INT8 matrix multiplication with exact INT32 accumulation.
TosaInt32ToInt8RescaleCase
A stable TOSA INT32-to-INT8 RESCALE with exact fixed-point rounding and saturation.
TosaPackedCase
A stable TOSA graph and packed low-precision oracle shared by host backends.
TosaRawCase
One mixed-type TOSA operator case with a numerical output oracle.

Enums§

Fp32TierTensor
Raw tensor storage used by the shared FP32-tier operator cases.
PackedDType
Packed scalar encoding used by a low-precision TOSA acceptance case.
RawTensor
Raw tensor storage used by mixed-type Hexagon operator-parity cases.

Constants§

ADD_FP16
Binary16 broadcast addition over [2, 1] and [1, 3] inputs.
CAST_FP8E4M3_TO_BF16
Explicit FP8 E4M3 → BF16 CAST spanning signed zeros, subnormals, ordinary values, finite maximum, and NaN. All 256 byte encodings repeat four times across one XDNA conversion tile.
CAST_FP8E5M2_TO_BF16
Explicit FP8 E5M2 → BF16 CAST spanning signed zeros, subnormals, one, finite maximum, infinity, and NaN. All 256 byte encodings repeat four times across one XDNA conversion tile.
FP32_BINARY_CASES
FP32 broadcast binary operators over [2, 1] and [1, 3] inputs producing [2, 3].
FP32_LOGICAL_CASES
FP32 comparisons, BOOL logic, and selection.
FP32_MOVEMENT_CASES
Static FP32 constants and data-movement operators.
FP32_OPERATOR_CASE_GROUPS
Every FP32-tier operator case group, for backends that iterate the whole tier.
FP32_REDUCTION_CASES
FP32 reductions and INT32 argmax over a [2, 3] input along axis 1.
FP32_UNARY_CASES
FP32 unary and activation operators over [0.5, 1, 2, 4].
HEXAGON_LOGICAL_CASES
Mixed BOOL/FP16 comparison, logical, and selection cases.
HEXAGON_MOVEMENT_CASES
Static constants and FP16 data-movement cases.
HEXAGON_REDUCTION_CASES
FP16 reductions and INT32 argmax over a two-row input.
HEXAGON_UNARY_FP16_CASES
FP16 unary and activation cases supported by the QNN HTP operator package.
IDENTITY_EDGES_FP16
Binary16 identity over NaN, infinities, signed zeros, a subnormal, and finite values.
IDENTITY_EDGES_FP32
FP32 identity over NaN, infinities, signed zeros, a subnormal, and ordinary finite values.
IDENTITY_FP8E4M3
FP8 E4M3 identity over signed zeros, subnormals, ordinary values, finite maximum, and NaN.
IDENTITY_FP8E5M2
FP8 E5M2 identity over signed zeros, subnormals, one, finite maximum, infinity, and NaN.
IDENTITY_INT4
Packed INT4 identity spanning the TOSA-defined finite range.
IDENTITY_INT8
INT8 identity spanning negative, zero, and positive values.
LINEAR_TANH_FP32
Three-operator FP32 graph: tanh(x · w + bias) with constant zero points and a [1, 1, 2] bias broadcast over the [1, 2, 2] product. Inputs are the mock classifier’s features and weights; the oracle is evaluated in binary64 and rounded.
MATMUL_FP16
Binary16 batched matrix multiplication with non-square operands.
MATMUL_FP32
FP32 batched matrix multiplication with non-square operands.
MATMUL_INT8
Non-square INT8 batched matrix multiplication with nonzero zero points and INT32 accumulation.
MAXIMUM_FP16
Binary16 broadcast maximum over [2, 1] and [1, 3] inputs.
MAX_POOL2D_BF16
BF16 two-channel NHWC max pooling with a 2x2 kernel, stride two, zero padding, and an exact integer-valued oracle.
MAX_POOL2D_FP16
Binary16 two-channel NHWC max pooling with a 2x2 kernel and stride two.
MAX_POOL2D_FP32
FP32 two-channel NHWC max pooling with a 2x2 kernel and stride two.
MINIMUM_FP16
Binary16 broadcast minimum over [2, 1] and [1, 3] inputs.
MOCK_LINEAR_CLASSIFIER_FP16
Two-sample, three-feature FP16 linear classifier with a direct-bound 3x2 weight matrix.
MUL_FP16
Binary16 broadcast multiplication with a compile-time zero shift.
POW_FP16
Binary16 broadcast power over [2, 1] and [1, 3] inputs.
QUANTIZED_CLASSIFIER_INT8
Exact INT8 quantized linear classifier: two samples, three features, two classes, INT32 logits with unambiguous per-sample argmax winners.
RESCALE_INT32_TO_INT8
Signed scale32 RESCALE covering ties, negative values, saturation, and a nonzero output point.
SUB_FP16
Binary16 broadcast subtraction over [2, 1] and [1, 3] inputs.