Skip to main content

f32_to_fp8_bits

Function f32_to_fp8_bits 

Source
pub fn f32_to_fp8_bits(format: Fp8Format, value: f32) -> u8
Expand description

The exact binary32 value of a binary16 bit pattern (host side of the kernels’ unpack: every binary16 value, subnormals included, is exactly representable in binary32; NaN payloads are preserved). Narrow binary32 to one FP8 encoding with round-to-nearest, ties-to-even.

Overflow policy. TOSA 1.0 through 1.2 leave float-to-FP8 overflow undefined, so this is the crate’s policy rather than the spec’s. A magnitude too large for E4M3 — which has no infinity — becomes NaN, not the finite maximum. The reason is that the alternatives are not symmetric: a consumer who wants saturation can CLAMP in a wider dtype before the CAST and get it exactly, while a consumer handed a saturated 448 cannot tell it from a value that was always 448. NaN preserves the choice; saturation destroys it. It also matches what this crate already does one format up, where narrow_f16 signals unrepresentability as infinity rather than clamping to 65504.

E5M2 needs no policy: it is IEEE-shaped, so overflow becomes infinity like any binary float. NaN in becomes a canonical quiet NaN out, sign preserved, for both encodings.