pub fn f32_to_fp8_bits(format: Fp8Format, value: f32) -> u8Expand description
The exact binary32 value of a binary16 bit pattern (host side of the kernels’ unpack: every binary16 value, subnormals included, is exactly representable in binary32; NaN payloads are preserved). Narrow binary32 to one FP8 encoding with round-to-nearest, ties-to-even.
Overflow policy. TOSA 1.0 through 1.2 leave float-to-FP8 overflow undefined, so this is
the crate’s policy rather than the spec’s. A magnitude too large for E4M3 — which has no
infinity — becomes NaN, not the finite maximum. The reason is that the alternatives are not
symmetric: a consumer who wants saturation can CLAMP in a wider dtype before the CAST
and get it exactly, while a consumer handed a saturated 448 cannot tell it from a value that
was always 448. NaN preserves the choice; saturation destroys it. It also matches what this
crate already does one format up, where narrow_f16 signals unrepresentability as infinity
rather than clamping to 65504.
E5M2 needs no policy: it is IEEE-shaped, so overflow becomes infinity like any binary float. NaN in becomes a canonical quiet NaN out, sign preserved, for both encodings.