pub const fn stream_matmul_shared_bytes() -> u32
Bytes of workgroup-shared memory the streaming kernel declares for its final reduction: every invocation’s STREAM_ROWS × lanes partial sums, at the widest lane count.
STREAM_ROWS × lanes