Skip to main content

scale_to_dtype

Function scale_to_dtype 

pub fn scale_to_dtype(scale: f32, dtype: ScaleDtype) -> f32
Expand description

Round a scale up to the smallest value representable by the scale dtype that is no smaller.

Backends that keep scales in f32 must apply this when quantizing, so that the scale they divide by is the one that will actually be stored. Otherwise a tensor dequantizes differently after a save/load round trip.

Up rather than to nearest, because a scale is derived from the largest magnitude it has to cover. Rounding down puts that value past the end of the quantized range, where it clips, which measured several times worse than the coarser step rounding up costs.