Struct QuantScheme
pub struct QuantScheme {
pub value: QuantValue,
pub store: QuantStore,
pub mode: QuantMode,
/* private fields */
}Expand description
Describes a quantization scheme/configuration.
Scales come at up to two levels, each an optional field set through
per_tensor and per_block in any order:
// One scale for the whole tensor, stored as f32. Also what a scheme with no level resolves to.
QuantScheme::default().per_tensor(ScaleDtype::F32);
// One scale per block of 32 values.
QuantScheme::default().per_block([32], ScaleDtype::F32);
// Two levels: ue4m3 block scales, normalized by a single per-tensor f32 scale.
QuantScheme::default()
.per_block([16], ScaleDtype::UE4M3)
.per_tensor(ScaleDtype::F32);A two-level scheme exists so block scales can live in a narrow type: the global per-tensor scale
absorbs the tensor’s dynamic range, and the block dtype only covers the spread between blocks.
That spread is still bounded: a block whose scale falls further below the largest one than the
block dtype can express is stored at that dtype’s smallest value, far too coarse for it, and
every value in the block quantizes to zero. ScaleDtype::UE4M3 spans about 2^18 this way, so
a tensor holding a genuine outlier can lose its ordinary values.
Fields§
§value: QuantValueThe logical data type of quantized input values (e.g., QuantValue::Q8F).
This defines how values are interpreted during computation, independent of how they’re stored.
store: QuantStoreData type used for storing quantized values.
mode: QuantModeQuantization mode (e.g., symmetric).
Implementations§
§impl QuantScheme
impl QuantScheme
pub fn nvfp4() -> QuantScheme
pub fn nvfp4() -> QuantScheme
The NVFP4 format: fp4 (e2m1) values in blocks of 16, with ue4m3 block scales normalized by one per-tensor f32 scale.
§impl QuantScheme
impl QuantScheme
pub fn with_mode(self, mode: QuantMode) -> QuantScheme
pub fn with_mode(self, mode: QuantMode) -> QuantScheme
Set the quantization mode.
pub fn with_value(self, value: QuantValue) -> QuantScheme
pub fn with_value(self, value: QuantValue) -> QuantScheme
Set the data type used for quantized values.
pub fn with_store(self, store: QuantStore) -> QuantScheme
pub fn with_store(self, store: QuantStore) -> QuantScheme
Set the data type used to store quantized values.
pub fn per_tensor(self, dtype: ScaleDtype) -> QuantScheme
pub fn per_tensor(self, dtype: ScaleDtype) -> QuantScheme
Set the per-tensor scale level, stored as dtype.
pub fn per_block(
self,
block: impl AsRef<[u8]>,
dtype: ScaleDtype,
) -> QuantScheme
pub fn per_block( self, block: impl AsRef<[u8]>, dtype: ScaleDtype, ) -> QuantScheme
Set the per-block scale level: one scale per block of block values, stored as dtype.
pub fn tensor_scale(&self) -> Option<ScaleDtype>
pub fn tensor_scale(&self) -> Option<ScaleDtype>
The per-tensor scale level, the global level when a block level is present.
A scheme storing no level at all resolves here to a per-tensor f32 scale; the resolution
is not stored, so such a scheme compares equal to Default, not to an explicit
per_tensor(F32).
pub fn block_scale(&self) -> Option<BlockScale>
pub fn block_scale(&self) -> Option<BlockScale>
The per-block scale level, the innermost when both levels are present.
pub fn num_levels(&self) -> usize
pub fn num_levels(&self) -> usize
The number of scale levels: as many scale tensors ride along with the values.
pub fn scale_dtype(&self) -> ScaleDtype
pub fn scale_dtype(&self) -> ScaleDtype
The innermost level’s scale dtype, the type the per-position scales are stored in.
pub fn block_size(&self) -> Option<BlockSize>
pub fn block_size(&self) -> Option<BlockSize>
The block level’s size, or None for per-tensor quantization.
pub fn swap_block_dims(&mut self, rank: usize, dim0: usize, dim1: usize)
pub fn swap_block_dims(&mut self, rank: usize, dim0: usize, dim1: usize)
Swap two tensor dimensions in the block level, mirroring shape.swap(dim0, dim1). The
per-tensor level is unaffected.
dim0/dim1 are bare indices on purpose, mirroring [T]::swap’s own signature.
pub fn permute_block_dims(&mut self, rank: usize, axes: &[usize])
pub fn permute_block_dims(&mut self, rank: usize, axes: &[usize])
Permute the block level, mirroring a permutation of the tensor’s axes. The per-tensor level is unaffected.
pub fn size_bits_stored(&self) -> usize
pub fn size_bits_stored(&self) -> usize
Returns the size of the quantization storage type in bits.
pub fn size_bits_value(&self) -> usize
pub fn size_bits_value(&self) -> usize
Returns the size of the quantization storage type in bits.
pub fn num_quants(&self) -> usize
pub fn num_quants(&self) -> usize
Returns the number of quantized values stored in a single element.
pub fn native_packing(&self) -> usize
pub fn native_packing(&self) -> usize
Returns the native packing factor for the values. When native packing > 1, the packed
representation stores num_quants elements grouped into packs of native_packing size.
pub fn packing_dim(&self) -> Option<usize>
pub fn packing_dim(&self) -> Option<usize>
Returns the packing dim for the store.
pub fn swap_packing_dim(&mut self, dim0: usize, dim1: usize)
pub fn swap_packing_dim(&mut self, dim0: usize, dim1: usize)
Swaps the packing dim if it’s either of dim0 or dim1.
Executes the corresponding update to shape.swap(dim0, dim1).
Trait Implementations§
§impl Clone for QuantScheme
impl Clone for QuantScheme
§fn clone(&self) -> QuantScheme
fn clone(&self) -> QuantScheme
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreimpl Copy for QuantScheme
§impl Debug for QuantScheme
impl Debug for QuantScheme
§impl Default for QuantScheme
impl Default for QuantScheme
§fn default() -> QuantScheme
fn default() -> QuantScheme
§impl<'de> Deserialize<'de> for QuantScheme
impl<'de> Deserialize<'de> for QuantScheme
§fn deserialize<__D>(
__deserializer: __D,
) -> Result<QuantScheme, <__D as Deserializer<'de>>::Error>where
__D: Deserializer<'de>,
fn deserialize<__D>(
__deserializer: __D,
) -> Result<QuantScheme, <__D as Deserializer<'de>>::Error>where
__D: Deserializer<'de>,
impl Eq for QuantScheme
§impl Hash for QuantScheme
impl Hash for QuantScheme
§impl Ord for QuantScheme
impl Ord for QuantScheme
§fn cmp(&self, other: &QuantScheme) -> Ordering
fn cmp(&self, other: &QuantScheme) -> Ordering
1.21.0 (const: unstable) · Source§fn max(self, other: Self) -> Selfwhere
Self: Sized,
fn max(self, other: Self) -> Selfwhere
Self: Sized,
§impl PartialEq for QuantScheme
impl PartialEq for QuantScheme
§impl PartialOrd for QuantScheme
impl PartialOrd for QuantScheme
§impl Serialize for QuantScheme
impl Serialize for QuantScheme
§fn serialize<__S>(
&self,
__serializer: __S,
) -> Result<<__S as Serializer>::Ok, <__S as Serializer>::Error>where
__S: Serializer,
fn serialize<__S>(
&self,
__serializer: __S,
) -> Result<<__S as Serializer>::Ok, <__S as Serializer>::Error>where
__S: Serializer,
impl StructuralPartialEq for QuantScheme
Auto Trait Implementations§
impl Freeze for QuantScheme
impl RefUnwindSafe for QuantScheme
impl Send for QuantScheme
impl Sync for QuantScheme
impl Unpin for QuantScheme
impl UnsafeUnpin for QuantScheme
impl UnwindSafe for QuantScheme
Blanket Implementations§
impl<T> Boilerplate for T
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
§impl<Q, K> Comparable<K> for Q
impl<Q, K> Comparable<K> for Q
§impl<K, Q> Comparable<Q> for K
impl<K, Q> Comparable<Q> for K
impl<T> DeserializeOwned for Twhere
T: for<'de> Deserialize<'de>,
§impl<Q, K> Equivalent<K> for Q
impl<Q, K> Equivalent<K> for Q
§fn equivalent(&self, key: &K) -> bool
fn equivalent(&self, key: &K) -> bool
key and return true if they are equal.§impl<Q, K> Equivalent<K> for Q
impl<Q, K> Equivalent<K> for Q
§fn equivalent(&self, key: &K) -> bool
fn equivalent(&self, key: &K) -> bool
§impl<K, Q> Equivalent<Q> for K
impl<K, Q> Equivalent<Q> for K
§fn equivalent(&self, key: &Q) -> bool
fn equivalent(&self, key: &Q) -> bool
key and return true if they are equal.impl<T> ErasedDestructor for Twhere
T: 'static,
§impl<T> Instrument for T
impl<T> Instrument for T
§fn instrument(self, span: Span) -> Instrumented<Self>
fn instrument(self, span: Span) -> Instrumented<Self>
§fn in_current_span(self) -> Instrumented<Self>
fn in_current_span(self) -> Instrumented<Self>
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self>
fn into_either(self, into_left: bool) -> Either<Self, Self>
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more§impl<T> Pointable for T
impl<T> Pointable for T
§impl<T> PolicyExt for Twhere
T: ?Sized,
impl<T> PolicyExt for Twhere
T: ?Sized,
impl<T> Read<Exclusive, BecauseExclusive> for Twhere
T: ?Sized,
Source§impl<R, P> ReadPrimitive<R> for P
impl<R, P> ReadPrimitive<R> for P
Source§fn read_from_little_endian(read: &mut R) -> Result<Self, Error>
fn read_from_little_endian(read: &mut R) -> Result<Self, Error>
ReadEndian::read_from_little_endian().