Module quantization
Expand description
Quantization data representation.
Structs§
- Block
Layout - Which block each element of a tensor falls in, for a block scheme. A block is a rectangle, so its members are a run of the flat storage only when it spans the trailing dimension; anything walking values against block scales asks here rather than chunking.
- Block
Scale - The per-block scale level of a
QuantScheme: one scale per block of values. - Block
Size - Copyable block size, specialized version of
SmallVec. - QParam
Tensor - A quantization parameter tensor descriptor.
- QParams
- The quantization tensor data parameters.
- Quant
Scheme - Describes a quantization scheme/configuration.
- Quantization
Parameters Primitive - The quantization parameters primitive.
- Quantized
Bytes - Quantized data bytes representation.
Enums§
- Calibration
- Calibration method used to compute the quantization range mapping.
- Quant
Mode - Strategy used to quantize values.
- Quant
Propagation - Specify if the output of an operation is quantized using the scheme of the input or returned unquantized.
- Quant
Store - Data type used to stored quantized values.
- Quant
Value - Data type used to represent quantized values.
- Scale
Dtype - The data type a scale level stores its scales in.
Constants§
- QPARAM_
ALIGN - Alignment (in bytes) for quantization parameters in serialized tensor data.
Functions§
- compute_
q_ params - Compute the quantization parameters.
- compute_
range - Compute the quantization range mapping.
- global_
scale_ dtype - The dtype of the per-tensor scale block scales are normalized against, for a two-level scheme.
- params_
shape - Calculate the shape of the block scale grid for a given tensor and scheme.
- quantizable
- Whether a backend can quantize against this scheme’s scales.
- scale_
to_ dtype - Round a scale up to the smallest value representable by the scale dtype that is no smaller.