Skip to main content

Module quantization

Module quantization 

Expand description

Quantization data representation.

Structs§

BlockLayout
Which block each element of a tensor falls in, for a block scheme. A block is a rectangle, so its members are a run of the flat storage only when it spans the trailing dimension; anything walking values against block scales asks here rather than chunking.
BlockScale
The per-block scale level of a QuantScheme: one scale per block of values.
BlockSize
Copyable block size, specialized version of SmallVec.
QParamTensor
A quantization parameter tensor descriptor.
QParams
The quantization tensor data parameters.
QuantScheme
Describes a quantization scheme/configuration.
QuantizationParametersPrimitive
The quantization parameters primitive.
QuantizedBytes
Quantized data bytes representation.

Enums§

Calibration
Calibration method used to compute the quantization range mapping.
QuantMode
Strategy used to quantize values.
QuantPropagation
Specify if the output of an operation is quantized using the scheme of the input or returned unquantized.
QuantStore
Data type used to stored quantized values.
QuantValue
Data type used to represent quantized values.
ScaleDtype
The data type a scale level stores its scales in.

Constants§

QPARAM_ALIGN
Alignment (in bytes) for quantization parameters in serialized tensor data.

Functions§

compute_q_params
Compute the quantization parameters.
compute_range
Compute the quantization range mapping.
global_scale_dtype
The dtype of the per-tensor scale block scales are normalized against, for a two-level scheme.
params_shape
Calculate the shape of the block scale grid for a given tensor and scheme.
quantizable
Whether a backend can quantize against this scheme’s scales.
scale_to_dtype
Round a scale up to the smallest value representable by the scale dtype that is no smaller.