Part of Language AI Handbook
Explains how weight quantization maps floating-point values to integers, reducing LLM memory by 4x. Topics include scale, zero-point.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Weight Quantization Basics: Scale, Zero-Point & Calibration
Large language models are memory hogs. A 7 billion parameter model stored in 16-bit floating point requires 14 GB just for the weights, and larger models scale proportionally: 70 billion parameters means roughly 140 GB of weight storage before you have loaded a single input token. The practical implication is stark. Even if you have a GPU powerful enough to run these models in principle, you may simply not have enough VRAM to hold the weights. The model never gets a chance to run. Memory capacity is the hard constraint that determines whether a model is deployable at all.
But even when you have enough memory, a second constraint governs inference speed. Modern GPUs are extraordinarily fast at computation, but moving data between GPU memory and the compute units costs time. Every forward pass must load billions of parameters across this memory bus, and this data movement dominates inference latency when the batch size is small. Most production LLM serving operates with small or unit batch sizes, because users arrive one at a time, and each request must be answered quickly. Under these conditions, the GPU spends most of its time waiting for data rather than computing. You have a sports car stuck in traffic.
Quantization provides a direct solution to both problems: represent model weights using fewer bits. Instead of 16 or 32 bits per parameter, we can use 8, 4, or even fewer bits. This shrinks memory footprint, reduces bandwidth requirements, and often enables hardware-accelerated integer arithmetic that runs faster than floating-point operations. A 4-bit quantized 70B model fits in roughly 35 GB instead of 140 GB, which makes it deployable on hardware that could never touch the original. The bandwidth reduction is proportional: 4-bit weights move four times less data across the memory bus compared to 16-bit weights, directly accelerating generation speed.
But quantization is not free. Cramming continuous values into discrete bins inevitably loses information. Floating-point numbers can represent a large range of values with fine granularity: a 32-bit float can distinguish values that differ by one part in ten million within a given range. When we compress that into 8 integers, we can only distinguish 256 different values across the same range. Some information is irretrievably lost. The goal of quantization is minimizing this information loss while maximizing compression, and it turns out that neural networks are remarkably tolerant of carefully applied quantization. The challenge lies in doing it carefully.
This chapter covers the basic concepts that every quantization technique builds upon. We start with the arithmetic of how floating-point values map to integers: what scale and zero-point mean, how to derive them, and what the resulting error looks like. We then examine the two main quantization schemes, symmetric and asymmetric, which trade computational simplicity against representation efficiency. We cover quantization granularity, from a single scale for an entire tensor to separate scales for each group of values, and explain why finer granularity dramatically improves accuracy at modest storage cost. Finally, we examine calibration: the process of determining optimal quantization parameters from data. Every advanced quantization method you will encounter in later chapters builds on these foundations. Understanding them deeply is the prerequisite for understanding everything that follows.
The use of reduced-precision arithmetic in neural networks predates the LLM era by several decades. Researchers in the 1990s explored fixed-point representations for neural network hardware because floating-point units were expensive. The modern resurgence of quantization research was driven by the intersection of two trends: the explosion in model size beginning around 2018, and the maturation of hardware accelerators with specialized integer compute units. Google's deployment of INT8 inference at scale in 2018, combined with NVIDIA's tensor core support for INT8, demonstrated that aggressive quantization could work in production. The subsequent development of 4-bit methods like GPTQ (2022), AWQ (2023), and NF4 (2023) showed that the accuracy-compression frontier could be pushed much further than anyone had previously believed practical. Today, quantization is as much a deployment tool as a research topic.
Why Memory Bandwidth Limits Inference
Understanding why quantization matters requires a brief look at the physics of modern GPU architecture. We often think of GPUs as compute devices, but for LLM inference, they behave primarily as memory devices. The compute units sit idle while waiting for data, and quantization is the most direct tool for addressing this basic bottleneck.
Modern GPUs have enormous compute throughput that vastly outpaces their memory systems. An NVIDIA A100 delivers 312 teraflops of FP16 compute but only 2 TB/s of memory bandwidth. These numbers sound abstract, so think of it this way: for every byte loaded from memory, the GPU can perform roughly 156 floating-point operations. If an operation requires loading 1 byte and doing 1 computation, the GPU sits idle for the other 155 operation slots. The compute is effectively free compared to the cost of moving data.
During autoregressive generation, each new token requires exactly one forward pass through the entire model, loading essentially all of the model's weight matrices once. A single matrix multiplication of shape , where is the model dimension and is the feed-forward dimension, loads parameters but performs only floating-point operations. This gives an arithmetic intensity of just 2 FLOPs per loaded element. When the hardware can sustain 156 FLOPs per byte, an arithmetic intensity of 2 means you are using roughly 1.3% of the available compute. The rest waits.
This situation, where memory bandwidth limits throughput rather than compute, is called memory-bound operation. It is the dominant regime for LLM inference at small batch sizes, which includes most real-world serving scenarios. Quantization directly attacks this bottleneck. Loading 4-bit weights instead of 16-bit weights moves four times less data across the memory bus. Even if dequantization adds some computational overhead, the memory savings dominate for memory-bound workloads. This explains why efficient LLM deployment now depends heavily on quantization.

The relationship between bit width and bandwidth is exactly linear: cutting bits in half cuts bandwidth in half. This linearity is what makes quantization so attractive as an optimization strategy. There is no engineering trick here, just arithmetic. Fewer bits per number means fewer bytes transferred, and fewer bytes transferred means faster inference.
The Quantization Mapping
The process of mapping values from a continuous (or high-precision) representation to a discrete (or lower-precision) representation. In neural network contexts, this typically means converting 32-bit or 16-bit floating-point numbers to 8-bit or smaller integers.
Quantization addresses a basic tension between precision and efficiency. Neural networks learn weights as floating-point numbers, which can represent an enormous range of values with fine granularity. However, this precision comes at a cost: each 32-bit float consumes four bytes of memory, and operations on floating-point numbers require specialized hardware circuits. Quantization works because neural networks can tolerate small perturbations in their weights. We do not need exact values; we need values that are close enough to preserve the network's learned behavior. The question is: how close is close enough, and how do we achieve the best approximation with a fixed bit budget?
Think of the process as converting a high-resolution photograph to a lower-resolution version. The original has fine detail, and the compressed version loses some of that detail. If you compress carefully, preserving the most important features, the result looks nearly identical at a casual glance. If you compress carelessly, throwing away structure that matters, the result looks blurry or distorted. Quantization is the same process applied to numbers instead of pixels: we want to choose our discrete grid points to best represent the actual values we have.
Consider a weight tensor with values ranging from to . We want to represent these using 8-bit signed integers, which can hold values from to . The mapping must preserve the relative relationships between values while fitting them into our target range. Think of this as creating a ruler where each tick mark represents one integer value, and we need to position values from the original continuous number line onto the nearest tick mark. The spacing of the tick marks (the scale) and where we place the zero mark (the zero-point) together determine the quality of the mapping.
The key insight is that this is a linear mapping problem. We are fitting a straight line between the real-valued domain and the integer domain, and the parameters of that line are the scale and zero-point. Different choices of these parameters distribute the available integer "buckets" differently across the real-valued range, and some distributions represent the actual data better than others.
The standard affine quantization formula maps a real value to a quantized integer :
where:
- : the quantized integer value
- : the input real value (floating-point)
- : the scale factor, determining the step size of the quantization grid
- : the zero-point, an integer value to which real zero is mapped
- : the rounding operation (typically to the nearest integer)
This formula captures the essence of what quantization does. First, we divide the real value by the scale, which converts from real units to "quantization units." The scale determines how much of the real number line each integer step represents. If the scale is 0.01, then adjacent integers represent values that differ by 0.01 in the original space. After scaling, we round to the nearest integer because integers cannot represent fractional values. Finally, we add the zero-point, which shifts the entire mapping so that real zero lands on a particular integer value rather than necessarily on integer zero.
The zero-point serves an important purpose: it allows the quantized representation to efficiently use the available integer range even when the original values are not centered at zero. If all your weights happen to be positive, you would waste half the signed integer range without a zero-point adjustment. The zero-point corrects this misalignment, sliding the entire quantization grid until it best overlaps with the data.
To recover the original value (approximately), we apply the inverse dequantization operation:
where:
- : the reconstructed real value (approximation of the original )
- : the scale factor
- : the quantized integer
- : the zero-point
The dequantization formula simply reverses the quantization process. We first subtract the zero-point to undo the shift, then multiply by the scale to convert back from integer units to real units. Notice the hat notation on , which signals that this is an approximation rather than the exact original value. The reconstructed value will not exactly equal due to rounding, and this difference constitutes quantization error. Our goal is choosing and to minimize this error across the tensor.

Determining Scale and Zero-Point
With the quantization formula established, we face a practical question: how do we choose the scale and zero-point? These parameters must be selected carefully because they determine both the range of representable values and the precision within that range. The intuition is straightforward: we want to map the full range of actual tensor values onto the full range of available integers, using every bit of precision the integer format provides.
Think of it as a translation and scaling problem. You have a thermometer calibrated from to degrees, and you want to map it onto a ruler with integer markings from to . If you scale correctly, the coldest temperature on the thermometer maps to and the warmest maps to , and every mark on the ruler corresponds to a well-defined temperature. If you scale poorly, the thermometer might only use a portion of the ruler's range, leaving most of the precision wasted.
We derive the quantization parameters by establishing a linear mapping where the real minimum maps to the integer minimum and maps to . This ensures that the smallest and largest values in our tensor precisely hit the boundaries of the integer range, leaving no representable integers unused.
where:
- : the minimum and maximum values in the real (floating-point) tensor
- : the minimum and maximum values in the target integer range (e.g., -128 and 127 for signed 8-bit)
- : the scale factor to be determined
- : the zero-point to be determined
These two equations represent the boundary conditions of our linear mapping. The first equation says that when we dequantize the minimum integer value, we should recover the minimum real value. The second says the same for the maximum. Together, they give us two equations with two unknowns, and , which we can solve algebraically.
Subtracting the first equation from the second eliminates , which is the key insight that makes this system tractable. The zero-point cancels because we are computing the difference between two dequantization equations, and the shift cancels out in the difference:
Solving for yields the scale factor:
where:
- : the scale factor representing the step size (real value per integer increment)
- : the total range of the continuous values
- : the total number of discrete steps available
This formula has a straightforward interpretation: the scale is simply the ratio of the real range to the integer range. If your real values span 1.0 units and you have 256 integers available, each integer step represents real units. The scale tells you exactly how much precision you have: smaller scales mean finer distinctions between values, while larger scales mean coarser quantization. A large scale arises when the real range is wide relative to the integer range, and it forces many real values to map to the same integer.
Substituting back into the first equation allows us to solve for the zero-point :
where:
- : the zero-point (in real arithmetic before rounding)
- : the minimum real value scaled to the integer domain
The zero-point equation determines where real zero falls in the integer range. If the real values are centered around zero, the zero-point will be near the middle of the integer range. If the real values are all positive, the zero-point will be at or near the minimum integer value. Geometrically, the zero-point specifies how many integer steps separate the integer origin from the point where real zero would land.
Since must be an integer, we round the result:
where:
- : the computed scale factor
- : the computed zero-point
- : the minimum value in the floating-point tensor
- : the minimum value of the target integer range
The rounding of the zero-point introduces a small complication: it means that real zero may not map exactly to integer , and the boundary conditions we started with may not be satisfied exactly. In practice, this slight imprecision is negligible compared to the rounding errors inherent in quantization itself. The zero-point rounding shifts the mapping by at most half an integer step, which is the same order as the rounding error affecting every value in the tensor.
For 8-bit unsigned integers, and , giving discrete steps. For 8-bit signed integers, and , giving the same 255 steps (note: not 256, because the range of -bit integers contains values but only steps between adjacent values). This means signed and unsigned 8-bit formats have the same quantization resolution; they differ only in where the range is centered.
Quantization Error
Every quantization operation introduces error because we are fundamentally discarding information. When we round a continuous value to the nearest integer, any fractional part of the scaled value is lost. Understanding the nature and magnitude of this error helps us make informed decisions about quantization parameters and bit widths.
The rounding operation introduces error that depends on where the original value falls relative to the quantization grid. Think of placing a set of points on a number line and then rounding each point to the nearest tick mark on a ruler. Points that fall very close to a tick mark move only slightly and incur small errors. Points that fall exactly halfway between two tick marks must be rounded to one or the other, incurring the maximum possible error of half a step.
For uniformly distributed values, the maximum error from rounding is . This occurs when a value falls exactly halfway between two adjacent quantization levels. On average, the rounding error follows a uniform distribution between and , with mean zero and variance:
where:
- : the variance of quantization rounding error
- : the scale (quantization step size)
This means smaller scales produce smaller errors but can only represent smaller ranges. There is a direct tradeoff: cutting the scale in half cuts the step size in half, which cuts the variance by a factor of four, but it also halves the representable range. If the true value lies outside the representable range, we must clip it to the nearest boundary value, potentially causing much larger errors for outliers.
The mean squared quantization error for a tensor depends on both the value distribution and the choice of and . Values near the center of the distribution incur only rounding error, bounded by . Values at the extremes are handled well when the scale is chosen to just accommodate them. But values beyond the representable range are clipped, and the error for a clipped value can be arbitrarily large. An outlier at gets clipped to , incurring an error of . This observation motivates many of the advanced quantization techniques we will explore in later chapters, which focus on handling outliers gracefully rather than letting them to corrupt the quantization of typical values.
The theoretical relationship between bit width and quantization error is elegant. Each additional bit doubles the number of representable levels, which halves the step size , which reduces the variance of the rounding error by a factor of four. In other words, quantization error decreases by roughly dB (a factor of 4 in variance, or 2 in amplitude) for each additional bit. This is why plots of quantization error versus bit width on a log scale show nearly straight lines with a slope of 4x per bit.
Symmetric vs Asymmetric Quantization
The two main quantization schemes differ in how they handle the zero-point. This choice significantly affects both accuracy and computational efficiency. Understanding when to use each scheme requires considering both the mathematical properties of the data and the practical constraints of the target hardware. The schemes represent different points on a tradeoff curve between representational flexibility and computational simplicity.
Symmetric Quantization
Symmetric quantization takes a simpler approach by forcing the zero-point to be zero (), which centers the quantized range around zero. This constraint means that real zero always maps to integer zero, creating a symmetric mapping where positive and negative values are treated identically. The scale is determined by the largest absolute value in the tensor:
where:
- : the scale factor
- : the maximum absolute value present in the tensor
- : the maximum positive integer in the target range
This formula ensures that the largest magnitude value, whether positive or negative, maps to the edge of the integer range. The symmetry comes from the fact that would map to , using the same scale factor. The entire representation is anchored at zero, which mirrors the typical shape of neural network weight distributions.
For 8-bit signed integers with , a tensor with values in would use . This means:
- maps to
- maps to
- maps to
The main advantage of symmetric quantization is computational simplicity. With , the dequantization becomes a simple multiplication without any addition:
where:
- : the reconstructed real value
- : the scale factor
- : the quantized integer
This simplification significantly improves computational efficiency. Matrix multiplications can stay in integer arithmetic longer before converting back to floating point. When multiplying two symmetrically quantized matrices and with scales and , the entire integer multiplication proceeds first, then the combined scale is applied to the integer result. This allows modern hardware to use its specialized integer multiply-accumulate units for the bulk of the computation. Many hardware accelerators, including tensor cores on NVIDIA GPUs, are heavily optimized for this pattern.
However, symmetric quantization wastes representable range when the tensor distribution is asymmetric. In our example, values only range up to 0.3, but we allocated integer range up to 76, leaving integers 77 through 127 completely unused. Those 51 integer levels sit idle, representing precision budget that we paid for in bits but never used. The scale of 0.00394 is larger than it needs to be given the actual positive value range, meaning each step in the positive direction is coarser than necessary.
This wastefulness matters more at lower bit widths. With 8 bits, wasting 20% of the range costs little. With 4 bits, you have only 16 distinct positive values, and wasting a few is expensive. As we push toward 4-bit quantization, the precision budget becomes tight enough that asymmetric schemes often become worth their added complexity.
Asymmetric Quantization
Asymmetric quantization removes the symmetry constraint, letting a non-zero and thereby using the full integer range for the actual value distribution. This flexibility comes at the cost of additional complexity but provides better precision when the data distribution is skewed.
The key insight is that real zero does not need to map to integer zero. It only needs to map to some integer, and we store that integer (the zero-point) so we can undo the shift during dequantization. By letting this shift, we can slide the integer range to perfectly bracket the actual range of values, with no wasted integers at either end.
For our tensor with values in , using signed 8-bit integers (, ):
Now the mapping looks like:
- maps to
- maps to
- maps to
The full integer range from to is utilized. The scale is 0.00314 instead of 0.00394, meaning each integer step represents a smaller real interval and quantization error is reduced. This improvement of approximately 20% in scale translates directly to a 20% reduction in maximum quantization error per value. For a 4-bit system where every level counts, this 20% improvement can make the difference between acceptable and unacceptable accuracy loss.
The tradeoff is computational complexity. The zero-point must be tracked and subtracted during dequantization or matrix operations. For a simple dequantization, this adds just one subtraction. But for matrix multiplications, the situation is more complex. When multiplying two asymmetrically quantized matrices with values and , the product expands into cross-terms involving the zero-points:
The term can be computed efficiently using integer hardware. The constant term can be precomputed. But the terms and require multiplying integers by constants, which adds overhead. This overhead does not affect the weight quantization use case as severely, since weights are constant and their zero-point contributions can be precomputed and cached. For activation quantization, the overhead is more significant.
Choosing Between Schemes
The choice depends on the tensor distribution and hardware constraints. This is not a one-size-fits-all decision, and different layers in the same model may benefit from different approaches.
Weights often follow approximately symmetric distributions centered near zero. Neural network initialization schemes like Xavier and Kaiming initialization deliberately produce symmetric distributions, and gradient descent tends to preserve this symmetry for weights in the interior of the network. Symmetric quantization works well here, with minimal range waste and lower computational overhead.
Activations frequently have asymmetric distributions, especially after ReLU (where all values are non-negative) or after certain normalization layers that shift the mean away from zero. Asymmetric quantization better captures these distributions by placing the integer range entirely over the non-zero region.
Hardware support varies considerably. Some accelerators handle symmetric quantization much more efficiently, which makes it preferable even when asymmetric would be theoretically better. The practical speedup depends on whether the hardware has efficient instructions for the zero-point operations. In environments where hardware only accelerates symmetric quantization, choosing asymmetric for its theoretical precision advantage may result in slower inference.
In practice, weight quantization commonly uses symmetric schemes while activation quantization uses asymmetric schemes, though this varies by implementation. The decision ultimately balances mathematical optimality against engineering constraints, and different deployment scenarios may favor different choices. When in doubt, benchmark both approaches on your target hardware with your target model before committing.


Per-Tensor vs Per-Channel Quantization
Beyond the symmetric/asymmetric choice, we must decide the granularity at which to compute quantization parameters. Should a single scale cover an entire weight matrix, or should different parts of the matrix have different scales? This decision affects both quantization accuracy and storage overhead. It also touches on a basic property of neural networks: different parts of a model have very different weight distributions, and treating them identically is often suboptimal.
Think of it like calibrating a thermometer for use in different environments. A single calibration might work reasonably well across the range from arctic to tropical temperatures. But if you need maximum accuracy in both environments separately, you calibrate one thermometer for cold conditions and another for hot conditions. The cost is two calibrations instead of one, but the accuracy improvement can be dramatic at the extremes. Quantization granularity works the same way.
Per-Tensor Quantization
The simplest approach computes one scale (and possibly one zero-point) for an entire tensor. All elements share the same quantization mapping, which means every value in the tensor is quantized using identical parameters regardless of where it appears in the matrix.
This is memory-efficient: a weight matrix with millions of elements needs only one or two extra parameters for quantization metadata. The overhead is essentially zero, making per-tensor quantization attractive for memory-constrained environments. However, if different regions of the tensor have very different value ranges, per-tensor quantization performs poorly. Outliers force a large scale, reducing precision for the majority of values.
Consider a weight matrix where most values fall in but a few outliers reach . Per-tensor symmetric 8-bit quantization uses . Values of 0.05 quantize to and dequantize to approximately 0.047. The relative error is 6%. Meanwhile, the outlier at 1.0 maps perfectly to 127 with negligible error. The common-case values suffer large relative errors to accommodate rare outliers.
This phenomenon illustrates a basic tradeoff in quantization: we must balance precision across the entire value distribution. When a single scale must serve all values, extreme values dictate the scale, and typical values pay the precision penalty. The few outliers hold the precision of the entire tensor hostage. This problem becomes severe at 4-bit widths, where the precision budget is so tight that misallocating it to accommodate outliers can cause perceptible quality degradation.
Per-Channel Quantization
Per-channel quantization addresses this limitation by computing separate scales for each output channel of a weight tensor. For a linear layer with weight matrix , where is the output dimension and is the input dimension, we compute different scales, one per row.
This allows each output neuron's weights to use the full integer range based on that neuron's specific value distribution. If neuron 0 has weights in and neuron 1 has weights in , they get scales of approximately 0.00079 and 0.00787 respectively. Each achieves good precision for its own distribution, without being penalized by the other's characteristics. The scale mismatch that causes problems in per-tensor quantization disappears when each channel has its own scale.
The cost is storing scale values instead of one. For typical transformer dimensions, this is negligible: a 4096-dimensional layer adds 4096 floats (16 KB at FP16) to store scales, while the weight matrix itself contains elements (8 MB at 4-bit). The overhead is well under 1%. This small cost buys significant accuracy improvements, particularly for weights with heterogeneous distributions across channels.
Per-channel quantization also integrates cleanly with existing hardware kernels. When dequantizing before or after a matrix multiplication, we can apply the per-channel scales as a vector scaling operation, which most hardware supports efficiently. The scale for each output channel is applied to the corresponding row of the weight matrix during dequantization, requiring one scaling operation per row rather than one per element.
Per-Group Quantization
An intermediate approach divides each channel into groups of fixed size (commonly 32, 64, or 128 elements) and computes scales per group. This handles within-channel variation while keeping overhead manageable. The group size represents a tunable parameter that balances precision against storage.
The motivation for per-group quantization becomes clear when you examine how modern LLM weight matrices distribute their values. Within a single output channel, values are not uniformly distributed: some sections of the input dimension may correspond to certain semantic features, and the weights attending to those features may have systematically different magnitudes. A single scale per channel treats all sections identically, potentially misallocating precision within the channel.
For a matrix with group size 128, we need scale values. At FP16, this adds 262 KB of overhead, roughly 3% of the 4-bit weight size. The finer granularity better handles within-channel distribution variations without significantly inflating storage.
Per-group quantization has become standard in aggressive 4-bit quantization schemes like GPTQ and AWQ, which we will cover in upcoming chapters. The additional precision from per-group scales often makes the difference between acceptable and unacceptable accuracy loss when pushing to very low bit widths. When you have only 16 possible values to represent a continuous distribution, every bit of precision matters, and per-group quantization provides a reliable way to recover precision lost to value distribution variations within a channel.
Worked Example: Computing Scale and Zero-Point
Let us work through a concrete numerical example to make the formulas tangible. We will quantize a small tensor using both symmetric and asymmetric approaches, computing every intermediate value by hand.
Suppose we have the following weight tensor (7 values for clarity):
We want to represent this using 4-bit signed integers, which range from to (giving , ). This tiny tensor will show the mechanics clearly.
Step 1: Find the range.
Step 2a: Compute the symmetric scale.
The symmetric approach uses the maximum absolute value:
Step 2b: Compute the asymmetric scale and zero-point.
Step 3: Quantize each value.
For symmetric ():
| Error | ||||
|---|---|---|---|---|
| -0.8 | -6.22 | -6 | -0.771 | 0.029 |
| -0.4 | -3.11 | -3 | -0.386 | 0.014 |
| -0.1 | -0.78 | -1 | -0.129 | 0.029 |
| 0.0 | 0.00 | 0 | 0.000 | 0.000 |
| 0.2 | 1.56 | 2 | 0.257 | 0.057 |
| 0.5 | 3.89 | 4 | 0.514 | 0.014 |
| 0.9 | 7.00 | 7 | 0.900 | 0.000 |
The maximum error occurs at 0.2, which maps to integer 2 but dequantizes to 0.257, an error of 0.057. This error arises because the symmetric scale is set by the extreme value of 0.9, which coarsens the step size throughout the tensor.
For asymmetric (, ):
| Error | ||||
|---|---|---|---|---|
| -0.8 | -7.06 + (-1) = -8.06 | -8 | -0.793 | 0.007 |
| -0.4 | -3.53 + (-1) = -4.53 | -5 | -0.453 | 0.053 |
| -0.1 | -0.88 + (-1) = -1.88 | -2 | -0.113 | 0.013 |
| 0.0 | 0.00 + (-1) = -1.00 | -1 | 0.000 | 0.000 |
| 0.2 | 1.77 + (-1) = 0.77 | 1 | 0.227 | 0.027 |
| 0.5 | 4.41 + (-1) = 3.41 | 3 | 0.453 | 0.047 |
| 0.9 | 7.94 + (-1) = 6.94 | 7 | 0.793 | 0.107 |
The dequantization here uses .
The key insight from this comparison is not that one scheme is uniformly better. The symmetric scheme perfectly represents 0.9 (the anchor of its scale) with zero error, while asymmetric cannot. The asymmetric scheme better represents -0.8 (small error) but assigns its maximum error to 0.9. The average error depends on the full distribution, and asymmetric generally wins when the distribution is not centered at zero. For this tensor, which has a slightly positive skew, asymmetric quantization gives smaller average error for most values.
Step 4: Verify dequantization.
For symmetric, dequantization is . For asymmetric, dequantization is .
At (asymmetric): . Compare to the original : error of 0.007. The asymmetric mapping distributes its integer range more efficiently across the actual value range, which is why its error for typical values is smaller.
Calibration
The process of determining optimal quantization parameters (scale, zero-point, or clipping thresholds) by analyzing the actual value distributions in a model, typically using representative input data.
We have assumed we know the minimum and maximum values of each tensor, but finding these optimally requires care. Simply taking the observed min/max may not be ideal, especially for activations that vary with input data. Calibration is the process of determining these parameters systematically, and the quality of calibration directly impacts quantized model accuracy. Poor calibration can turn a theoretically excellent quantization scheme into a practical failure.
Think of calibration as tuning an instrument before a performance. You could tune once using the manufacturer's default settings, but a skilled musician tunes to the specific acoustic conditions of the hall and the other instruments they will play with. Similarly, a quantization scheme calibrated on representative data for the target deployment scenario will outperform one that uses generic or theoretical parameter choices.
Weight Calibration
Weights are fixed after training, so their distributions are known exactly. For per-tensor quantization, we can directly compute global min/max. For per-channel or per-group quantization, we compute statistics for each partition. Since weights do not change during inference, calibration happens once and the parameters are stored with the model, adding negligible size overhead.
However, raw min/max may be skewed by rare extreme values. Some methods clip the range to exclude extreme outliers, accepting large errors on those values in exchange for better precision on typical values. The reasoning is that if only 0.1% of values are extreme, accepting large errors on those values while improving precision for the other 99.9% reduces overall error. Common approaches include:
- Percentile clipping: Use the -th percentile instead of the true maximum, sacrificing outlier accuracy for finer granularity on typical values. Common choices are the 99th or 99.9th percentile.
- MSE minimization: Search for the clipping threshold that minimizes mean squared quantization error over the entire tensor. This is optimal in the mean-squared-error sense but requires a search procedure.
- Entropy-based: Choose thresholds that minimize KL divergence between the original and quantized value distributions. This preserves the statistical character of the tensor rather than minimizing a single error metric.
The right choice depends on your application. For tasks sensitive to rare values (factual recall, where the model needs to retrieve specific numbers), clipping extreme values aggressively may hurt. For tasks dominated by typical values (general text generation), aggressive clipping often helps by improving precision for the common case.
Activation Calibration
Activations present a harder problem because their distributions depend on inputs. A model might produce activations ranging from to on one input and to on another. We need to find quantization parameters that work well across the expected input distribution, not just for any single input. Static activation quantization requires making a decision once that must generalize to all future inputs.
The standard approach runs the model on a calibration dataset: a small representative sample of inputs from the target domain. We collect activation statistics (min, max, percentiles, or full histograms) across this dataset, then compute quantization parameters from the aggregated statistics. The calibration dataset is a proxy for the true inference distribution, and its quality directly determines calibration quality.
Common aggregation strategies include:
- Running min/max: Track the most extreme values seen across all calibration samples. This ensures no calibration input will be clipped, but may produce overly large scales that waste precision on typical inputs.
- Moving average: Smooth the statistics across samples to reduce sensitivity to outliers in the calibration set itself. A commonly used formula is , with .
- Histogram-based: Build full histograms and choose optimal thresholds by minimizing reconstruction error. This is the most accurate but also the most expensive approach.
The calibration dataset should resemble actual inference inputs. It should cover similar content and sequence lengths while matching the expected writing style. Using the wrong distribution during calibration can lead to systematic clipping or range waste during deployment. A chatbot calibrated on encyclopedia text may have poorly tuned activations for conversational inputs. A few hundred to a few thousand samples typically suffice for stable statistics, though the exact number depends on how variable the activations are and how sensitive the model is to quantization.
Dynamic vs Static Quantization
Calibration determines whether quantization is static or dynamic, and this distinction has major practical implications for inference systems.
Static quantization fixes all quantization parameters at calibration time. Inference uses predetermined scales and zero-points, with no runtime overhead for parameter computation. This is efficient and predictable but may lose accuracy if inference distributions differ from calibration. Static quantization is the default for production deployment: calibrate once, deploy everywhere.
Dynamic quantization computes activation quantization parameters on-the-fly for each input. This adapts to the actual values in each tensor. This provides optimal parameters regardless of the input distribution. The cost is computing statistics (at minimum, a min/max scan) over each activation tensor during inference, adding latency and overhead. Dynamic quantization is commonly used when activation distributions are highly variable, when calibration data is unavailable, or when maximum accuracy is required with no deployment pipeline for offline calibration.
Weight quantization is almost always static since weights do not change. The question is whether to quantize weights permanently (storing integers) or to quantize them on-the-fly from stored floating-point weights (saving memory but doing the conversion at runtime). Permanent weight quantization is more common for deployed models. Activation quantization may be either static or dynamic, depending on the accuracy-efficiency tradeoff desired.
A third approach, quantization-aware training (QAT), trains the model while simulating quantization effects, letting the model to adapt its weights to minimize quantized error. QAT typically outperforms post-training quantization but requires access to training data and significant compute. We will cover QAT in detail in a later chapter.
Implementation
Let us implement the core quantization operations to solidify these concepts. We will start with basic symmetric quantization, then extend to asymmetric and per-channel variants. Working through concrete code helps build intuition for how the mathematical formulas translate to practice, and reveals the engineering choices that pure mathematical treatment leaves implicit.
import torch
# Create a sample weight tensor with a realistic distribution. The shared setup
# seed keeps every render variant deterministic.
weights = torch.randn(64, 128) * 0.02 # Typical initialization scaleWeight tensor shape: torch.Size([64, 128]) Value range: [-0.0767, 0.0689] Mean: 0.000130, Std: 0.0200
The weights follow a roughly symmetric distribution centered near zero, typical of neural network parameters after initialization or training. This distribution is well-suited to symmetric quantization, which we will implement first. The standard deviation of approximately 0.02 matches common initialization scales, and the range is several standard deviations wide, with occasional outliers from the tails of the Gaussian.

Symmetric Per-Tensor Quantization
The symmetric per-tensor implementation demonstrates the simplest form of quantization. We compute a single scale from the maximum absolute value, then apply the same quantization formula to every element in the tensor. The implementation is only a few lines of code, which reflects the mathematical simplicity of the symmetric scheme.
def quantize_symmetric(tensor, num_bits=8):
"""Symmetric per-tensor quantization."""
qmax = 2 ** (num_bits - 1) - 1 # For signed: 127 for 8-bit
# Compute scale from maximum absolute value
max_abs = tensor.abs().max()
scale = max_abs / qmax
# Quantize: round(x / scale)
q_tensor = torch.round(tensor / scale).clamp(-qmax - 1, qmax)
q_tensor = q_tensor.to(torch.int8)
return q_tensor, scale
def dequantize_symmetric(q_tensor, scale):
"""Dequantize symmetrically quantized tensor."""
return q_tensor.float() * scale# Quantize and dequantize
q_weights, scale = quantize_symmetric(weights, num_bits=8)
reconstructed = dequantize_symmetric(q_weights, scale)
# Compute quantization error
error = (weights - reconstructed).abs()Scale: 0.000604 Quantized values range: [-127, 114] Mean absolute error: 0.000149 Max absolute error: 0.000302 Relative error (vs std): 0.7448%
The mean absolute error is small compared to the weight standard deviation. 8-bit symmetric quantization preserves weights well when the distribution is roughly symmetric. The relative error of less than 1% indicates that the quantized weights closely approximate the originals, which bodes well for maintaining model accuracy. The scale parameter encodes exactly how much real-valued range each integer step represents, and comparing it to the weight standard deviation gives a sense of quantization resolution relative to the spread of the data.
Asymmetric Per-Tensor Quantization
The asymmetric implementation adds complexity by computing and tracking a zero-point. This extra parameter allows better use of the integer range when values are not centered at zero. The implementation shows how the zero-point changes both quantization and dequantization, and how careful handling is needed to avoid edge cases like division by zero for constant tensors.
def quantize_asymmetric(tensor, num_bits=8):
"""Asymmetric per-tensor quantization."""
qmin, qmax = 0, 2**num_bits - 1 # Unsigned: 0 to 255
# Compute scale and zero-point
x_min, x_max = tensor.min(), tensor.max()
scale = (x_max - x_min) / (qmax - qmin)
# Avoid division by zero for constant tensors
if scale == 0:
scale = torch.tensor(1.0)
zero_point = qmin - torch.round(x_min / scale)
zero_point = zero_point.clamp(qmin, qmax)
# Quantize
q_tensor = torch.round(tensor / scale) + zero_point
q_tensor = q_tensor.clamp(qmin, qmax).to(torch.uint8)
return q_tensor, scale, zero_point
def dequantize_asymmetric(q_tensor, scale, zero_point):
"""Dequantize asymmetrically quantized tensor."""
return (q_tensor.float() - zero_point) * scale# Quantize asymmetrically
q_weights_asym, scale_asym, zp = quantize_asymmetric(weights, num_bits=8)
reconstructed_asym = dequantize_asymmetric(q_weights_asym, scale_asym, zp)
error_asym = (weights - reconstructed_asym).abs()Scale: 0.000571 Zero-point: 134 Quantized values range: [0, 255] Mean absolute error: 0.000143 Relative error (vs std): 0.7129%
For this symmetric weight distribution, asymmetric quantization provides similar accuracy to symmetric. The zero-point is near 128, roughly the middle of the unsigned range, as expected for zero-centered data. When the data is already symmetric, asymmetric quantization offers little benefit but also causes no harm. The computation confirms our theoretical expectation: for symmetric data, both schemes use the integer range similarly and produce comparable errors.
Per-Channel Quantization
Per-channel quantization computes separate scales for each row of the weight matrix, letting each output channel to use its full precision budget independently. The implementation uses PyTorch's broadcasting-aware operations to handle all channels simultaneously.
def quantize_per_channel(tensor, num_bits=8, axis=0):
"""Symmetric per-channel quantization along specified axis."""
qmax = 2 ** (num_bits - 1) - 1
# Compute scale per channel (reduce across the non-channel dimensions)
# For 2D matrices: if axis=0 (rows), reduce dim 1. If axis=1 (cols), reduce dim 0.
reduce_dim = 1 if axis == 0 else 0
max_abs = tensor.abs().amax(dim=reduce_dim, keepdim=True)
scale = max_abs / qmax
scale = torch.where(scale == 0, torch.ones_like(scale), scale)
# Quantize
q_tensor = torch.round(tensor / scale).clamp(-qmax - 1, qmax)
q_tensor = q_tensor.to(torch.int8)
return q_tensor, scale
def dequantize_per_channel(q_tensor, scale, axis=0):
"""Dequantize per-channel quantized tensor."""
# Scale already has correct broadcasting shape from quantize
return q_tensor.float() * scaleTo illustrate the benefits of per-channel quantization, we create a tensor where different channels have dramatically different value ranges. This simulates the real-world situation where some neurons develop larger weights than others during training. In practice, weight magnitude heterogeneity is common and can be severe: some neurons act as feature detectors for rare but important patterns and develop large weights, while others handle common patterns with smaller weights.
# Create a tensor with varying scales per channel
weights_varied = weights.clone()
weights_varied[0] *= 10 # First channel has larger values
weights_varied[1] *= 0.1 # Second channel has smaller values
# Compare per-tensor vs per-channel
q_pt, s_pt = quantize_symmetric(weights_varied)
recon_pt = dequantize_symmetric(q_pt, s_pt)
q_pc, s_pc = quantize_per_channel(weights_varied, axis=0)
recon_pc = dequantize_per_channel(q_pc, s_pc, axis=0)
# Compute quantization errors
error_pt = (weights_varied - recon_pt).abs()
error_pc = (weights_varied - recon_pc).abs()Per-tensor quantization: Single scale: 0.003952 Channel 0 error (large weights): 0.001089 Channel 1 error (small weights): 0.000978 Per-channel quantization: Channel 0 scale: 0.003952 Channel 1 scale: 0.000039 Channel 0 error: 0.001089 Channel 1 error: 0.000010
Per-channel quantization substantially reduces error for channel 1. With per-tensor quantization, the large channel forces a big scale, causing the small channel to quantize coarsely. Per-channel quantization gives each channel an appropriate scale tailored to its own value range. The error reduction for channel 1 is several orders of magnitude. This shows concretely why per-channel quantization is needed for high-accuracy quantization when weight distributions vary across channels.


Visualizing Quantization Error
Visualization helps build intuition about how quantization errors distribute across values. The following plots show both the error distribution and the relationship between original and reconstructed values.
import matplotlib.pyplot as plt
use_book_style()
plt.rcParams["figure.figsize"] = (6.0, 4.4)
plt.figure()
ax = plt.gca()
ax.hist(error.flatten().numpy(), bins=50, edgecolor=PALETTE["ink"], alpha=0.7)
ax.axvline(
scale.item() / 2,
color=theme_color("red"),
linestyle="--",
label=f"Scale/2 = {scale.item() / 2:.5f}",
)
ax.set_xlabel("Absolute Error")
ax.set_ylabel("Frequency")
ax.set_title("Quantization Error Distribution")
ax.legend(
loc="upper center",
bbox_to_anchor=(0.5, -0.16),
borderaxespad=0,
)
polish_axes(ax)
plt.show()
import matplotlib.pyplot as plt
import numpy as np
use_book_style()
plt.rcParams["figure.figsize"] = (6.0, 4.0)
plt.figure()
sample_idx = np.random.choice(weights.numel(), 1000)
orig_sample = weights.flatten()[sample_idx].numpy()
recon_sample = reconstructed.flatten()[sample_idx].numpy()
plt.scatter(orig_sample, recon_sample, alpha=0.5, s=10)
plt.plot(
[orig_sample.min(), orig_sample.max()],
[orig_sample.min(), orig_sample.max()],
color=PALETTE["coral"],
linestyle="--",
label="Perfect reconstruction",
)
plt.xlabel("Original Weight")
plt.ylabel("Reconstructed Weight")
plt.title("Original vs Reconstructed Values")
plt.legend()
polish_axes(plt.gca())
plt.show()
The error histogram shows that quantization errors are bounded and mostly small. The maximum possible error is half the scale (when a value falls exactly between two quantization levels), and the red dashed line marks this theoretical bound. The scatter plot confirms tight correlation between original and reconstructed values, with points clustering closely around the diagonal line that represents perfect reconstruction. The slight discretization visible in the reconstructed values (they can only take values that are multiples of the scale) creates the characteristic "staircase" pattern visible in this plot.
Simulating Calibration
Calibration in practice involves running the model on representative data to collect statistics. The following code simulates this process for a simple model. This shows both minmax and percentile-based calibration approaches.
def calibrate_activations(
model_fn, calibration_data, num_bits=8, method="minmax"
):
"""
Calibrate quantization parameters for activations.
model_fn: Function that takes input and returns activations
calibration_data: List of input tensors
method: 'minmax' or 'percentile'
"""
all_activations = []
# Collect activations across calibration set
for data in calibration_data:
activations = model_fn(data)
all_activations.append(activations)
all_activations = torch.cat(all_activations, dim=0)
if method == "minmax":
x_min = all_activations.min()
x_max = all_activations.max()
elif method == "percentile":
# Use 99.9th percentile to exclude outliers
x_min = torch.quantile(all_activations.float(), 0.001)
x_max = torch.quantile(all_activations.float(), 0.999)
# Compute asymmetric quantization parameters
qmin, qmax = 0, 2**num_bits - 1
scale = (x_max - x_min) / (qmax - qmin)
zero_point = qmin - torch.round(x_min / scale)
return scale, zero_point, x_min, x_max# Define fixed weights for the mock model
input_dim = 128
hidden_dim = 64
W = torch.randn(input_dim, hidden_dim) * 0.1
b = torch.randn(hidden_dim) * 0.01
# Simulate a simple activation function
def mock_model(x):
# Simulates ReLU-like activations (mostly positive with some outliers)
# Using fixed weights ensures consistent behavior across calls
return torch.relu(x @ W + b)
# Generate calibration data
calibration_data = [torch.randn(32, 128) for _ in range(10)]
# Calibrate with both methods
scale_mm, zp_mm, min_mm, max_mm = calibrate_activations(
mock_model, calibration_data, method="minmax"
)
scale_pct, zp_pct, min_pct, max_pct = calibrate_activations(
mock_model, calibration_data, method="percentile"
)MinMax Calibration: Range: [0.0000, 4.1836] Scale: 0.016406 Percentile Calibration (99.9%): Range: [0.0000, 3.4641] Scale: 0.013585 Scale reduction from percentile clipping: 17.2%
Percentile-based calibration produces a smaller scale. This provides finer quantization resolution for the bulk of activations at the cost of clipping rare outliers. The scale reduction directly translates to smaller quantization errors for typical values, which usually improves overall model accuracy. The trade is explicit: we accept large errors for the 0.1% most extreme activation values in exchange for smaller errors for the other 99.9%. Whether this is a good trade depends on how much those extreme values matter for your application.

Comparing Bit Widths
The following experiment measures how quantization error changes with bit width. This shows the exponential relationship between precision and error. This relationship is basic to understanding why moving from 8-bit to 4-bit quantization is much more challenging than moving from 32-bit to 16-bit.
def measure_quantization_error(tensor, num_bits):
"""Measure mean squared error for given bit width."""
q, s = quantize_symmetric(tensor, num_bits)
recon = dequantize_symmetric(q, s)
mse = ((tensor - recon) ** 2).mean()
return mse.item()
# Test different bit widths
bit_widths = [2, 3, 4, 5, 6, 7, 8]
errors = [measure_quantization_error(weights, b) for b in bit_widths]import matplotlib.pyplot as plt
use_book_style()
plt.rcParams["figure.figsize"] = (6.0, 4.0)
plt.figure()
ax = plt.gca()
ax.plot(bit_widths, errors, "o-", markersize=5.5, linewidth=1.1)
use_book_log_scale(ax, axis="y")
plt.xlabel("Bit Width")
plt.ylabel("Mean Squared Error (log scale)")
plt.title("Quantization Error vs Bit Width")
plt.xticks(bit_widths)
polish_axes(ax)
# Add annotations
for bw, err in zip(bit_widths, errors):
plt.annotate(
f"{err:.2e}",
(bw, err),
textcoords="offset points",
xytext=(0, 10),
ha="center",
fontsize=8,
color=PALETTE["ink"],
bbox={
"boxstyle": "round,pad=0.18",
"facecolor": PALETTE["paper"],
"edgecolor": "none",
"alpha": 0.92,
},
)
plt.show()
Error drops roughly by a factor of 4 for each additional bit. This follows from quantization theory: each extra bit doubles the number of representable levels, halving the quantization step size, which reduces mean squared error by a factor of 4. On the log-scale plot, this appears as a straight line with a slope of 4x per bit. This 4x-per-bit rule is a useful mental model for reasoning about quantization tradeoffs: the difference between 8-bit and 4-bit quantization in error terms is approximately , a factor of 256 in mean squared error. This is a large change, which explains why aggressive 4-bit quantization requires much more sophisticated techniques than simple 8-bit quantization.
Key Parameters
The key parameters for the quantization functions are:
- num_bits: The bit width for the target integer representation (e.g., 8 for INT8, 4 for INT4). Lower bits increase compression but reduce precision. The 4x-per-bit rule predicts the error cost of reducing bit width.
- axis: The dimension along which quantization parameters (scale, zero-point) are computed. For per-channel quantization of linear layers, this is typically the output dimension (axis 0), letting each output neuron to have its own scale.
- method: The strategy for determining quantization ranges during calibration. The choice between minmax and percentile-based calibration significantly affects accuracy when activation distributions have heavy tails.
Limitations and Practical Considerations
Quantization's appeal lies in its simplicity, but several challenges arise in practice that require careful handling. Understanding these limitations is needed for successfully deploying quantized models and for understanding why the more sophisticated methods covered in subsequent chapters were developed.
Outlier sensitivity remains the primary obstacle to aggressive quantization. Transformer models frequently develop weight and activation outliers during training, particularly in attention layers and certain projection matrices. Research on large transformers has found that some activation dimensions can take values an order of magnitude larger than the mean, and these outliers become more pronounced as model scale increases. A single extreme value in a tensor forces a large scale, wasting precision for all other values. Per-channel and per-group quantization mitigate this for weights, but handling activation outliers often requires more sophisticated techniques. Methods like LLM.int8() detect outliers at runtime and process them in higher precision, while AWQ and GPTQ use careful calibration to minimize outlier impact. The challenge is that these outliers are not random: they tend to concentrate in specific channels and follow patterns that can be exploited, but exploiting them requires analysis beyond simple min/max calibration.
Accuracy degradation varies dramatically across models and tasks. Simple tasks like classification often tolerate aggressive quantization with minimal accuracy loss, because the model needs only to distinguish between a small number of output classes and can absorb substantial weight perturbations. Generation tasks, especially those requiring precise reasoning or factual recall, prove more sensitive. A model that achieves acceptable perplexity after 4-bit quantization may show subtle but significant degradation in multi-step reasoning, mathematical calculations, or factual question answering. The difficulty is that standard benchmarks may not reveal this degradation. A model might score well on perplexity while generating subtly incorrect or less coherent text. Always evaluate on task-specific metrics, not just generic benchmarks, and use human evaluation for generation quality when possible.
Calibration dataset selection significantly impacts results, and this dependence is often underestimated. Using the wrong distribution during calibration can cause systematic clipping or range waste during deployment. For general-purpose models, calibration data should span the expected content and input lengths while reflecting the likely writing styles. For specialized applications, calibration on domain-specific data often yields better results than generic calibration. A model deployed for medical record analysis should be calibrated on medical text, not Wikipedia. The mismatch between calibration distribution and deployment distribution is a common source of unexpected accuracy loss that is difficult to diagnose after deployment.
Hardware support determines practical speedup, and this is where theoretical analysis often diverges from empirical results. Integer operations are only faster if the hardware has efficient integer compute units and the software stack (CUDA kernels, runtime libraries) is optimized to use them. Modern GPUs offer accelerated INT8 and sometimes INT4 operations through tensor cores, but not all hardware supports all quantization formats, and not all quantization schemes are supported by the available kernels. Some quantization schemes that theoretically reduce compute find no actual speedup due to missing hardware support or kernel availability. Before committing to a quantization approach for production deployment, verify actual inference latency on the target hardware. The theoretical memory reduction is guaranteed, but the compute speedup depends on the specific hardware and software stack.
Despite these challenges, quantization has enabled an explosion in LLM accessibility. Models that required enterprise-grade hardware now run on laptops and smartphones. The memory savings cascade into reduced serving costs, making AI applications economically viable at scales previously impossible. The next chapters will explore specific quantization methods, INT8 and INT4 techniques, and formats like GPTQ and AWQ that push compression further while maintaining accuracy. The foundational concepts in this chapter, scale, zero-point, granularity, and calibration, underpin all of those methods.
Summary
Quantization maps high-precision floating-point values to lower-precision integers, reducing memory footprint and letting faster computation for memory-bound inference workloads. The core insight is that LLM inference is primarily limited by memory bandwidth, not compute, and loading fewer bytes per parameter directly translates to faster token generation.
The basic operation involves two parameters: the scale , which determines the step size between adjacent quantization levels, and the zero-point , which shifts the integer range to align with the actual distribution of values. Together they define a linear mapping from real values to integers, with inverse for dequantization.
Symmetric quantization forces , simplifying computation but potentially wasting integer range when data is not zero-centered. Asymmetric quantization allows , using the full integer range for any distribution at the cost of additional arithmetic during dequantization. Weights are typically quantized symmetrically, while activations benefit from asymmetric schemes.
Quantization granularity ranges from per-tensor (one scale for all values) through per-channel (one scale per output dimension) to per-group (multiple scales within each channel). Finer granularity better handles distribution variation at the cost of additional scale storage, with per-channel adding negligible overhead while giving substantial accuracy improvements.
Calibration determines optimal quantization parameters by analyzing actual tensor values. Weight calibration is straightforward because weights are fixed. Activation calibration requires running on representative data and choosing aggregation strategies that balance outlier handling against precision for typical values.
The basic tradeoff in quantization is information loss versus compression. Fewer bits mean smaller models and faster inference but increased quantization error, with each additional bit reducing error by approximately 4x. Understanding this tradeoff, and the tools for managing it, prepares you for the specific quantization techniques in the chapters that follow.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about weight quantization fundamentals.
Weight Quantization Basics
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!