Part of Language AI Handbook
Covers INT8 weight quantization with absmax and smooth quantization techniques. Solve the outlier problem in large language models.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
INT8 Quantization
In the previous chapter on weight quantization basics, we explored why representing neural network parameters with fewer bits can sharply reduce memory requirements and accelerate inference. We saw how the core ideas of scale factors and zero-points let us map floating-point values onto discrete integer grids. We now turn to the most widely deployed quantization format in production systems: 8-bit integers (INT8). This precision point is a sweet spot that the machine learning community discovered through years of experimentation: memory savings are substantial (2x compression from FP16, 4x from FP32), specialized hardware acceleration is widely available, and accuracy degradation remains manageable for most applications. Understanding INT8 quantization in depth means understanding every major LLM inference library used today, from bitsandbytes to TensorRT-LLM.
INT8 quantization is about storing numbers more compactly and running them faster. Modern GPUs and specialized AI accelerators include dedicated INT8 matrix multiplication units that can reach 2-4x higher throughput than their floating-point counterparts. NVIDIA's Tensor Cores, for example, perform INT8 operations at twice the rate of FP16 operations. Quantizing both weights and activations to INT8 therefore reduces memory use and can accelerate computation. The improvement comes from two compounding effects: smaller data types require less memory bandwidth to load, and dedicated integer arithmetic units can perform more operations per clock cycle than their floating-point equivalents. On memory-bandwidth-limited tasks like large-batch LLM inference, these effects can substantially increase throughput.
The challenge is that this sounds easier than it is. Squeezing the continuous range of floating-point values into just 256 discrete integers inevitably introduces quantization error. Every value you store must round to one of those 256 representable points, and the difference between the original and rounded value accumulates across millions of operations. For small models and simple tasks, this noise is negligible. For large language models with billions of parameters and complex dependencies between layers, naive INT8 quantization can cause large accuracy degradation. The field had to develop principled techniques for minimizing this error before INT8 LLM inference became practical.
The journey to effective INT8 quantization for LLMs involved a surprising discovery: the problem was not the weights but the activations. Weights are static tensors we can analyze carefully offline, adjusting our quantization scheme to fit their distribution. Activations are dynamic, varying with every input, and in large language models they develop pathological outliers that resist standard quantization approaches. A small number of hidden dimensions produce values sharply larger than all other dimensions, which makes it impossible to choose a quantization scale that works well for both the outliers and the typical values. This observation, and the elegant mathematical fix for it, forms the intellectual core of this chapter.
We will explore the mathematics of range mapping, the practical absmax quantization scheme that forms the foundation, the granularity choices between per-tensor and per-channel quantization, the emergent outlier problem in large language models, and the smooth quantization technique that makes INT8 viable even for models with hundreds of billions of parameters. By the end you will understand how to apply these techniques, why each design choice was made, and what happens when things go wrong.
INT8 quantization for neural networks dates to the late 2010s, initially developed for convolutional neural networks in computer vision (Jacob et al., 2018, "Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference"). These early techniques worked well for vision models but struggled with transformers. The outlier problem in large language models was described systematically by Dettmers et al. in 2022 with LLM.int8(), which showed that beyond 6.7B parameters, a small fraction of activation channels exhibit emergent outlier features. The SmoothQuant technique (Xiao et al., 2022) then provided a cleaner solution by migrating quantization difficulty from activations to weights through an offline transformation. Together these papers established the modern approach to INT8 LLM inference that is now standard in production systems.
The INT8 Representation Space
An 8-bit signed integer can stand for values from -128 to 127, giving us exactly 256 distinct values to work with. Think of this as having 256 labeled slots on a number line, and every floating-point weight or activation value must be assigned to the nearest slot. When we quantize a floating-point tensor, we create a mapping (technically a codebook) that associates each of these 256 integers with a specific floating-point value. The integers become the compact representation we store and compute with, and the codebook lets us convert back to floating-point whenever we need the full precision.
To understand this mapping intuitively, imagine you have a number line representing all possible floating-point values a weight might take. Quantization places 256 evenly-spaced "buckets" along this line, and every floating-point value must be assigned to its nearest bucket. The original value is then represented by that bucket's integer label, and reconstruction involves converting the label back to the bucket's center position. The spacing between these buckets, and where they are positioned along the number line, completely determines the quality of our approximation. Place them well and every value lands close to its bucket center. Place them poorly and some values are forced into buckets far from where they belong.
The set of representable values after quantization forms a uniform grid in the original floating-point space. The spacing between adjacent grid points is called the quantization step size or scale factor. Every floating-point value must round to the nearest grid point, and the maximum possible error is half the step size.
The basic challenge is choosing how to position this grid. Consider a weight tensor with values ranging from -0.5 to 1.2. We need to decide what floating-point value should -128 stand for, what floating-point value should 127 stand for, and how to handle values outside our chosen range (clipping them to the boundary). These choices define our quantization scheme and directly impact the reconstruction error. A poorly chosen grid might waste many of its 256 integers representing values that never occur in the tensor, while cramming the observed values into too few buckets. The goal is to align our quantization grid as closely as possible with the distribution of values we need to stand for.
The key insight is that the quality of quantization depends almost entirely on how well the grid matches the distribution of values in the tensor. If our weights happen to be uniformly distributed between -1 and 1, a grid spanning that range is nearly optimal. But if most weights cluster near zero with a long tail of larger values, we need to make a choice: use a wide grid that accommodates the tail but has poor resolution near zero, or use a narrow grid that captures the bulk well but clips the tail. Neither choice is free, and the right answer depends on which values matter more for the model's accuracy.

Symmetric Quantization with Absmax
The simplest and most common approach to INT8 quantization is symmetric quantization, where we center the quantization grid at zero. The "absmax" method finds the maximum absolute value in the tensor and maps it to the largest representable integer. This approach derives its name from the key statistic it computes: the absolute maximum, which determines the extent of our quantization range in both positive and negative directions simultaneously.
Think of absmax quantization as stretching or compressing a rubber band. Your floating-point values are marked on the band, and you resize it until the most extreme mark lines up with the INT8 boundary at -128 or 127. The scale factor tells you how much you stretched or compressed. Every other value gets rescaled by the same factor, preserving relative distances. The reconstruction step is just reversing this scaling operation. The elegance of this approach is that a single number, the scale factor, captures everything you need to know about the quantization.
The intuition behind absmax quantization is straightforward. We want to ensure that no value in our tensor falls outside the representable range, which would force us to clip it and introduce potentially large errors. By finding the largest magnitude value in the tensor and designing our grid to just barely accommodate it, we guarantee that every original value can be represented without clipping. At the same time, we make the grid as fine as possible given this constraint, since using a larger range than necessary would waste precision on integers that never get used.
Given a tensor with floating-point values, absmax quantization works as follows. First we compute the scale factor from the tensor's extreme value, then we apply rounding to obtain integer codes:
where:
- : the input floating-point tensor we want to quantize
- : the maximum absolute value in the input tensor, called the "absmax"
- : the scale factor that determines how much floating-point distance each integer step is
- : the maximum representable value in a signed 8-bit integer (we reserve -128 as an "overflow" value and map the positive side to )
- : the resulting tensor of quantized 8-bit integers, each in
Let's trace through this process step by step. First, we scan the entire tensor to find , the largest magnitude value regardless of sign. This becomes our reference point for setting the scale. Next, we compute the scale factor by dividing by 127, which tells us how much "floating-point distance" each integer step is. If , then , meaning each integer step covers about 0.005 in floating-point space. Finally, we quantize each value by dividing it by the scale factor and rounding to the nearest integer. This rounding step is where the actual information loss occurs, as we snap each continuous value to its nearest discrete representation.
To recover an approximation of the original values, we simply multiply by the scale:
where:
- : the reconstructed floating-point tensor, an approximation of the original
- : the scale factor used during quantization (must be stored alongside the quantized tensor)
- : the quantized integer tensor we stored
The hat notation indicates this is a reconstruction, not the exact original values. The difference is the quantization error. This error arises because the rounding operation discards fractional information. When we divided by and rounded, we lost the remainder, and multiplying the rounded result back by cannot restore what was lost. The beauty of this scheme is that the error is bounded: no reconstructed value can differ from its original by more than half a scale step, or . For a tensor with absmax of 1.0, the scale is , so the maximum error is about 0.004. This is roughly 0.4% of the full range, which is small enough to be acceptable for most neural network weights.
Why Symmetric Around Zero?
Centering the grid at zero has a important computational advantage: the integer zero maps exactly to floating-point zero. This property might seem like a minor mathematical nicety, but it has significant practical implications for neural network inference.
This matters because many neural network activations are zero (from ReLU, dropout, and padding), zero-valued weights do not contribute to computations, and preserving exact zeros avoids accumulating unnecessary rounding errors. Consider what happens during a matrix multiplication when many elements are zero. If zero maps exactly to zero, these elements contribute nothing to the output, exactly as they should. But if zero were represented by some non-zero integer due to an offset, every "zero" element would contribute a small spurious value to the computation. Over millions of operations, these spurious contributions could accumulate into real errors.
Symmetric quantization also simplifies the mathematics during matrix multiplication. When computing with quantized values, we only need to track scale factors, not additional offset terms. This simplification translates directly into faster inference, since the dequantization step requires only a single multiplication rather than a multiplication followed by an addition. When both scales are scalar constants (as they are for per-tensor quantization), this final rescaling is in effect free compared to the matrix multiplication itself.
The Clipping Trade-off
Absmax quantization clips all values to the range . If is dominated by a few outlier values, the quantization grid becomes coarse for the majority of values. This creates a basic tension in quantization design: we want to accommodate outliers to avoid clipping errors, but we also want a fine grid to minimize rounding errors for typical values.
Consider a tensor where 99% of values lie between -0.1 and 0.1, but one outlier is 5.0. Using absmax:
This scale means values around 0.1 get quantized to just . The reconstruction is , introducing significant relative error for these common values. To understand the magnitude of this problem, note that a typical value of 0.1 experiences an 18% relative error, while the outlier value of 5.0 would be reconstructed nearly perfectly. We have sacrificed precision where it matters most, on the bulk of our values, to perfectly preserve a single outlier.
One solution is to clip outliers: choose smaller than the true maximum, accepting that extreme values will be clipped. This trades outlier accuracy for better precision on the bulk of values. For example, if we set instead of 5.0, our scale becomes , and a value of 0.1 now quantizes to integer 25 with reconstruction , a much smaller relative error. The outlier would be clipped to 127, reconstructing to 0.5 instead of 5.0, but this might be an acceptable trade-off depending on the model's sensitivity to those extreme values.
Finding the optimal clipping threshold is the basis for calibration-based quantization methods. These methods analyze the value distribution across representative data to find the clipping point that minimizes total reconstruction error, balancing the clipping errors from outliers against the rounding errors for typical values. Calibration is often as important as the quantization scheme itself, and models that fail to quantize well with default settings often succeed after proper calibration.


Asymmetric Quantization
When tensor values are not centered around zero, symmetric quantization wastes representable range. ReLU activations, for example, are always non-negative. Mapping the range symmetrically would waste half our integers representing negative values that never occur. This inefficiency becomes severe when the actual data distribution is heavily skewed or shifted away from zero. With 256 integers available and only half the range needed, symmetric quantization effectively reduces your precision from 8 bits to 7 bits for these tensors.
Asymmetric quantization solves this by introducing a zero-point offset that shifts the quantization grid to the data's range. Think of it as sliding our 256-slot number line left or right until it lines up with the data distribution, rather than always keeping it centered at zero. The grid spacing remains uniform, but we can position it anywhere. This flexibility comes at a cost: more complex arithmetic during inference, as we now have two parameters (scale and zero-point) to track rather than one.
The asymmetric quantization formulas show how to derive both parameters from the data range:
where:
- : the input floating-point tensor
- : the scale factor, computed from the full data range (minimum to maximum)
- : the number of intervals in an 8-bit range ( intervals for 256 values)
- : the zero-point integer that shifts the grid to align with the data distribution
- : the resulting quantized tensor with values in
The key difference from symmetric quantization lies in how we compute the scale and the introduction of the zero-point. Instead of basing the scale on the maximum absolute value, we use the full range from minimum to maximum. This keeps all 256 integers can potentially be used. The zero-point is an integer that shifts the quantization grid so that the minimum value maps to -128 and the maximum to 127. This shift allows us to "slide" our quantization grid along the number line to wherever the actual data lies.
Dequantization reverses this process by first undoing the zero-point shift and then applying the scale:
where:
- : the reconstructed floating-point values
- : the scale factor
- : the quantized integer values
- : the zero-point offset, which is subtracted before scaling to recover the original range
While asymmetric quantization can stand for arbitrary ranges more efficiently, it complicates the mathematics during inference. Matrix multiplications now involve additional terms from the zero-point offsets. When we expand the matrix multiplication with asymmetric quantization, cross-terms involving the zero-points appear, requiring additional computation and memory access. For weights, which are typically centered around zero due to common initialization and regularization practices, symmetric quantization remains the standard choice. The computational overhead of asymmetric quantization is generally reserved for activations where the efficiency gains from better range utilization outweigh the additional complexity.


Per-Tensor vs Per-Channel Quantization
The granularity at which we compute scale factors materially impacts quantization quality. This design choice is a trade-off between simplicity and accuracy, with different granularities being appropriate for different situations. Understanding the trade-off is needed for deploying quantized models effectively, because a choice that works perfectly for one layer might fail badly for another.
Per-tensor quantization uses a single scale factor for an entire weight matrix or activation tensor. Think of it as fitting one set of clothing to an entire crowd: some people will be well-fitted, but anyone materially different from the average will have a poor fit. This approach is simple and introduces minimal overhead, but a single outlier anywhere in the tensor degrades precision everywhere. Imagine a weight matrix where one row has unusually large values: that single row would force a large scale factor that reduces precision for all other rows, potentially turning what should be fine-grained distinctions into coarse approximations.
Per-channel quantization computes separate scale factors for each output channel of a weight matrix. This is like fitting each person with their own set of clothing: everyone gets a good fit, at the cost of measuring each person individually. For a linear layer with weight matrix , we compute different scale factors, one for each row (corresponding to each output channel):
where:
- : the scale factor for the -th output channel (row)
- : the vector of weights in the -th output channel (row of the weight matrix)
- : the maximum absolute value in that channel's weights
- : the maximum value for a signed 8-bit integer
This allows each output channel to use its full INT8 range, accommodating the fact that different channels often have different magnitude distributions. Consider why this happens: during training, different output features may learn patterns of varying intensity. A channel that detects subtle statistical patterns might have small weights, while a channel that looks for strong, distinctive features might have large weights. Per-channel quantization respects these differences by giving each channel its own appropriately-sized quantization grid.
The overhead is storing scale factors instead of one, which is negligible compared to the weight tensor itself. For a layer with 4096 output channels and 4096 input features, we store 4096 scale factors (one per channel) versus over 16 million quantized weights. The scale factors add less than 0.1% to the storage requirements while potentially sharply improving accuracy.
Per-channel quantization for weights combined with per-tensor quantization for activations has emerged as the standard approach. It offers good accuracy with manageable complexity. The asymmetry makes sense: weights are static and can be analyzed offline to compute optimal per-channel scales, while activations vary with each input and would require expensive runtime scale computation for per-channel treatment. Performing per-channel scale computation on activations at every forward pass would add latency that could negate the inference speedup from quantization, defeating the purpose.
Worked Example: Quantizing a Small Weight Tensor
To build concrete intuition, let's trace through a complete quantization example with specific numbers. This will make the abstractions tangible and reveal where errors accumulate.
Suppose we have a small weight tensor (one row of a larger matrix):
Step 1: Compute the absmax. We scan the entire vector and find:
Step 2: Compute the scale factor. Dividing by 127 to cover the full INT8 range:
Step 3: Quantize each element. Divide by and round to the nearest integer:
The quantized tensor is: .
Step 4: Reconstruct and compute error. Multiply each integer by :
The maximum error is 0.0037, and the mean absolute error is about 0.0015. Notice that the largest element (1.12) reconstructs perfectly since it was the absmax and maps exactly to 127. The element with value 0.05 (the smallest magnitude) experiences the largest relative error: the rounding from 5.67 to 6 introduces a 5.9% relative error. This pattern is typical: absmax quantization gives best relative precision to large-magnitude values and worst to small ones.
The maximum possible error for any element is , which all our actual errors satisfy. Storing the entire vector requires just 8 bytes (8 integers) plus 4 bytes for the float32 scale factor, totaling 12 bytes instead of the original 32 bytes (8 float32 values). That is 2.67x compression with bounded error.
The Outlier Problem in Large Language Models
As language models scale to billions of parameters, a problematic pattern emerges: certain activation dimensions develop extreme outlier values. Research has shown that in models like OPT-175B and BLOOM-176B, a small number of hidden dimensions (sometimes called "large activations" or "emergent features") can have values 10-100x larger than typical activations. This phenomenon was not anticipated by early quantization research conducted on vision models and smaller transformers, and it poses a significant challenge to naive INT8 approaches.
The emergence of these outliers appears to be linked to how transformers process information at scale. Certain hidden dimensions seem to serve as "highways" for important information, developing consistently large activation magnitudes across many inputs. These are not random fluctuations but systematic features of the model's learned representations. The same channels show outlier behavior across different inputs and different layers, suggesting they encode something structurally important about how the model is information. While we are still working to fully understand why this happens, the practical implication is clear: any quantization strategy for large language models must account for these outliers.
The key insight is that outlier channels break the basic assumption behind per-tensor quantization. When we assume all channels have similar magnitude distributions and can share a single scale factor, we implicitly assume that no channel is sharply different from the others. The outlier phenomenon violates this assumption catastrophically. One channel's behavior forces a scale factor that is appropriate for that channel alone and wildly suboptimal for all others.
Consider a hidden state where 99.9% of values lie in , but dimension 1847 consistently produces values around 50. Per-tensor absmax quantization would set , meaning values of magnitude 1 quantize to just 2-3 integers. The reconstruction error for these common values becomes unacceptable. A value of 0.5, which might be important for the model's computation, would quantize to either 1 or 2, introducing errors of up to 20% in its representation. This is not a minor precision loss that can be ignored. It can cause the model to make qualitatively wrong predictions.
The naive solutions all have significant drawbacks:
- Clipping outliers: Destroys important model information encoded in those dimensions. The outlier channels are apparently important, not artifacts.
- Per-channel activation quantization: Requires computing new scale factors for every token at runtime, adding latency and complexity.
- Mixed precision: Keeping outlier channels in FP16 complicates kernels and reduces throughput, partially defeating the purpose.
Each of these approaches involves either losing accuracy, losing speed, or adding significant implementation complexity. What was needed was a technique that could tame the outlier problem without sacrificing the efficiency benefits of uniform INT8 quantization. Smooth quantization gives exactly this.
Smooth Quantization
The key insight behind smooth quantization is that while activations are hard to quantize (dynamic, contain outliers), weights are easy to quantize (static, relatively uniform). We can mathematically migrate the quantization difficulty from activations to weights through a channel-wise scaling transformation, making the problem tractable.
The intuition is straightforward. If one matrix is "spiky" with outliers and another is "smooth" and well-behaved, we can transfer some of the spikiness from the first to the second. As long as we do this in a mathematically consistent way that preserves the final computation, we have not changed the model's behavior, only redistributed the quantization difficulty to where it is easier to handle. Think of smooth quantization as a seesaw: we load more weight (quantization difficulty) onto the weight side, which can bear it, and unload the activation side, which was struggling under the strain.
The insight that makes this possible is that the product is unchanged if we multiply by some diagonal matrix and also multiply by . These operations are inverses and cancel out. But while they cancel mathematically, they can sharply change the magnitude distributions of both matrices, making each individually much easier to quantize. The question is just how to choose to optimally redistribute the difficulty.
The Mathematical Equivalence
Consider a linear layer computing , where is the activation input and is the weight matrix. We introduce a diagonal matrix with positive scaling factors, then insert between the two matrices:
where:
- : the output of the linear layer, unchanged by this transformation
- : the input activation tensor with outlier channels
- : the weight matrix with relatively uniform distribution
- : the diagonal smoothing matrix containing scaling factors for each input dimension
- : the smoothed activations, where each column is divided by
- : the adjusted weights, where each row is multiplied by
This transformation is mathematically exact since . We have not changed the computation, just how we factor it. The key mathematical property we exploit is that multiplying by a matrix and then its inverse produces the identity, so inserting between and changes nothing about the final result. We are free to group these factors however we like, and we choose to group with the activations and with the weights.
The transformation affects the data as follows: smoothed activations divide each column of by , which shrinks outlier channels by choosing large values. Meanwhile, adjusted weights multiply each row of by the same , increasing the magnitude of those rows to compensate. If channel has outlier activations, we choose a large . This shrinks the activations (making them easier to quantize) while enlarging the corresponding weights (which are already easy to quantize). We are transferring the outlier problem to a domain where it is more manageable.
Choosing the Smoothing Factors
How do we pick the values? If we make too large, we might shrink the activations so much that we lose precision in other ways, or we might inflate the weights to the point where they become hard to quantize. The optimal choice balances these competing concerns, and the SmoothQuant paper proposes a principled way to reach this balance.
The SmoothQuant paper proposes balancing the quantization difficulty between activations and weights using a migration strength parameter :
where:
- : the smoothing factor for the -th input channel
- : the maximum absolute value in the -th column of activations, measured across a calibration dataset
- : the maximum absolute value in the -th row of weights
- : the migration strength hyperparameter controlling how much difficulty shifts from activations to weights
This formula balances two competing goals. When , we have , which fully smooths the activations by dividing each channel by its maximum. This makes the activations perfectly balanced but potentially makes weights hard to quantize. When , we have , which leaves activations unchanged. The intermediate value gives the geometric mean, which ensures that after transformation, the maximum activation and maximum weight for each channel have roughly equal magnitude, spreading the quantization difficulty evenly.
To understand this formula intuitively, consider what "equal difficulty" means. After applying the transformation, the smoothed activation maximum for channel is , and the adjusted weight maximum is . For these to be equal:
Solving for :
This is exactly the case of the SmoothQuant formula, confirming that is the "equal difficulty" point. Values above 0.5 shift more difficulty to the weights, values below 0.5 leave more with the activations.
In practice, works well for most models, achieving roughly equal quantization ranges for both activations and weights after smoothing. However, some models may benefit from different values. Models with particularly severe activation outliers might use to more aggressively smooth the activations, while models with already-difficult-to-quantize weights might use to avoid making the weights worse.

Offline Computation and LayerNorm Fusion
A necessary practical advantage of smooth quantization is that the transformation can be computed entirely offline during a calibration phase. This means inference has no additional runtime cost compared to standard quantized inference. The calibration procedure works as follows:
First, run a small calibration dataset (typically a few hundred representative samples) through the model to collect activation statistics. During this calibration forward pass, we record the maximum absolute value seen in each channel across all calibration samples. Second, compute for each channel using these statistics, capturing how large the outliers tend to be. Third, calculate smoothing factors for each layer using the formula above. Fourth, apply the weight transformation and store the modified weights. Fifth, and most cleverly, fuse into the preceding layer's weights or normalization parameters.
Step 5 is important and deserves careful explanation. Rather than explicitly dividing activations by at runtime, we absorb this division into the bias of the previous layer or the scale/shift parameters of the preceding LayerNorm. This means inference has zero additional cost compared to standard quantized inference. The activation values arrive at the quantization step already "smoothed" because the previous layer was modified to produce smoothed outputs.
The fusion works because transformer architectures typically have a LayerNorm immediately before each linear layer. LayerNorm applies a learned scale () and shift () to each channel. We can incorporate the smoothing factor directly into these existing parameters. If the original LayerNorm scale for channel was and shift was , we replace them with and . The activations coming out of the modified LayerNorm are already "smoothed," and we never perform the division at runtime. This is mathematically equivalent to explicitly dividing at every forward pass but requires zero additional computation.
The offline nature of smooth quantization also means it is compatible with any existing INT8 quantization infrastructure. You do not need to modify the inference engine or write new kernels. Simply transform the weights and LayerNorm parameters before quantizing, then proceed with standard absmax or per-channel quantization. The resulting model can be deployed on any hardware that supports INT8 matrix multiplication.
Quantized Matrix Multiplication
Understanding how quantized operations execute helps clarify why INT8 gives speedup. The key mechanism is that we perform the heavy matrix multiplication entirely in integer arithmetic, then apply scale factors to recover floating-point values for the final output. This two-stage approach exploits the fact that integer operations are cheaper than floating-point operations, especially for the large matrix multiplications that dominate transformer inference.
When both weights and activations are in INT8, the matrix multiplication proceeds as follows. We multiply the integer tensors together, accumulating into 32-bit integers to prevent overflow, and then convert and rescale the result:
where:
- : the intermediate result matrix, accumulated in 32-bit integers
- : the quantized input activations in 8-bit integer format
- : the quantized weights in 8-bit integer format
The integer multiplication produces INT32 results to avoid overflow from accumulating many INT8 products. Each INT8 times INT8 multiplication produces a result up to , which fits in INT16. However, a matrix multiplication accumulates thousands or millions of such products, and these sums can easily exceed the INT16 range. INT32 gives enough headroom to accumulate the typical number of products found in neural network layers without overflow. The hardware handles this automatically: dedicated INT8 multiply-accumulate (MAC) units produce 32-bit partial sums that can be safely added together.
After the integer matrix multiplication, we convert back to floating-point and apply the scale factors to recover approximate floating-point outputs:
where:
- : the final floating-point output tensor
- : the quantization scale factor for the activations (scalar for per-tensor, vector for per-channel)
- : the quantization scale factor for the weights (scalar or per-channel vector)
- : the integer output from the matrix multiplication, converted to float before scaling
The scale factors and are multiplied together before being applied, since the quantization error from both sources combines multiplicatively. Since scale factors are scalar (or per-channel), this final rescaling is cheap compared to the matrix multiplication itself. The computational cost of the rescaling step is for an output matrix of size , while the matrix multiplication itself is . Since is typically large (thousands of features), the rescaling overhead is less than 0.1% of the total computation.
Hardware INT8 units execute this pattern efficiently because of several compounding advantages. INT8 multiply-accumulate operations use simpler circuits than floating-point, with no need to handle exponent alignment, mantissa normalization, or special cases like NaN and infinity. Smaller operands allow higher parallelism in the same silicon area: the same chip real estate that fits one FP32 multiplier can fit four INT8 multipliers. Memory bandwidth is halved compared to FP16, reducing the primary bottleneck for large models where the GPU must stream weights from memory for every token.
Implementation
Let's implement INT8 quantization from scratch to solidify these concepts. The implementation will proceed from basic absmax quantization through per-channel quantization, then show the outlier problem and show how smooth quantization resolves it.
import numpy as np
import torch
# Set random seed for reproducibility
torch.manual_seed(42)
np.random.seed(42)Absmax Quantization
We'll start with the basic absmax quantization and dequantization operations. The function returns both the integer tensor and the scale factor, since we need both to reconstruct the original values:
def absmax_quantize(x: torch.Tensor) -> tuple[torch.Tensor, float]:
"""
Quantize a tensor to INT8 using absmax symmetric quantization.
Returns:
quantized: INT8 tensor
scale: Scale factor for dequantization
"""
# Find the maximum absolute value
absmax = x.abs().max().item()
# Compute scale factor (avoid division by zero)
scale = absmax / 127.0 if absmax != 0 else 1.0
# Quantize: divide by scale and round to nearest integer
quantized = torch.round(x / scale).clamp(-128, 127).to(torch.int8)
return quantized, scale
def dequantize(quantized: torch.Tensor, scale: float) -> torch.Tensor:
"""Recover floating-point approximation from quantized tensor."""
return quantized.float() * scaleLet's test this on a simple tensor and observe the quantization error:
# Create a sample weight tensor
weights = torch.randn(4, 4) * 0.5
# Quantize
quantized, scale = absmax_quantize(weights)
# Dequantize
reconstructed = dequantize(quantized, scale)
# Compute error
error = (weights - reconstructed).abs()Original weights: [[ 0.9635 0.7436 0.4504 -1.0528] [ 0.3392 -0.6173 -0.0215 -0.8023] [-0.3761 0.8244 -0.1962 -0.7018] [-0.3639 -0.2797 -0.3844 0.3812]] Scale factor: 0.008289 Quantized (INT8): [[ 116 90 54 -127] [ 41 -74 -3 -97] [ -45 99 -24 -85] [ -44 -34 -46 46]] Reconstructed: [[ 0.9616 0.7461 0.4476 -1.0528] [ 0.3399 -0.6134 -0.0249 -0.8041] [-0.373 0.8207 -0.1989 -0.7046] [-0.3647 -0.2818 -0.3813 0.3813]] Absolute error: [[1.881e-03 2.409e-03 2.728e-03 0.000e+00] [6.580e-04 3.853e-03 3.335e-03 1.744e-03] [3.043e-03 3.705e-03 2.708e-03 2.800e-03] [7.950e-04 2.127e-03 3.105e-03 9.200e-05]] Max error: 0.003853 Mean error: 0.002186
The maximum error is bounded by half the quantization step size (), as values are rounded to the nearest integer. For this tensor with a scale around 0.01, our maximum possible error is about 0.005. The element with the largest magnitude (which determined the absmax) reconstructs with zero error, as expected from the theory.
Analyzing Quantization Error
Let's visualize how quantization error varies across different value ranges. This reveals the characteristic "sawtooth" pattern that emerges from uniform rounding:
# Generate values across a range
original_values = torch.linspace(-1.5, 1.5, 1000)
# Quantize
scale = 1.5 / 127
quantized_vals = torch.round(original_values / scale).clamp(-128, 127)
reconstructed_vals = quantized_vals * scale
# Compute errors
errors = original_values - reconstructed_vals

The error follows a uniform distribution between and , exactly what we expect from rounding to the nearest integer on a uniform grid. This uniformity has an important implication: the expected value of the quantization error is zero, meaning quantization introduces noise but not systematic bias. This property helps explain why quantized models often perform nearly as well as full-precision models despite the apparent information loss.
Per-Channel Quantization
Now let's implement per-channel quantization for weight matrices. This extends the basic absmax approach by computing a separate scale for each output channel:
def per_channel_quantize(
weight: torch.Tensor,
) -> tuple[torch.Tensor, torch.Tensor]:
"""
Quantize weight matrix with per-output-channel scales.
Args:
weight: Shape (out_features, in_features)
Returns:
quantized: INT8 weight matrix
scales: Scale factor for each output channel
"""
# Compute absmax for each row (output channel)
absmax_per_channel = weight.abs().max(dim=1).values
# Compute scales (avoid division by zero)
scales = absmax_per_channel / 127.0
scales = torch.where(scales == 0, torch.ones_like(scales), scales)
# Quantize each row by its scale
quantized = (
torch.round(weight / scales.unsqueeze(1))
.clamp(-128, 127)
.to(torch.int8)
)
return quantized, scales
def per_channel_dequantize(
quantized: torch.Tensor, scales: torch.Tensor
) -> torch.Tensor:
"""Dequantize using per-channel scales."""
return quantized.float() * scales.unsqueeze(1)Let's compare per-tensor vs per-channel quantization on a weight matrix with varying row magnitudes. We artificially scale different rows to simulate the kind of distribution variation you find in real neural networks:
# Create weight matrix where different rows have different magnitudes
# This simulates real neural network weights
weights_varied = torch.randn(8, 16)
# Artificially scale different rows
row_scales = torch.tensor([0.1, 0.2, 0.5, 1.0, 2.0, 0.3, 0.8, 0.15])
weights_varied = weights_varied * row_scales.unsqueeze(1)
# Per-tensor quantization
qt_tensor, scale_tensor = absmax_quantize(weights_varied)
recon_tensor = dequantize(qt_tensor, scale_tensor)
error_tensor = (weights_varied - recon_tensor).abs()
# Per-channel quantization
qt_channel, scales_channel = per_channel_quantize(weights_varied)
recon_channel = per_channel_dequantize(qt_channel, scales_channel)
error_channel = (weights_varied - recon_channel).abs()Row magnitude scales: [0.1 0.2 0.5 1. 2. 0.3 0.8 0.15] Per-tensor scale: 0.0271 Per-channel scales: [0.0013 0.0033 0.0073 0.0175 0.0271 0.0059 0.0092 0.0023] Per-tensor max error: 0.013417 Per-channel max error: 0.013251 Per-tensor mean error: 0.006463 Per-channel mean error: 0.002101 Mean error per row: Row | Magnitude | Per-tensor | Per-channel --------------------------------------------- 0 | 0.10 | 0.004851 | 0.000280 1 | 0.20 | 0.009060 | 0.000748 2 | 0.50 | 0.005872 | 0.002023 3 | 1.00 | 0.007094 | 0.003080 4 | 2.00 | 0.006667 | 0.006667 5 | 0.30 | 0.005832 | 0.001644 6 | 0.80 | 0.006436 | 0.001909 7 | 0.15 | 0.005891 | 0.000461
Per-channel quantization reaches substantially lower error, especially for rows with small magnitudes. The per-tensor approach is forced to use a scale dictated by the largest row, wasting precision for smaller rows. This matches our theoretical understanding: the row with magnitude 0.1 gets 20x worse relative precision under per-tensor quantization than the row with magnitude 2.0, while per-channel quantization gives both rows equally good relative precision.

Demonstrating the Outlier Problem
Let's simulate the outlier pattern seen in large language models. We inject a single high-magnitude channel into an otherwise normal activation tensor and observe the damage to quantization quality:
# Simulate activation tensor with outlier channel
batch_size, hidden_dim = 32, 128
activations = torch.randn(batch_size, hidden_dim)
# Inject outliers in channel 50 (simulating LLM behavior)
outlier_channel = 50
activations[:, outlier_channel] *= 20 # 20x larger than typical
# Per-tensor quantization
qt_act, scale_act = absmax_quantize(activations)
recon_act = dequantize(qt_act, scale_act)
error_act = (activations - recon_act).abs()
# Separate error for outlier vs normal channels
normal_mask = torch.ones(hidden_dim, dtype=torch.bool)
normal_mask[outlier_channel] = False
error_normal = error_act[:, normal_mask].mean().item()
error_outlier = error_act[:, outlier_channel].mean().item()Activation statistics: Normal channels max: 3.83 Outlier channel max: 38.07 Quantization scale: 0.2998 Mean absolute error: Normal channels: 0.075173 Outlier channel: 0.063882 Relative error (error/value): Normal channels: 9.4% Outlier channel: 0.4%
The outlier channel forces a large scale factor, causing severe relative error for normal channels. This is exactly the problem that smooth quantization solves. Despite having 20x larger values, the outlier channel reconstructs better in relative terms than the normal channels because the scale was set based on its magnitude. Meanwhile, the normal channels suffer because their values occupy only a tiny portion of the INT8 range.


Implementing Smooth Quantization
Now let's implement the smooth quantization transformation. The core logic fits in two functions: one that computes the smoothing scales from statistics, and one that applies the transformation to the tensors:
def compute_smooth_scales(
activation_absmax: torch.Tensor, # Per-channel activation max
weight_absmax: torch.Tensor, # Per-channel weight max
alpha: float = 0.5,
) -> torch.Tensor:
"""
Compute smoothing scale factors.
Args:
activation_absmax: Max absolute activation for each input channel
weight_absmax: Max absolute weight for each input channel (row of W)
alpha: Migration strength (0=no smoothing, 1=full smoothing)
Returns:
scales: Smoothing factors for each channel
"""
# s_j = (act_max^alpha) / (weight_max^(1-alpha))
scales = (activation_absmax**alpha) / (weight_absmax ** (1 - alpha))
# Clamp to avoid extreme values
scales = scales.clamp(min=1e-5)
return scales
def apply_smooth_transform(
activations: torch.Tensor, # (batch, in_features)
weights: torch.Tensor, # (in_features, out_features)
scales: torch.Tensor, # (in_features,)
) -> tuple[torch.Tensor, torch.Tensor]:
"""
Apply smoothing transformation to activations and weights.
Returns:
smoothed_activations: X @ diag(1/s)
smoothed_weights: diag(s) @ W
"""
# Divide activations by scales (broadcast across batch)
smoothed_activations = activations / scales.unsqueeze(0)
# Multiply weight rows by scales
smoothed_weights = weights * scales.unsqueeze(1)
return smoothed_activations, smoothed_weightsLet's apply smooth quantization to our outlier problem and measure the improvement:
# Create weight matrix for the same linear layer
weights_for_layer = torch.randn(hidden_dim, 64) * 0.1 # (in, out)
# Compute channel-wise statistics
act_absmax = activations.abs().max(dim=0).values # (hidden_dim,)
weight_absmax = weights_for_layer.abs().max(dim=1).values # (hidden_dim,)
# Compute smooth scales with alpha=0.5
smooth_scales = compute_smooth_scales(act_absmax, weight_absmax, alpha=0.5)
# Apply transformation
smoothed_act, smoothed_weights = apply_smooth_transform(
activations, weights_for_layer, smooth_scales
)
# Now quantize the smoothed tensors
qt_smooth_act, scale_smooth_act = absmax_quantize(smoothed_act)
recon_smooth_act = dequantize(qt_smooth_act, scale_smooth_act)
error_smooth_act = (smoothed_act - recon_smooth_act).abs()
# Transform error back to original space for comparison
# (multiply back by scales to compare apples-to-apples)
error_smooth_original = error_smooth_act * smooth_scales.unsqueeze(0)Before smoothing:
Activation max (normal): 3.83
Activation max (outlier): 38.07
Ratio: 9.9x
After smoothing:
Activation max (normal): 1.08
Activation max (outlier): 3.30
Ratio: 3.1x
Smooth scale for outlier channel: 11.52
Quantization error (in original space):
Without smoothing:
Normal channels: 0.075173
Outlier channel: 0.063882
With smoothing:
Normal channels: 0.019993
Outlier channel: 0.063882Smooth quantization sharply reduces the activation range ratio, making per-tensor quantization much more effective. The error for normal channels drops materially because they are no longer penalized by the outlier's presence. The smoothing scales absorb the outlier information into the weights, which can accommodate the increased magnitude without sacrificing precision.


End-to-End Quantized Linear Layer
Let's put everything together into a complete quantized linear layer. This class combines per-channel weight quantization with optional smooth quantization to handle outlier activations:
class QuantizedLinear:
"""
INT8 quantized linear layer with optional smooth quantization.
"""
def __init__(
self,
weight: torch.Tensor, # Original FP32 weights (out, in)
bias: torch.Tensor = None,
smooth_scales: torch.Tensor = None, # If provided, applies smoothing
):
self.bias = bias
self.smooth_scales = smooth_scales
# If smoothing, transform weights
if smooth_scales is not None:
weight = weight * smooth_scales.unsqueeze(0) # (out, in) * (in,)
# Quantize weights (per-channel)
self.weight_quantized, self.weight_scales = per_channel_quantize(weight)
def forward(self, x: torch.Tensor) -> torch.Tensor:
"""
Forward pass with INT8 computation.
In real implementations, this would use INT8 GEMM kernels.
Here we simulate the numerical behavior.
"""
# Apply activation smoothing if configured
if self.smooth_scales is not None:
x = x / self.smooth_scales.unsqueeze(0)
# Quantize activations (per-tensor)
x_quantized, x_scale = absmax_quantize(x)
# Simulate INT8 matrix multiplication
# In hardware: INT8 x INT8 -> INT32, accumulated
# Here we do it in float but with quantized values
output_int = x_quantized.float() @ self.weight_quantized.float().T
# Dequantize: multiply by both scales
output = output_int * x_scale * self.weight_scales.unsqueeze(0)
if self.bias is not None:
output = output + self.bias
return outputNow let's compare the quantized layer against full precision on inputs with outlier activations:
# Create a test linear layer
in_features, out_features = 128, 64
weight = torch.randn(out_features, in_features) * 0.1
bias = torch.randn(out_features) * 0.01
# Test input (with outlier)
test_input = torch.randn(16, in_features)
test_input[:, 50] *= 20 # Inject outlier
# Full precision output
fp_output = test_input @ weight.T + bias
# Quantized without smoothing
quant_layer = QuantizedLinear(weight, bias, smooth_scales=None)
quant_output = quant_layer.forward(test_input)
# Compute smooth scales from test input
act_max = test_input.abs().max(dim=0).values
weight_max = weight.abs().max(dim=0).values
smooth_s = compute_smooth_scales(act_max, weight_max, alpha=0.5)
# Quantized with smoothing
smooth_quant_layer = QuantizedLinear(weight, bias, smooth_scales=smooth_s)
smooth_quant_output = smooth_quant_layer.forward(test_input)
# Calculate errors for comparison
error_quant = (fp_output - quant_output).abs()
error_smooth = (fp_output - smooth_quant_output).abs()Linear layer output comparison:
Output shape: torch.Size([16, 64])
Without smoothing:
Max error: 0.224213
Mean error: 0.049900
Relative error: 3.33%
With smooth quantization:
Max error: 0.070532
Mean error: 0.017435
Relative error: 1.16%Smooth quantization reduces both maximum and mean error, making INT8 quantization viable even with outlier activations. The combination of techniques we have implemented here, per-channel weight quantization plus smooth activation quantization, is the current best practice for deploying large language models with INT8 precision.


Key Parameters
The key parameters for the INT8 quantization implementation are:
- alpha: The migration strength in smooth quantization, typically 0.5. Controls how much quantization difficulty is shifted from activations to weights. Higher values (closer to 1.0) shift more difficulty to the weights; lower values leave more with the activations.
- scale: The step size of the quantization grid. Determined by the maximum absolute value (absmax) in the tensor or channel. Smaller scales give finer precision but only stand for values in a smaller range.
- smooth_scales: The channel-wise scaling factors derived from calibration data to balance activation and weight magnitudes. These are computed once and stored as part of the model.
- Calibration dataset size: Typically 256 to 512 samples suffice for collecting reliable activation statistics. Too few samples may miss worst-case outlier channels.
INT8 Accuracy in Practice
The accuracy impact of INT8 quantization varies materially between model architectures and parameter counts. It also depends on the task. Understanding these patterns helps you set realistic expectations and choose the right approach for your deployment scenario.
Models under 1 billion parameters typically quantize to INT8 with minimal degradation, usually less than 1% accuracy loss on standard benchmarks. The weight distributions are well-behaved, and activation outliers are rare or mild. For these models, simple per-channel weight quantization with per-tensor activation quantization usually suffices without any smoothing.
Models from 1 billion to 10 billion parameters show increasing sensitivity. Without techniques like smooth quantization, accuracy can degrade 2-5% on certain tasks. Perplexity increases are noticeable but often acceptable for deployment. At this scale, smooth quantization or careful calibration-based clipping is typically necessary to maintain acceptable quality.
Models above 10 billion parameters exhibit the emergent outlier features described earlier, making naive INT8 quantization fail noticeably. Smooth quantization or similar techniques become needed at this scale, and even with best practices, a small accuracy gap compared to FP16 usually remains. The LLM.int8() technique, which keeps outlier channels in FP16 while quantizing the rest to INT8, was the original solution to this problem and remains useful when smooth quantization is insufficient.
Task sensitivity also matters. Simple classification tasks are reliable to quantization because the decision boundary does not require high numerical precision. Open-ended generation shows more degradation as errors compound across hundreds or thousands of autoregressive steps. Mathematical reasoning and code generation are particularly sensitive because small numerical errors can cascade into completely wrong answers.
The calibration dataset is an often-overlooked factor in INT8 accuracy. The statistics used to set scale factors and smooth quantization parameters must come from data representative of the deployment distribution. A model calibrated on English text may produce suboptimal quantization parameters for code completion or multilingual inputs. Choosing calibration data carefully, and potentially using task-specific calibration, can materially improve quantized model quality.
Limitations and Practical Considerations
INT8 quantization is a practical and widely-deployed optimization, but it comes with important limitations that you should understand before deploying in production. The limitations are not reasons to avoid quantization, but they are constraints that must be accounted for in system design.
The most basic limitation is that some model information is permanently lost during quantization. While smooth quantization and careful calibration minimize this loss, they cannot eliminate it entirely. For safety-necessary applications or tasks requiring the highest possible accuracy, the small degradation from INT8 may be unacceptable. This is why many production systems maintain FP16 or FP32 models for evaluation and fine-tuning while deploying INT8 for inference. The irreversibility is by design: if we could recover the original values losslessly from INT8, we would not have achieved any compression. The compression and the information loss are two sides of the same coin.
Calibration data sensitivity presents another practical challenge. Smooth quantization computes channel-wise statistics from a calibration dataset, and if this data does not stand for the actual deployment distribution, the smoothing factors may be suboptimal. A model calibrated on English text may perform poorly when quantized for code completion or multilingual inputs. In production, this means you need calibration pipelines that are maintained alongside the model and updated when the deployment distribution shifts. It also means that models with broad deployment contexts (such as general-purpose chatbots serving diverse user queries) may need larger and more diverse calibration sets to reach reliable quantization quality.
Hardware support, while widespread, is not universal. INT8 acceleration requires specific hardware features such as NVIDIA's Tensor Cores or ARM's NEON instructions, and the speedup varies by platform and use case. On CPUs without INT8 vectorization, quantized inference may be slower than FP32 due to the quantization and dequantization overhead. On GPUs, the speedup is most pronounced for large matrix multiplications (typical of large-batch inference) but may be minimal for small-batch or latency-sensitive inference where compute is not the bottleneck. Always benchmark on your target hardware before committing to a quantization strategy.
Attention layers present a special challenge that pure weight quantization does not solve. The key-value cache that accumulates during autoregressive generation stores activation values, not weights. As sequences grow longer, the KV cache grows proportionally and can become the dominant memory consumer. Quantizing the KV cache requires activation quantization, which is harder than weight quantization because activation distributions vary with the input. Dedicated techniques for KV cache quantization are an active area of development and are addressed in later chapters.
Finally, INT8 may not give sufficient compression for edge deployment scenarios. With 8 bits per weight, a 7B parameter model still requires 7 GB of storage, which exceeds the capabilities of most mobile devices and many edge accelerators. For mobile or embedded applications, the more aggressive INT4 quantization techniques we will explore in the next chapter become necessary, accepting greater accuracy trade-offs in exchange for smaller model sizes that fit in constrained hardware.
Summary
INT8 quantization gives a practical path to halving model memory and accelerating inference through hardware INT8 units. The full picture requires understanding several interconnected concepts, each building on the others:
- Absmax quantization maps floating-point values to INT8 using a single scale factor derived from the maximum absolute value, centering the quantization grid at zero to preserve exact zeros
- Asymmetric quantization introduces a zero-point offset to shift the grid for data distributions that are not centered at zero, improving efficiency for activations like ReLU outputs
- Per-channel quantization applies separate scales to each output channel, sharply reducing error when channels have different magnitude distributions while adding negligible storage overhead
- The outlier problem emerges in large language models where a small fraction of activation dimensions have values 10-100x larger than typical, making per-tensor quantization fail catastrophically for these models
- Smooth quantization solves the outlier problem by migrating quantization difficulty from activations to weights through a mathematically equivalent transformation, computable entirely offline and fusible into existing LayerNorm parameters at zero runtime cost
- The smoothing factor balances activation and weight ranges using a migration strength , with giving a balanced default that works well for most models
INT8 quantization with smooth quantization reaches near-FP16 accuracy on models up to hundreds of billions of parameters while giving 2x memory reduction and significant throughput improvements on compatible hardware. The techniques described here form the foundation for the inference infrastructure behind most large language model deployments today. The upcoming chapters on INT4 quantization, GPTQ, and AWQ explore techniques that push quantization to just 4 bits per weight, accepting somewhat more accuracy degradation in exchange for even greater memory efficiency.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about INT8 quantization techniques.
INT8 Quantization
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!