GPTQ: Optimizing 4-Bit Weight Quantization for LLMs

Michael BrenndoerferJanuary 13, 202649 min read

Part of Language AI Handbook

Explains how GPTQ optimizes weight quantization using Hessian-based error compensation to compress LLMs to 4 bits while maintaining near-FP16 accuracy.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

GPTQ

In the previous chapters on INT8 and INT4 quantization, we explored straightforward approaches to weight compression: round each weight to the nearest representable value in the target format. While these methods work reasonably well, they treat each weight independently, ignoring the complex interactions between weights in a neural network. A small rounding error in one weight might be catastrophic for model quality, while a larger error in another weight might barely matter. The round-to-nearest strategy has no mechanism to distinguish between these cases, and so it ends up treating a harmless rounding as equally important as a damaging one.

GPTQ (GPT Quantization) takes a fundamentally different approach. Rather than treating quantization as a simple rounding problem, GPTQ frames it as an optimization problem: given a layer's weights, find the quantized values that minimize the layer's output error. The key insight is that when you quantize one weight, you can partially compensate for the resulting error by adjusting the remaining weights before they too are quantized. This sequential compensation mechanism propagates corrections through the weight matrix, sharply reducing the accumulated error that plagues naive methods.

Think of GPTQ as a team of engineers managing a large construction project where every measurement must be rounded to the nearest centimeter. A naive approach would round each measurement independently and accept whatever cumulative error results. A smarter approach, the one GPTQ takes, observes each rounding error as it happens and adjusts subsequent measurements to keep the overall structure as close as possible to the original blueprint. Each small correction nudges the final result back toward accuracy, so that by the time you've finished rounding all measurements, the cumulative deviation is far smaller than if you had rounded everything in isolation.

This compensation mechanism, combined with algorithmic optimizations, allows GPTQ to reach quantization error far below what naive methods can manage. Models quantized with GPTQ to 4 bits often perform nearly as well as their full-precision counterparts, letting LLMs with tens of billions of parameters to run on consumer GPUs. GPTQ was one of the first methods to make models like LLaMA-65B practically usable on hardware that would otherwise be unable to load them, compressing 130 GB of FP16 weights into under 35 GB of INT4 data.

The mathematical foundation of GPTQ draws from classical work in neural network compression, particularly the Optimal Brain Surgeon framework developed in the early 1990s for pruning. GPTQ's authors recognized that the same machinery developed to remove weights from a network could be repurposed to quantize weights with minimal damage. This cross-pollination of ideas from pruning into quantization is a recurring theme in modern efficiency research: the core question in both settings is how to perturb a network's weights while preserving its behavior.

Understanding GPTQ deeply will serve you well beyond knowing how to call the library. The concepts here, particularly the Hessian matrix as a measure of weight importance and the closed-form compensation formula, appear throughout modern quantization research. AWQ, QuIP, and AQLM all build on ideas that GPTQ establishes. By the end of this chapter, you will have the theoretical grounding to understand GPTQ and the entire family of second-order quantization methods that have followed it.

Historical Context

GPTQ was introduced in the paper "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" by Frantar, Ashkboos, Hoefler, and Alistarh in 2022. It was a point where practice changed in the broader access of large language models. Before GPTQ, running a 65B parameter model required enterprise-grade multi-GPU setups. After GPTQ, a single consumer GPU with 24 GB of VRAM could load and run the same model. The paper built on Optimal Brain Quantization (OBQ), which itself descended from the Optimal Brain Surgeon (OBS) framework by Hassibi and Stork (1993) and the later Optimal Brain Compression (OBC) work. The lineage from 1990s pruning research to 2022 LLM quantization illustrates how basic theoretical ideas persist and resurface as new applications demand them.

The Layer-Wise Reconstruction Objective

To understand how GPTQ approaches quantization, we must first establish what it means to quantize well. The basic question is: what objective should we optimize? One natural answer might be to minimize the difference between original and quantized weights directly. However, this ignores a important insight: not all weight errors are equally harmful. What truly matters is how quantization affects the layer's output, because the output is what downstream computations depend upon. A weight that multiplies a small input contributes little to the output regardless of its quantization error; a weight that multiplies a large, frequently-occurring input is necessary and must be handled with care.

GPTQ operates on one layer at a time, treating each layer's quantization as an independent optimization problem. This layer-wise decomposition is both a practical necessity and a reasonable approximation. It is practical because computing gradients through the entire model for every quantization decision would require a full backward pass per weight, which is completely infeasible for models with billions of parameters. It is a reasonable approximation because the output of each layer is the input to the next, and minimizing the error introduced at each layer independently tends to minimize the total error propagated through the network.

Think of it as optimizing each department of a large company independently. You cannot simultaneously optimize every decision across every team, but if each department minimizes its own errors, the overall company performance tends to be much better than if no one cared about local quality. The layer-wise approach gives GPTQ a tractable decomposition of an otherwise intractable global problem.

Consider a linear layer with weights W∈Rdout×din\mathbf{W} \in \mathbb{R}^{d_{out} \times d_{in}} and inputs X∈Rdin×n\mathbf{X} \in \mathbb{R}^{d_{in} \times n}, where nn is the number of calibration tokens we use to estimate the layer's behavior. The goal is to find quantized weights W^\hat{\mathbf{W}} that minimize the reconstruction error:

L(W^)=∣∣WX−W^X∣∣F2\mathcal{L}(\hat{\mathbf{W}}) = ||\mathbf{W}\mathbf{X} - \hat{\mathbf{W}}\mathbf{X}||_F^2

where:

  • L(W^)\mathcal{L}(\hat{\mathbf{W}}): the reconstruction error (loss) for the quantized weights
  • W\mathbf{W}: the original high-precision weight matrix
  • W^\hat{\mathbf{W}}: the quantized weight matrix we are searching for
  • X\mathbf{X}: the input matrix containing calibration data (each column is one token's activations)
  • ∣∣⋅∣∣F||\cdot||_F: the Frobenius norm, equal to the square root of the sum of all squared elements

This loss function measures how much the layer's output changes due to quantization. We are comparing outputs, not weights directly. This distinction is needed: a weight that interacts with large input values will have a greater impact on the output than a weight that sees only small inputs. Two weights might differ by the same absolute amount from their quantized versions, but if one multiplies inputs that are ten times larger, it creates ten times more output distortion. The Frobenius norm aggregates all of these output differences into a single scalar that we can minimize.

The reconstruction error can be rewritten in a form that reveals important mathematical structure. This reformulation exposes the role of input correlations in determining which weights matter most. Expanding the Frobenius norm using the identity ∣∣A∣∣F2=Tr(AAT)||\mathbf{A}||_F^2 = \text{Tr}(\mathbf{A}\mathbf{A}^T):

L(W^)=∣∣(W−W^)X∣∣F2=Tr[((W−W^)X)((W−W^)X)T](apply identity: ∣∣A∣∣F2=Tr(AAT))=Tr[(W−W^)XXT(W−W^)T](expand transpose: (AB)T=BTAT)\begin{aligned} \mathcal{L}(\hat{\mathbf{W}}) &= ||(\mathbf{W} - \hat{\mathbf{W}})\mathbf{X}||_F^2 \\ &= \text{Tr}\left[ ((\mathbf{W} - \hat{\mathbf{W}})\mathbf{X}) ((\mathbf{W} - \hat{\mathbf{W}})\mathbf{X})^T \right] && \text{(apply identity: } ||\mathbf{A}||_F^2 = \text{Tr}(\mathbf{A}\mathbf{A}^T)\text{)} \\ &= \text{Tr}\left[(\mathbf{W} - \hat{\mathbf{W}})\mathbf{X}\mathbf{X}^T(\mathbf{W} - \hat{\mathbf{W}})^T\right] && \text{(expand transpose: } (\mathbf{AB})^T = \mathbf{B}^T\mathbf{A}^T\text{)} \end{aligned}

where:

  • Tr[⋅]\text{Tr}[\cdot]: the trace operator, which sums the diagonal elements of a square matrix
  • (W−W^)(\mathbf{W} - \hat{\mathbf{W}}): the weight error matrix, capturing how much each weight was shifted by quantization
  • XT\mathbf{X}^T: the transpose of the input matrix

The trace operation sums the diagonal elements of the resulting matrix. What emerges from this manipulation is that the loss depends on the weight errors not in isolation, but as weighted by the matrix XXT\mathbf{X}\mathbf{X}^T. This matrix captures how the inputs correlate with each other: if two input dimensions tend to activate together, errors in their corresponding weights will interact. The matrix XXT\mathbf{X}\mathbf{X}^T is in effect the unnormalized sample covariance of the layer's inputs, computed across all calibration tokens.

Define H=2XXT\mathbf{H} = 2\mathbf{X}\mathbf{X}^T, which we call the Hessian matrix (we'll explain why this name is appropriate in the next section). The loss becomes:

L(W^)=12Tr[(W−W^)H(W−W^)T]\mathcal{L}(\hat{\mathbf{W}}) = \frac{1}{2}\text{Tr}\left[(\mathbf{W} - \hat{\mathbf{W}})\mathbf{H}(\mathbf{W} - \hat{\mathbf{W}})^T\right]

where:

  • Tr[⋅]\text{Tr}[\cdot]: the trace operator
  • (W−W^)(\mathbf{W} - \hat{\mathbf{W}}): the weight error matrix
  • H\mathbf{H}: the Hessian matrix (2XXT2\mathbf{X}\mathbf{X}^T) that captures input correlations
  • 12\frac{1}{2}: a scaling factor arising from the definition of H\mathbf{H}

This quadratic form indicates that the loss is a weighted sum of squared errors, where the weighting comes from the Hessian matrix. Errors in weights that correspond to highly active or correlated inputs are penalized more heavily than errors in weights that rarely contribute to the output. The quadratic structure is not merely convenient notation: it is what lets GPTQ to compute closed-form solutions for optimal corrections rather than relying on iterative gradient descent.

Since the rows of W\mathbf{W} don't interact in this expression (each output dimension is computed independently as a dot product with the inputs), we can optimize each row separately. This further simplifies our problem: instead of optimizing over all dout×dind_{out} \times d_{in} weights simultaneously, we can solve doutd_{out} independent problems, each involving only dind_{in} weights. For a single row w∈Rdin\mathbf{w} \in \mathbb{R}^{d_{in}}:

L(w^)=12(w−w^)H(w−w^)T\mathcal{L}(\hat{\mathbf{w}}) = \frac{1}{2}(\mathbf{w} - \hat{\mathbf{w}})\mathbf{H}(\mathbf{w} - \hat{\mathbf{w}})^T

where:

  • w\mathbf{w}: a single row of the original weight matrix (one output neuron's incoming weights)
  • w^\hat{\mathbf{w}}: the corresponding row of quantized weights
  • H\mathbf{H}: the Hessian matrix, shared across all rows because all output neurons see the same inputs
  • (w−w^)(\mathbf{w} - \hat{\mathbf{w}}): the error vector for this row

This weighted squared error gives us the foundation for GPTQ's optimization strategy. The quadratic structure of this loss function is important because it lets closed-form solutions for optimal weight updates. We do not need to run gradient descent or perform any iterative optimization; we can compute the exact optimal correction in a single step, which is what makes GPTQ both theoretically clean and computationally efficient.

The Hessian and Its Role

The matrix H=2XXT\mathbf{H} = 2\mathbf{X}\mathbf{X}^T is called the Hessian because it equals the second derivative of the reconstruction loss with respect to the weights. This naming shows a deep connection to optimization theory. The Hessian matrix characterizes the local curvature of the loss landscape, telling us how rapidly the loss changes as we move in different directions through weight space. High curvature in a particular direction means that small weight changes in that direction cause large changes in the loss; low curvature means the loss is relatively insensitive to perturbations in that direction.

Think of the Hessian as a topographic map of the loss landscape around the current weights. A standard map shows elevation as a function of position; the Hessian shows how steeply the loss rises in every direction from the current point. Directions with high curvature correspond to "steep valleys" where we must be precise. Directions with low curvature correspond to "flat plains" where we have freedom to move without much consequence. GPTQ uses this map to prioritize accuracy in the steep valleys while accepting more error on the flat plains.

To see why this matrix is second derivatives, consider the loss for a single output neuron. We want to find weights that make the neuron's output match what it would produce with the original weights:

L(w)=∑i=1n(w⋅xi−yi)2\mathcal{L}(\mathbf{w}) = \sum_{i=1}^{n} (\mathbf{w} \cdot \mathbf{x}_i - y_i)^2

where:

  • w\mathbf{w}: the weight vector being optimized
  • xi\mathbf{x}_i: the ii-th input vector from the calibration set
  • yiy_i: the target output scalar, computed using the original weights w∗\mathbf{w}^* applied to xi\mathbf{x}_i
  • nn: the total number of calibration tokens

This is a standard least-squares objective. To find its minimum, we compute the gradient by differentiating with respect to each weight component:

∂L∂w=2∑i=1n(w⋅xi−yi)xi\frac{\partial \mathcal{L}}{\partial \mathbf{w}} = 2\sum_{i=1}^{n} (\mathbf{w} \cdot \mathbf{x}_i - y_i)\mathbf{x}_i

where:

  • ∂L∂w\frac{\partial \mathcal{L}}{\partial \mathbf{w}}: the gradient vector of the loss with respect to the weights
  • (w⋅xi−yi)(\mathbf{w} \cdot \mathbf{x}_i - y_i): the prediction error for the ii-th calibration sample
  • xi\mathbf{x}_i: the ii-th input vector, which scales how each prediction error affects each weight

Taking the derivative once more, we obtain the Hessian, which tells us how the gradient itself changes as we modify the weights:

H=∂2L∂w2=2∑i=1nxixiT=2XXT\mathbf{H} = \frac{\partial^2 \mathcal{L}}{\partial \mathbf{w}^2} = 2\sum_{i=1}^{n} \mathbf{x}_i \mathbf{x}_i^T = 2\mathbf{X}\mathbf{X}^T

where:

  • H\mathbf{H}: the Hessian matrix (second-order derivative of the loss)
  • ∂2L∂w2\frac{\partial^2 \mathcal{L}}{\partial \mathbf{w}^2}: the matrix of all second partial derivatives
  • xixiT\mathbf{x}_i \mathbf{x}_i^T: the outer product of each input with itself, a din×dind_{in} \times d_{in} matrix
  • X\mathbf{X}: the matrix of all calibration inputs stacked as columns

The Hessian captures the curvature of the loss landscape around the current weights. Each element of this matrix has a specific interpretation. The diagonal elements HqqH_{qq} indicate how sensitive the loss is to changes in weight wqw_q: a large value means that small perturbations to this weight cause large changes in the output, making this weight "important" in the sense that we must quantize it carefully. The off-diagonal elements HqjH_{qj} capture interactions between weights: they tell us how changes in one weight affect the optimal value of another. When two weights have a large off-diagonal element, their quantization errors are not independent, and an error in one can be partially offset by adjusting the other.

Hessian as Input Statistics

The Hessian H=2XXT\mathbf{H} = 2\mathbf{X}\mathbf{X}^T is simply a scaled version of the unnormalized sample covariance matrix of the layer's inputs. We don't need access to labels or backpropagation to compute it. We just need to observe what inputs flow through the layer during a forward pass on calibration data. This means computing the Hessian is cheap: a single forward pass suffices. The fact that a second-order quantity, which normally requires expensive Hessian-vector products, can be computed from raw activations is what makes GPTQ practical at scale.

The key insight here is that the Hessian tells us which weights are important and how their importance relates to the actual data distribution. A weight connected to an input feature that never activates has zero curvature; no amount of error there will hurt the model. A weight connected to a feature that activates strongly and frequently has high curvature; errors there propagate directly into output degradation. GPTQ uses these curvature measurements as a lens through which every quantization decision is filtered.

This observation has significant practical implications. Computing the Hessian requires only a forward pass through the network on calibration data. We never need to differentiate through the model or compute any target labels. The Hessian emerges naturally from the statistics of the inputs that the layer observes during normal operation. This makes GPTQ a true post-training quantization method: we take a pre-trained model, run it on some representative data, and use the resulting statistics to guide quantization. No fine-tuning, no gradient computation through the model, and no labeled data are required.

Optimal Brain Quantization

GPTQ builds on a framework called Optimal Brain Quantization (OBQ), which itself descends from classical work on neural network pruning. The historical connection is illuminating: pruning and quantization are closely related problems. In pruning, we set certain weights exactly to zero. In quantization, we round weights to the nearest value in a discrete set. Both operations introduce error, and in both cases we want to minimize the impact on the network's output. The mathematics that describes the optimal pruning correction turns out to describe the optimal quantization correction as well.

The core insight of OBQ is that when you quantize one weight, you can compute the optimal adjustment to all remaining weights that minimizes the resulting increase in loss. This is not a heuristic or an approximation: given the quadratic structure of our loss function, there exists a closed-form formula for the best possible compensation. The formula emerges from setting the derivative of the loss (after fixing one weight to its quantized value) to zero with respect to the remaining weights and solving analytically.

Think of OBQ as a chess player who, after being forced to sacrifice a piece, immediately recalculates the optimal position for all remaining pieces. The sacrifice introduces a loss, but the player compensates by repositioning everything else to minimize the damage. In GPTQ, the "sacrifice" is the rounding error introduced when a weight is quantized, and the "repositioning" is the adjustment applied to all remaining weights.

Suppose we have decided to quantize weight wqw_q to value w^q\hat{w}_q. We want to adjust the remaining weights wF\mathbf{w}_F (where FF denotes the set of weights not yet quantized) to minimize the resulting loss. Taking the derivative of the loss with respect to δwF\delta\mathbf{w}_F and setting it to zero yields the optimal adjustment:

δwF=−wq−w^q[H−1]qq⋅[H−1]F,q\delta\mathbf{w}_F = -\frac{w_q - \hat{w}_q}{[\mathbf{H}^{-1}]_{qq}} \cdot [\mathbf{H}^{-1}]_{F,q}

where:

  • δwF\delta\mathbf{w}_F: the optimal adjustment vector for the remaining unquantized weights
  • wq−w^qw_q - \hat{w}_q: the quantization error for the current weight qq, the difference between what we wanted and what we got after rounding
  • [H−1]qq[\mathbf{H}^{-1}]_{qq}: the diagonal element of the inverse Hessian corresponding to weight qq
  • [H−1]F,q[\mathbf{H}^{-1}]_{F,q}: the elements of the qq-th column of the inverse Hessian for all remaining weights FF

Let us unpack this formula carefully to build intuition for what each component contributes. The numerator wq−w^qw_q - \hat{w}_q is the quantization error for weight qq: simply the difference between what we wanted and what we got after rounding. Larger errors require larger compensations.

The denominator [H−1]qq[\mathbf{H}^{-1}]_{qq} is the inverse curvature, measuring how flexible the loss landscape is with respect to weight qq. When this value is large, the curvature of the loss with respect to wqw_q is small, meaning errors in this weight are relatively easy to absorb through small adjustments elsewhere. When this value is small, the curvature is high, meaning errors in this weight are costly and difficult to compensate.

The vector [H−1]F,q[\mathbf{H}^{-1}]_{F,q} determines how this error should be distributed across remaining weights. It acts like a routing map: weights that are strongly correlated with wqw_q through the inverse Hessian receive larger adjustments, because adjusting them has the greatest effect on undoing the damage caused by the rounding error in wqw_q.

The resulting increase in loss from quantizing weight qq (after optimal compensation) is:

ΔLq=(wq−w^q)22[H−1]qq\Delta\mathcal{L}_q = \frac{(w_q - \hat{w}_q)^2}{2[\mathbf{H}^{-1}]_{qq}}

where:

  • ΔLq\Delta\mathcal{L}_q: the unavoidable increase in reconstruction error caused by quantizing weight qq
  • (wq−w^q)2(w_q - \hat{w}_q)^2: the squared quantization error (larger errors cost more)
  • [H−1]qq[\mathbf{H}^{-1}]_{qq}: the inverse curvature, representing the "stiffness" of weight qq; smaller values mean more damage per unit of rounding error

This formula tells us exactly how much each quantization decision costs, accounting for optimal compensation. It has an elegant structure: cost is proportional to the squared rounding error and inversely proportional to the inverse curvature. Weights where [H−1]qq[\mathbf{H}^{-1}]_{qq} is large can be quantized with relatively low cost even if the rounding error is substantial. Conversely, weights with small [H−1]qq[\mathbf{H}^{-1}]_{qq} values are stiff: even small errors cause significant loss increases. This cost formula is exactly what OBQ uses to decide which weight to quantize next, always choosing the one with minimum ΔLq\Delta\mathcal{L}_q.

Worked Example: Error Compensation in Three Weights

To make the OBQ compensation formula concrete, let us walk through a tiny numerical example with just three weights. This example strips away all the complexity to show exactly how error compensation works step by step.

Suppose a row has three weights w=[0.7,0.3,−0.5]\mathbf{w} = [0.7, 0.3, -0.5] and the calibration data gives the following inverse Hessian:

H−1=(2.00.5−0.30.51.50.2−0.30.21.0)\mathbf{H}^{-1} = \begin{pmatrix} 2.0 & 0.5 & -0.3 \\ 0.5 & 1.5 & 0.2 \\ -0.3 & 0.2 & 1.0 \end{pmatrix}

We use 4-bit symmetric quantization with scale s=0.1s = 0.1 (so quantized values are integer multiples of 0.10.1 in the range [−0.8,0.7][-0.8, 0.7]).

Step 1: Quantize w0=0.7w_0 = 0.7.

Rounding 0.70.7 to the nearest multiple of 0.10.1 gives w^0=0.7\hat{w}_0 = 0.7. The error is e0=0.7−0.7=0.0e_0 = 0.7 - 0.7 = 0.0. Because the error is zero, no compensation is needed. The other weights remain unchanged: w=[0.7,0.3,−0.5]\mathbf{w} = [0.7, 0.3, -0.5].

Step 2: Quantize w1=0.3w_1 = 0.3.

Rounding 0.30.3 gives w^1=0.3\hat{w}_1 = 0.3. Again the error is zero. Still no compensation needed.

Step 3: Quantize w2=−0.5w_2 = -0.5.

Rounding −0.5-0.5 gives w^2=−0.5\hat{w}_2 = -0.5. Error is zero. This example happened to be perfectly aligned with the grid.

Let us now change the example slightly to see real compensation. Suppose instead w=[0.73,0.28,−0.53]\mathbf{w} = [0.73, 0.28, -0.53].

Step 1: Quantize w0=0.73w_0 = 0.73.

Rounding to the nearest 0.10.1 gives w^0=0.7\hat{w}_0 = 0.7. The error is e0=0.73−0.70=0.03e_0 = 0.73 - 0.70 = 0.03.

The optimal compensation for the remaining weights {w1,w2}\{w_1, w_2\} is:

δw{1,2}=−e0[H−1]00⋅[H−1]{1,2},0=−0.032.0⋅(0.5−0.3)=(−0.00750.0045)\delta\mathbf{w}_{\{1,2\}} = -\frac{e_0}{[\mathbf{H}^{-1}]_{00}} \cdot [\mathbf{H}^{-1}]_{\{1,2\},0} = -\frac{0.03}{2.0} \cdot \begin{pmatrix} 0.5 \\ -0.3 \end{pmatrix} = \begin{pmatrix} -0.0075 \\ 0.0045 \end{pmatrix}

After compensation, the adjusted weights are w1=0.28−0.0075=0.2725w_1 = 0.28 - 0.0075 = 0.2725 and w2=−0.53+0.0045=−0.5255w_2 = -0.53 + 0.0045 = -0.5255.

Step 2: Quantize adjusted w1=0.2725w_1 = 0.2725.

Rounding to the nearest 0.10.1 gives w^1=0.3\hat{w}_1 = 0.3. The error is e1=0.2725−0.30=−0.0275e_1 = 0.2725 - 0.30 = -0.0275.

The optimal compensation for {w2}\{w_2\} (the only remaining weight) is:

δw2=−−0.0275[H−1]11⋅[H−1]2,1=−−0.02751.5⋅0.2=0.00367\delta w_2 = -\frac{-0.0275}{[\mathbf{H}^{-1}]_{11}} \cdot [\mathbf{H}^{-1}]_{2,1} = -\frac{-0.0275}{1.5} \cdot 0.2 = 0.00367

After compensation, w2=−0.5255+0.00367=−0.52183w_2 = -0.5255 + 0.00367 = -0.52183.

Step 3: Quantize adjusted w2=−0.52183w_2 = -0.52183.

Rounding gives w^2=−0.5\hat{w}_2 = -0.5. The error is e2=−0.52183−(−0.50)=−0.02183e_2 = -0.52183 - (-0.50) = -0.02183.

No compensation is needed because there are no remaining weights.

Comparing the two approaches. With naive round-to-nearest, the quantized weights would be w^=[0.7,0.3,−0.5]\hat{\mathbf{w}} = [0.7, 0.3, -0.5] with total squared weight error (0.03)2+(0.02)2+(0.03)2=0.0022(0.03)^2 + (0.02)^2 + (0.03)^2 = 0.0022. With GPTQ, the quantized weights are [0.7,0.3,−0.5][0.7, 0.3, -0.5] with the same integer representations, but GPTQ's intermediate adjustments ensured that each rounding decision accounted for the correlations between weights in the Hessian. In a more complex example with finer grid spacing or larger off-diagonal Hessian elements, GPTQ would choose different integer values than naive rounding for weights w1w_1 and w2w_2 precisely because the compensation shifted their effective values before rounding. The key point is that the compensation changes the value of each weight before it is rounded, so later weights end up being rounded to different and better integer codes than they would have been under naive quantization.

The GPTQ Algorithm

The naive OBQ approach would quantize weights one at a time, choosing at each step the weight whose quantization causes the smallest loss increase ΔLq\Delta\mathcal{L}_q. While this greedy strategy seems sensible, its computational cost is prohibitive. Each step requires examining all remaining weights to find the best candidate, and after quantizing each weight, we must update the inverse Hessian. This requires O(din2)O(d_{in}^2) operations per weight, yielding O(din3)O(d_{in}^3) complexity per row and O(din4)O(d_{in}^4) for the entire layer. For a model where dind_{in} can be 4096 or larger, O(din4)O(d_{in}^4) operations are completely impractical.

GPTQ makes three key modifications that reduce the total complexity to O(din3)O(d_{in}^3) while maintaining nearly identical accuracy. The central observation is that the optimal quantization order matters less than one might expect. The compensation mechanism is powerful enough that processing weights in a fixed sequential order yields results almost as good as the optimal greedy order. What matters most is performing the compensation correctly after each quantization step.

Think of the three optimizations as different aspects of making the same core algorithm more efficient, like optimizing a factory assembly line. Column-wise processing batches the work across all output neurons simultaneously (vectorization). Lazy batch updates reduce the number of times we touch memory by accumulating changes before applying them (cache efficiency). Cholesky-based inverse updates replace a naive matrix inversion with a factorization that can be maintained cheaply as we process each column (algorithmic efficiency).

Column-wise processing. Instead of quantizing weights one at a time in optimal order, GPTQ processes all weights in a fixed left-to-right column order. The important insight is that all rows of the weight matrix share the same Hessian, because all output neurons see the same inputs. This means we can process all rows simultaneously for a given column, quantizing column qq of every output neuron at once. This vectorization turns a sequential loop over rows into a single matrix operation, achieving a large constant-factor speedup.

Lazy batch updates. Rather than updating all remaining weights after each individual column quantization, GPTQ accumulates updates in blocks (typically of size 128) and applies them all at once when the block is complete. This improves cache efficiency sharply because memory access patterns become more localized. Weight matrix rows are stored contiguously in memory, and processing a full block before writing updates means each memory location is accessed sequentially rather than randomly.

Cholesky-based inverse updates. Computing and maintaining the full inverse Hessian naively would cost O(din2)O(d_{in}^2) per column after the initial O(din3)O(d_{in}^3) inversion, yielding O(din3)O(d_{in}^3) total work just for the updates. GPTQ instead precomputes the Cholesky decomposition H−1=LLT\mathbf{H}^{-1} = \mathbf{L}\mathbf{L}^T where L\mathbf{L} is lower triangular. The triangular structure of L\mathbf{L} directly encodes the sequential dependencies between weights: the qq-th column of L\mathbf{L} gives the compensation coefficients needed when processing column qq. After processing column qq, we simply move to the next column of L\mathbf{L} without any additional update. This reduces the per-column cost from O(din2)O(d_{in}^2) to O(din)O(d_{in}) after the initial factorization.

Algorithm Steps

Here is the complete GPTQ procedure in detail:

  1. Collect calibration data. Run a small set of representative examples through the model, using forward hooks to record each layer's inputs X\mathbf{X} as the data flows through.

  2. Compute the Hessian. For each layer, compute H=2XXT\mathbf{H} = 2\mathbf{X}\mathbf{X}^T and add a small diagonal dampening term λI\lambda \mathbf{I} for numerical stability. Compute the inverse H−1\mathbf{H}^{-1} via Cholesky decomposition.

  3. Process each row of the weight matrix. For each row w\mathbf{w}:

    a. For each column qq from 0 to din−1d_{in} - 1:

    • Quantize: w^q=quantize(wq)\hat{w}_q = \text{quantize}(w_q)
    • Compute error: δq=wq−w^q\delta_q = w_q - \hat{w}_q
    • Update remaining weights: wq+1:←wq+1:−δq[H−1]qq⋅[H−1]q+1:,qw_{q+1:} \leftarrow w_{q+1:} - \frac{\delta_q}{[\mathbf{H}^{-1}]_{qq}} \cdot [\mathbf{H}^{-1}]_{q+1:,q}

    b. Store the quantized row w^\hat{\mathbf{w}} along with per-group scale and zero-point parameters.

  4. Replace the layer's weights with their quantized integer values plus the associated dequantization metadata.

The order in which columns are processed affects accuracy. The default left-to-right order works reasonably well, but ActOrder (processing columns in decreasing order of Hessian diagonal) often improves results by quantizing the most important weights first, when there are still many remaining weights available to absorb compensation. We will explore ActOrder in detail later in this chapter.

Mathematical Details of the Update

The inverse Hessian update is a necessary mathematical component that lets efficient processing. After quantizing column qq, we need to update the inverse Hessian to reflect that column qq is no longer a free variable. The weight at column qq has been fixed to its quantized value; it can no longer be adjusted to compensate for future quantization errors. Mathematically, we need to compute a new inverse Hessian that covers only the remaining unquantized weights F∖{q}F \setminus \{q\}.

Let HF−1\mathbf{H}_F^{-1} denote the inverse Hessian restricted to the remaining (unquantized) weights. After removing column and row qq, the new inverse is given by the Schur complement formula:

[HF∖q−1]ij=[HF−1]ij−[HF−1]iq[HF−1]qj[HF−1]qq[\mathbf{H}_{F \setminus q}^{-1}]_{ij} = [\mathbf{H}_F^{-1}]_{ij} - \frac{[\mathbf{H}_F^{-1}]_{iq} [\mathbf{H}_F^{-1}]_{qj}}{[\mathbf{H}_F^{-1}]_{qq}}

where:

  • [HF∖q−1]ij[\mathbf{H}_{F \setminus q}^{-1}]_{ij}: the updated inverse Hessian element at row ii, column jj, after removing weight qq
  • [HF−1]ij[\mathbf{H}_F^{-1}]_{ij}: the current inverse Hessian element before the update
  • [HF−1]iq,[HF−1]qj[\mathbf{H}_F^{-1}]_{iq}, [\mathbf{H}_F^{-1}]_{qj}: elements from the qq-th column and row of the current inverse Hessian
  • [HF−1]qq[\mathbf{H}_F^{-1}]_{qq}: the diagonal element corresponding to weight qq
  • i,ji, j: indices ranging over the remaining unquantized weights in F∖{q}F \setminus \{q\}

This formula is a rank-one update, subtracting an outer product that corresponds to removing the degree of freedom at position qq. The intuition is that by fixing weight qq, we eliminate one dimension from our optimization space. The remaining weights must "account for" the fact that they can no longer rely on wqw_q to absorb any future compensation. The Schur complement formula computes exactly how the correlations between remaining weights change when one weight is removed from consideration.

Computing this update naively as a dense matrix operation would cost O(d2)O(d^2) per column. Over all dd columns, this would total O(d3)O(d^3), which is the same as just re-inverting the Hessian from scratch each time. The Cholesky decomposition gives the escape from this cost.

The Cholesky factorization gives us H−1=LLT\mathbf{H}^{-1} = \mathbf{L}\mathbf{L}^T where L\mathbf{L} is lower triangular. This factorization is particularly useful because the triangular structure directly encodes the sequential dependencies between weights. For column qq, the required compensation coefficients [H−1]q+1:,q[\mathbf{H}^{-1}]_{q+1:,q} are exactly the entries in the qq-th column of L\mathbf{L} (below the diagonal). The update formula simplifies to reading off the next column of L\mathbf{L} rather than performing a full matrix update, reducing the per-column cost to O(d)O(d) after the initial O(d3)O(d^3) Cholesky factorization. This is why the total complexity of GPTQ is O(din3)O(d_{in}^3) per layer: dominated by the one-time Cholesky factorization, with all subsequent operations being linear in dind_{in}.

Calibration Data

GPTQ requires calibration data to estimate the Hessian. The quality and quantity of this data affect quantization results, yet practitioners often treat calibration data as an afterthought. Calibration is a significant design decision.

The calibration data does not need labels. GPTQ only performs forward passes through the model to collect layer activations, so any unlabeled text will do. What matters is that the calibration data covers the kinds of inputs the model will encounter at deployment. The Hessian is an empirical estimate of input statistics: run calibration on text that looks nothing like your deployment distribution, and the Hessian will mis-estimate which weights matter, leading to worse quantization choices.

Think of the calibration data as a "map of typical inputs." If you are quantizing a general-purpose language model, a diverse sample that combines web and book text with source code will give the Hessian a realistic picture of the model's inputs. If you are quantizing a specialized code-generation model, calibrating on web text means the Hessian will reflect a distribution very different from the model's actual use, and the quantization will be suboptimal for code generation tasks even if it looks fine on general language benchmarks.

The key parameters for calibration are:

  • Quantity. GPTQ typically uses 128 to 1,024 samples. More samples give a better Hessian estimate, reducing variance in the diagonal and off-diagonal elements. In practice, 128 samples of 2,048 tokens each (totaling 262,144 tokens) is sufficient for most models, with diminishing returns beyond 512 samples.

  • Representativeness. The calibration data should resemble the data the model will see at inference. Using random text to calibrate a code model yields worse results than using code samples, because the Hessian will reflect different activation patterns.

  • Sequence length. Longer sequences give more tokens per sample, leading to a more accurate Hessian estimate. Most implementations use 2,048 tokens per sequence.

Common choices include random samples from C4 (a large web text corpus), WikiText (cleaned Wikipedia text), or domain-specific data for specialized models. The C4 and WikiText defaults work well for general-purpose models. For specialized deployments, investing in domain-specific calibration data consistently improves results.

Group Quantization

Standard per-tensor or per-channel quantization uses a single scale and zero point for many weights. For weights with non-uniform distributions within a channel, a single scale factor forces the quantization grid to accommodate the full dynamic range, wasting precision on weights that are clustered in a small portion of that range. GPTQ often employs group quantization to address this limitation.

In group quantization, weights are divided into contiguous groups (commonly 128 weights each) that each have their own scale and zero-point parameters. This is a middle ground between per-tensor quantization (cheapest, least accurate) and per-weight quantization (most expensive, most accurate). The overhead of storing group-wise quantization parameters is small relative to the accuracy gains.

Think of group quantization as using local rulers instead of a single global ruler to measure distances. If you have a collection of objects ranging from 1 mm to 1 m in size, a single ruler calibrated to handle the full range will be very imprecise for the small objects. Multiple rulers, each calibrated for a smaller range, will be more precise throughout. Group quantization applies this same principle to weight distributions.

Group quantization gives a middle ground between accuracy and overhead:

  • Per-tensor quantization. One scale for all weights in the layer. Minimum storage overhead but potentially high quantization error if weight distributions vary across channels or weight groups.

  • Per-group quantization. Separate scales for groups of 128 weights. The overhead is small: one FP16 scale per 128 INT4 weights adds 16 bits for every 4 ×\times 128 = 512 bits of weight data, which is 0.125 bits per weight, increasing the effective bit-width from 4.0 to 4.125 bits. This tiny overhead often yields substantially better accuracy.

  • Per-weight quantization. Individual scales for each weight. Maximum accuracy but impractical overhead: storing one FP16 scale per INT4 weight doubles the memory usage, eliminating the benefit of quantization entirely.

With group size 128 and 4-bit weights, the effective storage is approximately 4.125 bits per weight, a tiny overhead for significant accuracy gains. In practice, group sizes of 64 to 128 offer the best tradeoff, with group size 128 being the most common choice in deployed systems.

Implementation with AutoGPTQ

Let us see GPTQ in practice. We will start by implementing the core algorithm from scratch to build understanding, then show how to use the AutoGPTQ library for production use.

First, we load a small model to examine before quantization:

In[5]:
Code
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load a small model for demonstration
model_name = "facebook/opt-125m"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.float32
)

Let us examine the original model's memory footprint to establish a baseline:

In[6]:
Code
def count_parameters(model):
    """Count total parameters and calculate memory in FP16."""
    total_params = sum(p.numel() for p in model.parameters())
    memory_fp16_mb = total_params * 2 / (1024**2)  # 2 bytes per FP16
    return total_params, memory_fp16_mb


total_params, memory_mb = count_parameters(model)
int4_memory_mb = memory_mb / 4
Out[7]:
Console
Total parameters: 125,239,296
FP16 memory: 238.88 MB
Estimated INT4 memory: 59.72 MB

The output shows that 4-bit quantization reduces the memory footprint to a quarter of its FP16 size. For this 125M parameter model, this means dropping from over 200 MB to roughly 60 MB. For a 7B parameter model, the same ratio shrinks 14 GB to 3.5 GB, making the model runnable on a laptop GPU. For a 65B model, it means fitting 130 GB into about 35 GB on a single high-end consumer GPU.

Now let us implement the core GPTQ components. We start with the Hessian computation:

In[8]:
Code
def compute_hessian(inputs: torch.Tensor) -> torch.Tensor:
    """
    Compute the Hessian matrix H = 2 * X @ X^T for GPTQ.

    Args:
        inputs: Layer inputs of shape (n_samples, seq_len, hidden_dim)
                or (n_tokens, hidden_dim)
    Returns:
        Hessian matrix of shape (hidden_dim, hidden_dim)
    """
    # Flatten to (n_tokens, hidden_dim)
    if inputs.dim() == 3:
        inputs = inputs.reshape(-1, inputs.shape[-1])

    # H = 2 * X^T @ X (note: we transpose so each row is a feature)
    H = 2 * inputs.T @ inputs

    # Add small diagonal for numerical stability
    H += 1e-4 * torch.eye(H.shape[0], device=H.device, dtype=H.dtype)

    return H

The 1e-4 diagonal dampening prevents numerical issues when the Hessian is nearly singular, which can happen when some input features are nearly linearly dependent or rarely active. In practice, the dampening factor is typically set as a fraction of the mean Hessian diagonal (0.1% to 1%), but the absolute value works well for our demonstration.

Next, we implement the quantization and dequantization primitives:

In[9]:
Code
def quantize_weight(
    weight: float, scale: float, zero_point: int, n_bits: int = 4
) -> int:
    """Quantize a single weight value to n-bit integer."""
    qmin, qmax = 0, (2**n_bits) - 1
    q = round(weight / scale + zero_point)
    return max(qmin, min(qmax, q))


def dequantize_weight(q: int, scale: float, zero_point: int) -> float:
    """Convert quantized integer back to float."""
    return scale * (q - zero_point)

Here is the core GPTQ algorithm for a single row of weights. The key loop quantizes each column, computes the error, and immediately applies the compensation to all remaining weights:

In[10]:
Code
def gptq_quantize_row(
    weights: torch.Tensor,  # Shape: (d_in,)
    H_inv: torch.Tensor,  # Inverse Hessian, shape: (d_in, d_in)
    n_bits: int = 4,
    group_size: int = 128,
) -> tuple[torch.Tensor, list]:
    """
    Apply GPTQ quantization to a single row of weights.

    Returns:
        Tuple of (quantized_weights, quantization_params)
    """
    d_in = weights.shape[0]
    weights = weights.clone().float()
    quantized = torch.zeros_like(weights, dtype=torch.int8)
    params = []  # Store (scale, zero_point) for each group

    # Process columns left to right
    for col in range(d_in):
        # Determine group for this column
        group_idx = col // group_size
        group_start = group_idx * group_size
        group_end = min(group_start + group_size, d_in)

        # Compute scale and zero point for this group (if at group boundary)
        if col == group_start:
            group_weights = weights[group_start:group_end]
            w_min, w_max = (
                group_weights.min().item(),
                group_weights.max().item(),
            )

            # Asymmetric quantization
            qmin, qmax = 0, (2**n_bits) - 1
            scale = (w_max - w_min) / (qmax - qmin) if w_max > w_min else 1.0
            zero_point = round(-w_min / scale) if scale > 0 else 0
            zero_point = max(qmin, min(qmax, zero_point))
            params.append((scale, zero_point))

        scale, zero_point = params[group_idx]

        # Quantize current weight
        w = weights[col].item()
        q = quantize_weight(w, scale, zero_point, n_bits)
        quantized[col] = q

        # Compute quantization error
        w_hat = dequantize_weight(q, scale, zero_point)
        error = w - w_hat

        # Update remaining weights to compensate for error
        if col < d_in - 1 and H_inv[col, col] > 1e-10:
            # Optimal update: delta_w_F = -error / H_inv[q,q] * H_inv[F,q]
            update = -error / H_inv[col, col] * H_inv[col + 1 :, col]
            weights[col + 1 :] += update

    return quantized, params

Notice how the compensation update modifies weights[col + 1:] in place. Each subsequent weight is adjusted before it is rounded, so the rounding decision for column q+1q+1 is made on the corrected value rather than the original. This sequential correction is the essence of GPTQ: every quantization step immediately propagates its error into the remaining weights, giving them a chance to compensate.

Let us show this on a sample weight matrix:

In[11]:
Code
# Create a sample weight matrix and random inputs
torch.manual_seed(42)
d_out, d_in = 128, 256
sample_weights = torch.randn(d_out, d_in) * 0.1
sample_inputs = torch.randn(512, d_in)  # 512 calibration tokens

# Compute Hessian and its inverse
H = compute_hessian(sample_inputs)
H_inv = torch.linalg.inv(H)
In[12]:
Code
# Quantize one row as demonstration
original_row = sample_weights[0]
quantized_row, quant_params = gptq_quantize_row(original_row, H_inv)

# Dequantize to compare
dequantized_row = torch.zeros_like(original_row)
group_size = 128
for col in range(d_in):
    group_idx = col // group_size
    scale, zero_point = quant_params[group_idx]
    dequantized_row[col] = dequantize_weight(
        quantized_row[col].item(), scale, zero_point
    )

# Compute reconstruction error
mse = ((original_row - dequantized_row) ** 2).mean().item()

# Compare with naive round-to-nearest
naive_scale = (original_row.max() - original_row.min()) / 15
naive_zp = round(-original_row.min().item() / naive_scale.item())
naive_q = torch.round(original_row / naive_scale + naive_zp).clamp(0, 15)
naive_deq = naive_scale * (naive_q - naive_zp)
naive_mse = ((original_row - naive_deq) ** 2).mean().item()
weight_mse_ratio = mse / naive_mse
Out[13]:
Console
GPTQ quantization MSE: 0.000135
Naive quantization MSE: 0.000086
GPTQ raw weight MSE is 1.56x naive MSE

GPTQ does not minimize unweighted error between the original and quantized weights. In this example, its raw weight MSE is higher than naive round-to-nearest. GPTQ instead uses the Hessian to spend the error budget where it has the least effect on layer outputs. The output-space comparison below is therefore the meaningful test of whether compensation worked.

Out[14]:
Visualization
Overlapping histograms of raw weight quantization errors. The orange GPTQ histogram spans a broader range than the blue naive-quantization histogram.
Distribution of raw weight quantization errors for naive and GPTQ methods. GPTQ (orange) permits a broader distribution than naive quantization (blue) because it optimizes Hessian-weighted output reconstruction rather than unweighted weight MSE.

Measuring Output Reconstruction Error

Layer output error is a better measure of quantization quality than weight error. Weight error measures how much the weights changed; output error measures how much the layer's predictions changed. Since the model's downstream behavior depends entirely on layer outputs rather than raw weights, output error is the relevant metric.

We want to verify that GPTQ's weight-space optimization translates into output-space improvements. If it did not, the Hessian-weighted objective would be clever mathematics without practical payoff. In practice, the improvement in output error is even more pronounced than the improvement in weight error, precisely because the Hessian weighting ensures we concentrate accuracy where the outputs are most sensitive.

In[15]:
Code
def measure_output_error(original_weights, quantized_weights, inputs):
    """Compute the output reconstruction error ||W @ X - W_hat @ X||."""
    original_out = inputs @ original_weights.T
    quantized_out = inputs @ quantized_weights.T
    mse = ((original_out - quantized_out) ** 2).mean().item()
    return mse
In[16]:
Code
# Quantize all rows of the sample weight matrix
quantized_weights = torch.zeros_like(sample_weights)

for i in range(d_out):
    q_row, params = gptq_quantize_row(sample_weights[i], H_inv)
    # Dequantize for error measurement
    for col in range(d_in):
        group_idx = col // 128
        scale, zero_point = params[group_idx]
        quantized_weights[i, col] = dequantize_weight(
            q_row[col].item(), scale, zero_point
        )
In[17]:
Code
# Measure output reconstruction error
gptq_output_error = measure_output_error(
    sample_weights, quantized_weights, sample_inputs
)

# Compare with naive quantization
naive_weights = torch.zeros_like(sample_weights)
for i in range(d_out):
    row = sample_weights[i]
    scale = (row.max() - row.min()) / 15
    zp = round(-row.min().item() / scale.item()) if scale > 0 else 0
    q = torch.round(row / scale + zp).clamp(0, 15)
    naive_weights[i] = scale * (q - zp)

naive_output_error = measure_output_error(
    sample_weights, naive_weights, sample_inputs
)
error_ratio = naive_output_error / gptq_output_error
Out[18]:
Console
GPTQ output reconstruction MSE: 0.022085
Naive output reconstruction MSE: 0.030633
GPTQ reaches 1.39x lower output error

The improvement in output reconstruction error is even more pronounced than the improvement in weight error, which is exactly what the Hessian-weighted loss function is designed to produce.

Out[19]:
Visualization
Line chart of per-dimension output MSE for naive and GPTQ quantization across 50 output dimensions. GPTQ line sits below the naive line throughout.
Per-dimension mean squared output reconstruction error for naive quantization versus GPTQ across the first 50 output dimensions. GPTQ consistently produces lower error across all dimensions, with the gap being especially large in dimensions where naive quantization struggles.

Using AutoGPTQ in Practice

For production use, the AutoGPTQ library gives an optimized implementation with GPU acceleration, including CUDA kernels for fast INT4 matrix multiplication during inference. The library handles all the details of Hessian computation, Cholesky factorization, and group quantization, exposing a clean API that takes a model and calibration data as inputs and produces a quantized model ready for deployment.

The workflow has three phases. In the first phase, you prepare calibration data by tokenizing representative text samples and collecting them into a list. In the second phase, you configure the quantization parameters (bits, group size, column ordering) and run the quantization algorithm. In the third phase, you save the quantized model to disk so it can be loaded quickly for inference without repeating the quantization process.

In[31]:
Code
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load model and tokenizer
model_name = "meta-llama/Llama-2-7b-hf"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.float16, device_map="auto"
)

# Prepare calibration data
calibration_data = load_dataset("c4", "en", split="train", streaming=True)
calibration_samples = []
for sample in calibration_data.take(128):
    tokenized = tokenizer(sample["text"], truncation=True, max_length=2048)
    calibration_samples.append(tokenized["input_ids"])

# Configure quantization
quantize_config = BaseQuantizeConfig(
    bits=4,  # 4-bit quantization
    group_size=128,  # Group size for scales
    desc_act=True,  # Enable ActOrder (activation ordering)
    damp_percent=0.01,  # Dampening for numerical stability
)

# Quantize the model
model = AutoGPTQForCausalLM.from_pretrained(model_name, quantize_config)
model.quantize(calibration_samples)

# Save the quantized model
model.save_quantized("llama-2-7b-gptq-4bit")

Key Parameters

Understanding the key AutoGPTQ parameters helps you make informed tradeoffs when quantizing a model for deployment:

  • bits: Target bit width, typically 4 or 8. 4-bit is the most common choice for deployment. It offers the best size reduction. 3-bit is possible but causes more accuracy degradation for most models.
  • group_size: Number of weights sharing quantization parameters. 128 is the standard choice. It offers a good balance between accuracy and overhead. Smaller group sizes (64, 32) improve accuracy at the cost of slightly more storage for scale parameters.
  • desc_act: Whether to use ActOrder (activation ordering), which processes columns by decreasing Hessian diagonal. Enabling this typically improves perplexity by 0.1 to 0.5 points but slightly increases quantization time.
  • damp_percent: Dampening factor added to the Hessian diagonal as a fraction of its mean value. This prevents numerical instability when the Hessian is nearly singular. The default of 0.01 (1%) works well for most models.

ActOrder: Activation-Based Column Ordering

The order in which GPTQ processes columns affects the final quantization error. The default left-to-right order treats all weights equally, but weights differ substantially in their importance as measured by the Hessian diagonal. Processing a highly important weight late in the sequence, after many others have already been quantized and their compensation capacity used up, is suboptimal. That important weight receives compensation from fewer remaining weights than it would have if processed earlier.

ActOrder (activation ordering) addresses this by processing columns in decreasing order of their Hessian diagonal values HqqH_{qq}. The intuition is simple: weights with larger HqqH_{qq} have greater impact on the output and should be quantized first, when the maximum number of other weights are still available to absorb compensation. As important weights are quantized early, their errors are distributed among many remaining weights, each absorbing a small fraction. Less important weights are quantized later, when fewer candidates remain for compensation, but this matters less because their errors have smaller impact on the output anyway.

Think of ActOrder as a surgeon strategy for high-stakes operations. A surgical team handles the most necessary, time-sensitive procedures first, while the available support staff is at its fullest. Routine procedures are handled later, when some staff may have moved on, because the consequences of imperfect execution are smaller.

In[20]:
Code
def get_actorder_permutation(H: torch.Tensor) -> torch.Tensor:
    """
    Compute column ordering based on Hessian diagonal.

    Returns permutation that processes largest diagonal elements first.
    """
    diag = torch.diag(H)
    perm = torch.argsort(diag, descending=True)
    return perm
In[21]:
Code
# Demonstrate ActOrder on our sample Hessian
perm = get_actorder_permutation(H)
first_10 = perm[:10].tolist()
last_10 = perm[-10:].tolist()

# Compare Hessian diagonal values
diag_important = H[perm[0], perm[0]]
diag_least = H[perm[-1], perm[-1]]
Out[22]:
Console
First 10 columns to process (most important): [38, 191, 192, 22, 200, 179, 198, 114, 162, 186]
Last 10 columns to process (least important): [62, 67, 208, 61, 78, 81, 232, 222, 230, 168]

Hessian diagonal for most important column: 1183.9625
Hessian diagonal for least important column: 876.6682

The Hessian diagonal values reveal that the "important" columns have much higher sensitivity (larger values) than the least important ones. Processing these sensitive columns first minimizes the accumulation of error in the output.

Out[23]:
Visualization
model.safetensors: reconstructing file:   0%|          |  0.00B /  251MB            
model.safetensors: downloading bytes:           |  0.00B            
Log-scale line chart of sorted Hessian diagonal values showing a steep initial decline followed by a long low tail across weight indices.
Hessian diagonal values sorted in descending order, plotted on a logarithmic scale. The steep initial drop followed by a long tail is characteristic of realistic transformer activations, where a small fraction of input dimensions carry most of the weight sensitivity. ActOrder exploits this structure by quantizing high-sensitivity weights first.

ActOrder typically improves perplexity by 0.1 to 0.5 points at 4-bit precision, with the benefit more pronounced for smaller models where each individual weight matters more. For large models like LLaMA-65B, the baseline GPTQ without ActOrder already performs well, and the marginal improvement from ActOrder is smaller. For 7B and 13B models, ActOrder is often worth enabling.

Perplexity Comparison

The ultimate test of quantization quality is model performance on downstream tasks. For language models, perplexity on held-out text gives a principled proxy: it measures how well the model predicts unseen text, capturing both fluency and factual accuracy in a single number. Lower perplexity means better predictions.

Perplexity is defined as the exponentiated average negative log-likelihood per token:

PPL=exp⁡(−1N∑i=1Nlog⁡P(xi∣x<i))\text{PPL} = \exp\left(-\frac{1}{N} \sum_{i=1}^{N} \log P(x_i \mid x_{<i})\right)

where:

  • PPL\text{PPL}: perplexity, where lower values indicate better language modeling
  • NN: the total number of tokens in the evaluation set
  • xix_i: the ii-th token in the sequence
  • P(xi∣x<i)P(x_i \mid x_{<i}): the model's predicted probability for token xix_i given all preceding tokens
  • exp⁡(⋅)\exp(\cdot): the exponential function that converts average log-probability into a perplexity score

A perplexity of 5 means the model is, on average, as uncertain as if it had to choose uniformly among 5 equally likely options for each token. A perplexity of 4.5 instead of 5.0 is a real improvement in language understanding. The following table shows typical GPTQ performance for LLaMA models evaluated on WikiText-2:

Perplexity comparison between FP16 and GPTQ 4-bit on WikiText-2.
ModelFP16 PPLGPTQ 4-bit PPLDegradation
LLaMA-7B5.685.85+0.17
LLaMA-13B5.095.20+0.11
LLaMA-30B4.774.84+0.07
LLaMA-65B4.534.58+0.05

The perplexity degradation shrinks with model size, and this is not a coincidence. Larger models have more parameters, which means more redundancy in the weight space. More redundancy means that the compensation mechanism has more "slack" to work with: when one weight is rounded away from its optimal value, there are more other weights available to absorb the error. At 65B parameters, the difference between FP16 and GPTQ 4-bit is barely measurable on most benchmarks. The effective information density of a quantized 65B model exceeds that of a full-precision 7B model, making GPTQ quantization a worthwhile trade at every scale.

The relationship between model size and quantization robustness has an important practical implication: if you must choose between a smaller full-precision model and a larger quantized model of similar memory size, the larger quantized model usually wins on accuracy. A GPTQ 4-bit LLaMA-13B fits in roughly the same GPU memory as an FP16 LLaMA-3B, but performs substantially better on almost every benchmark.

Limitations and Practical Tradeoffs

GPTQ is a major advance over naive quantization, but understanding its limitations is needed for deploying it effectively. Each limitation also suggests where further research has gone, and where you might need to supplement GPTQ with additional techniques.

Calibration sensitivity. Quantization quality depends on how well the calibration data matches the model's deployment inputs. If you calibrate a code-generation model on web text, the Hessian will reflect the statistics of web text activations, not code activations. Code has more structured patterns, longer token dependencies, and different vocabulary distributions. A model quantized with mismatched calibration data will precisely quantize weights that matter for web text while representing weights important for code less accurately. For general-purpose models used on general text, the standard C4 or WikiText calibration works well. For specialized deployments, invest in domain-specific calibration data.

Computational cost. GPTQ quantization is materially slower than naive round-to-nearest. Quantizing a 7B model takes around 10 to 30 minutes on a modern GPU, and the time scales roughly as O(d3)O(d^3) per layer with the Hessian computation and Cholesky factorization. Quantizing a 65B model can take several hours. This is acceptable for one-time quantization, but makes GPTQ unsuitable for anything requiring on-the-fly quantization during inference. You quantize once and deploy; you do not requantize per request. The computational cost is also why GPTQ is usually run on a GPU: CPU-based Cholesky factorization for large hidden dimensions would be impractically slow.

Layer-wise approximation. By optimizing each layer independently, GPTQ ignores how quantization errors compound across layers. In a deep transformer, a small error introduced at layer 1 might be amplified by subsequent attention and FFN computations, reaching layer 20 in a magnified form. GPTQ's per-layer optimization minimizes each layer's individual error without accounting for this cross-layer amplification. This approximation is empirically reasonable (the results in the perplexity table above show this), but it is not exact. Methods that incorporate cross-layer information, while more expensive, can in principle do better.

Activation quantization not included. GPTQ only quantizes weights, not activations. For inference on specialized hardware accelerators that require both weight and activation quantization (such as certain edge devices or custom chips designed for INT4 compute), GPTQ alone is insufficient. Additional quantization-aware training or separate activation quantization techniques are needed to reach full INT4 inference on such hardware.

Outlier sensitivity. The Hessian estimation can be skewed by outlier activations. As discussed in the chapter on INT8 quantization, transformer models sometimes produce extreme activation values in certain dimensions that can be orders of magnitude larger than typical values. These outliers disproportionately influence the Hessian: a few very large activations in one dimension will make the Hessian diagonal for that dimension enormous, leading GPTQ to treat it as critically sensitive and focus excessive precision there, potentially neglecting other dimensions. The dampening factor helps by reducing the relative influence of outliers, but does not eliminate it. Techniques like SmoothQuant, which pre-scales activations to reduce outlier magnitude, can be combined with GPTQ to address this issue.

Despite these limitations, GPTQ remains one of the most effective post-training quantization methods for LLMs. Its combination of theoretical foundation (optimal error compensation from second-order information) and practical efficiency (Cholesky-based linear updates) set the standard for weight quantization. The success of GPTQ inspired subsequent work like AWQ (Activation-aware Weight Quantization), which we will explore in the next chapter. AWQ takes a complementary approach: instead of compensating for quantization errors, it identifies and protects the most important weights from quantization entirely, often achieving even better results on models with strong activation outliers.

Summary

GPTQ turns weight quantization from a simple rounding problem into a principled optimization problem grounded in second-order information theory. By compensating for each weight's quantization error through Hessian-guided updates to remaining weights, GPTQ reaches reconstruction error far below what naive quantization can manage.

The key elements that make GPTQ work are:

  • The layer-wise reconstruction objective. Optimizing the difference between original and quantized layer outputs (rather than weights directly) minimizes the metric relevant to downstream model performance.
  • The Hessian matrix. Computed cheaply from a single forward pass on calibration data, the Hessian captures which weights are important by measuring the curvature of the loss with respect to each weight. Weights connected to frequently active input features have high curvature and must be quantized precisely.
  • The optimal compensation formula. When a weight is quantized, a closed-form formula distributes the rounding error across all remaining weights proportionally to their correlations with the quantized weight via the inverse Hessian. This redistribution recovers much of the accuracy lost to rounding.
  • Cholesky-based efficiency. Precomputing the Cholesky factorization of the inverse Hessian reduces the per-column update cost from O(d2)O(d^2) to O(d)O(d), making the entire quantization procedure tractable at scale with O(d3)O(d^3) total complexity per layer.
  • Group quantization. Using separate scale parameters for groups of 128 weights captures local weight distribution variations, adding only 0.125 bits per weight in overhead while substantially improving accuracy.
  • ActOrder. Processing high-sensitivity columns first, when more remaining weights are available for compensation, improves perplexity by 0.1 to 0.5 points at 4-bit precision.

For you, GPTQ means that 4-bit quantization is practical for production LLM deployment. Models lose minimal accuracy while fitting in a fraction of the original memory. A 65B parameter model that required a multi-GPU server in FP16 now runs on a single consumer GPU. The calibration data requirements and quantization time are modest one-time costs for large and permanent efficiency gains.

Understanding GPTQ also gives insight into the broader space of quantization methods. The principle of error compensation, rather than pure rounding, appears in various forms across modern quantization techniques. The Hessian as a measure of weight importance recurs in AWQ, QuIP, and AQLM. The layer-wise decomposition appears in nearly every practical post-training quantization method. Whether you are deploying a quantized model or developing new efficiency methods, the foundations established by GPTQ remain needed knowledge for anyone working in this space.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about GPTQ.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026gptqoptimizing, author = {Michael Brenndoerfer}, title = {GPTQ: Optimizing 4-Bit Weight Quantization for LLMs}, year = {2026}, url = {https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). GPTQ: Optimizing 4-Bit Weight Quantization for LLMs. Retrieved from https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide
MLAAcademic
Michael Brenndoerfer. "GPTQ: Optimizing 4-Bit Weight Quantization for LLMs." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide>.
CHICAGOAcademic
Michael Brenndoerfer. "GPTQ: Optimizing 4-Bit Weight Quantization for LLMs." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide.
HARVARDAcademic
Michael Brenndoerfer (2026) 'GPTQ: Optimizing 4-Bit Weight Quantization for LLMs'. Available at: https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). GPTQ: Optimizing 4-Bit Weight Quantization for LLMs. https://mbrenndoerfer.com/writing/gptq-4bit-weight-quantization-llm-guide

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.