Part of Language AI Handbook
Explains how gradient accumulation simulates large batch sizes on limited GPU memory by splitting batches into micro-batches.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Gradient Accumulation
Modern language models are voracious consumers of memory. A single forward pass through a large transformer can require tens of gigabytes of GPU memory just to store the activations needed for backpropagation, and that figure grows linearly with batch size. Yet the batch size you use during training is not a mere implementation detail. As we explored in the Large Batch Training chapter, the effective batch size directly shapes how gradients estimate the true loss gradient, how well learning rate scaling rules apply, and how quickly the model converges. When your GPU can only fit 4 samples at a time but your training recipe calls for an effective batch size of 512, hardware limits conflict with optimization requirements.
Gradient accumulation resolves this conflict with an elegant observation: the gradient of a loss computed over a large batch is equal to the average of the gradients computed over smaller sub-batches. Instead of requiring all samples to be processed simultaneously, you can process them sequentially and sum up their gradients before applying a single weight update. The effect is mathematically equivalent to using the full large batch, but the memory footprint at any given moment reflects only the small sub-batch currently being processed.
This technique has become a standard component of the training toolkit for large language models. When training GPT-style models on single nodes, when fine-tuning large models with limited GPU memory, or when experimenting on consumer hardware, gradient accumulation provides the bridge between what your hardware can hold and what your optimizer needs to work effectively.
The Mathematical Foundation
The insight behind gradient accumulation predates deep learning. It is a direct consequence of the linearity of differentiation: the gradient of a sum is the sum of the gradients, and the gradient of a mean is the mean of the gradients. This property means you can split a batch into any number of smaller chunks, compute gradients independently on each chunk, add the results together, and obtain the same answer as if you had processed the entire batch at once.
To see why this works, consider a dataset and a model with parameters . The standard training objective is to minimize the empirical loss over a batch :
where:
- : the per-sample loss for input and label
- : the number of samples in the batch
- : the model parameters
The gradient of this loss is:
Now suppose we partition into non-overlapping micro-batches of equal size . The gradient of the loss on micro-batch is:
If we compute the average of these per-micro-batch gradients:
Since the micro-batches partition and each has size :
The average of the micro-batch gradients equals the full-batch gradient exactly. This identity is not an approximation or a statistical trick. It is an algebraic consequence of the distributive property of addition over summation.
Deep learning frameworks make this easy to exploit. In PyTorch, each call to .backward() adds gradients into the parameter's .grad attribute rather than overwriting it. Gradient accumulation is simply using this additive behavior intentionally: you call .backward() multiple times without calling optimizer.zero_grad() between calls, allowing gradients to accumulate. The only thing you need to get right is the scaling factor, which ensures the accumulated sum equals the mean rather than the sum of the micro-batch gradients.
The Accumulation Procedure
Gradient accumulation works by splitting a logical large batch into a series of smaller micro-batches, running forward and backward passes on each one without applying the optimizer step, and accumulating the resulting gradients. Only after all micro-batches have been processed does the optimizer perform a single weight update using the accumulated gradients.
The procedure follows a precise sequence:
- Zero the gradients at the start of each accumulation cycle. This is the same
optimizer.zero_grad()call you know from ordinary training, but it now occurs once per logical batch rather than once per micro-batch. - Process each micro-batch: run the forward pass, compute the loss, scale the loss by where is the number of accumulation steps, and run the backward pass. The backward pass adds gradients to the gradient buffers without resetting them.
- Step the optimizer once after all micro-batches have been processed. The gradient buffers now contain the accumulated sum, which after the scaling applied to each loss equals the mean gradient over the full logical batch.
- Repeat for the next logical batch.
The memory benefit arises from step 2: at any point in time, only a single micro-batch occupies activation memory. The gradient buffers themselves must still hold the full model's gradients, but activation memory, which is proportional to batch size, is bounded by the micro-batch size.
The key structural difference from ordinary training is the placement of zero_grad() and step(). In ordinary training, both calls appear at every iteration. With gradient accumulation, zero_grad() appears at the start of every -th iteration, and step() appears at the end of every -th iteration. The iterations in between only call backward(), allowing gradients to accumulate.
The Scaling Requirement
The loss scaling in step 2 deserves careful attention. Standard training computes the mean loss over a batch and then backpropagates that mean. When you accumulate gradients across micro-batches, each micro-batch computes its own mean loss. If you simply sum these mean losses and backpropagate each one without scaling, your accumulated gradient is times larger than the gradient you would have computed from the full batch. This means the optimizer takes a step times too large, corrupting training.
To preserve mathematical equivalence with full-batch training, you must divide each micro-batch loss by before calling .backward(). This ensures that the accumulated gradients match the gradient from the combined batch exactly.
Let denote the gradient vector computed from the -th micro-batch with loss scaling applied. The scaled gradient for micro-batch is:
where:
- : the total number of accumulation steps (micro-batches per logical batch)
- : the model parameters
- : the mean cross-entropy loss computed over micro-batch
- : the -th micro-batch, a non-overlapping subset of the full logical batch
- : the gradient operator with respect to
After backward passes, PyTorch's autograd engine has accumulated the sum of all gradient contributions into the parameter gradient buffers:
where:
- : the accumulated gradient stored in the parameter buffers after all backward passes
- : the gradient contribution from micro-batch , already scaled by
If the micro-batches partition evenly (each of size ), the accumulated gradient equals the gradient computed directly on the full batch :
This equality holds because gradient computation is linear: the gradient of a mean is the mean of the gradients. Formally, if is a partition of the full batch with for all , then:
where:
- : the per-sample loss for sample
- : the full logical batch size
- : the size of each micro-batch
The accumulated result is an unbiased estimate of the full-batch gradient, and the optimizer update proceeds exactly as if the large batch had been processed in one shot.
What happens if you forget the scaling? The accumulated gradient becomes times larger than intended, so the optimizer takes an update step times too large. This is equivalent to multiplying the learning rate by , which typically causes training to diverge. This mistake is easy to make, especially when adapting a standard training loop. Always check that the loss passed to .backward() is divided by accumulation_steps before the call.
Accumulation Steps: How Many?
The number of accumulation steps is a key design choice. It directly determines the effective batch size seen by the optimizer:
where:
- : the effective batch size, the number of samples the optimizer "sees" per weight update
- : the number of accumulation steps
- : the micro-batch size, the number of samples in each forward-backward pass
If your GPU can process 8 samples per forward pass and you want an effective batch size of 256, you set .

Choosing K for Memory Constraints
The primary driver for choosing is the target effective batch size from your training recipe. As discussed in the Large Batch Training chapter, large language models typically train with effective batch sizes ranging from 256 to 4 million tokens per step, depending on the model size and phase of training. If your target effective batch size is fixed by empirical best practices (such as the linear scaling rule for learning rates), then is determined by:
where:
- : the ceiling function, rounding up to the nearest integer
- : the target effective batch size from your training recipe
- : the maximum micro-batch size that fits in GPU memory
A secondary consideration is the computational overhead of accumulation. Each micro-batch requires a complete forward and backward pass. With instead of , you perform eight forward-backward passes per optimizer step rather than one. On a single GPU, this overhead is negligible in terms of floating-point operations (you perform exactly the same number of multiplications and additions), but it can affect memory bandwidth utilization and synchronization overhead in distributed settings.
A practical way to find the maximum micro-batch size is to start with a small batch (for instance, 1 or 2 samples) and double it until you hit an out-of-memory error. The last successful size is your . Then set , rounding up.
The Effective Batch Size Trade-off
Increasing is not always beneficial beyond a certain point. Once the effective batch size exceeds the critical batch size for a given task, adding more samples per update produces diminishing returns in gradient quality while increasing the number of steps needed to reach the same number of epochs. The Large Batch Training chapter covers this trade-off in detail, including the Chinchilla and GPT-3 effective batch size recipes. The key point here is that gradient accumulation is a tool for reaching a target effective batch size, not an excuse to make that target arbitrarily large.
A useful mental model is that gradient quality saturates. With very few samples (say ), the gradient from a single sample can point in nearly any direction because it reflects only one data point's contribution. As batch size grows, the gradient direction becomes more stable and more representative of the true loss landscape. But beyond a certain size, adding samples barely changes the direction. Gradient accumulation lets you climb out of the noisy, small-batch regime without requiring hardware that can hold the entire logical batch simultaneously.
To see this saturation quantitatively, consider the variance of the stochastic gradient estimator. If each sample contributes a gradient with variance , the variance of the mean over independent samples is . Doubling from 64 to 128 halves the variance, a clear benefit. Doubling from 8192 to 16384 again halves the variance, but the remaining variance is already small. The signal-to-noise ratio of the gradient direction improves much faster at small batch sizes than at large ones. The concept of the critical batch size formalizes where this diminishing return becomes severe enough to matter in practice.
Gradient Accumulation for Memory
The memory savings from gradient accumulation are real and significant, but they apply only to a specific category of memory: activation memory. Understanding which types of memory are affected shows when accumulation helps.
Memory Taxonomy
GPU memory during training is occupied by several distinct components:
- Model parameters: The weights and biases themselves. For a model with parameters stored in float32, this costs bytes. In mixed precision training, a float16 copy costs bytes.
- Optimizer states: Adam maintains first and second moment estimates, costing roughly bytes in float32 (two copies of the parameter count). This dominates for large models.
- Gradient buffers: One gradient tensor per parameter, typically matching the parameter precision. Costs bytes in float32 or in float16.
- Activation memory: Intermediate tensors computed during the forward pass and retained for the backward pass. Unlike the above, this scales with batch size. For a transformer, activation memory grows roughly as where is the number of layers, is the model dimension, and is the batch size.
Gradient accumulation reduces only activation memory, since at any point only one micro-batch is being processed. Parameters, optimizer states, and gradient buffers remain constant regardless of .
It is worth being specific about why activation memory scales with batch size. During the forward pass, the model computes intermediate representations at each layer: attention scores, attention outputs, feed-forward intermediate activations, and so on. These are not thrown away after the forward pass because the backward pass needs them to compute gradients via the chain rule. Each sample in the batch has its own set of these intermediate tensors. If you double the batch size, you double the number of samples in flight, and thus double the total activation storage. Gradient accumulation sidesteps this by processing one micro-batch at a time: the activations for micro-batch 1 are computed, used for the backward pass, and then freed before micro-batch 2 begins.
The gradient buffers, by contrast, are not freed between micro-batches. They must persist across the entire accumulation cycle so that gradients can be summed. The size of these buffers is proportional to the number of parameters, not the batch size. For a 125-million parameter model in float32, the gradient buffers occupy roughly 500 MB regardless of whether you accumulate 1 or 64 micro-batches.


When Accumulation Helps Most
The benefit of accumulation is largest when activation memory is the binding constraint. For very large models, the optimizer state alone can exceed available memory, and accumulation provides no relief. In that regime, techniques like mixed precision training (covered in the Training Infrastructure part), ZeRO optimizer sharding, and activation checkpointing are necessary.
For models where activation memory dominates, gradient accumulation is immediately effective. A model that fits only 2 samples per forward pass can use to reach an effective batch size of 128, with the only cost being extra sequential forward-backward passes.
For a GPT-2 medium model (345 million parameters), the parameter memory in float32 is approximately 1.3 GB. The optimizer state in Adam adds another 2.6 GB, for a total of nearly 4 GB that is independent of batch size. Activation memory at batch size 32 for sequence length 1024 can exceed 8 GB. Reducing the micro-batch to 4 drops activation memory to roughly 1 GB, and gradient accumulation with recovers the effective batch size of 32. Total memory drops from over 12 GB to under 6 GB, making training feasible on a single 8 GB GPU.
In practice, the combination of micro-batch size and accumulation steps is often the first knob to turn when a training run fails with an out-of-memory error. Before resorting to more complex techniques like activation checkpointing or model parallelism, halving the micro-batch size and doubling can free enough memory to continue training without any change to the effective batch size or optimizer behavior.
Throughput and Wall-Clock Time
Gradient accumulation trades sequential computation for memory efficiency. Understanding the throughput implications helps you decide whether accumulation is the right tool or whether another memory-saving technique would be more efficient for your workload.
Computational Equivalence
In terms of raw floating-point operations, gradient accumulation is exactly as expensive as processing the full effective batch in a single forward-backward pass. The total number of multiply-accumulate operations performed across all micro-batches is identical to what a single large batch would require. This means that if your GPU were perfectly utilized at every batch size, the wall-clock time per optimizer step would not change at all.
In practice, GPU utilization is not constant across batch sizes. GPUs achieve their peak throughput when the workload is large enough to saturate all compute units and hide memory access latency. Very small micro-batches can leave compute units idle while waiting for data to transfer from global memory, resulting in lower effective throughput. For this reason, processing 8 micro-batches of size 4 can be slower in wall-clock time than processing one batch of size 32, even though the arithmetic work is identical.
The impact depends heavily on the model architecture and the hardware. For large models with high parameter counts relative to batch size, the model weights themselves dominate memory access time, and the overhead from sequential micro-batch processing is small. For smaller models where activations are proportionally larger, the inefficiency is more pronounced.
Throughput Implications in Practice
A useful way to think about accumulation overhead is in terms of kernel launch efficiency. Each forward and backward pass through a model involves hundreds to thousands of individual GPU kernel launches. Each kernel has a fixed overhead for scheduling and dispatch, regardless of the batch size. With accumulation steps, you incur times as many kernel launches per optimizer step. For large models where each kernel processes a substantial amount of data, this overhead is negligible. For very small models on fast GPUs, it can become noticeable.
The other throughput consideration is memory allocation and deallocation. Each micro-batch allocates activation tensors in the forward pass and frees them during or after the backward pass. The allocator must satisfy these requests efficiently. PyTorch's CUDA memory allocator uses a caching strategy that largely eliminates allocation overhead after the first accumulation cycle, so this is rarely a bottleneck in practice.

The practical takeaway is to choose the largest micro-batch size that fits in memory. Using micro-batch size 1 with achieves the same effective batch size as micro-batch size 8 with , but is likely to be slower due to the overhead of 32 kernel launch cycles versus 4. The correct workflow is to maximize within memory constraints, then set to hit the target effective batch size.
Accumulation Correctness
Gradient accumulation introduces several subtle correctness issues that can silently undermine training if not addressed carefully. The mathematical equivalence to full-batch training holds only under specific conditions.
Normalization Layers
Batch normalization computes statistics (mean and variance) over the batch dimension. When you use gradient accumulation, each micro-batch has different statistics from the full logical batch. This means that:
- The forward pass statistics used during training will differ from what they would be on the full batch.
- The running mean and variance estimates updated during the forward pass will reflect micro-batch statistics rather than full-batch statistics.
This difference goes beyond numerical precision. Batch normalization layers compute a different function when the batch size changes, and the model's behavior during inference (which uses the running statistics) will reflect whichever batch size was used during training.
The magnitude of this effect depends on how small the micro-batches are. With micro-batches of 64, the batch statistics are already a reasonable estimate of the population statistics, and the deviation from full-batch behavior is small. With micro-batches of 2 or 4, the statistics are highly variable, and the normalization can be erratic.
Layer normalization, which normalizes over the feature dimension rather than the batch dimension, is not affected by this problem. This is one reason why transformer architectures, which use layer normalization, are more amenable to gradient accumulation than convolutional architectures that rely on batch normalization.
If you must use batch normalization with gradient accumulation, you can use ghost batch normalization: compute batch norm statistics as if the full logical batch were present by accumulating statistics across micro-batches before the normalization step. This is more complex to implement but preserves the intended behavior.
Dropout
Dropout applies a random binary mask to activations during the forward pass, and the masked elements are excluded from the backward pass. When accumulating gradients across micro-batches, each micro-batch sees a different dropout mask. This is the correct behavior: the full logical batch processed in one shot would also apply dropout independently to each sample. The stochasticity is per-sample, not per-batch, so accumulation preserves the intended regularization.
However, a subtle issue arises if you set the random seed deterministically per micro-batch index rather than per sample. Ensure that your random state is not reset between micro-batches within an accumulation cycle.
Gradient Clipping
As discussed in the gradient clipping material from the Neural Network Foundations part, gradient clipping constrains the gradient norm before the optimizer step to prevent exploding gradients. When using gradient accumulation, you must clip the gradient after all micro-batches have been processed, not after each one. Clipping after each micro-batch would artificially reduce the effective gradient magnitude and could prevent gradients from flowing correctly.
To understand why, consider a scenario where the true gradient norm is 2.0 and you are using a clip threshold of 1.0. With micro-batches, each micro-batch contributes a gradient of norm approximately 0.5 (the full gradient divided by ). If you clip each micro-batch gradient at 1.0, no clipping occurs, but the accumulated gradient has norm 2.0, which would have been clipped. The accumulated gradient is twice the intended magnitude.
Conversely, if the true gradient norm is 0.25 and each micro-batch gradient has norm approximately 0.0625, clipping at 1.0 per micro-batch has no effect, but the result after accumulation (norm 0.25) is also fine. The correctness problem only manifests when the true full-batch gradient norm exceeds the clip threshold.
The correct procedure involves three steps performed in sequence:
- Accumulate gradients across all micro-batches.
- After the final backward pass, call gradient clipping on the accumulated gradients.
- Then call the optimizer step.

Mixed Precision and Loss Scaling
Mixed precision training uses float16 (or bfloat16) for forward and backward passes to reduce memory usage and increase throughput. Float16 has limited dynamic range, and gradients that are very small can underflow to zero. To prevent this, a loss scaler multiplies the loss by a large factor before the backward pass and then divides the resulting gradients by before the optimizer step.
When combining gradient accumulation with mixed precision, the loss scale and the accumulation scaling factor interact. The combined scaling applied to each micro-batch loss before backpropagation is:
where:
- : the scaled loss passed to
.backward()for micro-batch - : the loss scale factor used by the mixed precision scaler (typically a large power of 2)
- : the number of accumulation steps
- : the unscaled mean loss over micro-batch
After all micro-batches, the gradients are divided by before the optimizer step to reverse the loss scaling. Most modern training frameworks handle this automatically when you use their built-in gradient accumulation support. If implementing manually, ensure that both scaling factors are applied correctly and that the loss scaler's overflow detection operates on the accumulated gradients rather than on individual micro-batch gradients.
Bfloat16 (the 16-bit brain float format used in TPUs and supported by A100 GPUs) has wider dynamic range than float16 and often eliminates the need for loss scaling entirely. When using bfloat16, you can simplify the combined scaling: only the accumulation factor is required.
Distributed Training Interaction
In data-parallel distributed training, each GPU processes a different mini-batch and gradients are averaged across GPUs before the optimizer step. This all-reduce communication typically happens automatically during the backward pass via hooks attached to each parameter.
When using gradient accumulation in a distributed setting, you want to suppress the all-reduce communication for the first micro-batches and allow it only for the final micro-batch. This avoids unnecessary communication rounds that would each consume network bandwidth and introduce synchronization overhead.
To understand the cost savings, consider a training run with 8 GPUs and accumulation steps. Without the no_sync() pattern, each backward pass triggers an all-reduce, for a total of 8 all-reduce operations per optimizer step. With no_sync(), only the final backward pass triggers the all-reduce, reducing communication to 1 operation per optimizer step. For models where communication is a significant fraction of step time (common for large models on fast GPUs connected by slower interconnects), this can yield substantial speedups.
PyTorch's DistributedDataParallel provides the no_sync() context manager for exactly this purpose. When a model is wrapped in no_sync(), gradients are accumulated locally on each GPU without triggering all-reduce. On the final micro-batch (outside no_sync()), the backward pass triggers the all-reduce as usual, averaging the accumulated gradients across all GPUs.
A Worked Example
To make the mechanics concrete, consider a simple linear regression scenario that illustrates the mathematical equivalence between full-batch training and gradient accumulation.
Suppose you have a dataset of 8 samples and a model parameterized by a single weight , with mean squared error loss:
where:
- : the scalar model parameter
- : the batch of samples
- : the -th input-output pair
- : the batch size
The full-batch gradient over all 8 samples is:
Now split the 8 samples into micro-batches of size 2 each. For micro-batch (samples and ), the mean loss is:
The scaled gradient contribution from micro-batch (after dividing by ) is:
Summing across all 4 micro-batches to obtain the accumulated gradient:
The accumulated gradient equals the full-batch gradient exactly. This confirms that the scaling is necessary and sufficient to achieve mathematical equivalence.
The same derivation applies to any differentiable loss and any model architecture. The linearity of gradient computation means that you can always decompose a full-batch gradient into a sum of micro-batch gradients. The only requirement is that the micro-batches partition the full batch without overlap, and that each micro-batch uses the same model parameters.
What Changes When Micro-batches Are Unequal?
In some settings, the last micro-batch of an epoch may be smaller than the others, because the dataset size is not perfectly divisible by the full logical batch size. For example, if you have 1000 samples, a logical batch of 64, and micro-batch size 16, then 15 logical batches will have all four micro-batches of size 16, but the final logical batch may have only 1000 mod 64 = 40 samples, requiring some micro-batches of size 10 or fewer.
In this case, the correct approach is to weight each micro-batch by its actual size rather than using uniform scaling. The properly weighted contribution of micro-batch with samples is:
where is the total number of samples in the logical batch. This produces an unbiased estimate of the full-batch gradient even when micro-batches have different sizes.
Most training frameworks simplify this by dropping the last incomplete batch entirely, which avoids the unequal weighting problem. Whether to drop or include it is a hyperparameter choice: dropping it is simpler and slightly wastes data, while including it with proper weighting is more data-efficient but requires additional bookkeeping.
Code Implementation
We will build a complete gradient accumulation training loop in PyTorch that demonstrates the correct implementation, including the loss scaling, optimizer synchronization, and gradient clipping interaction.
Setting Up the Training Environment
Let us start by defining the model and data:
import random
import numpy as np
import torch
import torch.nn as nn
# Reproducibility
torch.manual_seed(42)
np.random.seed(42)
random.seed(42)
# Simple transformer-like language model for demonstration
class SmallTransformer(nn.Module):
def __init__(
self, vocab_size=1000, d_model=128, nhead=4, num_layers=2, seq_len=32
):
super().__init__()
self.embedding = nn.Embedding(vocab_size, d_model)
encoder_layer = nn.TransformerEncoderLayer(
d_model=d_model,
nhead=nhead,
dim_feedforward=256,
dropout=0.0,
batch_first=True,
)
self.transformer = nn.TransformerEncoder(
encoder_layer, num_layers=num_layers
)
self.head = nn.Linear(d_model, vocab_size)
self.d_model = d_model
def forward(self, x):
emb = self.embedding(x) * (self.d_model**0.5)
out = self.transformer(emb)
return self.head(out)
# Synthetic token dataset
VOCAB_SIZE = 1000
SEQ_LEN = 32
DATASET_SIZE = 512
tokens = torch.randint(0, VOCAB_SIZE, (DATASET_SIZE, SEQ_LEN + 1))
inputs = tokens[:, :-1] # shape: (512, 32)
targets = tokens[:, 1:] # shape: (512, 32)
dataset = torch.utils.data.TensorDataset(inputs, targets)Baseline: Standard Training Loop
First, let us implement the standard training loop with a batch size of 32, to establish the baseline gradient behavior. The demonstration model sets dropout to zero so that both paths compute deterministic gradients; the earlier dropout discussion still applies to real training runs.
from torch.utils.data import DataLoader
BATCH_SIZE = 32
MICRO_BATCH = 8
ACCUM_STEPS = BATCH_SIZE // MICRO_BATCH # = 4
model_baseline = SmallTransformer(vocab_size=VOCAB_SIZE)
optimizer_baseline = torch.optim.Adam(model_baseline.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
loader_baseline = DataLoader(dataset, batch_size=BATCH_SIZE, shuffle=False)
# Run one step to capture gradient norms
model_baseline.train()
optimizer_baseline.zero_grad()
x_full, y_full = next(iter(loader_baseline))
logits = model_baseline(x_full) # (32, 32, 1000)
loss = loss_fn(logits.reshape(-1, VOCAB_SIZE), y_full.reshape(-1))
loss.backward()
# Record gradient norms before optimizer step
grad_norms_baseline = {}
for name, param in model_baseline.named_parameters():
if param.grad is not None:
grad_norms_baseline[name] = param.grad.norm().item()
total_grad_norm_baseline = (
sum(v**2 for v in grad_norms_baseline.values()) ** 0.5
)Baseline batch size: 32 Baseline loss: 7.0769 Total gradient norm (baseline): 1.075834 Number of parameter tensors: 27
The baseline run processes 32 samples simultaneously and computes the mean loss across all of them. The gradient norm reflects the average gradient direction over the full batch.
Gradient Accumulation Loop
Now let us implement the equivalent training step using gradient accumulation with micro-batches of size 8 and accumulation steps:
# Create a fresh model with the same initial weights as baseline
model_accum = SmallTransformer(vocab_size=VOCAB_SIZE)
model_accum.load_state_dict(model_baseline.state_dict())
optimizer_accum = torch.optim.Adam(model_accum.parameters(), lr=1e-3)
loader_micro = DataLoader(dataset, batch_size=MICRO_BATCH, shuffle=False)
micro_iter = iter(loader_micro)
model_accum.train()
optimizer_accum.zero_grad() # Zero ONCE per logical batch
accum_loss = 0.0
for step in range(ACCUM_STEPS):
x_micro, y_micro = next(micro_iter)
logits_micro = model_accum(x_micro) # (8, 32, 1000)
# Compute mean loss over micro-batch, then scale by 1/K
loss_micro = loss_fn(
logits_micro.reshape(-1, VOCAB_SIZE), y_micro.reshape(-1)
)
scaled_loss = loss_micro / ACCUM_STEPS # CRITICAL: divide by K
scaled_loss.backward() # Accumulates into gradient buffers
accum_loss += loss_micro.item()
# Record accumulated gradient norms (before optimizer step)
grad_norms_accum = {}
for name, param in model_accum.named_parameters():
if param.grad is not None:
grad_norms_accum[name] = param.grad.norm().item()
total_grad_norm_accum = sum(v**2 for v in grad_norms_accum.values()) ** 0.5
mean_loss_accum = accum_loss / ACCUM_STEPSAccumulation steps (K): 4 Micro-batch size: 8 Effective batch size: 32 Mean loss (accumulated): 7.0769 Total gradient norm (accumulated): 1.075834
With identical initial weights and the same data ordering, the accumulated gradient norm should match the baseline gradient norm closely. The small numerical differences arise from floating-point rounding order, which varies when summing gradients in a different sequence.
Comparing Gradient Equivalence
Let us verify the gradient equivalence numerically by comparing per-parameter gradient norms:
# Collect per-layer gradient differences
param_names = list(grad_norms_baseline.keys())
grad_diffs = {}
for name in param_names:
if name in grad_norms_accum:
diff = abs(grad_norms_baseline[name] - grad_norms_accum[name])
rel_diff = diff / (grad_norms_baseline[name] + 1e-12)
grad_diffs[name] = {"abs_diff": diff, "rel_diff": rel_diff}
max_rel_diff = max(v["rel_diff"] for v in grad_diffs.values())
mean_rel_diff = sum(v["rel_diff"] for v in grad_diffs.values()) / len(
grad_diffs
)Gradient norm comparison (baseline vs accumulated): Total norm baseline: 1.07583426 Total norm accumulated: 1.07583422 Absolute difference: 4.22e-08 Max relative diff (per layer): 1.16e-07 Mean relative diff: 2.82e-08
The relative differences should be on the order of or smaller, which corresponds to floating-point rounding errors rather than any systematic bias. This confirms the mathematical equivalence between full-batch training and gradient accumulation.

Correct Gradient Clipping Integration
When gradient clipping is part of the training recipe, it must be applied after accumulation but before the optimizer step:
model_clip = SmallTransformer(vocab_size=VOCAB_SIZE)
model_clip.load_state_dict(model_baseline.state_dict())
optimizer_clip = torch.optim.Adam(model_clip.parameters(), lr=1e-3)
MAX_GRAD_NORM = 1.0
loader_clip = DataLoader(dataset, batch_size=MICRO_BATCH, shuffle=False)
clip_iter = iter(loader_clip)
model_clip.train()
optimizer_clip.zero_grad()
for step in range(ACCUM_STEPS):
x_c, y_c = next(clip_iter)
logits_c = model_clip(x_c)
loss_c = loss_fn(logits_c.reshape(-1, VOCAB_SIZE), y_c.reshape(-1))
(loss_c / ACCUM_STEPS).backward()
# Clip AFTER accumulation, BEFORE optimizer step
grad_norm_before_clip = nn.utils.clip_grad_norm_(
model_clip.parameters(), max_norm=MAX_GRAD_NORM
)
optimizer_clip.step() # Now apply the clipped accumulated gradientsGradient norm before clipping: 1.0758 Max gradient norm allowed: 1.0000 Clipping applied: True Gradients were scaled by: 0.9295
The clip_grad_norm_ function returns the gradient norm before clipping, which is useful for logging. If the norm exceeds MAX_GRAD_NORM, it scales all gradients down proportionally. This also is a health monitor: persistently high pre-clip norms can signal the training instabilities discussed in the Training Stability chapter.
Distributed Training with no_sync
In multi-GPU training with PyTorch's DistributedDataParallel, you should suppress gradient all-reduce during the accumulation micro-batches and allow it only on the final step:
# Pattern for DDP + gradient accumulation (not executed: requires DDP setup)
# model must be wrapped: model = DDP(base_model, ...)
# optimizer and loss_fn as before
optimizer.zero_grad()
for step in range(ACCUM_STEPS):
x_micro, y_micro = next(micro_iter)
# Suppress all-reduce for all but the last micro-batch
if step < ACCUM_STEPS - 1:
with model.no_sync():
logits = model(x_micro)
loss = loss_fn(logits.reshape(-1, VOCAB_SIZE), y_micro.reshape(-1))
(loss / ACCUM_STEPS).backward()
else:
# Final micro-batch: allow the all-reduce to fire
logits = model(x_micro)
loss = loss_fn(logits.reshape(-1, VOCAB_SIZE), y_micro.reshape(-1))
(loss / ACCUM_STEPS).backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()Without no_sync(), each of the backward passes would trigger an all-reduce across all GPUs, performing times as many communication rounds as necessary. For a large model with , this would add 15 unnecessary all-reduce operations per optimizer step.
Memory Footprint Comparison
Let us measure the actual memory savings from reducing the micro-batch size:
import gc
def measure_peak_memory(
batch_size, model_fn, vocab_size=VOCAB_SIZE, seq_len=SEQ_LEN
):
"""Measure peak GPU memory (or estimate via tensor sizes) for one forward-backward."""
gc.collect()
torch.cuda.empty_cache() if torch.cuda.is_available() else None
model = model_fn()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
x = torch.randint(0, vocab_size, (batch_size, seq_len))
y = torch.randint(0, vocab_size, (batch_size, seq_len))
# Estimate activation memory: each layer stores O(batch * seq * d_model) floats
d_model = model.d_model
num_layers = 2 # from SmallTransformer definition
bytes_per_float = 4 # float32
# Rough activation estimate: 12 tensors per layer (QKV, attention, FFN intermediates)
activation_bytes = (
batch_size * seq_len * d_model * num_layers * 12 * bytes_per_float
)
return activation_bytes / (1024**2) # MB
batch_sizes = [1, 2, 4, 8, 16, 32]
memory_estimates = []
for bs in batch_sizes:
mem_mb = measure_peak_memory(
bs, lambda: SmallTransformer(vocab_size=VOCAB_SIZE)
)
memory_estimates.append(mem_mb)Estimated activation memory by micro-batch size: Micro-batch Activation (MB) Effective batch (K=32/bs) ---------------------------------------------------------- 1 0.4 32 2 0.8 32 4 1.5 32 8 3.0 32 16 6.0 32 32 12.0 32
The activation memory scales linearly with micro-batch size. A micro-batch of 1 uses 1/32 the activation memory of a micro-batch of 32, while gradient accumulation with recovers the same effective batch size of 32. Parameters, gradients, and optimizer states (not shown) remain constant across all rows.

Key Parameters
The key parameters for gradient accumulation are:
- accumulation_steps (): The number of micro-batches per optimizer step. Set to
effective_batch_size // micro_batch_size. Higher values reduce activation memory but increase the number of forward-backward passes per update. - micro_batch_size: The number of samples processed per forward-backward pass. Limited by GPU memory capacity. Smaller values consume less activation memory.
- loss_scaling: Always divide each micro-batch loss by before calling
.backward(). Failure to do this scales the effective learning rate by . - gradient_clipping: Apply
clip_grad_norm_()after the final backward pass and beforeoptimizer.step(). Never clip inside the accumulation loop. - no_sync(): In distributed training, wrap the first micro-batch backward passes in
model.no_sync()to suppress unnecessary all-reduce communication.
Gradient Accumulation in Practice
Beyond the mechanics, gradient accumulation has shaped how practitioners approach training at every scale. Understanding how it is used in real systems helps you apply it correctly and diagnose problems when they arise.
The BERT Fine-Tuning Legacy
Gradient accumulation became widely known in the NLP community through the original BERT fine-tuning recipes released by Google in 2018. The BERT paper specified training with batch sizes of 32 or 64, but the reference implementation included a gradient_accumulation_steps parameter in its command-line interface. Researchers running BERT fine-tuning on single GPUs would set accumulation steps to 4 or 8 to match the intended effective batch size while working within 16 GB GPU memory limits.
This pattern set a template that virtually every subsequent fine-tuning library has followed. Hugging Face's Transformers library, which became the dominant interface for working with pre-trained language models, exposes gradient_accumulation_steps as a top-level parameter in its TrainingArguments class. The Accelerate library, designed to make distributed training accessible, wraps PyTorch's no_sync() pattern automatically when accumulation is enabled, so users don't need to manage the context manager manually.
GPT-2 Community Replication
When OpenAI released the GPT-2 model weights in 2019, a wave of community replication efforts followed. Researchers who wanted to train GPT-2 scale models from scratch often lacked access to the multi-GPU clusters used in the original training. Gradient accumulation was a key tool for these efforts: by running with micro-batch sizes of 4 or 8 on consumer GPUs and accumulating to effective batch sizes of 512 or more, researchers could reproduce the training dynamics of the original run without requiring the full hardware stack.
The combination of gradient accumulation, mixed precision training, and gradient checkpointing made it possible to train billion-parameter models on clusters far smaller than the original research infrastructure. This democratization of large-model training has had lasting effects on the research community.
Token-Level Effective Batch Size
Language model training recipes often specify effective batch sizes in terms of tokens rather than sequences, because sequence length varies. A recipe that calls for 524,288 tokens per step (a common choice for large language models) translates to different numbers of sequences depending on the context length. For a context length of 2048 tokens, this is 256 sequences per step.
If your GPU can fit 4 sequences of length 2048 per forward pass, you need accumulation steps to reach 256 sequences. If you can fit 16 sequences (perhaps by reducing sequence length to 512), you only need . The token-level framing makes it easier to think about the relationship between , , and the target effective batch size across different hardware configurations.
Interaction with Learning Rate Schedules
Gradient accumulation changes the relationship between gradient steps and training epochs. With , each optimizer step consumes 16 forward-backward passes worth of data. For a learning rate schedule specified in terms of training steps (not epochs), this means the schedule progresses 16 times more slowly in terms of data seen per unit of time.
When adapting a training recipe that was developed with one effective batch size to a different effective batch size using accumulation, you should verify how the learning rate schedule is parameterized. If the warmup and decay schedule is specified as a fraction of total steps, and you change , the total number of optimizer steps changes, which changes when the warmup ends and when the decay begins relative to the amount of data seen. Always specify schedules in terms of the optimizer step count, and verify that the step count is consistent with your setting.
Combining Accumulation with Other Memory Techniques
Gradient accumulation is rarely used in isolation. It combines naturally with other memory-saving techniques, and understanding how they interact helps you design an efficient training configuration.
Activation Checkpointing
Activation checkpointing (also called gradient checkpointing) reduces activation memory by recomputing some activations during the backward pass instead of storing them from the forward pass. Where gradient accumulation reduces activation memory proportional to micro-batch size, activation checkpointing reduces it proportional to the number of checkpointed layers, at the cost of one additional forward pass per backward pass.
The two techniques are orthogonal: you can use both simultaneously. With activation checkpointing, even a micro-batch of size 8 can have dramatically reduced activation memory, which means you can potentially increase before resorting to accumulation. The tradeoff is that checkpointing increases computation time (roughly by 33%, since you recompute one forward pass), while accumulation does not increase the total computation per token but does add kernel launch overhead.
For large models where even single-sample activation memory is significant, activation checkpointing allows you to use meaningful micro-batch sizes that maintain good GPU utilization. Combining both techniques (checkpointing plus accumulation) gives you control over both computation overhead and communication overhead.
ZeRO Optimizer Sharding
The ZeRO (Zero Redundancy Optimizer) family of techniques shards optimizer states, gradients, and parameters across GPUs to reduce per-GPU memory. ZeRO-1 shards optimizer states, ZeRO-2 additionally shards gradients, and ZeRO-3 shards all three including parameters.
When using ZeRO-2 or ZeRO-3, gradient accumulation still works but interacts with the sharding pattern. With ZeRO-2, gradients are gathered and reduced across GPUs at each backward pass. The no_sync() pattern for suppressing all-reduce during accumulation needs to be coordinated with ZeRO's gradient reduction hooks. Libraries like DeepSpeed handle this automatically through their gradient accumulation integration.
The combination of ZeRO-3 and gradient accumulation is particularly powerful for very large models: ZeRO-3 eliminates the need to hold full parameter, gradient, and optimizer state copies on each GPU, while gradient accumulation controls the activation memory and allows matching the effective batch size to the training recipe.
Mixed Precision Training
Mixed precision training, as discussed in the context of the mixed precision section above, reduces parameter and activation memory by using float16 or bfloat16 for most computations. The main-copy parameters are maintained in float32 for numerical stability.
Mixed precision roughly halves activation memory (since float16 activations are half the size of float32), which means you can often double the micro-batch size. If doubling the micro-batch is enough to reach the target effective batch size without accumulation, you can potentially eliminate accumulation entirely, trading the overhead of kernel launches for the overhead of the mixed precision scaler.
When both techniques are used together, the combined effect is multiplicative: mixed precision halves activation memory, and gradient accumulation by a factor of further reduces the effective activation footprint to of the full-batch float32 baseline.
Limitations and Impact
Gradient accumulation is not a universal solution to the memory-batch-size trade-off, and its limitations are worth understanding clearly.
The most significant limitation is that accumulation does not reduce the memory consumed by model parameters, gradients, or optimizer states. For very large models, these components often dominate. A 13-billion parameter model in float32 requires 52 GB just for parameters and another 104 GB for Adam optimizer states, totaling 156 GB that no amount of accumulation can reduce. In this regime, distributed optimization techniques like ZeRO-3 (which shards optimizer states across GPUs), model parallelism, or gradient checkpointing are required in addition to or instead of accumulation.
Training speed is a subtler limitation. While the total number of floating-point operations is identical whether you use one large batch or small micro-batches, the wall-clock time can differ. On a single GPU, the overhead is minimal. In a distributed setting, the no_sync() pattern reduces communication overhead, but the synchronization that does occur on the final micro-batch still blocks all GPUs from proceeding to the next logical batch until they have exchanged gradients. Sequential micro-batch processing also reduces the ability to overlap computation and communication.
Batch normalization imposes a concrete constraint. Production systems that rely heavily on batch normalization (common in vision models but less so in NLP) require additional engineering effort to ensure statistical consistency across micro-batches. For transformer-based language models, which use layer normalization throughout, this problem does not arise.
Another subtle limitation involves training diagnostics. When monitoring training health, metrics like gradient norm and loss should be computed at the logical batch level, not the micro-batch level. A loss spike visible in individual micro-batch losses may disappear after averaging across all micro-batches; conversely, an instability might appear less severe when viewed at the micro-batch granularity. Always log the fully accumulated gradient norm and the mean loss over the logical batch to get a consistent picture of training dynamics.
Despite these limitations, gradient accumulation has had substantial practical impact. It enabled researchers and practitioners with access to only a few GPUs to reproduce and extend experiments that originally required entire server clusters. The ability to target any effective batch size regardless of hardware constraints made gradient accumulation a standard component of BERT fine-tuning recipes, GPT-2 replication efforts, and instruction tuning pipelines. Hugging Face's Trainer class, which is the most widely used fine-tuning interface in the NLP community, exposes gradient accumulation as a first-class parameter because it is so frequently needed.
Looking ahead to the Training Stability chapter, gradient accumulation interacts with loss spike detection in an important way. When monitoring gradient norms for stability, you should log the gradient norm after full accumulation rather than per micro-batch. A spike in the per-micro-batch gradient norm that resolves after averaging across all micro-batches is not a true instability. Only a spike in the fully accumulated gradient warrants intervention.
Summary
Gradient accumulation splits a logical large batch into smaller micro-batches, runs forward and backward passes on each without updating weights, and applies a single optimizer step after all micro-batches are processed. The key results are:
- The accumulated gradient equals the full-batch gradient when each micro-batch loss is divided by before backpropagation.
- Activation memory is bounded by the micro-batch size, not the effective batch size.
- Parameters, gradients, and optimizer states are unaffected by accumulation.
- Gradient clipping must be applied after the full accumulation cycle, not inside the loop.
- Batch normalization statistics are inconsistent across micro-batches; layer normalization is unaffected.
- Distributed training requires the
no_sync()pattern to avoid redundant all-reduce calls during accumulation. - The technique combines naturally with activation checkpointing, mixed precision training, and ZeRO optimizer sharding.
- Throughput is best preserved by using the largest micro-batch size that fits in memory, then setting to reach the target effective batch size.
The effective batch size is the quantity that governs optimization behavior. Gradient accumulation gives you control over this quantity independently of GPU memory, making it an essential tool for matching your training recipe to your hardware.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about gradient accumulation.
Gradient Accumulation Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
1 comment
This topic covered really well !!