Part of Language AI Handbook
Explains how pipeline parallelism splits deep models across devices, manages bubble overhead with micro-batching, and compares GPipe vs 1F1B schedules.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Pipeline Parallelism: Stages, Micro-Batching, GPipe, and 1F1B
Training large language models requires distributing computation across multiple devices, and you have already seen two complementary strategies for doing this. Data parallelism splits the training dataset across workers so each device processes a different mini-batch while sharing gradient updates. Tensor parallelism splits individual weight matrices across devices so each GPU holds a slice of a single layer's parameters. Both strategies are powerful, but neither directly addresses a fundamental bottleneck: when a model is so large that a single layer or a modest contiguous block of layers cannot fit in one GPU's memory, and when the sequential depth of the network creates idle time because each layer must wait for the previous layer to finish before it can begin.
Pipeline parallelism takes a different approach. Instead of splitting data or splitting weight matrices, it splits the model itself along its depth, assigning contiguous groups of layers to different devices. The first device holds layers 1 through 8, the second holds layers 9 through 16, the third holds layers 17 through 24, and so on. This lets the model's total parameter count scale linearly with the number of pipeline stages, independent of any single device's memory capacity. The challenge, as we will see, is that a naive implementation leaves most devices idle most of the time, waiting for activations to arrive from the previous stage. Understanding how to minimize this idle time is the central engineering problem of pipeline parallelism, and the history of solutions to it has shaped how the largest models in existence are trained today.
The intuition behind pipeline parallelism is borrowed from industrial assembly lines. A car factory does not wait for one vehicle to be fully assembled before starting the next. Instead, different vehicles are in different stages of assembly simultaneously: the welding station, the painting station, and the final inspection station are all busy at the same time, each working on a different car. Pipeline parallelism applies exactly this logic to deep neural networks. While one micro-batch is being processed by the second stage of the model, the first stage can already be working on the next micro-batch. The goal is to keep every stage busy simultaneously, maximizing the fraction of time all devices are performing useful computation rather than sitting idle waiting for their neighbor to finish.
The history of pipeline parallelism in deep learning spans several important papers. The GPipe paper from Google Brain (2019) introduced a clean flush-based schedule that proved pipeline parallelism was practical at scale. The PipeDream paper from Microsoft Research (2019) demonstrated asynchronous variants with weight stashing. Megatron-LM from NVIDIA refined the approach with the 1F1B schedule and later the interleaved schedule, which became the standard for training hundreds-of-billions-parameter models. Understanding the progression from GPipe through 1F1B to interleaved scheduling is not just historical curiosity: each step was motivated by a concrete engineering constraint, and knowing those constraints helps you understand which schedule to choose for your own training runs.
Model Partitioning into Stages
The first design decision in pipeline parallelism is how to divide the model into stages. A stage is a contiguous sequence of layers that runs on a single device. The boundary between stages is called a pipeline cut point, and it determines both the computational load per device and the volume of data that must be communicated between devices at each boundary.
A pipeline stage is a contiguous block of model layers assigned to a single device. During forward propagation, each stage receives activations from the previous stage, processes them through its layers, and passes the result to the next stage. During backward propagation, gradient flow reverses this path, with each stage receiving gradients from the next stage and passing gradients to the previous stage.
Choosing where to place cut points involves balancing several competing factors. If the stages are imbalanced in computational cost, the slowest stage becomes the bottleneck: faster stages finish quickly but cannot proceed until they receive either activations (during the forward pass) or gradients (during the backward pass) from the slow stage. The goal is to partition layers such that each stage takes approximately the same wall-clock time per sample. This is called load balancing, and it is more subtle than it might initially appear. Two stages with the same number of layers can have very different computational costs if one contains attention layers operating on long sequences and another contains feed-forward layers with large hidden dimensions. Similarly, the first stage that includes the embedding lookup and the last stage that includes the output projection and softmax may have compute profiles that differ from the middle transformer blocks.
For a transformer model with layers and devices, the simplest strategy assigns consecutive layers to each device. This achieves perfect load balance if every layer is identical in compute cost, which is approximately true for standard transformer blocks. In practice, the embedding layer and the output projection layer may have different costs than interior transformer blocks, and some architectures include routing layers (as in mixture-of-experts models) that behave very differently from standard dense layers. A profiling-driven partition, where you measure the actual wall-clock time per layer on representative inputs and then use a dynamic programming algorithm to find the partition that minimizes the maximum stage duration, is more reliable than a uniform assignment. The problem of finding the optimal partition is formally equivalent to a bin-packing variant, and it can be solved exactly in polynomial time given per-layer timing measurements.
Communication Volume at Stage Boundaries
The communication cost between stages is determined by the size of the activation tensor passed at each cut point. For a transformer with sequence length , batch size , and hidden dimension , the tensor passed between stages has shape , so the communication volume per forward pass is:
where:
- : micro-batch size (number of sequences per micro-batch)
- : sequence length in tokens
- : hidden dimension (model width)
- : 2 for bfloat16, 4 for float32
For a GPT-3-scale model with , , and , using bfloat16, the communication tensor has size MB per stage boundary. This must flow across the interconnect linking the two pipeline stages. On NVLink, which provides roughly 600 GB/s of bidirectional bandwidth between GPUs in the same node, this transfer takes about 0.33 milliseconds. On InfiniBand HDR (200 Gb/s), the same transfer takes closer to 8 milliseconds. The compute time per stage for this same configuration is several hundred milliseconds on a single A100. The ratio of communication to compute time determines whether the pipeline is compute-bound or communication-bound: on NVLink it is clearly compute-bound (the interconnect is fast enough to overlap communication with computation), while on slower networks the communication overhead can become a meaningful fraction of total time.
This is typically much smaller than the communication volume in data parallelism, which involves gradient tensors proportional to total parameter count. A GPT-3-scale model with 175 billion parameters requires 350 GB of gradients (in bfloat16) per all-reduce operation in data parallelism. Compared to 201 MB per pipeline boundary per micro-batch, the pipeline communication overhead is orders of magnitude smaller. Pipeline stage-to-stage communication is also point-to-point: only two devices are involved at each boundary, and only one tensor flows per forward pass and one tensor flows per backward pass. This contrasts with the all-reduce patterns of data parallelism, which involve all devices simultaneously.
Memory Implications of Staging
Each device in a pipeline stores the parameters and buffers for its assigned layers, plus the activation tensors needed for the backward pass. The parameter memory per device scales with the total model size divided by the number of stages:
where:
- : total number of model parameters (in bytes, accounting for numerical precision)
- : number of pipeline stages
This fraction is the primary motivation for pipeline parallelism: models that cannot fit on one device can be spread across multiple devices, with each device using only of the total parameter memory. A model with 175 billion parameters in 16-bit precision requires roughly 350 GB of memory, which exceeds the capacity of any single GPU available as of 2024. Spreading it across 8 pipeline stages reduces per-device parameter memory to about 44 GB, which fits comfortably on modern 80 GB A100 or H100 GPUs. With 16 stages, per-device parameter memory drops to about 22 GB.
Beyond raw parameter storage, the optimizer state memory must also be considered. With the Adam optimizer, each parameter requires two additional floating-point values (the first and second moment estimates). When using mixed-precision training (model weights in bfloat16 but optimizer states in float32), the optimizer state memory is larger than the parameter memory: each parameter requires 4 bytes for the float32 master weight, 4 bytes for the first moment, and 4 bytes for the second moment, totaling 12 bytes per parameter in optimizer state alone. For a stage holding parameters, the total memory requirement including optimizer state is:
where the factor of 2 comes from the bfloat16 model parameters and the factor of 12 comes from float32 optimizer states. For a 175 billion parameter model split across 8 stages, this gives roughly per stage, which is tighter but still feasible on 80 GB GPUs.
The activation memory requires more careful analysis. During the forward pass, each stage must retain its intermediate activations until the corresponding backward pass arrives. In a naive sequential pipeline, stage must hold its activations in memory from the moment it finishes the forward pass until the backward pass through all later stages has completed and the gradient has been returned to stage . This can be a long wait, especially for early stages in a deep pipeline. The key challenge therefore extends beyond the per-stage parameter and optimizer state memory, which are fixed and predictable, to the activation memory, which depends on the schedule and can vary by an order of magnitude depending on implementation choices.
The Pipeline Bubble Problem
The most important challenge in pipeline parallelism is not communication or memory but time: the problem of stages sitting idle while waiting for work. To understand why this happens, it is worth tracing through exactly what happens when you run a model through a naive pipeline.
Imagine four pipeline stages labeled S1 through S4. A single input mini-batch enters S1 at time step 1. S2 cannot begin until S1 finishes, S3 cannot begin until S2 finishes, and so on. The result is a startup pattern where:
- At time step 1: S1 is working, S2, S3, and S4 are idle.
- At time step 2: S2 is working on the result from S1, while S1 is idle (finished its pass and waiting), and S3 and S4 are still idle.
- At time step 3: S3 is working, while S1, S2, and S4 are all idle.
- At time step 4: S4 is working, while S1, S2, and S3 are all idle.
After the forward pass completes, the backward pass flows in reverse: S4 computes gradients first, then passes them to S3, and so on back to S1. This also creates a sequential chain of idle time, now in reverse order. The combined result is that the startup and teardown phases of the pipeline are almost entirely wasted: at any given moment, only one stage is doing useful work.
A pipeline bubble is the idle time that accumulates across pipeline stages due to the sequential dependency between stages. Bubbles appear at the beginning and end of every micro-batch sequence. For a pipeline with stages, the naive single-batch approach wastes of total device-time on bubbles. With four stages, three-quarters of all device-time is wasted.
To quantify the bubble fraction precisely, consider pipeline stages processing micro-batches. The startup ramp contributes idle time steps before all stages can participate (the first stage begins immediately, but the -th stage cannot begin until time steps have passed because it must wait for activations from earlier stages). The drain phase at the end is symmetric, contributing another idle time steps as gradients move backward through the pipeline. Each stage performs useful operations— forward and backward—and the full schedule lasts time steps. The bubble fraction is therefore:
where:
- : number of pipeline stages (equal to the number of devices)
- : number of micro-batches processed in each pipeline run
- : total idle time slots from the startup ramp and teardown drain
For large , the denominator is dominated by , so the bubble fraction approaches , which converges to zero as grows. This asymptotic behavior is the key insight: the bubble overhead becomes negligible when is much larger than .
To see this concretely with numbers: for stages and micro-batch, the bubble fraction is . You are paying for four GPUs but using only about 25% of their combined compute capacity. Increasing to 8 micro-batches brings the bubble fraction to . Increasing to 32 micro-batches brings it to . The improvement continues, but each additional doubling of produces diminishing returns: going from to reduces the bubble from about 9% to about 4.5%.
It is worth understanding why the bubble cannot be eliminated entirely with any finite number of micro-batches. The startup ramp is unavoidable: the first micro-batch must travel through all stages sequentially before the pipeline can be considered "full." During this ramp-up, early stages are idle after finishing their portion of the first micro-batch and before the second micro-batch arrives. Similarly, the drain ramp at the end is unavoidable: after the last forward pass, backward passes flow back through the pipeline, and late stages finish first while early stages wait for gradients to arrive from later stages. As grows, these fixed startup and drain costs become a smaller fraction of the schedule and approach zero asymptotically.

Micro-Batching: Filling the Bubble
The solution to the bubble problem is to process multiple micro-batches simultaneously. A micro-batch is a small chunk of the full mini-batch. Instead of waiting for one full mini-batch to complete its forward pass before starting the next one, we split the mini-batch into equal micro-batches and inject them into the pipeline as fast as the first stage can produce them.
A micro-batch is a subdivision of the full training mini-batch. When using pipeline parallelism, the mini-batch is split into equal parts. Each part flows through the pipeline independently, and the gradients from all micro-batches are accumulated before the optimizer step is applied. From the optimizer's perspective, the result is identical to processing the full mini-batch at once.
The pipeline schedule looks quite different with multiple micro-batches in flight. While S1 is processing micro-batch , S2 is simultaneously processing micro-batch . While S1 processes , S2 processes and S3 processes . The stages are kept busy by overlapping the processing of different micro-batches on different stages. The bubble still exists at the beginning and end of the sequence (there is no way to avoid the startup ramp for a strictly sequential pipeline), but the bubble fraction decreases as increases because the useful work time grows while the ramp time stays constant.
The key insight is that all micro-batches share the same model weights and contribute to the same gradient update. At the end of the pipeline run, all stages have produced their partial gradients, these are accumulated across micro-batches, and the optimizer takes a single parameter update step. From the optimizer's perspective, the pipeline processed one mini-batch of size . The decomposition into micro-batches is an implementation detail for managing pipeline utilization. It is transparent to the loss function and the optimizer.
One important subtlety is the choice of micro-batch size. Smaller micro-batches mean more micro-batches for the same total mini-batch size, which reduces the bubble fraction. But very small micro-batches can underutilize the GPU's tensor cores. Modern GPU tensor cores achieve peak efficiency when matrix dimensions are large multiples of 8 or 16 (and ideally 64 or 128 for the best hardware utilization on A100s and H100s). A micro-batch with too few sequences may result in matrix multiplications with tiny leading dimensions, causing the GPU to spend most of its time on memory access rather than computation. There is a sweet spot where micro-batches are large enough to keep the GPU's arithmetic units busy but small enough that is large enough to suppress the bubble. Finding this sweet spot requires profiling on the specific hardware configuration, though a rule of thumb is that the micro-batch size should be at least 8 sequences for standard transformer models.
Gradient Accumulation and Weight Updates
The relationship between micro-batches, mini-batches, and gradient accumulation is worth understanding precisely, because it affects both memory usage and training stability.
During the forward pass for micro-batch , each stage computes the activations and passes them to stage . Each stage also stores for use during the backward pass. During the backward pass for micro-batch , stage receives gradients from stage , computes gradients with respect to its own parameters, and passes to stage .
The parameter gradients are accumulated across micro-batches. After all micro-batches have completed their backward passes, stage holds the accumulated gradient:
where:
- : the accumulated gradient used for the parameter update at stage
- : the contribution to the gradient from micro-batch at stage
- : the number of micro-batches per optimizer step
This accumulated gradient is mathematically equivalent to computing the gradient over the full mini-batch simultaneously, assuming the micro-batch splits are uniform and all micro-batches use the same parameter values. The optimizer step then proceeds exactly as it would in single-device training. This equivalence is what makes synchronous pipeline parallelism a valid distributed training strategy: the model sees the same gradient signal it would see in non-pipelined training, just computed in parallel across devices.
One practical consideration is that different micro-batches could in principle encounter different parameter values if weights are updated between micro-batches. The standard synchronous pipeline parallelism approach avoids this by processing all micro-batches before taking any weight update step. This ensures gradient coherence: all micro-batches see the same model weights, so their gradients can be averaged. Asynchronous variants that update weights between micro-batches introduce a form of weight staleness, where later micro-batches in the same optimizer step observe slightly different weights than earlier ones. Asynchronous schedules can potentially achieve better device utilization, but the staleness introduces gradient inconsistency that tends to harm convergence for large models. For this reason, synchronous pipeline parallelism is the standard approach in practice.
The Micro-Batch Size and Total Batch Size Relationship
Understanding the relationship between micro-batch size, the number of micro-batches, and the total effective batch size is essential for correctly configuring pipeline parallelism training runs.
If each micro-batch contains sequences and there are micro-batches per optimizer step, the total batch size processed per step per pipeline replica is:
When this pipeline is also combined with data parallelism, where replicas of the full pipeline are run in parallel, the global effective batch size is:
Large-scale training runs often target a specific global batch size for training stability reasons (the learning rate and its schedule are typically tuned for a specific batch size). Pipeline parallelism and data parallelism together give you multiple levers to hit the target batch size: you can increase to reduce the bubble fraction, increase to add more data-parallel replicas, and adjust to satisfy hardware utilization constraints. The flexibility to tune these independently, rather than treating the batch size as a single monolithic parameter, is one of the practical benefits of the 3D parallelism approach used by modern large model training frameworks.
Pipeline Schedules
The order in which forward and backward passes are scheduled across stages and micro-batches is called the pipeline schedule. Different schedules make different tradeoffs between bubble fraction, activation memory, and implementation complexity. The two most important schedules in practice are GPipe and 1F1B, with the interleaved 1F1B schedule providing a further refinement.
The schedule determines how busy each device is, how much memory it requires, and how long the pipeline takes to produce a gradient update. These tradeoffs are non-trivial: a schedule that minimizes the bubble fraction might require enormous activation memory, making it impractical for large models. A schedule that minimizes memory might introduce communication complexity that reduces efficiency. Understanding the specific tradeoffs of each schedule is essential for selecting the right one for a given training configuration, cluster topology, and model architecture.
GPipe: Flush-Based Schedule
GPipe (Google's Pipeline for Machine Learning) was one of the earliest systematic frameworks for pipeline parallelism in deep learning. The paper, published by researchers at Google Brain in 2019, demonstrated that models with billions of parameters could be trained reliably using pipeline parallelism with gradient checkpointing. Its schedule is conceptually straightforward: all micro-batches complete their forward passes before any backward passes begin. This "flush" design is easy to reason about, simple to implement, and was used to train some of the earliest large language models.
The full GPipe schedule for stages and micro-batches proceeds as follows. In the forward phase, micro-batches through are injected sequentially into the pipeline. Stage 1 processes , then , then , and so on, passing activations downstream as each micro-batch completes. By the time stage has processed , all activations for all micro-batches are sitting in GPU memory across all stages. Only then does the backward phase begin. Stage computes gradients for , then , and so on in reverse order. Gradients flow backward through the pipeline until all micro-batches have completed their backward passes and all parameter gradients have been accumulated.
The bubble fraction for GPipe is:
where:
- : number of pipeline stages
- : number of micro-batches
Note that this formula uses in the numerator rather than . This is because in the GPipe flush-based schedule, the startup bubble and the drain bubble are shared: the same pipeline states that are idle during the forward-phase ramp-up become idle again during the backward-phase drain, but the two phases are not counted separately in the same way as a fully interleaved schedule. The effective bubble fraction formula is , which is equivalent to the general formula when you account for the specific timing of the flush.
For and , this gives . For and , it gives . As grows large, the bubble fraction approaches zero.
The critical weakness of GPipe is its peak activation memory. Because all micro-batch activations across all stages must be stored simultaneously (they are all needed for the eventual backward phase), the total activation memory scales as:
where:
- : number of micro-batches
- : number of pipeline stages
- : activation memory per layer per micro-batch
For large and , this can easily exceed available GPU memory. Consider a GPT-3-scale model with stages and micro-batches. Each activation tensor has shape in bfloat16, which is about 201 MB. The total activation memory is , which on top of the parameter and optimizer state memory would exceed any single GPU.
GPipe addresses this through gradient checkpointing (also called activation recomputation). Instead of storing all intermediate activations during the forward pass, only certain checkpoint activations are kept. When the backward pass needs an intermediate activation that was not stored, the forward pass is rerun from the nearest checkpoint to recompute it. This reduces activation memory dramatically at the cost of additional computation.
Gradient checkpointing is a technique that trades compute for memory. Instead of storing all intermediate activations during the forward pass, only a subset of checkpoint activations is saved. When the backward pass needs a discarded activation, it recomputes the forward pass from the nearest checkpoint. This reduces activation memory at the cost of approximately one additional forward pass per backward pass, roughly doubling the compute cost for the forward-pass portion of training.
In GPipe, gradient checkpointing is applied at the stage granularity: each stage stores only its input activation (the tensor it received from the previous stage) and recomputes all intermediate activations within the stage during the backward pass. With this approach, the activation memory per stage is just regardless of how many layers are in the stage or how many micro-batches are in flight. The recomputation cost is roughly one additional forward pass for every backward pass, which adds about 30-33% to total compute time (since the backward pass typically costs about twice as much as the forward pass, and you are adding one more forward pass).
The simplicity of GPipe made it the natural first implementation choice, and it demonstrated that pipeline parallelism was practical at scale. The activation memory limitations, combined with the compute overhead of gradient checkpointing, motivated the development of more memory-efficient schedules.
1F1B: One Forward, One Backward
The 1F1B (One Forward, One Backward) schedule, introduced by the Megatron-LM project at NVIDIA, addresses the memory problem of GPipe by interleaving forward and backward passes rather than separating them into two distinct phases. The core idea is that a device should begin its backward pass as soon as it has accumulated enough in-flight forward passes to make the backward pass possible, rather than waiting for all micro-batches to complete their forward phases first.
In the steady-state of a 1F1B schedule, each device alternates between one forward-pass micro-batch and one backward-pass micro-batch. At any point in time, each device is always working on one computation, and the peak number of activations stored in memory is bounded by the number of stages rather than the number of micro-batches .
The 1F1B schedule proceeds in three distinct phases.
The warmup phase fills the pipeline: stage 1 processes micro-batches through in the forward direction, and stage performs its first backward pass after receiving from all prior stages. This warmup phase is where the pipeline bubble accumulates, since early stages are busy on their forward passes while late stages are still waiting for activations to arrive. The warmup phase lasts for time steps of startup latency.
The steady-state phase then alternates forward and backward passes, maintaining full pipeline utilization. Every stage is either processing a forward pass or a backward pass at every time step. In the steady state, stage forwards micro-batch while simultaneously backward-passing micro-batch . The stage is never idle during the steady state, as there is always either a new forward micro-batch arriving or a backward gradient ready to process.
The drain phase at the end completes the remaining backward passes as the pipeline empties. After the last forward pass ( reaching stage ), no new forward passes enter, and the pipeline drains as gradients flow backward through each stage. The drain phase mirrors the startup bubble in duration: it also occupies time steps.
The bubble fraction for 1F1B is identical to GPipe:
The advantage of 1F1B is not a smaller bubble but a much smaller peak activation memory. Because each stage begins processing backward passes before accumulating many pending forward-pass activations, the maximum number of in-flight activations at any stage is bounded by :
where:
- : number of pipeline stages
- : activation memory per layer per micro-batch
Comparing to GPipe's , the 1F1B schedule saves a factor of in activation memory. For a typical training run with and , 1F1B uses 32 times less activation memory than naive GPipe without checkpointing. This allows 1F1B to handle significantly larger without running out of memory, which in turn enables smaller micro-batches (for better tensor core utilization) while still achieving a low bubble fraction.
The reason for the memory difference becomes clear when you examine the schedule visually. In GPipe, by the time the backward phase begins, every stage has already processed all micro-batches in the forward direction and is holding all of their activations in memory simultaneously. In 1F1B, a stage begins its backward pass much sooner. Stage begins its first backward pass after receiving only its -th micro-batch's activations, not after receiving all micro-batches. As each backward pass completes at stage , the corresponding forward activations can be freed. The memory at any stage is therefore bounded by the maximum number of micro-batches simultaneously "in flight" at that stage, which is , not .
This is the key algorithmic insight of 1F1B: by scheduling backward passes early, you free activation memory early, keeping the peak memory footprint proportional to pipeline depth rather than to the number of micro-batches. For practical training runs where can be tens or hundreds of micro-batches, this memory savings is what makes 1F1B the default schedule for large model training.
The Interleaved 1F1B Schedule
The interleaved schedule is a further refinement of 1F1B introduced by Megatron-LM to reduce the bubble fraction at the cost of increased communication. Standard 1F1B assigns a contiguous block of layers to each device. The interleaved schedule instead assigns smaller, non-contiguous blocks of layers each to every device. A single device might hold layers 1-4 and layers 33-36, with another device holding layers 5-8 and 37-40, and so on.
This arrangement means that each micro-batch traverses the pipeline times instead of once, because it must visit each device multiple times, once per model chunk assigned to that device. Each traversal is called a pipeline pass. The bubble fraction for the interleaved schedule improves by a factor of compared to standard 1F1B:
where:
- : number of model chunks per device (also called virtual pipeline stages)
- : number of devices
- : number of micro-batches
With , the interleaved schedule halves the bubble fraction. With , it reduces the bubble to a quarter. The intuition is that by splitting the model into more, smaller chunks, the pipeline fills up faster relative to the total amount of work. The startup ramp covers the same number of time steps, but those time steps represent a smaller fraction of the total work because each "pass" through the pipeline covers only of the model's layers. The effective depth of each pipeline pass is shorter, so the bubble overhead is smaller relative to the useful computation.
The tradeoff is that each device now sends and receives activation tensors times per micro-batch instead of once, so inter-device communication volume increases by a factor of . For , the device sends activations at four different points in the forward pass (instead of two: once when passing from chunk 1 to the next device's chunk 2, once again when passing from chunk 3 to the next device's chunk 4, and similarly in the backward pass). Whether this tradeoff is favorable depends critically on the ratio of communication bandwidth to compute throughput. On systems with fast interconnects, such as NVLink between GPUs within the same node, the extra communication overhead may be small relative to the bubble reduction gains. On systems with slower interconnects, such as InfiniBand between nodes, the added communication can become a bottleneck that negates the bubble benefits.
The interleaved schedule also imposes a constraint that the number of micro-batches must be at least for full pipeline efficiency, and clean scheduling requires to be a multiple of . In large-scale training runs where is already large, these constraints are not binding. In smaller experiments or when memory constraints limit , the constraint can be inconvenient.


Comparing the Three Schedules
The three schedules form a clear progression in terms of capabilities and implementation complexity. Standard GPipe establishes the baseline: full forward phase followed by full backward phase, simple to implement, large activation memory. Standard 1F1B matches GPipe's bubble fraction but dramatically reduces activation memory by interleaving. Interleaved 1F1B goes further, reducing the bubble fraction at the cost of more communication.
Choosing between them in practice comes down to the constraints of the specific hardware setup. For a cluster with fast NVLink interconnects where communication cost is low, interleaved 1F1B with or is often the best choice, as the bubble reduction improves hardware utilization without a significant communication penalty. For clusters with slower interconnects where communication is a bottleneck, standard 1F1B may be preferable. GPipe is primarily used today when simplicity is the priority, such as for initial prototyping or for smaller models where activation memory is not a binding constraint.
Worked Example: Four-Stage 1F1B Pipeline
To make the scheduling mechanics concrete, let us trace a 4-stage pipeline with micro-batches through a standard 1F1B schedule. Label the stages S1 through S4 and the micro-batches B1 through B4. Use "F" to denote the forward pass of micro-batch and "B" to denote the backward pass of micro-batch .
The timeline proceeds as follows.
During the warmup phase, S1 runs F1. After S1 completes F1, S2 can start F1 while S1 starts F2. After S2 completes F1, S3 can start F1; and so on. The warmup expands the pipeline: after time steps, all stages are simultaneously occupied with forward passes. At this point, S4 completes F1 and can begin B1, which marks the start of the steady-state phase.
During the steady state, S4 alternates: after completing B1, it runs F2, then B2, then F3, then B3, then F4, then B4. S3 runs its forward passes one step behind S4 and its backward passes one step after S4. S1 and S2 follow the same pattern at their respective offsets.
During the drain phase, after S4 completes F4 (the last forward pass), the pipeline empties as backward passes drain from S4 back to S1. No new forward passes enter, and stages become idle one by one as their remaining backward passes complete.
The total number of time steps for the complete run is:
where:
- : forward and backward work time steps for four micro-batches
- : bubble time steps from the startup ramp and drain
Of the 14 total time steps, 8 represent useful forward or backward work and 6 are pipeline overhead, giving a bubble fraction of . The example is deliberately small to make the schedule traceable; in practice , which makes the bubble fraction much smaller.
The activation memory at any stage in this 1F1B schedule is bounded by the maximum number of concurrent forward-pass activations. In standard 1F1B, this maximum is micro-batches' worth of activations per stage, regardless of . At the peak of the warmup phase, S1 has processed F1, F2, F3, and F4 in the forward direction and is holding all four sets of activations. It then frees each set as the corresponding backward pass returns from the later stages. The peak is activations, confirming the memory bound.

Code Implementation
This section implements a simplified pipeline parallelism simulation that illustrates the scheduling mechanics without requiring multiple physical GPUs. We model each stage as a Python object, route micro-batches through the pipeline following GPipe and 1F1B schedules, and measure utilization metrics. The simulation is intentionally simplified to highlight the scheduling logic rather than actual distributed computing mechanics.
Imports and Setup
Stage and Micro-Batch Definitions
We define a lightweight data structure for micro-batches and a class that simulates a single pipeline stage. Each stage tracks how many forward and backward passes it has completed, which lets us compute utilization statistics afterward.
from dataclasses import dataclass
from typing import Optional
import numpy as np
@dataclass
class MicroBatch:
"""Represents a single micro-batch moving through the pipeline."""
batch_id: int
activations: np.ndarray
gradients: Optional[np.ndarray] = None
class PipelineStage:
"""Simulates one stage of a pipeline-parallel model."""
def __init__(self, stage_id: int, num_layers: int, hidden_size: int):
self.stage_id = stage_id
self.num_layers = num_layers
self.hidden_size = hidden_size
# Simple weight matrix to simulate layer computation
self.weights = np.random.randn(hidden_size, hidden_size) * 0.01
self.forward_count = 0
self.backward_count = 0
self.idle_steps = 0
self.active_steps = 0
def forward(self, micro_batch: MicroBatch) -> MicroBatch:
"""Run forward pass: apply a linear transformation to simulate layer computation."""
self.forward_count += 1
self.active_steps += 1
out = micro_batch.activations @ self.weights
return MicroBatch(batch_id=micro_batch.batch_id, activations=out)
def backward(self, micro_batch: MicroBatch) -> MicroBatch:
"""Run backward pass: simulate gradient computation."""
self.backward_count += 1
self.active_steps += 1
# Simplified gradient: propagate a scaled version backward
grad_input = micro_batch.gradients @ self.weights.T
return MicroBatch(
batch_id=micro_batch.batch_id,
activations=micro_batch.activations,
gradients=grad_input,
)GPipe Schedule Simulation
The GPipe simulation processes all micro-batches through the forward pass across all stages, then processes all micro-batches through the backward pass in reverse order. This sequential structure mirrors the flush-based schedule.
from typing import Dict
def simulate_gpipe(
num_stages: int, num_microbatches: int, hidden_size: int
) -> Dict:
"""Simulate the GPipe (flush-based) pipeline schedule."""
stages = [
PipelineStage(i, num_layers=4, hidden_size=hidden_size)
for i in range(num_stages)
]
# Initialize micro-batches with random activations
micro_batches = [
MicroBatch(m, np.random.randn(8, hidden_size))
for m in range(num_microbatches)
]
total_time_steps = 0
schedule_log = []
# Forward phase: all micro-batches complete forward pass
for stage_idx in range(num_stages):
for mb in micro_batches:
result = stages[stage_idx].forward(mb)
mb.activations = result.activations
total_time_steps += 1
schedule_log.append(("F", stage_idx, mb.batch_id))
# Attach synthetic gradients for backward pass
for mb in micro_batches:
mb.gradients = np.ones_like(mb.activations)
# Backward phase: all micro-batches complete backward pass
for stage_idx in reversed(range(num_stages)):
for mb in reversed(micro_batches):
result = stages[stage_idx].backward(mb)
mb.gradients = result.gradients
total_time_steps += 1
schedule_log.append(("B", stage_idx, mb.batch_id))
useful_steps = 2 * num_stages * num_microbatches
bubble_steps = total_time_steps - useful_steps
utilization = (
useful_steps / total_time_steps if total_time_steps > 0 else 0.0
)
return {
"schedule": "GPipe",
"total_time_steps": total_time_steps,
"bubble_steps": bubble_steps,
"utilization": utilization,
"stages": stages,
"schedule_log": schedule_log,
}1F1B Schedule Simulation
The 1F1B simulation interleaves forward and backward passes. After the warmup phase fills the pipeline, the steady state alternates between processing one new forward micro-batch and completing one backward pass for the oldest pending micro-batch.
from collections import defaultdict
def simulate_1f1b(
num_stages: int, num_microbatches: int, hidden_size: int
) -> Dict:
"""Simulate the 1F1B pipeline schedule.
In steady state, each stage alternates: one forward pass followed by
one backward pass. This bounds activation memory to O(D) rather than O(M*D).
"""
stages = [
PipelineStage(i, num_layers=4, hidden_size=hidden_size)
for i in range(num_stages)
]
# Track activation states for each micro-batch at each stage
activations = {
m: np.random.randn(8, hidden_size) for m in range(num_microbatches)
}
gradients = {}
schedule_log = []
total_time_steps = 0
in_flight_activations = defaultdict(list)
# Warmup phase: fill pipeline with forward passes
warmup_count = min(num_stages, num_microbatches)
for warmup_step in range(warmup_count):
mb_id = warmup_step
for s in range(warmup_step + 1):
result = stages[s].forward(MicroBatch(mb_id, activations[mb_id]))
activations[mb_id] = result.activations
in_flight_activations[s].append(mb_id)
total_time_steps += 1
schedule_log.append(("F", s, mb_id))
# Steady state: interleave F and B passes
next_forward = warmup_count
next_backward = 0
while next_backward < num_microbatches:
# Forward pass for next micro-batch (if available)
if next_forward < num_microbatches:
mb_id = next_forward
for s in range(num_stages):
result = stages[s].forward(
MicroBatch(mb_id, activations[mb_id])
)
activations[mb_id] = result.activations
total_time_steps += 1
schedule_log.append(("F", s, mb_id))
next_forward += 1
# Backward pass for oldest pending micro-batch
mb_id = next_backward
grad = np.ones_like(activations[mb_id])
for s in reversed(range(num_stages)):
result = stages[s].backward(
MicroBatch(mb_id, activations[mb_id], grad)
)
grad = result.gradients
total_time_steps += 1
schedule_log.append(("B", s, mb_id))
gradients[mb_id] = grad
next_backward += 1
useful_steps = 2 * num_stages * num_microbatches
bubble_steps = max(0, total_time_steps - useful_steps)
utilization = (
useful_steps / total_time_steps if total_time_steps > 0 else 0.0
)
return {
"schedule": "1F1B",
"total_time_steps": total_time_steps,
"bubble_steps": bubble_steps,
"utilization": utilization,
"stages": stages,
"schedule_log": schedule_log,
}Running the Comparison
We run both schedules across a range of micro-batch counts and compare their utilization against the theoretical prediction.
# Configuration
num_stages = 4
hidden_size = 64
# Run both schedules across different micro-batch counts
microbatch_counts = [1, 2, 4, 8, 16, 32]
gpipe_results = []
fb1b_results = []
for M in microbatch_counts:
gp = simulate_gpipe(num_stages, M, hidden_size)
fb = simulate_1f1b(num_stages, M, hidden_size)
gpipe_results.append(gp)
fb1b_results.append(fb)
# Compute theoretical bubble fractions
theoretical_bubble = [
(num_stages - 1) / (M + num_stages - 1) for M in microbatch_counts
]Pipeline configuration: 4 stages, hidden_size=64 M | GPipe Util | 1F1B Util | Theoretical Bubble ----------------------------------------------------- 1 | 100.0% | 160.0% | 75.0% 2 | 100.0% | 145.5% | 60.0% 4 | 100.0% | 123.1% | 42.9% 8 | 100.0% | 110.3% | 27.3% 16 | 100.0% | 104.9% | 15.8% 32 | 100.0% | 102.4% | 8.6%
As the number of micro-batches grows, utilization rises steadily toward 100% for both schedules. With only 1 micro-batch, the pipeline bubble is severe: the bubble fraction is for stages, leaving most device time idle. With 16 micro-batches, utilization climbs above 80%, and with 32 micro-batches, it exceeds 90%. The simulation confirms the theoretical prediction: increasing is the primary lever for reducing pipeline bubble overhead.
Note that the 1F1B simulation shows similar utilization to GPipe in this simple sequential model, as expected from the formula. In a real distributed system, however, 1F1B would allow much larger before running out of activation memory, because its activation memory bound is (constant in ) rather than (linear in ). The ability to use large without hitting memory limits is where 1F1B provides its practical advantage.
Activation Memory Comparison
The activation memory calculation makes the 1F1B advantage tangible. For realistic LLM configurations, the difference between GPipe and 1F1B activation memory can be the difference between a training run that fits in GPU memory and one that crashes with an out-of-memory error.
def estimate_activation_memory(
num_stages: int,
num_microbatches: int,
micro_batch_size: int,
seq_len: int,
hidden_size: int,
bytes_per_element: int = 2, # bfloat16
) -> Dict:
"""Estimate peak activation memory for GPipe vs 1F1B."""
# Activation size per micro-batch per stage (bytes)
activation_per_mb_per_stage = (
micro_batch_size * seq_len * hidden_size * bytes_per_element
)
# GPipe: must store all M micro-batches' activations across all D stages
# simultaneously (they are all needed for the eventual backward pass)
gpipe_peak = num_microbatches * num_stages * activation_per_mb_per_stage
# 1F1B: at most D micro-batches' activations in flight at any stage
# (because backward passes interleave with forward passes, freeing memory early)
fb1b_peak = num_stages * activation_per_mb_per_stage
return {
"gpipe_gb": gpipe_peak / (1024**3),
"fb1b_gb": fb1b_peak / (1024**3),
"reduction_factor": num_microbatches, # 1F1B is M times cheaper
}# Realistic LLM dimensions: GPT-3 scale
realistic_config = {
"num_stages": 8,
"num_microbatches": 16,
"micro_batch_size": 4,
"seq_len": 2048,
"hidden_size": 12288,
"bytes_per_element": 2, # bfloat16
}
memory_est = estimate_activation_memory(**realistic_config)Activation memory estimate (GPT-3 scale, 8 stages, M=16 micro-batches) GPipe peak activation memory : 24.0 GB 1F1B peak activation memory : 1.5 GB Memory reduction factor : 16x GPipe requires storing all M*D micro-batch activations simultaneously. 1F1B bounds activation memory to D stages regardless of M.
The memory numbers favor 1F1B. For a GPT-3-scale model partitioned into 8 stages and processing 16 micro-batches, GPipe requires 16 times more activation memory than 1F1B. In practical training, where every gigabyte matters for fitting the largest possible model, this difference often determines which schedule is feasible at all.

Key Parameters
The key parameters for pipeline parallelism are:
- num_stages (D): Number of pipeline stages, equal to the number of devices. Larger increases memory efficiency (each stage holds fewer parameters) but also increases the bubble fraction for a fixed . Adding more stages therefore requires a proportionally larger to maintain the same device utilization.
- num_microbatches (M): Number of micro-batches per optimizer step. Larger reduces the bubble fraction but increases per-step latency (you must process more micro-batches before taking a weight update) and, for GPipe without checkpointing, increases activation memory linearly.
- micro_batch_size: Size of each individual micro-batch, in sequences. Smaller values reduce per-stage activation memory but may underutilize GPU tensor cores, which need large matrix dimensions to operate efficiently. The total effective batch size equals .
- schedule: GPipe or 1F1B. GPipe separates forward and backward phases cleanly; 1F1B interleaves them to bound activation memory at stages' worth of data, enabling much larger values without running out of memory.
- V (virtual stages): For interleaved 1F1B, the number of non-contiguous model chunks per device. Larger reduces the bubble fraction further (by factor ) but increases inter-device communication volume by factor .
Limitations and Practical Considerations
Pipeline parallelism solves the problem of fitting very deep models across multiple devices, but it introduces its own set of engineering challenges that practitioners must manage carefully.
The Irreducible Pipeline Bubble
The pipeline bubble remains the most fundamental limitation. Even with a large number of micro-batches, the bubble fraction is never zero. For a 16-stage pipeline with 128 micro-batches, the bubble fraction is still . This unavoidable overhead means that pipeline parallelism alone cannot achieve linear scaling. Doubling the number of pipeline stages doubles the parameter capacity but does not double throughput, because the bubble grows proportionally with while must also grow proportionally to maintain the same bubble fraction.
In practice, the sweet spot for pipeline depth is typically 4 to 16 stages. Below 4 stages, the parameter capacity per device may be insufficient for the target model. Above 16 stages, the bubble becomes difficult to amortize without a very large number of micro-batches, which increases per-step latency, reduces the frequency of weight updates, and can degrade training stability for certain learning rate schedules. Systems like Megatron-LM typically use 8 to 16 pipeline stages for models in the 100-billion-to-trillion-parameter range.
Weight Staleness in Asynchronous Variants
Weight staleness is a subtler problem that arises in asynchronous variants of pipeline parallelism. In synchronous pipeline parallelism (the standard approach), all micro-batches see the same parameter values throughout the optimizer step, and the weight update happens only after all micro-batches have completed. This is clean and mathematically well-defined, but it requires waiting for the entire pipeline to drain before updating weights.
Some asynchronous schedules, like PipeDream's weight stashing approach, attempt to overlap weight updates with ongoing micro-batch processing. The idea is that while stage is processing forward passes for micro-batch , stage could be using an updated weight matrix that incorporates gradients from micro-batch . This means different stages of the pipeline see different versions of the model weights simultaneously, introducing a form of gradient inconsistency. The updated weights must be "stashed" so that the backward pass for each micro-batch uses the same weights that were used during its forward pass. Weight stashing adds memory overhead (each in-flight micro-batch requires its own copy of the stage's weights) and makes checkpointing more complex. Asynchronous schedules can potentially achieve better device utilization in some configurations, but the gradient staleness generally requires additional stabilization measures, and for large models the weight stashing memory overhead can be prohibitive.
Interaction with Tensor and Data Parallelism
The interaction between pipeline parallelism and other forms of parallelism requires careful orchestration. In production, large-scale training systems combine all three strategies simultaneously. Pipeline parallelism partitions layers across devices along the depth dimension. Tensor parallelism (covered in the previous chapter) partitions individual weight matrices within each layer across devices along the width dimension. Data parallelism replicates the full pipeline-and-tensor-parallel group across multiple independent replicas, each processing different mini-batches.
This three-dimensional parallelism, popularized by the Megatron-LM framework, distributes computation across all available dimensions. The devices are organized into a 3D mesh, where one dimension corresponds to tensor-parallel degree, another to pipeline-parallel degree, and the third to data-parallel degree. A cluster with 3,072 GPUs might be configured as 8-way tensor parallelism times 12-way pipeline parallelism times 32-way data parallelism. Each dimension of the mesh requires different communication patterns: tensor parallelism requires all-reduce within a node (typically using NVLink), pipeline parallelism requires point-to-point activations across nodes (using InfiniBand), and data parallelism requires all-reduce of gradients across replicas (also using InfiniBand, but less frequently due to gradient accumulation).
Configuring these dimensions correctly requires understanding the communication topology of the cluster, the memory requirements of the model, and the target throughput. There is no universal formula: the best configuration depends on model size, cluster size, interconnect bandwidth, and training batch size. Tools like Megatron-LM's automatic parallelism configuration search can help find a good starting point.
Load Imbalance Across Stages
Load imbalance across stages is a practical concern that can significantly degrade throughput. If one stage takes 20% longer than the others due to a larger layer or a different computation pattern, that stage becomes the bottleneck and all other stages must wait for it. The pipeline utilization is limited by the slowest stage, not the average stage.
For transformer models, the embedding layer at the beginning and the LM head (output projection plus softmax) at the end can have different compute profiles than the middle transformer blocks. The embedding layer involves a large vocabulary lookup, which is memory-bound. The LM head involves a large matrix multiplication from the hidden dimension to the vocabulary size, which can be compute-intensive for large vocabularies. Placing these special layers in stages with fewer transformer blocks can help equalize stage durations.
Profiling-driven partitioning is more reliable than uniform partitioning. You measure the actual wall-clock time per layer for the specific model and hardware configuration, then solve the partition problem by minimizing the maximum stage duration. Dynamic programming solves this exactly in time, where is the number of layers and is the number of stages.
Debugging and Reproducibility
Pipeline parallelism adds engineering complexity that is absent from single-device training. The model state at any point in training is distributed across devices, so inspecting the full model requires gathering tensors from multiple GPUs. Debugging gradient correctness requires comparing activations at stage boundaries, which involves coordinating logging across multiple processes.
A subtle class of bugs arises from incorrect synchronization. If one stage falls behind due to a scheduling issue or hardware variability, the entire pipeline stalls. Unlike a crash (which produces an obvious error), a subtle timing issue may manifest only as reduced throughput, making it hard to detect. Monitoring per-stage throughput and the fraction of time each stage spends idle can reveal these issues.
Numerical reproducibility is also harder to ensure in pipeline parallelism than in single-device training. The order in which gradients are accumulated across micro-batches may vary slightly depending on hardware timing, leading to different floating-point summation orders and therefore slightly different numerical results. For research purposes (where exact reproducibility is important for comparing runs), this nondeterminism must be controlled by fixing seeds and using deterministic algorithms throughout.
The Path to 3D Parallelism in Production
Despite these challenges, pipeline parallelism has become a standard component of large-scale training infrastructure. The models that define the modern era of large language models, including GPT-3, Gopher, PaLM, Megatron-Turing NLG, and their successors, all relied on pipeline parallelism combined with tensor and data parallelism to distribute training across thousands of GPUs. The pipeline parallel technique has been refined from GPipe's original flush-based approach through 1F1B's memory-efficient interleaving to the sophisticated interleaved schedule with virtual stages, each refinement motivated by real engineering constraints encountered in training progressively larger models.
The next chapter extends this distributed training picture by examining the communication optimization strategies, specifically gradient compression, all-reduce algorithms, and network topology-aware scheduling, that make cross-device synchronization fast enough to avoid bottlenecking on network bandwidth as the device count scales into the thousands.
Summary
Pipeline parallelism splits a deep model vertically along its layer dimension, assigning contiguous groups of layers to separate devices. Each device holds a fraction of the total parameters, enabling models that far exceed a single GPU's memory capacity to be trained in practice.
The core challenge is the pipeline bubble: the idle time that accumulates because later stages must wait for earlier stages to produce activations, and because gradients must drain backward through all stages before the next optimizer step. The bubble fraction is , where is the number of stages and is the number of micro-batches. Increasing is the primary strategy for driving the bubble fraction toward zero, at the cost of a larger effective batch size and more time between weight updates.
GPipe addresses the bubble by processing all micro-batches through the complete forward pass before beginning any backward passes. This achieves good device utilization for large but requires storing all micro-batch activations simultaneously, leading to activation memory that scales as . GPipe uses gradient checkpointing to manage this memory at the cost of roughly 33% additional compute.
The 1F1B schedule interleaves forward and backward passes to bound activation memory at stages' worth of data regardless of . The bubble fraction is identical to GPipe, but the memory efficiency of 1F1B makes it the standard choice for large model training: it allows much larger values without running out of activation memory, which in turn allows smaller micro-batches and better tensor core utilization.
The interleaved 1F1B variant assigns non-contiguous layer chunks to each device, reducing the bubble fraction by a factor equal to the number of chunks per device. This further reduces idle time at the cost of proportionally more inter-device communication volume. Whether the tradeoff is favorable depends on the interconnect bandwidth relative to compute throughput.
Key takeaways:
- Pipeline parallelism assigns consecutive layer groups to separate devices, allowing parameter count to scale linearly with device count while keeping per-device memory fixed
- The pipeline bubble is unavoidable but shrinks as the ratio grows; practical training runs use to keep bubble fractions below 10%
- GPipe and 1F1B achieve the same bubble fraction; 1F1B uses times less activation memory, making it the preferred schedule for large models
- The interleaved schedule reduces the bubble fraction by factor at the cost of times more communication
- In production systems, pipeline parallelism is combined with tensor parallelism and data parallelism into 3D parallelism, allowing training of models with hundreds of billions of parameters across thousands of GPUs
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about pipeline parallelism.
Pipeline Parallelism Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!