Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B

Michael BrenndoerferJanuary 20, 202654 min read

Part of Language AI Handbook

Explains how pipeline parallelism splits deep models across devices, manages bubble overhead with micro-batching, and compares GPipe vs 1F1B schedules.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Pipeline Parallelism: Stages, Micro-Batching, GPipe, and 1F1B

Training large language models requires distributing computation across multiple devices, and you have already seen two complementary strategies for doing this. Data parallelism splits the training dataset across workers so each device processes a different mini-batch while sharing gradient updates. Tensor parallelism splits individual weight matrices across devices so each GPU holds a slice of a single layer's parameters. Both strategies are powerful, but neither directly addresses a fundamental bottleneck: when a model is so large that a single layer or a modest contiguous block of layers cannot fit in one GPU's memory, and when the sequential depth of the network creates idle time because each layer must wait for the previous layer to finish before it can begin.

Pipeline parallelism takes a different approach. Instead of splitting data or splitting weight matrices, it splits the model itself along its depth, assigning contiguous groups of layers to different devices. The first device holds layers 1 through 8, the second holds layers 9 through 16, the third holds layers 17 through 24, and so on. This lets the model's total parameter count scale linearly with the number of pipeline stages, independent of any single device's memory capacity. The challenge, as we will see, is that a naive implementation leaves most devices idle most of the time, waiting for activations to arrive from the previous stage. Understanding how to minimize this idle time is the central engineering problem of pipeline parallelism, and the history of solutions to it has shaped how the largest models in existence are trained today.

The intuition behind pipeline parallelism is borrowed from industrial assembly lines. A car factory does not wait for one vehicle to be fully assembled before starting the next. Instead, different vehicles are in different stages of assembly simultaneously: the welding station, the painting station, and the final inspection station are all busy at the same time, each working on a different car. Pipeline parallelism applies exactly this logic to deep neural networks. While one micro-batch is being processed by the second stage of the model, the first stage can already be working on the next micro-batch. The goal is to keep every stage busy simultaneously, maximizing the fraction of time all devices are performing useful computation rather than sitting idle waiting for their neighbor to finish.

The history of pipeline parallelism in deep learning spans several important papers. The GPipe paper from Google Brain (2019) introduced a clean flush-based schedule that proved pipeline parallelism was practical at scale. The PipeDream paper from Microsoft Research (2019) demonstrated asynchronous variants with weight stashing. Megatron-LM from NVIDIA refined the approach with the 1F1B schedule and later the interleaved schedule, which became the standard for training hundreds-of-billions-parameter models. Understanding the progression from GPipe through 1F1B to interleaved scheduling is not just historical curiosity: each step was motivated by a concrete engineering constraint, and knowing those constraints helps you understand which schedule to choose for your own training runs.

Model Partitioning into Stages

The first design decision in pipeline parallelism is how to divide the model into stages. A stage is a contiguous sequence of layers that runs on a single device. The boundary between stages is called a pipeline cut point, and it determines both the computational load per device and the volume of data that must be communicated between devices at each boundary.

Pipeline Stage

A pipeline stage is a contiguous block of model layers assigned to a single device. During forward propagation, each stage receives activations from the previous stage, processes them through its layers, and passes the result to the next stage. During backward propagation, gradient flow reverses this path, with each stage receiving gradients from the next stage and passing gradients to the previous stage.

Choosing where to place cut points involves balancing several competing factors. If the stages are imbalanced in computational cost, the slowest stage becomes the bottleneck: faster stages finish quickly but cannot proceed until they receive either activations (during the forward pass) or gradients (during the backward pass) from the slow stage. The goal is to partition layers such that each stage takes approximately the same wall-clock time per sample. This is called load balancing, and it is more subtle than it might initially appear. Two stages with the same number of layers can have very different computational costs if one contains attention layers operating on long sequences and another contains feed-forward layers with large hidden dimensions. Similarly, the first stage that includes the embedding lookup and the last stage that includes the output projection and softmax may have compute profiles that differ from the middle transformer blocks.

For a transformer model with LL layers and DD devices, the simplest strategy assigns L/DL/D consecutive layers to each device. This achieves perfect load balance if every layer is identical in compute cost, which is approximately true for standard transformer blocks. In practice, the embedding layer and the output projection layer may have different costs than interior transformer blocks, and some architectures include routing layers (as in mixture-of-experts models) that behave very differently from standard dense layers. A profiling-driven partition, where you measure the actual wall-clock time per layer on representative inputs and then use a dynamic programming algorithm to find the partition that minimizes the maximum stage duration, is more reliable than a uniform assignment. The problem of finding the optimal partition is formally equivalent to a bin-packing variant, and it can be solved exactly in polynomial time given per-layer timing measurements.

Communication Volume at Stage Boundaries

The communication cost between stages is determined by the size of the activation tensor passed at each cut point. For a transformer with sequence length TT, batch size BB, and hidden dimension HH, the tensor passed between stages has shape (B,T,H)(B, T, H), so the communication volume per forward pass is:

Vcomm=B⋅T⋅H⋅bytes_per_elementV_{\text{comm}} = B \cdot T \cdot H \cdot \text{bytes\_per\_element}

where:

  • BB: micro-batch size (number of sequences per micro-batch)
  • TT: sequence length in tokens
  • HH: hidden dimension (model width)
  • bytes_per_element\text{bytes\_per\_element}: 2 for bfloat16, 4 for float32

For a GPT-3-scale model with H=12288H = 12288, T=2048T = 2048, and B=4B = 4, using bfloat16, the communication tensor has size 4×2048×12288×2=2014 \times 2048 \times 12288 \times 2 = 201 MB per stage boundary. This must flow across the interconnect linking the two pipeline stages. On NVLink, which provides roughly 600 GB/s of bidirectional bandwidth between GPUs in the same node, this transfer takes about 0.33 milliseconds. On InfiniBand HDR (200 Gb/s), the same transfer takes closer to 8 milliseconds. The compute time per stage for this same configuration is several hundred milliseconds on a single A100. The ratio of communication to compute time determines whether the pipeline is compute-bound or communication-bound: on NVLink it is clearly compute-bound (the interconnect is fast enough to overlap communication with computation), while on slower networks the communication overhead can become a meaningful fraction of total time.

This is typically much smaller than the communication volume in data parallelism, which involves gradient tensors proportional to total parameter count. A GPT-3-scale model with 175 billion parameters requires 350 GB of gradients (in bfloat16) per all-reduce operation in data parallelism. Compared to 201 MB per pipeline boundary per micro-batch, the pipeline communication overhead is orders of magnitude smaller. Pipeline stage-to-stage communication is also point-to-point: only two devices are involved at each boundary, and only one tensor flows per forward pass and one tensor flows per backward pass. This contrasts with the all-reduce patterns of data parallelism, which involve all devices simultaneously.

Memory Implications of Staging

Each device in a pipeline stores the parameters and buffers for its assigned layers, plus the activation tensors needed for the backward pass. The parameter memory per device scales with the total model size divided by the number of stages:

param memory per stage=PtotalD\text{param memory per stage} = \frac{P_{\text{total}}}{D}

where:

  • PtotalP_{\text{total}}: total number of model parameters (in bytes, accounting for numerical precision)
  • DD: number of pipeline stages

This fraction is the primary motivation for pipeline parallelism: models that cannot fit on one device can be spread across multiple devices, with each device using only 1/D1/D of the total parameter memory. A model with 175 billion parameters in 16-bit precision requires roughly 350 GB of memory, which exceeds the capacity of any single GPU available as of 2024. Spreading it across 8 pipeline stages reduces per-device parameter memory to about 44 GB, which fits comfortably on modern 80 GB A100 or H100 GPUs. With 16 stages, per-device parameter memory drops to about 22 GB.

Beyond raw parameter storage, the optimizer state memory must also be considered. With the Adam optimizer, each parameter requires two additional floating-point values (the first and second moment estimates). When using mixed-precision training (model weights in bfloat16 but optimizer states in float32), the optimizer state memory is larger than the parameter memory: each parameter requires 4 bytes for the float32 master weight, 4 bytes for the first moment, and 4 bytes for the second moment, totaling 12 bytes per parameter in optimizer state alone. For a stage holding Ptotal/DP_{\text{total}}/D parameters, the total memory requirement including optimizer state is:

Mstage=PtotalD⋅(2+12)=14⋅PtotalDM_{\text{stage}} = \frac{P_{\text{total}}}{D} \cdot (2 + 12) = \frac{14 \cdot P_{\text{total}}}{D}

where the factor of 2 comes from the bfloat16 model parameters and the factor of 12 comes from float32 optimizer states. For a 175 billion parameter model split across 8 stages, this gives roughly (14×350 GB)/8≈613 GB/8≈76.6 GB(14 \times 350 \text{ GB}) / 8 \approx 613 \text{ GB} / 8 \approx 76.6 \text{ GB} per stage, which is tighter but still feasible on 80 GB GPUs.

The activation memory requires more careful analysis. During the forward pass, each stage must retain its intermediate activations until the corresponding backward pass arrives. In a naive sequential pipeline, stage kk must hold its activations in memory from the moment it finishes the forward pass until the backward pass through all later stages has completed and the gradient has been returned to stage kk. This can be a long wait, especially for early stages in a deep pipeline. The key challenge therefore extends beyond the per-stage parameter and optimizer state memory, which are fixed and predictable, to the activation memory, which depends on the schedule and can vary by an order of magnitude depending on implementation choices.

The Pipeline Bubble Problem

The most important challenge in pipeline parallelism is not communication or memory but time: the problem of stages sitting idle while waiting for work. To understand why this happens, it is worth tracing through exactly what happens when you run a model through a naive pipeline.

Imagine four pipeline stages labeled S1 through S4. A single input mini-batch B1B_1 enters S1 at time step 1. S2 cannot begin until S1 finishes, S3 cannot begin until S2 finishes, and so on. The result is a startup pattern where:

  • At time step 1: S1 is working, S2, S3, and S4 are idle.
  • At time step 2: S2 is working on the result from S1, while S1 is idle (finished its pass and waiting), and S3 and S4 are still idle.
  • At time step 3: S3 is working, while S1, S2, and S4 are all idle.
  • At time step 4: S4 is working, while S1, S2, and S3 are all idle.

After the forward pass completes, the backward pass flows in reverse: S4 computes gradients first, then passes them to S3, and so on back to S1. This also creates a sequential chain of idle time, now in reverse order. The combined result is that the startup and teardown phases of the pipeline are almost entirely wasted: at any given moment, only one stage is doing useful work.

Pipeline Bubble

A pipeline bubble is the idle time that accumulates across pipeline stages due to the sequential dependency between stages. Bubbles appear at the beginning and end of every micro-batch sequence. For a pipeline with DD stages, the naive single-batch approach wastes (D−1)/D(D-1)/D of total device-time on bubbles. With four stages, three-quarters of all device-time is wasted.

To quantify the bubble fraction precisely, consider DD pipeline stages processing MM micro-batches. The startup ramp contributes D−1D - 1 idle time steps before all stages can participate (the first stage begins immediately, but the kk-th stage cannot begin until k−1k - 1 time steps have passed because it must wait for activations from earlier stages). The drain phase at the end is symmetric, contributing another D−1D - 1 idle time steps as gradients move backward through the pipeline. Each stage performs 2M2M useful operations—MM forward and MM backward—and the full schedule lasts 2M+2(D−1)2M + 2(D - 1) time steps. The bubble fraction is therefore:

bubble fraction=2(D−1)2M+2(D−1)=D−1M+D−1\text{bubble fraction} = \frac{2(D-1)}{2M + 2(D-1)} = \frac{D-1}{M + D-1}

where:

  • DD: number of pipeline stages (equal to the number of devices)
  • MM: number of micro-batches processed in each pipeline run
  • 2(D−1)2(D-1): total idle time slots from the startup ramp and teardown drain

For large MM, the denominator is dominated by MM, so the bubble fraction approaches (D−1)/M(D-1)/M, which converges to zero as MM grows. This asymptotic behavior is the key insight: the bubble overhead becomes negligible when MM is much larger than DD.

To see this concretely with numbers: for D=4D = 4 stages and M=1M = 1 micro-batch, the bubble fraction is 3/4=0.753/4 = 0.75. You are paying for four GPUs but using only about 25% of their combined compute capacity. Increasing MM to 8 micro-batches brings the bubble fraction to 3/11≈0.273/11 \approx 0.27. Increasing to 32 micro-batches brings it to 3/35≈0.0863/35 \approx 0.086. The improvement continues, but each additional doubling of MM produces diminishing returns: going from M=32M = 32 to M=64M = 64 reduces the bubble from about 9% to about 4.5%.

It is worth understanding why the bubble cannot be eliminated entirely with any finite number of micro-batches. The startup ramp is unavoidable: the first micro-batch must travel through all DD stages sequentially before the pipeline can be considered "full." During this ramp-up, early stages are idle after finishing their portion of the first micro-batch and before the second micro-batch arrives. Similarly, the drain ramp at the end is unavoidable: after the last forward pass, backward passes flow back through the pipeline, and late stages finish first while early stages wait for gradients to arrive from later stages. As MM grows, these fixed startup and drain costs become a smaller fraction of the schedule and approach zero asymptotically.

Out[3]:
Visualization
Line chart showing decreasing bubble fraction as micro-batch count increases for D=2,4,8,16 stages.
Pipeline bubble fraction as a function of micro-batch count M for different pipeline depths D. All curves follow the formula (D-1)/(M+D-1), converging toward zero as M grows. With M=1, a 4-stage pipeline wastes 75% of device time; increasing to M=32 drops waste below 10%. Deeper pipelines require larger M to reach the same efficiency level.

Micro-Batching: Filling the Bubble

The solution to the bubble problem is to process multiple micro-batches simultaneously. A micro-batch is a small chunk of the full mini-batch. Instead of waiting for one full mini-batch to complete its forward pass before starting the next one, we split the mini-batch into MM equal micro-batches and inject them into the pipeline as fast as the first stage can produce them.

Micro-Batch

A micro-batch is a subdivision of the full training mini-batch. When using pipeline parallelism, the mini-batch is split into MM equal parts. Each part flows through the pipeline independently, and the gradients from all MM micro-batches are accumulated before the optimizer step is applied. From the optimizer's perspective, the result is identical to processing the full mini-batch at once.

The pipeline schedule looks quite different with multiple micro-batches in flight. While S1 is processing micro-batch B2B_2, S2 is simultaneously processing micro-batch B1B_1. While S1 processes B3B_3, S2 processes B2B_2 and S3 processes B1B_1. The stages are kept busy by overlapping the processing of different micro-batches on different stages. The bubble still exists at the beginning and end of the sequence (there is no way to avoid the startup ramp for a strictly sequential pipeline), but the bubble fraction decreases as MM increases because the useful work time grows while the ramp time stays constant.

The key insight is that all MM micro-batches share the same model weights and contribute to the same gradient update. At the end of the pipeline run, all stages have produced their partial gradients, these are accumulated across micro-batches, and the optimizer takes a single parameter update step. From the optimizer's perspective, the pipeline processed one mini-batch of size Btotal=M×BmicroB_{\text{total}} = M \times B_{\text{micro}}. The decomposition into micro-batches is an implementation detail for managing pipeline utilization. It is transparent to the loss function and the optimizer.

One important subtlety is the choice of micro-batch size. Smaller micro-batches mean more micro-batches for the same total mini-batch size, which reduces the bubble fraction. But very small micro-batches can underutilize the GPU's tensor cores. Modern GPU tensor cores achieve peak efficiency when matrix dimensions are large multiples of 8 or 16 (and ideally 64 or 128 for the best hardware utilization on A100s and H100s). A micro-batch with too few sequences may result in matrix multiplications with tiny leading dimensions, causing the GPU to spend most of its time on memory access rather than computation. There is a sweet spot where micro-batches are large enough to keep the GPU's arithmetic units busy but small enough that MM is large enough to suppress the bubble. Finding this sweet spot requires profiling on the specific hardware configuration, though a rule of thumb is that the micro-batch size should be at least 8 sequences for standard transformer models.

Gradient Accumulation and Weight Updates

The relationship between micro-batches, mini-batches, and gradient accumulation is worth understanding precisely, because it affects both memory usage and training stability.

During the forward pass for micro-batch mm, each stage kk computes the activations ak,m\mathbf{a}_{k,m} and passes them to stage k+1k+1. Each stage also stores ak,m\mathbf{a}_{k,m} for use during the backward pass. During the backward pass for micro-batch mm, stage kk receives gradients ∇ak+1,m\nabla \mathbf{a}_{k+1,m} from stage k+1k+1, computes gradients with respect to its own parameters, and passes ∇ak,m\nabla \mathbf{a}_{k,m} to stage k−1k-1.

The parameter gradients are accumulated across micro-batches. After all MM micro-batches have completed their backward passes, stage kk holds the accumulated gradient:

∇Wk=1M∑m=1M∇Wk,m\nabla W_k = \frac{1}{M} \sum_{m=1}^{M} \nabla W_{k,m}

where:

  • ∇Wk\nabla W_k: the accumulated gradient used for the parameter update at stage kk
  • ∇Wk,m\nabla W_{k,m}: the contribution to the gradient from micro-batch mm at stage kk
  • MM: the number of micro-batches per optimizer step

This accumulated gradient is mathematically equivalent to computing the gradient over the full mini-batch simultaneously, assuming the micro-batch splits are uniform and all micro-batches use the same parameter values. The optimizer step then proceeds exactly as it would in single-device training. This equivalence is what makes synchronous pipeline parallelism a valid distributed training strategy: the model sees the same gradient signal it would see in non-pipelined training, just computed in parallel across devices.

One practical consideration is that different micro-batches could in principle encounter different parameter values if weights are updated between micro-batches. The standard synchronous pipeline parallelism approach avoids this by processing all MM micro-batches before taking any weight update step. This ensures gradient coherence: all micro-batches see the same model weights, so their gradients can be averaged. Asynchronous variants that update weights between micro-batches introduce a form of weight staleness, where later micro-batches in the same optimizer step observe slightly different weights than earlier ones. Asynchronous schedules can potentially achieve better device utilization, but the staleness introduces gradient inconsistency that tends to harm convergence for large models. For this reason, synchronous pipeline parallelism is the standard approach in practice.

The Micro-Batch Size and Total Batch Size Relationship

Understanding the relationship between micro-batch size, the number of micro-batches, and the total effective batch size is essential for correctly configuring pipeline parallelism training runs.

If each micro-batch contains BμB_{\mu} sequences and there are MM micro-batches per optimizer step, the total batch size processed per step per pipeline replica is:

Btotal=M×BμB_{\text{total}} = M \times B_{\mu}

When this pipeline is also combined with data parallelism, where RR replicas of the full pipeline are run in parallel, the global effective batch size is:

Bglobal=R×M×BμB_{\text{global}} = R \times M \times B_{\mu}

Large-scale training runs often target a specific global batch size for training stability reasons (the learning rate and its schedule are typically tuned for a specific batch size). Pipeline parallelism and data parallelism together give you multiple levers to hit the target batch size: you can increase MM to reduce the bubble fraction, increase RR to add more data-parallel replicas, and adjust BμB_{\mu} to satisfy hardware utilization constraints. The flexibility to tune these independently, rather than treating the batch size as a single monolithic parameter, is one of the practical benefits of the 3D parallelism approach used by modern large model training frameworks.

Pipeline Schedules

The order in which forward and backward passes are scheduled across stages and micro-batches is called the pipeline schedule. Different schedules make different tradeoffs between bubble fraction, activation memory, and implementation complexity. The two most important schedules in practice are GPipe and 1F1B, with the interleaved 1F1B schedule providing a further refinement.

The schedule determines how busy each device is, how much memory it requires, and how long the pipeline takes to produce a gradient update. These tradeoffs are non-trivial: a schedule that minimizes the bubble fraction might require enormous activation memory, making it impractical for large models. A schedule that minimizes memory might introduce communication complexity that reduces efficiency. Understanding the specific tradeoffs of each schedule is essential for selecting the right one for a given training configuration, cluster topology, and model architecture.

GPipe: Flush-Based Schedule

GPipe (Google's Pipeline for Machine Learning) was one of the earliest systematic frameworks for pipeline parallelism in deep learning. The paper, published by researchers at Google Brain in 2019, demonstrated that models with billions of parameters could be trained reliably using pipeline parallelism with gradient checkpointing. Its schedule is conceptually straightforward: all MM micro-batches complete their forward passes before any backward passes begin. This "flush" design is easy to reason about, simple to implement, and was used to train some of the earliest large language models.

The full GPipe schedule for DD stages and MM micro-batches proceeds as follows. In the forward phase, micro-batches B1B_1 through BMB_M are injected sequentially into the pipeline. Stage 1 processes B1B_1, then B2B_2, then B3B_3, and so on, passing activations downstream as each micro-batch completes. By the time stage DD has processed BMB_M, all activations for all micro-batches are sitting in GPU memory across all stages. Only then does the backward phase begin. Stage DD computes gradients for BMB_M, then BM−1B_{M-1}, and so on in reverse order. Gradients flow backward through the pipeline until all micro-batches have completed their backward passes and all parameter gradients have been accumulated.

The bubble fraction for GPipe is:

bubble fractionGPipe=D−1M+D−1\text{bubble fraction}_{\text{GPipe}} = \frac{D - 1}{M + D - 1}

where:

  • DD: number of pipeline stages
  • MM: number of micro-batches

Note that this formula uses D−1D-1 in the numerator rather than 2(D−1)2(D-1). This is because in the GPipe flush-based schedule, the startup bubble and the drain bubble are shared: the same pipeline states that are idle during the forward-phase ramp-up become idle again during the backward-phase drain, but the two phases are not counted separately in the same way as a fully interleaved schedule. The effective bubble fraction formula is (D−1)/(M+D−1)(D-1)/(M+D-1), which is equivalent to the general formula when you account for the specific timing of the flush.

For D=4D = 4 and M=8M = 8, this gives (4−1)/(8+4−1)=3/11≈0.27(4-1)/(8+4-1) = 3/11 \approx 0.27. For D=4D = 4 and M=32M = 32, it gives 3/35≈0.0863/35 \approx 0.086. As MM grows large, the bubble fraction approaches zero.

The critical weakness of GPipe is its peak activation memory. Because all MM micro-batch activations across all DD stages must be stored simultaneously (they are all needed for the eventual backward phase), the total activation memory scales as:

activation memoryGPipe=M×D×Alayer\text{activation memory}_{\text{GPipe}} = M \times D \times A_{\text{layer}}

where:

  • MM: number of micro-batches
  • DD: number of pipeline stages
  • AlayerA_{\text{layer}}: activation memory per layer per micro-batch

For large MM and DD, this can easily exceed available GPU memory. Consider a GPT-3-scale model with D=8D = 8 stages and M=32M = 32 micro-batches. Each activation tensor has shape (Bμ,T,H)=(4,2048,12288)(B_{\mu}, T, H) = (4, 2048, 12288) in bfloat16, which is about 201 MB. The total activation memory is 32×8×201 MB≈51 GB32 \times 8 \times 201 \text{ MB} \approx 51 \text{ GB}, which on top of the parameter and optimizer state memory would exceed any single GPU.

GPipe addresses this through gradient checkpointing (also called activation recomputation). Instead of storing all intermediate activations during the forward pass, only certain checkpoint activations are kept. When the backward pass needs an intermediate activation that was not stored, the forward pass is rerun from the nearest checkpoint to recompute it. This reduces activation memory dramatically at the cost of additional computation.

Gradient Checkpointing

Gradient checkpointing is a technique that trades compute for memory. Instead of storing all intermediate activations during the forward pass, only a subset of checkpoint activations is saved. When the backward pass needs a discarded activation, it recomputes the forward pass from the nearest checkpoint. This reduces activation memory at the cost of approximately one additional forward pass per backward pass, roughly doubling the compute cost for the forward-pass portion of training.

In GPipe, gradient checkpointing is applied at the stage granularity: each stage stores only its input activation (the tensor it received from the previous stage) and recomputes all intermediate activations within the stage during the backward pass. With this approach, the activation memory per stage is just Astage-inputA_{\text{stage-input}} regardless of how many layers are in the stage or how many micro-batches are in flight. The recomputation cost is roughly one additional forward pass for every backward pass, which adds about 30-33% to total compute time (since the backward pass typically costs about twice as much as the forward pass, and you are adding one more forward pass).

The simplicity of GPipe made it the natural first implementation choice, and it demonstrated that pipeline parallelism was practical at scale. The activation memory limitations, combined with the compute overhead of gradient checkpointing, motivated the development of more memory-efficient schedules.

1F1B: One Forward, One Backward

The 1F1B (One Forward, One Backward) schedule, introduced by the Megatron-LM project at NVIDIA, addresses the memory problem of GPipe by interleaving forward and backward passes rather than separating them into two distinct phases. The core idea is that a device should begin its backward pass as soon as it has accumulated enough in-flight forward passes to make the backward pass possible, rather than waiting for all MM micro-batches to complete their forward phases first.

In the steady-state of a 1F1B schedule, each device alternates between one forward-pass micro-batch and one backward-pass micro-batch. At any point in time, each device is always working on one computation, and the peak number of activations stored in memory is bounded by the number of stages DD rather than the number of micro-batches MM.

The 1F1B schedule proceeds in three distinct phases.

The warmup phase fills the pipeline: stage 1 processes micro-batches B1B_1 through BDB_D in the forward direction, and stage DD performs its first backward pass after receiving BDB_D from all prior stages. This warmup phase is where the pipeline bubble accumulates, since early stages are busy on their forward passes while late stages are still waiting for activations to arrive. The warmup phase lasts for D−1D - 1 time steps of startup latency.

The steady-state phase then alternates forward and backward passes, maintaining full pipeline utilization. Every stage is either processing a forward pass or a backward pass at every time step. In the steady state, stage kk forwards micro-batch Bk+DB_{k+D} while simultaneously backward-passing micro-batch BkB_k. The stage is never idle during the steady state, as there is always either a new forward micro-batch arriving or a backward gradient ready to process.

The drain phase at the end completes the remaining backward passes as the pipeline empties. After the last forward pass (BMB_M reaching stage DD), no new forward passes enter, and the pipeline drains as gradients flow backward through each stage. The drain phase mirrors the startup bubble in duration: it also occupies D−1D - 1 time steps.

The bubble fraction for 1F1B is identical to GPipe:

bubble fraction1F1B=D−1M+D−1\text{bubble fraction}_{\text{1F1B}} = \frac{D - 1}{M + D - 1}

The advantage of 1F1B is not a smaller bubble but a much smaller peak activation memory. Because each stage begins processing backward passes before accumulating many pending forward-pass activations, the maximum number of in-flight activations at any stage is bounded by DD:

activation memory1F1B=D×Alayer\text{activation memory}_{\text{1F1B}} = D \times A_{\text{layer}}

where:

  • DD: number of pipeline stages
  • AlayerA_{\text{layer}}: activation memory per layer per micro-batch

Comparing to GPipe's M×D×AlayerM \times D \times A_{\text{layer}}, the 1F1B schedule saves a factor of MM in activation memory. For a typical training run with M=32M = 32 and D=8D = 8, 1F1B uses 32 times less activation memory than naive GPipe without checkpointing. This allows 1F1B to handle significantly larger MM without running out of memory, which in turn enables smaller micro-batches (for better tensor core utilization) while still achieving a low bubble fraction.

The reason for the memory difference becomes clear when you examine the schedule visually. In GPipe, by the time the backward phase begins, every stage has already processed all MM micro-batches in the forward direction and is holding all of their activations in memory simultaneously. In 1F1B, a stage begins its backward pass much sooner. Stage DD begins its first backward pass after receiving only its DD-th micro-batch's activations, not after receiving all MM micro-batches. As each backward pass completes at stage DD, the corresponding forward activations can be freed. The memory at any stage is therefore bounded by the maximum number of micro-batches simultaneously "in flight" at that stage, which is DD, not MM.

This is the key algorithmic insight of 1F1B: by scheduling backward passes early, you free activation memory early, keeping the peak memory footprint proportional to pipeline depth rather than to the number of micro-batches. For practical training runs where MM can be tens or hundreds of micro-batches, this memory savings is what makes 1F1B the default schedule for large model training.

The Interleaved 1F1B Schedule

The interleaved schedule is a further refinement of 1F1B introduced by Megatron-LM to reduce the bubble fraction at the cost of increased communication. Standard 1F1B assigns a contiguous block of L/DL/D layers to each device. The interleaved schedule instead assigns VV smaller, non-contiguous blocks of L/(D×V)L/(D \times V) layers each to every device. A single device might hold layers 1-4 and layers 33-36, with another device holding layers 5-8 and 37-40, and so on.

This arrangement means that each micro-batch traverses the pipeline VV times instead of once, because it must visit each device multiple times, once per model chunk assigned to that device. Each traversal is called a pipeline pass. The bubble fraction for the interleaved schedule improves by a factor of VV compared to standard 1F1B:

bubble fractioninterleaved=1V⋅D−1M+D−1\text{bubble fraction}_{\text{interleaved}} = \frac{1}{V} \cdot \frac{D - 1}{M + D - 1}

where:

  • VV: number of model chunks per device (also called virtual pipeline stages)
  • DD: number of devices
  • MM: number of micro-batches

With V=2V = 2, the interleaved schedule halves the bubble fraction. With V=4V = 4, it reduces the bubble to a quarter. The intuition is that by splitting the model into more, smaller chunks, the pipeline fills up faster relative to the total amount of work. The startup ramp covers the same number of time steps, but those time steps represent a smaller fraction of the total work because each "pass" through the pipeline covers only 1/V1/V of the model's layers. The effective depth of each pipeline pass is shorter, so the bubble overhead is smaller relative to the useful computation.

The tradeoff is that each device now sends and receives activation tensors VV times per micro-batch instead of once, so inter-device communication volume increases by a factor of VV. For V=2V = 2, the device sends activations at four different points in the forward pass (instead of two: once when passing from chunk 1 to the next device's chunk 2, once again when passing from chunk 3 to the next device's chunk 4, and similarly in the backward pass). Whether this tradeoff is favorable depends critically on the ratio of communication bandwidth to compute throughput. On systems with fast interconnects, such as NVLink between GPUs within the same node, the extra communication overhead may be small relative to the bubble reduction gains. On systems with slower interconnects, such as InfiniBand between nodes, the added communication can become a bottleneck that negates the bubble benefits.

The interleaved schedule also imposes a constraint that the number of micro-batches MM must be at least D×VD \times V for full pipeline efficiency, and clean scheduling requires MM to be a multiple of DD. In large-scale training runs where MM is already large, these constraints are not binding. In smaller experiments or when memory constraints limit MM, the constraint can be inconvenient.

Out[4]:
Visualization
Line plot comparing GPipe and 1F1B bubble fractions, showing identical overlapping curves that decrease with M.
Bubble fraction comparison between GPipe and standard 1F1B for D=4 pipeline stages across varying micro-batch counts. Both schedules share identical bubble fractions (D-1)/(M+D-1), confirming that 1F1B's primary advantage is memory efficiency rather than reduced bubble time.
Bar chart showing bubble fraction decreasing as virtual stages V increases from 1 to 8.
Bubble fraction for the interleaved 1F1B schedule at D=4, M=16, as a function of virtual stages V. Each doubling of V halves the bubble fraction, with diminishing absolute gains as V grows and communication overhead becomes a larger concern.

Comparing the Three Schedules

The three schedules form a clear progression in terms of capabilities and implementation complexity. Standard GPipe establishes the baseline: full forward phase followed by full backward phase, simple to implement, large activation memory. Standard 1F1B matches GPipe's bubble fraction but dramatically reduces activation memory by interleaving. Interleaved 1F1B goes further, reducing the bubble fraction at the cost of more communication.

Choosing between them in practice comes down to the constraints of the specific hardware setup. For a cluster with fast NVLink interconnects where communication cost is low, interleaved 1F1B with V=2V = 2 or V=4V = 4 is often the best choice, as the bubble reduction improves hardware utilization without a significant communication penalty. For clusters with slower interconnects where communication is a bottleneck, standard 1F1B may be preferable. GPipe is primarily used today when simplicity is the priority, such as for initial prototyping or for smaller models where activation memory is not a binding constraint.

Worked Example: Four-Stage 1F1B Pipeline

To make the scheduling mechanics concrete, let us trace a 4-stage pipeline with M=4M = 4 micro-batches through a standard 1F1B schedule. Label the stages S1 through S4 and the micro-batches B1 through B4. Use "Fkk" to denote the forward pass of micro-batch kk and "Bkk" to denote the backward pass of micro-batch kk.

The timeline proceeds as follows.

During the warmup phase, S1 runs F1. After S1 completes F1, S2 can start F1 while S1 starts F2. After S2 completes F1, S3 can start F1; and so on. The warmup expands the pipeline: after D−1=3D - 1 = 3 time steps, all stages are simultaneously occupied with forward passes. At this point, S4 completes F1 and can begin B1, which marks the start of the steady-state phase.

During the steady state, S4 alternates: after completing B1, it runs F2, then B2, then F3, then B3, then F4, then B4. S3 runs its forward passes one step behind S4 and its backward passes one step after S4. S1 and S2 follow the same pattern at their respective offsets.

During the drain phase, after S4 completes F4 (the last forward pass), the pipeline empties as backward passes drain from S4 back to S1. No new forward passes enter, and stages become idle one by one as their remaining backward passes complete.

The total number of time steps for the complete run is:

ttotal=2M+2(D−1)=2(4)+2(4−1)=8+6=14t_{\text{total}} = 2M + 2(D - 1) = 2(4) + 2(4 - 1) = 8 + 6 = 14

where:

  • 2M=82M = 8: forward and backward work time steps for four micro-batches
  • 2(D−1)=62(D-1) = 6: bubble time steps from the startup ramp and drain

Of the 14 total time steps, 8 represent useful forward or backward work and 6 are pipeline overhead, giving a bubble fraction of 6/14≈0.436/14 \approx 0.43. The example is deliberately small to make the schedule traceable; in practice M≫DM \gg D, which makes the bubble fraction much smaller.

The activation memory at any stage in this 1F1B schedule is bounded by the maximum number of concurrent forward-pass activations. In standard 1F1B, this maximum is D=4D = 4 micro-batches' worth of activations per stage, regardless of MM. At the peak of the warmup phase, S1 has processed F1, F2, F3, and F4 in the forward direction and is holding all four sets of activations. It then frees each set as the corresponding backward pass returns from the later stages. The peak is DD activations, confirming the D×AlayerD \times A_{\text{layer}} memory bound.

Out[5]:
Visualization
Schedule grid with four pipeline stages and 14 time steps, showing forward passes in green, backward passes in orange, and idle slots in gray.
Gantt chart showing a valid 1F1B pipeline schedule for 4 stages and 4 micro-batches across 14 time steps. Green cells represent forward passes, orange cells backward passes, and gray cells idle time. The startup bubble appears in the lower-left, the drain bubble in the lower-right, and the small M=D example retains additional idle slots while backward dependencies propagate between stages.

Code Implementation

This section implements a simplified pipeline parallelism simulation that illustrates the scheduling mechanics without requiring multiple physical GPUs. We model each stage as a Python object, route micro-batches through the pipeline following GPipe and 1F1B schedules, and measure utilization metrics. The simulation is intentionally simplified to highlight the scheduling logic rather than actual distributed computing mechanics.

Imports and Setup

Stage and Micro-Batch Definitions

We define a lightweight data structure for micro-batches and a class that simulates a single pipeline stage. Each stage tracks how many forward and backward passes it has completed, which lets us compute utilization statistics afterward.

In[7]:
Code
from dataclasses import dataclass
from typing import Optional

import numpy as np


@dataclass
class MicroBatch:
    """Represents a single micro-batch moving through the pipeline."""

    batch_id: int
    activations: np.ndarray
    gradients: Optional[np.ndarray] = None


class PipelineStage:
    """Simulates one stage of a pipeline-parallel model."""

    def __init__(self, stage_id: int, num_layers: int, hidden_size: int):
        self.stage_id = stage_id
        self.num_layers = num_layers
        self.hidden_size = hidden_size
        # Simple weight matrix to simulate layer computation
        self.weights = np.random.randn(hidden_size, hidden_size) * 0.01
        self.forward_count = 0
        self.backward_count = 0
        self.idle_steps = 0
        self.active_steps = 0

    def forward(self, micro_batch: MicroBatch) -> MicroBatch:
        """Run forward pass: apply a linear transformation to simulate layer computation."""
        self.forward_count += 1
        self.active_steps += 1
        out = micro_batch.activations @ self.weights
        return MicroBatch(batch_id=micro_batch.batch_id, activations=out)

    def backward(self, micro_batch: MicroBatch) -> MicroBatch:
        """Run backward pass: simulate gradient computation."""
        self.backward_count += 1
        self.active_steps += 1
        # Simplified gradient: propagate a scaled version backward
        grad_input = micro_batch.gradients @ self.weights.T
        return MicroBatch(
            batch_id=micro_batch.batch_id,
            activations=micro_batch.activations,
            gradients=grad_input,
        )

GPipe Schedule Simulation

The GPipe simulation processes all micro-batches through the forward pass across all stages, then processes all micro-batches through the backward pass in reverse order. This sequential structure mirrors the flush-based schedule.

In[8]:
Code
from typing import Dict


def simulate_gpipe(
    num_stages: int, num_microbatches: int, hidden_size: int
) -> Dict:
    """Simulate the GPipe (flush-based) pipeline schedule."""
    stages = [
        PipelineStage(i, num_layers=4, hidden_size=hidden_size)
        for i in range(num_stages)
    ]

    # Initialize micro-batches with random activations
    micro_batches = [
        MicroBatch(m, np.random.randn(8, hidden_size))
        for m in range(num_microbatches)
    ]

    total_time_steps = 0
    schedule_log = []

    # Forward phase: all micro-batches complete forward pass
    for stage_idx in range(num_stages):
        for mb in micro_batches:
            result = stages[stage_idx].forward(mb)
            mb.activations = result.activations
            total_time_steps += 1
            schedule_log.append(("F", stage_idx, mb.batch_id))

    # Attach synthetic gradients for backward pass
    for mb in micro_batches:
        mb.gradients = np.ones_like(mb.activations)

    # Backward phase: all micro-batches complete backward pass
    for stage_idx in reversed(range(num_stages)):
        for mb in reversed(micro_batches):
            result = stages[stage_idx].backward(mb)
            mb.gradients = result.gradients
            total_time_steps += 1
            schedule_log.append(("B", stage_idx, mb.batch_id))

    useful_steps = 2 * num_stages * num_microbatches
    bubble_steps = total_time_steps - useful_steps
    utilization = (
        useful_steps / total_time_steps if total_time_steps > 0 else 0.0
    )

    return {
        "schedule": "GPipe",
        "total_time_steps": total_time_steps,
        "bubble_steps": bubble_steps,
        "utilization": utilization,
        "stages": stages,
        "schedule_log": schedule_log,
    }

1F1B Schedule Simulation

The 1F1B simulation interleaves forward and backward passes. After the warmup phase fills the pipeline, the steady state alternates between processing one new forward micro-batch and completing one backward pass for the oldest pending micro-batch.

In[9]:
Code
from collections import defaultdict


def simulate_1f1b(
    num_stages: int, num_microbatches: int, hidden_size: int
) -> Dict:
    """Simulate the 1F1B pipeline schedule.

    In steady state, each stage alternates: one forward pass followed by
    one backward pass. This bounds activation memory to O(D) rather than O(M*D).
    """
    stages = [
        PipelineStage(i, num_layers=4, hidden_size=hidden_size)
        for i in range(num_stages)
    ]

    # Track activation states for each micro-batch at each stage
    activations = {
        m: np.random.randn(8, hidden_size) for m in range(num_microbatches)
    }
    gradients = {}

    schedule_log = []
    total_time_steps = 0
    in_flight_activations = defaultdict(list)

    # Warmup phase: fill pipeline with forward passes
    warmup_count = min(num_stages, num_microbatches)
    for warmup_step in range(warmup_count):
        mb_id = warmup_step
        for s in range(warmup_step + 1):
            result = stages[s].forward(MicroBatch(mb_id, activations[mb_id]))
            activations[mb_id] = result.activations
            in_flight_activations[s].append(mb_id)
            total_time_steps += 1
            schedule_log.append(("F", s, mb_id))

    # Steady state: interleave F and B passes
    next_forward = warmup_count
    next_backward = 0
    while next_backward < num_microbatches:
        # Forward pass for next micro-batch (if available)
        if next_forward < num_microbatches:
            mb_id = next_forward
            for s in range(num_stages):
                result = stages[s].forward(
                    MicroBatch(mb_id, activations[mb_id])
                )
                activations[mb_id] = result.activations
                total_time_steps += 1
                schedule_log.append(("F", s, mb_id))
            next_forward += 1

        # Backward pass for oldest pending micro-batch
        mb_id = next_backward
        grad = np.ones_like(activations[mb_id])
        for s in reversed(range(num_stages)):
            result = stages[s].backward(
                MicroBatch(mb_id, activations[mb_id], grad)
            )
            grad = result.gradients
            total_time_steps += 1
            schedule_log.append(("B", s, mb_id))
        gradients[mb_id] = grad
        next_backward += 1

    useful_steps = 2 * num_stages * num_microbatches
    bubble_steps = max(0, total_time_steps - useful_steps)
    utilization = (
        useful_steps / total_time_steps if total_time_steps > 0 else 0.0
    )

    return {
        "schedule": "1F1B",
        "total_time_steps": total_time_steps,
        "bubble_steps": bubble_steps,
        "utilization": utilization,
        "stages": stages,
        "schedule_log": schedule_log,
    }

Running the Comparison

We run both schedules across a range of micro-batch counts and compare their utilization against the theoretical prediction.

In[10]:
Code
# Configuration
num_stages = 4
hidden_size = 64

# Run both schedules across different micro-batch counts
microbatch_counts = [1, 2, 4, 8, 16, 32]
gpipe_results = []
fb1b_results = []

for M in microbatch_counts:
    gp = simulate_gpipe(num_stages, M, hidden_size)
    fb = simulate_1f1b(num_stages, M, hidden_size)
    gpipe_results.append(gp)
    fb1b_results.append(fb)

# Compute theoretical bubble fractions
theoretical_bubble = [
    (num_stages - 1) / (M + num_stages - 1) for M in microbatch_counts
]
Out[11]:
Console
Pipeline configuration: 4 stages, hidden_size=64

   M |  GPipe Util |  1F1B Util |  Theoretical Bubble
-----------------------------------------------------
   1 |     100.0% |    160.0% |              75.0%
   2 |     100.0% |    145.5% |              60.0%
   4 |     100.0% |    123.1% |              42.9%
   8 |     100.0% |    110.3% |              27.3%
  16 |     100.0% |    104.9% |              15.8%
  32 |     100.0% |    102.4% |               8.6%

As the number of micro-batches grows, utilization rises steadily toward 100% for both schedules. With only 1 micro-batch, the pipeline bubble is severe: the bubble fraction is (D−1)/D=75%(D-1)/D = 75\% for D=4D = 4 stages, leaving most device time idle. With 16 micro-batches, utilization climbs above 80%, and with 32 micro-batches, it exceeds 90%. The simulation confirms the theoretical prediction: increasing MM is the primary lever for reducing pipeline bubble overhead.

Note that the 1F1B simulation shows similar utilization to GPipe in this simple sequential model, as expected from the formula. In a real distributed system, however, 1F1B would allow much larger MM before running out of activation memory, because its activation memory bound is DD (constant in MM) rather than M×DM \times D (linear in MM). The ability to use large MM without hitting memory limits is where 1F1B provides its practical advantage.

Activation Memory Comparison

The activation memory calculation makes the 1F1B advantage tangible. For realistic LLM configurations, the difference between GPipe and 1F1B activation memory can be the difference between a training run that fits in GPU memory and one that crashes with an out-of-memory error.

In[12]:
Code
def estimate_activation_memory(
    num_stages: int,
    num_microbatches: int,
    micro_batch_size: int,
    seq_len: int,
    hidden_size: int,
    bytes_per_element: int = 2,  # bfloat16
) -> Dict:
    """Estimate peak activation memory for GPipe vs 1F1B."""

    # Activation size per micro-batch per stage (bytes)
    activation_per_mb_per_stage = (
        micro_batch_size * seq_len * hidden_size * bytes_per_element
    )

    # GPipe: must store all M micro-batches' activations across all D stages
    # simultaneously (they are all needed for the eventual backward pass)
    gpipe_peak = num_microbatches * num_stages * activation_per_mb_per_stage

    # 1F1B: at most D micro-batches' activations in flight at any stage
    # (because backward passes interleave with forward passes, freeing memory early)
    fb1b_peak = num_stages * activation_per_mb_per_stage

    return {
        "gpipe_gb": gpipe_peak / (1024**3),
        "fb1b_gb": fb1b_peak / (1024**3),
        "reduction_factor": num_microbatches,  # 1F1B is M times cheaper
    }
In[13]:
Code
# Realistic LLM dimensions: GPT-3 scale
realistic_config = {
    "num_stages": 8,
    "num_microbatches": 16,
    "micro_batch_size": 4,
    "seq_len": 2048,
    "hidden_size": 12288,
    "bytes_per_element": 2,  # bfloat16
}

memory_est = estimate_activation_memory(**realistic_config)
Out[14]:
Console
Activation memory estimate (GPT-3 scale, 8 stages, M=16 micro-batches)
  GPipe peak activation memory : 24.0 GB
  1F1B peak activation memory  : 1.5 GB
  Memory reduction factor      : 16x

GPipe requires storing all M*D micro-batch activations simultaneously.
1F1B bounds activation memory to D stages regardless of M.

The memory numbers favor 1F1B. For a GPT-3-scale model partitioned into 8 stages and processing 16 micro-batches, GPipe requires 16 times more activation memory than 1F1B. In practical training, where every gigabyte matters for fitting the largest possible model, this difference often determines which schedule is feasible at all.

Out[15]:
Visualization
Line chart showing GPipe activation memory increasing linearly from 1.5 to 48 GiB as M grows from 1 to 32, while 1F1B remains flat at 1.5 GiB, below an 80 GiB GPU reference line.
Estimated peak activation memory for GPipe versus 1F1B as a function of micro-batch count M, using 8 stages, hidden size 12288, sequence length 2048, micro-batch size 4, and bfloat16. GPipe grows linearly from 1.5 GiB at M=1 to 48 GiB at M=32, while 1F1B remains constant at 1.5 GiB. The 80 GiB reference line shows the remaining capacity margin in the displayed range.

Key Parameters

The key parameters for pipeline parallelism are:

  • num_stages (D): Number of pipeline stages, equal to the number of devices. Larger DD increases memory efficiency (each stage holds fewer parameters) but also increases the bubble fraction for a fixed MM. Adding more stages therefore requires a proportionally larger MM to maintain the same device utilization.
  • num_microbatches (M): Number of micro-batches per optimizer step. Larger MM reduces the bubble fraction but increases per-step latency (you must process more micro-batches before taking a weight update) and, for GPipe without checkpointing, increases activation memory linearly.
  • micro_batch_size: Size of each individual micro-batch, in sequences. Smaller values reduce per-stage activation memory but may underutilize GPU tensor cores, which need large matrix dimensions to operate efficiently. The total effective batch size equals M×micro_batch_sizeM \times \text{micro\_batch\_size}.
  • schedule: GPipe or 1F1B. GPipe separates forward and backward phases cleanly; 1F1B interleaves them to bound activation memory at DD stages' worth of data, enabling much larger MM values without running out of memory.
  • V (virtual stages): For interleaved 1F1B, the number of non-contiguous model chunks per device. Larger VV reduces the bubble fraction further (by factor 1/V1/V) but increases inter-device communication volume by factor VV.

Limitations and Practical Considerations

Pipeline parallelism solves the problem of fitting very deep models across multiple devices, but it introduces its own set of engineering challenges that practitioners must manage carefully.

The Irreducible Pipeline Bubble

The pipeline bubble remains the most fundamental limitation. Even with a large number of micro-batches, the bubble fraction is never zero. For a 16-stage pipeline with 128 micro-batches, the bubble fraction is still (15)/(128+15)≈10.5%(15)/(128 + 15) \approx 10.5\%. This unavoidable overhead means that pipeline parallelism alone cannot achieve linear scaling. Doubling the number of pipeline stages doubles the parameter capacity but does not double throughput, because the bubble grows proportionally with DD while MM must also grow proportionally to maintain the same bubble fraction.

In practice, the sweet spot for pipeline depth is typically 4 to 16 stages. Below 4 stages, the parameter capacity per device may be insufficient for the target model. Above 16 stages, the bubble becomes difficult to amortize without a very large number of micro-batches, which increases per-step latency, reduces the frequency of weight updates, and can degrade training stability for certain learning rate schedules. Systems like Megatron-LM typically use 8 to 16 pipeline stages for models in the 100-billion-to-trillion-parameter range.

Weight Staleness in Asynchronous Variants

Weight staleness is a subtler problem that arises in asynchronous variants of pipeline parallelism. In synchronous pipeline parallelism (the standard approach), all MM micro-batches see the same parameter values throughout the optimizer step, and the weight update happens only after all micro-batches have completed. This is clean and mathematically well-defined, but it requires waiting for the entire pipeline to drain before updating weights.

Some asynchronous schedules, like PipeDream's weight stashing approach, attempt to overlap weight updates with ongoing micro-batch processing. The idea is that while stage kk is processing forward passes for micro-batch mm, stage k−1k-1 could be using an updated weight matrix that incorporates gradients from micro-batch m−1m-1. This means different stages of the pipeline see different versions of the model weights simultaneously, introducing a form of gradient inconsistency. The updated weights must be "stashed" so that the backward pass for each micro-batch uses the same weights that were used during its forward pass. Weight stashing adds memory overhead (each in-flight micro-batch requires its own copy of the stage's weights) and makes checkpointing more complex. Asynchronous schedules can potentially achieve better device utilization in some configurations, but the gradient staleness generally requires additional stabilization measures, and for large models the weight stashing memory overhead can be prohibitive.

Interaction with Tensor and Data Parallelism

The interaction between pipeline parallelism and other forms of parallelism requires careful orchestration. In production, large-scale training systems combine all three strategies simultaneously. Pipeline parallelism partitions layers across devices along the depth dimension. Tensor parallelism (covered in the previous chapter) partitions individual weight matrices within each layer across devices along the width dimension. Data parallelism replicates the full pipeline-and-tensor-parallel group across multiple independent replicas, each processing different mini-batches.

This three-dimensional parallelism, popularized by the Megatron-LM framework, distributes computation across all available dimensions. The devices are organized into a 3D mesh, where one dimension corresponds to tensor-parallel degree, another to pipeline-parallel degree, and the third to data-parallel degree. A cluster with 3,072 GPUs might be configured as 8-way tensor parallelism times 12-way pipeline parallelism times 32-way data parallelism. Each dimension of the mesh requires different communication patterns: tensor parallelism requires all-reduce within a node (typically using NVLink), pipeline parallelism requires point-to-point activations across nodes (using InfiniBand), and data parallelism requires all-reduce of gradients across replicas (also using InfiniBand, but less frequently due to gradient accumulation).

Configuring these dimensions correctly requires understanding the communication topology of the cluster, the memory requirements of the model, and the target throughput. There is no universal formula: the best configuration depends on model size, cluster size, interconnect bandwidth, and training batch size. Tools like Megatron-LM's automatic parallelism configuration search can help find a good starting point.

Load Imbalance Across Stages

Load imbalance across stages is a practical concern that can significantly degrade throughput. If one stage takes 20% longer than the others due to a larger layer or a different computation pattern, that stage becomes the bottleneck and all other stages must wait for it. The pipeline utilization is limited by the slowest stage, not the average stage.

For transformer models, the embedding layer at the beginning and the LM head (output projection plus softmax) at the end can have different compute profiles than the middle transformer blocks. The embedding layer involves a large vocabulary lookup, which is memory-bound. The LM head involves a large matrix multiplication from the hidden dimension to the vocabulary size, which can be compute-intensive for large vocabularies. Placing these special layers in stages with fewer transformer blocks can help equalize stage durations.

Profiling-driven partitioning is more reliable than uniform partitioning. You measure the actual wall-clock time per layer for the specific model and hardware configuration, then solve the partition problem by minimizing the maximum stage duration. Dynamic programming solves this exactly in O(L×D)O(L \times D) time, where LL is the number of layers and DD is the number of stages.

Debugging and Reproducibility

Pipeline parallelism adds engineering complexity that is absent from single-device training. The model state at any point in training is distributed across DD devices, so inspecting the full model requires gathering tensors from multiple GPUs. Debugging gradient correctness requires comparing activations at stage boundaries, which involves coordinating logging across multiple processes.

A subtle class of bugs arises from incorrect synchronization. If one stage falls behind due to a scheduling issue or hardware variability, the entire pipeline stalls. Unlike a crash (which produces an obvious error), a subtle timing issue may manifest only as reduced throughput, making it hard to detect. Monitoring per-stage throughput and the fraction of time each stage spends idle can reveal these issues.

Numerical reproducibility is also harder to ensure in pipeline parallelism than in single-device training. The order in which gradients are accumulated across micro-batches may vary slightly depending on hardware timing, leading to different floating-point summation orders and therefore slightly different numerical results. For research purposes (where exact reproducibility is important for comparing runs), this nondeterminism must be controlled by fixing seeds and using deterministic algorithms throughout.

The Path to 3D Parallelism in Production

Despite these challenges, pipeline parallelism has become a standard component of large-scale training infrastructure. The models that define the modern era of large language models, including GPT-3, Gopher, PaLM, Megatron-Turing NLG, and their successors, all relied on pipeline parallelism combined with tensor and data parallelism to distribute training across thousands of GPUs. The pipeline parallel technique has been refined from GPipe's original flush-based approach through 1F1B's memory-efficient interleaving to the sophisticated interleaved schedule with virtual stages, each refinement motivated by real engineering constraints encountered in training progressively larger models.

The next chapter extends this distributed training picture by examining the communication optimization strategies, specifically gradient compression, all-reduce algorithms, and network topology-aware scheduling, that make cross-device synchronization fast enough to avoid bottlenecking on network bandwidth as the device count scales into the thousands.

Summary

Pipeline parallelism splits a deep model vertically along its layer dimension, assigning contiguous groups of layers to separate devices. Each device holds a fraction of the total parameters, enabling models that far exceed a single GPU's memory capacity to be trained in practice.

The core challenge is the pipeline bubble: the idle time that accumulates because later stages must wait for earlier stages to produce activations, and because gradients must drain backward through all stages before the next optimizer step. The bubble fraction is (D−1)/(M+D−1)(D-1)/(M+D-1), where DD is the number of stages and MM is the number of micro-batches. Increasing MM is the primary strategy for driving the bubble fraction toward zero, at the cost of a larger effective batch size and more time between weight updates.

GPipe addresses the bubble by processing all MM micro-batches through the complete forward pass before beginning any backward passes. This achieves good device utilization for large MM but requires storing all micro-batch activations simultaneously, leading to activation memory that scales as M×DM \times D. GPipe uses gradient checkpointing to manage this memory at the cost of roughly 33% additional compute.

The 1F1B schedule interleaves forward and backward passes to bound activation memory at DD stages' worth of data regardless of MM. The bubble fraction is identical to GPipe, but the memory efficiency of 1F1B makes it the standard choice for large model training: it allows much larger MM values without running out of activation memory, which in turn allows smaller micro-batches and better tensor core utilization.

The interleaved 1F1B variant assigns non-contiguous layer chunks to each device, reducing the bubble fraction by a factor equal to the number of chunks per device. This further reduces idle time at the cost of proportionally more inter-device communication volume. Whether the tradeoff is favorable depends on the interconnect bandwidth relative to compute throughput.

Key takeaways:

  • Pipeline parallelism assigns consecutive layer groups to separate devices, allowing parameter count to scale linearly with device count while keeping per-device memory fixed
  • The pipeline bubble is unavoidable but shrinks as the ratio M/DM/D grows; practical training runs use M≫DM \gg D to keep bubble fractions below 10%
  • GPipe and 1F1B achieve the same bubble fraction; 1F1B uses MM times less activation memory, making it the preferred schedule for large models
  • The interleaved schedule reduces the bubble fraction by factor 1/V1/V at the cost of VV times more communication
  • In production systems, pipeline parallelism is combined with tensor parallelism and data parallelism into 3D parallelism, allowing training of models with hundreds of billions of parameters across thousands of GPUs

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about pipeline parallelism.

Pipeline Parallelism Quiz

Question 1 of 80 of 8 completed
What does pipeline parallelism split across multiple devices?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026pipelineparallelism, author = {Michael Brenndoerfer}, title = {Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B}, year = {2026}, url = {https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-06} }
APAAcademic
Michael Brenndoerfer (2026). Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B. Retrieved from https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b
MLAAcademic
Michael Brenndoerfer. "Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B." 2026. Web. October 6, 2026. <https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b>.
CHICAGOAcademic
Michael Brenndoerfer. "Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B." Accessed October 6, 2026. https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B'. Available at: https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b (Accessed: October 6, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Pipeline Parallelism: Stages, Micro-Batching, GPipe, 1F1B. https://mbrenndoerfer.com/writing/pipeline-parallelism-stages-micro-batching-gpipe-1f1b

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.