Part of Language AI Handbook
Explains how learning rate schedules improve neural network training. Topics include step decay, exponential decay, inverse square root with warmup.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Learning Rate Decay
Training a neural network is not a static process. The learning rate that works well early in training, when the model is far from any good solution and needs large steps to make progress, often becomes a liability later, when the model is close to a good solution and large steps cause it to overshoot. Learning rate decay, also called learning rate scheduling, addresses this mismatch by reducing the learning rate over the course of training.
The intuition is simple: think of gradient descent as hiking down a mountain in heavy fog. Early on, you can stride confidently in roughly the right direction. But as you approach the valley floor, you need to take smaller, more careful steps to avoid stumbling past the lowest point. A fixed learning rate cannot adapt to this changing landscape. A decaying schedule adjusts automatically, enabling fast initial progress and precise final convergence.
This chapter covers the main decay strategies: step decay, exponential decay, inverse square root decay, and cosine annealing. For each, we examine the mathematical formulation, the practical intuition, and when to prefer one over another. We also look at how these schedules interact with adaptive optimizers like Adam, and how modern training pipelines compose multiple schedules together. Along the way, we build concrete intuition for what happens inside optimization when the learning rate changes, and we look at the historical context that shaped each approach.
Why the Learning Rate Must Change
Before examining specific schedules, it is worth understanding precisely why a fixed learning rate causes problems at both ends of training. The issue is not merely cosmetic. A fixed learning rate is a fundamental mismatch with the geometry of neural loss surfaces, and understanding that mismatch explains why every major training recipe today uses some form of scheduling.
The Two-Regime Problem
At the start of training, the model parameters are randomly initialized and far from any good solution. The loss surface at this point is often steep and the gradients are large. A moderate learning rate causes the model to take meaningful steps and make rapid progress. If the learning rate is too small at this stage, training is unnecessarily slow: the model makes tiny moves when large moves are available and beneficial. You are walking when you could be running.
Later in training, the model has moved into a region near a local minimum (or saddle point) where the loss landscape is much flatter. Gradients become smaller and the curvature of the landscape matters more. A learning rate that was appropriate for the steep early landscape now causes the model to bounce back and forth across a shallow basin rather than settling into it. The model never fully converges: it oscillates around the minimum with a loss that fluctuates rather than decreasing smoothly. You are trying to settle into a shallow depression with steps large enough to carry you across it entirely.
The learning rate controls how large a step gradient descent takes in the direction of the negative gradient. At each update, parameters move by . A larger makes faster but less precise updates; a smaller makes slower but more precise updates.
A decaying schedule interpolates between these two needs by starting large and shrinking over time. The key design choices are: how quickly to decay, what functional form the decay follows, and what minimum learning rate to allow.
Loss Landscape Geometry and Why It Changes
The loss landscape is not uniform throughout parameter space. Near initialization, the landscape is typically dominated by large-scale structure: the model is trying to learn basic patterns (for a language model, that might mean learning that certain words co-occur frequently, or that sentence structure follows certain patterns), and the gradients that point toward these basic patterns are large and relatively consistent across batches. A high learning rate is safe because the direction of improvement is clear.
As training proceeds and the model captures these large-scale patterns, the remaining signal becomes more subtle. The model is now trying to refine fine-grained distinctions, and the gradients that carry this information are smaller. The loss surface near a good local minimum looks like a shallow bowl, and convergence requires increasingly precise steps toward its center. With a fixed step size that is too large, the model overshoots the center every step, landing on the far wall, bouncing back, and oscillating indefinitely.
The mathematical reason for this oscillation comes from the relationship between the learning rate and the curvature. If the maximum eigenvalue of the Hessian (the matrix of second derivatives) at a point is , then gradient descent will converge stably only if the learning rate satisfies . Near a minimum where curvature is high, this constraint is tight. Early in training, where curvature is lower and the landscape is smoother, the same constraint is much looser and larger learning rates are safe.
Historical Context: Where Scheduling Came From
The practice of reducing the learning rate during training predates deep learning. Stochastic gradient descent theory, developed for convex optimization in the 1950s and 1960s, established convergence guarantees under conditions that require the learning rate to satisfy the Robbins-Monro conditions:
The first condition ensures the optimizer can reach any part of the parameter space; the second ensures that individual steps eventually become small enough to allow convergence. A constant learning rate satisfies the first condition but not the second (the sum of its squares diverges). The inverse square root schedule, which decays as , satisfies both conditions, which is one reason it carries such theoretical weight.
Practical deep learning inherited this intuition. Early neural network training papers in the 1980s and 1990s routinely described reducing the learning rate toward the end of training. The specific form varied: some halved it manually, others used exponential decay. The systematic study and naming of these schedules as distinct techniques emerged with the larger-scale experiments of the 2010s, when training runs on ImageNet made hyperparameter sensitivity clearly visible.
Step Decay
Step decay is the simplest approach and the one historically used in early deep learning. The learning rate is held constant for a fixed number of epochs, then multiplied by a decay factor, then held constant again, and so on. This produces a staircase pattern in the learning rate over time.
Formulation
The update rule at epoch is:
where:
- : the initial learning rate
- : the decay factor, typically between 0.1 and 0.5 (so the learning rate is halved, quartered, or reduced by 10x at each step)
- : the step size in epochs (how many epochs between drops)
- : the floor of , which counts how many complete steps have elapsed
For example, with , , and , the learning rate starts at 0.1, drops to 0.01 at epoch 30, drops to 0.001 at epoch 60, and so on. Each drop reduces the learning rate by exactly a factor of ten.
The floor function is what creates the staircase: it stays constant for consecutive epochs, then increments by exactly one. This means the learning rate at epoch 29 is identical to the learning rate at epoch 0, and the learning rate at epoch 30 is exactly times that.
Intuition
Step decay appeals to practitioners because it is easy to understand and diagnose. You can look at a training loss curve and see exactly when each drop happened. If the loss plateaus before a drop, that suggests the drop might be too late. If the loss jumps up briefly after a drop, the new learning rate might be too small. This interpretability makes step decay a useful debugging tool. Every phase of training is clearly delineated by the drop, and you can reason about what the model was doing in each phase.
The staircase structure also means training can be naturally organized into phases. Each phase trains at a fixed learning rate until the loss stops improving, then the next phase begins at a lower rate to refine the solution. Many early ResNet and VGG training recipes used this pattern: train for 30 epochs at 0.1, drop to 0.01, train 30 more, drop to 0.001, and finish. He et al.'s original ResNet paper for ImageNet used exactly this recipe, and the robustness of this approach across many architectures is what established its dominance in computer vision through the early 2010s.
The main limitation is rigidity. The schedule is specified in advance and does not adapt to the observed training progress. If a particular run converges faster or slower than expected, the drops happen at the wrong time. A model that has already plateaued two epochs before the scheduled drop is waiting unnecessarily. A model that is still making rapid progress when the drop arrives gets its learning rate cut during a productive phase. Some practitioners address this by monitoring validation loss and dropping the rate when it plateaus rather than at fixed epochs, a variant called ReduceLROnPlateau. This adaptive version maintains the staircase structure while letting training dynamics determine when each step occurs.
The Multiplicative Factor and Its Consequences
The choice of has a significant impact on the training dynamics. A very aggressive factor like (ten-fold reduction) means that after the drop, the model is taking steps one-tenth as large. This can effectively freeze learning on some parameters: if the gradient for a particular parameter is small at the pre-drop learning rate, it becomes negligible after a ten-fold reduction. The drop essentially shifts the focus of training to parameters with larger gradients.
A gentler factor like (halving) is less dramatic. The model slows down but does not stop making meaningful progress on any parameter. Multiple gentler drops can sometimes achieve the same end result as one aggressive drop while maintaining smoother loss curves.
The step size determines how long the model trains at each level. Setting too small means the model does not have time to exploit each learning rate before it drops. Setting too large means the model may have plateaued long before the next drop. The optimal is essentially the number of epochs it takes for the model to plateau at each learning rate level, which depends on the dataset, architecture, and optimizer.
Exponential Decay
Exponential decay smooths out the staircase by applying decay at every step rather than at discrete intervals. The learning rate decreases continuously as a smooth exponential curve.
Formulation
At step , the learning rate is:
where:
- : the initial learning rate
- : the decay rate, a positive constant controlling how quickly the rate falls
- : the current training step (or epoch)
An equivalent formulation uses a decay factor per step:
where . The two are mathematically identical; the second is often more convenient because you can set directly (e.g., for slow decay or for faster decay). The relationship means , so a decay factor of 0.99 corresponds to a rate of approximately 0.01.
Choosing the Decay Rate
The relationship between and training duration matters enormously. After steps with per-step factor , the learning rate is . If you want the learning rate to end at, say, 1% of its initial value after 100 epochs, you need , which gives . You can work backward from your desired endpoint to choose .
This backward calculation is a useful habit. Before setting , decide what fraction of the initial learning rate you want at the end of training, and how many steps training will run. The required follows directly:
This ensures the schedule hits your target endpoint rather than decaying too fast or too slow.
Intuition and Limitations
The exponential function has a special property: its rate of change is proportional to its current value. This means exponential decay reduces the learning rate by the same multiplicative factor at every step. Early in training, the absolute reduction is large (because the learning rate is high); later in training, the absolute reduction is small. The relative reduction is constant throughout.
In practice, the effect of this is that exponential decay can shrink the learning rate too quickly, especially with even moderately large . After steps, the learning rate has shrunk by a factor of , which can become extremely small well before training ends. Practitioners often add a floor: a minimum learning rate below which the schedule does not go, preventing the rate from becoming so small that training effectively stops.
Without a floor, exponential decay with per epoch would reach after 100 epochs starting from , which is effectively zero. With a floor of , the schedule plateaus at that value and training continues to make at least some progress.
The continuous nature of exponential decay is both a strength and a weakness. There are no abrupt transitions to cause loss spikes, but the schedule also never provides the sharp signal of a step drop that can sometimes help the model escape a local plateau. Empirically, exponential decay tends to work well when training is well-behaved and the main concern is smooth, gradual refinement rather than dramatic transitions between phases.
import numpy as np
def step_decay(t, eta0, gamma, step_size):
return eta0 * (gamma ** (t // step_size))
def exponential_decay(t, eta0, gamma):
return eta0 * (gamma**t)
def inv_sqrt_decay(t, eta0):
return eta0 / np.sqrt(t + 1)
def cosine_decay(t, eta_max, eta_min, T_max):
return eta_min + 0.5 * (eta_max - eta_min) * (1 + np.cos(np.pi * t / T_max))
epochs = np.arange(0, 100)
eta0 = 0.1
step_lrs = np.array(
[step_decay(t, eta0=eta0, gamma=0.1, step_size=30) for t in epochs]
)
exp_lrs = np.array(
[exponential_decay(t, eta0=eta0, gamma=0.955) for t in epochs]
)
inv_sqrt_lrs = np.array([inv_sqrt_decay(t, eta0=eta0) for t in epochs])
cosine_lrs = np.array(
[cosine_decay(t, eta_max=eta0, eta_min=1e-4, T_max=100) for t in epochs]
)
The plot shows the four schedules on a logarithmic scale. Step decay drops abruptly at epochs 30 and 60. Exponential decay falls smoothly but steeply. Inverse square root decay drops quickly at first and then flattens. Cosine annealing curves smoothly from its maximum to its minimum.
Inverse Square Root Decay
Inverse square root decay, popularized by the original Transformer paper "Attention Is All You Need," has a distinctive profile: it decreases quickly at first and then levels off gradually. This behavior makes it well-matched to how optimization often proceeds: rapid early progress followed by slower refinement.
Mathematical Foundation
The schedule is named for its decay rate: the learning rate is proportional to the inverse of the square root of the step count. This means the rate falls rapidly at first (the difference between step 1 and step 4 is large) and then changes slowly at large step counts (the difference between step 10,000 and step 10,001 is tiny). On a log-log scale, the decay appears as a straight line with slope .
The schedule is:
where:
- : the peak learning rate reached after warmup
- : the number of warmup steps
- : the current step
This form describes the decay phase while capping the rate at for . In practice, it is usually paired with a separate linear warmup that increases the learning rate from zero to the peak. After warmup, the term equals , and the schedule follows a true inverse square root curve.
A common combined form that includes the linear warmup is the original Transformer formulation:
where:
- : the model hidden dimension (used as a scaling factor in the original Transformer)
- The first argument governs the decay phase
- The second argument governs the linear warmup phase
The minimum of the two quantities ensures a smooth transition: during warmup the linear term is smaller (it is small because is small), and after warmup the inverse square root term takes over (it is smaller because shrinks faster than the linear term grows). The crossover happens exactly at , where both terms equal .
Comparing Decay Rates Across Schedules
To understand why inverse square root is preferred over exponential for long training runs, it helps to compare how aggressively each schedule reduces the learning rate. With exponential decay at per step, after 100,000 steps the learning rate has fallen to , which is effectively zero. With inverse square root, after 100,000 steps from a warmup of 4,000 steps, the learning rate is times the peak, a reduction to 20% of peak. For pre-training runs that span hundreds of thousands of steps, this gentler decay means the model is still making meaningful learning rate-scaled gradient steps even late in training.
This is particularly important for large language model pre-training, where the model is learning from a dataset so large that many distinct learning signals appear throughout training. An exponential schedule would reduce the learning rate to near zero long before the dataset is exhausted, effectively stopping learning. The inverse square root schedule allows the model to continue learning from later portions of the data with a non-negligible step size.
Warmup: Why Start Small?
The warmup period addresses a specific training instability. At the very start of training, the model parameters are random, the gradient estimates are highly noisy, and the loss surface is poorly characterized. A large learning rate in this regime can cause the loss to explode or send the model into a region of parameter space that is very difficult to escape.
Linear warmup starts the learning rate near zero and increases it steadily over the first several thousand steps. By the time the learning rate reaches its peak value, the model has already made some initial progress, gradient estimates have stabilized, and the optimizer's momentum and variance estimates (in Adam) have had time to accumulate. The peak learning rate can then safely be much larger than if training had started at full speed.
The warmup is especially important for Adam. Adam estimates the first and second moments of the gradient using exponential moving averages. At step one, these estimates are initialized to zero and have not yet had time to reflect the true gradient distribution. Early updates with a large learning rate are based on these poorly calibrated estimates, which can cause wild oscillations. Warmup gives the moment estimates time to converge before the learning rate reaches its full value.
A warmup period is a phase at the beginning of training during which the learning rate is increased from a small value to the target rate. Linear warmup increases the rate by a fixed amount each step. Warmup prevents instability caused by noisy gradients and poorly calibrated optimizer statistics early in training.
The warmup length is typically set to 4,000 to 10,000 steps for transformer training, though the optimal value depends on batch size and model size. Larger models and larger batches generally benefit from longer warmup. With very large batch sizes (as in distributed training), the signal-to-noise ratio in gradient estimates is higher because more samples are averaged, but the learning rate is also larger due to linear scaling, so the need for warmup does not disappear.
Warmup Length and the Batch Size Relationship
An underappreciated aspect of warmup is how it interacts with batch size. When you double the batch size, the linear scaling rule (introduced by Goyal et al. in 2017 for ImageNet training with SGD) recommends doubling the learning rate to maintain the same relative update magnitude. But doubling the learning rate without extending warmup often causes instability. The convention in large-batch training is to also scale the warmup duration: if you multiply batch size by , multiply warmup steps by as well. This gives the optimizer time to accumulate reliable statistics before taking the larger steps that the scaled learning rate demands.
def transformer_schedule(t, d_model, warmup_steps):
t = max(t, 1)
return d_model ** (-0.5) * min(t ** (-0.5), t * warmup_steps ** (-1.5))
d_model = 512
warmup_steps = 4000
steps = np.arange(1, 40001)
transformer_lrs = np.array(
[transformer_schedule(t, d_model, warmup_steps) for t in steps]
)
peak_step = int(np.argmax(transformer_lrs))
peak_lr = float(transformer_lrs[peak_step])Peak learning rate: 0.000699 Peak reached at step: 4000 LR at step 10000: 0.000442 LR at step 40000: 0.000221
The peak learning rate is reached exactly at the warmup step, after which the schedule decays as . The learning rate at step 40,000 is roughly 32% of its peak value, showing how slowly the inverse square root falls compared to exponential decay.

Cosine Annealing
Cosine annealing uses a half-cosine curve to reduce the learning rate from a maximum value to a minimum value over a fixed number of steps. It became widely adopted after Loshchilov and Hutter introduced it in 2017 and combined it with warm restarts to create SGDR (Stochastic Gradient Descent with Warm Restarts).
Why Cosine?
The choice of a cosine curve rather than a simple straight line or exponential is motivated by the shape of the decay, not just convenience. At the start of the cosine curve (near ), the cosine function changes slowly, meaning the learning rate stays near its maximum for a while and decreases gently at first. Near the end (as approaches ), the cosine function is near and also changing slowly, meaning the learning rate settles near its minimum and finishes smoothly. In the middle, the cosine passes through its steepest section, driving the sharpest decline.
This shape is intuitively appealing: the model starts at full speed, gradually shifts into a slower and slower rate, and comes to a gentle stop. Contrast this with linear decay, which reduces the learning rate at a constant rate regardless of where training is, and step decay, which drops abruptly. Cosine annealing provides a natural deceleration curve.
The cosine also has a nice mathematical property: the derivative of the cosine schedule is zero at both endpoints. This means the learning rate changes most slowly exactly at the start and end of the schedule, which is often where training is most sensitive to abrupt changes.
Formulation
The basic cosine annealing schedule is:
where:
- : the minimum learning rate (often 0 or a small positive value like )
- : the maximum learning rate
- : the number of steps for one half-cosine period
- : the current step within the period
At , the cosine term is , giving . At , the cosine term is , giving . In between, the rate follows the smooth cosine curve. The factor of ensures the output lies in the range .
Worked Numerical Example
Let us trace through a specific case to build intuition. Suppose , , and .
At :
At (one quarter through):
Since , this gives approximately .
At (halfway):
At (end):
Halfway through training, the learning rate is approximately half of the initial value. The curve is symmetric around the halfway point in a certain sense: the first quarter (epochs 0-25) sees a relatively small drop from 0.1 to 0.086, while the third quarter (epochs 50-75) sees a larger drop from 0.0505 to about 0.016. The decline accelerates in the middle of the schedule.
Warm Restarts
The standard cosine schedule ends at and stays there. Loshchilov and Hutter proposed restarting the schedule after reaching the minimum, cycling back to and annealing again. This is cosine annealing with warm restarts, sometimes written as SGDR or CosineAnnealingWarmRestarts.
The restart strategy has an appealing justification: near the end of each cycle, the low learning rate drives the model into a local minimum. The restart then kicks the model out of that minimum with a higher learning rate, potentially allowing it to find a flatter, more generalizable minimum in the next cycle. The intuition is that flat minima generalize better than sharp minima, and warm restarts encourage exploration of the loss landscape.
Restarts can also be paired with a multiplicative increase in the cycle length. Each successive cycle is made longer by a factor , so the model spends more and more time at low learning rates as training progresses. A typical choice is , doubling the cycle length each restart.
A learning rate schedule that follows a cosine curve from maximum to minimum, then resets to the maximum and repeats. The restart allows the optimizer to escape sharp local minima and explore flatter regions of the loss landscape that tend to generalize better.
The warm restart also provides an elegant approach to model selection. At the end of each cosine cycle, the model has converged to a local minimum with a very small learning rate. This is a natural checkpoint: the model at the end of each cycle can be saved and compared. Since each restart nudges the model toward a different minimum, the ensemble of cycle-end checkpoints can sometimes outperform any single checkpoint. This practice, called snapshot ensembling, was introduced alongside SGDR as a way to get multiple trained models for the cost of one training run.
def cosine_with_restarts(t, eta_max, eta_min, T0, T_mult=1):
T_cur = T0
t_remaining = t
while t_remaining >= T_cur:
t_remaining -= T_cur
T_cur = int(T_cur * T_mult)
return eta_min + 0.5 * (eta_max - eta_min) * (
1 + np.cos(np.pi * t_remaining / T_cur)
)
steps_restart = np.arange(0, 200)
cosine_fixed = [
cosine_with_restarts(t, eta_max=0.1, eta_min=1e-4, T0=50, T_mult=1)
for t in steps_restart
]
cosine_double = [
cosine_with_restarts(t, eta_max=0.1, eta_min=1e-4, T0=25, T_mult=2)
for t in steps_restart
]

The doubling schedule is useful when you are uncertain about total training time: early cycles provide frequent opportunities to explore, while later cycles devote progressively longer intervals to annealing within a basin. In this 200-epoch window, the 200-epoch cycle has only just begun at epoch 175.
Decay Scheduling in Practice
The main schedules covered above can be augmented with additional techniques for practical training pipelines. This section covers the remaining standard variants, discusses how to compose multiple schedules, and addresses practical choices that come up repeatedly.
Linear Decay
Linear decay reduces the learning rate in equal steps across training:
where is the total number of training steps. Linear decay is easy to understand and widely used in language model fine-tuning. Many Hugging Face training scripts use linear warmup followed by linear decay as their default schedule.
The key property of linear decay is predictability: the learning rate at any step is a straightforward linear interpolation between the starting and ending values. This makes it easy to reason about training progress. If you are at 50% of total steps, the learning rate is exactly 50% of the initial value. If you extend or shorten training, you can immediately compute the new schedule without any complex parameter relationships.
Linear decay is also well-matched to fine-tuning pre-trained models. In fine-tuning, the goal is to adapt the model to a new task while preserving the general knowledge from pre-training. The learning rate needs to start small enough to not disrupt pre-trained weights too violently, and then decrease further to ensure precise convergence. A linear schedule that starts at a modest learning rate like and decreases to zero over, say, three epochs is exactly this kind of gentle, predictable refinement.
Polynomial Decay
Polynomial decay generalizes both linear and inverse square root decay:
where is the polynomial power and is the final learning rate. Setting gives linear decay. Setting gives a schedule that drops quickly initially and then slows, resembling inverse square root. Setting gives a quadratic curve that drops slowly at first and quickly near the end, the opposite behavior.
The power is rarely tuned in practice because linear and cosine are simpler and well-understood. Polynomial decay appears in some large-scale training recipes where precise control over the decay shape is desired. The key insight is that the exponent controls the skew of the decay: smaller means more reduction happens early, larger means more reduction happens late.
The Floor Matters
Every schedule should have a non-zero minimum learning rate. Without a floor, exponential or inverse square root schedules approach zero asymptotically. At very small learning rates, training effectively stops: gradients are multiplied by a negligible factor and weights barely move. A floor of or ensures that training continues to make at least some progress.
The floor is especially important for schedules used in fine-tuning over many epochs. If you accidentally set too aggressively for an exponential schedule, the learning rate could reach the floor very early in training, wasting compute on steps where the parameters barely change. Monitor the learning rate reached during training as well as the schedule configuration.
Composing Warmup with Decay
Most modern training pipelines compose a warmup phase with one of the main decay schedules. The warmup brings the learning rate from zero to its peak value over the first few hundred or thousand steps, and then one of the decay schedules takes it from the peak down to the floor.
The transition from warmup to decay is handled differently depending on the framework. In some implementations, warmup and decay are separate schedule objects that are chained together. In PyTorch's get_cosine_schedule_with_warmup from the Hugging Face transformers library, the two phases are unified into a single function that changes behavior at the warmup boundary.
def linear_warmup_then_decay(
t, eta_max, warmup_steps, total_steps, decay="linear", eta_end=0.0
):
if t < warmup_steps:
return eta_max * t / warmup_steps
progress = (t - warmup_steps) / max(total_steps - warmup_steps, 1)
if decay == "linear":
return eta_end + (eta_max - eta_end) * max(0.0, 1 - progress)
elif decay == "cosine":
return eta_end + 0.5 * (eta_max - eta_end) * (
1 + np.cos(np.pi * progress)
)
return eta_max
total_steps = 1000
warmup = 100
steps_fine = np.arange(0, total_steps + 1)
linear_schedule = [
linear_warmup_then_decay(
t,
eta_max=0.001,
warmup_steps=warmup,
total_steps=total_steps,
decay="linear",
)
for t in steps_fine
]
cosine_schedule_fine = [
linear_warmup_then_decay(
t,
eta_max=0.001,
warmup_steps=warmup,
total_steps=total_steps,
decay="cosine",
)
for t in steps_fine
]
The two schedules look similar during the main decay phase, but cosine decay is slightly more gradual at the start of decay and slightly steeper near the end. In practice, the difference between them is often smaller than the impact of the peak learning rate or warmup length.
Schedules and Adaptive Optimizers
Learning rate decay was designed primarily for SGD, where the learning rate directly controls step size. Adaptive optimizers like Adam complicate this picture considerably, and understanding the interaction is important for using schedules correctly.
How Adam Uses the Learning Rate
Adam maintains a per-parameter estimate of the learning rate through its second moment estimate . The effective step size for parameter is approximately , where tracks the historical squared gradients. Parameters with large, consistent gradients get small effective steps, and parameters with small or variable gradients get larger steps.
The global learning rate in Adam functions as an overall scaling factor on top of this per-parameter adaptation. When you apply a learning rate schedule on top of Adam, you are scaling this already-adapted step size. The schedule affects all parameters equally by a common multiplier, while Adam's internal adaptation continues to differentiate between them.
This combination generally works well in practice. The schedule handles the global progression from fast to slow learning, while Adam handles the per-parameter optimization. Most practitioners apply standard decay schedules to Adam without modification and find that this works reliably.
Why Schedules Still Matter for Adam
One might wonder whether schedules are necessary for Adam at all, given that Adam already adapts. The answer is yes, for several reasons.
First, Adam's adaptive mechanism operates on the ratio of first to second moments, not on the absolute magnitude of the learning rate. The global still controls the overall scale of updates, and a well-calibrated schedule ensures this scale shrinks appropriately as the model approaches convergence.
Second, Adam can get stuck oscillating around a minimum with a fixed learning rate for the same geometric reason as SGD. The curvature of the loss surface near the minimum sets a stability limit on the step size, and Adam's internal adaptation does not automatically reduce steps below this limit. A decaying schedule that brings below the stability threshold is what allows the model to settle.
Third, for fine-tuning, the schedule interacts with the model's pre-trained weights in an important way. Starting with a large Adam learning rate for fine-tuning can corrupt pre-trained representations, even with Adam's adaptation. A careful warmup followed by a decaying schedule protects the pre-trained weights by keeping updates small early on.
AdamW and Decoupled Decay
One subtlety arises with weight decay. Standard Adam with L2 regularization applies a decay term that is scaled by the adaptive learning rate, meaning parameters with large gradients receive less weight decay. This is mathematically inconsistent: weight decay is supposed to pull all parameters toward zero uniformly, but coupling it to the adaptive learning rate means it is applied differently to different parameters.
AdamW, introduced by Loshchilov and Hutter in 2019, fixes this by decoupling the weight decay from the gradient update: weight decay is applied at a fixed rate regardless of the adaptive step size. When using AdamW, the learning rate schedule affects gradient-based updates, and weight decay is controlled separately.
This decoupling matters for schedule design because it means the learning rate schedule in AdamW controls only the gradient step, not the regularization strength. You can set the learning rate schedule aggressively without worrying about undermining the weight decay regularization. The two can be tuned independently, which simplifies hyperparameter search.
Implementation
Let us implement all four schedules using PyTorch's scheduler API and verify they behave as expected.
import torch.nn as nn
import torch.optim as optim
class TinyModel(nn.Module):
def __init__(self):
super().__init__()
self.fc = nn.Linear(10, 1)
def forward(self, x):
return self.fc(x)
model = TinyModel()
optimizer_step = optim.SGD(model.parameters(), lr=0.1)
scheduler_step = optim.lr_scheduler.StepLR(
optimizer_step, step_size=30, gamma=0.1
)
optimizer_exp = optim.SGD(model.parameters(), lr=0.1)
scheduler_exp = optim.lr_scheduler.ExponentialLR(optimizer_exp, gamma=0.955)
optimizer_cos = optim.SGD(model.parameters(), lr=0.1)
scheduler_cos = optim.lr_scheduler.CosineAnnealingLR(
optimizer_cos, T_max=100, eta_min=1e-4
)
optimizer_cosr = optim.SGD(model.parameters(), lr=0.1)
scheduler_cosr = optim.lr_scheduler.CosineAnnealingWarmRestarts(
optimizer_cosr, T_0=50, T_mult=1
)def record_lr_schedule(optimizer, scheduler, n_steps=100):
lrs = []
for _ in range(n_steps):
lrs.append(optimizer.param_groups[0]["lr"])
scheduler.step()
return lrs
torch_step_lrs = record_lr_schedule(optimizer_step, scheduler_step)
torch_exp_lrs = record_lr_schedule(optimizer_exp, scheduler_exp)
torch_cos_lrs = record_lr_schedule(optimizer_cos, scheduler_cos)
torch_cosr_lrs = record_lr_schedule(optimizer_cosr, scheduler_cosr)Step decay LRs at epochs [0, 30, 60, 90]: Epoch 0: 0.100000 Epoch 30: 0.010000 Epoch 60: 0.001000 Epoch 90: 0.000100 Cosine LRs at epochs [0, 25, 50, 75, 99]: Epoch 0: 0.100000 Epoch 25: 0.085370 Epoch 50: 0.050050 Epoch 75: 0.014730 Epoch 99: 0.000125
The step decay drops exactly at epoch 30 and 60 as configured. The cosine decay smoothly interpolates from 0.1 to near 0.0001 over 100 epochs.
For the inverse square root schedule with warmup, PyTorch does not provide a built-in, but it is straightforward to implement using LambdaLR:
def make_inv_sqrt_lambda(warmup_steps):
def lr_lambda(t):
if t < warmup_steps:
return t / warmup_steps
return (warmup_steps**0.5) / (t**0.5)
return lr_lambda
model_inv = TinyModel()
optimizer_inv = optim.AdamW(model_inv.parameters(), lr=0.001)
scheduler_inv = optim.lr_scheduler.LambdaLR(
optimizer_inv, lr_lambda=make_inv_sqrt_lambda(warmup_steps=100)
)
torch_inv_lrs = record_lr_schedule(optimizer_inv, scheduler_inv, n_steps=500)Inverse sqrt LRs at steps [0, 50, 100, 200, 499]: Step 0: 0.000000 Step 50: 0.000500 Step 100: 0.001000 Step 200: 0.000707 Step 499: 0.000448
The learning rate climbs linearly during warmup (steps 0 to 100), peaks at step 100, and then decays as the inverse square root of the step count.
Key Parameters
The key parameters to configure for each schedule are:
- Step decay:
step_size(epochs between drops),gamma(multiplicative factor, typically 0.1 or 0.5) - Exponential decay:
gamma(per-step multiplier, typically 0.99 to 0.999 for step-wise or 0.1 to 0.5 per epoch) - Inverse square root:
warmup_steps(typically 1 to 10 percent of total steps),eta_peak(the peak learning rate) - Cosine annealing:
T_max(period length),eta_min(minimum value, often ) - Warm restarts:
T_0(initial period),T_mult(cycle growth factor, 1 for fixed cycles, 2 for doubling)
Using Hugging Face Schedulers
The Hugging Face transformers library provides convenient schedule constructors that are particularly useful for fine-tuning language models:
from transformers import (
get_cosine_schedule_with_warmup,
get_linear_schedule_with_warmup,
)
model_hf = TinyModel()
optimizer_hf = optim.AdamW(model_hf.parameters(), lr=2e-5)
hf_warmup_steps = 100
hf_total_steps = 1000
scheduler_hf_linear = get_linear_schedule_with_warmup(
optimizer_hf,
num_warmup_steps=hf_warmup_steps,
num_training_steps=hf_total_steps,
)
model_hf2 = TinyModel()
optimizer_hf2 = optim.AdamW(model_hf2.parameters(), lr=2e-5)
scheduler_hf_cosine = get_cosine_schedule_with_warmup(
optimizer_hf2,
num_warmup_steps=hf_warmup_steps,
num_training_steps=hf_total_steps,
)
hf_linear_lrs = record_lr_schedule(
optimizer_hf, scheduler_hf_linear, n_steps=hf_total_steps
)
hf_cosine_lrs = record_lr_schedule(
optimizer_hf2, scheduler_hf_cosine, n_steps=hf_total_steps
)HuggingFace linear schedule at key steps: Step 0: 0.00000000 Step 50: 0.00001000 Step 100: 0.00002000 Step 500: 0.00001111 Step 999: 0.00000002
These functions handle the warmup transition automatically. The num_warmup_steps parameter sets the warmup duration, and num_training_steps sets the total. The schedule is then applied via scheduler.step() after each optimizer step, exactly like any PyTorch scheduler.

Worked Example: Effect on Convergence
To see concretely how schedule choice affects training, let us train a small neural network on a synthetic regression task using four different schedules and compare convergence.
torch.manual_seed(42)
np.random.seed(42)
n_samples = 400
X_data = torch.randn(n_samples, 10)
true_weights = torch.randn(10)
y_data = X_data @ true_weights + 0.1 * torch.randn(n_samples)
dataset = torch.utils.data.TensorDataset(X_data, y_data)
dataloader = torch.utils.data.DataLoader(dataset, batch_size=32, shuffle=True)def train_with_schedule(scheduler_type, n_epochs=60, lr=0.05):
net = nn.Linear(10, 1)
optimizer = optim.SGD(net.parameters(), lr=lr, momentum=0.9)
criterion = nn.MSELoss()
if scheduler_type == "step":
scheduler = optim.lr_scheduler.StepLR(
optimizer, step_size=20, gamma=0.1
)
elif scheduler_type == "exponential":
gamma_e = (1e-4 / lr) ** (1 / n_epochs)
scheduler = optim.lr_scheduler.ExponentialLR(optimizer, gamma=gamma_e)
elif scheduler_type == "cosine":
scheduler = optim.lr_scheduler.CosineAnnealingLR(
optimizer, T_max=n_epochs, eta_min=1e-5
)
else:
scheduler = None
losses = []
for epoch in range(n_epochs):
epoch_loss = 0.0
for X_batch, y_batch in dataloader:
optimizer.zero_grad()
pred = net(X_batch).squeeze()
loss = criterion(pred, y_batch)
loss.backward()
optimizer.step()
epoch_loss += loss.item()
losses.append(epoch_loss / len(dataloader))
if scheduler is not None:
scheduler.step()
return losses
n_epochs = 60
losses_step = train_with_schedule("step", n_epochs=n_epochs)
losses_exp = train_with_schedule("exponential", n_epochs=n_epochs)
losses_cosine = train_with_schedule("cosine", n_epochs=n_epochs)
losses_constant = train_with_schedule("constant", n_epochs=n_epochs)Final training loss (epoch 60): Step decay: 0.00985 Exponential: 0.01002 Cosine annealing: 0.00997 Constant LR: 0.01206

The constant learning rate model shows the highest final loss because its updates remain large enough to keep it moving around the minimum. All three decay schedules finish lower in this run. The logarithmic axis makes those late-stage differences visible without hiding the much larger loss reduction during the first few epochs.
Interpreting the Convergence Curves
The convergence plot reveals more than just final loss values. Notice how each schedule's characteristic shape shows up in the training loss.
For step decay, scheduled cuts at epochs 20 and 40 reduce the size of subsequent updates. In this noisy mini-batch run, those transitions are subtler than the learning-rate staircase itself; the clearest effect is the tighter band of losses after the first cut. A large loss drop at a schedule boundary can indicate that the previous rate was causing oscillation, but it is not guaranteed in every run.
For exponential decay, the loss falls smoothly throughout. There are no phase transitions visible in the loss curve, which is a double-edged sword: smooth training is easy to monitor, but you lose the visual feedback that phase boundaries provide. If exponential decay is decaying too fast, you will see the loss plateau early, while the learning rate is still non-negligible, with no obvious cause.
For cosine annealing, the loss curves similarly smoothly. The cosine schedule's gradual start means the learning rate at the very beginning is nearly identical to the initial value, so there is no initial benefit from decay. The benefit accumulates through the middle and end of training.
Comparing Schedules at a Glance
The table below summarizes the key properties of each schedule to guide selection.
| Schedule | Shape | Best for | Limitations |
|---|---|---|---|
| Step decay | Staircase | CV training, interpretability | Rigid timing, requires epoch tuning |
| Exponential | Smooth monotone | Short runs with a floor | Decays too fast without tuning |
| Inverse sqrt | Fast drop then plateau | Transformer pre-training | Needs warmup configuration |
| Cosine | Smooth S-curve | Fine-tuning, general use | Requires total steps known upfront |
| Cosine + restarts | Cyclic | Exploratory training | Cycle length sensitive |
Choosing a Schedule
No single schedule is universally best. The right choice depends on the model architecture, optimizer, dataset size, and training budget. What follows is a practical guide to making this choice, organized by use case.
Training Large Transformers from Scratch
For training large transformer models from scratch, the inverse square root schedule with linear warmup is the canonical default. It was validated on the original Transformer and remains widely used for pre-training. The slow later decay means the model continues to see meaningful learning rate values even late in training.
The warmup length for large models is typically 4,000 to 10,000 steps. The peak learning rate is usually in the range of to for model dimensions around 512-1024. For larger models, the peak is often scaled inversely with the square root of the model dimension, following the Transformer paper's formula.
Some recent large language model training runs have moved away from the inverse square root schedule in favor of cosine annealing with a fixed total step count. The argument is that cosine annealing with a known endpoint is more predictable and allows planning the training budget in advance. The inverse square root schedule is "endless" in the sense that there is no natural stopping point, which complicates budget planning.
Fine-Tuning Pre-Trained Models
For fine-tuning pre-trained models, linear or cosine decay is preferred because fine-tuning involves much shorter training runs. The learning rate needs to reach a very small value by the end of training to avoid disrupting the pre-trained weights. Cosine decay provides a smooth, predictable endpoint.
The peak learning rate for fine-tuning is typically much smaller than for pre-training. For BERT-style models, common values are to . For larger models like GPT-style architectures, the fine-tuning learning rate is often even smaller, sometimes as low as , to prevent catastrophic forgetting of pre-trained knowledge.
The warmup period for fine-tuning is shorter than for pre-training, typically 5 to 10 percent of total steps. Since the model's weights are already in a reasonable part of parameter space (thanks to pre-training), the risk of early instability is lower. The warmup mainly serves to let Adam's moment estimates calibrate.
Computer Vision with SGD
For computer vision with SGD, step decay remains common. The explicit phase structure matches the training recipes for ResNet and related architectures, and the abrupt drops are known to produce good convergence behavior.
The standard three-phase recipe for ImageNet training is: 90 epochs total, starting learning rate of 0.1, drops of 10x at epochs 30 and 60. This recipe has been so successful and widely validated that it has become a community reference point. Many papers compare their training protocols against this baseline.
When you switch from SGD to AdamW for computer vision, the step decay recipe often needs adjustment. Adam's internal adaptation changes the effective sensitivity to the global schedule, and the ten-fold drops that work well for SGD can be too aggressive for Adam. A cosine schedule with AdamW tends to transfer more reliably across different model scales and training durations.
When in Doubt
When in doubt between linear and cosine decay, cosine decay is slightly preferred because it is slower to start falling (reducing the learning rate less aggressively early in decay) and ends more smoothly. In practice, the difference is small compared to the impact of other hyperparameters. If you are setting up a new experiment and have no prior experience with the task, cosine annealing with linear warmup using 5 to 10 percent of total steps is a reasonable starting point.
The most important hyperparameter to tune is the peak learning rate, not the schedule shape. A cosine schedule with a poorly tuned peak learning rate will underperform a linear schedule with the right peak learning rate. Spend your tuning budget on the peak learning rate first, and only refine the schedule shape once that is set.
Limitations and Practical Considerations
Learning rate scheduling is powerful but introduces additional hyperparameters that must be set correctly. A poorly chosen schedule can be worse than no schedule at all. If the learning rate decays too quickly, the model gets stuck in a suboptimal region before it has had time to find a good basin. If it decays too slowly, the final convergence is poor and the training may appear to have plateaued.
Schedule Sensitivity and Hyperparameter Overhead
Each schedule introduces at least one and often several additional hyperparameters beyond the base learning rate. Step decay needs both the step size and the gamma. Cosine annealing needs the minimum learning rate and the total period length. The inverse square root schedule needs the warmup duration. All of these interact with the peak learning rate in non-trivial ways.
In practice, these hyperparameters are often set based on rules of thumb derived from prior experiments rather than being tuned from scratch. This is a practical necessity: running a full hyperparameter sweep over schedule parameters for every new task is computationally infeasible. The standard recipes that have emerged from community experience (like linear warmup to 10% of steps followed by cosine decay) are valuable precisely because they generalize across many settings without requiring explicit tuning.
The Interaction with Adaptive Optimizers
The interaction between schedules and adaptive optimizers like Adam is not fully understood. Adam's internal adaptation already provides a form of per-parameter learning rate control, and layering a global schedule on top creates a complex combined behavior. Some researchers have argued that with properly tuned Adam, the sensitivity to the global learning rate schedule is lower than with SGD, though in practice schedules still affect performance.
One specific concern is that Adam's second moment estimate accumulates history according to its decay factor . When the learning rate is decayed and then kept constant for a long time (as in step decay), the effective step size changes both because of the explicit decay and because continues to evolve. The combined effect is not easy to predict analytically. This is part of why cosine and linear schedules, which change the learning rate smoothly and predictably, tend to behave more reliably with Adam than step decay does.
Warmup in Distributed Training
Warmup deserves particular attention in distributed training. When training with large batch sizes across many GPUs, the effective batch size is the product of the per-GPU batch size and the number of GPUs. The linear scaling rule suggests scaling the learning rate proportionally to the batch size, which can push the learning rate very high. Warmup becomes critical in this regime: starting at the full scaled learning rate with a large batch and random initialization can cause catastrophic early instability. Longer warmup periods are used for very large distributed training runs.
The challenge is that the standard warmup durations (4,000 to 10,000 steps) were designed for moderate batch sizes. When batch size is scaled by a factor of 64x for distributed training, the effective number of samples seen per step increases dramatically, and the notion of "a step" in the warmup duration becomes less meaningful. Some large-scale training recipes specify warmup in terms of tokens or samples processed rather than steps, which remains invariant to batch size scaling.
Cycle-Based Schedules in Practice
Cycle-based schedules like cosine annealing with warm restarts have an appealing theoretical motivation but require careful setting of the cycle length. A cycle that is too short means the model restarts before it has converged within each basin, wasting each cycle. A cycle that is too long is effectively just cosine annealing without restarts. In practice, the cycle length is often tuned as a hyperparameter or set based on prior experience with similar models.
The restart mechanism can also cause difficulties for checkpoint selection. With a standard monotone schedule, the best checkpoint is usually the final one. With warm restarts, the best checkpoint might be at the end of any cycle. You need to track cycle-end checkpoints explicitly and compare them, rather than simply saving the final model. This adds infrastructure complexity to training pipelines.
Generalization and the Learning Rate at Convergence
The relationship between schedules and generalization is subtle. It is well established empirically that decaying the learning rate to a small value at the end of training improves generalization, not just training loss. The intuition is that a small learning rate at convergence means the model has found a region where the loss surface is locally flat, and flat minima tend to generalize better. But this connection is not a tight theoretical guarantee, and counter-examples exist.
More precisely, the "flat minima generalize better" hypothesis has been the subject of ongoing debate. Some theoretical work argues that the measure of flatness depends on the parameterization and can be an artifact of scale. Others have shown empirically that the correlation between flatness and generalization holds across a wide range of architectures and tasks. In practice, ending training at a small learning rate reliably improves results, even if the theoretical mechanism is not fully settled.
Summary
Learning rate decay is a fundamental component of neural network training that bridges the gap between fast early learning and precise final convergence. The core problem is that a fixed learning rate cannot simultaneously provide the large steps needed for rapid progress early in training and the small steps needed for stable convergence later. Schedules resolve this by continuously adjusting the learning rate over time.
The key schedules are:
- Step decay: multiplies the learning rate by a factor at fixed epoch intervals, producing a staircase. Easy to interpret and historically common in computer vision. The sharp drops provide clear phase boundaries in training.
- Exponential decay: applies a constant multiplicative factor at every step, producing a smooth exponential curve. Simple but can decay too aggressively without a floor or careful gamma selection.
- Inverse square root decay: falls quickly at first and then levels off, making it well-matched to optimization dynamics. The standard choice for training transformers from scratch, especially combined with linear warmup.
- Cosine annealing: follows a smooth half-cosine curve from maximum to minimum. Widely used for fine-tuning and naturally extended with warm restarts for exploratory training where escaping local minima matters.
All schedules benefit from linear warmup, which prevents instability caused by noisy gradients and uncalibrated optimizer statistics at the start of training. The warmup length is typically 1 to 10 percent of total steps, with longer warmup for larger models and batch sizes.
The choice of schedule interacts with the optimizer. AdamW with cosine annealing and linear warmup has become the default for most NLP fine-tuning work. Inverse square root with warmup remains the standard for large-scale pre-training. SGD with step decay persists in computer vision training recipes that have been validated at scale.
In practice, the peak learning rate is more important to tune than the schedule shape. A well-tuned peak learning rate with any reasonable schedule will outperform a poorly tuned peak with the optimal schedule. Start by finding a good peak learning rate, then refine the schedule if further improvement is needed. The differences between well-tuned schedules are often small compared to the impact of the peak learning rate itself.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about learning rate decay and scheduling.
Learning Rate Decay Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!