Learning Rate Decay: Step, Exponential

Michael BrenndoerferJanuary 24, 202651 min read

Part of Language AI Handbook

Explains how learning rate schedules improve neural network training. Topics include step decay, exponential decay, inverse square root with warmup.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Learning Rate Decay

Training a neural network is not a static process. The learning rate that works well early in training, when the model is far from any good solution and needs large steps to make progress, often becomes a liability later, when the model is close to a good solution and large steps cause it to overshoot. Learning rate decay, also called learning rate scheduling, addresses this mismatch by reducing the learning rate over the course of training.

The intuition is simple: think of gradient descent as hiking down a mountain in heavy fog. Early on, you can stride confidently in roughly the right direction. But as you approach the valley floor, you need to take smaller, more careful steps to avoid stumbling past the lowest point. A fixed learning rate cannot adapt to this changing landscape. A decaying schedule adjusts automatically, enabling fast initial progress and precise final convergence.

This chapter covers the main decay strategies: step decay, exponential decay, inverse square root decay, and cosine annealing. For each, we examine the mathematical formulation, the practical intuition, and when to prefer one over another. We also look at how these schedules interact with adaptive optimizers like Adam, and how modern training pipelines compose multiple schedules together. Along the way, we build concrete intuition for what happens inside optimization when the learning rate changes, and we look at the historical context that shaped each approach.

Why the Learning Rate Must Change

Before examining specific schedules, it is worth understanding precisely why a fixed learning rate causes problems at both ends of training. The issue is not merely cosmetic. A fixed learning rate is a fundamental mismatch with the geometry of neural loss surfaces, and understanding that mismatch explains why every major training recipe today uses some form of scheduling.

The Two-Regime Problem

At the start of training, the model parameters are randomly initialized and far from any good solution. The loss surface at this point is often steep and the gradients are large. A moderate learning rate causes the model to take meaningful steps and make rapid progress. If the learning rate is too small at this stage, training is unnecessarily slow: the model makes tiny moves when large moves are available and beneficial. You are walking when you could be running.

Later in training, the model has moved into a region near a local minimum (or saddle point) where the loss landscape is much flatter. Gradients become smaller and the curvature of the landscape matters more. A learning rate that was appropriate for the steep early landscape now causes the model to bounce back and forth across a shallow basin rather than settling into it. The model never fully converges: it oscillates around the minimum with a loss that fluctuates rather than decreasing smoothly. You are trying to settle into a shallow depression with steps large enough to carry you across it entirely.

Learning Rate

The learning rate η\eta controls how large a step gradient descent takes in the direction of the negative gradient. At each update, parameters move by θ←θ−η⋅∇θL\theta \leftarrow \theta - \eta \cdot \nabla_\theta \mathcal{L}. A larger η\eta makes faster but less precise updates; a smaller η\eta makes slower but more precise updates.

A decaying schedule interpolates between these two needs by starting large and shrinking over time. The key design choices are: how quickly to decay, what functional form the decay follows, and what minimum learning rate to allow.

Loss Landscape Geometry and Why It Changes

The loss landscape is not uniform throughout parameter space. Near initialization, the landscape is typically dominated by large-scale structure: the model is trying to learn basic patterns (for a language model, that might mean learning that certain words co-occur frequently, or that sentence structure follows certain patterns), and the gradients that point toward these basic patterns are large and relatively consistent across batches. A high learning rate is safe because the direction of improvement is clear.

As training proceeds and the model captures these large-scale patterns, the remaining signal becomes more subtle. The model is now trying to refine fine-grained distinctions, and the gradients that carry this information are smaller. The loss surface near a good local minimum looks like a shallow bowl, and convergence requires increasingly precise steps toward its center. With a fixed step size that is too large, the model overshoots the center every step, landing on the far wall, bouncing back, and oscillating indefinitely.

The mathematical reason for this oscillation comes from the relationship between the learning rate and the curvature. If the maximum eigenvalue of the Hessian (the matrix of second derivatives) at a point is λmax⁡\lambda_{\max}, then gradient descent will converge stably only if the learning rate satisfies η<2/λmax⁡\eta < 2 / \lambda_{\max}. Near a minimum where curvature is high, this constraint is tight. Early in training, where curvature is lower and the landscape is smoother, the same constraint is much looser and larger learning rates are safe.

Historical Context: Where Scheduling Came From

The practice of reducing the learning rate during training predates deep learning. Stochastic gradient descent theory, developed for convex optimization in the 1950s and 1960s, established convergence guarantees under conditions that require the learning rate to satisfy the Robbins-Monro conditions:

∑t=1∞ηt=∞and∑t=1∞ηt2<∞\sum_{t=1}^{\infty} \eta_t = \infty \quad \text{and} \quad \sum_{t=1}^{\infty} \eta_t^2 < \infty

The first condition ensures the optimizer can reach any part of the parameter space; the second ensures that individual steps eventually become small enough to allow convergence. A constant learning rate satisfies the first condition but not the second (the sum of its squares diverges). The inverse square root schedule, which decays as t−1/2t^{-1/2}, satisfies both conditions, which is one reason it carries such theoretical weight.

Practical deep learning inherited this intuition. Early neural network training papers in the 1980s and 1990s routinely described reducing the learning rate toward the end of training. The specific form varied: some halved it manually, others used exponential decay. The systematic study and naming of these schedules as distinct techniques emerged with the larger-scale experiments of the 2010s, when training runs on ImageNet made hyperparameter sensitivity clearly visible.

Step Decay

Step decay is the simplest approach and the one historically used in early deep learning. The learning rate is held constant for a fixed number of epochs, then multiplied by a decay factor, then held constant again, and so on. This produces a staircase pattern in the learning rate over time.

Formulation

The update rule at epoch tt is:

ηt=η0⋅γ⌊t/s⌋\eta_t = \eta_0 \cdot \gamma^{\lfloor t / s \rfloor}

where:

  • η0\eta_0: the initial learning rate
  • γ\gamma: the decay factor, typically between 0.1 and 0.5 (so the learning rate is halved, quartered, or reduced by 10x at each step)
  • ss: the step size in epochs (how many epochs between drops)
  • ⌊t/s⌋\lfloor t / s \rfloor: the floor of t/st / s, which counts how many complete steps have elapsed

For example, with η0=0.1\eta_0 = 0.1, γ=0.1\gamma = 0.1, and s=30s = 30, the learning rate starts at 0.1, drops to 0.01 at epoch 30, drops to 0.001 at epoch 60, and so on. Each drop reduces the learning rate by exactly a factor of ten.

The floor function ⌊t/s⌋\lfloor t / s \rfloor is what creates the staircase: it stays constant for ss consecutive epochs, then increments by exactly one. This means the learning rate at epoch 29 is identical to the learning rate at epoch 0, and the learning rate at epoch 30 is exactly γ\gamma times that.

Intuition

Step decay appeals to practitioners because it is easy to understand and diagnose. You can look at a training loss curve and see exactly when each drop happened. If the loss plateaus before a drop, that suggests the drop might be too late. If the loss jumps up briefly after a drop, the new learning rate might be too small. This interpretability makes step decay a useful debugging tool. Every phase of training is clearly delineated by the drop, and you can reason about what the model was doing in each phase.

The staircase structure also means training can be naturally organized into phases. Each phase trains at a fixed learning rate until the loss stops improving, then the next phase begins at a lower rate to refine the solution. Many early ResNet and VGG training recipes used this pattern: train for 30 epochs at 0.1, drop to 0.01, train 30 more, drop to 0.001, and finish. He et al.'s original ResNet paper for ImageNet used exactly this recipe, and the robustness of this approach across many architectures is what established its dominance in computer vision through the early 2010s.

The main limitation is rigidity. The schedule is specified in advance and does not adapt to the observed training progress. If a particular run converges faster or slower than expected, the drops happen at the wrong time. A model that has already plateaued two epochs before the scheduled drop is waiting unnecessarily. A model that is still making rapid progress when the drop arrives gets its learning rate cut during a productive phase. Some practitioners address this by monitoring validation loss and dropping the rate when it plateaus rather than at fixed epochs, a variant called ReduceLROnPlateau. This adaptive version maintains the staircase structure while letting training dynamics determine when each step occurs.

The Multiplicative Factor and Its Consequences

The choice of γ\gamma has a significant impact on the training dynamics. A very aggressive factor like γ=0.1\gamma = 0.1 (ten-fold reduction) means that after the drop, the model is taking steps one-tenth as large. This can effectively freeze learning on some parameters: if the gradient for a particular parameter is small at the pre-drop learning rate, it becomes negligible after a ten-fold reduction. The drop essentially shifts the focus of training to parameters with larger gradients.

A gentler factor like γ=0.5\gamma = 0.5 (halving) is less dramatic. The model slows down but does not stop making meaningful progress on any parameter. Multiple gentler drops can sometimes achieve the same end result as one aggressive drop while maintaining smoother loss curves.

The step size ss determines how long the model trains at each level. Setting ss too small means the model does not have time to exploit each learning rate before it drops. Setting ss too large means the model may have plateaued long before the next drop. The optimal ss is essentially the number of epochs it takes for the model to plateau at each learning rate level, which depends on the dataset, architecture, and optimizer.

Exponential Decay

Exponential decay smooths out the staircase by applying decay at every step rather than at discrete intervals. The learning rate decreases continuously as a smooth exponential curve.

Formulation

At step tt, the learning rate is:

ηt=η0⋅e−λt\eta_t = \eta_0 \cdot e^{-\lambda t}

where:

  • η0\eta_0: the initial learning rate
  • λ\lambda: the decay rate, a positive constant controlling how quickly the rate falls
  • tt: the current training step (or epoch)

An equivalent formulation uses a decay factor γ\gamma per step:

ηt=η0⋅γt\eta_t = \eta_0 \cdot \gamma^t

where γ=e−λ\gamma = e^{-\lambda}. The two are mathematically identical; the second is often more convenient because you can set γ\gamma directly (e.g., γ=0.999\gamma = 0.999 for slow decay or γ=0.99\gamma = 0.99 for faster decay). The relationship γ=e−λ\gamma = e^{-\lambda} means λ=−ln⁡(γ)\lambda = -\ln(\gamma), so a decay factor of 0.99 corresponds to a rate of approximately 0.01.

Choosing the Decay Rate

The relationship between γ\gamma and training duration matters enormously. After TT steps with per-step factor γ\gamma, the learning rate is η0⋅γT\eta_0 \cdot \gamma^T. If you want the learning rate to end at, say, 1% of its initial value after 100 epochs, you need γ100=0.01\gamma^{100} = 0.01, which gives γ=0.011/100≈0.955\gamma = 0.01^{1/100} \approx 0.955. You can work backward from your desired endpoint to choose γ\gamma.

This backward calculation is a useful habit. Before setting γ\gamma, decide what fraction of the initial learning rate you want at the end of training, and how many steps training will run. The required γ\gamma follows directly:

γ=(ηfinalη0)1/T\gamma = \left(\frac{\eta_{\text{final}}}{\eta_0}\right)^{1/T}

This ensures the schedule hits your target endpoint rather than decaying too fast or too slow.

Intuition and Limitations

The exponential function has a special property: its rate of change is proportional to its current value. This means exponential decay reduces the learning rate by the same multiplicative factor at every step. Early in training, the absolute reduction is large (because the learning rate is high); later in training, the absolute reduction is small. The relative reduction is constant throughout.

In practice, the effect of this is that exponential decay can shrink the learning rate too quickly, especially with even moderately large λ\lambda. After tt steps, the learning rate has shrunk by a factor of e−λte^{-\lambda t}, which can become extremely small well before training ends. Practitioners often add a floor: a minimum learning rate below which the schedule does not go, preventing the rate from becoming so small that training effectively stops.

Without a floor, exponential decay with γ=0.9\gamma = 0.9 per epoch would reach 0.9100≈2.6×10−50.9^{100} \approx 2.6 \times 10^{-5} after 100 epochs starting from η0=0.1\eta_0 = 0.1, which is effectively zero. With a floor of 10−510^{-5}, the schedule plateaus at that value and training continues to make at least some progress.

The continuous nature of exponential decay is both a strength and a weakness. There are no abrupt transitions to cause loss spikes, but the schedule also never provides the sharp signal of a step drop that can sometimes help the model escape a local plateau. Empirically, exponential decay tends to work well when training is well-behaved and the main concern is smooth, gradual refinement rather than dramatic transitions between phases.

In[5]:
Code
import numpy as np


def step_decay(t, eta0, gamma, step_size):
    return eta0 * (gamma ** (t // step_size))


def exponential_decay(t, eta0, gamma):
    return eta0 * (gamma**t)


def inv_sqrt_decay(t, eta0):
    return eta0 / np.sqrt(t + 1)


def cosine_decay(t, eta_max, eta_min, T_max):
    return eta_min + 0.5 * (eta_max - eta_min) * (1 + np.cos(np.pi * t / T_max))


epochs = np.arange(0, 100)
eta0 = 0.1
step_lrs = np.array(
    [step_decay(t, eta0=eta0, gamma=0.1, step_size=30) for t in epochs]
)
exp_lrs = np.array(
    [exponential_decay(t, eta0=eta0, gamma=0.955) for t in epochs]
)
inv_sqrt_lrs = np.array([inv_sqrt_decay(t, eta0=eta0) for t in epochs])
cosine_lrs = np.array(
    [cosine_decay(t, eta_max=eta0, eta_min=1e-4, T_max=100) for t in epochs]
)
Out[6]:
Visualization
Line plot comparing four learning rate schedules over 100 epochs on a log scale.
Four common learning rate schedules over 100 training epochs, plotted on a logarithmic scale. Step decay creates a staircase pattern with abrupt drops at epochs 30 and 60. Exponential decay falls smoothly at a constant multiplicative rate. Inverse square root decay drops quickly early and then flattens, staying higher at late epochs than exponential decay. Cosine annealing follows a smooth curve from maximum to near-zero.

The plot shows the four schedules on a logarithmic scale. Step decay drops abruptly at epochs 30 and 60. Exponential decay falls smoothly but steeply. Inverse square root decay drops quickly at first and then flattens. Cosine annealing curves smoothly from its maximum to its minimum.

Inverse Square Root Decay

Inverse square root decay, popularized by the original Transformer paper "Attention Is All You Need," has a distinctive profile: it decreases quickly at first and then levels off gradually. This behavior makes it well-matched to how optimization often proceeds: rapid early progress followed by slower refinement.

Mathematical Foundation

The schedule is named for its decay rate: the learning rate is proportional to the inverse of the square root of the step count. This means the rate falls rapidly at first (the difference between step 1 and step 4 is large) and then changes slowly at large step counts (the difference between step 10,000 and step 10,001 is tiny). On a log-log scale, the decay appears as a straight line with slope −1/2-1/2.

The schedule is:

ηt=ηpeak⋅twarmupmax⁡(t, twarmup)\eta_t = \eta_{\text{peak}} \cdot \frac{\sqrt{t_{\text{warmup}}}}{\sqrt{\max(t,\, t_{\text{warmup}})}}

where:

  • ηpeak\eta_{\text{peak}}: the peak learning rate reached after warmup
  • twarmupt_{\text{warmup}}: the number of warmup steps
  • tt: the current step

This form describes the decay phase while capping the rate at ηpeak\eta_{\text{peak}} for t≤twarmupt \leq t_{\text{warmup}}. In practice, it is usually paired with a separate linear warmup that increases the learning rate from zero to the peak. After warmup, the max⁡(t,twarmup)\max(t, t_{\text{warmup}}) term equals tt, and the schedule follows a true inverse square root curve.

A common combined form that includes the linear warmup is the original Transformer formulation:

ηt=dmodel−0.5⋅min⁡ ⁣(t−0.5,  t⋅twarmup−1.5)\eta_t = d_{\text{model}}^{-0.5} \cdot \min\!\left(t^{-0.5},\; t \cdot t_{\text{warmup}}^{-1.5}\right)

where:

  • dmodeld_{\text{model}}: the model hidden dimension (used as a scaling factor in the original Transformer)
  • The first argument t−0.5t^{-0.5} governs the decay phase
  • The second argument t⋅twarmup−1.5t \cdot t_{\text{warmup}}^{-1.5} governs the linear warmup phase

The minimum of the two quantities ensures a smooth transition: during warmup the linear term is smaller (it is small because tt is small), and after warmup the inverse square root term takes over (it is smaller because t−0.5t^{-0.5} shrinks faster than the linear term grows). The crossover happens exactly at t=twarmupt = t_{\text{warmup}}, where both terms equal twarmup−0.5t_{\text{warmup}}^{-0.5}.

Comparing Decay Rates Across Schedules

To understand why inverse square root is preferred over exponential for long training runs, it helps to compare how aggressively each schedule reduces the learning rate. With exponential decay at γ=0.999\gamma = 0.999 per step, after 100,000 steps the learning rate has fallen to 0.999100000≈4.5×10−440.999^{100000} \approx 4.5 \times 10^{-44}, which is effectively zero. With inverse square root, after 100,000 steps from a warmup of 4,000 steps, the learning rate is 4000/100000≈0.2\sqrt{4000} / \sqrt{100000} \approx 0.2 times the peak, a reduction to 20% of peak. For pre-training runs that span hundreds of thousands of steps, this gentler decay means the model is still making meaningful learning rate-scaled gradient steps even late in training.

This is particularly important for large language model pre-training, where the model is learning from a dataset so large that many distinct learning signals appear throughout training. An exponential schedule would reduce the learning rate to near zero long before the dataset is exhausted, effectively stopping learning. The inverse square root schedule allows the model to continue learning from later portions of the data with a non-negligible step size.

Warmup: Why Start Small?

The warmup period addresses a specific training instability. At the very start of training, the model parameters are random, the gradient estimates are highly noisy, and the loss surface is poorly characterized. A large learning rate in this regime can cause the loss to explode or send the model into a region of parameter space that is very difficult to escape.

Linear warmup starts the learning rate near zero and increases it steadily over the first several thousand steps. By the time the learning rate reaches its peak value, the model has already made some initial progress, gradient estimates have stabilized, and the optimizer's momentum and variance estimates (in Adam) have had time to accumulate. The peak learning rate can then safely be much larger than if training had started at full speed.

The warmup is especially important for Adam. Adam estimates the first and second moments of the gradient using exponential moving averages. At step one, these estimates are initialized to zero and have not yet had time to reflect the true gradient distribution. Early updates with a large learning rate are based on these poorly calibrated estimates, which can cause wild oscillations. Warmup gives the moment estimates time to converge before the learning rate reaches its full value.

Warmup Steps

A warmup period is a phase at the beginning of training during which the learning rate is increased from a small value to the target rate. Linear warmup increases the rate by a fixed amount each step. Warmup prevents instability caused by noisy gradients and poorly calibrated optimizer statistics early in training.

The warmup length is typically set to 4,000 to 10,000 steps for transformer training, though the optimal value depends on batch size and model size. Larger models and larger batches generally benefit from longer warmup. With very large batch sizes (as in distributed training), the signal-to-noise ratio in gradient estimates is higher because more samples are averaged, but the learning rate is also larger due to linear scaling, so the need for warmup does not disappear.

Warmup Length and the Batch Size Relationship

An underappreciated aspect of warmup is how it interacts with batch size. When you double the batch size, the linear scaling rule (introduced by Goyal et al. in 2017 for ImageNet training with SGD) recommends doubling the learning rate to maintain the same relative update magnitude. But doubling the learning rate without extending warmup often causes instability. The convention in large-batch training is to also scale the warmup duration: if you multiply batch size by kk, multiply warmup steps by kk as well. This gives the optimizer time to accumulate reliable statistics before taking the larger steps that the scaled learning rate demands.

In[7]:
Code
def transformer_schedule(t, d_model, warmup_steps):
    t = max(t, 1)
    return d_model ** (-0.5) * min(t ** (-0.5), t * warmup_steps ** (-1.5))


d_model = 512
warmup_steps = 4000
steps = np.arange(1, 40001)
transformer_lrs = np.array(
    [transformer_schedule(t, d_model, warmup_steps) for t in steps]
)
peak_step = int(np.argmax(transformer_lrs))
peak_lr = float(transformer_lrs[peak_step])
Out[8]:
Console
Peak learning rate: 0.000699
Peak reached at step: 4000
LR at step 10000: 0.000442
LR at step 40000: 0.000221

The peak learning rate is reached exactly at the warmup step, after which the schedule decays as t−0.5t^{-0.5}. The learning rate at step 40,000 is roughly 32% of its peak value, showing how slowly the inverse square root falls compared to exponential decay.

Out[9]:
Visualization
Line plot of Transformer learning rate with linear warmup and inverse square root decay over 40000 steps.
Transformer learning rate schedule for d_model=512 with 4,000 warmup steps. The learning rate rises linearly during warmup, peaks at step 4,000, and then decays as the inverse square root of the step count. The slow decay after the peak keeps the learning rate meaningfully large even at step 40,000, which is important for long pre-training runs.

Cosine Annealing

Cosine annealing uses a half-cosine curve to reduce the learning rate from a maximum value to a minimum value over a fixed number of steps. It became widely adopted after Loshchilov and Hutter introduced it in 2017 and combined it with warm restarts to create SGDR (Stochastic Gradient Descent with Warm Restarts).

Why Cosine?

The choice of a cosine curve rather than a simple straight line or exponential is motivated by the shape of the decay, not just convenience. At the start of the cosine curve (near t=0t = 0), the cosine function changes slowly, meaning the learning rate stays near its maximum for a while and decreases gently at first. Near the end (as tt approaches Tmax⁡T_{\max}), the cosine function is near −1-1 and also changing slowly, meaning the learning rate settles near its minimum and finishes smoothly. In the middle, the cosine passes through its steepest section, driving the sharpest decline.

This shape is intuitively appealing: the model starts at full speed, gradually shifts into a slower and slower rate, and comes to a gentle stop. Contrast this with linear decay, which reduces the learning rate at a constant rate regardless of where training is, and step decay, which drops abruptly. Cosine annealing provides a natural deceleration curve.

The cosine also has a nice mathematical property: the derivative of the cosine schedule is zero at both endpoints. This means the learning rate changes most slowly exactly at the start and end of the schedule, which is often where training is most sensitive to abrupt changes.

Formulation

The basic cosine annealing schedule is:

ηt=ηmin⁡+12(ηmax⁡−ηmin⁡) ⁣(1+cos⁡ ⁣(πtTmax⁡))\eta_t = \eta_{\min} + \frac{1}{2}\left(\eta_{\max} - \eta_{\min}\right)\!\left(1 + \cos\!\left(\frac{\pi t}{T_{\max}}\right)\right)

where:

  • ηmin⁡\eta_{\min}: the minimum learning rate (often 0 or a small positive value like 10−510^{-5})
  • ηmax⁡\eta_{\max}: the maximum learning rate
  • Tmax⁡T_{\max}: the number of steps for one half-cosine period
  • tt: the current step within the period

At t=0t = 0, the cosine term is cos⁡(0)=1\cos(0) = 1, giving η0=ηmax⁡\eta_0 = \eta_{\max}. At t=Tmax⁡t = T_{\max}, the cosine term is cos⁡(π)=−1\cos(\pi) = -1, giving ηTmax⁡=ηmin⁡\eta_{T_{\max}} = \eta_{\min}. In between, the rate follows the smooth cosine curve. The factor of 1/21/2 ensures the output lies in the range [ηmin⁡,ηmax⁡][\eta_{\min}, \eta_{\max}].

Worked Numerical Example

Let us trace through a specific case to build intuition. Suppose ηmax⁡=0.1\eta_{\max} = 0.1, ηmin⁡=0.001\eta_{\min} = 0.001, and Tmax⁡=100T_{\max} = 100.

At t=0t = 0:

η0=0.001+12(0.1−0.001)(1+cos⁡(0))=0.001+12(0.099)(2)=0.1\eta_0 = 0.001 + \frac{1}{2}(0.1 - 0.001)(1 + \cos(0)) = 0.001 + \frac{1}{2}(0.099)(2) = 0.1

At t=25t = 25 (one quarter through):

η25=0.001+12(0.099)(1+cos⁡ ⁣(π⋅25100))=0.001+12(0.099)(1+cos⁡(0.25π))\eta_{25} = 0.001 + \frac{1}{2}(0.099)\left(1 + \cos\!\left(\frac{\pi \cdot 25}{100}\right)\right) = 0.001 + \frac{1}{2}(0.099)(1 + \cos(0.25\pi))

Since cos⁡(0.25π)=cos⁡(45°)≈0.707\cos(0.25\pi) = \cos(45°) \approx 0.707, this gives approximately 0.001+0.0495×1.707≈0.0860.001 + 0.0495 \times 1.707 \approx 0.086.

At t=50t = 50 (halfway):

η50=0.001+12(0.099)(1+cos⁡(0.5π))=0.001+12(0.099)(1+0)=0.001+0.0495≈0.0505\eta_{50} = 0.001 + \frac{1}{2}(0.099)(1 + \cos(0.5\pi)) = 0.001 + \frac{1}{2}(0.099)(1 + 0) = 0.001 + 0.0495 \approx 0.0505

At t=100t = 100 (end):

η100=0.001+12(0.099)(1+cos⁡(π))=0.001+12(0.099)(0)=0.001\eta_{100} = 0.001 + \frac{1}{2}(0.099)(1 + \cos(\pi)) = 0.001 + \frac{1}{2}(0.099)(0) = 0.001

Halfway through training, the learning rate is approximately half of the initial value. The curve is symmetric around the halfway point in a certain sense: the first quarter (epochs 0-25) sees a relatively small drop from 0.1 to 0.086, while the third quarter (epochs 50-75) sees a larger drop from 0.0505 to about 0.016. The decline accelerates in the middle of the schedule.

Warm Restarts

The standard cosine schedule ends at ηmin⁡\eta_{\min} and stays there. Loshchilov and Hutter proposed restarting the schedule after reaching the minimum, cycling back to ηmax⁡\eta_{\max} and annealing again. This is cosine annealing with warm restarts, sometimes written as SGDR or CosineAnnealingWarmRestarts.

The restart strategy has an appealing justification: near the end of each cycle, the low learning rate drives the model into a local minimum. The restart then kicks the model out of that minimum with a higher learning rate, potentially allowing it to find a flatter, more generalizable minimum in the next cycle. The intuition is that flat minima generalize better than sharp minima, and warm restarts encourage exploration of the loss landscape.

Restarts can also be paired with a multiplicative increase in the cycle length. Each successive cycle is made longer by a factor TmultT_{\text{mult}}, so the model spends more and more time at low learning rates as training progresses. A typical choice is Tmult=2T_{\text{mult}} = 2, doubling the cycle length each restart.

Cosine Annealing with Warm Restarts

A learning rate schedule that follows a cosine curve from maximum to minimum, then resets to the maximum and repeats. The restart allows the optimizer to escape sharp local minima and explore flatter regions of the loss landscape that tend to generalize better.

The warm restart also provides an elegant approach to model selection. At the end of each cosine cycle, the model has converged to a local minimum with a very small learning rate. This is a natural checkpoint: the model at the end of each cycle can be saved and compared. Since each restart nudges the model toward a different minimum, the ensemble of cycle-end checkpoints can sometimes outperform any single checkpoint. This practice, called snapshot ensembling, was introduced alongside SGDR as a way to get multiple trained models for the cost of one training run.

In[10]:
Code
def cosine_with_restarts(t, eta_max, eta_min, T0, T_mult=1):
    T_cur = T0
    t_remaining = t
    while t_remaining >= T_cur:
        t_remaining -= T_cur
        T_cur = int(T_cur * T_mult)
    return eta_min + 0.5 * (eta_max - eta_min) * (
        1 + np.cos(np.pi * t_remaining / T_cur)
    )


steps_restart = np.arange(0, 200)
cosine_fixed = [
    cosine_with_restarts(t, eta_max=0.1, eta_min=1e-4, T0=50, T_mult=1)
    for t in steps_restart
]
cosine_double = [
    cosine_with_restarts(t, eta_max=0.1, eta_min=1e-4, T0=25, T_mult=2)
    for t in steps_restart
]
Out[11]:
Visualization
Line plot of cosine annealing with fixed 50-epoch cycles on log scale over 200 epochs.
Cosine annealing with four 50-epoch cycles over 200 epochs. The schedule returns to the maximum learning rate after epochs 50, 100, and 150. Each restart gives the optimizer a chance to escape the current local minimum and explore the loss landscape.
Line plot of cosine annealing with doubling cycle lengths on log scale over 200 epochs.
Cosine annealing with doubling cycle lengths starting at 25 epochs. Restarts at epochs 25, 75, and 175 begin cycles of 50, 100, and 200 epochs, so the final 25 visible epochs show only the beginning of the longest cycle.

The doubling schedule is useful when you are uncertain about total training time: early cycles provide frequent opportunities to explore, while later cycles devote progressively longer intervals to annealing within a basin. In this 200-epoch window, the 200-epoch cycle has only just begun at epoch 175.

Decay Scheduling in Practice

The main schedules covered above can be augmented with additional techniques for practical training pipelines. This section covers the remaining standard variants, discusses how to compose multiple schedules, and addresses practical choices that come up repeatedly.

Linear Decay

Linear decay reduces the learning rate in equal steps across training:

ηt=η0⋅(1−tT)\eta_t = \eta_0 \cdot \left(1 - \frac{t}{T}\right)

where TT is the total number of training steps. Linear decay is easy to understand and widely used in language model fine-tuning. Many Hugging Face training scripts use linear warmup followed by linear decay as their default schedule.

The key property of linear decay is predictability: the learning rate at any step is a straightforward linear interpolation between the starting and ending values. This makes it easy to reason about training progress. If you are at 50% of total steps, the learning rate is exactly 50% of the initial value. If you extend or shorten training, you can immediately compute the new schedule without any complex parameter relationships.

Linear decay is also well-matched to fine-tuning pre-trained models. In fine-tuning, the goal is to adapt the model to a new task while preserving the general knowledge from pre-training. The learning rate needs to start small enough to not disrupt pre-trained weights too violently, and then decrease further to ensure precise convergence. A linear schedule that starts at a modest learning rate like 2×10−52 \times 10^{-5} and decreases to zero over, say, three epochs is exactly this kind of gentle, predictable refinement.

Polynomial Decay

Polynomial decay generalizes both linear and inverse square root decay:

ηt=(η0−ηend)⋅(1−tT)p+ηend\eta_t = \left(\eta_0 - \eta_{\text{end}}\right) \cdot \left(1 - \frac{t}{T}\right)^p + \eta_{\text{end}}

where pp is the polynomial power and ηend\eta_{\text{end}} is the final learning rate. Setting p=1p = 1 gives linear decay. Setting p=0.5p = 0.5 gives a schedule that drops quickly initially and then slows, resembling inverse square root. Setting p=2p = 2 gives a quadratic curve that drops slowly at first and quickly near the end, the opposite behavior.

The power pp is rarely tuned in practice because linear and cosine are simpler and well-understood. Polynomial decay appears in some large-scale training recipes where precise control over the decay shape is desired. The key insight is that the exponent pp controls the skew of the decay: smaller pp means more reduction happens early, larger pp means more reduction happens late.

The Floor Matters

Every schedule should have a non-zero minimum learning rate. Without a floor, exponential or inverse square root schedules approach zero asymptotically. At very small learning rates, training effectively stops: gradients are multiplied by a negligible factor and weights barely move. A floor of 10−510^{-5} or 10−610^{-6} ensures that training continues to make at least some progress.

The floor is especially important for schedules used in fine-tuning over many epochs. If you accidentally set γ\gamma too aggressively for an exponential schedule, the learning rate could reach the floor very early in training, wasting compute on steps where the parameters barely change. Monitor the learning rate reached during training as well as the schedule configuration.

Composing Warmup with Decay

Most modern training pipelines compose a warmup phase with one of the main decay schedules. The warmup brings the learning rate from zero to its peak value over the first few hundred or thousand steps, and then one of the decay schedules takes it from the peak down to the floor.

The transition from warmup to decay is handled differently depending on the framework. In some implementations, warmup and decay are separate schedule objects that are chained together. In PyTorch's get_cosine_schedule_with_warmup from the Hugging Face transformers library, the two phases are unified into a single function that changes behavior at the warmup boundary.

In[12]:
Code
def linear_warmup_then_decay(
    t, eta_max, warmup_steps, total_steps, decay="linear", eta_end=0.0
):
    if t < warmup_steps:
        return eta_max * t / warmup_steps
    progress = (t - warmup_steps) / max(total_steps - warmup_steps, 1)
    if decay == "linear":
        return eta_end + (eta_max - eta_end) * max(0.0, 1 - progress)
    elif decay == "cosine":
        return eta_end + 0.5 * (eta_max - eta_end) * (
            1 + np.cos(np.pi * progress)
        )
    return eta_max


total_steps = 1000
warmup = 100
steps_fine = np.arange(0, total_steps + 1)
linear_schedule = [
    linear_warmup_then_decay(
        t,
        eta_max=0.001,
        warmup_steps=warmup,
        total_steps=total_steps,
        decay="linear",
    )
    for t in steps_fine
]
cosine_schedule_fine = [
    linear_warmup_then_decay(
        t,
        eta_max=0.001,
        warmup_steps=warmup,
        total_steps=total_steps,
        decay="cosine",
    )
    for t in steps_fine
]
Out[13]:
Visualization
Two overlapping line plots showing learning rate for linear and cosine decay after warmup over 1000 steps.
Linear warmup combined with linear decay vs. cosine decay, both over 1,000 steps with 100 warmup steps. The shaded gray region shows the warmup phase where the learning rate rises to its peak. After warmup, cosine decay falls more gradually initially and more steeply near the end, while linear decay decreases at a constant rate.

The two schedules look similar during the main decay phase, but cosine decay is slightly more gradual at the start of decay and slightly steeper near the end. In practice, the difference between them is often smaller than the impact of the peak learning rate or warmup length.

Schedules and Adaptive Optimizers

Learning rate decay was designed primarily for SGD, where the learning rate directly controls step size. Adaptive optimizers like Adam complicate this picture considerably, and understanding the interaction is important for using schedules correctly.

How Adam Uses the Learning Rate

Adam maintains a per-parameter estimate of the learning rate through its second moment estimate vtv_t. The effective step size for parameter θi\theta_i is approximately η/vt+ϵ\eta / \sqrt{v_t + \epsilon}, where vtv_t tracks the historical squared gradients. Parameters with large, consistent gradients get small effective steps, and parameters with small or variable gradients get larger steps.

The global learning rate η\eta in Adam functions as an overall scaling factor on top of this per-parameter adaptation. When you apply a learning rate schedule on top of Adam, you are scaling this already-adapted step size. The schedule affects all parameters equally by a common multiplier, while Adam's internal adaptation continues to differentiate between them.

This combination generally works well in practice. The schedule handles the global progression from fast to slow learning, while Adam handles the per-parameter optimization. Most practitioners apply standard decay schedules to Adam without modification and find that this works reliably.

Why Schedules Still Matter for Adam

One might wonder whether schedules are necessary for Adam at all, given that Adam already adapts. The answer is yes, for several reasons.

First, Adam's adaptive mechanism operates on the ratio of first to second moments, not on the absolute magnitude of the learning rate. The global η\eta still controls the overall scale of updates, and a well-calibrated schedule ensures this scale shrinks appropriately as the model approaches convergence.

Second, Adam can get stuck oscillating around a minimum with a fixed learning rate for the same geometric reason as SGD. The curvature of the loss surface near the minimum sets a stability limit on the step size, and Adam's internal adaptation does not automatically reduce steps below this limit. A decaying schedule that brings η\eta below the stability threshold is what allows the model to settle.

Third, for fine-tuning, the schedule interacts with the model's pre-trained weights in an important way. Starting with a large Adam learning rate for fine-tuning can corrupt pre-trained representations, even with Adam's adaptation. A careful warmup followed by a decaying schedule protects the pre-trained weights by keeping updates small early on.

AdamW and Decoupled Decay

One subtlety arises with weight decay. Standard Adam with L2 regularization applies a decay term that is scaled by the adaptive learning rate, meaning parameters with large gradients receive less weight decay. This is mathematically inconsistent: weight decay is supposed to pull all parameters toward zero uniformly, but coupling it to the adaptive learning rate means it is applied differently to different parameters.

AdamW, introduced by Loshchilov and Hutter in 2019, fixes this by decoupling the weight decay from the gradient update: weight decay is applied at a fixed rate regardless of the adaptive step size. When using AdamW, the learning rate schedule affects gradient-based updates, and weight decay is controlled separately.

This decoupling matters for schedule design because it means the learning rate schedule in AdamW controls only the gradient step, not the regularization strength. You can set the learning rate schedule aggressively without worrying about undermining the weight decay regularization. The two can be tuned independently, which simplifies hyperparameter search.

Implementation

Let us implement all four schedules using PyTorch's scheduler API and verify they behave as expected.

In[14]:
Code
import torch.nn as nn
import torch.optim as optim


class TinyModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc = nn.Linear(10, 1)

    def forward(self, x):
        return self.fc(x)


model = TinyModel()

optimizer_step = optim.SGD(model.parameters(), lr=0.1)
scheduler_step = optim.lr_scheduler.StepLR(
    optimizer_step, step_size=30, gamma=0.1
)

optimizer_exp = optim.SGD(model.parameters(), lr=0.1)
scheduler_exp = optim.lr_scheduler.ExponentialLR(optimizer_exp, gamma=0.955)

optimizer_cos = optim.SGD(model.parameters(), lr=0.1)
scheduler_cos = optim.lr_scheduler.CosineAnnealingLR(
    optimizer_cos, T_max=100, eta_min=1e-4
)

optimizer_cosr = optim.SGD(model.parameters(), lr=0.1)
scheduler_cosr = optim.lr_scheduler.CosineAnnealingWarmRestarts(
    optimizer_cosr, T_0=50, T_mult=1
)
In[15]:
Code
def record_lr_schedule(optimizer, scheduler, n_steps=100):
    lrs = []
    for _ in range(n_steps):
        lrs.append(optimizer.param_groups[0]["lr"])
        scheduler.step()
    return lrs


torch_step_lrs = record_lr_schedule(optimizer_step, scheduler_step)
torch_exp_lrs = record_lr_schedule(optimizer_exp, scheduler_exp)
torch_cos_lrs = record_lr_schedule(optimizer_cos, scheduler_cos)
torch_cosr_lrs = record_lr_schedule(optimizer_cosr, scheduler_cosr)
Out[16]:
Console
Step decay LRs at epochs [0, 30, 60, 90]:
  Epoch   0: 0.100000
  Epoch  30: 0.010000
  Epoch  60: 0.001000
  Epoch  90: 0.000100

Cosine LRs at epochs [0, 25, 50, 75, 99]:
  Epoch   0: 0.100000
  Epoch  25: 0.085370
  Epoch  50: 0.050050
  Epoch  75: 0.014730
  Epoch  99: 0.000125

The step decay drops exactly at epoch 30 and 60 as configured. The cosine decay smoothly interpolates from 0.1 to near 0.0001 over 100 epochs.

For the inverse square root schedule with warmup, PyTorch does not provide a built-in, but it is straightforward to implement using LambdaLR:

In[17]:
Code
def make_inv_sqrt_lambda(warmup_steps):
    def lr_lambda(t):
        if t < warmup_steps:
            return t / warmup_steps
        return (warmup_steps**0.5) / (t**0.5)

    return lr_lambda


model_inv = TinyModel()
optimizer_inv = optim.AdamW(model_inv.parameters(), lr=0.001)
scheduler_inv = optim.lr_scheduler.LambdaLR(
    optimizer_inv, lr_lambda=make_inv_sqrt_lambda(warmup_steps=100)
)

torch_inv_lrs = record_lr_schedule(optimizer_inv, scheduler_inv, n_steps=500)
Out[18]:
Console
Inverse sqrt LRs at steps [0, 50, 100, 200, 499]:
  Step    0: 0.000000
  Step   50: 0.000500
  Step  100: 0.001000
  Step  200: 0.000707
  Step  499: 0.000448

The learning rate climbs linearly during warmup (steps 0 to 100), peaks at step 100, and then decays as the inverse square root of the step count.

Key Parameters

The key parameters to configure for each schedule are:

  • Step decay: step_size (epochs between drops), gamma (multiplicative factor, typically 0.1 or 0.5)
  • Exponential decay: gamma (per-step multiplier, typically 0.99 to 0.999 for step-wise or 0.1 to 0.5 per epoch)
  • Inverse square root: warmup_steps (typically 1 to 10 percent of total steps), eta_peak (the peak learning rate)
  • Cosine annealing: T_max (period length), eta_min (minimum value, often 10−510^{-5})
  • Warm restarts: T_0 (initial period), T_mult (cycle growth factor, 1 for fixed cycles, 2 for doubling)

Using Hugging Face Schedulers

The Hugging Face transformers library provides convenient schedule constructors that are particularly useful for fine-tuning language models:

In[19]:
Code
from transformers import (
    get_cosine_schedule_with_warmup,
    get_linear_schedule_with_warmup,
)

model_hf = TinyModel()
optimizer_hf = optim.AdamW(model_hf.parameters(), lr=2e-5)

hf_warmup_steps = 100
hf_total_steps = 1000

scheduler_hf_linear = get_linear_schedule_with_warmup(
    optimizer_hf,
    num_warmup_steps=hf_warmup_steps,
    num_training_steps=hf_total_steps,
)

model_hf2 = TinyModel()
optimizer_hf2 = optim.AdamW(model_hf2.parameters(), lr=2e-5)
scheduler_hf_cosine = get_cosine_schedule_with_warmup(
    optimizer_hf2,
    num_warmup_steps=hf_warmup_steps,
    num_training_steps=hf_total_steps,
)

hf_linear_lrs = record_lr_schedule(
    optimizer_hf, scheduler_hf_linear, n_steps=hf_total_steps
)
hf_cosine_lrs = record_lr_schedule(
    optimizer_hf2, scheduler_hf_cosine, n_steps=hf_total_steps
)
Out[20]:
Console
HuggingFace linear schedule at key steps:
  Step    0: 0.00000000
  Step   50: 0.00001000
  Step  100: 0.00002000
  Step  500: 0.00001111
  Step  999: 0.00000002

These functions handle the warmup transition automatically. The num_warmup_steps parameter sets the warmup duration, and num_training_steps sets the total. The schedule is then applied via scheduler.step() after each optimizer step, exactly like any PyTorch scheduler.

Out[21]:
Visualization
Line plot of HuggingFace linear and cosine warmup schedules over 1000 training steps.
Hugging Face linear and cosine schedules with warmup over 1,000 steps. Both schedules use 100 warmup steps to reach the peak learning rate of 2e-5. After warmup, linear decay falls at a constant rate while cosine decay follows the characteristic smooth curve, ending more gradually than linear.

Worked Example: Effect on Convergence

To see concretely how schedule choice affects training, let us train a small neural network on a synthetic regression task using four different schedules and compare convergence.

In[22]:
Code
torch.manual_seed(42)
np.random.seed(42)

n_samples = 400
X_data = torch.randn(n_samples, 10)
true_weights = torch.randn(10)
y_data = X_data @ true_weights + 0.1 * torch.randn(n_samples)

dataset = torch.utils.data.TensorDataset(X_data, y_data)
dataloader = torch.utils.data.DataLoader(dataset, batch_size=32, shuffle=True)
In[23]:
Code
def train_with_schedule(scheduler_type, n_epochs=60, lr=0.05):
    net = nn.Linear(10, 1)
    optimizer = optim.SGD(net.parameters(), lr=lr, momentum=0.9)
    criterion = nn.MSELoss()

    if scheduler_type == "step":
        scheduler = optim.lr_scheduler.StepLR(
            optimizer, step_size=20, gamma=0.1
        )
    elif scheduler_type == "exponential":
        gamma_e = (1e-4 / lr) ** (1 / n_epochs)
        scheduler = optim.lr_scheduler.ExponentialLR(optimizer, gamma=gamma_e)
    elif scheduler_type == "cosine":
        scheduler = optim.lr_scheduler.CosineAnnealingLR(
            optimizer, T_max=n_epochs, eta_min=1e-5
        )
    else:
        scheduler = None

    losses = []
    for epoch in range(n_epochs):
        epoch_loss = 0.0
        for X_batch, y_batch in dataloader:
            optimizer.zero_grad()
            pred = net(X_batch).squeeze()
            loss = criterion(pred, y_batch)
            loss.backward()
            optimizer.step()
            epoch_loss += loss.item()
        losses.append(epoch_loss / len(dataloader))
        if scheduler is not None:
            scheduler.step()

    return losses


n_epochs = 60
losses_step = train_with_schedule("step", n_epochs=n_epochs)
losses_exp = train_with_schedule("exponential", n_epochs=n_epochs)
losses_cosine = train_with_schedule("cosine", n_epochs=n_epochs)
losses_constant = train_with_schedule("constant", n_epochs=n_epochs)
Out[24]:
Console
Final training loss (epoch 60):
  Step decay:       0.00985
  Exponential:      0.01002
  Cosine annealing: 0.00997
  Constant LR:      0.01206
Out[25]:
Visualization
Line plot comparing training loss over 60 epochs for step, exponential, cosine, and constant learning rate schedules.
Training loss for a linear regression model under four learning rate schedules over 60 epochs, shown on a logarithmic scale. All three decay schedules settle below the constant rate by the end. The log scale preserves the rapid initial descent while making the small late-stage differences and the constant rate's greater variability visible.

The constant learning rate model shows the highest final loss because its updates remain large enough to keep it moving around the minimum. All three decay schedules finish lower in this run. The logarithmic axis makes those late-stage differences visible without hiding the much larger loss reduction during the first few epochs.

Interpreting the Convergence Curves

The convergence plot reveals more than just final loss values. Notice how each schedule's characteristic shape shows up in the training loss.

For step decay, scheduled cuts at epochs 20 and 40 reduce the size of subsequent updates. In this noisy mini-batch run, those transitions are subtler than the learning-rate staircase itself; the clearest effect is the tighter band of losses after the first cut. A large loss drop at a schedule boundary can indicate that the previous rate was causing oscillation, but it is not guaranteed in every run.

For exponential decay, the loss falls smoothly throughout. There are no phase transitions visible in the loss curve, which is a double-edged sword: smooth training is easy to monitor, but you lose the visual feedback that phase boundaries provide. If exponential decay is decaying too fast, you will see the loss plateau early, while the learning rate is still non-negligible, with no obvious cause.

For cosine annealing, the loss curves similarly smoothly. The cosine schedule's gradual start means the learning rate at the very beginning is nearly identical to the initial value, so there is no initial benefit from decay. The benefit accumulates through the middle and end of training.

Comparing Schedules at a Glance

The table below summarizes the key properties of each schedule to guide selection.

Learning rate schedule comparison.
ScheduleShapeBest forLimitations
Step decayStaircaseCV training, interpretabilityRigid timing, requires epoch tuning
ExponentialSmooth monotoneShort runs with a floorDecays too fast without tuning
Inverse sqrtFast drop then plateauTransformer pre-trainingNeeds warmup configuration
CosineSmooth S-curveFine-tuning, general useRequires total steps known upfront
Cosine + restartsCyclicExploratory trainingCycle length sensitive

Choosing a Schedule

No single schedule is universally best. The right choice depends on the model architecture, optimizer, dataset size, and training budget. What follows is a practical guide to making this choice, organized by use case.

Training Large Transformers from Scratch

For training large transformer models from scratch, the inverse square root schedule with linear warmup is the canonical default. It was validated on the original Transformer and remains widely used for pre-training. The slow later decay means the model continues to see meaningful learning rate values even late in training.

The warmup length for large models is typically 4,000 to 10,000 steps. The peak learning rate is usually in the range of 10−410^{-4} to 5×10−45 \times 10^{-4} for model dimensions around 512-1024. For larger models, the peak is often scaled inversely with the square root of the model dimension, following the Transformer paper's formula.

Some recent large language model training runs have moved away from the inverse square root schedule in favor of cosine annealing with a fixed total step count. The argument is that cosine annealing with a known endpoint is more predictable and allows planning the training budget in advance. The inverse square root schedule is "endless" in the sense that there is no natural stopping point, which complicates budget planning.

Fine-Tuning Pre-Trained Models

For fine-tuning pre-trained models, linear or cosine decay is preferred because fine-tuning involves much shorter training runs. The learning rate needs to reach a very small value by the end of training to avoid disrupting the pre-trained weights. Cosine decay provides a smooth, predictable endpoint.

The peak learning rate for fine-tuning is typically much smaller than for pre-training. For BERT-style models, common values are 1×10−51 \times 10^{-5} to 5×10−55 \times 10^{-5}. For larger models like GPT-style architectures, the fine-tuning learning rate is often even smaller, sometimes as low as 10−610^{-6}, to prevent catastrophic forgetting of pre-trained knowledge.

The warmup period for fine-tuning is shorter than for pre-training, typically 5 to 10 percent of total steps. Since the model's weights are already in a reasonable part of parameter space (thanks to pre-training), the risk of early instability is lower. The warmup mainly serves to let Adam's moment estimates calibrate.

Computer Vision with SGD

For computer vision with SGD, step decay remains common. The explicit phase structure matches the training recipes for ResNet and related architectures, and the abrupt drops are known to produce good convergence behavior.

The standard three-phase recipe for ImageNet training is: 90 epochs total, starting learning rate of 0.1, drops of 10x at epochs 30 and 60. This recipe has been so successful and widely validated that it has become a community reference point. Many papers compare their training protocols against this baseline.

When you switch from SGD to AdamW for computer vision, the step decay recipe often needs adjustment. Adam's internal adaptation changes the effective sensitivity to the global schedule, and the ten-fold drops that work well for SGD can be too aggressive for Adam. A cosine schedule with AdamW tends to transfer more reliably across different model scales and training durations.

When in Doubt

When in doubt between linear and cosine decay, cosine decay is slightly preferred because it is slower to start falling (reducing the learning rate less aggressively early in decay) and ends more smoothly. In practice, the difference is small compared to the impact of other hyperparameters. If you are setting up a new experiment and have no prior experience with the task, cosine annealing with linear warmup using 5 to 10 percent of total steps is a reasonable starting point.

The most important hyperparameter to tune is the peak learning rate, not the schedule shape. A cosine schedule with a poorly tuned peak learning rate will underperform a linear schedule with the right peak learning rate. Spend your tuning budget on the peak learning rate first, and only refine the schedule shape once that is set.

Limitations and Practical Considerations

Learning rate scheduling is powerful but introduces additional hyperparameters that must be set correctly. A poorly chosen schedule can be worse than no schedule at all. If the learning rate decays too quickly, the model gets stuck in a suboptimal region before it has had time to find a good basin. If it decays too slowly, the final convergence is poor and the training may appear to have plateaued.

Schedule Sensitivity and Hyperparameter Overhead

Each schedule introduces at least one and often several additional hyperparameters beyond the base learning rate. Step decay needs both the step size and the gamma. Cosine annealing needs the minimum learning rate and the total period length. The inverse square root schedule needs the warmup duration. All of these interact with the peak learning rate in non-trivial ways.

In practice, these hyperparameters are often set based on rules of thumb derived from prior experiments rather than being tuned from scratch. This is a practical necessity: running a full hyperparameter sweep over schedule parameters for every new task is computationally infeasible. The standard recipes that have emerged from community experience (like linear warmup to 10% of steps followed by cosine decay) are valuable precisely because they generalize across many settings without requiring explicit tuning.

The Interaction with Adaptive Optimizers

The interaction between schedules and adaptive optimizers like Adam is not fully understood. Adam's internal adaptation already provides a form of per-parameter learning rate control, and layering a global schedule on top creates a complex combined behavior. Some researchers have argued that with properly tuned Adam, the sensitivity to the global learning rate schedule is lower than with SGD, though in practice schedules still affect performance.

One specific concern is that Adam's second moment estimate vtv_t accumulates history according to its decay factor β2\beta_2. When the learning rate is decayed and then kept constant for a long time (as in step decay), the effective step size changes both because of the explicit decay and because vtv_t continues to evolve. The combined effect is not easy to predict analytically. This is part of why cosine and linear schedules, which change the learning rate smoothly and predictably, tend to behave more reliably with Adam than step decay does.

Warmup in Distributed Training

Warmup deserves particular attention in distributed training. When training with large batch sizes across many GPUs, the effective batch size is the product of the per-GPU batch size and the number of GPUs. The linear scaling rule suggests scaling the learning rate proportionally to the batch size, which can push the learning rate very high. Warmup becomes critical in this regime: starting at the full scaled learning rate with a large batch and random initialization can cause catastrophic early instability. Longer warmup periods are used for very large distributed training runs.

The challenge is that the standard warmup durations (4,000 to 10,000 steps) were designed for moderate batch sizes. When batch size is scaled by a factor of 64x for distributed training, the effective number of samples seen per step increases dramatically, and the notion of "a step" in the warmup duration becomes less meaningful. Some large-scale training recipes specify warmup in terms of tokens or samples processed rather than steps, which remains invariant to batch size scaling.

Cycle-Based Schedules in Practice

Cycle-based schedules like cosine annealing with warm restarts have an appealing theoretical motivation but require careful setting of the cycle length. A cycle that is too short means the model restarts before it has converged within each basin, wasting each cycle. A cycle that is too long is effectively just cosine annealing without restarts. In practice, the cycle length is often tuned as a hyperparameter or set based on prior experience with similar models.

The restart mechanism can also cause difficulties for checkpoint selection. With a standard monotone schedule, the best checkpoint is usually the final one. With warm restarts, the best checkpoint might be at the end of any cycle. You need to track cycle-end checkpoints explicitly and compare them, rather than simply saving the final model. This adds infrastructure complexity to training pipelines.

Generalization and the Learning Rate at Convergence

The relationship between schedules and generalization is subtle. It is well established empirically that decaying the learning rate to a small value at the end of training improves generalization, not just training loss. The intuition is that a small learning rate at convergence means the model has found a region where the loss surface is locally flat, and flat minima tend to generalize better. But this connection is not a tight theoretical guarantee, and counter-examples exist.

More precisely, the "flat minima generalize better" hypothesis has been the subject of ongoing debate. Some theoretical work argues that the measure of flatness depends on the parameterization and can be an artifact of scale. Others have shown empirically that the correlation between flatness and generalization holds across a wide range of architectures and tasks. In practice, ending training at a small learning rate reliably improves results, even if the theoretical mechanism is not fully settled.

Summary

Learning rate decay is a fundamental component of neural network training that bridges the gap between fast early learning and precise final convergence. The core problem is that a fixed learning rate cannot simultaneously provide the large steps needed for rapid progress early in training and the small steps needed for stable convergence later. Schedules resolve this by continuously adjusting the learning rate over time.

The key schedules are:

  • Step decay: multiplies the learning rate by a factor at fixed epoch intervals, producing a staircase. Easy to interpret and historically common in computer vision. The sharp drops provide clear phase boundaries in training.
  • Exponential decay: applies a constant multiplicative factor at every step, producing a smooth exponential curve. Simple but can decay too aggressively without a floor or careful gamma selection.
  • Inverse square root decay: falls quickly at first and then levels off, making it well-matched to optimization dynamics. The standard choice for training transformers from scratch, especially combined with linear warmup.
  • Cosine annealing: follows a smooth half-cosine curve from maximum to minimum. Widely used for fine-tuning and naturally extended with warm restarts for exploratory training where escaping local minima matters.

All schedules benefit from linear warmup, which prevents instability caused by noisy gradients and uncalibrated optimizer statistics at the start of training. The warmup length is typically 1 to 10 percent of total steps, with longer warmup for larger models and batch sizes.

The choice of schedule interacts with the optimizer. AdamW with cosine annealing and linear warmup has become the default for most NLP fine-tuning work. Inverse square root with warmup remains the standard for large-scale pre-training. SGD with step decay persists in computer vision training recipes that have been validated at scale.

In practice, the peak learning rate is more important to tune than the schedule shape. A well-tuned peak learning rate with any reasonable schedule will outperform a poorly tuned peak with the optimal schedule. Start by finding a good peak learning rate, then refine the schedule if further improvement is needed. The differences between well-tuned schedules are often small compared to the impact of the peak learning rate itself.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about learning rate decay and scheduling.

Learning Rate Decay Quiz

Question 1 of 70 of 7 completed
Why does a fixed learning rate cause problems late in training?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026learningrate, author = {Michael Brenndoerfer}, title = {Learning Rate Decay: Step, Exponential}, year = {2026}, url = {https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-06} }
APAAcademic
Michael Brenndoerfer (2026). Learning Rate Decay: Step, Exponential. Retrieved from https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling
MLAAcademic
Michael Brenndoerfer. "Learning Rate Decay: Step, Exponential." 2026. Web. October 6, 2026. <https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling>.
CHICAGOAcademic
Michael Brenndoerfer. "Learning Rate Decay: Step, Exponential." Accessed October 6, 2026. https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Learning Rate Decay: Step, Exponential'. Available at: https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling (Accessed: October 6, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Learning Rate Decay: Step, Exponential. https://mbrenndoerfer.com/writing/learning-rate-decay-step-exponential-inverse-sqrt-scheduling

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.