Cosine Learning Rate Schedule: Decay, Restarts, and Warmup

Michael BrenndoerferJanuary 26, 202665 min read

Part of Language AI Handbook

Covers cosine learning rate schedule used in GPT and LLaMA training: the decay formula, warm restarts, key parameters, and comparison with linear decay.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Cosine Learning Rate Schedule

The learning rate is one of the most consequential hyperparameters in training neural networks. Set it too high and the model diverges; set it too low and it crawls toward a poor solution. As we saw in the Learning Rate Warmup chapter, starting with a small learning rate and gradually ramping it up helps the optimizer find a stable initial direction. And as covered in Learning Rate Decay, reducing the learning rate over time lets the model settle into the fine-grained structure of the loss landscape. The cosine learning rate schedule combines both ideas into a single, smooth function that has become the dominant choice for training large language models and vision transformers alike.

The central intuition is elegant. Instead of dropping the learning rate in abrupt steps or along a straight line from start to finish, the cosine schedule follows the shape of a cosine curve: it starts high, descends gradually at first, accelerates through the middle of training, and then slows again to arrive gently near zero. This smooth trajectory avoids the shock of sudden step reductions while also being faster than a linear decay in the early and late phases where the model benefits most from larger and then smaller steps respectively.

Why does the cosine curve, of all curves, suit this purpose so well? The answer comes down to the geometry of the loss landscape and the characteristic phases of deep learning optimization. Early in training, the model is far from any meaningful minimum, and a learning rate that stays moderately high lets the optimizer take bigger exploratory steps. In the middle phase, the model is making substantial parameter adjustments as it locates the broad basin of a good solution. Near the end, the model is refining its parameters within that basin, and a very small learning rate lets it settle precisely without oscillating around the minimum. The cosine curve allocates the most aggressive decay to this middle phase and decelerates gently at both ends, which closely mirrors what practitioners discovered empirically through years of trial and error with step decay schedules.

This chapter covers the cosine decay formula in full detail, explains when and why to use cosine annealing with warm restarts, walks through the role each schedule parameter plays, and compares cosine decay against linear and step decay alternatives. We will implement these schedules from scratch and then show how to integrate them with PyTorch's built-in scheduler API. We will also examine how this schedule appears in the published training configurations of GPT, LLaMA, Mistral, and other prominent language models.

Historical Context and Origin

The cosine schedule did not appear from nowhere. It has a clear intellectual lineage rooted in optimization theory and the physics of annealing.

The broader concept of annealing in optimization descends from the metallurgical process of slowly cooling a material to reduce crystalline defects. In physical annealing, a metal is heated until atoms can move freely, then cooled gradually so they have time to settle into a low-energy, ordered arrangement. Rapid cooling traps defects; gradual cooling allows the system to explore its energy landscape and find a more globally stable configuration. Simulated annealing, introduced by Kirkpatrick, Gelatt, and Vecchi in 1983, imported this intuition into numerical optimization: explore the solution space while "temperature" is high, then exploit as temperature cools. The learning rate in gradient descent plays the role of temperature, and its gradual reduction mirrors the cooling schedule of simulated annealing.

Early neural network training used fixed learning rates or manually designed step-drop schedules because practitioners did not yet have the compute resources to run enough experiments to compare schedule shapes systematically. The AlexNet paper from 2012 used a step decay schedule that reduced the learning rate by a factor of 10 whenever the validation error stopped improving, with three manual reductions over the course of training. This "reduce on plateau" approach worked well for relatively short training runs where a practitioner could monitor validation curves and decide when to drop the rate, but it did not scale to the longer, unsupervised pretraining runs that would come later.

The cosine annealing schedule was formalized in a neural network context by Loshchilov and Hutter in their 2016 ICLR paper "SGDR: Stochastic Gradient Descent with Warm Restarts." That paper's primary contribution was the warm restart mechanism, but it also established the cosine annealing formula itself as a principled choice for the decay within each cycle. The smooth interpolation using the cosine function offered a natural improvement over step decay without the manual threshold-setting. The paper demonstrated that cosine annealing with restarts improved performance on CIFAR-10 and CIFAR-100 relative to both step schedules and simple cosine without restarts.

Large language model training adopted the cosine schedule without restarts, dropping the restart component but retaining the smooth cosine decay. GPT-2 (2019) used a warmup-plus-cosine design. GPT-3 (2020) formalized it with explicit parameters. By the time LLaMA appeared in 2023, the warmup-cosine recipe was so standard that its use needed almost no justification in the paper. What began as an ablation result from an optimization paper became the de facto default for training language models at scale.

Understanding this history matters because it reveals that the cosine schedule's dominance is partly empirical and partly path-dependent. It has been tested and validated at every scale from millions to hundreds of billions of parameters. Its smooth shape avoids the worst failure modes of step decay. Its simplicity makes it easy to reason about and reproduce. None of this means that alternative schedules could not perform as well or better, and active research continues to explore this space. But it does mean that when you train a language model with a cosine schedule, you are following a well-trodden path with a large body of existing knowledge to guide your parameter choices.

The Loss Landscape and Why Decay Shape Matters

To develop real intuition for why the cosine schedule works, it helps to think carefully about what the loss landscape looks like during different phases of training and what kind of learning rate behavior each phase calls for.

The Phases of Optimization

Deep neural network loss surfaces are high-dimensional and highly non-convex. They contain many local minima, saddle points, and flat regions where gradients are small. Empirical research over the past decade has produced a rough picture of how gradient descent moves through this landscape during training.

In the earliest phase, immediately after random initialization, the model's parameters are far from any interesting region of the loss surface. The loss is high, and gradients point roughly in directions that reduce it. At this point, the exact step size matters less than staying stable and making clear directional progress. This is why warmup is so valuable: large learning rates on random initializations cause the optimizer to take wild, unstable steps that can overshoot into regions of catastrophically high loss. A small initial rate lets the optimizer establish a sensible direction before committing to larger steps.

Once the model is making meaningful progress, it enters a phase of broad exploration. The optimizer is traversing large regions of parameter space, and the loss is decreasing substantially with each step. The learning rate should be high enough to cover ground quickly but not so high that the optimizer overshoots the broad basin it is trying to find. This phase benefits from a learning rate that stays near its peak and only decreases slowly, which is exactly what the cosine schedule provides in the first quarter of training.

As training continues, the model begins to find the broad basin of a good solution. The loss is now decreasing more slowly, and gradients are pointing more consistently downward rather than in a noisy, exploratory fashion. The optimizer needs to make progressively finer adjustments. A learning rate that decreases at an accelerating pace during this phase, matching the cosine curve's steepest descent at the midpoint, aligns well with this transition from exploration to exploitation.

In the final phase, the model is already near a local minimum and needs to settle in. The gradient directions are consistent, but the useful step sizes are tiny. A very small learning rate lets the optimizer make these final micro-adjustments without overshooting the minimum. The cosine schedule's gentle approach to ηmin⁡\eta_{\min} in the final quarter of training mirrors this behavior perfectly.

Sharp vs. Flat Minima

The geometry of the minimum matters for generalization. A sharp minimum is one where the loss surface has steep walls on all sides, meaning the loss increases quickly as you move away from the optimal point. A flat minimum is one where the loss surface is wide and shallow, and the loss increases only slowly in all directions.

Sharp minima are problematic for generalization. When a model trained on one dataset is deployed on a slightly different distribution (test data, slightly different user inputs), its parameters are never exactly at the training minimum. A flat minimum tolerates this imprecision gracefully because the loss changes slowly. A sharp minimum punishes it severely because even a small perturbation takes the model to a point where loss is much higher.

This observation, developed rigorously by Hochreiter and Schmidhuber in the 1990s and revisited extensively in the modern deep learning era, has direct implications for learning rate scheduling. Larger learning rates tend to escape sharp minima because the optimizer takes steps large enough to hop over the steep walls. Smaller learning rates can get trapped in them because each step is too small to escape. The cosine schedule's slow early decay keeps the learning rate in a regime that favors flat minima for longer, before gradually reducing to the fine-grained regime that allows precise convergence.

This is also why warm restarts improve generalization: the sudden increase in learning rate at each restart gives the optimizer a chance to escape any sharp minimum it has found and continue searching for flatter, wider basins before the next round of slow convergence.

The Role of Gradient Noise

In mini-batch stochastic gradient descent, gradients are noisy estimates of the true gradient computed from a small sample of the training data. The noise level in gradients interacts with the learning rate in an important way: a large learning rate magnifies gradient noise, potentially causing the optimizer to take steps that are primarily noise-driven rather than signal-driven. A small learning rate attenuates noise but also attenuates the useful signal.

This interaction creates a natural argument for learning rate decay. Early in training, with the model far from any minimum, signal-to-noise ratios are generally high: gradient directions are consistent across mini-batches because they all point away from the same high-loss initialization. As the model approaches a minimum, gradients become smaller in magnitude and more noisy relative to their mean direction, because there is less clear "which way is down" once you are in a local basin. The optimal learning rate should therefore be smaller when the model is close to a minimum than when it is far away, which is exactly what decay schedules provide.

The cosine schedule does not adapt to gradient noise directly; it decays as a function of training steps regardless of what the gradients are doing. But because training steps and proximity to a minimum are loosely correlated (more steps means more progress toward a minimum), the step-based decay works reasonably well in practice. Adaptive learning rate methods like AdamW already handle per-parameter gradient noise through their moment estimates, but even they benefit from a global learning rate schedule that provides a coarse adjustment to the overall step magnitude.

The Cosine Decay Formula

The cosine schedule derives its shape from the cosine function, which decreases smoothly from 1 to −1-1 over a half-period. By remapping this interval to span from the initial learning rate down to a minimum value, we get a decay profile that is gentle near the boundaries and steeper in the middle.

Cosine Annealing

Cosine annealing refers to any learning rate schedule that uses the cosine function to interpolate between a maximum and minimum value. The term "annealing" comes from the metallurgical process of slowly cooling a material to reduce defects, analogizing the cooling of the learning rate to stabilize model parameters.

Basic Cosine Decay

We want a function that starts at ηmax⁡\eta_{\max} when t=0t = 0 and ends at ηmin⁡\eta_{\min} when t=Tt = T, following a smooth cosine curve in between. The standard cosine decay formula achieves this as follows:

η(t)=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡(πtT))\eta(t) = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(1 + \cos\left(\frac{\pi t}{T}\right)\right)

where:

  • η(t)\eta(t): the learning rate at step tt
  • ηmax⁡\eta_{\max}: the peak learning rate at the start of the schedule (after any warmup period)
  • ηmin⁡\eta_{\min}: the floor learning rate, often set to 0 or a small fraction of ηmax⁡\eta_{\max} such as 10−510^{-5}
  • tt: the current step index, ranging from 0 to TT
  • TT: the total number of decay steps
  • cos⁡(πt/T)\cos(\pi t / T): the cosine term that drives the smooth interpolation from 1 (at t=0t=0) to −1-1 (at t=Tt=T)

The expression 12(1+cos⁡(⋅))\frac{1}{2}(1 + \cos(\cdot)) is sometimes called the "raised cosine" because it shifts and scales the cosine from its natural range [−1,1][-1, 1] up into the range [0,1][0, 1]. Multiplying by (ηmax⁡−ηmin⁡)(\eta_{\max} - \eta_{\min}) stretches this unit interval to the desired learning rate range, and adding ηmin⁡\eta_{\min} lifts the floor.

To see why this particular construction works, think of it as a two-step transformation. The raw cosine cos⁡(πt/T)\cos(\pi t / T) takes values in [−1,1][-1, 1]: it equals +1+1 at t=0t = 0, passes through 00 at t=T/2t = T/2, and reaches −1-1 at t=Tt = T. We want a multiplier that goes from 1 to 0 (not from 1 to −1-1), so we add 1 and divide by 2 to get 1+cos⁡(πt/T)2\frac{1 + \cos(\pi t / T)}{2}. This quantity ranges from 1 down to 0. Multiplying by (ηmax⁡−ηmin⁡)(\eta_{\max} - \eta_{\min}) scales that unit range to match the desired learning rate range, and adding ηmin⁡\eta_{\min} ensures we never go below the floor.

Verification at boundary values. Let us confirm the formula produces the right values at t=0t = 0 and t=Tt = T:

At t=0t = 0:

η(0)=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡(0))=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+1)=ηmin⁡+(ηmax⁡−ηmin⁡)=ηmax⁡\begin{aligned} \eta(0) &= \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(1 + \cos(0)\right) \\ &= \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})(1 + 1) \\ &= \eta_{\min} + (\eta_{\max} - \eta_{\min}) \\ &= \eta_{\max} \end{aligned}

At t=Tt = T:

η(T)=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡(π))=ηmin⁡+12(ηmax⁡−ηmin⁡)(1−1)=ηmin⁡+0=ηmin⁡\begin{aligned} \eta(T) &= \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(1 + \cos(\pi)\right) \\ &= \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})(1 - 1) \\ &= \eta_{\min} + 0 \\ &= \eta_{\min} \end{aligned}

The schedule begins at ηmax⁡\eta_{\max} and ends exactly at ηmin⁡\eta_{\min}, as intended. At the midpoint t=T/2t = T/2, the cosine evaluates to cos⁡(π/2)=0\cos(\pi/2) = 0, giving a learning rate of exactly the average (ηmax⁡+ηmin⁡)/2(\eta_{\max} + \eta_{\min}) / 2.

The Decay Rate at Each Phase

One of the most instructive ways to understand the cosine schedule is to look at the derivative of η(t)\eta(t) with respect to tt. The derivative tells us how fast the learning rate is falling at each step:

dηdt=−π2T(ηmax⁡−ηmin⁡)sin⁡(πtT)\frac{d\eta}{dt} = -\frac{\pi}{2T}(\eta_{\max} - \eta_{\min})\sin\left(\frac{\pi t}{T}\right)

The sine function starts at zero, rises to its maximum at t=T/2t = T/2, and falls back to zero at t=Tt = T. This means the learning rate is decreasing most rapidly at the midpoint of training, and barely changing at all at the very beginning and very end. Compare this with linear decay, where the derivative is a constant −(ηmax⁡−ηmin⁡)/T-(\eta_{\max} - \eta_{\min}) / T at every step: the rate of decrease is the same regardless of how far along training has progressed.

This derivative structure is exactly what we want. At the start of training, gradients are large and noisy, and a slowly-changing learning rate gives the optimizer time to orient itself without the additional disruption of a rapidly shifting step size. Near the end of training, the learning rate itself is already very small, and a nearly flat derivative means the optimizer is making barely-perceptible adjustments to the rate, allowing for very fine-grained convergence. The middle phase, where the learning rate is most actively shrinking, corresponds to the period when the optimizer is making the transition from broad exploration to local refinement.

Why Smooth Decay Matters

Step decay schedules drop the learning rate by a fixed factor at discrete checkpoints. This creates a jarring discontinuity each time the rate drops: the optimizer is cruising along at one scale and suddenly the effective step size contracts by an order of magnitude. In practice, these drops often coincide with sudden sharp decreases in training loss followed by a plateau at the new rate. While this works, it makes scheduling sensitive to the choice of when to drop.

After each step-decay drop, the loss tends to spike briefly as the optimizer adjusts to the new scale, then quickly drops to a lower floor than before. This "shock and reset" behavior works, but it suggests the model is doing a kind of forced convergence at each checkpoint rather than settling smoothly. The timing of drops has to be chosen carefully: too early and the model has not finished benefiting from the higher rate; too late and it has been oscillating around a minimum without making progress.

Linear decay, covered in the Learning Rate Decay chapter, avoids this problem with a straight-line decrease. However, linear decay is fastest right at the start, when the model may still be in a phase where slightly larger steps would help exploration, and it decreases at a uniform rate that ignores the characteristic phases of learning.

Cosine decay allocates decay budget differently. The rate decreases slowly at first, giving the model time to make larger initial progress. It decelerates more rapidly through the middle of training when parameters are converging broadly. Near the end, the rate again slows its descent, allowing the model to settle delicately into a local minimum without overshooting. This pacing matches the natural phases of gradient descent quite well.

There is also a practical advantage: cosine decay requires no manual checkpoint selection. With step decay, you choose specific milestones, and if they are wrong for your architecture, problem, or batch size, performance suffers. The cosine schedule adapts its budget continuously and automatically based solely on the current step fraction t/Tt/T.

Out[3]:
Visualization
Line chart comparing cosine, linear, and step decay learning rate schedules over 1000 steps.
Comparison of cosine, linear, and step decay learning rate schedules over 1,000 training steps. Cosine decay stays higher than linear in the first half of training, supporting broader exploration, and drops below linear in the second half, enabling finer convergence. Step decay creates three discrete drops at fixed checkpoints, visible as vertical discontinuities. All schedules start at the same peak and finish at the same floor.

The Raised Cosine: A Worked Numerical Example

To make the formula concrete, let us trace through the exact values at several checkpoints for a schedule with ηmax⁡=3×10−4\eta_{\max} = 3 \times 10^{-4}, ηmin⁡=3×10−5\eta_{\min} = 3 \times 10^{-5}, and T=1000T = 1000 steps. These values are typical for training a medium-sized language model with AdamW.

At t=0t = 0: cos⁡(0)=1\cos(0) = 1, so η=3×10−5+12(2.7×10−4)(2)=3×10−4\eta = 3 \times 10^{-5} + \frac{1}{2}(2.7 \times 10^{-4})(2) = 3 \times 10^{-4}. This is the peak, as expected.

At t=250t = 250 (25% through training): cos⁡(π⋅250/1000)=cos⁡(π/4)≈0.707\cos(\pi \cdot 250 / 1000) = \cos(\pi/4) \approx 0.707. So:

η(250)=3×10−5+12(2.7×10−4)(1+0.707)=3×10−5+12(2.7×10−4)(1.707)=3×10−5+2.30×10−4≈2.60×10−4\begin{aligned} \eta(250) &= 3\times10^{-5} + \frac{1}{2}(2.7\times10^{-4})(1 + 0.707) \\ &= 3\times10^{-5} + \frac{1}{2}(2.7\times10^{-4})(1.707) \\ &= 3\times10^{-5} + 2.30\times10^{-4} \\ &\approx 2.60\times10^{-4} \end{aligned}

The learning rate has decayed from 3×10−43 \times 10^{-4} to approximately 2.60×10−42.60 \times 10^{-4}, a reduction of only about 13% despite being 25% through training. This is the "slow start" effect at work.

At t=500t = 500 (midpoint): cos⁡(π/2)=0\cos(\pi/2) = 0, so η=3×10−5+12(2.7×10−4)(1)=1.65×10−4\eta = 3\times10^{-5} + \frac{1}{2}(2.7\times10^{-4})(1) = 1.65\times10^{-4}. This is exactly the midpoint between ηmax⁡\eta_{\max} and ηmin⁡\eta_{\min}.

At t=750t = 750 (75% through training): cos⁡(3π/4)≈−0.707\cos(3\pi/4) \approx -0.707. So:

η(750)=3×10−5+12(2.7×10−4)(1−0.707)=3×10−5+12(2.7×10−4)(0.293)=3×10−5+3.96×10−5≈6.96×10−5\begin{aligned} \eta(750) &= 3\times10^{-5} + \frac{1}{2}(2.7\times10^{-4})(1 - 0.707) \\ &= 3\times10^{-5} + \frac{1}{2}(2.7\times10^{-4})(0.293) \\ &= 3\times10^{-5} + 3.96\times10^{-5} \\ &\approx 6.96\times10^{-5} \end{aligned}

By 75% of training, the learning rate has fallen to about 6.96×10−56.96 \times 10^{-5}, already very close to the minimum of 3×10−53 \times 10^{-5}. The remaining 25% of training will use a very small, slowly-changing learning rate, which is ideal for final refinement.

At t=1000t = 1000: cos⁡(π)=−1\cos(\pi) = -1, so η=3×10−5+0=3×10−5=ηmin⁡\eta = 3\times10^{-5} + 0 = 3\times10^{-5} = \eta_{\min}.

This worked example illustrates the curve's pacing clearly: the cosine schedule changes slowly near its peak, drops most rapidly around the midpoint, and slows again near the floor. A linear schedule also reaches the midpoint value of 1.65×10−41.65\times10^{-4} at step 500, but it reaches 2.60×10−42.60\times10^{-4} at about step 148 rather than step 250. Cosine therefore spends roughly 70% more steps above 2.60×10−42.60\times10^{-4} in this example; both schedules still spend half of training above the midpoint value.

Cosine Annealing with Warm Restarts

Training on a fixed cosine schedule marches the learning rate all the way to ηmin⁡\eta_{\min} and stops. This risks getting trapped in a sharp local minimum. Loshchilov and Hutter's 2016 paper "SGDR: Stochastic Gradient Descent with Warm Restarts" introduced a powerful extension: periodically reset the learning rate back to its maximum value, triggering a new cosine decay cycle.

Warm Restart

A warm restart is a technique where the learning rate is suddenly increased back toward its peak value after it has decayed toward its minimum. Unlike a cold start that reinitializes model weights, a warm restart only changes the learning rate while keeping the current parameter values. This allows the optimizer to escape sharp local minima and explore different regions of the loss landscape.

The Restart Schedule

In SGDR, each restart begins a new cosine annealing cycle with a potentially longer period. The period of the ii-th cycle is given by multiplying the base cycle length T0T_0 by a multiplicative factor raised to the cycle index:

Ti=T0⋅TmultiT_i = T_0 \cdot T_{\text{mult}}^i

where:

  • TiT_i: the number of steps in the ii-th cycle (starting from i=0i = 0)
  • T0T_0: the number of steps in the first (base) cycle
  • TmultT_{\text{mult}}: a multiplicative factor that lengthens each successive cycle. Common values are 1 (equal-length cycles) or 2 (doubling cycles)
  • ii: the zero-based cycle index

Within each cycle, the learning rate follows standard cosine decay from ηmax⁡\eta_{\max} down to ηmin⁡\eta_{\min}. Letting tcyclet_{\text{cycle}} denote the step count within the current cycle (reset to 0 at each restart):

η(t)=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡(π⋅tcycleTi))\eta(t) = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(1 + \cos\left(\frac{\pi \cdot t_{\text{cycle}}}{T_i}\right)\right)

where:

  • tcyclet_{\text{cycle}}: the number of steps elapsed since the last restart (or since training began for the first cycle)
  • TiT_i: the total length of the current cycle

When Tmult=1T_{\text{mult}} = 1, all cycles have the same length T0T_0, giving a periodic saw-tooth pattern. When Tmult=2T_{\text{mult}} = 2, each cycle is twice as long as the previous, so the schedule allocates progressively more training time to each refinement phase.

Out[4]:
Visualization
Cosine schedule with 4 equal-length restart cycles over 1000 steps.
Cosine annealing with equal-length warm restarts (T_mult=1, T0=250 steps). The learning rate resets to the peak value at steps 250, 500, and 750, creating four identical decay cycles over 1,000 total steps. Each restart can help the optimizer escape sharp local minima.
Cosine schedule with 250-step and 500-step cycles followed by the first quarter of a 1000-step cycle.
Cosine annealing with doubling warm restarts (T_mult=2, T0=250 steps). Complete cycles last 250 and 500 steps, with restarts at steps 250 and 750. The third 1,000-step cycle is only one quarter complete when this 1,000-step view ends, showing how later cycles allocate progressively more time to refinement.

Why Restarts Help

A sharp local minimum is characterized by a loss surface with steep walls: the loss is low at the minimum but increases rapidly in any direction. Models stuck in such minima tend to generalize poorly because sharp minima are sensitive to small perturbations in parameters. When the learning rate suddenly jumps back up at a restart, the optimizer takes larger steps that can escape the steep walls and find a path into a flatter, wider basin.

To understand this geometrically, imagine the parameter space as a landscape with many valleys. Some valleys are narrow and deep, corresponding to sharp minima where the Hessian of the loss has large eigenvalues. Others are wide and shallow, corresponding to flat minima where the loss changes slowly in all directions. With a monotonically decreasing learning rate, once the optimizer enters a narrow valley it tends to stay there because the small learning rate prevents it from climbing back out. With warm restarts, the sudden surge in learning rate gives the optimizer enough momentum to escape the narrow valley and potentially discover a wider, more generalizable one.

Flat minima are associated with better generalization in practice. The parameter values in a flat minimum remain nearly optimal over a broader range of perturbations, meaning that the model will perform well even under the slight distributional shifts that occur between training and test conditions. Warm restarts act as a form of implicit regularization by repeatedly biasing the optimizer toward broader basins.

There is also a practical benefit for model averaging. The models saved at the end of each cosine cycle, when the learning rate is near ηmin⁡\eta_{\min} and parameters have converged locally, represent good checkpoints for ensembling. Averaging several such checkpoints often outperforms using the last checkpoint alone, a technique sometimes called snapshot ensembling. Each cycle's endpoint represents a locally good solution reached from a different exploratory trajectory, and their average often achieves better generalization than any individual checkpoint.

When to Use Restarts

Warm restarts are most beneficial when the loss landscape is known to contain many suboptimal sharp minima, which is often the case in computer vision tasks and some classification problems. For large language model pretraining, practitioners have largely converged on single-cycle cosine decay without restarts. The reason is scale: LLMs are trained for hundreds of billions of tokens, and the cost of each "exploratory" restart is measured in millions of compute dollars. At that scale, the uncertainty about whether a restart will improve or hurt performance makes the single-cycle cosine the safer, more predictable choice.

Restarts are still useful in fine-tuning scenarios, where the model may get stuck in a local minimum quickly due to the small dataset size. They are also useful when you want multiple model checkpoints for ensembling without re-running training from scratch. If your task has a clear evaluation signal and you can afford the extra exploration, the doubling schedule (Tmult=2T_{\text{mult}} = 2) often offers a good balance between exploration and refinement.

Snapshot Ensembling

One of the most practical applications of warm restarts is snapshot ensembling, introduced by Huang et al. in the paper "Snapshot Ensembles: Train 1, Get M for Free" (2017). The main insight is that when the learning rate resets to ηmax⁡\eta_{\max} at each restart, the model takes a fresh exploratory path through the loss field. Each time the learning rate decays back toward ηmin⁡\eta_{\min}, the model converges to a locally good solution. By saving model weights at the end of each cycle, you collect a set of diverse models, all trained on the same data in a single training run.

Averaging the predictions of these snapshot models at inference time produces an ensemble that typically outperforms any individual snapshot. Each cycle's endpoint represents a distinct region of parameter space, reached after exploring from a different starting point (the previous cycle's endpoint, after a large learning rate jump displaced it). This is qualitatively different from averaging checkpoints from the same monotonic training run, where consecutive checkpoints are often very similar because the learning rate has been small for many steps.

The computational cost of snapshot ensembling is low compared to training separate models from scratch. You pay for MM cycle-endpoints worth of storage and MM forward passes at inference, but the total training compute is the same as a single model with warm restarts. For tasks where inference throughput is not critical, this trade-off is often worthwhile.

In practice, the quality of snapshot ensembles degrades if cycle lengths are very short, because the model has not converged within a single cycle and the snapshots are not sufficiently distinct from each other. Cycles of at least 10-20% of total training steps tend to produce better-quality individual snapshots. The doubling schedule, where later cycles are longer, naturally provides better snapshots from later cycles because each subsequent restart gives the model progressively more time to settle into its local minimum before being saved.

For large language models, snapshot ensembling is rarely used in pretraining because the cost of storing and running multiple large model copies at inference time is prohibitive. However, in settings like fine-tuning a model for a classification or retrieval task, where model size is more manageable and warm restarts are commonly used anyway, snapshot ensembling can provide meaningful quality improvements without any additional training cost.

Cosine vs. Linear Decay

The choice between cosine and linear decay often appears in ablation studies and reproducibility discussions. Both decay from ηmax⁡\eta_{\max} to ηmin⁡\eta_{\min}, but they allocate the budget differently across training steps.

Linear decay decreases the learning rate by the same absolute amount at every step. Starting from ηmax⁡\eta_{\max} at t=0t = 0 and ending at ηmin⁡\eta_{\min} at t=Tt = T, the linear schedule computes:

ηlinear(t)=ηmax⁡−(ηmax⁡−ηmin⁡)⋅tT\eta_{\text{linear}}(t) = \eta_{\max} - \frac{(\eta_{\max} - \eta_{\min}) \cdot t}{T}

where:

  • ηmax⁡−ηmin⁡\eta_{\max} - \eta_{\min}: the total decrease in learning rate over training
  • t/Tt / T: the fraction of training completed, ranging from 0 to 1

At any step tt, the learning rate has decreased by exactly the fraction t/Tt/T of the total budget.

Cosine decay is concave (curves downward) in the first half of training and convex (curves upward) in the second half. The difference between cosine and linear at any step tt is:

Δ(t)=ηcosine(t)−ηlinear(t)=12(ηmax⁡−ηmin⁡)(cos⁡(πtT)−1+2tT)\Delta(t) = \eta_{\text{cosine}}(t) - \eta_{\text{linear}}(t) = \frac{1}{2}(\eta_{\max} - \eta_{\min})\left(\cos\left(\frac{\pi t}{T}\right) - 1 + \frac{2t}{T}\right)

This difference Δ(t)\Delta(t) is positive in the first half of training (cosine is above linear) and negative in the second half (cosine is below linear), with Δ(0)=Δ(T)=0\Delta(0) = \Delta(T) = 0. The cosine schedule thus "front-loads" exploration by keeping the rate higher early on, then "back-loads" convergence by bringing the rate below the linear trajectory in the final phase.

The practical consequence is that cosine training often reaches a slightly lower final loss on the same number of steps compared to linear decay. The GPT-3, LLaMA, Mistral, and Gemma model families all use cosine schedules in their reported training configurations.

Out[5]:
Visualization
Area chart showing the signed difference between cosine and linear learning rates across 1000 training steps.
Difference between cosine and linear learning rate at each training step, expressed as a percentage of the total decay range. Positive values (first half) indicate that cosine keeps the learning rate higher than linear, favoring exploration. Negative values (second half) show cosine decaying below linear, applying finer updates during convergence. The symmetric shape reflects the cosine function's geometry.

Illustrative Comparison with Simulated Loss Curves

The schedule geometry alone does not guarantee that cosine will outperform linear decay on a real model. A controlled simulation can still illustrate the mechanism: higher rates early reduce optimization bias faster, while lower rates late reduce noise around the minimum. The next example averages 512 paired runs on the same noisy quadratic objective so that both schedules see identical gradient perturbations.

In[6]:
Code
import numpy as np

# Simulate paired SGD runs on a noisy quadratic: L(x) = 0.5 * x^2.
T_sim = 500
num_trials = 512
noise_scale = 0.5
rng_sim = np.random.default_rng(42)
gradient_noise = rng_sim.normal(0, noise_scale, size=(num_trials, T_sim))


def simulate_sgd(schedule_fn, noise_samples, x_init=2.0):
    x = np.full(noise_samples.shape[0], x_init, dtype=float)
    losses = []
    for t in range(noise_samples.shape[1]):
        lr = schedule_fn(t)
        grad = x + noise_samples[:, t]
        x = x - lr * grad
        losses.append(0.5 * np.mean(x**2))
    return np.asarray(losses)


cosine_fn = lambda t: cosine_lr(2.5e-2, 2.5e-3, t, T_sim)
linear_fn = lambda t: linear_lr(2.5e-2, 2.5e-3, t, T_sim)

cosine_losses = simulate_sgd(cosine_fn, gradient_noise)
linear_losses = simulate_sgd(linear_fn, gradient_noise)
Out[7]:
Console
Mean simulated losses across 512 paired noisy-quadratic runs:
  Cosine schedule (avg last 20 steps): 0.000410
  Linear schedule (avg last 20 steps): 0.000459
  Cosine improvement over linear: 10.7%
Out[8]:
Visualization
Log-scale mean loss curves for cosine and linear schedules across 512 paired noisy-quadratic simulations.
Mean loss across 512 paired noisy-quadratic optimization runs over 500 steps, shown on a logarithmic scale. Both schedules receive the same gradient perturbations. Cosine reduces loss faster early because it stays above linear, and its lower late-stage rate produces an average loss about 11% lower over the final 20 steps in this illustrative setup.

Cosine Schedule Parameters

A cosine schedule has four tunable parameters. Understanding what each controls helps you set them deliberately rather than by guessing.

Maximum Learning Rate

The maximum learning rate ηmax⁡\eta_{\max} determines the initial optimization speed. In a warmup-plus-cosine schedule (by far the most common configuration in large model training), ηmax⁡\eta_{\max} is the peak the warmup ramps toward. Setting this value follows the same guidelines as setting a fixed learning rate: too large and gradients explode, too small and training stalls. Linear scaling with batch size (multiply by a factor of B/BrefB / B_{\text{ref}} when increasing batch size from BrefB_{\text{ref}} to BB) is a common heuristic, though the right value for a given model ultimately requires some experimentation.

For AdamW-based LLM training, values in the range [1×10−4,3×10−4][1 \times 10^{-4}, 3 \times 10^{-4}] have been widely used at scales from a few hundred million parameters up to the tens of billions. Smaller learning rates around 1×10−51 \times 10^{-5} are typical for fine-tuning, where the model is already near a good solution and large steps risk destroying the pretrained representations.

Minimum Learning Rate

The minimum learning rate ηmin⁡\eta_{\min} is the floor the cosine curve decays toward. In many implementations this is simply 0, meaning the learning rate completely vanishes at the end of training. However, setting ηmin⁡\eta_{\min} to a small positive value like 10−610^{-6} or ηmax⁡/100\eta_{\max} / 100 can improve final performance by allowing the optimizer to keep making small corrections through the last steps. In LLaMA, for example, ηmin⁡\eta_{\min} is set to ηmax⁡/10\eta_{\max} / 10, meaning the final learning rate is 10% of the peak.

The choice of ηmin⁡\eta_{\min} has a subtle but real effect on the final model quality. Setting it to exactly zero creates a situation where the optimizer effectively stops learning in the final steps, which can be desirable if you want the model to fully converge to a minimum. Setting it to a non-zero value keeps a residual update signal alive, which helps when the training data has been partially exhausted and the model is in danger of overfitting. Many practitioners use ηmin⁡=ηmax⁡/10\eta_{\min} = \eta_{\max} / 10 as a safe default that avoids both extremes.

Total Steps and Decay Horizon

The parameter TT represents the total number of steps over which the cosine decay runs. In a single-cycle schedule this equals the full training duration. In a warmup-plus-cosine schedule, TT typically spans from the end of warmup to the end of training, though it can extend slightly beyond the actual training duration as a practical trick. If TT is set smaller than the actual training length, the learning rate will linger at ηmin⁡\eta_{\min} for the remaining steps, which is effectively the same as turning off the optimizer. Setting TT slightly larger than training means the schedule has not fully decayed at the end, keeping the learning rate from reaching zero.

One subtle design choice is whether to count warmup steps as part of TT. In some frameworks, TT refers to the total training steps including warmup, and the cosine decay starts from zero alongside the warmup. In others, warmup is treated as a separate phase and cosine decay begins after warmup ends, with TT counting only the post-warmup steps. The difference matters: if you set T=10000T = 10000 but warmup takes 1000 steps, you want to be explicit about whether the cosine curve should span 10000 or 9000 steps. Many practitioners prefer setting TT equal to the total training budget and letting the warmup logic handle the first phase separately, which is the pattern we follow in our implementation below.

Cycle Length for Warm Restarts

When using warm restarts, T0T_0 (the first cycle length) and TmultT_{\text{mult}} together determine the restart pattern. Short cycles with Tmult=1T_{\text{mult}} = 1 create frequent restarts that help exploration but may prevent the model from converging deeply into any minimum. Long initial cycles with Tmult=2T_{\text{mult}} = 2 start with a substantial refinement phase and progressively lengthen subsequent cycles, which often performs better for large models. A common starting point for language model training is to skip restarts entirely and use a single long cosine cycle, reserving restarts for computer vision tasks where snapshot ensembling benefits are more established.

When you do use restarts, the cycle length T0T_0 should be at least long enough for the model to make meaningful progress within each cycle. For typical LLM fine-tuning runs of 5,000-20,000 steps, T0T_0 values between 500 and 2000 are common. For pretraining runs of hundreds of thousands of steps, T0T_0 values of 10,000-50,000 steps are more appropriate.

How Real Models Use Cosine Schedules

The cosine schedule has been adopted by nearly every large language model developed in the past several years. Examining the published training recipes of prominent models reveals both the universality of the cosine approach and the specific parameter choices that practitioners have converged on.

GPT-3 (Brown et al., 2020) uses cosine decay to reduce the learning rate from ηmax⁡\eta_{\max} to 10%10\% of its peak value over the full training run. The warmup lasts for the first 375 million training tokens, which represents roughly 0.05% of the 300 billion token training corpus. The model uses a peak learning rate of 6×10−56 \times 10^{-5} for the 175-billion-parameter version. GPT-3 trained smaller and larger models with correspondingly adjusted peak learning rates, following a rough inverse square-root scaling with model size: larger models use lower peak learning rates, showing the intuition that wider parameter spaces benefit from more conservative steps.

LLaMA (Touvron et al., 2023) follows a nearly identical pattern: cosine decay with ηmin⁡=ηmax⁡/10\eta_{\min} = \eta_{\max} / 10, a warmup over the first 2000 steps, and a peak learning rate of 3×10−43 \times 10^{-4} for the 7-billion-parameter model. The 65-billion-parameter model uses a lower peak of 1.5×10−41.5 \times 10^{-4}. The LLaMA paper provides some of the clearest documentation of cosine schedule usage among published large models, explicitly noting that the minimum learning rate rule of ηmin⁡=ηmax⁡/10\eta_{\min} = \eta_{\max} / 10 was adopted from earlier empirical guidance.

LLaMA 2 (Touvron et al., 2023) extended the original LLaMA recipe with longer training horizons (up to 2 trillion tokens for the 7B model) while retaining the same schedule structure. One practical detail from LLaMA 2's training: the model's context length was increased midway through training, and the schedule was adjusted accordingly. The cosine decay was restarted from its current value at the point of context length expansion, treating the extended training as a continuation of the original schedule. This illustrates an important practical consideration: even with a planned cosine decay, real training runs often require ad-hoc adjustments that the rigid schedule does not account for.

Mistral-7B uses cosine decay with a warmup of 2000 steps and a peak learning rate of 3×10−43 \times 10^{-4}, trained on 1 trillion tokens. Like LLaMA, the minimum learning rate is 10% of the maximum. Mistral's architecture differs from LLaMA in using sliding window attention and grouped query attention, but the training schedule is nearly identical, suggesting the cosine recipe generalizes across these architectural variations.

PaLM (Chowdhery et al., 2022) uses cosine decay from 1×10−21 \times 10^{-2} down to 1×10−41 \times 10^{-4} for the 540-billion-parameter model, with linear warmup over 10,000 steps. The higher initial learning rate reflects PaLM's use of the Adafactor optimizer rather than Adam. Adafactor normalizes gradient magnitudes differently from Adam, allowing it to work well with higher raw learning rates. The cosine schedule structure is identical; only the scale differs.

Gemma (Team et al., 2024) from Google DeepMind also adopts the cosine schedule with warmup, using a peak learning rate of 1×10−31 \times 10^{-3} and a cosine decay to 1×10−41 \times 10^{-4} over the course of training. The relatively high peak learning rate, higher than LLaMA's, reflects the use of smaller batch sizes and different gradient accumulation strategies, again illustrating how the cosine schedule's structure transfers across configurations while the specific values need tuning.

Falcon (Penedo et al., 2023) trains with a cosine schedule from 3×10−43 \times 10^{-4} down to 3×10−53 \times 10^{-5}, trained on up to 1 trillion tokens. Falcon's training run is notable for its use of large batch sizes (up to 2 million tokens per batch), which required careful adjustment of the peak learning rate using linear scaling. Even with this scaling, the cosine decay structure remained unchanged.

The consistency across these very different models (in scale, architecture, training data, and optimizer variant) suggests that the cosine schedule tolerates these variations. The shared pattern is: a short warmup (a few hundred to a few thousand steps), a cosine decay over the full training run, and a minimum learning rate at roughly 10% of the peak. This recipe transfers well across scales and datasets.

Peak Learning Rate Scaling with Model Size

One pattern that emerges from these model comparisons is a rough relationship between model size and peak learning rate. Larger models tend to use lower peak learning rates, and this is not purely empirical. The Chinchilla scaling paper (Hoffmann et al., 2022) identified that the optimal number of training tokens scales linearly with model parameters. For a given compute budget, larger models are trained on fewer tokens (though still many), and their parameter updates span a much larger space. These two factors together suggest that larger models benefit from more conservative per-step updates.

A commonly cited approximation is that peak learning rate scales with model size NN roughly as ηmax⁡∝N−0.5\eta_{\max} \propto N^{-0.5}, though empirical fits vary across different research groups and architectures. In practice, the exact scaling exponent matters less than the general direction: if you double your model size, experimenting with a learning rate of ηmax⁡/2\eta_{\max} / \sqrt{2} rather than ηmax⁡\eta_{\max} is a reasonable first step. The cosine schedule's decay structure remains unchanged; you are adjusting only the peak value.

This relationship between model size and learning rate also interacts with the choice of ηmin⁡\eta_{\min}. If ηmin⁡\eta_{\min} is defined as a fixed fraction of ηmax⁡\eta_{\max}, as in the LLaMA convention of ηmin⁡=ηmax⁡/10\eta_{\min} = \eta_{\max} / 10, then reducing ηmax⁡\eta_{\max} also reduces ηmin⁡\eta_{\min} proportionally. This is generally the right behavior: finer final updates correspond to a finer resolution of parameter space, which is appropriate for larger, more expressive models settling into more precise local minima.

Code Implementation

We will build the cosine schedule from scratch to make the formula concrete, then show how to use PyTorch's built-in CosineAnnealingLR and CosineAnnealingWarmRestarts schedulers.

From-Scratch Implementation

Let us start by implementing the cosine decay function directly, without any framework.

In[9]:
Code
import math

import numpy as np


def cosine_schedule(eta_max, eta_min, t, T):
    """Compute the cosine-decayed learning rate at step t out of T total steps."""
    return eta_min + 0.5 * (eta_max - eta_min) * (1 + math.cos(math.pi * t / T))


def linear_schedule(eta_max, eta_min, t, T):
    """Compute the linearly-decayed learning rate at step t out of T total steps."""
    return eta_max - (eta_max - eta_min) * t / T


# Schedule parameters
eta_max = 3e-4
eta_min = 3e-5
T = 1000

steps = np.arange(0, T + 1)
cosine_lrs = np.array([cosine_schedule(eta_max, eta_min, t, T) for t in steps])
linear_lrs = np.array([linear_schedule(eta_max, eta_min, t, T) for t in steps])
Out[10]:
Console
Schedule parameters:
  eta_max = 3.0e-04
  eta_min = 3.0e-05
  T (total steps) = 1000

Learning rates at key steps:
  Step 0 (start):    cosine = 3.00e-04, linear = 3.00e-04
  Step 250 (25%):   cosine = 2.60e-04, linear = 2.32e-04
  Step 500 (50%):   cosine = 1.65e-04, linear = 1.65e-04
  Step 750 (75%): cosine = 6.95e-05, linear = 9.75e-05
  Step 1000 (end):    cosine = 3.00e-05, linear = 3.00e-05

Both schedules agree at the start and end, but they diverge significantly in between. At 25% through training, cosine decays more slowly than linear, preserving a higher learning rate for early exploration. At 75% through training, cosine is already below the linear trajectory, giving finer updates for the convergence phase.

Cosine with Warm Restarts

Next, let us implement the warm restart schedule from the SGDR paper.

In[11]:
Code
def cosine_with_restarts(eta_max, eta_min, t, T0, T_mult=1):
    """
    Compute the cosine-annealed learning rate with warm restarts at step t.

    Args:
        eta_max: Peak learning rate at the start of each cycle.
        eta_min: Floor learning rate at the end of each cycle.
        t: Current step index (global, across all cycles).
        T0: Length of the first cycle in steps.
        T_mult: Multiplicative factor for cycle lengths (1 = equal, 2 = doubling).
    """
    if T_mult == 1:
        # All cycles have length T0
        t_cycle = t % T0
        T_current = T0
    else:
        # Find which cycle we are in
        T_current = T0
        t_remaining = t
        while t_remaining >= T_current:
            t_remaining -= T_current
            T_current = int(T_current * T_mult)
        t_cycle = t_remaining

    return eta_min + 0.5 * (eta_max - eta_min) * (
        1 + math.cos(math.pi * t_cycle / T_current)
    )


# Compute two restart schedules
T_total = 1000
T0 = 250  # First cycle length

restart_equal = np.array(
    [cosine_with_restarts(eta_max, eta_min, t, T0, T_mult=1) for t in steps]
)
restart_double = np.array(
    [cosine_with_restarts(eta_max, eta_min, t, T0, T_mult=2) for t in steps]
)
Out[12]:
Console
Warm restart schedule with T0=250 steps:

  Equal cycles (T_mult=1): 4 restarts of 250 steps each

  Restart steps (T_mult=1): [250, 500, 750, 1000]

  Cycle lengths (T_mult=2): [250, 500]
  Restart steps (T_mult=2): [250]

The equal-length restart schedule creates four identical cycles of 250 steps, with the learning rate returning to ηmax⁡\eta_{\max} at steps 250, 500, and 750. The doubling schedule creates progressively longer cycles, giving later cycles more time to refine the model.

PyTorch Scheduler Integration

In practice, you will use PyTorch's built-in schedulers rather than reimplementing from scratch. Let us demonstrate how to configure both options.

In[13]:
Code
import torch.nn as nn
import torch.optim as optim

# Create a minimal model and optimizer for demonstration
model = nn.Linear(10, 1)
optimizer = optim.AdamW(model.parameters(), lr=eta_max)

# CosineAnnealingLR: single cycle, no restarts
scheduler_cosine = optim.lr_scheduler.CosineAnnealingLR(
    optimizer,
    T_max=T,  # Number of steps in the half-period (= total decay steps)
    eta_min=eta_min,  # Floor learning rate
)

# Reset optimizer for second scheduler
optimizer2 = optim.AdamW(model.parameters(), lr=eta_max)

# CosineAnnealingWarmRestarts: multiple cycles with optional doubling
scheduler_restarts = optim.lr_scheduler.CosineAnnealingWarmRestarts(
    optimizer2,
    T_0=T0,  # Steps in the first cycle
    T_mult=1,  # Equal-length cycles
    eta_min=eta_min,  # Floor learning rate
)

# Collect learning rates for visualization
pytorch_cosine_lrs = []
pytorch_restart_lrs = []

for step in range(T + 1):
    pytorch_cosine_lrs.append(
        scheduler_cosine.get_last_lr()[0]
        if step > 0
        else optimizer.param_groups[0]["lr"]
    )
    pytorch_restart_lrs.append(
        scheduler_restarts.get_last_lr()[0]
        if step > 0
        else optimizer2.param_groups[0]["lr"]
    )

    if step < T:
        scheduler_cosine.step()
        scheduler_restarts.step()

pytorch_cosine_lrs = np.array(pytorch_cosine_lrs)
pytorch_restart_lrs = np.array(pytorch_restart_lrs)
Out[14]:
Console
PyTorch CosineAnnealingLR:
  Starting LR: 3.00e-04
  LR at step 500: 1.65e-04
  Final LR: 3.00e-05

PyTorch CosineAnnealingWarmRestarts:
  Starting LR: 3.00e-04
  LR at step 250 (first restart): 3.00e-04
  LR at step 500 (second restart): 3.00e-04
  Final LR: 3.00e-04

The PyTorch CosineAnnealingLR takes T_max as the half-period of the cosine, not the full decay duration. When you want a single decay from ηmax⁡\eta_{\max} to ηmin⁡\eta_{\min} over TT steps, pass T_max=T. When you want restarts, use CosineAnnealingWarmRestarts with T_0 matching your desired cycle length.

One important note about CosineAnnealingLR: PyTorch's implementation uses T_max as a half-period, meaning that after Tmax⁡T_{\max} steps the learning rate reaches ηmin⁡\eta_{\min}, but if you continue calling scheduler.step() beyond that point, the learning rate will start increasing again toward ηmax⁡\eta_{\max}. In practice, you want to stop calling the scheduler after the final step, or set the training loop to end exactly at Tmax⁡T_{\max} steps.

Combining Warmup with Cosine Decay

The most common configuration in large-scale training combines a linear warmup phase with a subsequent cosine decay. Let us implement this combined schedule.

In[15]:
Code
def warmup_cosine_schedule(eta_max, eta_min, t, T_warmup, T_total):
    """
    Compute the learning rate at step t for a warmup-then-cosine schedule.

    Args:
        eta_max: Peak learning rate reached at end of warmup.
        eta_min: Floor learning rate at end of training.
        t: Current step index.
        T_warmup: Number of warmup steps.
        T_total: Total training steps.
    """
    if t < T_warmup:
        # Linear warmup from 0 to eta_max
        return eta_max * t / T_warmup
    else:
        # Cosine decay from eta_max to eta_min
        t_decay = t - T_warmup
        T_decay = T_total - T_warmup
        return eta_min + 0.5 * (eta_max - eta_min) * (
            1 + math.cos(math.pi * t_decay / T_decay)
        )


# Configuration matching typical LLM training (proportions)
T_warmup_steps = 100  # 10% of training for warmup
T_total_steps = 1000

warmup_cosine_lrs = np.array(
    [
        warmup_cosine_schedule(
            eta_max, eta_min, t, T_warmup_steps, T_total_steps
        )
        for t in range(T_total_steps + 1)
    ]
)
Out[16]:
Console
Warmup + Cosine schedule (100 warmup steps, 1000 total):

  Step 0 (cold start): 0.00e+00
  Step 50 (mid warmup): 1.50e-04
  Step 100 (peak, warmup end): 3.00e-04
  Step 550 (midpoint decay): 1.65e-04
  Step 1000 (end): 3.00e-05
Out[17]:
Visualization
Line chart of warmup-cosine schedule with shaded warmup region and annotated peak learning rate.
The warmup-plus-cosine learning rate schedule over 1,000 training steps with a 100-step linear warmup. The learning rate rises linearly from zero to the peak during warmup (shaded region), then follows a cosine decay through the remaining 900 steps. This two-phase design is the standard configuration for training large language models including LLaMA and GPT.

This combined schedule is what you will see in virtually every modern LLM training recipe. The warmup period prevents the unstable early-training loss spikes discussed in the Learning Rate Warmup chapter, and the cosine decay efficiently brings the learning rate to its minimum over the bulk of training.

Key Parameters

The key parameters for cosine learning rate schedules are:

  • eta_max (peak learning rate): The starting learning rate after warmup. Typical values range from 1e-4 to 3e-4 for large language models using AdamW.
  • eta_min (minimum learning rate): The floor learning rate at the end of training. A common choice is eta_max / 10 or a fixed small value like 1e-6.
  • T_max / T_total (decay steps): Total steps for the cosine decay. Usually set to the full training duration minus warmup steps, or slightly longer to avoid the zero-LR degenerate case.
  • T_warmup (warmup steps): Steps for linear warmup. Commonly 1-5% of total training steps.
  • T_0 (cycle length for restarts): Number of steps in the first restart cycle. Only relevant when using warm restarts.
  • T_mult (cycle multiplier): Multiplicative factor for restart cycle lengths. Common values are 1 (equal cycles) and 2 (doubling cycles).

Visualizing the Effect of Minimum Learning Rate

One underappreciated aspect of the cosine schedule is how dramatically the minimum learning rate changes the character of the decay. Let us compare several ηmin⁡\eta_{\min} values to see the differences clearly.

Out[18]:
Visualization
Multiple cosine decay curves with different minimum learning rates, ranging from 0 to 30% of peak.
Effect of minimum learning rate on cosine schedule shape. When eta_min is zero, the schedule decays all the way to zero and the optimizer effectively stops near the end of training. With a larger floor (10% of peak), the schedule retains a meaningful update signal throughout. The curve shape changes noticeably in the final 20% of training depending on this choice.

The visual differences are most pronounced in the final third of training. With ηmin⁡=0\eta_{\min} = 0, the learning rate decays all the way to zero, which can cause the optimizer to stall. With ηmin⁡=0.30⋅ηmax⁡\eta_{\min} = 0.30 \cdot \eta_{\max} (30% of peak), the decay is much more modest, and the model retains a useful update signal even at the end. The LLaMA choice of 10% sits in a practical middle ground: the floor is small enough that the parameter updates approach convergence, but large enough to prevent complete optimizer shutdown.

Using torch.optim.lr_scheduler.LambdaLR for Custom Schedules

If you need a warmup-plus-cosine schedule in PyTorch and want full control over the exact formula, LambdaLR lets you define the multiplier function directly. This approach is also compatible with Hugging Face's transformers library, which uses LambdaLR internally.

In[19]:
Code
import torch.nn as nn
import torch.optim as optim


def get_cosine_schedule_with_warmup(
    optimizer, num_warmup_steps, num_training_steps, eta_min_fraction=0.1
):
    """
    Create a cosine schedule with warmup using LambdaLR.

    Args:
        optimizer: The optimizer to schedule.
        num_warmup_steps: Steps for linear warmup.
        num_training_steps: Total training steps.
        eta_min_fraction: Minimum LR as fraction of peak LR (default: 10%).

    Returns:
        LambdaLR scheduler with cosine-plus-warmup multiplier.
    """

    def lr_lambda(current_step):
        if current_step < num_warmup_steps:
            return float(current_step) / float(max(1, num_warmup_steps))
        progress = float(current_step - num_warmup_steps) / float(
            max(1, num_training_steps - num_warmup_steps)
        )
        cosine_factor = 0.5 * (1.0 + math.cos(math.pi * progress))
        # Scale to [eta_min_fraction, 1.0]
        return eta_min_fraction + (1.0 - eta_min_fraction) * cosine_factor

    return optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)


# Test the custom schedule
model_demo = nn.Linear(10, 1)
opt_demo = optim.AdamW(model_demo.parameters(), lr=3e-4)
sched_demo = get_cosine_schedule_with_warmup(
    opt_demo, num_warmup_steps=100, num_training_steps=1000
)

lambda_lrs = []
for step in range(1001):
    lambda_lrs.append(opt_demo.param_groups[0]["lr"])
    if step < 1000:
        sched_demo.step()

lambda_lrs = np.array(lambda_lrs)
Out[20]:
Console
Custom LambdaLR warmup-cosine schedule:
  Step 0 (start): 0.00e+00
  Step 50 (mid warmup): 1.50e-04
  Step 100 (peak): 3.00e-04
  Step 550 (midpoint decay): 1.65e-04
  Step 1000 (end): 3.00e-05
  Ratio eta_end / eta_peak: 0.10

The LambdaLR approach gives you a single, auditable function for the entire schedule. The multiplier function returns values between eta_min_fraction and 1.0, and these multiply against the optimizer's base learning rate. If the base rate is 3e-4 and eta_min_fraction=0.1, the floor will be 3e-5, consistent with the LLaMA configuration.

Trapezoidal Schedules: A Modern Alternative

A more recent alternative that has gained attention for large-scale language model training is the trapezoidal (or "trapezoid") schedule. Rather than decaying from the very beginning, the trapezoidal schedule keeps the learning rate at its peak for most of training, then drops it sharply in a brief cooldown phase at the end.

The schedule has three phases:

  1. Warmup: Linear ramp from 0 to ηmax⁡\eta_{\max} over TwarmupT_{\text{warmup}} steps.
  2. Stable: Constant learning rate at ηmax⁡\eta_{\max} for TstableT_{\text{stable}} steps (typically 80-90% of total training).
  3. Cooldown: Cosine or linear decay from ηmax⁡\eta_{\max} to ηmin⁡\eta_{\min} over TcooldownT_{\text{cooldown}} steps.

The motivation for this schedule is compute-budget flexibility. With a cosine schedule, you must decide the total training duration TT in advance, because the decay rate depends on TT. If you want to train for 20% longer (perhaps because the model is still improving), the entire cosine curve needs to be rescaled. With a trapezoidal schedule, you can extend the stable phase arbitrarily without affecting the decay curve: just train for longer at the constant peak rate, then apply the same cooldown at the end.

Recent work including the Chinchilla analysis and follow-up scaling papers has noted that the final cooldown phase is disproportionately important for model quality. A model trained at constant learning rate for a very long time, then cooled down over a few thousand steps, can match or exceed the quality of a model trained with a cosine schedule for the same total compute. This property makes the trapezoidal schedule particularly useful for continual or incremental training, where the exact endpoint is not known in advance.

Out[21]:
Visualization
Comparison of trapezoidal and cosine learning rate schedules showing the stable plateau phase of the trapezoidal approach.
Trapezoidal learning rate schedule compared to cosine decay over 1,000 steps. The trapezoidal schedule maintains its peak learning rate through 80% of training, then decays steeply in a short cooldown window. This contrasts with the cosine schedule, which begins decaying immediately. The trapezoidal approach offers more flexibility for extending training runs without redesigning the entire schedule.

Choosing the Right Schedule for Your Task

With the core schedules covered, it helps to have a practical framework for deciding which one to apply to a given training scenario. The choice depends on your compute budget, whether the training duration is fixed, whether you need multiple model checkpoints, and how much hyperparameter tuning you can afford.

Fixed Budget Pretraining

If you are training a model from scratch with a fixed compute budget and you know the total number of steps in advance, the warmup-plus-cosine schedule is the safest default. It matches the training recipes of most publicly released LLMs, which means you can borrow hyperparameters from similar models as starting points. The warmup duration should cover 1-4% of total steps (2000 warmup steps out of 100,000 total is typical), and the minimum learning rate should be set to 10% of the peak. Starting from a peak learning rate that worked for a similar-sized model and adjusting up or down based on early loss curves is a reliable strategy.

The main risk with cosine decay in fixed-budget pretraining is choosing the wrong peak learning rate. If the peak is too high, training instabilities or loss spikes will appear early. If it is too low, the model will converge more slowly than necessary, wasting compute. Running a short learning rate range test (training for a few thousand steps at several candidate peak rates and comparing early loss trajectories) before committing to a full run saves significant compute at scale.

Variable or Extensible Training

If you are not certain how long training will run, perhaps because you are monitoring a validation metric and might stop early or extend the run depending on what you observe, the trapezoidal schedule handles this far more gracefully than cosine decay. You commit only to the cooldown duration and shape, and the stable phase can be extended arbitrarily. The cosine schedule's decay is defined relative to the total budget TT, so changing TT mid-run requires either rescaling the decay or accepting that the learning rate at the current step no longer matches the intended schedule.

For continual learning scenarios, where a model is periodically updated on new data without restarting from scratch, the trapezoidal schedule is also preferred. Each update cycle consists of a brief warmup from the current learning rate (not from zero), a stable phase on the new data, and a cooldown. The cosine schedule does not naturally accommodate this pattern because its decay rate is coupled to the total training horizon.

Fine-Tuning

Fine-tuning a pretrained model introduces different constraints from pretraining. The model is already near a good solution in its parameter space, and the risk of catastrophic forgetting (discussed in detail in the Finetuning Fundamentals chapter) means you want smaller, more careful updates throughout training. The peak learning rate should be roughly 10-100 times smaller than what you would use for pretraining, with typical values in the range of 1×10−51 \times 10^{-5} to 1×10−41 \times 10^{-4}.

With these smaller learning rates, warm restarts can be useful because the risk of a restart disturbing the pretrained representations is lower when the restart amplitude is small. A restart from 1×10−51 \times 10^{-5} is far less disruptive than a restart from 3×10−43 \times 10^{-4}. If you are fine-tuning on a small dataset where the model can overfit quickly, short restart cycles help the optimizer escape any sharp minima that form early in training.

The warmup phase is typically shorter in fine-tuning, often just 50-100 steps, because the pretrained model's parameters are already well-initialized and the optimizer does not need as much time to establish stable gradient directions.

Hyperparameter Tuning Budget

If you have limited computational resources for hyperparameter tuning, the cosine schedule's main advantage is that its shape is largely fixed once you specify ηmax⁡\eta_{\max}, ηmin⁡\eta_{\min}, and TT. Relative to an adaptive schedule like ReduceLROnPlateau, which requires specifying a validation metric, a patience threshold, a reduction factor, and a minimum learning rate, the cosine schedule is simpler to configure and more predictable.

For practitioners who can afford systematic hyperparameter sweeps, the most impactful parameters to tune in a cosine schedule are, roughly in order of importance: the peak learning rate, the warmup duration, and the minimum learning rate. The schedule shape itself (cosine vs. linear) is typically not the most sensitive hyperparameter; the absolute scale of the learning rate matters far more. This means you can fix the cosine schedule structure and focus your tuning budget on the peak value and warmup, which is an efficient allocation of experimental resources.

Limitations and Practical Considerations

The cosine schedule is not a silver bullet, and understanding its failure modes helps you adapt it appropriately.

One important limitation is sensitivity to the total training budget. The cosine curve is defined relative to the total number of steps TT, which means you must fix your training duration before you begin. If you need to stop training early or extend it, the schedule behaves unexpectedly: stopping early leaves the learning rate higher than intended at the end, while extending training pins the rate at ηmin⁡\eta_{\min} for the extra steps. This contrasts with some more adaptive schedules that respond to validation metrics rather than step counts. In practice, most large-scale training runs fix the compute budget ahead of time, so this limitation is manageable, but it requires forethought.

Another subtlety is that the cosine schedule decays most rapidly in the middle of training. If your model has not found a reasonable region of parameter space by the midpoint, the accelerating decay phase can be disruptive rather than helpful. This is why the warmup phase is so commonly paired with cosine decay: warmup ensures the model has made initial progress before the cosine curve begins decelerating. Models trained without warmup into a cosine schedule sometimes exhibit a "plateau then collapse" pattern, where the flat early-training behavior is followed by a sudden rapid descent as the cosine derivative peaks.

Warm restarts add a new dimension of complexity. Each restart risks disturbing a converging model and requires the optimizer to re-explore the loss landscape. If cycle lengths are too short, the model never fully converges within a cycle before being disturbed again. If they are too long, restarts offer little benefit over a single cosine decay. Tuning T0T_0 and TmultT_{\text{mult}} is often done empirically, which adds experimental overhead. For most language model training runs, practitioners have moved toward single-cycle cosine decay without restarts, reserving restarts for cases where snapshot ensembling is a goal.

The combination of warmup and cosine decay introduces two more hyperparameters (TwarmupT_{\text{warmup}} and ηmin⁡\eta_{\min}) on top of the base learning rate, making the full schedule four-dimensional. Each parameter interacts with the others. A very long warmup followed by aggressive cosine decay can starve the model of a useful exploratory phase, while a very short warmup into a gentle decay may fail to stabilize early training. Standard guidance (1-4% of steps for warmup, ηmin⁡=ηmax⁡/10\eta_{\min} = \eta_{\max} / 10) provides good starting points, but some tuning is usually required for novel architectures.

The trapezoidal schedule, discussed earlier, addresses the budget-sensitivity problem by decoupling the decay from the total training duration. However, it introduces its own sensitivity: the quality of training depends heavily on the cooldown duration and shape. Too short a cooldown and the model does not fully converge; too gradual and the plateau phase ends earlier than necessary. Research comparing trapezoidal and cosine schedules at large scale is still ongoing, and the field has not yet converged on a clear winner.

Debugging Training Runs with Schedule Issues

A number of common training pathologies can be traced to schedule misconfiguration, and recognizing these patterns early can save significant compute.

Loss spikes at training start almost always indicate that the peak learning rate is too high, or that the warmup is too short. The optimizer takes massive steps on the random initialization before establishing a stable gradient direction. Reducing ηmax⁡\eta_{\max} by a factor of 3-5 or doubling the warmup duration usually resolves this. A useful diagnostic: if loss spikes resolve within 50-100 steps and the loss curve then continues normally, the spike was a minor warmup artifact. If loss continues oscillating or diverges, the peak learning rate is the likely culprit.

Stalling in mid-training (loss plateauing well above the expected final value, for many steps) can indicate that the cosine decay has progressed too quickly to a regime where updates are too small to escape the current region. This often happens when ηmax⁡\eta_{\max} was set too conservatively. One remedy is to restart training with a higher peak or a longer warmup that allows more time at the peak learning rate. Another sign of this failure: the cosine curve is already 60-70% decayed but the loss has not yet entered its expected rapid descent phase.

Loss oscillation in the final phase (the last 10-20% of training) can indicate that ηmin⁡\eta_{\min} is too large. If the minimum learning rate is 30% or 50% of the peak, the optimizer is still taking substantial steps when it should be settling. Reducing ηmin⁡\eta_{\min} to 5-10% of the peak typically smooths the final convergence.

Unexpected loss increase after extending training is a sign that the cosine schedule has already decayed to ηmin⁡\eta_{\min} and the extra steps are essentially running with the optimizer frozen at the minimum rate. In some optimizers and loss surfaces, running at a near-zero learning rate for many steps can cause the model to drift slightly as momentum terms interact with nearly-zero updates in unexpected ways. The safest approach when extending training is to either restart the schedule from the current learning rate (effectively treating the extension as a new cosine cycle starting from wherever the schedule left off) or switch to the trapezoidal approach from the beginning.

Interaction with Batch Size

The cosine schedule interacts with batch size in a way that is easy to overlook. As covered in the context of learning rate warmup, the linear scaling rule suggests that if you multiply your batch size by kk, you should multiply your peak learning rate by kk as well to maintain the same effective optimization trajectory. This rule applies directly to the cosine schedule: both ηmax⁡\eta_{\max} and ηmin⁡\eta_{\min} should scale with batch size, keeping their ratio constant.

However, the warmup duration also needs to adjust when batch size changes. With a larger batch, each step covers more data, and the variance of gradient estimates is lower. The optimizer orients itself more quickly, meaning it needs fewer warmup steps in terms of iterations (though about the same number of data samples). A practical rule of thumb is to scale warmup steps inversely with batch size: if you double the batch size, halve the number of warmup steps, so the total number of samples seen during warmup remains constant.

The total number of decay steps TT also interacts with batch size through the total token count. If your budget is fixed in tokens rather than in steps, doubling the batch size halves the number of steps for the same token count. The cosine schedule's TT should be set to the token-equivalent step count, not an arbitrary step count, to ensure the decay profile is consistent regardless of batch size.

Despite these considerations, the cosine schedule performs reliably across a wide range of scales and architectures. The GPT series from OpenAI, the LLaMA family from Meta, Mistral, and countless smaller research models all report using cosine decay as their default schedule. Its widespread adoption has produced substantial empirical evidence about good default values, making it far easier to configure than more complex adaptive schedules that change behavior based on validation metrics.

Looking ahead, the next chapter on large batch training discusses how schedule parameters interact with batch size, and why the linear scaling rule for learning rates must be applied carefully when the cosine schedule is in use.

Summary

The cosine learning rate schedule applies the cosine function to interpolate the learning rate smoothly from a peak value ηmax⁡\eta_{\max} down to a minimum ηmin⁡\eta_{\min} over a fixed number of training steps. Its key properties are:

  • The decay rate is slow at the start and end of training and fastest in the middle, matching the natural phases of gradient descent.
  • Cosine decay consistently outperforms step decay (no abrupt drops) and slightly outperforms linear decay in final model quality, because it front-loads exploration and back-loads convergence.
  • The derivative of the cosine schedule is a scaled sine function, which is zero at the boundaries and maximal at the midpoint. This gives the schedule its characteristic pacing.
  • Warm restarts (SGDR) periodically reset the learning rate, helping the optimizer escape sharp minima and enabling snapshot ensembling. They are most useful in computer vision and fine-tuning scenarios, and are less common in large-scale LLM pretraining.
  • The most common production configuration combines a short linear warmup (1-5% of total steps) with a long single-cycle cosine decay.
  • Critical hyperparameters are ηmax⁡\eta_{\max}, ηmin⁡\eta_{\min} (typically ηmax⁡/10\eta_{\max}/10), total decay steps TT, and (when using restarts) cycle length T0T_0 and multiplier TmultT_{\text{mult}}.
  • The trapezoidal schedule tolerates variable training durations better, at the cost of a concentrated cooldown phase.

The cosine schedule is now the de facto standard for training large language models, not because it is theoretically optimal but because it is smooth, interpretable, and empirically reliable across a broad range of training scenarios. The worked numerical examples above make clear exactly how the formula allocates the learning rate budget across training, and the PyTorch implementations show how to apply these principles directly in practice.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about cosine learning rate schedules.

Cosine Learning Rate Schedule Quiz

Question 1 of 80 of 8 completed
What does the cosine learning rate schedule compute at step t=0 out of T total steps?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026cosinelearning, author = {Michael Brenndoerfer}, title = {Cosine Learning Rate Schedule: Decay, Restarts, and Warmup}, year = {2026}, url = {https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Cosine Learning Rate Schedule: Decay, Restarts, and Warmup. Retrieved from https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup
MLAAcademic
Michael Brenndoerfer. "Cosine Learning Rate Schedule: Decay, Restarts, and Warmup." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup>.
CHICAGOAcademic
Michael Brenndoerfer. "Cosine Learning Rate Schedule: Decay, Restarts, and Warmup." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Cosine Learning Rate Schedule: Decay, Restarts, and Warmup'. Available at: https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Cosine Learning Rate Schedule: Decay, Restarts, and Warmup. https://mbrenndoerfer.com/writing/cosine-learning-rate-schedule-decay-restarts-warmup

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.