Hyperparameter Selection: Search, Transfer, Default Recipes

Michael BrenndoerferJanuary 29, 202658 min read

Part of Language AI Handbook

Covers systematic strategies for hyperparameter search, how muP enables transfer across model scales.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Hyperparameter Selection

Training a language model involves far more choices than just the architecture. The learning rate, batch size, weight decay coefficient, warmup schedule, dropout rate, and dozens of other knobs all need values before a single gradient flows. Getting these wrong means wasted compute at best, divergent training at worst. Getting them right can mean the difference between a model that generalizes beautifully and one that memorizes noise.

In the chapters of Part XXXII, we have worked through individual training components in detail: learning rate warmup, decay schedules, large batch training, weight decay, gradient accumulation, and training stability. Now we pull these threads together. Hyperparameter selection is the discipline of choosing sensible values for all of these components jointly, understanding which choices matter most, and building a workflow that lets you iterate efficiently without re-running experiments from scratch every time.

This chapter covers four interconnected ideas. First, how to search for good hyperparameter configurations systematically. Second, how knowledge about good hyperparameters transfers between models of different sizes, enabling you to start large runs with confidence rather than guesswork. Third, which hyperparameters are critical (small deviations cause large performance differences) versus which are robust (performance is nearly flat across a wide range). Fourth, practical default recipes that have emerged from the research community and production practice, giving you a solid starting point for any new training run.

Why Hyperparameter Selection is Hard

The challenge is not simply that there are many hyperparameters. The challenge is that they interact. A learning rate that works well with a batch size of 256 may cause divergence at batch size 4096. Weight decay that prevents overfitting at one scale may cripple convergence at another. Warmup duration that is appropriate for a cosine schedule may be inadequate for a constant rate. The search space is high-dimensional, the objective function is expensive to evaluate (each evaluation requires training a model), and the interactions between dimensions create a landscape with local optima and flat plateaus that are hard to search.

Compounding this, the feedback loop is slow. Unlike tuning a software system where you can run hundreds of tests per hour, training a large language model takes days or weeks. Evaluating 100 hyperparameter configurations on a model with 7 billion parameters is not feasible. This computation cost has driven researchers toward three complementary strategies: systematic search on small proxy models, theoretical understanding of how hyperparameters scale, and empirical default recipes derived from the collective experience of many training runs.

There is also a subtler challenge: the objective landscape you observe during short proxy runs may not match the landscape you care about. A configuration that minimizes loss after 5,000 steps may not be the one that minimizes loss after 100,000 steps. Regularization effects take time to manifest. The optimal weight decay for a long run is often higher than the optimal weight decay for a short run, because the model has more time to overfit. These temporal dependencies make proxy experiments informative but not perfectly predictive of full-scale results.

Hyperparameter search is the process of finding configurations that perform well by evaluating multiple candidates. The right search strategy depends on your computational budget, the number of hyperparameters you are tuning, and how smooth the loss landscape is as a function of those parameters.

Grid search evaluates every point on a regular grid over the hyperparameter space. If you want to test three learning rates and four weight decay values, grid search runs twelve experiments covering all combinations. It is exhaustive and reproducible, with results that are easy to reason about.

The problem is combinatorial explosion. Adding one more dimension to search multiplies the number of experiments by the number of values in that dimension. A grid over six hyperparameters with five values each requires 56=15,6255^6 = 15{,}625 experiments. This is practical only for one or two hyperparameters or when each experiment is very cheap.

Grid search also wastes resources on less important dimensions. If learning rate matters a lot and weight decay barely matters, half your grid experiments vary weight decay while holding learning rate fixed at a bad value. You get many evaluations at useless configurations. The structure that makes grid search appealing, its exhaustiveness over the chosen axes, becomes a liability when the axes are not equally important, which is virtually always the case in practice.

Random search samples configurations uniformly at random from the hyperparameter space. Instead of a regular grid, you get scattered points. This sounds less systematic, but it is typically more efficient than grid search for the following reason: most hyperparameter landscapes have low intrinsic dimensionality. A few dimensions matter a lot and the rest matter little. In a grid, you spend 1/N1/N of your evaluations on each unique value of an important dimension; in random search, every evaluation uses a different value of that important dimension, so you explore it more thoroughly.

The seminal result from Bergstra and Bengio (2012) showed that random search finds equally good or better configurations than grid search using far fewer evaluations, especially when the number of relevant hyperparameters is small relative to the total number being tuned. For practical LLM training, where you might have fifteen hyperparameters but only three or four that substantially affect the outcome, random search is a natural choice.

Random Search Efficiency

The efficiency gain of random search over grid search is proportional to the ratio of total hyperparameters to important hyperparameters. If 3 out of 10 hyperparameters are important, random search explores the important dimensions roughly 3x more efficiently than grid search at the same compute budget.

Random search is embarrassingly parallel: all trials are independent and can run simultaneously. With a fixed budget of, say, 50 experiments, you can evaluate all 50 in parallel if you have the hardware. This property is invaluable in large-scale LLM work, where parallelizing across many GPUs is the norm.

The main limitation of random search is that it does not learn from previous evaluations. Each new sample is drawn independently of all past results, so random search cannot refine its search around promising regions. If the loss landscape has a sharp, narrow optimum, you may need many samples to hit it by chance. This observation motivates the next step in the progression: adaptive search methods that use past results to guide future queries.

Bayesian Optimization

Bayesian optimization treats hyperparameter search as a sequential decision problem. After each experiment, it updates a probabilistic model of how performance depends on the hyperparameter configuration and uses that model to choose the next configuration to evaluate. The key idea is to balance exploitation (trying configurations near the best found so far) with exploration (trying configurations in regions we know little about).

The probabilistic model is typically a Gaussian process or a tree-structured Parzen estimator (TPE). The model outputs a predicted performance and an uncertainty estimate. The next configuration is chosen by maximizing an acquisition function that combines expected performance with uncertainty. Common acquisition functions include:

  • Expected Improvement (EI): the expected amount by which the next trial improves over the current best
  • Upper Confidence Bound (UCB): predicted performance plus a scaled uncertainty term
  • Thompson Sampling: sample one function from the posterior and maximize it

Each of these acquisition functions encodes a different strategy for balancing exploitation and exploration. EI focuses on configurations likely to beat the current best. UCB adds a bonus for poorly-explored regions. Thompson Sampling introduces randomness by sampling a plausible objective function and optimizing it, which naturally diversifies exploration without needing a manually tuned exploration coefficient.

Bayesian optimization tends to find good configurations in fewer evaluations than random search, particularly for expensive objectives. The cost is that each selection step requires fitting a surrogate model and solving an inner optimization. For problems where each evaluation costs hours, this overhead is negligible.

The limitation is sequential dependency: Bayesian optimization selects one configuration at a time based on previous results. In practice, most teams run experiments in parallel batches, which breaks the strict sequential assumption. Asynchronous variants like hyperopt and Optuna handle this by selecting multiple candidates using the current surrogate model without waiting for all running trials to finish. They use a "lie" strategy: they pretend that currently running trials have already returned some expected result and use this fictitious observation to select the next batch of candidates.

For LLM training, the high cost of individual evaluations makes Bayesian optimization appealing in principle, but the small number of feasible trials (often fewer than 30 at a given scale) limits the surrogate model's accuracy. In this regime, the advantage over random search is modest, and many practitioners use random search for simplicity with Bayesian optimization reserved for settings where even a few evaluations is prohibitively expensive.

Early Stopping and Successive Halving

Even with efficient search strategies, each trial trains a full model to completion before reporting its final performance. This is wasteful: a configuration with a learning rate ten times too large will diverge within the first thousand steps, yet a naive approach spends a full training run confirming that fact.

Early stopping cuts this waste by terminating runs that fall behind early. At fixed intervals during training, you compare all running configurations; those that perform significantly below the current best are terminated, and their compute is reallocated to the more promising survivors. This is not always safe (some good configurations start slowly before finding their footing) but works well when early performance correlates with final performance, which is generally true for learning rate sensitivity.

Successive halving formalizes early stopping into a systematic elimination protocol. It dramatically reduces cost by training all candidates for a short period, discarding the worst half, doubling the resources for the survivors, and repeating until one winner remains.

Formally, given nn initial configurations and a total budget BB, successive halving proceeds as:

Round 1: train all n configs for B/(nlog⁡2n) steps, keep top n/2Round 2: train surviving n/2 configs for 2B/(nlog⁡2n) steps, keep top n/4⋮Round log⁡2n: train final 2 configs, return winner\begin{aligned} &\text{Round 1: train all } n \text{ configs for } B/(n \log_2 n) \text{ steps, keep top } n/2 \\ &\text{Round 2: train surviving } n/2 \text{ configs for } 2B/(n \log_2 n) \text{ steps, keep top } n/4 \\ &\vdots \\ &\text{Round } \log_2 n: \text{ train final 2 configs, return winner} \end{aligned}

where:

  • nn: the number of initial configurations
  • BB: total compute budget in steps (or flops)
  • log⁡2n\log_2 n: number of elimination rounds

The assumption is that good configurations are identifiable early in training. This holds reasonably well for learning rate (divergent runs fail quickly) but less well for regularization (the effect of weight decay often only appears after substantial training). When performance rankings correlate between early and late training, successive halving finds nearly-optimal configurations with a fraction of the total compute.

Hyperband extends successive halving by running it with multiple values of nn, trading off between getting more configurations (wide, shallow search) and training each longer (narrow, deep search). With a large nn, you screen many candidates quickly but eliminate promising slow-starters too early. With a small nn, each survivor gets more training but you start with fewer candidates. Hyperband hedges by running multiple brackets simultaneously, each with a different nn, then selecting the best result across all brackets. This makes Hyperband more robust to the assumption that early performance predicts final performance.

The practical significance of Hyperband is substantial. In a controlled comparison, Hyperband with a total compute budget of BB frequently outperforms running random search for the same BB on a single configuration, because it can explore far more of the hyperparameter space in the early rounds while still committing full resources to the best candidates. For LLM pretraining where individual runs are expensive, Hyperband-style thinking should guide how you allocate your search budget.

Practical Search Protocol for LLMs

For large language model training, a practical search protocol takes into account that individual runs are expensive and that a small number of hyperparameters dominate performance:

  1. Identify proxy tasks: Use a small model (a few hundred million parameters at most) trained on a representative data sample for a small number of steps. If the proxy correlates with full-scale performance, this drastically reduces search cost.

  2. Fix cheap defaults first: Set architecture-linked hyperparameters (layer count, hidden size, attention heads) by scaling laws. Fix dropout to 0 for the pretraining phase. These are not search dimensions.

  3. Search learning rate: This is the most critical hyperparameter. Run random search over two to three orders of magnitude around a reasonable prior (discussed below in defaults). Use the proxy model and early stopping.

  4. Search weight decay: After finding a good learning rate, search weight decay in a one-dimensional sweep.

  5. Search batch size: Batch size interacts with learning rate. If you change batch size substantially, re-verify the learning rate.

  6. Verify on target scale: Confirm the best configuration found at proxy scale transfers to the target scale before committing to a full training run.

In[3]:
Code
from itertools import product

import numpy as np

# Simulate a grid search vs random search comparison
# We model the loss landscape as a function of learning rate and weight decay
# where only learning rate matters significantly

rng = np.random.default_rng(42)


def simulated_loss(lr, wd, noise_scale=0.05):
    """
    Simulated loss that depends strongly on lr and weakly on weight decay.
    Optimal lr is around 3e-4, optimal wd is around 1e-2.
    """
    lr_log = np.log10(lr)
    wd_log = np.log10(wd)
    lr_component = (lr_log - np.log10(3e-4)) ** 2 * 4.0
    wd_component = (wd_log - np.log10(1e-2)) ** 2 * 0.1  # weak dependence
    noise = rng.normal(0, noise_scale)
    return 2.5 + lr_component + wd_component + noise


# Grid search: 5 lrs x 5 wds = 25 evaluations
lr_grid = np.logspace(-5, -2, 5)
wd_grid = np.logspace(-4, -1, 5)

grid_results = []
for lr, wd in product(lr_grid, wd_grid):
    loss = simulated_loss(lr, wd)
    grid_results.append((lr, wd, loss))

grid_results.sort(key=lambda x: x[2])
best_grid = grid_results[0]
unique_lrs_explored_grid = len(set(r[0] for r in grid_results))

# Random search: 25 evaluations, same budget
lr_samples = np.power(10, rng.uniform(-5, -2, 25))
wd_samples = np.power(10, rng.uniform(-4, -1, 25))

random_results = []
for lr, wd in zip(lr_samples, wd_samples):
    loss = simulated_loss(lr, wd)
    random_results.append((lr, wd, loss))

random_results.sort(key=lambda x: x[2])
best_random = random_results[0]
unique_lrs_explored_random = len(set(f"{r[0]:.2e}" for r in random_results))
Out[4]:
Console
Grid search (25 evaluations):
  Best LR: 3.16e-04, Best WD: 3.16e-03, Loss: 2.5304
  Unique LR values explored: 5

Random search (25 evaluations):
  Best LR: 2.67e-04, Best WD: 4.96e-03, Loss: 2.5436
  Unique LR values explored: 24

Random search found a comparable configuration
Loss difference: 0.0132

Grid search fixes the values of each dimension at predetermined points. When the important dimension (learning rate) has its optimum between grid points, grid search can miss it entirely. Random search spreads evaluations across continuous ranges of both dimensions, exploring more of the learning rate axis and finding configurations closer to the true optimum. In practice, this advantage grows with the number of unimportant hyperparameters: each grid dimension you add multiplies compute requirements without helping you find the optimal value for the important dimensions.

Out[5]:
Visualization
Contour map of simulated validation loss with 25 grid-search evaluations arranged in five columns, a star marking the best grid point, and a vertical line at the true learning-rate optimum.
Grid search evaluates 25 combinations but samples only five distinct learning rates. The simulated optimum at log10(LR)=-3.2 falls between columns, so the best grid point remains visibly above the minimum loss.
Out[6]:
Visualization
Contour map of simulated validation loss with 25 randomly scattered evaluations, a star marking the best random point, and a vertical line at the true learning-rate optimum.
Random search uses the same 25-evaluation budget but samples 25 distinct learning rates. Its best point lands close to the simulated optimum and reaches a lower loss than the coarse grid.

Hyperparameter Interactions

Before diving into transfer and defaults, it is worth pausing on a nuance that the simple sensitivity analyses above tend to obscure: hyperparameters interact with each other in non-trivial ways. The optimal learning rate is not independent of batch size. The appropriate warmup duration depends on the schedule type. The effective regularization from weight decay depends on the total number of training steps. Understanding these interactions prevents you from fixing one hyperparameter at a suboptimal value while searching for another, which would cause you to consistently underestimate the true achievable performance.

Learning Rate and Batch Size

When you scale the batch size, the gradient at each step becomes an average over more samples. The variance of this gradient estimate decreases proportionally. A larger batch gives a more accurate estimate of the true gradient direction, which allows taking a larger step without overshooting. This motivates the linear scaling rule: if you double the batch size, doubling the learning rate maintains approximately the same effective noise-to-signal ratio in the gradient updates.

The linear scaling rule works reliably for moderate batch sizes (up to roughly 32,000 tokens per batch in LLM training) but breaks down at very large batches. When the batch is so large that the gradient estimate is nearly noise-free, the model is in the data-limited regime and increasing the batch further does not help, so scaling the learning rate linearly would make steps too large. The square root scaling rule, which multiplies the learning rate by k\sqrt{k} when scaling the batch by kk, is more conservative and tends to be safer in this regime. Most practitioners use the linear rule up to the critical batch size identified by Kaplan et al. (2020), then switch to the square root rule for larger batches.

Warmup Duration and Schedule Type

The role of warmup is to stabilize the Adam optimizer's second moment estimate before taking large steps. At the beginning of training, the second moment estimate is initialized to zero and has seen very few gradient samples. If the learning rate is already at its peak value during this initialization phase, the effective step sizes for parameters with small gradient history can be wildly large, causing early instability. Warmup suppresses this by holding the learning rate low while the second moment estimate builds up.

The appropriate warmup duration depends on 1/(1−β2)1/(1 - \beta_2), the effective window of the second moment estimate. With β2=0.999\beta_2 = 0.999, the window spans approximately 1000 steps, so roughly 1000 steps of warmup are needed for the estimate to stabilize. With β2=0.95\beta_2 = 0.95, the window spans only 20 steps, so much shorter warmup suffices. Using β2=0.95\beta_2 = 0.95 (the LLaMA default) with a 2000-step warmup provides comfortable margin; using β2=0.999\beta_2 = 0.999 with only 500 steps of warmup risks early instability.

This interaction is often overlooked when copying hyperparameters from one recipe to another. If you take the LLaMA recipe (β2=0.95\beta_2 = 0.95, 2000 steps warmup) and change only β2\beta_2 to 0.9990.999 (perhaps because your framework uses it as default), you should also increase the warmup duration to match.

Weight Decay and Training Duration

Weight decay applies a shrinkage force proportional to each parameter's current value at every step. Over a training run of TT steps with learning rate η\eta and weight decay coefficient λ\lambda, a parameter that receives no gradient update will be multiplied by (1−ηλ)T(1 - \eta \lambda)^T total. With η=3×10−4\eta = 3 \times 10^{-4}, λ=0.1\lambda = 0.1, and T=100,000T = 100{,}000 steps:

(1−ηλ)T=(1−3×10−5)100000≈e−3≈0.050(1 - \eta \lambda)^T = (1 - 3 \times 10^{-5})^{100000} \approx e^{-3} \approx 0.050

A weight that receives no gradient signal will shrink to about 5% of its initial value over this training run. This is a powerful shrinkage. For longer runs or higher weight decay, the shrinkage is even more dramatic. If you extend a training run from 100K to 500K steps without adjusting weight decay, the effective regularization becomes five times stronger. This can suppress the model's capacity to represent certain patterns. The implication is that when you change training duration significantly, you should re-examine the weight decay setting. A run five times longer than a reference run may benefit from weight decay reduced by roughly a factor of five.

Hyperparameter Transfer

Even efficient search strategies are costly when experiments take days. Hyperparameter transfer sidesteps the problem by asking: can we find good hyperparameters on a small model and use them for a large model without re-searching?

The intuitive hope is that the optimal learning rate for a 125M parameter model should tell us something about the optimal learning rate for a 1.3B parameter model. This hope is sometimes true and sometimes badly wrong, depending on the hyperparameter and the scaling strategy.

The Transfer Problem

Standard training wisdom has been that hyperparameters do not transfer cleanly across model scales. The commonly observed phenomenon is that larger models require smaller learning rates. Practitioners scaling from 125M to 1.3B to 6.7B parameters would reduce the learning rate at each step, often by trial and error. This made scaling expensive: every new model size required its own hyperparameter search.

The root cause is that standard parameter initialization does not have stable statistics across width. When you double the width of a layer (its hidden dimension), the total gradient flowing through that layer changes in a way that depends on the initialization. If initialization variance does not scale correctly with width, the effective learning rate at a neuron level changes even when the optimizer's nominal learning rate stays the same.

To see why, think about what happens when a layer receives a gradient and updates its weights. The parameter update changes the layer's outputs, and the size of that output change determines how much the model changes at each step. If the output change is too large relative to the current activations, training becomes unstable. If it is too small, learning is inefficiently slow. Maintaining a stable ratio between the update magnitude and the activation magnitude is what makes the learning rate transferable.

Maximal Update Parametrization (muP)

Maximal Update Parametrization (muP), introduced by Yang et al. (2022), addresses this directly. It is a specific set of initialization scales and learning rate scaling rules that makes feature updates approximately constant as model width increases. The key insight is:

Maximal Update Parametrization (muP)

muP is a reparametrization of neural network weights that ensures the optimal learning rate is the same across model widths. With standard parametrization (SP), the optimal learning rate decreases as width grows. With muP, the optimal learning rate is width-independent, enabling hyperparameter transfer from small to large models.

To understand why this matters, consider a single linear layer with input dimension dind_{in}, output dimension doutd_{out}, and weight matrix W∈Rdout×dinW \in \mathbb{R}^{d_{out} \times d_{in}}. In standard Xavier initialization, each weight is drawn from:

Wij∼N(0,2din+dout)W_{ij} \sim \mathcal{N}\left(0, \frac{2}{d_{in} + d_{out}}\right)

where:

  • WijW_{ij}: the (i,j)(i,j) entry of the weight matrix
  • dind_{in}: input dimension of the layer
  • doutd_{out}: output dimension of the layer

When a gradient update of magnitude gg is applied, the change in the output activations is approximately:

Δh≈η⋅g⋅din\Delta \mathbf{h} \approx \eta \cdot g \cdot \sqrt{d_{in}}

where:

  • Δh\Delta \mathbf{h}: change in output activations
  • η\eta: learning rate
  • gg: gradient magnitude per weight

As dind_{in} grows, the effective activation update magnitude grows as din\sqrt{d_{in}}, even at the same nominal learning rate. To keep activation updates stable, the learning rate must shrink as 1/din1/\sqrt{d_{in}}. This is the source of the need to retune learning rates when changing width.

muP resolves this by modifying the initialization and the effective learning rate applied to each parameter type:

  • Input layer weights: Scale initialization by 1/din1/d_{in}, use learning rate η\eta
  • Hidden layer weights: Scale initialization by 1/din1/\sqrt{d_{in}}, scale learning rate by 1/dwidth1/d_{width} where dwidthd_{width} is the hidden dimension
  • Output layer weights: Scale initialization by 1/dout1/d_{out}, use a smaller learning rate η/dwidth\eta / d_{width}

Under muP, the contribution of each weight to the output remains constant in expectation as width grows, and the optimal learning rate becomes width-invariant. This means:

  1. Tune hyperparameters (especially learning rate) on a small proxy model using muP
  2. Scale up to the target model width using muP
  3. Use the same learning rate without retuning
In[7]:
Code
import numpy as np

# Demonstrate how initialization variance changes with width under SP vs muP


def sp_init_std(d_in, d_out):
    """Standard Xavier initialization std."""
    return np.sqrt(2.0 / (d_in + d_out))


def mup_hidden_init_std(d_in, d_out):
    """muP hidden layer initialization: 1/sqrt(d_in)."""
    return 1.0 / np.sqrt(d_in)


def effective_activation_change_sp(d_in, d_out, lr=1e-3, grad_magnitude=1.0):
    """
    Approximate magnitude of activation change per step under SP.
    This simplified scaling model isolates the sqrt(d_in) growth described
    above; it is not a full optimizer simulation.
    """
    return lr * grad_magnitude * np.sqrt(d_in)


def effective_activation_change_mup(
    d_in, d_out, base_width, lr=1e-3, grad_magnitude=1.0
):
    """
    Approximate width-invariant activation change under muP, anchored to the
    same value as SP at base_width.
    """
    return lr * grad_magnitude * np.sqrt(base_width)


# Vary width from 128 to 4096
widths = [128, 256, 512, 1024, 2048, 4096]
base_width = 128
lr = 1e-3

sp_changes = [effective_activation_change_sp(w, w, lr=lr) for w in widths]
mup_changes = [
    effective_activation_change_mup(w, w, base_width, lr=lr) for w in widths
]
Out[8]:
Console
   Width   SP activation change   muP activation change
--------------------------------------------------------
     128               0.011314                0.011314
     256               0.016000                0.011314
     512               0.022627                0.011314
    1024               0.032000                0.011314
    2048               0.045255                0.011314
    4096               0.064000                0.011314

SP ratio (width 4096 vs 128): 5.66x
muP ratio (width 4096 vs 128): 1.00x

Under standard parametrization, the effective activation change grows with width even at the same learning rate. Under muP, it stays constant regardless of width, which is why the same learning rate works across scales.

Out[9]:
Visualization
Line plot on a logarithmic width axis. The standard-parametrization curve rises from below one to about 5.7, while the muP curve remains flat at one.
A simplified scaling model compares activation updates relative to width 128. Standard parametrization grows with the square root of width, reaching 5.7 times the baseline at width 4096, while maximal update parametrization (muP) remains constant by construction.

A Worked muP Transfer Example

To make the muP workflow concrete, consider a team training a 1.3B parameter model and wanting to find the optimal learning rate efficiently. Without muP, they would need to run several training experiments directly at 1.3B scale, each costing substantial GPU time. With muP, they can work as follows.

First, they select a proxy width of 256 (versus the target hidden dimension of 2048, a ratio of 8). They train the proxy for 5,000 steps (proportional to the proxy's compute-optimal budget) with learning rates spanning {3×10−4, 6×10−4, 1×10−3, 2×10−3, 4×10−3}\{3 \times 10^{-4},\ 6 \times 10^{-4},\ 1 \times 10^{-3},\ 2 \times 10^{-3},\ 4 \times 10^{-3}\}. They initialize using muP rules: hidden layer weights scaled by 1/din1/\sqrt{d_{in}}, output layer weights scaled by 1/dwidth1/d_{width}.

After the proxy sweep, the optimal learning rate at width 256 is found to be 1×10−31 \times 10^{-3}. Under muP, this is the same optimal learning rate that should work at width 2048. They set the 1.3B model's learning rate to 1×10−31 \times 10^{-3} and begin the full training run. The first few thousand steps confirm that loss decreases smoothly, consistent with the proxy's behavior. No retuning at the 1.3B scale was needed.

Without muP, scaling from width 256 to 2048 (an 8x increase) under standard Xavier initialization would require reducing the learning rate by roughly 8≈2.8×\sqrt{8} \approx 2.8 \times to account for the activation update scaling. The team would either guess this scaling (risking a suboptimal learning rate) or run additional full-scale experiments (expensive). muP eliminates this uncertainty.

What Transfers and What Does Not

Hyperparameter transfer via muP works primarily for learning rate across model width. Other hyperparameters have different transfer properties, and understanding these differences prevents false confidence about which settings you can simply copy from a proxy run.

  • Learning rate (with muP): transfers almost perfectly across width. Transfer across depth is less reliable but generally acceptable within a 2x depth range. If you double both width and depth, transfer the learning rate for the width change and monitor carefully for the depth change.
  • Batch size: transfers well when adjusting learning rate proportionally. Two scaling rules are in common use. The linear scaling rule (double batch size, double learning rate) holds in the data-rich, short-horizon regime and was empirically validated for ImageNet training by Goyal et al. (2017). The square root scaling rule (double batch size, multiply learning rate by 2\sqrt{2}) is more conservative and tends to be safer for very large batch sizes or longer training runs where noise reduction from larger batches matters more.
  • Warmup steps: does not transfer directly as an absolute count. A 1000-step warmup for a 10,000-step proxy run represents 10% warmup, far too long for a 100,000-step target run. Express warmup as a fraction of total steps (e.g., 2% warmup) so it transfers naturally.
  • Weight decay: generally transfers well. The optimal weight decay coefficient is relatively insensitive to model width and moderately insensitive to scale. It is more sensitive to dataset size: when fine-tuning or domain-adapting on a small dataset, the optimal weight decay is often lower than the pretraining default.
  • Beta parameters (β1\beta_1, β2\beta_2 for Adam): very robust; near-universal defaults of 0.90.9 and values between 0.950.95 and 0.9990.999 apply almost everywhere.
  • Gradient clipping threshold: transfers moderately. Larger models can produce larger gradient norms during early training, especially without muP. The gradient norm typically stabilizes after warmup; monitoring it during the first few thousand steps reveals whether your clipping threshold is too aggressive or too permissive.

The practical implication is that for a new model, you should:

  1. Use muP if you plan to search learning rate (it is increasingly the default in research settings)
  2. Search learning rate and weight decay on a proxy model at 1/10th to 1/100th the target parameter count
  3. Use principled scaling rules for batch size
  4. Keep beta parameters at universal defaults
  5. Monitor gradient norms to set a clipping threshold

Critical vs. Robust Hyperparameters

Not all hyperparameters are equal in their impact on model quality. Understanding which hyperparameters are critical (meaning small deviations from the optimal value cause large performance degradation) versus robust (meaning performance is nearly flat across a wide range) tells you where to invest search effort.

A Taxonomy of Sensitivity

Critical hyperparameters are those where the performance landscape has a sharp peak. The loss drops steeply on either side of the optimum. For these, you need a precise value or at least a value within a narrow band.

Reliable hyperparameters are those where the performance field is nearly flat. The loss changes little across orders of magnitude. For these, a reasonable default suffices and extensive search is wasteful.

Between these extremes are hyperparameters that matter moderately: performance degrades gradually as you move away from the optimum, but you can tolerate some imprecision.

The taxonomy is:

Critical (search carefully):

  • Learning rate: The most important hyperparameter in practice. Performance is highly sensitive to learning rate, with an optimum that can be 2-3 orders of magnitude wide in log-space but still has a factor-of-2 sensitivity near the edges. A learning rate that is 5x too large often causes divergence; one that is 5x too small trains stably but reaches poor final performance.
  • Learning rate schedule: The shape of the decay curve significantly affects final performance. Cosine decay outperforms step decay and linear decay in most settings; using no decay at all loses 5-15% relative perplexity.

Moderate sensitivity (use informed defaults, search if budget allows):

  • Weight decay: Performance degrades smoothly outside the optimal range. Values between 0.010.01 and 0.10.1 are generally safe; values below 0.0010.001 often lead to overfitting and values above 0.30.3 can hurt convergence.
  • Warmup duration: A small amount of warmup (1-5% of total steps) matters; the precise value within this range does not.
  • Batch size: Affects sample efficiency but not final performance dramatically if learning rate is adjusted. Too small (below effective batch size) hurts; too large wastes compute without benefit.

Robust (use universal defaults):

  • Adam β1\beta_1: Values between 0.850.85 and 0.950.95 all work nearly identically. Default 0.90.9 is fine everywhere.
  • Adam β2\beta_2: Values between 0.950.95 and 0.9990.999 all work nearly identically. Default 0.9990.999 is the most common; 0.950.95 is used for training stability.
  • Adam ε\varepsilon: Has essentially no effect on performance across a wide range (10−810^{-8} to 10−610^{-6}) except at extreme values.
  • Gradient clipping threshold (when using a reasonable value): Values between 0.50.5 and 5.05.0 give nearly identical results.
In[10]:
Code
import numpy as np

rng = np.random.default_rng(0)


def model_perplexity(lr, wd, beta2, warmup_frac, noise=0.02):
    """
    Simulate how final model perplexity depends on different hyperparameters.
    Returns a perplexity value (lower is better).

    Critical: lr (sharp optimum at log10=-3.5)
    Moderate: wd (optimum at log10=-2, gradual)
    Robust: beta2 (nearly flat from 0.95 to 0.999)
    Moderate: warmup_frac (flat from 0.01 to 0.05)
    """
    # Learning rate component: sharp quadratic in log space
    lr_opt = np.log10(3e-4)
    lr_comp = 15.0 * (np.log10(lr) - lr_opt) ** 2

    # Weight decay component: moderate quadratic in log space
    wd_opt = np.log10(1e-2)
    wd_comp = 2.0 * (np.log10(wd) - wd_opt) ** 2

    # Beta2 component: very flat
    beta2_comp = 0.3 * (beta2 - 0.999) ** 2 * 100

    # Warmup component: moderate but only for very low values
    warmup_comp = 1.5 * max(0, 0.005 - warmup_frac) ** 2 * 10000

    base_perplexity = 20.0
    noise_term = rng.normal(0, noise)
    return (
        base_perplexity
        + lr_comp
        + wd_comp
        + beta2_comp
        + warmup_comp
        + noise_term
    )


# Vary each hyperparameter while keeping others at their optimum
lr_values = np.logspace(-5, -2, 50)
wd_values = np.logspace(-4, -1, 50)
beta2_values = np.linspace(0.90, 0.999, 50)
warmup_values = np.linspace(0.001, 0.10, 50)

lr_perplexities = [
    model_perplexity(lr=v, wd=1e-2, beta2=0.999, warmup_frac=0.02)
    for v in lr_values
]
wd_perplexities = [
    model_perplexity(lr=3e-4, wd=v, beta2=0.999, warmup_frac=0.02)
    for v in wd_values
]
beta2_perplexities = [
    model_perplexity(lr=3e-4, wd=1e-2, beta2=v, warmup_frac=0.02)
    for v in beta2_values
]
warmup_perplexities = [
    model_perplexity(lr=3e-4, wd=1e-2, beta2=0.999, warmup_frac=v)
    for v in warmup_values
]

# Compute sensitivity: range of perplexity change
lr_range = max(lr_perplexities) - min(lr_perplexities)
wd_range = max(wd_perplexities) - min(wd_perplexities)
beta2_range = max(beta2_perplexities) - min(beta2_perplexities)
warmup_range = max(warmup_perplexities) - min(warmup_perplexities)
Out[11]:
Console
Hyperparameter sensitivity (perplexity range across full sweep):

  Learning rate:   34.79 perplexity points  -> CRITICAL
  Weight decay:    8.03 perplexity points  -> MODERATE
  Beta2:           0.34 perplexity points  -> ROBUST
  Warmup fraction: 0.28 perplexity points  -> MODERATE

Search priority: focus compute on critical hyperparameters.

The sensitivity analysis confirms the intuition: learning rate has the largest impact on performance, weight decay and warmup fraction have moderate impacts, and beta2 has minimal impact. This tells you how to allocate your search budget.

Out[12]:
Visualization
Two perplexity curves over log10 hyperparameter values. The learning-rate curve forms a steep U shape around minus 3.5, while the weight-decay curve forms a broader, shallower U around minus 2.
Learning rate produces a sharp perplexity minimum near 3e-4, while weight decay has a much shallower basin near 1e-2. Moving the learning rate tenfold from its optimum costs roughly 15 perplexity points, making it the higher-priority search dimension.
Out[13]:
Visualization
Dual-axis line plot comparing beta2 and warmup sweeps. The beta2 curve changes gradually by less than half a perplexity point, and the warmup curve quickly flattens after 0.5 percent.
Adam beta2 changes simulated perplexity by less than half a point across the displayed range, while warmup incurs a penalty only below 0.5 percent of training. The broad flat regions explain why defaults are usually sufficient for both parameters.

Sensitivity and Model Scale

Sensitivity patterns are not always constant across model scale. Some observations from large-scale training:

  • Learning rate sensitivity increases with model scale. Larger models are more sensitive to suboptimal learning rates, partly because their optimization landscapes have more complex structure.
  • Weight decay sensitivity also increases somewhat with scale. Larger models have more capacity to overfit, making regularization more important.
  • Beta2 sensitivity is scale-invariant in practice. The same defaults work from small models to the largest trained.
  • Dropout sensitivity follows a U-curve: too little and the model overfits; too much and capacity is wasted. The optimal dropout rate tends to increase with model size but is relatively robust within a factor of 2.

The implication for your search workflow is to concentrate your compute on the learning rate first, then weight decay, and treat everything else as fixed defaults. Within a 5x learning rate range around your prior, you should be able to find a near-optimal value in ten to twenty proxy experiments. Within a 3x weight decay range around 0.10.1, five experiments usually suffice. You will recover far more performance from careful learning rate tuning than from fine-grained optimization of beta2 or epsilon.

Default Recipes

The research community has converged on several default recipes for LLM training. These represent the distilled experience of many training runs and are reliable starting points for new models. You will still need to tune learning rate and verify on your specific setup, but using these defaults dramatically narrows the search space.

The GPT-Style Training Recipe

The configuration used in GPT-3 and refined in subsequent work has become the de facto standard for autoregressive LLM pretraining:

  • Optimizer: Adam with β1=0.9\beta_1 = 0.9, β2=0.95\beta_2 = 0.95, ε=10−8\varepsilon = 10^{-8}
  • Learning rate: peak LR in the range [10−4,3×10−4][10^{-4}, 3 \times 10^{-4}] depending on model size, scaled by 1/dmodel1/\sqrt{d_{model}} approximately
  • Learning rate schedule: cosine decay to 10% of peak LR
  • Warmup: linear warmup for 1% to 4% of total training steps (up to a few thousand steps)
  • Weight decay: 0.10.1 applied to all non-bias, non-embedding parameters
  • Gradient clipping: 1.0 (global gradient norm)
  • Batch size: large, typically 0.5M to 4M tokens per batch
  • Dropout: 0 during pretraining (all regularization via weight decay and data diversity)

Note the choice of β2=0.95\beta_2 = 0.95 rather than the Adam default of 0.9990.999. This reflects the observation that large-batch LLM training benefits from a shorter memory in the second moment estimate. With β2=0.999\beta_2 = 0.999, the second moment estimate is dominated by a long history of gradients, which can cause the optimizer to respond too slowly to sudden changes in gradient magnitude during training. The value 0.950.95 gives a more responsive effective window of approximately 1/(1−0.95)=201/(1 - 0.95) = 20 steps rather than 1000 steps.

The decision to set dropout to zero during pretraining is deliberate. Dropout was a critical regularizer in the pre-LLM era, when models were trained on relatively small, task-specific datasets. For LLM pretraining on trillions of tokens from diverse sources, data diversity itself provides strong implicit regularization. Adding dropout during pretraining hurts sample efficiency without providing meaningful overfitting protection. Dropout is sometimes reintroduced during fine-tuning, where the dataset is smaller and overfitting becomes a real concern.

The Chinchilla Scaling Recipe

The Chinchilla paper (Hoffmann et al., 2022) contributed scaling laws and practical training defaults for compute-optimal models:

  • Dataset size: train for roughly 20×N20 \times N tokens where NN is the model size in parameters
  • Learning rate: similar to GPT-style but with slightly lower weight decay (0.010.01 to 0.10.1 depending on scale)
  • Cooldown: explicit cooldown phase of 10% of total steps before linear decay ends

The important takeaway from Chinchilla for hyperparameter selection is that dataset size and model size are themselves hyperparameters that interact. Under-training a large model (fewer tokens than compute-optimal) or over-training a small model both give worse performance per flop.

The Chinchilla result says both the number of parameters NN and the number of training tokens DD should scale equally with compute. Empirically, D≈20ND \approx 20N (train on 20 tokens per parameter). The compute cost of training a transformer is approximately C≈6NDC \approx 6ND FLOPs (roughly 6 multiply-accumulates per parameter per token). Substituting D=20ND = 20N:

C≈6ND=6N⋅20N=120N2N∗=C120\begin{aligned} C &\approx 6 N D = 6 N \cdot 20N = 120 N^2 \\ N^* &= \sqrt{\frac{C}{120}} \end{aligned}

where:

  • N∗N^*: optimal model size in parameters for a given compute budget
  • D∗D^*: optimal number of training tokens, approximately 20N∗20N^*
  • CC: total compute budget in floating-point operations (FLOPs)

(The exact constant varies by source; the Chinchilla paper uses coefficients derived from empirical fits. The key insight is that N∗N^* scales as C\sqrt{C}, meaning you should grow both model size and data together as compute increases.)

For a given compute budget, training a model with N∗N^* parameters is more efficient than training a larger model for fewer steps or a smaller model for more steps.

The Chinchilla compute-optimal recipe optimizes training loss at a fixed compute budget. In deployment, the cost of inference often dominates the cost of training: a model that you run millions of times at inference may justify training for longer (more tokens than Chinchilla-optimal) to push down the model size needed to achieve a target quality. This trade-off has led several organizations to train smaller models for longer than the Chinchilla recipe recommends, accepting higher training cost in exchange for lower inference cost.

LLaMA-Style Training

The LLaMA family (Touvron et al., 2023) refined the GPT recipe with several changes that have become widely adopted:

  • Optimizer: AdamW (weight decay decoupled from the gradient update, as discussed in the Weight Decay chapter)
  • Learning rate: 3×10−43 \times 10^{-4} for 7B, 1.5×10−41.5 \times 10^{-4} for 65B
  • Schedule: cosine decay to 10−510^{-5} (nearly to zero)
  • Warmup: 2000 steps (roughly 0.4% of total training steps for a 500B token run)
  • Weight decay: 0.10.1
  • Gradient clipping: 1.01.0
  • Batch size: 4M tokens for 65B model

The shift to AdamW from Adam with standard weight decay matters primarily at this scale. In standard Adam, weight decay interacts with the second moment estimate, effectively reducing decay for parameters with large gradients. AdamW decouples the two, applying weight decay uniformly regardless of gradient history. This produces more consistent regularization.

LLaMA's decision to decay to nearly zero (final LR of 10−510^{-5}, which is about 3% of the peak 3×10−43 \times 10^{-4}) rather than the GPT convention of stopping at 10% of peak represents a more aggressive cooldown. The motivation is that the final phase of cosine decay, where the learning rate is very small, allows the model to make fine adjustments without large perturbations. In practice, the difference between decaying to 5% and 10% of peak is modest, but decaying to zero or near-zero has been consistently observed to improve final performance slightly.

Learning Rate Priors by Model Size

A practical prior for learning rate based on observed compute-optimal configurations:

Peak learning rate priors by model size for autoregressive pretraining with cosine schedule.
Model sizeTypical peak LRRange
125M parameters6×10−46 \times 10^{-4}[3×10−4, 1×10−3][3 \times 10^{-4},\ 1 \times 10^{-3}]
350M parameters4×10−44 \times 10^{-4}[2×10−4, 8×10−4][2 \times 10^{-4},\ 8 \times 10^{-4}]
1.3B parameters3×10−43 \times 10^{-4}[1×10−4, 5×10−4][1 \times 10^{-4},\ 5 \times 10^{-4}]
6.7B parameters2×10−42 \times 10^{-4}[8×10−5, 4×10−4][8 \times 10^{-5},\ 4 \times 10^{-4}]
30B parameters1.5×10−41.5 \times 10^{-4}[5×10−5, 3×10−4][5 \times 10^{-5},\ 3 \times 10^{-4}]
65B parameters1×10−41 \times 10^{-4}[4×10−5, 2×10−4][4 \times 10^{-5},\ 2 \times 10^{-4}]

The trend is clear: optimal learning rate decreases slowly with model size in log-scale, roughly halving for every 10x increase in parameters. This relationship emerges from the combination of increased gradient sensitivity and the need for stable optimization at large scale.

In[14]:
Code
import numpy as np

# Model sizes and their typical optimal learning rates (from literature)
model_sizes_billions = np.array([0.125, 0.35, 1.3, 6.7, 30.0, 65.0])
optimal_lrs = np.array([6e-4, 4e-4, 3e-4, 2e-4, 1.5e-4, 1e-4])

# Fit a power law: lr = a * N^b in log-log space
log_sizes = np.log10(model_sizes_billions * 1e9)  # convert to parameter count
log_lrs = np.log10(optimal_lrs)

# Linear regression in log-log space
coeffs = np.polyfit(log_sizes, log_lrs, 1)
slope, intercept = coeffs

# Predictions from the fit
predicted_lrs = 10 ** (slope * log_sizes + intercept)

# Estimate LR for hypothetical 175B model
lr_175b = 10 ** (slope * np.log10(175e9) + intercept)
Out[15]:
Console
Power law fit: LR ≈ 7.6818e-02 × N^-0.2634
(where N is model size in parameters)

Fit quality:
   0.125B params: observed 6.0e-04, predicted 5.7e-04, ratio 1.06
   0.350B params: observed 4.0e-04, predicted 4.3e-04, ratio 0.93
   1.300B params: observed 3.0e-04, predicted 3.1e-04, ratio 0.98
   6.700B params: observed 2.0e-04, predicted 2.0e-04, ratio 1.01
  30.000B params: observed 1.5e-04, predicted 1.3e-04, ratio 1.12
  65.000B params: observed 1.0e-04, predicted 1.1e-04, ratio 0.92

Extrapolation to 175B: estimated LR ≈ 8.40e-05

The power law relationship between model size and optimal learning rate is not a fundamental law, but an empirical regularity from training runs across many model sizes. Using it as a prior saves you from searching a completely uninformed range when starting with a new model size.

Out[16]:
Visualization
Log-log scatter plot of peak learning-rate priors against model parameter count, with a descending power-law fit and a shaded factor-of-three search band.
Illustrative peak learning-rate priors decline with model size and follow an approximate power law. The shaded factor-of-three band is a practical initial search range rather than a confidence interval.
Out[17]:
Visualization
Learning-rate curve that rises linearly during a 2,000-step warmup and then follows a smooth cosine decay to 10 percent of its peak by step 100,000.
A 100,000-step cosine schedule warms linearly to 3e-4 over the first 2,000 steps, then decays smoothly to 10 percent of the peak. The marked boundaries make the warmup and minimum-rate choices explicit.

The Weight Decay Default

The choice of weight decay 0.10.1 in GPT-style training deserves explanation. Weight decay at this scale is applied to weight matrices only (not biases, not embedding tables, not layer norm parameters). At this level:

  • It provides moderate regularization without strongly suppressing parameter magnitudes
  • It interacts well with the cosine schedule (parameters shrink toward zero as the learning rate decays)
  • It is large enough to provide a meaningful regularization signal on large models

Lower values (0.010.01) are used when the dataset is very large relative to model size (less risk of overfitting). Higher values (0.30.3) are sometimes used for smaller models or fine-tuning scenarios.

The question of which parameters receive weight decay has a principled answer that goes beyond arbitrary convention. Biases and layer normalization parameters are scalar offsets and scales that adjust the mean and variance of activations. Applying weight decay to them would bias the model toward activations with zero mean and unit variance regardless of what the data suggests, which can impede the model's ability to represent the appropriate activation statistics for a given layer. Embedding parameters represent specific token identities; shrinking them toward zero would make all token embeddings converge and lose their distinctiveness. By restricting weight decay to weight matrices, you regularize the relational structure of the model (how inputs are transformed) without constraining its representational offsets.

Fine-Tuning Hyperparameter Recipes

Pretraining and fine-tuning have fundamentally different hyperparameter needs, and applying pretraining defaults to fine-tuning is one of the most common mistakes in practice. Understanding why they differ helps you adapt appropriately.

During pretraining, the model is learning from scratch on a vast and diverse dataset. The learning rate can be relatively high because the model has no existing knowledge to protect, and large gradient steps efficiently move parameters from their random initial positions to useful regions of weight space. Weight decay at 0.10.1 is safe because with trillions of tokens, the risk of catastrophic overfitting is low.

During fine-tuning on a task-specific dataset (which might contain thousands to millions of examples rather than trillions), the situation is inverted. The model already has useful representations from pretraining. Large gradient steps destroy these representations rather than refining them. High weight decay accelerates this destruction by actively pushing weights toward zero.

The community consensus for supervised fine-tuning (SFT) and similar adaptation tasks:

  • Learning rate: 1/10 to 1/100 of the pretraining learning rate. For a model pretrained at 3×10−43 \times 10^{-4}, use 3×10−53 \times 10^{-5} to 3×10−63 \times 10^{-6} for fine-tuning. Lower rates better preserve pretraining knowledge.
  • Schedule: cosine or linear decay over the fine-tuning steps. Total steps are often much fewer (1-3 epochs over the fine-tuning dataset rather than a fixed token budget).
  • Warmup: shorter than pretraining, often just 1-3% of fine-tuning steps or a fixed 50-100 steps. The model is already well-initialized, so the second moment estimate stabilizes quickly.
  • Weight decay: lower than pretraining, often 0.010.01 or even 0.00.0. Aggressive weight decay during fine-tuning can erase the pretrained representations that make the model useful.
  • Dropout: may be reintroduced at low rates (0.050.05 to 0.10.1) when the fine-tuning dataset is very small and overfitting is a concern.
  • Batch size: smaller than pretraining, often 16 to 128 samples. Smaller batches increase gradient noise, which provides implicit regularization and can help the model adapt to the new distribution without memorizing.

For parameter-efficient fine-tuning methods like LoRA, the learning rate for the trainable adapters is typically higher than for full fine-tuning, often back in the 1×10−41 \times 10^{-4} range, because the adapters are small and randomly initialized relative to the frozen pretrained weights. The frozen base model provides stable feature extraction while the adapters learn task-specific modifications.

Implementing a Hyperparameter Search Workflow

A complete hyperparameter search workflow for a new LLM training run combines the elements above into a practical sequence:

In[18]:
Code
from dataclasses import dataclass
from typing import List

import numpy as np


@dataclass
class HPConfig:
    """Complete hyperparameter configuration for an LLM training run."""

    # Optimizer
    learning_rate: float
    beta1: float = 0.9
    beta2: float = 0.95
    epsilon: float = 1e-8
    weight_decay: float = 0.1

    # Schedule
    warmup_steps: int = 2000
    total_steps: int = 100000
    min_lr_ratio: float = 0.1  # min_lr = peak_lr * min_lr_ratio

    # Gradient clipping
    max_grad_norm: float = 1.0

    # Architecture-linked (not searched)
    model_size_billions: float = 1.3

    def validate(self) -> List[str]:
        """Check configuration for common mistakes."""
        warnings = []

        # LR sanity check
        size_params = self.model_size_billions * 1e9
        expected_lr = 1.2e-3 * (size_params**-0.07)  # rough power law
        lr_ratio = self.learning_rate / expected_lr
        if lr_ratio > 5:
            warnings.append(
                f"LR {self.learning_rate:.2e} is {lr_ratio:.1f}x above size-expected {expected_lr:.2e}"
            )
        elif lr_ratio < 0.2:
            warnings.append(
                f"LR {self.learning_rate:.2e} is {1 / lr_ratio:.1f}x below size-expected {expected_lr:.2e}"
            )

        # Warmup as fraction of total
        warmup_frac = self.warmup_steps / self.total_steps
        if warmup_frac < 0.005:
            warnings.append(
                f"Warmup fraction {warmup_frac:.3f} is very short; recommend at least 0.01"
            )
        elif warmup_frac > 0.10:
            warnings.append(
                f"Warmup fraction {warmup_frac:.3f} is unusually long; typical is 0.01-0.05"
            )

        # Weight decay range
        if self.weight_decay < 0.001:
            warnings.append(
                f"Weight decay {self.weight_decay} is very small; may cause overfitting"
            )
        elif self.weight_decay > 0.5:
            warnings.append(
                f"Weight decay {self.weight_decay} is very large; may hurt convergence"
            )

        return warnings


def cosine_lr_schedule(step: int, config: HPConfig) -> float:
    """Compute learning rate at a given step."""
    if step < config.warmup_steps:
        return config.learning_rate * (step / config.warmup_steps)

    progress = (step - config.warmup_steps) / (
        config.total_steps - config.warmup_steps
    )
    cosine_factor = 0.5 * (1 + np.cos(np.pi * progress))
    min_lr = config.learning_rate * config.min_lr_ratio
    return min_lr + (config.learning_rate - min_lr) * cosine_factor


# Example: configure a 1.3B model run
config_1b = HPConfig(
    learning_rate=3e-4,
    model_size_billions=1.3,
    warmup_steps=2000,
    total_steps=100000,
)

# Test LR schedule across training
steps = np.arange(0, config_1b.total_steps + 1, 100)
lr_values = [cosine_lr_schedule(s, config_1b) for s in steps]

warnings_1b = config_1b.validate()
Out[19]:
Console
HPConfig for 1.3B model:
  Learning rate:   3.00e-04
  Weight decay:    0.1
  Beta1, Beta2:    0.9, 0.95
  Warmup steps:    2000 (2.0% of total)
  Total steps:     100,000
  Min LR:          3.00e-05

Configuration: no warnings.

LR at step 0:                    0.00e+00
LR at end of warmup (2000): 3.00e-04
LR at 50% of training:           1.69e-04
LR at 90% of training:           3.69e-05
LR at final step:                3.00e-05

The configuration object makes hyperparameter relationships explicit and enables validation of common mistakes before training begins. The schedule shows the warmup phase rising linearly to the peak, followed by cosine decay to 10% of the peak value.

Key Parameters

The primary hyperparameters for an LLM training run and their recommended defaults are:

  • learning_rate: The most important parameter. Start with the size-based prior from Table lr-by-size; search within a factor of 3x around this prior.
  • beta2: Use 0.950.95 for large-scale training; 0.9990.999 for smaller models. This is robust within each setting.
  • weight_decay: Default 0.10.1 for pretraining. Lower (0.010.01) for data-rich scenarios; higher (0.30.3) for fine-tuning on small datasets.
  • warmup_steps: 1-4% of total steps, or a fixed 2000 steps for long runs. Very robust to exact value within this range.
  • max_grad_norm: 1.0. Monitor actual gradient norms during training; if they regularly exceed 2x the clip threshold, investigate instability.
  • min_lr_ratio: 0.1 (cosine decays to 10% of peak). Some recipes decay to near-zero; this is acceptable but slightly more aggressive.

Worked Example: Hyperparameter Search for a 350M Model

To illustrate the full workflow, consider setting up training for a 350M parameter model with a 100B token dataset.

Step 1: Set architecture-determined hyperparameters.

For a 350M model, a typical architecture is 24 layers, hidden dimension 1024, 16 attention heads, FFN ratio 4. These are fixed by the scaling recipe and not searched.

Step 2: Set the compute-optimal training length.

Using the Chinchilla recipe, a 350M model should train on roughly 20×350×106=7×10920 \times 350 \times 10^6 = 7 \times 10^9 tokens. With a batch size of 0.5M tokens per step, that is 7×109/(5×105)=14,0007 \times 10^9 / (5 \times 10^5) = 14{,}000 steps. We will use 15,000 steps rounded up.

Step 3: Proxy search on a 40M model.

Build a 40M parameter proxy using the same architecture ratios (fewer layers, smaller hidden dimension). Train it for 2,000 steps (proportional to the proxy's compute-optimal budget) at 7 different learning rates:

lr∈{3×10−5, 1×10−4, 3×10−4, 6×10−4, 1×10−3, 3×10−3, 1×10−2}\text{lr} \in \{3 \times 10^{-5},\ 1 \times 10^{-4},\ 3 \times 10^{-4},\ 6 \times 10^{-4},\ 1 \times 10^{-3},\ 3 \times 10^{-3},\ 1 \times 10^{-2}\}

Evaluate validation loss at the end of each run. Identify the best configuration and the second-best; the optimal is likely between these two.

Step 4: Refine on the proxy.

Run a finer search between the two best values. With 5 additional evaluations, you can narrow down to within 20% of the true optimum.

Step 5: Transfer to the 350M target.

If using muP, transfer the proxy learning rate directly. If using standard parametrization, scale by the width ratio: lr350M=lr40M×d40M/d350M\text{lr}_{350\text{M}} = \text{lr}_{40\text{M}} \times \sqrt{d_{40\text{M}} / d_{350\text{M}}} approximately.

Step 6: Run and monitor.

During the full training run, monitor:

  • Training loss curve: should decrease smoothly with no spikes
  • Gradient norm: should stabilize around 1.0-3.0 after warmup
  • Learning rate schedule: confirm the schedule matches expectations
  • Validation loss: check against expected scaling law prediction

If validation loss diverges from prediction by more than 10%, investigate before continuing. Common causes are suboptimal learning rate, data quality issues, or hardware bugs.

Troubleshooting Hyperparameter Problems

Even with careful selection, training runs sometimes go wrong. Diagnosing whether a problem is hyperparameter-related versus data-related versus hardware-related is a skill that comes with experience. Here are the most common failure patterns and their likely causes.

Training loss diverges early (within the first 5% of steps). This almost always indicates a learning rate that is too large. The loss may spike suddenly or oscillate with increasing amplitude. Reduce the learning rate by 3-5x. If the problem persists even at very low learning rates, check for data corruption (NaN values in the training data, or sequences with degenerate token distributions) and verify that gradient clipping is enabled.

Training loss decreases but plateaus far above expected scaling law predictions. This indicates a learning rate that is too small, an inappropriate schedule, or data quality issues. If gradient norms are consistently near zero, the learning rate is likely too small. If gradient norms are healthy but the loss plateau is unexpectedly high, the data may contain too many duplicates or low-quality examples that dilute the signal.

Training is stable but validation loss begins increasing after many steps. This is overfitting, most common when the training dataset is small relative to model size or training duration. Increase weight decay, reduce the number of training steps, or augment the dataset. For fine-tuning scenarios specifically, also consider reducing the learning rate.

Loss spikes periodically during mid-training. Periodic spikes are often caused by gradient outliers from specific training examples (corrupted documents, very long sequences that overflow attention), or by numerical precision issues with mixed-precision training. Setting gradient clipping to a lower threshold (0.5 instead of 1.0) and checking the data pipeline for anomalies usually resolves this.

The proxy model's optimal learning rate does not work at the target scale. If you are not using muP, this is expected: you need to scale the learning rate by the width ratio. If you are using muP, verify that the muP initialization and learning rate scaling rules were applied correctly to all parameter types, including the output projection and embeddings, which are easy to miss.

Limitations and Impact

Hyperparameter search remains expensive despite methodological advances. Even with muP enabling transfer, you must still run experiments on the proxy model, which takes time and compute. For very large models (100B parameters and above), even the proxy is expensive: a "small" proxy at 1/100th the parameter count might still be a 1B parameter model. In practice, most large-scale training runs at frontier scale use informed defaults with light verification, not exhaustive search. The real value of systematic search methods is at the medium scale (100M to 10B parameters) where you have enough budget to explore but not so much that defaults always suffice.

The assumption underlying hyperparameter transfer is that the proxy model behaves qualitatively like the target model. This holds for width scaling under muP but is less reliable for depth scaling, for qualitative differences in data distribution, or for different training objectives. Pretraining and fine-tuning have different optimal hyperparameters: fine-tuning typically benefits from lower learning rates (often 1/10th to 1/100th of the pretraining rate), shorter schedules, and lower weight decay. When these assumptions break down, configurations that seemed optimal on the proxy can underperform on the target by a substantial margin, and you will not know until you have committed significant compute.

A deeper limitation is that muP is not yet universally adopted. Implementing muP correctly requires modifying the initialization and learning rate scaling for each parameter type, which is non-trivial in existing codebases and easy to get wrong. Some publicly available implementations have subtle bugs. Until muP is a standard feature of training frameworks, many practitioners work with standard parametrization and rely on the empirical learning rate priors described in Table lr-by-size combined with a small amount of verification at the target scale.

Default recipes embed assumptions about the training regime. The GPT-style recipe assumes long-horizon training on diverse text data. Applying it to a domain-specific model trained for far fewer steps or on narrow data may give suboptimal results. Weight decay of 0.10.1 is appropriate when the dataset is large relative to model capacity. It can over-regularize when fine-tuning a large model on a small dataset, suppressing the model's ability to adapt to the target domain. In such cases, reducing weight decay to 0.010.01 or even 0.00.0 for the fine-tuning phase is often beneficial.

Despite these limitations, the development of systematic search strategies, muP for transfer, and community-validated default recipes has substantially reduced the cost of training new models. Researchers today can configure a training run with much more confidence than those five years ago, not because the search problem is solved, but because the prior over good configurations has narrowed dramatically. The field has moved from "search everything" to "know what to search, use defaults for everything else." The next chapters move to Part XLVII, where we explore specialized training for code generation models, which inherits this foundation but adds new considerations around tokenization and training objectives specific to programming languages.

Summary

Hyperparameter selection in LLM training requires choosing configurations for learning rate, schedule, optimizer parameters, regularization, and batch size simultaneously. The key ideas from this chapter are:

  • Grid search covers all combinations but scales poorly with dimensionality. Random search explores the important dimensions more efficiently at the same compute budget.
  • Bayesian optimization finds good configurations in fewer evaluations than random search by building a surrogate model of the objective.
  • Successive halving and Hyperband reduce total compute by eliminating poor configurations early in training, allocating more resources to promising candidates.
  • Hyperparameter interactions matter: learning rate and batch size scale together, warmup duration depends on β2\beta_2, and weight decay interacts with training duration. Fix these jointly rather than independently.
  • Hyperparameter transfer exploits the observation that good configurations on small models predict good configurations on large models. This works reliably under muP for learning rate across model widths.
  • Critical hyperparameters (learning rate, schedule shape) require careful search. Robust hyperparameters (Adam betas, gradient clipping) work with universal defaults across nearly all settings.
  • Default recipes (GPT-style, Chinchilla, LLaMA-style) provide validated starting points. Use the size-based learning rate prior from Table lr-by-size and adjust based on your specific setup.
  • Fine-tuning requires lower learning rates, shorter warmup, and lower weight decay than pretraining, because the model is adapting an existing representation rather than building one from scratch.
  • The recommended workflow is: set architecture hyperparameters by scaling laws, search learning rate on a proxy model, transfer to the target, and monitor gradient norms and loss curves during training.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about hyperparameter selection for LLM training.

Hyperparameter Selection Quiz

Question 1 of 80 of 8 completed
Why does random search typically outperform grid search for hyperparameter tuning when only a few hyperparameters are critical?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026hyperparameterselection, author = {Michael Brenndoerfer}, title = {Hyperparameter Selection: Search, Transfer, Default Recipes}, year = {2026}, url = {https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-06} }
APAAcademic
Michael Brenndoerfer (2026). Hyperparameter Selection: Search, Transfer, Default Recipes. Retrieved from https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training
MLAAcademic
Michael Brenndoerfer. "Hyperparameter Selection: Search, Transfer, Default Recipes." 2026. Web. October 6, 2026. <https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training>.
CHICAGOAcademic
Michael Brenndoerfer. "Hyperparameter Selection: Search, Transfer, Default Recipes." Accessed October 6, 2026. https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Hyperparameter Selection: Search, Transfer, Default Recipes'. Available at: https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training (Accessed: October 6, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Hyperparameter Selection: Search, Transfer, Default Recipes. https://mbrenndoerfer.com/writing/hyperparameter-selection-search-transfer-recipes-llm-training

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.