Part of Language AI Handbook
Explains how DeepMind's Chinchilla scaling laws changed LLM training by proving models should use 20 tokens per parameter for compute-optimal performance.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Chinchilla Scaling Laws
In March 2022, a paper from DeepMind titled "Training Compute-Optimal Large Language Models" upended conventional wisdom about how to train large language models. The paper introduced Chinchilla, a 70 billion parameter model that outperformed the 280 billion parameter Gopher despite using the same computational budget. This result challenged the prevailing approach of building ever-larger models and revealed that the field had been systematically undertraining its models on data.
As we discussed in the Kaplan Scaling Laws chapter, the original OpenAI scaling work suggested that model size should increase faster than dataset size as compute budgets grow. This led labs to prioritize larger models, sometimes training them on relatively small amounts of data. Chinchilla demonstrated that this approach was suboptimal, showing that model size and training tokens should scale together in roughly equal proportion.
Think of the situation this way: imagine you are studying for an exam and you have a fixed number of study hours. One strategy is to buy a thick textbook with thousands of pages and read as much of it as possible. Another strategy is to buy a concise guide and study it deeply and repeatedly, covering every concept thoroughly. The Kaplan-era LLM community favored the first strategy: buy the biggest textbook (build the largest model) and read as much as you can. Chinchilla showed that the second strategy is better: a well-studied smaller model consistently outperforms a poorly studied larger one given the same time investment.
This insight seems obvious in retrospect, but it ran against powerful incentives in the field. Larger models were easier to announce, more impressive in press releases, and more likely to attract funding. The race to build the biggest model obscured a more fundamental question: given a fixed computational budget, what is the optimal way to spend it? The Chinchilla paper reframed LLM development from a size race into a compute efficiency problem, with far-reaching consequences for how organizations now design and train language models.
The impact of Chinchilla extended far beyond the specific model it introduced. The paper's methodology, using multiple independent experimental approaches to triangulate the same conclusion, set a new standard for empirical scaling research. Its finding that the 20 tokens-per-parameter ratio provides a practical guideline for training gave practitioners a concrete tool for planning training runs. And its demonstration that a smaller model could beat a much larger one shifted the conversation toward smarter resource allocation rather than raw scale. Understanding why Chinchilla reached its conclusions, and what assumptions underlie them, is essential context for anyone working in or studying LLMs today.
The DeepMind Experiments
The Chinchilla paper is notable for its careful methods. Rather than relying on a single experimental approach, the DeepMind team used three independent methods to estimate optimal scaling, finding very consistent results across all three. This agreement between different approaches gives us strong confidence in the conclusions. When three different methods of measuring the same phenomenon all point to the same answer, we can trust that answer far more than if it came from a single experiment. We examine each approach in detail to understand both the methodology and the insights each reveals.
The scale of the experimental effort behind Chinchilla is worth appreciating. The team trained over 400 models in total, ranging from 70 million to 16 billion parameters. Each training run consumed real GPU hours and required careful coordination to ensure the results were comparable across configurations. This was not a theoretical analysis or a small proof-of-concept study: it was a systematic empirical program designed to answer a precise question about the optimal allocation of compute between model size and training data. That breadth, covering many model sizes and training durations rather than just a few flagship experiments, let the team test the allocation question across a wide range.
Understanding why the DeepMind team chose three separate methods, rather than just one definitive experiment, also illuminates something about the nature of scaling research. Each method captures a different slice of the data-generating process. The first method asks: given a model of fixed size, how does more data help? The second method asks: given a fixed compute budget, how should we split it between capacity and data? The third method asks: can we fit a mathematical model to all the data simultaneously and read off the optimal relationship analytically? The fact that all three converge on the same answer is not a coincidence but a deliberate experimental design choice intended to rule out artifacts from any single methodology.
Before Chinchilla, the dominant view in the community came from the 2020 Kaplan et al. paper from OpenAI, which had studied scaling behavior extensively but with models clustered at smaller scales and using training configurations that did not always optimize learning rate schedules for the full training duration. The Kaplan paper's conclusion that model size should scale faster than data influenced virtually every large model trained in 2020 and 2021, including GPT-3 (175B parameters), Gopher (280B), and Megatron-Turing NLG (530B). All of these models were trained on far fewer tokens than the Chinchilla analysis would later show to be optimal. DeepMind's team, in designing the Chinchilla experiments, deliberately tried to address methodological gaps in the Kaplan analysis, particularly the limited compute range studied and the sensitivity of power-law fits to training configuration choices.
Approach 1: Fixing Model Sizes
The first approach trained over 400 models ranging from 70 million to 16 billion parameters. For each model size, the team varied the number of training tokens and measured the resulting loss. This produced loss curves as a function of training tokens for each model size.
The intuition behind this approach is straightforward: if we hold model size constant and simply vary how much data we train on, we can see exactly how each model size benefits from additional training. Some models might saturate quickly, suggesting they've learned all they can from the available data. Others might continue improving steadily, indicating they could benefit from even more training. By mapping these curves across many model sizes, we can compare how different architectures respond to different data quantities.
From these curves, the team extracted the minimum loss achievable at each compute budget. Here compute for parameters and training tokens, and the factor of 6 accounts for both forward and backward passes across all operations. By examining which model size achieved the lowest loss at each compute level, they could determine the optimal allocation between parameters and data.
The key insight from this approach was that for any fixed compute budget, there exists an optimal model size, and training either too large or too small a model results in worse performance. Models that were too large couldn't be trained on enough data given the compute constraint, while models that were too small couldn't capture sufficient complexity regardless of training duration. This finding shows a key trade-off in neural network training: we must balance the model's capacity to represent complex patterns against its opportunity to learn those patterns from data. Neither capacity nor data alone determines performance. Their interaction does.
Notice that this approach requires a commitment to carefully controlled experiments. Simply reporting the loss after a fixed number of training steps would conflate two effects: the model's capacity and the amount of data it has seen. By explicitly varying training tokens for each model size and tracking the resulting loss curves, the team disentangled these factors. This careful separation is what makes the conclusions interpretable. When they later say that model X is "optimal" for compute budget C, they mean specifically: among all models tested, X achieves the lowest loss when given the tokens that its compute budget allows it to train on.
In practice, this means that the loss curves revealed something subtle: at any fixed compute level, the model that achieves the best loss is not necessarily the one that would eventually achieve the best loss if trained for much longer. A small model might reach excellent loss quickly and appear optimal at moderate compute, but a larger model, given the same total compute, might still be catching up. The IsoFLOP analysis in Approach 2 addresses this precisely, but the loss curve data from Approach 1 is what makes it possible to identify where the curves for different sizes cross as a function of total compute invested.

Approach 2: IsoFLOP Curves
The second approach defined several fixed compute budgets (measured in FLOPs) and trained many models at each budget. For a given compute budget , the team explored different allocations: some configurations used more parameters with fewer training tokens, while others used fewer parameters with more training tokens.
This approach inverts the question from the first method: rather than asking "how does performance vary as we give a fixed model more data?", it asks "given a fixed computational budget, how should we divide that budget between model capacity and training duration?" This framing directly addresses the practical question for any team with a finite compute budget: should we build a bigger model and train it less, or build a smaller model and train it more?
This created "iso-FLOP" curves showing how loss varies with model size at fixed compute. Each curve exhibited a clear minimum corresponding to the compute-optimal allocation. As the compute budget increased, this optimal point shifted to larger models, but importantly, the optimal dataset size grew at a similar rate. The shape of these curves, rising on both ends with a clear valley in the middle, visually demonstrates that extremes are costly: training a model that is too small wastes compute on redundant passes over data the model cannot fully use, while training a model that is too large wastes compute on parameters that never get enough gradient updates to learn useful representations.
The iso-FLOP methodology directly answers the question a practitioner faces when planning a training run. You cannot train everything: you have a GPU cluster, a time budget, and a target FLOPs count. Given that constraint, you need to decide how to configure your model. The iso-FLOP curves give you exactly that answer: find your budget on the x-axis, find the minimum of the corresponding curve, and read off the optimal model size. The fact that the curve has a clear, unambiguous minimum at every compute level tested means the answer is not sensitive to minor variations in experimental conditions.
The key insight from the iso-FLOP analysis is the shift in the optimal point as compute grows. If you plot the optimal model size at each compute budget, you get a curve that rises. If you plot the optimal token count, you get another curve that also rises. The question is: how fast does each curve rise relative to the other? This is exactly the question of whether the scaling exponents and are equal or unequal. The Chinchilla analysis found that both curves rise at approximately the same rate, meaning both the optimal parameter count and the optimal token count grow at the same pace as compute increases. This equal-rate growth is the mathematical foundation of the 20:1 token-to-parameter ratio and the practical prescription that model size and data should scale together.

Approach 3: Parametric Loss Fitting
The third approach fit a parametric function to predict loss as a function of both model size and training tokens:
where:
- : the predicted loss for a model with parameters trained on tokens
- : the number of model parameters
- : the number of training tokens
- : the irreducible entropy of natural language, the minimum loss achievable with infinite model size and infinite data
- , : learned scaling coefficients that determine how quickly loss decreases with more parameters or data
- , : learned exponents that control the rate of diminishing returns for parameters and data respectively
This functional form provides a way to understand scaling behavior. Rather than treating each data point in isolation, this approach assumes that loss follows a specific mathematical structure, then fits that structure to all observed data simultaneously. The advantage is that once we have reliable estimates of the parameters, we can extrapolate to compute regimes we haven't yet explored and predict what loss we might achieve with a model ten times larger than any we've trained.
This functional form captures three contributions to the loss:
- represents the irreducible entropy of natural language, the minimum loss achievable with infinite model size and infinite data. This term sets a floor that no amount of scaling can overcome. Even a perfect model with unlimited training cannot predict language better than the inherent randomness in human text allows. This irreducible component reflects the unpredictability of language. sometimes multiple valid words could follow a given context, and no model can know which one a human will choose.
- captures how loss decreases as model capacity increases. The power law form reflects diminishing returns. Each doubling of parameters yields a fixed percentage improvement, not a fixed absolute improvement. Moving from 1 billion to 2 billion parameters provides the same relative gain as moving from 100 billion to 200 billion parameters. The exponent controls how steep this diminishing return curve is. Larger values of mean faster initial improvement but quicker saturation.
- captures how loss decreases as training data increases. Similarly, this term shows that more data helps, but with diminishing returns governed by the exponent . The first million tokens teach the model basic language structure. The next million refine that understanding. Eventually, additional tokens provide only marginal improvements as the model has already captured most learnable patterns.
By fitting this function to the experimental data, the team estimated the exponents and along with the constants , , and . These parameters then allow computation of the optimal allocation for any compute budget. This approach has predictive power: rather than running expensive experiments at every scale we care about, we can fit the function to data from smaller, cheaper experiments and then mathematically derive what should happen at larger scales.

The Chinchilla Scaling Equations
All three experimental approaches yielded consistent conclusions about optimal scaling. The Chinchilla team derived that for compute-optimal training, both model size and dataset size should scale as power laws of the compute budget. This convergence across three independent methods is the strongest form of empirical evidence: not only did the team find a result, they found it three times, using three different experimental designs, and it came out the same each time.
The mathematical form of these relationships is a power law, the same functional family that the Kaplan paper found. Such relationships also appear throughout physics and economics, as well as biology, when quantities scale with each other. A power law relationship says that when increases by a factor of 10, increases by a factor of . If , they scale together linearly. If , grows more slowly than . The scaling exponent therefore encodes the efficiency of converting one resource into another, and finding that exponent precisely is the core empirical contribution of the Chinchilla paper.
Before seeing the equations, notice an important structural constraint: compute is approximately equal to , where is parameter count and is training tokens. This means that if we increase and by some factors, compute increases by the product of those factors. Any claim about how and should grow with must be consistent with this multiplicative relationship. This constraint will force the scaling exponents and to satisfy a specific relationship, which we will work out below.
The Chinchilla team derived that for compute-optimal training, both model size and dataset size should scale as power laws of the compute budget:
where:
- : the optimal number of model parameters for a given compute budget
- : the optimal number of training tokens for a given compute budget
- : the total compute budget (in FLOPs)
- , : scaling exponents that determine how parameters and data should grow with compute
- The constraint follows from (compute scales with the product of parameters and tokens)
These power law relationships mean that as your compute budget grows by some factor, both the optimal model size and dataset size grow by predictable amounts determined by the exponents and . The proportionality symbol () indicates that while we don't specify the exact constants, the scaling relationship holds. If you increase compute by a factor of 100, the optimal parameter count increases by and the optimal token count increases by .
The constraint that is important because it comes from basic arithmetic rather than experiments. Since compute is approximately proportional to the product of parameters and tokens (), and since and , we have . For this equation to hold, we need . This constraint means that the exponents and cannot be chosen independently. They must sum to one, representing different ways of dividing the same computational pie.
The Equal Scaling Exponents
The central finding was that , meaning model size and dataset size should grow at roughly the same rate with compute. This stands in stark contrast to the Kaplan scaling laws, which suggested and , implying model size should grow much faster than dataset size.
Consider what this means in practice. With Chinchilla scaling (equal exponents around 0.5), if you increase your compute budget by a factor of 100, you should increase both your model size and your dataset size by a factor of 10 (since ). With Kaplan scaling (unequal exponents), the same 100x compute increase would suggest increasing model size by about 50x () but dataset size by only about 2x (). The difference matters: Kaplan scaling funnels most additional resources into model capacity, while Chinchilla scaling distributes resources evenly between capacity and training data.


Fitting the parametric loss model to experimental data yielded:
where:
- : the exponent governing how loss decreases with model size (from the term)
- : the exponent governing how loss decreases with training data (from the term)
These exponents determine how quickly each source of error decreases as we add parameters or data. The relatively similar values explain why balanced scaling is optimal: neither parameters nor data provide much faster diminishing returns compared to the other. If were much larger than , it would mean that increasing model size reduces loss much faster than increasing data, and we should favor parameters. If were much larger, we should favor data. But with and , both resources contribute roughly equally to loss reduction per unit of compute invested, so we should invest roughly equally in both.
Optimal Tokens Per Parameter
The most practical finding from Chinchilla is the optimal ratio of training tokens to model parameters. The analysis suggests:
where:
- : the optimal number of training tokens
- : the optimal number of model parameters
- The ratio of 20 emerges from the balanced scaling exponents ()
This means a compute-optimal model should be trained on approximately 20 tokens per parameter. A 1 billion parameter model should see roughly 20 billion tokens, a 10 billion parameter model should see 200 billion tokens, and so on.
This ratio provides a practical guideline: before training any model, you can estimate whether your training corpus is appropriately sized. Multiply your planned parameter count by 20, and that's roughly how many tokens you need for compute-optimal training. If you have fewer tokens available, you're training a model too large for your data. If you have many more tokens, you might benefit from a larger model.
This ratio is remarkably consistent across scales. The team found that deviating significantly from this ratio leads to compute inefficiency: either the model is too small to fully utilize the available training signal, or too large to be trained adequately given the data budget. The intuition is that each parameter represents a "slot" for storing learned knowledge, and 20 tokens provides enough gradient signal to fill that slot effectively. Too few tokens means the slot remains partially empty; too many tokens means the model cannot store all the patterns it could learn.
Chinchilla vs Kaplan: Understanding the Discrepancy
The Chinchilla and Kaplan scaling laws arrive at different conclusions despite both being grounded in extensive experiments. Understanding why helps clarify the details of scaling analysis. This is not a case of one paper being right and the other wrong in some absolute sense. Both papers conducted large-scale experiments and fit power laws to their data. The discrepancy arises from differences in scope, methodology, and the sensitivity of power-law fitting to choices that might seem inconsequential but turn out to matter quite a lot at large scales.
Think of fitting a scaling law as like estimating the slope of a road by measuring its rise and run over a short stretch. If you measure over 100 meters, you get an estimate. If you later measure over 10 kilometers, you might find the slope is slightly different because the road curves. Neither measurement is wrong. They describe the road in different local regions. Power-law exponents for neural network scaling are similar: they are estimated from a finite range of experiments, and extrapolating beyond that range introduces uncertainty. The Kaplan and Chinchilla papers both estimated these exponents, but from different ranges and under different experimental conditions, and the resulting exponents diverged in ways that matter enormously at the scales where frontier models are trained.
The discrepancy had enormous practical consequences. Researchers and organizations that followed Kaplan's guidance trained models like GPT-3 (175B parameters, 300B tokens) and Gopher (280B parameters, 300B tokens) with parameter-to-token ratios far below what Chinchilla would recommend. The resources spent training those models were not wasted, since the models performed well, but they were not as efficiently spent as they could have been. Understanding the source of the discrepancy is therefore not merely academic: it speaks directly to how future scaling research should be designed and interpreted.
Different Compute Ranges
The Kaplan experiments focused primarily on smaller models (up to a few hundred million parameters) with relatively limited compute budgets. The Chinchilla experiments covered a broader range, including models up to 16 billion parameters trained with significantly more compute.
Power law exponents estimated from one regime may not extrapolate to others. The Chinchilla team's larger-scale experiments revealed behavior that wasn't apparent in the smaller-scale Kaplan analysis.
Training Details Matter
The two studies used different training configurations:
| Aspect | Kaplan | Chinchilla |
|---|---|---|
| Learning rate schedule | Shorter warmup | Cosine decay with longer training |
| Number of training tokens | Fixed, relatively short | Varied extensively |
| Model architectures | GPT-style | Similar but with some modifications |
| Stopping criteria | Earlier stopping | Training to convergence |
The Kaplan analysis suggested that large models can be "early stopped" effectively, meaning you don't need to train them to convergence to achieve good performance. However, this early stopping analysis appears to have underestimated the benefits of longer training.
The Learning Rate Schedule Effect
One key difference involves learning rate scheduling. The Kaplan experiments often trained models for fewer steps than optimal, and learning rate schedules were not always tuned for the full training duration. The Chinchilla experiments used carefully tuned cosine learning rate schedules that decayed properly over the intended training length.
When you tune your learning rate schedule specifically for longer training runs, the benefit of additional training data increases. Models can extract more signal from the data because the optimization process is better matched to the training duration.
The Loss Decomposition Interpretation
Both analyses fit loss as a function of parameters and data, but with different assumptions about the functional form. The Kaplan decomposition allowed for interaction terms and considered how loss varies in suboptimal regimes differently than Chinchilla.
When fitting power laws to noisy data, small differences in the fitting procedure can lead to different exponent estimates, especially when extrapolating beyond the observed range.
Implications of Chinchilla Scaling
The Chinchilla results had immediate implications for how the field approaches language model training. The paper did not just publish a set of equations and leave practitioners to work out the consequences. The results landed in the middle of an active period of model development, and labs that had spent significant resources building very large models suddenly had to reckon with evidence that their approach had been suboptimal. The response was rapid and visible: within a year, the field had largely shifted toward training smaller models more extensively on higher-quality data.
The implications spread across several dimensions simultaneously. The immediate technical implication was that future models should apply the 20:1 tokens-per-parameter ratio. The economic implication was that smaller models are cheaper to deploy at inference time, meaning Chinchilla-optimal training could deliver both better performance and lower serving costs. The data implication was that achieving the optimal ratio at large scales requires enormous amounts of data. This pushed the field to invest heavily in collecting and filtering data while improving its quality. Each of these consequences reshaped how teams thought about model development, and together they shifted the focus from parameter count toward compute efficiency.
The key insight here is that the Chinchilla paper changed what models were built and how success was defined. Before Chinchilla, the implicit metric was often "biggest model that can be trained within a fixed budget." After Chinchilla, the relevant question became "most capable model per FLOP of training compute." This shift in framing had downstream effects on benchmarking, on how model announcements were structured, and on what kinds of architectural innovations were considered worth pursuing. Understanding these implications requires examining each area where the results made contact with practice.
Many Models Were Undertrained
At the time of publication, most large language models had been trained well below the Chinchilla-optimal data budget:
| Model | Parameters | Training Tokens | Tokens/Parameter |
|---|---|---|---|
| GPT-3 | 175B | 300B | 1.7 |
| Gopher | 280B | 300B | 1.1 |
| Megatron-Turing NLG | 530B | 270B | 0.5 |
| Chinchilla | 70B | 1.4T | 20 |
GPT-3, despite its impressive capabilities, was trained on only 1.7 tokens per parameter, far below the Chinchilla-optimal 20 tokens. The same compute budget could have trained a smaller model to significantly better performance.

The Shift to Smaller, Better-Trained Models
Following Chinchilla, the field rapidly shifted toward training smaller models on more data. LLaMA 1 (7B to 65B parameters) trained on 1-1.4T tokens. LLaMA 2 continued this trend. The Chinchilla insight provided both economic and capability motivation: smaller models are cheaper to deploy while potentially achieving better performance than larger but undertrained alternatives.
The LLaMA family from Meta became the most visible embodiment of this shift. The 7B LLaMA model, trained on 1 trillion tokens, achieved performance on many benchmarks that rivaled or exceeded much larger models trained under the older paradigm. This was not a marginal improvement. It was a qualitative demonstration that the Chinchilla framework was correct in practice, not just in theory. The field took note. Subsequent model families from Mistral and Falcon adopted similar philosophies of training smaller models more extensively, as did Qwen and models from other groups.
There is also an important consequence for accessibility. A 7B model that matches the performance of a 70B undertrained model is cheaper to serve and fundamentally more accessible. It can run on consumer hardware, can be fine-tuned by individual researchers, and can be deployed in contexts where the computational infrastructure for larger models simply does not exist. The Chinchilla insight therefore democratized high-quality language models in a real sense by improving efficiency metrics and lowering the hardware requirements for achieving a given capability level.
Data Becomes the Bottleneck
If models should be trained on 20 tokens per parameter, then a 100 billion parameter model needs 2 trillion training tokens. At this scale, data availability becomes a serious constraint. High-quality text data is finite, and the Chinchilla scaling laws pushed the field toward solutions like:
- More sophisticated data filtering and deduplication
- Synthetic data generation
- Multi-epoch training on high-quality subsets
- Multimodal data incorporation
We'll explore these data-constrained scenarios in the upcoming chapter on Data-Constrained Scaling, where the Chinchilla optimum cannot be achieved due to limited data availability.
The data constraint has turned out to be one of the most consequential unsolved problems in language model development. The Common Crawl, which is a large-scale web scrape that forms the backbone of most training datasets, contains roughly 15-20 trillion tokens of text in its complete form. At the Chinchilla-optimal ratio, this is sufficient to train a model of about one trillion parameters at most, assuming the data could all be used at high quality. In practice, after deduplication, quality filtering, and safety filtering, the usable portion shrinks substantially. This means that for models above roughly 100-200B parameters, the Chinchilla-optimal training regime is already bumping up against the limits of available high-quality internet text.
The response to this constraint has taken several forms. Organizations have invested heavily in data curation pipelines that extract more signal from the available raw data. Some have turned to proprietary text sources: books, scientific papers, code repositories, and licensed news archives. Others have begun using large language models themselves to generate synthetic training data, an approach that raises interesting questions about the circularity of training models on model-generated text. All of these strategies represent the field grappling with the consequence that Chinchilla scaling, by making data as important as model size, created a demand for training data that now exceeds the supply of naturally occurring high-quality human text.
Inference Cost Considerations
The Chinchilla framing focuses purely on training compute. However, models deployed at scale accumulate inference costs that eventually dominate training costs. A smaller Chinchilla-optimal model costs less per inference request than a larger undertrained model.
This inference advantage compounds: if a 70B Chinchilla-optimal model matches the quality of a 280B undertrained model, the inference cost reduction is roughly 4x for each query. For models serving millions of requests, this adds up to large savings.
The economics become even more favorable when you consider that inference runs continuously after training, while training is a one-time cost. A model trained over three months might then serve queries for two or three years. If inference costs are ten times higher than training costs over the model's lifetime (a plausible order-of-magnitude estimate for widely deployed models), then the efficiency of inference matters more than the efficiency of training. Chinchilla's recommendation of smaller, better-trained models aligns naturally with this economic reality: you spend somewhat more on data (which is cheap) and somewhat less on model parameters (which are expensive to store and serve), and the deployment economics strongly favor the result.
In practice, this reasoning has led some organizations to deliberately overtrain smaller models beyond the Chinchilla compute-optimal point. If you know you will be serving many queries, it can make sense to train a 7B model on 2 trillion tokens rather than the Chinchilla-optimal 140 billion tokens, because the inference savings from running a 7B model instead of a 70B model are substantial enough to justify the extra training cost. This logic is sometimes called inference-optimal scaling, and it represents an important extension of the Chinchilla framework that acknowledges the real economic environment in which language models operate.
Worked Example: Computing Chinchilla-Optimal Allocations
This section shows how to compute Chinchilla-optimal model sizes and dataset sizes for a given compute budget. This shows how the scaling equations translate into concrete numbers. Seeing the algebra work through a specific example is important because the abstract power-law relationships can feel distant from the practical decisions that model developers make. When you compute that a FLOP budget corresponds to a specific parameter count and token count, the equations stop being theoretical constructs and become engineering guidelines you can apply to your own training plans.
We will work through the calculation step by step, explaining the role of each factor and why the algebra takes the form it does. The derivation is short but worth following carefully, because understanding where the square root comes from and why the factor of 120 appears will help you apply the formula correctly in other contexts and catch errors when something does not look right.
Suppose we have a compute budget of FLOPs. Using the approximation that training FLOPs equals (where is parameters and is training tokens), and applying the Chinchilla ratio :
where:
- : the total compute budget in FLOPs
- : the number of model parameters (what we want to solve for)
- : substituting the Chinchilla-optimal token count
- The factor combines the compute constant with the optimal ratio
This substitution works because if (the Chinchilla-optimal ratio), then the compute formula simplifies to a function of alone. This constrains us to move along the Chinchilla-optimal line in the parameter-token space and determines where along that line our compute budget places us. The quadratic relationship () emerges because both and grow together. Increasing automatically increases by the same factor, so compute grows as the square.
Solving for :
where:
- : the optimal number of model parameters we're solving for
- : our target compute budget in FLOPs
- The factor 120 comes from (the compute constant 6 multiplied by the Chinchilla-optimal token ratio 20)
- The result means approximately 9 billion parameters
The square root appears because we're inverting the quadratic relationship: if , then . This means that to double our optimal model size, we need to quadruple our compute budget. Similarly, a 10x increase in compute allows only about a 3.2x increase in model size (since ).
So approximately 9 billion parameters. The optimal token count is:
where:
- : the optimal number of training tokens
- : the optimal parameter count computed above
- The factor 20 is the Chinchilla-optimal tokens-per-parameter ratio
Let's verify the compute calculation:
This verification step confirms our arithmetic and shows that the Chinchilla-optimal allocation exactly consumes the available compute budget. Nothing is wasted. Every FLOP goes toward training a model that is neither too large (which would mean insufficient training) nor too small (which would mean wasted training iterations).
For comparison, the Kaplan scaling laws would suggest allocating more to model size. Using Kaplan's and (with appropriate constants), the same compute budget would suggest roughly 30-50 billion parameters trained on only 30-50 billion tokens. The Chinchilla allocation produces better loss for the same compute because the Kaplan allocation produces a severely undertrained model. With only about 1 token per parameter, the model has insufficient opportunity to learn from data, and many of its parameters remain poorly optimized.
Code Implementation
Let's implement functions to compute Chinchilla-optimal allocations and compare them with Kaplan allocations. The implementation will help you build intuition for how the two frameworks diverge as compute scales, and give you a practical tool for planning training runs under either framework. We will start with the core allocation functions, then compute and visualize the results across a wide range of compute budgets to see how the predictions diverge at large scales.
import numpy as np
# Set random seed for reproducibility
np.random.seed(42)
def chinchilla_optimal_allocation(compute_flops: float) -> dict:
"""
Compute Chinchilla-optimal model size and dataset size.
Uses the approximation C = 6*N*D and D/N ≈ 20.
"""
# C = 6 * N * D = 6 * N * 20 * N = 120 * N^2
# Therefore N = sqrt(C / 120)
optimal_params = np.sqrt(compute_flops / 120)
optimal_tokens = 20 * optimal_params
return {
"parameters": optimal_params,
"tokens": optimal_tokens,
"tokens_per_param": optimal_tokens / optimal_params,
"compute_flops": compute_flops,
}
def kaplan_optimal_allocation(compute_flops: float) -> dict:
"""
Compute Kaplan-style allocation with N scaling faster than D.
Uses approximate exponents a=0.73, b=0.27 from original paper.
Calibrated so both methods agree at C = 10^20 for comparison.
"""
# Calibration constants (tuned to match similar scale)
alpha_n = 0.73
alpha_d = 0.27
# Reference point for calibration
c_ref = 1e20
n_ref = np.sqrt(c_ref / 120) # Match Chinchilla at reference
d_ref = 20 * n_ref
k_n = n_ref / (c_ref**alpha_n)
k_d = d_ref / (c_ref**alpha_d)
optimal_params = k_n * (compute_flops**alpha_n)
optimal_tokens = k_d * (compute_flops**alpha_d)
return {
"parameters": optimal_params,
"tokens": optimal_tokens,
"tokens_per_param": optimal_tokens / optimal_params,
"compute_flops": compute_flops,
}Now let's compute allocations across a range of compute budgets and visualize the differences:
# Define compute budgets from 10^18 to 10^25 FLOPs
compute_budgets = np.logspace(18, 25, 50)
chinchilla_allocations = [
chinchilla_optimal_allocation(c) for c in compute_budgets
]
kaplan_allocations = [kaplan_optimal_allocation(c) for c in compute_budgets]
# Extract arrays for plotting
chinchilla_params = np.array([a["parameters"] for a in chinchilla_allocations])
chinchilla_tokens = np.array([a["tokens"] for a in chinchilla_allocations])
kaplan_params = np.array([a["parameters"] for a in kaplan_allocations])
kaplan_tokens = np.array([a["tokens"] for a in kaplan_allocations])

The left panel shows that as compute budgets increase, Kaplan scaling recommends much larger models than Chinchilla scaling. At FLOPs, the gap is nearly an order of magnitude. The right panel reveals the consequence: Kaplan allocates far fewer tokens to training, resulting in undertrained models. Chinchilla's balanced approach maintains the 20:1 token-to-parameter ratio across all scales.
The divergence becomes more pronounced at higher compute budgets. We can quantify this by looking at specific examples:
def format_number(n: float) -> str:
"""Format large numbers with appropriate suffixes."""
if n >= 1e12:
return f"{n / 1e12:.1f}T"
elif n >= 1e9:
return f"{n / 1e9:.1f}B"
elif n >= 1e6:
return f"{n / 1e6:.1f}M"
else:
return f"{n:.0f}"
# Compare at several compute scales
compute_examples = [1e20, 1e22, 1e24]
comparison_results = []
for c in compute_examples:
chin = chinchilla_optimal_allocation(c)
kap = kaplan_optimal_allocation(c)
comparison_results.append((c, chin, kap))Chinchilla vs Kaplan Optimal Allocations ---------------------------------------------------------------------- Compute Budget: 1e+20 FLOPs Method Parameters Tokens Tokens/Param ------------------------------------------------------ Chinchilla 912.9M 18.3B 20.0 Kaplan 912.9M 18.3B 20.0 Compute Budget: 1e+22 FLOPs Method Parameters Tokens Tokens/Param ------------------------------------------------------ Chinchilla 9.1B 182.6B 20.0 Kaplan 26.3B 63.3B 2.4 Compute Budget: 1e+24 FLOPs Method Parameters Tokens Tokens/Param ------------------------------------------------------ Chinchilla 91.3B 1.8T 20.0 Kaplan 759.3B 219.5B 0.3
At FLOPs, Kaplan scaling would suggest training a model with over 20x more parameters than Chinchilla recommends, but on far fewer tokens. The Chinchilla model would likely outperform significantly because it receives adequate training data.
The comparison table reveals a clear pattern. While both methods scale up with compute, they diverge sharply in how they allocate resources. At FLOPs, Kaplan scaling recommends a 2.4T parameter model trained on only 700B tokens (0.3 tokens per parameter), while Chinchilla recommends a 91B parameter model trained on 1.8T tokens (20 tokens per parameter). The Kaplan-style allocation would produce a severely undertrained model.
Now we visualize the tokens-per-parameter ratio across compute scales:

The flat blue line at 20 tokens per parameter represents the Chinchilla prescription: regardless of how much compute you have, maintain the same data intensity. The declining red line shows the problematic Kaplan implication: as compute grows, train on proportionally less data per parameter. At FLOPs, Kaplan scaling would suggest training on fewer than 5 tokens per parameter, a recipe for severe underfitting.
This visualization makes the fundamental difference clear: Chinchilla scaling maintains a constant data intensity regardless of scale, while Kaplan scaling implies training larger models on proportionally less data.
The Loss Surface Perspective
We can also visualize why balanced scaling is optimal by plotting the loss surface as a function of parameter count and token count for a fixed compute budget. This perspective offers geometric intuition for what the equations tell us algebraically. The optimal allocation lies not at the extremes but in a balanced middle region.
Think of the loss surface as a map of possible allocations. The x-axis represents how many parameters we put in our model, and the y-axis represents how many tokens we train on. For a fixed compute budget, we cannot be at every point on this surface: we are constrained to a curve (the iso-FLOP curve) where the product is constant. Moving along this curve trades more parameters for fewer tokens, or fewer parameters for more tokens, while keeping compute fixed. The question is: which point on the curve gives the lowest loss?
What the Chinchilla analysis shows is that the iso-FLOP ridge has a valley in the middle, not at the extremes. If we walk too far to the left (large model, few tokens), we end up in a region where the model has enormous capacity but has not been trained enough to use it. The loss is high because parameters are underutilized. If we walk too far to the right (small model, many tokens), we end up in a region where the model has seen plenty of data but lacks the capacity to capture all the patterns it has been shown. The loss is high because the model is overloaded. The lowest point, the optimal allocation, sits in between, where capacity and training data are in balance.
def estimate_loss(n_params: float, n_tokens: float) -> float:
"""
Estimate loss using Chinchilla-style parametric form.
L(N, D) = E + A/N^α + B/D^β
Uses approximate fitted values from the paper.
"""
E = 1.69 # Irreducible entropy of natural language (nats)
A = 406.4
B = 410.7
alpha = 0.34
beta = 0.28
return E + A / (n_params**alpha) + B / (n_tokens**beta)The loss function implemented above directly encodes the Chinchilla parametric form. Each call evaluates the predicted loss for a given parameter-token combination, allowing us to map loss across parameter-token combinations and identify where optimal performance lies.

Optimal allocation at C = 1e+22 FLOPs: Parameters: 5.0B Tokens: 334.9B Tokens per parameter: 67.3 Estimated loss: 2.139 nats
The computed optimal allocation of approximately 9B parameters and 180B tokens closely matches our earlier analytical calculation, validating the loss function approach. The estimated loss of around 2.0 nats represents a large reduction from what either extreme allocation would achieve. Moving toward the "Undertrained" region (upper left) or "Underparameterized" region (lower right) would increase loss significantly.
The optimal point lies in the balanced region, confirming the Chinchilla insight that neither extreme allocation (very large model with few tokens, or small model with many tokens) achieves the best performance. The loss surface visualization reveals the geometry underlying the scaling equations: the iso-FLOP curve traces out all possible ways to spend a fixed compute budget, and the optimal point marks where loss is minimized along this curve. The U-shaped loss profile along the curve shows that deviating in either direction, toward larger models or toward more data, incurs a penalty.

Key Parameters
The key parameters for Chinchilla scaling calculations are:
- compute_flops: Total compute budget in FLOPs, the primary input for determining optimal allocations
- tokens_per_param ratio (≈20): The Chinchilla-optimal ratio of training tokens to model parameters
- E, A, B, α, β: Fitted constants in the parametric loss function that determine the loss surface shape
- Scaling exponents (a, b ≈ 0.5): Exponents determining how optimal parameters and tokens scale with compute (, )
Caveats and Extensions
While Chinchilla scaling provides valuable guidance, several caveats apply. The Chinchilla paper represents a major empirical contribution, but it is not the final word on optimal scaling. Subsequent research has tested its conclusions and modified them in some settings. Understanding these caveats is essential for applying the Chinchilla framework correctly, rather than treating the 20:1 ratio as a universal constant that applies identically in all situations.
The broader lesson from these caveats is that scaling laws are empirical approximations derived from specific experimental conditions, not fundamental laws of nature. The power-law fits are excellent over the ranges studied, but the exponents are not constant: they depend on architecture, data distribution, training configuration, and the capability being measured. Using scaling laws well requires understanding their assumptions and being alert to when those assumptions might break down. With that context in mind, we examine each major caveat in detail.
The 20:1 Ratio Isn't Universal
The exact optimal ratio depends on data quality, model architecture, and training setup. Different analyses have found ratios ranging from 15:1 to 25:1. The key point is that the ratio should remain roughly constant across scales, not that 20:1 is precisely optimal for all situations.
The quality of training data matters substantially for this ratio. The Chinchilla experiments used a dataset called MassiveText, a carefully curated mix of web text and books. GitHub code, Wikipedia articles, and several other sources completed the mixture. If your training data contains a higher fraction of low-quality or repetitive content, the effective information per token is lower, and you may need more tokens to achieve the same loss reduction. Conversely, very high-quality data, such as scientific papers or well-edited books, may provide more signal per token, potentially reducing the required token count somewhat. The 20:1 ratio is a good starting point, but it should be validated against your specific data mixture before being treated as exact.
Beyond Single-Epoch Training
The Chinchilla analysis assumes each token is seen once during training. When data is limited, multi-epoch training adds complexity. Repeated data provides diminishing returns, so the optimal strategy shifts toward smaller models. We'll explore this in detail in the upcoming Data-Constrained Scaling chapter.
The mathematics of multi-epoch training are subtle. When a model sees the same token a second time, the gradient update it receives carries less new information than the first pass. The model has already updated its parameters in response to that data point, so repeating it provides a smaller marginal signal. This diminishing return from repeated data means that if you are forced to train in a data-limited regime, you should compensate by reducing model size: a smaller model can extract more of the available signal from a limited corpus before overfitting. The Chinchilla framework handles the single-epoch case elegantly, but practitioners working with constrained datasets need to apply corrections or use the extensions developed in subsequent work.
Downstream Task Performance
Chinchilla scaling optimizes for pre-training loss, not necessarily downstream task performance. Some evidence suggests that larger models show emergent capabilities absent in smaller but better-trained models. The loss-capability relationship is not always linear, as we discuss in the chapter on Emergence in Neural Networks.
The challenge here is fundamental: pre-training loss is a smooth, continuous quantity that is easy to optimize and measure, but downstream task performance is often discrete and can exhibit sudden jumps as model size crosses certain thresholds. A model with slightly better perplexity does not always score higher on a reasoning benchmark. Certain tasks seem to require a minimum model capacity before any progress is made, after which performance improves rapidly. This non-monotonic relationship between loss and task performance means that a Chinchilla-optimal model may be suboptimal for specific target applications. If you are training a model for a particular use case with known benchmark requirements, it may be worth investigating whether those benchmarks respond more to model size or data quantity before applying the 20:1 ratio mechanically.
Inference Considerations
As noted earlier, Chinchilla scaling ignores inference costs. For models that will serve billions of queries, the accumulated inference cost may justify training a somewhat larger model to reduce the number of sequential operations at serving time. The upcoming Inference Scaling chapter addresses these trade-offs.
More broadly, speculative decoding and quantization have changed inference economics in ways that the original Chinchilla framework could not anticipate. Sparse attention introduces another change. A 70B model with 4-bit quantization can run on hardware that would not support a full-precision 70B model, changing the cost calculus significantly. As these serving-side optimizations mature, the "right" model size for a given deployment may diverge further from the Chinchilla compute-optimal size, favoring either larger models (if serving optimizations can compensate) or smaller models (if deployment constraints are severe).
Limitations and Impact
The Chinchilla scaling laws changed how the field approaches language model development, but they come with important limitations worth understanding. Any scaling law derived from empirical experiments inherits the assumptions and constraints of those experiments, and the Chinchilla results are no exception. Recognizing these limitations does not diminish the contribution of the paper. Rather, it enables us to apply its insights more precisely and to understand where follow-up research has extended or modified the original framework.
The most significant limitation is the focus on training compute while excluding other costs. While a Chinchilla-optimal 70B model achieves better loss than an undertrained 280B model for the same training compute, the larger model might be preferable if inference costs are negligible (for example, in research settings with limited deployment). The optimal allocation depends on the full lifecycle of the model, not just training.
Additionally, the Chinchilla analysis assumes access to unlimited unique training data. In practice, high-quality data is finite, and training on 20 tokens per parameter for a 100B+ model requires trillions of tokens. This has pushed the field toward synthetic data generation, multi-epoch training strategies, and careful data curation. These topics extend beyond the original Chinchilla analysis, which was designed for the single-epoch regime where unique tokens are plentiful enough to avoid repetition entirely. Using perplexity as the optimization target also has limits. Some capabilities appear to require model scale beyond what perplexity improvements would suggest. A 10B model trained to Chinchilla-optimal perplexity may still lack capabilities present in a 100B undertrained model, even if the smaller model has better loss. The relationship between loss and emergent capabilities is still an active research area.
A subtler limitation is that the Chinchilla experiments focused exclusively on a decoder-only transformer architecture trained with next-token prediction. Different architectures or training objectives might have different optimal ratios. Mixture-of-experts models, for instance, have more total parameters than their effective compute cost suggests, which complicates applying the Chinchilla ratio directly. Similarly, models trained with masked language modeling or encoder-decoder objectives process tokens differently, so the 6ND compute estimate may not apply as directly. When extending Chinchilla reasoning to non-standard architectures, practitioners should verify that the compute accounting and scaling exponents still apply in their setting.
Despite these limits, Chinchilla's impact has been significant. The paper showed that careful empirical analysis of scaling can overturn conventional wisdom and produce immediately useful insights. Labs rapidly adopted the 20:1 guideline, leading to models like LLaMA that achieve strong performance with far fewer parameters than GPT-3 era models. The shift toward data-efficient architectures and high-quality data curation traces directly to the Chinchilla insight that data matters as much as model size.
The paper's methodological contribution may be as important as its specific findings. The use of three independent experimental approaches to triangulate the same result established a template for rigorous scaling research. Subsequent papers on topics like data quality, instruction tuning, and reasoning capabilities have adopted similar multi-method validation approaches, making the field's empirical base more reliable. Chinchilla also made the conversation about compute efficiency precise enough to be actionable: instead of vague claims about the importance of data, practitioners now have a specific quantitative target (20 tokens per parameter) that they can check their training runs against. This concreteness, grounded in careful experiments rather than theoretical arguments, is what enabled rapid adoption across the industry and established Chinchilla scaling as the standard reference for compute-optimal training.
Summary
The Chinchilla scaling laws changed how we think about how to allocate compute between model size and training data. The key findings are:
- Balanced scaling: Model parameters and training tokens should grow at roughly equal rates as compute increases, not with the parameter-heavy allocation suggested by Kaplan
- The 20:1 ratio: Compute-optimal training uses approximately 20 tokens per parameter, meaning a 10B parameter model should see about 200B tokens
- Previous models were undertrained: GPT-3 and similar models trained on far fewer tokens per parameter than optimal, leaving performance gains unrealized
- Three independent methods: The Chinchilla team verified their findings using fixed model sizes, iso-FLOP curves, and parametric loss fitting, all yielding consistent results
- The parametric loss form: The three-term equation captures how parameters and data contribute diminishing returns while an irreducible entropy floor sets the minimum achievable loss
The practical implications were immediate: smaller models trained on more data could match or exceed the performance of much larger undertrained models, while being cheaper to deploy. This shift toward compute-optimal training continues to influence model development, though new considerations around data constraints and inference scaling add important detail to the original Chinchilla prescription.
Looking ahead, the Chinchilla framework provides a foundation but not a complete answer. Subsequent research has addressed data-constrained regimes where the 20:1 ratio cannot be achieved, inference-optimal training that accounts for deployment costs, and the question of whether emergent capabilities obey the same scaling relationships as perplexity. The original Chinchilla result will likely be refined further as new architectures and training methods are explored. What will not change is the fundamental lesson: training compute is a resource to be allocated carefully between model capacity and data exposure, and the two must scale together to achieve optimal results. Every future scaling analysis builds on this insight, whether it confirms the 20:1 ratio, refines it, or identifies conditions under which it breaks down.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Chinchilla scaling laws and compute-optimal training.
Chinchilla Scaling Laws
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!