Power Laws in Deep Learning: Understanding Neural Scaling

Michael BrenndoerferOctober 21, 202569 min read

Part of Language AI Handbook

Explains how power laws govern neural network scaling. Topics include log-log analysis, fitting techniques, and how to predict model performance at any scale.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Power Laws in Deep Learning

Throughout this book, we've built increasingly advanced language models, from simple n-gram models through BERT and GPT architectures to modern designs like LLaMA and Mistral. At each step, we made architectural choices: how many layers, how large a hidden dimension, how much training data to use. Most of those choices were made through intuition, ablation studies, or by following prior work. But a more basic question lurks beneath all of them: what governs model performance as we scale up? If we double the parameters, how much better does the model get? If we multiply the data tenfold, what happens to the loss?

For most of machine learning's history, these questions had no principled answer. Researchers knew bigger was often better, but "often" and "better" were vague. You could train a model with 100M parameters and then one with 1B parameters, and you could observe that the larger model performed better, but you couldn't predict by how much. Every large training run was a costly gamble. Teams at large research labs would commit enormous compute budgets to experiments whose outcomes were unknown, because they lacked a mathematical framework for forming an informed expectation. The result was a field that advanced partly through systematic experimentation, but also substantially through luck, intuition, and the willingness to spend money on gambles.

The answer, it turns out, lies in a deceptively simple mathematical relationship that governs neural network scaling: the power law. Power laws are not a new invention specific to deep learning. They appear in physics and economics, in biology and linguistics, and even in city population distributions. What makes them remarkable in the context of language models is how consistently and reliably they hold across many orders of magnitude of model scale. A relationship you observe at 10 million parameters predicts, with startling accuracy, what will happen at 100 billion parameters. This predictive power has transformed how teams approach large-scale model development, converting expensive experiments into measured extrapolations and turning resource allocation from guesswork into principled optimization.

Understanding power laws transforms how we think about training language models. Instead of guessing at the right model size or dataset, we can predict performance across orders of magnitude, make informed resource allocation decisions, and understand the basic constraints that govern what's possible with current approaches. Before you invest millions of dollars and months of compute time in a training run, you can run small-scale experiments, fit a power law, and read off a prediction for the large model's loss. The prediction is not perfect, but it's accurate enough to be useful. The ability to test an idea at 1% of the cost and then extrapolate reliably to full scale has changed what it means to do empirical research in large-scale ML.

Power laws also clarify the economics of model development in a way that informs strategic decisions. If the loss-versus-parameters relationship has a small exponent, doubling model size buys relatively little improvement, and you should think carefully about whether scaling is the right lever to pull. If the exponent is large, scaling is highly efficient, and additional investment in parameters or data will yield substantial returns. This framing moves scaling decisions from "spend as much as you can and hope" to something closer to quantitative investment analysis.

This chapter establishes the mathematical foundations of power laws, explores why they appear so universally in deep learning, and develops the intuition you'll need to interpret the scaling laws we'll examine in subsequent chapters. We'll move from the basic mathematical definition through the key log-log transformation trick, connect power laws to Zipf's law in language, and build practical skills for fitting and applying these relationships. By the end, you'll have the vocabulary and intuition to read scaling law papers with real understanding, not just familiarity with the formulas.

Historical Context

Power laws were first studied systematically by the Italian economist Vilfredo Pareto in the 1890s, who noticed that roughly 80% of land in Italy was owned by 20% of the population. This "Pareto principle" is a consequence of an underlying power law distribution. The physicist George Kingsley Zipf applied similar analysis to word frequencies in 1935, discovering what we now call Zipf's law. In deep learning, the systematic study of neural scaling laws began with the work of Kaplan et al. in 2020, who published "Scaling Laws for Neural Language Models," and was substantially revised by Hoffmann et al. in 2022 with "Training Compute-Optimal Large Language Models." These papers established that power law relationships hold across many orders of magnitude as compute and data increase and parameter counts grow, making rigorous prediction possible for the first time in large-scale model training. The Kaplan et al. paper was particularly important because it gave the field a common mathematical language for discussing scale: instead of saying "bigger is better," researchers could now say "the parameter-scaling exponent is approximately 0.076, and each order of magnitude of parameters reduces loss by a factor of 10−0.076≈0.8410^{-0.076} \approx 0.84."

What is a Power Law?

Before we can understand how neural networks scale, we need to establish the mathematical vocabulary that describes scaling relationships. The concept we're after is the power law, a particular type of functional relationship that, as you'll soon see, captures something needed about how complex systems behave as they grow. The power law is a mathematical framework for expressing empirical scaling laws, and fluency with it is a prerequisite for understanding every major result in the scaling laws literature.

A power law describes a relationship where one quantity varies as a power of another. This might sound abstract at first, so let's ground it with an intuition. Think of a power law as the mathematical description of how "bang for your buck" changes with scale. When you're small, every additional resource buys a lot of improvement. As you grow larger, each additional resource buys less improvement, but the relationship between resources and improvement follows a predictable, consistent pattern. The power law makes that pattern precise. It's like the difference between a vague promise ("you'll get better with practice") and a quantitative guarantee ("for every tenfold increase in practice hours, your error rate drops by 30%"). The power law gives the quantitative guarantee.

Imagine you're measuring how some output changes as you adjust an input. In the simplest case of a linear relationship, doubling the input doubles the output. The world of power laws is different and more interesting. If you double the input, the output doesn't simply double. Instead, it changes by a fixed multiplicative factor that depends on a special number called the exponent. This exponent is the heart of a power law: it tells you exactly how sensitive the output is to changes in the input. If the exponent is −0.1-0.1, doubling the input multiplies the output by 2−0.1≈0.9332^{-0.1} \approx 0.933. If the exponent is −0.5-0.5, doubling the input multiplies the output by 2−0.5≈0.7072^{-0.5} \approx 0.707. The magnitude of the exponent directly governs how quickly the relationship changes.

Power Law

A power law is a functional relationship of the form y=axby = ax^b, where aa is a constant (the coefficient), bb is the exponent (or scaling exponent), and xx and yy are the related quantities. The exponent bb determines how steeply yy changes as xx increases.

The general form is:

y=axby = ax^b

where:

  • yy: the dependent variable (output quantity). This is what we're measuring or predicting. In deep learning contexts, it's often the loss of our model.
  • xx: the independent variable (input quantity). This is what we control or vary: parameters, data, compute, or training steps.
  • aa: the coefficient, which sets the overall scale of the relationship. Think of aa as calibrating where the curve sits vertically. If aa is large, the entire curve shifts upward; if aa is small, it shifts downward. The coefficient captures baseline performance characteristics specific to your architecture, task, or data. The coefficient aa tells you the value of yy when x=1x = 1, since a⋅1b=aa \cdot 1^b = a for any bb.
  • bb: the exponent (or scaling exponent), which determines how steeply yy changes as xx increases. This is the most important parameter for understanding scaling behavior. The exponent tells us whether improvements come quickly or slowly as we add resources.

When bb is negative, we often write the relationship as:

y=axαy = \frac{a}{x^{\alpha}}

where α=−b>0\alpha = -b > 0 is the positive scaling exponent. This alternative notation proves convenient in deep learning contexts because it makes the relationship more immediately interpretable. When written this way, we can see at a glance that as xx increases, yy decreases, which matches our intuition that adding more parameters or data should improve, that is, lower, our loss. The positive exponent α\alpha directly tells us the rate of this improvement. A larger α\alpha means faster improvement; a smaller α\alpha means slower improvement. In the scaling laws literature, you'll frequently see expressions like L∝N−0.076L \propto N^{-0.076} or L∝C−0.050L \propto C^{-0.050}, where NN is parameter count, CC is compute, and the reported number is the positive exponent α\alpha.

Comparing Functional Forms

Now that we have the basic form, let's understand what makes power laws distinctive compared to other mathematical relationships you might encounter. The key distinction is in how each type of function responds to proportional changes in the input. Consider three common functional forms and how they behave as we scale xx from small to large:

Linear: y=ax+cy = ax + c. Doubling xx adds a fixed amount to yy. If you go from 100 to 200 parameters, you get the same absolute improvement as going from 1 billion to 1 billion and 100 parameters. This kind of relationship treats all additions equally, regardless of scale. However, this doesn't match how neural networks behave: there are clearly regions of diminishing returns. Adding one neuron to a 10-neuron network is a much bigger deal than adding one neuron to a billion-neuron network.

Exponential: y=a⋅ebxy = a \cdot e^{bx}. Doubling xx multiplies yy by a fixed factor that grows explosively. Small changes in xx can produce enormous changes in yy. Exponential relationships describe radioactive decay, compound interest, and viral spread, situations where growth feeds on itself. Neural network scaling doesn't show this kind of explosive behavior. If performance scaled exponentially with parameters, a 2×\times increase in model size would produce the same proportional benefit as any other 2×\times increase, regardless of the starting scale. We don't observe this.

Power law: y=axby = ax^b. Multiplying xx by any factor kk multiplies yy by kbk^b. The key insight is that power laws respond to proportional changes, not absolute changes. Whether you're scaling from 10 million to 100 million parameters or from 10 billion to 100 billion parameters, the proportional change is the same (a factor of 10), and therefore the proportional effect on the output is the same.

Think of it this way: a linear relationship is like a salary increase where you always get $10,000 more per year of experience. An exponential relationship is like compound interest, where each year's earnings are reinvested. A power law is like a negotiated raise where each additional year of experience buys a fixed percentage improvement in salary, but that percentage shrinks as you become more experienced. The shrinking-percentage character of power law returns is exactly what we observe in neural network scaling: each additional order of magnitude of compute or parameters helps, but by a bit less than the previous order of magnitude, and the reduction follows a completely predictable mathematical rule.

Out[3]:
Visualization
Line chart on linear axes comparing linear, power law, and exponential growth curves.
Linear scale comparison of three functional forms. The exponential function dominates quickly, while the power law and linear functions grow more slowly.
The same three curves on a logarithmic y-axis revealing their different growth rates.
Same data on a logarithmic y-axis reveals growth rates. The exponential appears as a straight line on the semi-log plot, while the power law curves, and the linear function bends toward flat.

Scale Invariance: The Core Property

The distinctive feature of power laws is scale invariance: the relationship between proportional changes stays constant regardless of where you are on the curve. A 10×\times increase in model parameters produces the same proportional decrease in loss whether you're going from 100M to 1B parameters or from 1B to 10B parameters. This property matters because it means power laws look the same at every scale.

Think of scale invariance like a fractal: zoom in or zoom out, and the mathematical structure is unchanged. If you look at the power law relationship between compute and loss from one vantage point, doubling your compute buys you the same percentage improvement as doubling your compute from any other vantage point. This self-similarity at all scales is not just mathematically elegant. It is the foundation of practical predictability.

Scale invariance is what lets us to make predictions about enormous models based on experiments with much smaller ones. When we run a series of experiments from 1M to 100M parameters, we're not just learning about small models. We're measuring the exponent of a power law that will continue to hold at 100B parameters. The small experiments are effectively sampling from the same underlying mathematical structure that governs the large ones. This is a remarkable fact: it means that studying a 10-million parameter model tells you something quantitatively real about a 100-billion parameter model, even though those models differ by a factor of 10,000 in size and are almost certainly learning qualitatively different internal representations.

Why does this property hold? A formal characterization is that power laws are the only functions that are scale-invariant in the following sense: if f(kx)=g(k)⋅f(x)f(kx) = g(k) \cdot f(x) for all values of kk and xx, then ff must be a power law. No other functional form satisfies this multiplicative scaling property exactly. This makes power laws mathematically special, not just empirically convenient. The requirement is that proportional changes in the input always produce proportional changes in the output, with the proportionality factor depending only on the scale factor kk and not on the absolute value of xx. Proving that power laws are the unique family satisfying this condition is a straightforward exercise in functional equations, but the conceptual point is more important: scale invariance is not a coincidence. It is the defining property of power laws, and any relationship we observe in nature that turns out to be scale-invariant will necessarily be a power law.

Log-Log Linear Relationships

Understanding the mathematical form of power laws is only the beginning. In practice, a transformation helps you identify and validate power laws, then work with them by converting their curved relationships into straight lines. This transformation bridges theory and practice, turning a specialized nonlinear analysis into the familiar, well-understood territory of linear regression. If you understand only one technical tool from this chapter, let it be this one: the log-log transformation is the lens through which all scaling law analysis is conducted.

Taking the logarithm of both sides of y=axby = ax^b reveals the hidden linear structure. Let's step through this derivation carefully, as each step illuminates something important:

log⁡y=log⁡(axb)=log⁡a+log⁡(xb)(product rule: log⁡(AB)=log⁡A+log⁡B)=log⁡a+blog⁡x(power rule: log⁡(xb)=blog⁡x)\begin{aligned} \log y &= \log(ax^b) \\ &= \log a + \log(x^b) \quad \text{(product rule: } \log(AB) = \log A + \log B \text{)} \\ &= \log a + b \log x \quad \text{(power rule: } \log(x^b) = b \log x \text{)} \end{aligned}

where:

  • log⁡y\log y: the logarithm of the dependent variable
  • log⁡a\log a: a constant term that becomes the y-intercept in the transformed space
  • bb: the power law exponent, which becomes the slope of the line
  • log⁡x\log x: the logarithm of the independent variable

To see the linear structure most clearly, introduce substitutions Y=log⁡yY = \log y and X=log⁡xX = \log x. The equation becomes:

Y=(log⁡a)+b⋅XY = (\log a) + b \cdot X

This is unmistakably the equation of a straight line Y=mX+cY = mX + c in slope-intercept form, where the slope m=bm = b is the power law exponent and the intercept c=log⁡ac = \log a encodes the coefficient. Each component has a clear geometric interpretation:

  • log⁡y\log y: the logarithm of the dependent variable, which becomes our new vertical axis in the transformed space
  • log⁡a\log a: a constant term that becomes the y-intercept, the value where the line crosses the vertical axis when log⁡x=0\log x = 0, that is, when x=1x = 1
  • bb: the power law exponent, which becomes the slope of our line, telling us how many units log⁡y\log y changes for each unit change in log⁡x\log x
  • log⁡x\log x: the logarithm of the independent variable, which becomes our new horizontal axis

Why does this matter so much? Because straight lines are easy to work with. We have centuries of well-developed tools for analyzing linear relationships, fitting them to data, and extrapolating from the result. By transforming power laws into lines, we gain access to this entire toolkit: linear regression, confidence intervals, residual analysis, hypothesis testing. The slope of a line is one of the most intuitive quantities in mathematics: it tells you the rate of change, and we have immediate intuitions about what different slopes look like and what they mean.

The slope interpretation is particularly useful for comparing scaling behaviors. If architecture A has a parameter-scaling slope of −0.10-0.10 in log-log space and architecture B has a slope of −0.07-0.07, architecture A's loss drops faster with increasing parameters. The difference between these slopes represents a 43% advantage in scaling efficiency (0.10 vs. 0.07) that compounds sharply over large scale ranges.

This transformation has three concrete practical implications that you'll use repeatedly when working with scaling data:

  • Detection: If your data forms a straight line on log-log axes, where both axes are logarithmic, you've found a power law. This gives you a simple visual test. Plot your scaling data with logarithmic scales on both axes and look for linearity. No complex statistical test is needed at first.
  • Parameter estimation: The slope gives you the exponent directly. You don't need specialized nonlinear optimization. Basic linear regression on log-changed data is enough and more reliable.
  • Extrapolation: Linear extrapolation in log space corresponds to power law extrapolation in linear space. We can extend a straight line with confidence. The log-log transformation lets us translate that confidence to power law predictions across orders of magnitude beyond our observed data.

Worked Example: Computing Scale from Exponent

Let's make each of these steps concrete with a worked numerical example. Suppose loss LL decreases with compute CC according to:

L=5.4⋅C−0.05L = 5.4 \cdot C^{-0.05}

where:

  • LL: the model's loss (lower is better, no units)
  • CC: compute budget in FLOPs (floating-point operations)
  • 5.45.4: the coefficient, representing the loss we would observe at C=1C = 1 FLOP (an entirely hypothetical anchor point with no physical interpretation, but a useful mathematical reference)
  • −0.05-0.05: the scaling exponent, with the negative sign showing loss decreases as compute increases

Before diving into the calculation, consider what this equation is telling us qualitatively. The coefficient 5.4 anchors the curve but isn't real on its own, since no real model trains with one floating-point operation. The exponent −0.05-0.05 carries the prediction: loss decreases slowly as compute increases, and the small magnitude (0.05) tells us the improvement per order of magnitude is modest. A large lab running this kind of experiment would typically see the curve covering 10 or more orders of magnitude of compute, from small ablation runs to the largest training runs they can afford.

Taking base-10 logarithms of both sides:

log⁡10L=log⁡10(5.4⋅C−0.05)=log⁡10(5.4)+(−0.05)log⁡10C=0.73−0.05⋅log⁡10C\begin{aligned} \log_{10} L &= \log_{10}(5.4 \cdot C^{-0.05}) \\ &= \log_{10}(5.4) + (-0.05) \log_{10} C \\ &= 0.73 - 0.05 \cdot \log_{10} C \end{aligned}

In log space, our line has y-intercept 0.73 and slope −0.05-0.05. Now we can answer a concrete question: if we increase compute by a factor of 10 (one order of magnitude), meaning log⁡10C\log_{10} C increases by exactly 1, how much does loss change?

Δlog⁡10L=−0.05×1=−0.05\Delta \log_{10} L = -0.05 \times 1 = -0.05

What does a change of −0.05-0.05 in log⁡10L\log_{10} L mean for actual loss? We work backwards: if log⁡10L\log_{10} L decreases by 0.050.05, then LL is multiplied by:

10−0.05≈0.89110^{-0.05} \approx 0.891

The new loss is approximately 89.1% of its previous value. This is an 11% reduction in loss for every tenfold increase in compute. Each order of magnitude of compute reduces loss by about 11%. This fixed proportional improvement per order of magnitude is the signature of power law scaling, and it's the kind of prediction that makes planning possible.

Why does this formula make sense? Notice that the exponent −0.05-0.05 appears as a conversion factor between "orders of magnitude of compute" and "fractional reduction in loss." The negative sign enforces the inverse relationship. The magnitude (0.05) determines the efficiency of the conversion. If you had a larger exponent in magnitude, say −0.10-0.10, each order of magnitude of compute would buy you more improvement. Comparing scaling laws from different architectures or training configurations therefore amounts to comparing these exponents.

Let's extend this to a multi-step planning exercise. Suppose you have access to a compute budget that allows you to do a training run that costs 102110^{21} FLOPs. You've run small experiments at 101710^{17} and 101810^{18} FLOPs, measured losses of approximately 3.0 and 2.8 respectively, and fit the power law above. What loss should you expect?

Step 1: Verify the fit. Using the formula L=5.4⋅C−0.05L = 5.4 \cdot C^{-0.05}, compute predicted losses at your observed scales.

At C=1017C = 10^{17}: L=5.4⋅(1017)−0.05=5.4⋅10−0.85=5.4⋅0.141≈0.76L = 5.4 \cdot (10^{17})^{-0.05} = 5.4 \cdot 10^{-0.85} = 5.4 \cdot 0.141 \approx 0.76.

Wait, that doesn't match the observed 3.0. This means our example power law L=5.4⋅C−0.05L = 5.4 \cdot C^{-0.05} needs re-anchoring to the observed data. Let's instead work with the observed data directly. From the two observations:

log⁡10(3.0)=log⁡10a+b⋅17log⁡10(2.8)=log⁡10a+b⋅18\begin{aligned} \log_{10}(3.0) &= \log_{10} a + b \cdot 17 \\ \log_{10}(2.8) &= \log_{10} a + b \cdot 18 \end{aligned}

Subtracting: log⁡10(2.8)−log⁡10(3.0)=b⋅1\log_{10}(2.8) - \log_{10}(3.0) = b \cdot 1, so b=log⁡10(2.8/3.0)=log⁡10(0.933)≈−0.030b = \log_{10}(2.8/3.0) = \log_{10}(0.933) \approx -0.030.

Then log⁡10a=log⁡10(3.0)−(−0.030)⋅17=0.477+0.510=0.987\log_{10} a = \log_{10}(3.0) - (-0.030) \cdot 17 = 0.477 + 0.510 = 0.987, so a=100.987≈9.7a = 10^{0.987} \approx 9.7.

At C=1021C = 10^{21}: L=9.7⋅(1021)−0.030=9.7⋅10−0.63=9.7⋅0.234≈2.27L = 9.7 \cdot (10^{21})^{-0.030} = 9.7 \cdot 10^{-0.63} = 9.7 \cdot 0.234 \approx 2.27.

The predicted loss at your large training run is approximately 2.27. You've made a specific, quantitative, checkable prediction before spending a cent on the expensive run. This is the practical power of power law scaling.

Let's visualize this relationship concretely:

In[4]:
Code
import numpy as np

# Generate power law data: L = 5.4 * C^(-0.05)
compute = np.logspace(15, 24, 50)  # From 10^15 to 10^24 FLOPs
loss = 5.4 * np.power(compute, -0.05)
Out[5]:
Visualization
Curved line on linear plot showing loss decreasing with compute.
Power law relationship on linear axes shows a characteristic convex curve. The steep initial decline flattens as compute grows, which makes it difficult to read off precise scaling behavior.
Straight line on log-log plot with negative slope.
The same data on log-log axes reveals the underlying linear structure. The perfectly straight line confirms the power law form and its slope directly encodes the exponent -0.05.

The left plot shows the curved relationship typical of power laws on linear axes, while the right plot reveals the elegant simplicity of the log-log representation. The slope of the line in log-log space gives us the exponent −0.05-0.05 directly. Notice how the curve that appeared complex on linear axes becomes a perfectly straight line once we transform both axes to logarithmic scales. This visual simplicity shows the mathematical simplicity we derived above: power laws are linear relationships in disguise, and the log-log transformation reveals that hidden linearity.

Log-log plots are a mathematical lens rather than a visualization choice. They make the structure of the data visible. When researchers at AI labs plot training curves or scaling results, they almost universally use log-log or semi-log axes because linear axes obscure the simple relationship that makes the data interpretable and extrapolable.

Power Laws in Language

Power laws aren't new to language modeling. In fact, you've already encountered one of the most famous power laws in language earlier in this book: Zipf's law. This connection is more than historical curiosity. It reveals that power law behavior is woven into the very fabric of language itself, predating neural networks by decades and suggesting something deep about how natural language is structured. The fact that language itself exhibits power law structure is both a fascinating empirical observation and a clue about why language models might also scale according to power laws.

When we discussed word frequency distributions in the tokenization chapters, we noted that the frequency of words follows a remarkably consistent pattern. If you rank words by frequency, the rr-th most common word appears with frequency proportional to 1/r1/r. Formally:

f(r)∝1rαf(r) \propto \frac{1}{r^{\alpha}}

where:

  • f(r)f(r): the frequency of the word at rank rr, meaning how many times it appears per million tokens
  • rr: the rank of the word (1 = most common, 2 = second most common, and so on)
  • α\alpha: Zipf's exponent, approximately 1 for most natural language corpora
  • ∝\propto: "proportional to," meaning the relationship holds up to a constant factor that depends on corpus size

The word "the" appears vastly more often than "cat," which appears vastly more often than "syzygy." This isn't a gradual decline. The most common words dominate text, while the long tail of rare words contributes relatively little to token counts. What's striking is that this distribution follows the same mathematical form across every language, every domain, every time period that linguists have studied. English novels, French newspapers, scientific papers, social media posts: all Zipfian. The universality is itself remarkable and suggests that Zipf's law shows something basic about how humans use language for communication, rather than being an artifact of any particular language or domain.

The exponent α≈1\alpha \approx 1 has a striking implication that we can trace through arithmetically. If the most common word appears 1 million times per million tokens (by definition, it accounts for 100% of tokens, which is impossible, so let's say it appears 70,000 times per million tokens as a more realistic estimate), then:

  • Rank 10 word: 70,000/101=7,00070,000 / 10^1 = 7,000 occurrences
  • Rank 100 word: 70,000/1001=70070,000 / 100^1 = 700 occurrences
  • Rank 1,000 word: 70,000/10001=7070,000 / 1000^1 = 70 occurrences

Each factor of 10 in rank produces a factor of 10 decrease in frequency. This is scale invariance in action: the same proportional pattern repeating across the entire frequency spectrum. There is nothing special about rank 100 versus rank 1000: both follow the same law, and the law is the same whether you're looking at the top 10 words or the top 10 million words.

Zipf's law connects directly to the vocabulary problem we explored in the tokenization chapters. The heavy tail of rare words means that no matter how large your vocabulary, you'll encounter words not seen in training. This motivated subword tokenization approaches like BPE and WordPiece, which handle rare words by decomposing them into more frequent subunits. In this way, understanding the power law structure of language directly informed the design of modern tokenization systems. The mathematical structure we're studying now has practical engineering consequences that extend far beyond the abstract mathematical formalism.

There's an even deeper connection worth considering. If language itself exhibits power law structure, and language models are trained to model language, it's perhaps not so surprising that the models' learning dynamics also follow power laws. The statistical structure of the training data and the statistical structure of the model's performance improvements may be related at a basic level. A model trained to predict the next token in a Zipfian distribution faces a problem where common patterns are easy to learn and rare patterns are hard, and the difficulty of prediction tasks is itself distributed as a power law. The aggregate loss over this distribution, as the model improves, could plausibly follow a power law as a consequence.

Connections to Information Theory

Zipf's law has a deep connection to information theory. The famous result of information theory states that optimal codes assign shorter code words to more probable symbols and longer code words to less probable symbols. A language with Zipfian word frequencies is, in a specific technical sense, a language that has been optimized for efficient communication: common words are short (think "a", "I", "the") and rare words are long (think "syzygy", "sesquipedalian"). This near-optimal structure suggests that natural language has been shaped by evolutionary and cultural pressures toward communicative efficiency. The same information-theoretic considerations that explain why language is Zipfian may also explain why language models scale as power laws: they are progressively extracting structure from a data source that is itself organized according to information-theoretic principles.

In[6]:
Code
from collections import Counter

# Generate a synthetic corpus with Zipfian word-frequency distribution.
# True Zipf's law: frequency(rank) = C / rank^alpha, alpha ~ 1.
# We use 5000 distinct word types with counts drawn from this distribution,
# then add mild Gaussian noise so the log-log fit is not artificially perfect.
np.random.seed(17)
vocab_size = 5000
ranks_gen = np.arange(1, vocab_size + 1)
# Approximate English: most-common word ~70,000 occurrences per million tokens
base_freq = 70_000
true_alpha = 1.0
true_counts = (base_freq / ranks_gen**true_alpha).astype(float)
# Add ~8% multiplicative noise to simulate real corpus variation
noise = np.random.lognormal(mean=0.0, sigma=0.08, size=vocab_size)
noisy_counts = np.maximum(1, np.round(true_counts * noise)).astype(int)

# Re-rank by observed frequency (as we would after counting a real corpus)
sort_idx = np.argsort(noisy_counts)[::-1]
frequencies = noisy_counts[sort_idx]
ranks = np.arange(1, vocab_size + 1)

# Build a Counter-compatible structure for the downstream print cell
ranked_words = [
    (f"word_{i + 1}", int(frequencies[i])) for i in range(vocab_size)
]
word_counts = Counter(dict(ranked_words))
Out[7]:
Visualization
Scatter plot showing word frequencies following Zipf's law pattern.
Word frequency versus rank on log-log axes shows Zipf's law across a 5,000-word synthetic vocabulary. The approximately linear scatter with fitted slope near -1.0 confirms the power law form, and the slight noise shows realistic corpus variation.
Out[8]:
Console
Top 5 words by frequency:
  word_1: 71564
  word_2: 30174
  word_3: 24528
  word_4: 19179
  word_5: 15211

Fitted Zipf exponent: 1.00

The output shows the most frequent word appearing far more often than words in the long tail, exactly as Zipf's law predicts. The fitted Zipf exponent close to 1 confirms Zipf's law holds across this 5,000-word vocabulary distribution. This same mathematical pattern, a power law relationship, governs how neural network performance scales with resources.

Other Power Laws in Language and NLP

Zipf's law is the most famous, but it's not the only power law in language. The Heap's law (also called Herdan's law) describes how vocabulary size grows with corpus size: as you add more documents to a corpus, the number of distinct word types grows as a power law of the total token count, with an exponent typically between 0.4 and 0.6. This has direct implications for tokenization strategy: no finite vocabulary can fully cover a growing corpus, which is part of the motivation for subword and character-level approaches.

Sentence length distributions, paragraph length distributions, document length distributions in web crawls: all exhibit approximately power law behavior over substantial ranges. The same pattern appears at multiple levels of linguistic organization, from individual characters up to full documents. This multi-scale power law structure in language suggests that the statistical regularities neural networks must learn are themselves organized hierarchically and scale-invariantly. The model that can handle short-range patterns must also handle long-range patterns, and the transition between scales follows the same mathematical rule at every level.

Fitting Power Laws

To use power laws predictively, we need to estimate the parameters aa and bb from data. This is where the log-log transformation proves most useful in practice. Instead of wrestling with nonlinear curve fitting, which is sensitive to initialization, prone to local minima, and statistically tricky, we can use the simple, well-understood tools of linear regression on log-transformed data.

The strategy is straightforward: take logs of both the inputs and outputs, then apply ordinary linear regression. The slope of the resulting fit is the exponent bb, and the intercept is log⁡a\log a. This approach has several practical advantages over directly fitting y=axby = ax^b with nonlinear methods. Most importantly, it is numerically stable across the wide ranges typical of scaling experiments (compute budgets spanning 15 orders of magnitude, model sizes spanning 5 or more orders of magnitude). Nonlinear optimization applied directly to these ranges would require careful initialization and often fails to converge reliably.

Given data points (xi,yi)(x_i, y_i), we transform to (log⁡xi,log⁡yi)(\log x_i, \log y_i) and fit the linear model:

log⁡y=log⁡a+blog⁡x\log y = \log a + b \log x

where:

  • log⁡y\log y: the log-transformed dependent variable, computed from our measured loss values
  • log⁡x\log x: the log-transformed independent variable, computed from our resource measurements (parameters, data, compute)
  • log⁡a\log a: the y-intercept in log space, recovered from the linear regression
  • bb: the slope, which equals the power law exponent directly

Comparing this to the standard linear form Y=mX+cY = mX + c makes the connection explicit. The slope mm in our transformed regression equals bb. The intercept cc equals log⁡a\log a, which we convert back to the original coefficient via exponentiation: a=eca = e^c if using natural logarithms, or a=10ca = 10^c if using base-10 logarithms.

This approach works because ordinary least squares regression minimizes the sum of squared residuals. When we apply it to log-transformed data, we're effectively minimizing:

∑i(log⁡yi−log⁡a−blog⁡xi)2=∑i(log⁡yiaxib)2\sum_i \left( \log y_i - \log a - b \log x_i \right)^2 = \sum_i \left( \log \frac{y_i}{a x_i^b} \right)^2

This objective function penalizes relative errors, not absolute errors. A prediction that's off by 10% at loss 3.0 contributes the same squared residual as a prediction that's off by 10% at loss 1.0. This is exactly the right behavior when values span many orders of magnitude: relative errors are the real measure of fit quality.

Pitfalls in Power Law Fitting

Fitting power laws correctly requires care. Several subtle issues can trip up even experienced analysts, and understanding them will help you interpret published scaling results critically.

Use log-transformed data for regression, not raw data. This point deserves emphasis because it's counterintuitive. You might think, "I want to fit y=axby = ax^b, so I should use nonlinear regression to minimize the sum of squared errors in yy." The problem is that this approach overweights large values. If your losses range from 10 at small scales to 0.1 at large scales, the squared errors at the small end (potentially hundreds) will dwarf squared errors at the large end (fractions of 0.01). The optimizer focuses on matching the high-loss region while ignoring the low-loss region. By contrast, fitting in log space minimizes relative errors: a 10% miss at loss 10 contributes as much to the objective as a 10% miss at loss 0.1. This is usually what we want when values span many orders of magnitude.

Consider weighted regression for heteroskedastic data. If measurement uncertainty varies across your data range, weighted regression can improve estimates. For instance, if small-scale experiments are run once while large-scale experiments are run multiple times with averaged results, the large-scale points have lower variance and deserve more weight in the fit. In practice, this issue is common because large training runs are expensive and researchers typically replicate small-scale experiments more than large-scale ones.

Validate with residual analysis. After fitting, examine residuals in log space. If the residuals show systematic patterns, such as being consistently positive at small scales and negative at large scales, this indicates the power law model may not be appropriate. True power law data should produce residuals that scatter randomly around zero with no systematic trend. A curved residual pattern suggests you might be dealing with multiple regimes or a more complex functional form. This residual analysis step is often skipped in published work but is needed for building confidence in the fit.

Check for outliers before fitting. A single outlier in log space can have a disproportionate effect on the estimated slope. When fitting scaling laws, examine whether any data points deviate substantially from the trend before including them in the regression. Common sources of outliers include training runs that were stopped early, runs with suboptimal hyperparameters, or runs where something went wrong with the data pipeline.

Beware of the data selection effect. Researchers typically publish scaling experiments where the power law fit is clean. Experiments where the relationship is noisy or non-power-law are often not published. This means the published literature may systematically overestimate how cleanly power laws hold in practice. When you run your own scaling experiments, expect more noise than the published papers suggest.

Let's implement power law fitting and see these principles in action:

In[9]:
Code
def fit_power_law(x, y):
    """
    Fit a power law y = a * x^b using linear regression in log space.

    Returns: a, b, r_squared
    """
    log_x = np.log(x)
    log_y = np.log(y)

    # Linear regression in log space
    slope, intercept = np.polyfit(log_x, log_y, 1)

    # Convert back to power law parameters
    b = slope
    a = np.exp(intercept)

    # Calculate R-squared
    y_pred_log = intercept + slope * log_x
    ss_res = np.sum((log_y - y_pred_log) ** 2)
    ss_tot = np.sum((log_y - np.mean(log_y)) ** 2)
    r_squared = 1 - ss_res / ss_tot

    return a, b, r_squared

Let's test our fitting function on synthetic data where we know the true parameters. This makes possible us verify that the fitting procedure recovers the correct exponent and coefficient before applying it to real data:

In[10]:
Code
# Generate synthetic scaling data with noise
np.random.seed(42)
model_sizes = np.array([10e6, 50e6, 100e6, 500e6, 1e9, 5e9, 10e9])  # Parameters
true_a = 10.0
true_b = -0.076
noise_factor = 0.02  # 2% relative noise

losses = true_a * np.power(model_sizes, true_b)
losses_noisy = losses * (1 + np.random.randn(len(losses)) * noise_factor)
In[11]:
Code
# Fit power law to noisy data
a_fit, b_fit, r_squared = fit_power_law(model_sizes, losses_noisy)
Out[12]:
Console
True parameters: a = 10.000, b = -0.0760
Fitted parameters: a = 9.867, b = -0.0748
R-squared: 0.9934

The fitted parameters closely match the true values, with the coefficient recovered accurately and the exponent nearly identical. The R-squared value above 0.99 indicates an excellent fit. This shows that our log-space regression approach effectively recovers power law parameters even with noisy data. The coefficient may deviate slightly because noise interacts with the curve differently at different scales, but the exponent, which controls the slope of the log-log line, is recovered accurately.

Out[13]:
Visualization
Log-log plot with data points and fitted power law line.
Power law fitted to noisy scaling data on log-log axes. The fitted curve (red) closely matches the true relationship (green dashed) within the observed range and extrapolates into the blue-shaded region well beyond any observed data points, illustrating the predictive power of the method.

The fit recovers the true parameters with reasonable accuracy. More importantly, it lets extrapolation: we can predict performance for model sizes we haven't trained yet. The blue shaded region shows where we're extrapolating well beyond observed data, yet the fitted curve closely tracks the true relationship. This predictive power is what makes scaling laws so useful for planning large language model training runs.

To validate that our power law fit is appropriate, we examine the residuals, the differences between observed and predicted values in log space. A good fit produces residuals that scatter randomly around zero with no systematic pattern:

Out[14]:
Visualization
Scatter plot of residuals in log space versus model size, showing random scatter around a zero reference line.
Residual analysis for the power law fit shows log-space residuals versus log model size. The random scatter around zero without any systematic trend confirms the power law model is appropriate for this data range.

The residuals scatter randomly within a narrow band around zero, confirming that our power law model is appropriate for this data. If we saw a curved pattern in the residuals, for example, positive at both extremes and negative in the middle, it would suggest the true relationship isn't a simple power law. The residual plot is your first diagnostic when a scaling law fit looks suspicious.

Power Law Universality

One of the most striking aspects of deep learning scaling is how consistently power laws appear. Empirical studies across different model architectures, different datasets, and different tasks repeatedly find power law relationships. This consistency is remarkable and demands explanation. It would be easy to dismiss if power laws showed up only for transformer language models, or only for a narrow range of scales. The remarkable fact is that they show up in effect everywhere: vision models, code models, protein structure prediction models, reinforcement learning agents, even models that are structurally very different from standard transformers.

Loss scales as a power law with:

  • Number of parameters
  • Amount of training data
  • Total compute budget
  • Training steps (in certain regimes)

This universality goes beyond what you might naively expect. A transformer trained on code, a recurrent network trained on books, and a mixture-of-experts model trained on web text are architecturally very different systems with different inductive biases and training dynamics. Yet they all show power law scaling with similar exponents. Understanding why requires some theoretical investigation, and while the theoretical picture is still incomplete, several perspectives offer partial explanations.

Theoretical Perspectives

Several theoretical perspectives offer partial explanations for power law universality, each illuminating a different facet of the phenomenon. None fully explains power law universality on its own, but together they build a clear picture of why this mathematical form might be basic rather than accidental.

Statistical mechanics analogy. Complex systems with many interacting components often exhibit power law behavior as they approach necessary points, which are phase transitions where system behavior changes qualitatively. Think of water near its boiling point: fluctuations of all scales appear simultaneously, and correlations extend over arbitrarily long distances. This scale-free behavior is a hallmark of criticality. Neural networks, with billions of parameters interacting through nonlinear dynamics, may inhabit a regime analogous to a necessary point in statistical mechanics. The large number of interacting units, combined with training dynamics that push the system toward a particular region of parameter space, may naturally give rise to scale-invariant behavior. Gradient descent, by minimizing loss, may be implicitly seeking necessary points where the network balances between over-fitting to the training data and under-fitting, and criticality naturally produces power law statistics.

Random matrix theory. The Hessian of neural network loss functions, the matrix of second derivatives that describes the local curvature of the loss landscape, shows spectral properties connected to random matrix theory. In large random matrices, eigenvalue distributions often follow power laws. If the curvature of the loss landscape has power law structure in its eigenvalue distribution, this can translate to power law scaling in learning dynamics: some directions in parameter space are learned quickly, others slowly, and the distribution of learning rates follows a power law. The slowly-learned directions correspond to rare, complex patterns in the data, and the distribution of pattern complexities in natural language is itself approximately a power law.

Feature learning dynamics. As networks train, they progressively learn features of increasing complexity. Simple patterns like common words, basic syntactic structures, and high-frequency collocations are learned first. Complex patterns like logical reasoning, rare idiomatic expressions, and cross-sentence coreference come later. The marginal difficulty of learning additional features may increase in a way that produces power law returns on additional capacity. Each new feature is harder to learn than the previous one, but not impossibly harder. This gradual increase in difficulty, distributed across the entire feature space, manifests as power law scaling at the aggregate level. Think of it as a hierarchy of tasks: the model must master each level before benefiting from the next, and the difficulty of each level is greater by a fixed multiplicative factor.

Bias-variance decomposition. The total prediction error of a model can be decomposed into irreducible noise (inherent randomness in labels), approximation error (mismatch between model family and true function), and estimation error (mismatch due to finite training data). Each of these components scales differently with model size and data size. Under certain distributional assumptions about the structure of the function being learned, the sum of these components exhibits power law scaling even if the individual components don't. This perspective suggests that power law behavior might be an emergent property of the interaction among multiple error sources, and that the exponent of the aggregate power law shows a combination of the scaling rates of the individual components.

The key insight from all these perspectives is that power laws emerge when you have many interacting components, each contributing a small but nonzero amount to the overall behavior. Neural networks, with their billions of parameters trained on billions of tokens, satisfy this condition in abundance. Whether the ultimate explanation is rooted in statistical mechanics, random matrix theory, or information theory, all of these frameworks point to the same conclusion: power law scaling is not a coincidence. It is a consequence of the deep mathematical structure of high-dimensional optimization over structured data.

Empirical Evidence Across Architectures

While no single theory fully explains power law universality, the empirical record is reliable. Kaplan et al. (2020) documented power law scaling across five decades of compute for transformer language models. Their measurements showed clean power law relationships with R2R^2 values above 0.99 across many orders of magnitude. Subsequent work verified similar relationships in vision models, protein structure prediction, code generation, and even reinforcement learning settings. The universality is not limited to language.

One particularly striking demonstration is that models with very different architectures, trained on different data, for different tasks, sometimes yield nearly identical scaling exponents. This suggests the exponent captures something basic about the difficulty of the learning problem, rather than being an artifact of a specific architectural choice. When two very different models, trained by different teams on different data, yield the same scaling exponent, it strongly implies that exponent shows a property of the task itself, not of either particular model. We'll see in the next chapter how Kaplan et al. measured these exponents precisely for transformer language models.

The consistency of scaling exponents across architectures also has a practical implication: architectural search conducted at small scale is more likely to transfer to large scale if the two architectures have similar scaling exponents. An architecture that performs better than a baseline at 100M parameters but has a slightly smaller scaling exponent may be overtaken by the baseline at 10B parameters. This interaction between absolute performance and scaling efficiency is one of the subtle but important considerations in modern architecture design.

Building Power Law Intuition

Understanding power laws conceptually helps you reason about scaling decisions without constantly returning to the equations. Mathematical formulas are needed, but developing intuition about what those formulas mean allows you to think fluently about scaling in conversation, when reading papers, or when making resource allocation decisions. Here are the key intuitions to internalize.

The most important shift in perspective that power law thinking demands is moving from absolute quantities to ratios and proportions. When you think about linear relationships, you naturally think about how much you're adding: "we added 10 billion parameters." When you think about power law relationships, you should think about the ratio: "we multiplied our parameter count by 10." The absolute quantity is almost never the needed one. The ratio is what the power law responds to, and training your intuition to think in ratios is the single most practical habit you can develop for scaling law work.

Diminishing but Persistent Returns

A power law of the form L=aN−αL = aN^{-\alpha} with α>0\alpha > 0 means each additional parameter contributes less than the previous one, but never zero. Let's make each term explicit:

  • LL: the loss we're trying to minimize
  • NN: the number of model parameters
  • aa: a constant that sets the overall scale
  • α\alpha: the positive scaling exponent, with α>0\alpha > 0 so diminishing returns

The negative exponent ensures that as NN increases, LL decreases, but the rate of decrease slows as NN grows larger. The millionth parameter is worth less than the first, but it's not worthless. This gradual decline contrasts sharply with exponential decay, where returns plummet rapidly after an initial period. Power law diminishing returns are persistent. There is always some benefit to scaling further, even if the marginal gain shrinks.

Think of it as a highway that has no speed limit but where the speed of traffic (your loss improvement) follows a known curve. You can always go faster by adding more lanes (parameters), but each additional lane contributes a bit less than the previous one. You never reach a point where additional lanes are completely useless. You just get diminishing returns forever. This means that the question "how much should we scale?" has no absolute answer based on the power law alone. It depends on your cost function: how much compute you can afford, how much improvement is worth paying for, and what alternative uses exist for your resources.

Out[15]:
Visualization
Semi-log plot showing loss declining smoothly as model parameter count grows from millions to trillions.
Loss decreases as model parameter count grows from millions to trillions. The semi-log plot shows the smooth, persistent decline predicted by power law scaling with no sudden plateau or drop-off.
Log-log plot of marginal improvement per parameter, showing a steadily declining curve that never reaches zero.
Marginal improvement per additional parameter on a log-log plot. The steadily declining curve never reaches zero, confirming that there is always some benefit to adding more parameters, even if it shrinks with scale.

Multiplicative Thinking

Power law thinking is multiplicative rather than additive. Going from 1B to 2B parameters gives the same proportional improvement as going from 10B to 20B parameters. In both cases, you're doubling the parameter count, and power laws respond to proportional changes, not absolute ones.

When planning scale-ups, think in terms of 2×\times, 10×\times, or 100×\times increases rather than adding fixed numbers of parameters. Asking "what if we add another billion parameters?" uses the wrong framing. It describes an absolute change that becomes increasingly irrelevant as models grow larger. Asking "what if we double our model size?" aligns with power-law behavior and produces consistent predictions regardless of starting scale. A team planning to scale from 10B to 20B parameters has made the same proportional commitment as a team scaling from 1B to 2B, and the power law predicts the same proportional improvement in both cases.

This multiplicative mindset extends to how you think about data and compute. If you have a 100B token dataset and consider doubling it to 200B tokens, power law thinking says: what exponent governs data scaling, and what does a 2×\times increase in data buy us? The answer to that question is the same whether you're going from 100B to 200B tokens or from 1T to 2T tokens. The proportional change is identical, and power laws respond to proportional changes.

The multiplicative framing also clarifies comparisons between different resource types. If doubling model size reduces loss by 5% and doubling data reduces loss by 7%, you should prefer data scaling, all else being equal. But you need to think about what doubling each resource costs, because the cost of doubling parameters (manufacturing more compute, memory, communication bandwidth) and the cost of doubling data (crawling, filtering, tokenizing more text) are very different. The power law tells you the return; cost accounting tells you the investment; together they tell you where to allocate resources.

The Exponent is Everything

The exponent is the most important number in any power law for a given scaling regime. Small differences in exponents compound into large differences at scale.

Consider two models with different parameter scaling exponents. If L∝N−0.07L \propto N^{-0.07}, doubling parameters multiplies loss by:

2−0.07≈0.9532^{-0.07} \approx 0.953

This is a 4.7% reduction. If instead L∝N−0.1L \propto N^{-0.1}, doubling parameters multiplies loss by:

2−0.1≈0.9332^{-0.1} \approx 0.933

This is a 6.7% reduction. A difference of 0.03 in the exponent might seem minor, but these differences compound sharply over large scale ranges. Over a 1000×\times scale-up:

  • Exponent −0.07-0.07: loss becomes 1000−0.07≈0.631000^{-0.07} \approx 0.63, a 37% reduction
  • Exponent −0.10-0.10: loss becomes 1000−0.1≈0.501000^{-0.1} \approx 0.50, a 50% reduction

A difference of 0.03 in exponent turns into a 13 percentage point difference in total improvement over three orders of magnitude of scaling. If you're planning a training run that spans six orders of magnitude, the difference would be even more large. Understanding your specific power law exponent is important for efficient resource allocation.

The exponent is also the quantity that distinguishes architectures. When researchers claim that architecture A scales better than architecture B, what they often mean is that A has a larger-magnitude exponent: a given proportional increase in parameters buys more improvement for A than for B. The coefficient aa matters less, because the coefficient affects where you start, but the exponent determines where you end up after scaling. Two architectures with the same coefficient but different exponents will produce different models after a 1000×\times scale-up, even if they started from the same loss. The architecture with the larger exponent will win decisively at large scale even if it's slightly worse at small scale.

This has a concrete implication for architecture search: when evaluating a new architectural idea, measure its absolute performance at a convenient scale and how its scaling exponent compares to the baseline. An architectural change that looks promising at small scale may have a slightly worse exponent and underperform at large scale. Conversely, an architectural change that looks modest or even slightly harmful at small scale may have a better exponent and sharply outperform at large scale.

Prediction from Small-Scale Experiments

Power laws make prediction possible in a way that arbitrary curves don't. Unlike a function specified by many parameters, a power law is fully specified by just two numbers: aa and bb. If you measure performance at two scales, you have enough information to determine both parameters and therefore the entire curve.

In practice, you'd want more than two points to get a reliable estimate with confidence intervals. But the main point is that small-scale experiments inform large-scale predictions. If you train a 100M parameter model and a 1B parameter model, you can fit a power law and predict performance for a 100B parameter model. That's extrapolating 100×\times beyond your largest observation, based on just two data points. The prediction is not exact, but in the literature, researchers have been surprised by how accurate these extrapolations turn out to be.

This predictive power lets rational planning for expensive training runs. Instead of committing to a 1B parameter model and hoping for the best, you can train a series of models at 10M, 30M, 100M, and 300M parameters, fit a power law, and read off the predicted performance at 1B. If the prediction is promising, you proceed. If not, you investigate whether your architecture or data pipeline can improve the exponent before committing to the large run. This iterative "measure small, predict large, decide" workflow has become standard practice at research labs doing large-scale training.

The workflow also applies to intermediate-scale decisions. Should you train at 1B parameters or 3B parameters? If your power law predicts that 3B will reduce loss by an additional 7% compared to 1B, and you've already determined that this 7% improvement translates to a real performance gain on your downstream task, the decision is quantitatively grounded rather than intuitive. The power law is the bridge between the cost you pay (3×\times more compute) and the benefit you receive (7% lower loss).

Out[16]:
Visualization
Multiple power law curves showing steeper slopes with larger exponents.
Impact of different power law exponents on scaling trajectories from 1 million to 1 trillion parameters. A steeper exponent (larger magnitude, shown in red) delivers substantially more improvement per order of magnitude than a shallow exponent (green), and the gap widens sharply at large scales.
Out[17]:
Console
Improvement from 1B to 100B parameters:
  α = 0.03: 12.9% loss reduction
  α = 0.05: 20.6% loss reduction
  α = 0.07: 27.6% loss reduction
  α = 0.10: 36.9% loss reduction

These numbers quantify the large impact of the scaling exponent. With a modest exponent of 0.03, scaling by 100×\times yields only a 13% loss reduction. With an exponent of 0.10, the same scale-up delivers nearly three times the benefit at 37% reduction. The exponent makes an enormous difference in the economics of scaling.

Out[18]:
Visualization
Semi-log chart showing cumulative percentage loss reduction for four different power law exponents as scale factor increases from 1x to one million times.
Cumulative loss reduction as models scale from 1x to 1 million times baseline. The gap between exponents grows nonlinearly: by the time you reach a million-fold scale-up, the difference between α=0.03 and α=0.10 is the difference between a 34% improvement and a 75% improvement.

Practical Considerations in Power Law Analysis

When applying power law analysis to real neural network scaling data, several practical issues arise that pure theory doesn't address. Knowing these issues in advance helps you avoid common pitfalls and interpret results with appropriate confidence. The gap between textbook power law analysis and real-world scaling experiments is wider than it might appear from the clean figures in published papers.

Finite-Size Effects and the Small-Model Regime

Very small models may not follow the same power law as larger ones. There is often a "small-model regime" where performance is worse than the power law would predict. This makes sense intuitively. A model with 100 parameters cannot possibly learn the complex patterns that a transformer needs, regardless of what the power law formula suggests at that scale. The model is too small to represent the input data adequately, let alone generalize. Think of it as a minimum viable threshold: below it, the model doesn't even have enough capacity to attempt the task in a real way, and its performance shows architectural limitations rather than the statistical difficulty of learning.

The power law describes behavior in a regime where the model has enough capacity to engage in real learning. The threshold between "too small to fit the pattern" and "large enough to follow the power law" is an empirical question that varies with the task and architecture. When fitting scaling laws, it's common practice to exclude the smallest models from the fit to avoid contaminating the exponent estimate with finite-size effects. If you include them, the fitted exponent will be too small, because the performance improvement from the tiny regime to the power-law regime looks like a steeper slope than the underlying power law has.

Practically, this means you should be skeptical of power law predictions that extrapolate far below your smallest training point. If your smallest experiment used 10M parameters, predicting performance at 1M parameters using the fitted power law may be unreliable.

Saturation and the Irreducible Loss Floor

Power laws cannot continue forever. Eventually, models approach irreducible limits: noise in the data, inherent task ambiguity, the entropy of the underlying distribution, or the information-theoretic ceiling imposed by the task itself.

If the true labels in your dataset have inherent randomness, such as annotator disagreement on subjective tasks or fundamentally ambiguous examples, no model can reach zero loss. There is a floor below which you cannot go, regardless of how much you scale. In the language modeling literature, this floor is related to the entropy of human language, the irreducible uncertainty in predicting the next token given all prior context. The best possible language model could only reach perplexity equal to the true entropy of language; it could not do better.

When saturation effects are present, the observed scaling curve will deviate from the power law at large scales, bending upward and flattening toward the floor. If you fit a power law to data that includes the saturation regime, you'll underestimate the magnitude of the true exponent. The curve will appear to have a shallower slope than the underlying power law in the unsaturated regime, because the flattening at large scales pulls the slope downward.

One practical way to handle this is to use a modified power law model that incorporates an explicit floor. Instead of L=aN−αL = aN^{-\alpha}, use L=Lmin⁡+aN−αL = L_{\min} + aN^{-\alpha}, where Lmin⁡L_{\min} is the irreducible loss. This three-parameter model can be fit with slightly more data but gives a better characterization of scaling behavior near the saturation regime. The Kaplan et al. (2020) paper used exactly this form, estimating Lmin⁡L_{\min} for transformer language models on a specific text corpus.

Multiple Regimes and Regime Changes

Some scaling relationships show different power law exponents in different regions. For example, parameter scaling might have one exponent below 1B parameters and another above. These regime changes can correspond to qualitative shifts in what the model is learning.

One well-documented example involves emergent capabilities. For certain tasks, model performance appears flat or near-chance across a wide range of scales, then suddenly improves sharply over a narrow range. This phase transition-like behavior produces a scaling curve that doesn't fit a single power law. The flat region and the steep improvement region have very different effective exponents.

Multiple regimes also arise when the nature of the bottleneck changes with scale. A small model might be bottlenecked by its inability to stand for the needed features at all. A medium model might be bottlenecked by having too little data for its capacity. A large model might be bottlenecked by the information content of the training distribution. Each bottleneck has a different characteristic exponent. When you see a regime change in scaling data, the most productive question to ask is: "what bottleneck changed, and why did it change at this particular scale?"

Correlations Between Variables

Model size, data size, and compute are often increased together in practice. Separating their individual contributions requires careful experimental design. If you always double your data when you double your parameters, you cannot determine from the aggregate curve whether improvements come from more parameters, more data, or both. You're measuring a confounded combination of multiple power laws.

Controlled experiments that vary one factor while holding others constant are needed for understanding individual contributions. This is easier said than done at large scale, where a single training run consumes significant resources. Researchers address this by training a grid of models at different parameter counts and different data amounts, then fitting a joint model that separates the contributions of each. We'll see how this is done in the Chinchilla scaling laws chapter. The challenge is that the number of experiments required for a full factorial design grows multiplicatively with the number of factors, making exhaustive separation difficult at the frontier of compute.

Out[19]:
Visualization
Plot showing power law with small-scale and large-scale deviations.
Real scaling data often shows multiple regimes: a small-scale deviation where the model is too small for the power law to hold, a clean power law in the middle range, and potential saturation at large scales approaching the irreducible loss floor. The realistic curve (red) diverges from the idealized power law (blue dashed) at both extremes.

These deviations don't invalidate power law analysis. They inform its proper use. Power laws describe intermediate regimes well but should be applied cautiously at the extremes of scale. When you see a scaling law paper, check what range of model sizes they used and think carefully about whether their conclusions extrapolate reliably to your target scale.

Limitations and Impact

Power laws give remarkably accurate descriptions of neural network scaling, but they have important limitations that prevent them from being a complete theory of scaling. Understanding these limitations is as important as understanding the mathematics: it tells you when to trust scaling predictions and when to be skeptical.

What Power Laws Cannot Tell You

The most basic limitation is that power laws describe what happens but not why. A power law fit tells you the exponent but not the underlying mechanism creating it. Two completely different processes can yield identical power laws. Knowing that loss decreases as N−0.07N^{-0.07} with parameters tells you the shape of the scaling curve, but nothing about the internal dynamics driving it: what features the model is learning, why adding more parameters helps, or what would happen if you changed the architecture fundamentally.

This limitation has practical consequences. If you observe a power law in your scaling data and want to improve the exponent, the power law itself gives no guidance on how to do so. You need to understand the underlying learning dynamics. Is the bottleneck the model's representational capacity, its optimization dynamics, the quality of the data, the nature of the task? The power law is agnostic to all of these. Two architectures with very different internal mechanisms can have the same scaling exponent. Conversely, a small architectural change can sometimes shift the exponent materially without any obvious reason that the change should matter for scaling. Power laws describe the symptom, not the disease.

Power laws also offer no guidance on optimal resource allocation. Knowing that loss decreases as N−0.07N^{-0.07} with parameters and as D−0.095D^{-0.095} with data tokens tells you how much improvement to expect from each resource independently, but not how to optimally trade them off against each other. If you have a fixed compute budget, should you spend it on a larger model trained for fewer steps, or a smaller model trained longer? That question, the central question of compute-optimal scaling, requires understanding how multiple power laws interact. We'll address this directly in the Chinchilla chapter.

Another important limitation is that power laws say nothing about capability thresholds. Suppose a task requires performing multi-step logical reasoning. Performance on that task might be near-random for models below a certain scale and then jump sharply as the model crosses some necessary parameter count. This emergent capability doesn't violate the power law for next-token prediction loss, but it's not predicted by it either. The aggregate loss metric that power laws describe can be flat even while specific capabilities are developing, and can remain flat even after significant capabilities emerge, if those capabilities require only a small fraction of the model's total prediction mass. A model that suddenly gains the ability to do multi-step arithmetic contributes that ability to very few of the next-token predictions it makes, so its aggregate loss barely budges.

Power laws also don't tell you which scaling variable matters most for your specific application. In general, the optimal allocation of compute between parameters and training tokens depends on both the scaling exponents for each resource and the cost of each resource. Different applications may have very different optimal allocations. A model deployed for inference at scale may benefit from fewer parameters (to reduce inference cost) even if more parameters would reduce training loss further. A model trained on a domain with abundant data may benefit from more training steps relative to parameters. The power law is a building block for these decisions, not a complete decision framework on its own.

The Extrapolation Risk

Extrapolation carries risk that's easy to underestimate. Power laws fit within the observed range can break down outside it. The successes of scaling predictions, such as researchers accurately forecasting GPT-4's performance from smaller experiments, should not obscure the basic uncertainty in extrapolating across orders of magnitude.

The history of physics is full of examples where a law that seemed universal turned out to be an approximation valid within a limited regime. Newtonian mechanics is an approximation valid at low velocities; it breaks down near the speed of light. Ideal gas laws are approximations valid at moderate temperatures and pressures; they break down at extremes. Power law scaling in deep learning may be similarly limited. The regimes accessible to today's experiments are narrow relative to the full range of possible model scales. We simply don't know whether the power laws we've measured will continue to hold at models with 10 trillion or 100 trillion parameters.

This uncertainty is compounded by the fact that we cannot easily distinguish a true power law from other slowly-varying functions over a limited range. A log-linear relationship, for example, might fit scaling data over two orders of magnitude almost as well as a power law, but would predict very different behavior at extreme scales. When someone claims to have validated a power law over "five orders of magnitude," it's worth examining exactly what range that covers and whether the data quality is consistent across that range, since early data points at small scale and recent data points at large scale may have very different levels of measurement noise.

The practical implication is to treat scaling law predictions as probabilistic estimates with uncertainty bounds, not as deterministic forecasts. A well-fit power law should come with confidence intervals on both the exponent and the coefficient, and those confidence intervals should be propagated through any extrapolation. An extrapolation 2 orders of magnitude beyond the observed data range should carry substantially wider uncertainty than an interpolation within the observed range. Papers and blog posts often present scaling law predictions without these uncertainty bounds, which can create a misleading impression of precision.

The Impact of Scaling Laws on the Field

Despite these limitations, power laws have fundamentally changed how the field approaches large-scale training. Before scaling laws were understood empirically, large training runs were expensive experiments with uncertain outcomes. A team might spend millions of dollars training a large model with little advance knowledge of whether the result would be good, mediocre, or a failure. The field advanced through costly experiments, guided partly by intuition and luck.

Scaling laws introduced a new paradigm: empirical prediction before commitment. Teams can now run a series of small experiments to estimate scaling exponents, then use those exponents to predict with reasonable confidence how a much larger model will perform. This makes possible rational compute allocation decisions, risk assessment for large training investments, and systematic architecture comparison at equivalent compute. The ability to answer "how much better will a 10×\times larger model be?" with quantitative precision, rather than a shrug and an educated guess, has restructured the economics of large-scale ML research.

The impact has been concrete and measurable. OpenAI's decision to invest heavily in large GPT models was informed by early scaling law measurements. DeepMind's Chinchilla work, which revised the estimated optimal token-to-parameter ratio and led to materially more efficient training, was directly enabled by power law analysis. Meta's LLaMA models achieved remarkable performance at modest sizes partly because of careful analysis of scaling laws under different training regimes. In each case, the power law framework allowed researchers to make quantitative predictions that guided expensive decisions, and those predictions turned out to be accurate enough to justify the framework.

Looking forward, power laws give a framework for thinking about the limits of current approaches. If loss follows L∝N−0.07L \propto N^{-0.07}, and the irreducible loss floor is at some value Lmin⁡L_{\min}, we can estimate how much scaling would be needed to approach that floor. If the required scale is astronomically large, it suggests that architectural improvements or data quality improvements, which could shift the coefficient or the exponent, may be more important than raw scaling. Power laws don't just tell us how well current approaches scale. They implicitly tell us where their limits lie, and they give researchers a common language for discussing whether a new approach is "better at any scale" (lower coefficient), "scales better" (larger exponent), or both.

The next chapters build directly on this foundation. We'll examine the specific scaling laws discovered by Kaplan et al. and later revised by Hoffmann et al. (Chinchilla), exploring how these power law relationships translate into practical training decisions and what they imply for the design of language models at every scale. The abstract mathematics of this chapter will crystallize into concrete predictions about how many parameters and how many tokens a model trained on a given compute budget should use.

Summary

Power laws describe relationships where one quantity varies as a power of another: y=axby = ax^b. In deep learning, these relationships appear ubiquitously, governing how model performance scales with parameter count and data volume under a given compute budget. They have transformed how the field approaches large-scale training by making prediction possible before commitment.

The key takeaways from this chapter:

  • Functional form: A power law y=axby = ax^b is characterized by its coefficient aa (setting overall scale) and exponent bb (determining the rate of change). For improving loss, the exponent is negative and usually written L=aN−αL = aN^{-\alpha} with α>0\alpha > 0.
  • Log-log linearity: Power laws become linear when plotted on logarithmic axes. This transformation turns nonlinear analysis into simple linear regression. This makes parameter estimation and visual validation straightforward.
  • Scale invariance: The proportional change in output for a given proportional change in input is constant, regardless of absolute scale. A 10×\times increase in parameters produces the same proportional loss improvement at any scale.
  • Parameter estimation: Linear regression in log space recovers power law coefficients and lets prediction beyond observed ranges. The slope of the log-log fit directly gives the exponent.
  • Universality: Power laws appear consistently across model architectures and tasks, as well as across resource types. Several theoretical frameworks (statistical mechanics, random matrix theory, feature learning dynamics) offer partial explanations for this universality.
  • Practical intuition: Power laws imply diminishing but persistent returns. Think multiplicatively, not additively. The exponent is the most useful number for predicting scaling efficiency, and small differences in exponents compound materially at large scale.
  • Limitations: Power laws describe what happens but not why. They don't specify optimal resource allocation, don't capture emergent capabilities, carry extrapolation risk beyond observed scales, and don't account for irreducible loss floors or regime changes.

This mathematical framework gives the basis for understanding the empirical scaling laws that have reshaped how we train large language models. In the next chapter, we'll examine the Kaplan scaling laws, which first quantified these relationships for transformer language models, and see how the abstract power law mathematics translates into concrete training recommendations.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about power laws in deep learning.

Power Laws in Deep Learning

Question 1 of 80 of 8 completed
In the power law equation y=axby=ax^b, what does the exponent bb determine?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025powerlaws, author = {Michael Brenndoerfer}, title = {Power Laws in Deep Learning: Understanding Neural Scaling}, year = {2025}, url = {https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Power Laws in Deep Learning: Understanding Neural Scaling. Retrieved from https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling
MLAAcademic
Michael Brenndoerfer. "Power Laws in Deep Learning: Understanding Neural Scaling." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling>.
CHICAGOAcademic
Michael Brenndoerfer. "Power Laws in Deep Learning: Understanding Neural Scaling." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Power Laws in Deep Learning: Understanding Neural Scaling'. Available at: https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Power Laws in Deep Learning: Understanding Neural Scaling. https://mbrenndoerfer.com/writing/power-laws-deep-learning-neural-network-scaling

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.