NTK-aware Scaling: Extending Context Length in LLMs

Michael BrenndoerferUpdated July 1, 202557 min read

Part of Language AI Handbook

NTK-aware scaling extends transformer context windows by adjusting rotary frequencies. Covers position preservation, longer sequences, and tradeoffs.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

NTK-aware Scaling

Position Interpolation, as we saw in the previous chapter, extends context length by simply scaling down all rotation frequencies. While this works, it comes with a hidden cost: by compressing all frequencies equally, we lose the fine-grained positional distinctions that high-frequency components provide. Consider a model trained with 4,096 tokens. Position Interpolation at 8x scale compresses everything to fit 32,768 positions, but now adjacent tokens that were once easily distinguishable become nearly identical in their high-frequency dimensions.

NTK-aware scaling takes a more surgical approach. Instead of treating all frequencies equally, it recognizes that different frequency bands serve different purposes. High frequencies distinguish nearby tokens, while low frequencies capture long-range structure. By scaling frequencies non-uniformly, NTK-aware methods preserve local precision while still extending the effective context window.

Think of it like a camera lens. If you want to photograph a larger scene, you can zoom out uniformly, but then the small details you cared about become blurry. A smarter approach would be to change only the focal length for objects far away while keeping close-up clarity intact. NTK-aware scaling applies exactly this logic to the rotary frequency spectrum, compressing the "distant" low-frequency dimensions while leaving the "close-up" high-frequency ones at full resolution.

The motivation for this approach came from an observation about how neural networks use their positional information. During training, models learn to use different frequency bands for different purposes. High-frequency dimensions encode fine-grained local relationships, things like "this word is immediately followed by that word" or "this is a possessive construction." Low-frequency dimensions encode structural and global relationships, things like "this sentence is in the first half of the document" or "this is part of the concluding paragraph." Position Interpolation disrupts both kinds of encoding equally, even though they have very different sensitivity requirements.

The mathematical insight underlying NTK-aware scaling is elegant: rather than modifying the output frequencies directly, we can achieve dimension-dependent scaling by modifying the base constant of the RoPE formula. Because frequencies decrease exponentially with dimension index, a change to the base has a larger effect on low-frequency dimensions than on high-frequency ones. This gives us the non-uniform scaling we want almost for free, as a mathematical consequence of the exponential structure.

This chapter develops the intuition behind NTK-aware scaling, derives the mathematical formula from first principles, and implements both static and dynamic variants. You will see exactly why this approach often outperforms linear interpolation, and you will leave with a concrete understanding of when to use static versus dynamic scaling in production systems.

The Problem with Uniform Scaling

To understand why uniform scaling falls short, we need to revisit how RoPE frequencies work and appreciate what each part of the frequency spectrum contributes to the model's ability to encode position. Uniform scaling causes measurable degradation. Understanding that degradation helps us design better solutions.

Recall that RoPE assigns each dimension pair ii a base frequency that determines how fast the rotation angle changes with position:

θi=1b2i/d\theta_i = \frac{1}{b^{2i/d}}

where:

  • θi\theta_i: the base rotation frequency (in radians per position) for dimension pair ii
  • bb: the base constant, typically 10000, which controls the overall frequency range
  • ii: the dimension pair index, ranging from 0 to d/2−1d/2 - 1
  • dd: the total embedding dimension (must be even since we work in pairs)
  • 2i/d2i/d: the exponent that creates a geometric progression of frequencies across dimensions

The first dimension pair (i=0i = 0) has frequency θ0=1/100000=1\theta_0 = 1/10000^0 = 1, meaning it rotates by 1 radian per position and completes a full rotation every 2π≈6.282\pi \approx 6.28 positions. The last pair (i=d/2−1i = d/2 - 1) has a much smaller frequency, completing a rotation only after thousands of positions. This spread creates a multi-scale representation:

  • High-frequency pairs (ii near 0): Distinguish nearby tokens with precision. Tokens at positions 5 and 6 look very different.
  • Low-frequency pairs (ii near d/2d/2): Capture coarse position information. Tokens far apart have noticeably different rotations.

Notice that this multi-scale design is deliberate and important. RoPE's ability to encode position without explicit positional embeddings depends on the entire spectrum working together. Each frequency band plays a role in shaping the inner products between queries and keys in the attention mechanism. When you attend from position mm to position nn, the resulting attention score is influenced by the rotation difference Δ=n−m\Delta = n - m across every dimension pair simultaneously. High-frequency pairs create a rapidly oscillating component that is sensitive to small values of Δ\Delta, while low-frequency pairs create a slowly varying component that only changes significantly when Δ\Delta is large.

Position Interpolation addresses context extension by scaling all frequencies uniformly:

θi′=θis\theta_i' = \frac{\theta_i}{s}

where:

  • θi′\theta_i': the scaled frequency for dimension pair ii after Position Interpolation
  • θi\theta_i: the original frequency for dimension pair ii
  • ss: the context extension factor, computed as s=Ltarget/Ltrains = L_{\text{target}} / L_{\text{train}} (target length divided by training length)

If we want to extend from 4,096 to 32,768 tokens, s=32768/4096=8s = 32768 / 4096 = 8. Every frequency gets divided by 8. The high-frequency pair that previously rotated 1 radian per position now rotates only 1/8=0.1251/8 = 0.125 radians. Adjacent tokens, which were once clearly distinguishable, become nearly indistinguishable in these dimensions.

The key insight is that this is a lossy compression. The model was trained to use the high-frequency dimensions to encode fine-grained neighborhood relationships. Tokens at positions 100 and 101 should look clearly different in those dimensions, because the attention mechanism needs to distinguish "the word immediately to my right" from "the word two positions to my right." After Position Interpolation with s=8s = 8, those two positions differ by only 0.1250.125 radians instead of 1 radian, a factor of 8 reduction in separability. The model has no way to recover this information; it simply cannot see a distinction that was previously obvious.

In practice, this translates to measurable performance degradation. Tasks that depend heavily on local attention, such as identifying named entity boundaries, tracking pronoun references within a sentence, or parsing syntactic dependencies, are especially vulnerable. The model may still produce reasonable outputs because other mechanisms compensate, but the positional signal that was supposed to guide local attention has become weaker.

Let's visualize this problem:

Out[4]:
Visualization
Line plot showing rotation angles increasing with position for original and scaled RoPE.
Rotation angles at each position for the high-frequency dimension (i=0). Original RoPE increases steeply while Position Interpolation compresses to a shallow slope.
Bar chart comparing angular separation between adjacent tokens.
Angular separation between adjacent tokens. Position Interpolation with 8x scaling reduces separation from 1 radian to 0.125 radians.

The compression is dramatic. With Position Interpolation at 8x scale, adjacent tokens are separated by only 0.125 radians in the highest-frequency dimension, compared to 1 radian originally. This means the model has far less "room" to distinguish nearby positions, potentially harming tasks that require precise local attention.

One might ask: why does this matter so much if the model can still distinguish all positions from one another? The answer is that attention is computed via inner products, and inner product differences encode similarity. When adjacent tokens are close together in the rotation space, their inner products with any given query become nearly identical, which means the softmax attention distribution becomes nearly uniform over nearby tokens. The model loses the sharp, precise attention to individual nearby tokens that it learned during training. This loss shows up in perplexity benchmarks and in downstream tasks that require local syntactic sensitivity.

From Intuition to Formula: The NTK-aware Approach

Now that we've seen the problem with uniform scaling, let's develop a solution from first principles. The journey from intuition to a working formula requires answering three questions: What property do we want? How can we achieve it mathematically? And does the result match our expectations?

Neural Tangent Kernel (NTK)

The Neural Tangent Kernel describes how neural networks behave during training in the infinite-width limit. In this context, "NTK-aware" refers to preserving the high-frequency components that the network needs to learn fine-grained distinctions between nearby positions. The term was coined in a 2023 Reddit post by user "bloc97" who made the connection between neural network frequency sensitivity and the mathematical properties of the NTK, motivating a scaling strategy that respects the model's learned frequency preferences.

Historical Context

NTK-aware scaling emerged not from a formal research paper but from a Reddit post in May 2023, written by a researcher working independently to extend LLaMA context windows. The poster observed that Position Interpolation's uniform frequency compression conflicted with what NTK theory predicts about how neural networks use frequency information. The insight spread quickly through the open-source community and was rapidly validated experimentally. Within weeks it became a standard technique incorporated into LLaMA derivative models including Code Llama and Mistral. This is a striking example of the open community driving fundamental technique improvements in parallel with formal academic channels.

The Core Insight: Frequency Bands Serve Different Purposes

The key realization is that not all frequencies are equal in importance. Think of RoPE's multi-frequency design like a ruler with both centimeter and millimeter markings. The millimeter marks (high frequencies) give you precision for small measurements, while the centimeter marks (low frequencies) help you measure longer distances quickly. Position Interpolation is like shrinking the entire ruler by 8x, making even the millimeter marks too small to read. What we really want is to keep the millimeter precision while adjusting the centimeter scale.

Another useful analogy: think of the frequency dimensions as instruments in an orchestra. The high-frequency dimensions are like the piccolo and violins, playing rapid, precise notes that define the immediate rhythmic detail of the music. The low-frequency dimensions are like the bass and cello. This provides the slow harmonic foundation that gives the piece its overall structure. Position Interpolation is like telling every instrument to play at 1/8 tempo, which ruins the fast parts even though the slow parts could tolerate it. A better approach would be to slow only the bass section while keeping the violins at full tempo.

Translating this intuition into concrete requirements:

  1. Preserve high frequencies: The fastest-rotating dimension pairs (small ii) should remain unchanged. These are our "millimeter marks" for local position distinctions.

  2. Scale low frequencies: The slowest-rotating dimension pairs (large ii) can be compressed by the full factor ss. These are our "centimeter marks" that need adjustment for longer contexts.

  3. Smooth transition: Intermediate dimensions should scale gradually between these extremes, not abruptly jump.

These three requirements are not arbitrary design choices. They follow from thinking about what the model learned during training. For a model trained on sequences up to LtrainL_{\text{train}} tokens, every positional relationship it has seen involves positions within that range. High-frequency dimensions complete many full rotations within that range and therefore carry meaningful, learned patterns. Low-frequency dimensions may only complete a fraction of a rotation, so extending the context just means continuing the pattern they were already set up to handle.

Designing the Solution: Modify the Base

How can we achieve dimension-dependent scaling? The original RoPE frequency formula is θi=1/b2i/d\theta_i = 1/b^{2i/d}. Position Interpolation modifies the output by dividing: θi′=θi/s\theta_i' = \theta_i / s. This gives uniform scaling because every frequency is divided by the same constant.

Instead, we'll modify the input: the base bb. This is the key creative step in the derivation. By operating on the base rather than the output, we can exploit the mathematical structure of the exponential to get dimension-dependent effects automatically.

If we replace bb with a larger base b′b', all frequencies decrease (slower rotation). The key insight is that the exponent 2i/d2i/d causes this decrease to affect different dimensions differently:

  • When i=0i = 0: θ0′=1/(b′)0=1\theta_0' = 1/(b')^0 = 1. No matter what b′b' is, the highest frequency remains 1.
  • When ii is large: θi′=1/(b′)2i/d\theta_i' = 1/(b')^{2i/d} shrinks more because the exponent is larger.

This is exactly the property we want. By choosing the right b′b', we can leave high frequencies untouched while scaling low frequencies by whatever factor we need. The mathematical structure of the exponential gives us the smooth, dimension-dependent transition between these extremes as a natural consequence, without any ad hoc weighting or interpolation scheme.

Notice the elegance here: we do not need to specify how each individual dimension should be scaled. We just need to find one number, b′b', and the exponential structure of the formula takes care of distributing the scaling correctly across all dimensions. This is a beautiful example of how choosing the right mathematical representation of a problem can make the solution fall out almost automatically.

Out[5]:
Visualization
2D scatter plot showing token positions rotated by high-frequency dimension under different scaling methods.
High-frequency dimension (i=0): Original RoPE rotates by 1 radian per position, creating widely spaced points. Position Interpolation compresses to 0.125 rad/pos, making adjacent positions nearly overlap. NTK-aware preserves the original spacing.
2D scatter plot showing token positions rotated by low-frequency dimension under different scaling methods.
Low-frequency dimension (i=31): All methods produce similar results here. NTK-aware and Position Interpolation both apply full compression to low-frequency dimensions.

This geometric view makes the difference concrete. In the high-frequency dimension (left), original RoPE spreads positions around the circle, allowing the model to distinguish them easily. Position Interpolation bunches them together near the starting point, reducing distinguishability. NTK-aware scaling preserves the original spread. In the low-frequency dimension (right), all methods behave similarly since both Position Interpolation and NTK-aware apply full compression to these slow-rotating dimensions.

Deriving the Formula

Let's work backward from our requirements to find the exact formula for b′b'. We want a scaling function where:

  • scale0=1\text{scale}_0 = 1 (no scaling at highest frequency)
  • scaled/2−1=s\text{scale}_{d/2-1} = s (full scaling at lowest frequency)

We write the modified base as b′=b⋅αb' = b \cdot \alpha for some multiplier α\alpha to be determined. This parameterization is useful because it lets us express the new base as a proportional increase over the old one, and we can reason about the multiplier α\alpha independently. Substituting into the frequency formula:

θi′=1(b⋅α)2i/d\theta_i' = \frac{1}{(b \cdot \alpha)^{2i/d}}

Using the exponent rule (ab)n=an⋅bn(ab)^n = a^n \cdot b^n, we can separate this:

θi′=1b2i/d⋅1α2i/d=θi⋅α−2i/d\theta_i' = \frac{1}{b^{2i/d}} \cdot \frac{1}{\alpha^{2i/d}} = \theta_i \cdot \alpha^{-2i/d}

The effective scaling factor for dimension ii becomes:

scalei=θiθi′=α2i/d\text{scale}_i = \frac{\theta_i}{\theta_i'} = \alpha^{2i/d}

Now we apply our constraints. At i=0i = 0, we need scale0=1\text{scale}_0 = 1:

scale0=α2⋅0/d=α0=1(as required)\text{scale}_0 = \alpha^{2 \cdot 0/d} = \alpha^0 = 1 \quad \text{(as required)}

This is automatically satisfied for any α>0\alpha > 0, since any number raised to zero equals one. Our first requirement is met regardless of what α\alpha we choose.

At i=d/2−1i = d/2 - 1 (the last dimension pair), we need scaled/2−1=s\text{scale}_{d/2-1} = s:

scaled/2−1=α2(d/2−1)/d=s\text{scale}_{d/2-1} = \alpha^{2(d/2 - 1)/d} = s

Simplifying the exponent step by step:

α(d−2)/d=s\alpha^{(d - 2)/d} = s

To solve for α\alpha, we raise both sides to the power d/(d−2)d/(d-2):

α=sd/(d−2)\alpha = s^{d/(d-2)}

This gives us the complete NTK-aware formula:

b′=b⋅sdd−2b' = b \cdot s^{\frac{d}{d - 2}}

where:

  • b′b': the modified base for NTK-aware scaling
  • bb: the original base (typically 10000)
  • ss: the context extension factor (e.g., s=8s = 8 to extend from 4k to 32k tokens)
  • dd: the embedding dimension
  • d/(d−2)d/(d-2): an exponent slightly greater than 1 (for d=64d = 64, this equals 64/62≈1.03264/62 \approx 1.032)

The exponent d/(d−2)d/(d-2) is worth examining closely. For large dd, this converges toward 1, meaning the formula approaches b′=b⋅sb' = b \cdot s. For small dd, it is significantly larger than 1. In practice, the head dimension dd used in modern transformers is typically 64 or 128, so the exponent is either ≈1.032\approx 1.032 or ≈1.016\approx 1.016. These are small but meaningful corrections that ensure the boundary condition at i=d/2−1i = d/2 - 1 is met exactly. If you were to mistakenly use b′=b⋅sb' = b \cdot s instead of the correct formula, the lowest-frequency dimension would be scaled by slightly less than ss, producing a small but systematic mismatch at the low-frequency end of the spectrum.

Understanding the Resulting Frequencies

Substituting the new base into the frequency formula reveals how each dimension is affected. Starting from the frequency definition, replacing bb with b⋅sd/(d−2)b \cdot s^{d/(d-2)}, and simplifying:

θi′=1(b′)2i/d=1(b⋅sd/(d−2))2i/d=θi⋅s−2i/(d−2)\theta_i' = \frac{1}{(b')^{2i/d}} = \frac{1}{(b \cdot s^{d/(d-2)})^{2i/d}} = \theta_i \cdot s^{-2i/(d-2)}

The effective scaling factor for each dimension, defined as the ratio of the original frequency to the new frequency, is:

scalei=θiθi′=s2id−2\text{scale}_i = \frac{\theta_i}{\theta_i'} = s^{\frac{2i}{d - 2}}

Let's verify this matches our design goals:

  • When i=0i = 0 (highest frequency): scale0=s0=1\text{scale}_0 = s^0 = 1. The highest frequencies are not scaled at all. ✓
  • When i=d/2−1i = d/2 - 1 (lowest frequency): scaled/2−1=s(d−2)/(d−2)=s\text{scale}_{d/2-1} = s^{(d-2)/(d-2)} = s. The lowest frequencies are scaled by the full factor. ✓
  • Intermediate dimensions: Scaling increases smoothly from 1 to ss as ii increases. ✓

This is precisely what we wanted: high frequencies preserved, low frequencies scaled, smooth transition between.

Notice also that the scaling exponent 2i/(d−2)2i/(d-2) grows roughly linearly with dimension index ii. This means the scaling factor is exponential in the dimension index, which mirrors the original exponential spacing of the frequencies. In other words, NTK-aware scaling respects the logarithmic structure of RoPE. It does not arbitrarily compress some frequencies more than others; it applies compression in a way that is proportional to the frequency's original "slowness."

A Worked Example

Let's make this concrete with typical values: d=64d = 64 dimensions, s=8s = 8 extension factor (4k to 32k tokens). This is a realistic scenario corresponding to extending LLaMA-2, which was trained with a 4,096-token context, to a 32,768-token context window.

First, compute the NTK exponent:

dd−2=6462≈1.032\frac{d}{d-2} = \frac{64}{62} \approx 1.032

The exponent is only slightly above 1, which means the modified base is only slightly larger than what you would get by simply multiplying bb by ss. This small correction is what makes the last dimension pair satisfy the boundary condition exactly.

The modified base becomes:

b′=10000⋅81.032=10000⋅8.72≈87,200b' = 10000 \cdot 8^{1.032} = 10000 \cdot 8.72 \approx 87,200

In practice, this means the model will behave as if it was trained with a base of 87,200 rather than 10,000. You can think of this as a kind of "stretching" of the frequency spectrum: by raising the base, we make all frequencies slightly smaller (slower rotation), but the effect is much stronger at the low-frequency end because those dimensions have a large exponent in the denominator.

Now we can compute scaling factors for specific dimensions:

NTK-aware scaling factors for d=64d=64, s=8s=8. High-frequency dimensions (small ii) are preserved while low-frequency dimensions receive full scaling.
Dimension iiExponent 2i/(d−2)2i/(d-2)scalei=8exp\text{scale}_i = 8^{\text{exp}}Interpretation
001.00No scaling (preserved)
80.2581.72Slight scaling
160.5162.95Moderate scaling
240.7745.06Significant scaling
311.008.00Full scaling

The gradient is smooth and continuous. High-frequency pairs (small ii) stay close to their original values, while low-frequency pairs (large ii) are compressed to fit the extended context. Notice that the "moderate scaling" region, around dimension pairs 12 to 20, corresponds to the intermediate frequencies that provide mid-range positional encoding, useful for tracking structure at the paragraph level rather than the word or document level. NTK-aware scaling applies intermediate compression here, which is intuitively the right thing to do: these dimensions can tolerate some compression, but not the full factor.

Out[6]:
Visualization
Line plot showing uniform scaling for Position Interpolation and progressive scaling for NTK-aware method across dimension indices.
Scaling factors across dimension pairs for Position Interpolation vs NTK-aware scaling. Position Interpolation applies uniform scaling (flat line), while NTK-aware scaling preserves high frequencies (left) and only compresses low frequencies (right).

The visualization confirms our mathematical analysis. The NTK-aware curve (squares) starts at 1 for the highest-frequency pairs and smoothly increases to the full scale factor of 8 for the lowest-frequency pairs. Meanwhile, Position Interpolation (circles) maintains a flat line at 8, treating all dimensions identically.

Implementation

With the formula derived and verified, let's translate it into code. We'll build the implementation in layers, starting with the core frequency computation and progressively adding the full RoPE transformation. The implementation is remarkably simple, which is one of the key practical advantages of NTK-aware scaling over more complex methods. The entire context extension reduces to a single change in one parameter.

In practice, you would integrate this into a model's RoPE layer by patching the frequency computation before inference. In frameworks like Hugging Face Transformers, this means replacing the inv_freq buffer inside the rotary embedding module. Libraries like llama.cpp and vLLM implement NTK-aware scaling as a configurable option that can be toggled with a single flag.

Computing Frequencies

The foundation is a function that computes RoPE frequencies for any base value. This same function serves both the original frequencies and NTK-aware frequencies, since the only difference is the base we pass in. By keeping the function generic, we avoid code duplication and make it easy to swap between methods:

In[7]:
Code
def compute_rope_frequencies(d, base=10000):
    """Compute original RoPE frequencies for d/2 dimension pairs."""
    dim_pairs = np.arange(d // 2)
    frequencies = 1.0 / (base ** (2 * dim_pairs / d))
    return frequencies

Position Interpolation simply divides all frequencies by the scale factor:

In[8]:
Code
def position_interpolation_frequencies(d, base=10000, scale=1.0):
    """Compute Position Interpolation frequencies (uniform scaling)."""
    original = compute_rope_frequencies(d, base)
    return original / scale

Position Interpolation is the simplest possible extension: just divide every frequency by the scale factor. This is a useful baseline that is easy to understand and easy to implement, but as we have seen, it does not handle all frequency bands well.

NTK-aware scaling modifies the base according to our derived formula, then computes frequencies using the modified base:

In[9]:
Code
def ntk_aware_frequencies(d, base=10000, scale=1.0):
    """Compute NTK-aware scaled frequencies (non-uniform scaling)."""
    # Modify the base according to NTK formula: b' = b * s^(d/(d-2))
    ntk_base = base * (scale ** (d / (d - 2)))
    return compute_rope_frequencies(d, ntk_base)

The elegance of this implementation is worth appreciating: the entire NTK-aware scaling logic is a one-liner. Computing the modified base and passing it to the same frequency function as before handles everything. This simplicity is a major practical advantage: you can retrofit NTK-aware scaling onto any existing RoPE implementation with a single line change.

Comparing the Approaches

Let's compute frequencies for all three methods and compare them side by side. Looking at the raw numbers helps build intuition before we see the visualizations:

In[10]:
Code
# Compare the two approaches
d = 64
base = 10000
scale = 8  # Extend context by 8x

freq_original = compute_rope_frequencies(d, base)
freq_pi = position_interpolation_frequencies(d, base, scale)
freq_ntk = ntk_aware_frequencies(d, base, scale)
Out[11]:
Console
Context extension factor: 8x
Embedding dimension: 64

Frequency comparison (first 5 dimension pairs):
 Dim     Original           PI          NTK   PI ratio  NTK ratio
--------------------------------------------------------------
   0     1.000000     0.125000     1.000000       8.00       1.00
   1     0.749894     0.093737     0.701242       8.00       1.07
   2     0.562341     0.070293     0.491741       8.00       1.14
   3     0.421697     0.052712     0.344829       8.00       1.22
   4     0.316228     0.039528     0.241809       8.00       1.31

Frequency comparison (last 5 dimension pairs):
 Dim     Original           PI          NTK   PI ratio  NTK ratio
--------------------------------------------------------------
  27     0.000422     0.000053     0.000069       8.00       6.12
  28     0.000316     0.000040     0.000048       8.00       6.54
  29     0.000237     0.000030     0.000034       8.00       7.00
  30     0.000178     0.000022     0.000024       8.00       7.48
  31     0.000133     0.000017     0.000017       8.00       8.00

The numbers confirm what we derived mathematically. Look at the "PI ratio" column: it's a constant 8.00 across all dimensions. This reflects Position Interpolation's uniform scaling. Now look at "NTK ratio": it starts near 1.00 for dimension 0 and gradually increases toward 8.00 for the highest dimensions. This is the dimension-dependent scaling in action.

The NTK ratio column tells you something important about what the model will experience at inference time. For dimension 0, the frequencies are essentially unchanged, so the model's learned patterns for detecting adjacent tokens are fully intact. For the highest dimensions, the frequencies are reduced by the full factor of 8, which is the necessary price of extending the context: you need the slow-rotating dimensions to cover 8x more positions, so they must rotate 8x more slowly. The difference from Position Interpolation is that this cost is paid entirely by the dimensions that are designed to handle long-range structure, not by the dimensions that handle local patterns.

Visualizing the Frequency Spectrum

A log-scale plot reveals the full picture across all 32 dimension pairs. The logarithmic scale is important here because the frequencies themselves span many orders of magnitude:

Out[12]:
Visualization
Line plot comparing angular separation between adjacent tokens for original, Position Interpolation, and NTK-aware scaling across dimension pairs.
Angular separation between adjacent tokens across all dimension pairs. NTK-aware scaling preserves the original separation for high-frequency dimensions (left side), while Position Interpolation compresses all dimensions uniformly.

On the log scale, observe how NTK-aware frequencies (triangles) track the original (circles) closely for low dimension indices, where high frequencies live. As we move right toward higher dimension indices (lower frequencies), the NTK curve gradually diverges to match Position Interpolation (squares). This is exactly the "preserve high frequencies, scale low frequencies" behavior we designed.

The crossover point, where NTK-aware frequencies visibly diverge from the original, corresponds roughly to the middle of the dimension range. In a model with d=64d = 64 head dimension, this is around dimension pair 15 to 20. Dimensions below this threshold receive essentially no scaling and behave as if the context window were not extended. Dimensions above this threshold receive progressively more scaling and are responsible for encoding the longer-range positional structure needed for the extended context.

Rotation Angle Heatmaps

To visualize how different scaling methods affect the rotation patterns, let's create heatmaps showing the rotation angles across positions and dimension pairs. In these heatmaps, each row corresponds to a dimension pair and each column to a position. The color represents the rotation angle modulo 2π2\pi, so rapid color cycling indicates a high-frequency dimension while slow color change indicates a low-frequency one:

Out[13]:
Visualization
Heatmap of original rotation angles across 20 positions and 32 dimension pairs.
Original RoPE rotation angles. High-frequency dimensions (bottom) show rapid variation, while low-frequency dimensions (top) change slowly.
Heatmap of Position Interpolation rotation angles showing uniform compression.
Position Interpolation uniformly compresses all rotation angles, reducing variation across all dimensions equally.
Heatmap of NTK-aware rotation angles showing preserved high frequencies.
NTK-aware scaling preserves high-frequency patterns (bottom) while compressing only low-frequency dimensions (top).

The heatmaps reveal the key difference between methods. In the original (left), the bottom rows (high-frequency dimensions) show rapid color cycling as position increases, while top rows (low-frequency dimensions) change slowly. Position Interpolation (center) uniformly slows down all cycling, making all rows look more like the slow top rows. NTK-aware scaling (right) preserves the rapid cycling in the bottom rows while slowing only the top rows. This visual confirms that NTK-aware scaling selectively modifies frequencies based on dimension.

Looking at the NTK heatmap versus the original, a key observation is that the bottom half of the rows looks nearly identical in both. This is the visual representation of "high frequencies preserved." The top half of the rows is where you see the difference, and even there the NTK-aware version looks more similar to the original than the Position Interpolation version does. The original's top rows complete a fraction of a full rotation over the 20 positions shown, and the NTK-aware version matches this better than Position Interpolation, which further compresses them.

Applying RoPE with NTK-aware Scaling

Now that we can compute the frequencies, let's implement the complete RoPE transformation and measure its effect on token distinguishability. This section is where theory meets practice: we apply the frequencies to actual vectors and observe whether NTK-aware scaling delivers on its promise of better local discriminability.

The RoPE Transformation

RoPE works by rotating each pair of embedding dimensions by a position-dependent angle. The rotation angle for position mm and dimension pair ii is simply m⋅θim \cdot \theta_i, where θi\theta_i is the frequency we computed above. This is a standard 2D rotation matrix applied to consecutive pairs of dimensions. Importantly, the rotation is applied to both the query and key vectors at each position, so that the inner product between a query at position mm and a key at position nn depends only on the relative offset n−mn - m through the rotation difference:

In[14]:
Code
def apply_rope(x, frequencies):
    """Apply RoPE to a sequence of vectors.

    Args:
        x: Input vectors of shape (seq_len, d)
        frequencies: Base frequencies for each dimension pair, shape (d/2,)

    Returns:
        Rotated vectors of same shape as x
    """
    seq_len, d = x.shape
    positions = np.arange(seq_len)

    # Compute rotation angles: (seq_len, d/2)
    # Each position m gets angle m * theta_i for each dimension pair i
    angles = np.outer(positions, frequencies)

    # Split input into pairs for 2D rotation
    x_pairs = x.reshape(seq_len, -1, 2)  # (seq_len, d/2, 2)

    # Precompute sin and cos for efficiency
    cos_angles = np.cos(angles)  # (seq_len, d/2)
    sin_angles = np.sin(angles)

    # Apply 2D rotation: [x, y] -> [x*cos - y*sin, x*sin + y*cos]
    x_rotated = np.zeros_like(x_pairs)
    x_rotated[:, :, 0] = (
        x_pairs[:, :, 0] * cos_angles - x_pairs[:, :, 1] * sin_angles
    )
    x_rotated[:, :, 1] = (
        x_pairs[:, :, 0] * sin_angles + x_pairs[:, :, 1] * cos_angles
    )

    return x_rotated.reshape(seq_len, d)

Measuring Token Distinguishability

The ultimate test of our scaling method is whether adjacent tokens remain distinguishable after rotation. If tokens at positions mm and m+1m+1 become too similar, the model loses its ability to attend precisely to nearby positions.

Think of it this way: the attention mechanism computes scores by taking inner products between queries and keys. If all nearby keys produce nearly the same score for a given query, the softmax distribution becomes flat over those positions. The model cannot selectively attend to "the word right before me" versus "the word two steps back," because both produce the same attention score. This is the practical consequence of high-frequency compression, and it is what we want to measure.

We'll measure this using cosine similarity between adjacent token embeddings. Lower similarity means more distinguishable tokens. We use random embeddings here as a proxy for actual model outputs, which is valid because the distinguishability depends primarily on the rotations applied, not on the specific content of the vectors:

In[15]:
Code
# Create one deterministic sample after the shared setup seed.
seq_len = 10
d = 64
embeddings = np.random.randn(seq_len, d) * 0.1

# Apply RoPE with different scaling methods
rotated_original = apply_rope(embeddings, freq_original)
rotated_pi = apply_rope(embeddings, freq_pi)
rotated_ntk = apply_rope(embeddings, freq_ntk)
In[16]:
Code
# Compute cosine similarity between adjacent tokens
def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))


# Calculate similarities for each method
similarities = {"original": [], "pi": [], "ntk": []}

for i in range(seq_len - 1):
    similarities["original"].append(
        cosine_similarity(rotated_original[i], rotated_original[i + 1])
    )
    similarities["pi"].append(
        cosine_similarity(rotated_pi[i], rotated_pi[i + 1])
    )
    similarities["ntk"].append(
        cosine_similarity(rotated_ntk[i], rotated_ntk[i + 1])
    )
Out[17]:
Console
Cosine similarity between adjacent token embeddings:
   Positions   Original         PI        NTK
--------------------------------------------
       0 & 1    -0.0311     0.0152    -0.0237
       1 & 2    -0.2305    -0.2127    -0.2279
       2 & 3    -0.0370    -0.0254    -0.0345
       3 & 4     0.2245     0.2120     0.2135
       4 & 5    -0.0194     0.0068    -0.0210
       5 & 6    -0.0112     0.0196    -0.0044
       6 & 7    -0.1081    -0.0881    -0.1076
       7 & 8     0.1659     0.1631     0.1591
       8 & 9     0.0982     0.1182     0.1049
--------------------------------------------
     Average     0.0057     0.0232     0.0065
Out[18]:
Visualization
Bar chart comparing average cosine similarity between adjacent tokens for Original, Position Interpolation, and NTK-aware scaling methods.
Average cosine similarity between adjacent tokens for each scaling method. Lower similarity indicates more distinguishable positions. NTK-aware scaling preserves token distinguishability closer to the original, while Position Interpolation compresses the representation.

The visualization makes the pattern clear: Position Interpolation significantly increases similarity between adjacent tokens (making them harder to distinguish), while NTK-aware scaling keeps the similarity close to the original. This preservation of token distinguishability is the practical benefit of frequency-dependent scaling.

In practice, the gap between NTK-aware and Position Interpolation is most pronounced in tasks where local syntactic structure matters. Question answering, where the model must identify a precise span of text, benefits from the preserved local discriminability. Summarization, where the model reads long text and produces a concise summary, is somewhat less affected because the model mostly needs to understand what is there rather than to precisely track where it is. This is an important calibration point when choosing a context extension method for your specific use case.

Dynamic NTK Scaling

Static NTK-aware scaling uses a fixed extension factor ss. But what if the actual sequence length varies? A model configured for 32k tokens shouldn't apply aggressive scaling when processing only 2k tokens. Ideally, the scaling should adapt to the current context: no modification for short sequences, progressive scaling as sequences grow longer.

A model serving a chatbot receives requests of widely varying lengths. Short questions, paragraph-length responses, and multi-page code reviews all pass through the same model. If you apply aggressive 8x static scaling to a short 200-token prompt, you are compressing frequencies that do not need to be compressed at all. The model's fine-tuned patterns for short-context reasoning may degrade unnecessarily.

Dynamic NTK scaling adapts the base in real-time based on the current sequence length. The formula is identical to static scaling, but the scale factor is computed from the actual sequence length at inference time rather than being fixed in advance:

b′(L)=b⋅(LLtrain)dd−2b'(L) = b \cdot \left(\frac{L}{L_{\text{train}}}\right)^{\frac{d}{d-2}}

where:

  • b′(L)b'(L): the dynamically computed base, now a function of sequence length LL
  • bb: the original base (typically 10000)
  • LL: the current sequence length being processed
  • LtrainL_{\text{train}}: the context length the model was trained with
  • L/LtrainL/L_{\text{train}}: the effective extension factor, computed on-the-fly
  • d/(d−2)d/(d-2): the same exponent as in static NTK-aware scaling

The formula is identical to static NTK scaling, except that s=L/Ltrains = L/L_{\text{train}} is computed dynamically rather than fixed in advance.

When L≤LtrainL \leq L_{\text{train}}, the ratio L/Ltrain≤1L/L_{\text{train}} \leq 1, and we clamp it to 1 (no scaling). When LL exceeds the training length, scaling kicks in proportionally. A sequence of 8k tokens (2x the training length of 4k) gets s=2s = 2; a sequence of 32k tokens (8x) gets s=8s = 8.

The clamping behavior is important. Without it, short sequences would have a scale factor less than 1, which would make the base smaller than the original and effectively increase frequencies. This would make adjacent tokens rotate faster than they did during training, which is a different kind of distortion. The clamp ensures dynamic NTK is always either "no change" or "compress," never "expand."

In[19]:
Code
def dynamic_ntk_frequencies(d, base, train_length, current_length):
    """Compute dynamically scaled NTK-aware frequencies.

    Args:
        d: Embedding dimension
        base: Original RoPE base (typically 10000)
        train_length: Training context length
        current_length: Current sequence length

    Returns:
        Frequencies for each dimension pair
    """
    if current_length <= train_length:
        # No scaling needed
        return compute_rope_frequencies(d, base)

    # Dynamic scaling factor
    scale = current_length / train_length
    ntk_base = base * (scale ** (d / (d - 2)))
    return compute_rope_frequencies(d, ntk_base)
In[20]:
Code
# Demonstrate dynamic scaling at different sequence lengths
d = 64
base = 10000
train_length = 4096
test_lengths = [2048, 4096, 8192, 16384, 32768]

freq_by_length = {
    L: dynamic_ntk_frequencies(d, base, train_length, L) for L in test_lengths
}
Out[21]:
Console
Dynamic NTK scaling: effective base at different sequence lengths
Training length: 4096

  Seq Length    Scale   Effective Base
--------------------------------------
        2048     1.00            10000
        4096     1.00            10000
        8192     2.00            20452
       16384     4.00            41829
       32768     8.00            85550

The effective base scales smoothly with sequence length. At 32k tokens (8x the training length), the base increases to approximately 87k, matching our earlier static calculation. At 4,096 tokens or below (within the training distribution), the base remains at 10,000, meaning the model behaves exactly as it did during training. This is a clean and desirable property: dynamic NTK is a strict generalization of the original model, falling back to the original behavior whenever the sequence is short enough.

Notice that the effective base does not jump discontinuously at the training-length boundary. There is a smooth transition that begins exactly at LtrainL_{\text{train}}, which avoids any sudden behavioral change as sequences cross this threshold. This is important in applications like streaming generation, where the context window grows one token at a time; the model's behavior changes gradually rather than abruptly as the context crosses the training boundary.

Out[22]:
Visualization
Line plot showing how RoPE frequencies change with sequence length under dynamic NTK scaling, with different curves for different sequence lengths.
Dynamic NTK-aware frequencies adapt based on sequence length. Short sequences (at or below training length) use original frequencies, while longer sequences progressively modify frequencies to extend context.

Attention Score Analysis

To understand the practical impact, let's examine how attention scores behave under different scaling methods. Frequency comparisons tell us about the mechanics, but attention scores tell us about the downstream effect on the model's ability to focus on relevant tokens. We'll create synthetic queries and keys at various distances and measure how the attention weight between them changes with distance under each method.

The expected behavior is a gradual decay: tokens closer to the query should receive higher attention than tokens far away, all else being equal. This is because RoPE encodes relative position directly into the inner product, and vectors at small relative offsets have more similar rotations than vectors at large relative offsets. If context extension distorts this decay pattern, the model may inappropriately up-weight distant tokens or fail to distinguish near from far.

In[23]:
Code
def compute_attention_scores(q, k):
    """Compute scaled dot-product attention scores."""
    d = q.shape[-1]
    scores = q @ k.T / np.sqrt(d)
    return scores


def create_position_aware_vectors(positions, d, frequencies):
    """Create vectors with RoPE applied at specified positions."""
    n = len(positions)
    # The shared setup seed keeps this example deterministic across variants.
    base_vectors = np.random.randn(n, d) * 0.1

    # Create rotation matrices for each position
    rotated = np.zeros_like(base_vectors)
    for idx, pos in enumerate(positions):
        angles = pos * frequencies
        cos_a = np.cos(angles)
        sin_a = np.sin(angles)

        for i in range(d // 2):
            x, y = base_vectors[idx, 2 * i], base_vectors[idx, 2 * i + 1]
            rotated[idx, 2 * i] = x * cos_a[i] - y * sin_a[i]
            rotated[idx, 2 * i + 1] = x * sin_a[i] + y * cos_a[i]

    return rotated
In[24]:
Code
# Examine attention between a query and keys at varying distances
d = 64
base = 10000
train_length = 4096
extended_length = 32768
scale = extended_length / train_length

# Different scaling approaches
freq_orig = compute_rope_frequencies(d, base)
freq_pi = position_interpolation_frequencies(d, base, scale)
freq_ntk = ntk_aware_frequencies(d, base, scale)

# Query at position 1000, keys at various distances
query_pos = 1000
key_distances = [1, 2, 5, 10, 50, 100, 500, 1000]
key_positions = [query_pos + dist for dist in key_distances]
Out[25]:
Visualization
Line plot showing attention score decay with distance for original, Position Interpolation, and NTK-aware scaling methods.
Relative attention scores as a function of distance from the query token. NTK-aware scaling maintains a decay pattern closer to the original, preserving the model's learned attention behavior for nearby tokens.

The attention decay pattern reveals a subtle but important distinction. While Position Interpolation flattens the attention curve (making nearby and distant tokens more similar in score), NTK-aware scaling preserves more of the original decay structure, especially at short distances.

The most significant differences appear at very short distances, within a few positions of the query. This is where the high-frequency dimensions dominate the inner product, because they are the ones that change fastest with relative position. Since NTK-aware scaling preserves these high-frequency dimensions, the attention behavior at short distances stays closer to the original. At long distances, both methods produce similar behavior because the long-range structure is dominated by low-frequency dimensions, which both methods compress to a similar degree.

Comparing NTK to Position Interpolation

Let's directly compare how well each method preserves the relative position encoding properties. The key metric here is how closely the scaled methods track the original RoPE behavior as a function of relative position offset. A method that perfectly preserved all properties would have zero deviation everywhere:

Out[26]:
Visualization
Line plot of dot product values vs relative position for original, Position Interpolation, and NTK-aware methods.
Dot product between rotated vectors as a function of relative position. The original RoPE curve shows strong sensitivity to position differences.
Line plot showing deviation from original dot product values for both scaling methods.
Deviation from original behavior. NTK-aware scaling (triangles) stays closer to zero, especially at small relative positions, indicating better preservation of learned patterns.

The deviation plot (right) makes the advantage clear: NTK-aware scaling maintains smaller deviations from the original behavior, particularly for small relative positions where local attention patterns matter most.

The deviation for NTK-aware scaling starts near zero at relative position 0 and grows slowly, while Position Interpolation's deviation is larger throughout. This quantifies the intuition we built earlier: NTK-aware scaling is more faithful to the original model's behavior for nearby tokens, which is exactly what we need for precise local attention.

One way to think about this: Position Interpolation is like translating a document with Google Translate and then running spell-check. The translation changes every word, and the spell-checker doesn't know which original word each new word was supposed to correspond to. NTK-aware scaling is like translating only the technical jargon while leaving everyday words unchanged, a more targeted intervention that preserves more of the original meaning.

Interpolation Factor Analysis

An alternative perspective on NTK-aware scaling is to analyze the interpolation factor for each frequency. Rather than thinking in terms of modified bases, we can ask: for each dimension pair, how far along the spectrum from "no change" to "fully interpolated" does NTK-aware scaling move us? This framing reveals that NTK-aware scaling is implicitly doing something very natural: it applies a smooth, dimension-dependent blend of the two extremes.

We can express NTK-aware scaling as a blend between "no interpolation" and "full interpolation":

In[27]:
Code
def compute_interpolation_factors(d, scale):
    """Compute effective interpolation factor for each dimension pair.

    Returns values between 0 (no interpolation, high freq preserved)
    and 1 (full interpolation).
    """
    dim_pairs = np.arange(d // 2)

    # NTK scaling formula: scale^(2i / (d-2))
    ntk_factors = scale ** (2 * dim_pairs / (d - 2))

    # Normalize to [0, 1] range: 1 means no scaling, scale means full scaling
    # interpolation = 0 when factor = 1, interpolation = 1 when factor = scale
    interpolation = (ntk_factors - 1) / (scale - 1)

    return interpolation
Out[28]:
Visualization
Line plot showing interpolation factor from 0 to 1 across dimension indices, with smooth S-curve transition.
Interpolation factor across dimension pairs under NTK-aware scaling. Low-frequency dimensions (right) receive full interpolation, while high-frequency dimensions (left) are preserved with minimal modification.

This interpolation perspective provides an intuitive understanding: NTK-aware scaling smoothly transitions from "preserve high frequencies" (interpolation factor near 0) to "fully scale low frequencies" (interpolation factor near 1).

The S-shaped curve in the interpolation factor plot reflects the exponential growth of the scaling factor. Initially, for the highest-frequency dimensions, the interpolation factor grows very slowly because the scaling factor is close to 1 and changing little. As you move to lower-frequency dimensions, the factor accelerates and approaches 1 more quickly. The exact shape depends on the scale factor ss: larger values of ss make the S-curve steeper, because the transition from near-zero to near-one must be completed in the same number of dimension pairs.

This perspective also helps you reason about what happens when you need an extension factor larger than what was originally intended. If a model was designed for s=4s = 4 but you want to push to s=16s = 16, the interpolation factors at intermediate dimensions will be significantly larger, meaning more distortion in the middle frequencies. This is one reason why very large extension factors often require fine-tuning even when smaller extension factors can work zero-shot.

Limitations and Practical Considerations

NTK-aware scaling works better than Position Interpolation, but it comes with its own trade-offs and limitations. Understanding these is essential for making good engineering decisions about when to use it, how to configure it, and when to look for something better.

The core tension in context extension is that you cannot perfectly preserve all the properties of the original position encoding while also extending the context window. NTK-aware scaling prioritizes high-frequency preservation at the cost of more aggressive low-frequency compression. For tasks that rely heavily on precise long-range position information, this trade-off may not be ideal. Document retrieval, where finding a specific passage requires accurate global positioning, might suffer compared to summarization tasks where local coherence matters more. The right calibration depends on your workload, and it is worth profiling both methods on your specific task before committing to one.

Also, NTK-aware scaling, like Position Interpolation, typically requires fine-tuning to achieve good performance. While models may work "out of the box" with NTK scaling applied at inference time, the attention patterns learned during training assumed different frequency relationships. Fine-tuning on longer sequences helps the model adapt to the new frequency set. The amount of fine-tuning needed is generally comparable to Position Interpolation: a few hundred to a few thousand steps on representative long-context data often suffices. Without fine-tuning, performance at very long contexts (near the maximum extended length) tends to degrade faster than performance at shorter lengths within the original training distribution.

One important subtlety is that NTK-aware scaling makes a specific assumption about the model's frequency usage that may not hold perfectly for all architectures or training regimes. The derivation assumes that the high-frequency dimensions are the most sensitive to scaling and that the boundary conditions (no scaling at i=0i = 0, full scaling at i=d/2−1i = d/2 - 1) are exactly right. In practice, different layers of a transformer may use frequencies differently. Shallow layers, which process raw token sequences, might rely more on high-frequency dimensions, while deep layers, which process highly abstract representations, might be less sensitive to frequency distortion. NTK-aware scaling applies the same modification to every layer, which is a simplification.

The choice between static and dynamic NTK scaling depends on deployment constraints. Static scaling is simpler to implement and has deterministic behavior, but requires knowing the maximum context length in advance. Dynamic scaling adapts gracefully to varying sequence lengths but adds computational overhead, as frequencies must be recomputed based on current sequence length. There is also a subtlety with KV-cache reuse: when you use static scaling, the cached key/value pairs from previous steps remain valid as the context grows. With dynamic scaling, each new sequence length changes the frequencies, which means cached key/value pairs from earlier positions were computed with different frequencies than the ones now being used for the current position. This KV-cache incompatibility can be a significant practical concern for efficient inference, and is one reason why static scaling is often preferred in production systems despite dynamic scaling's theoretical advantages.

NTK-aware scaling is also specific to RoPE and does not transfer to other position encoding schemes. Models using absolute position embeddings, ALiBi, or other mechanisms require different extension strategies. This limits the generality of the approach, though the dominance of RoPE in modern architectures (LLaMA and Mistral, alongside Falcon and Qwen, all use RoPE) means NTK-aware scaling is broadly applicable in practice.

Finally, there is the question of the maximum achievable extension factor. In principle you can set ss as large as you want, but performance degrades when ss is very large. The reason is that even low-frequency dimensions eventually become too compressed: a scale factor of s=128s = 128 would reduce the lowest-frequency dimension to 1/128 of its original speed, making it essentially useless for encoding position over the long range. Empirically, extension factors of s≤16s \leq 16 tend to work well with modest fine-tuning, while larger factors require either more fine-tuning or a more sophisticated method such as YaRN, which we cover in the next chapter. The practical ceiling for NTK-aware scaling without significant fine-tuning is roughly a 4x to 8x extension.

In Practice: Deploying NTK-aware Scaling

Understanding how NTK-aware scaling behaves in theory is valuable, but deploying it effectively requires knowing some practical details about how to integrate it with real models and frameworks.

The simplest deployment path is to use a library that already supports NTK-aware scaling as a built-in option. In Hugging Face Transformers, you can enable it by setting the rope_scaling configuration parameter in the model config. For a LLaMA-based model, this looks like passing {"type": "dynamic", "factor": 8.0} or {"type": "linear", "factor": 8.0} to the rope_scaling argument when loading the model. The library handles the frequency recomputation automatically based on the type specified.

If you need to implement it manually, the key integration point is the inv_freq buffer inside the rotary embedding module. This buffer stores the base frequencies θi\theta_i for each dimension pair. NTK-aware scaling requires replacing the computation inv_freq = 1.0 / (base ** (2 * dim_pairs / d)) with inv_freq = 1.0 / (ntk_base ** (2 * dim_pairs / d)) where ntk_base = base * (scale ** (d / (d - 2))). For dynamic scaling, this computation needs to be re-executed whenever the sequence length changes.

One commonly encountered issue when first deploying NTK-aware scaling is forgetting to account for the model's head dimension versus the total model dimension. The variable dd in the NTK formula refers to the per-head embedding dimension, not the total hidden size. For a model with hidden size 4096 and 32 attention heads, d=128d = 128, and the NTK exponent is 128/126≈1.016128/126 \approx 1.016. Using the wrong dd will produce a different modified base and different scaling factors across dimension pairs.

Another practical consideration is evaluation. When testing a context-extended model, it is important to evaluate on sequences that are longer than the training context. Evaluating only on short sequences will not reveal whether the extension is working correctly, because short sequences do not exercise the modified low-frequency dimensions. A good evaluation suite includes sequences at 1x1x, 2x2x, 4x4x, and 8x8x the training context length, measuring both perplexity and task performance at each length. If perplexity starts rising steeply before the target length, or if task performance drops suddenly at a certain threshold, more fine-tuning or a larger extension factor adjustment is needed.

Key Parameters

When implementing NTK-aware scaling, several parameters control the behavior of the context extension. Getting these right is important for correct behavior and good performance.

  • base (default: 10000): The original RoPE base constant. This value comes from the pretrained model and should match what was used during training. Common values are 10000 (LLaMA, Mistral) or 500000 (some newer models with extended context). Using the wrong base will produce frequencies that are scaled from the wrong starting point, degrading performance in unpredictable ways. Always check the model card or configuration file for the original base.

  • scale (ss): The context extension factor, computed as target length divided by training length. For example, extending from 4k to 32k tokens gives s=8s = 8. Larger values enable longer contexts but increase the distortion from original frequencies. In practice, ss between 2 and 8 works well with minimal fine-tuning, while ss above 16 typically requires significant adaptation training.

  • d (embedding dimension): The model's embedding dimension, used in computing the NTK exponent d/(d−2)d/(d-2). This is fixed by the model architecture and typically ranges from 64 to 128 for the head dimension in modern transformers. Note that this is the head dimension, not the total model dimension; if the model has a total embedding dimension of 4096 with 32 attention heads, each head has dimension 4096/32=1284096 / 32 = 128, and you should use d=128d = 128 in the formula.

  • train_length (for dynamic scaling): The context length the model was originally trained with. This is the threshold below which no scaling is applied. Using the correct value is important for dynamic NTK to work properly. If you set this too low, the model will apply scaling even for short sequences that do not need it. If you set it too high, the model will fail to scale up for sequences that exceed the training context.

When choosing between static and dynamic scaling, consider your deployment scenario. Static scaling with a fixed ss is simpler and more predictable, suitable when you know the maximum sequence length in advance. Dynamic scaling adapts to varying inputs but adds computational overhead for recomputing frequencies. More importantly, dynamic scaling has implications for KV-cache reuse, as described in the Limitations section. For most production deployments, static scaling is recommended unless your application has highly variable context lengths that would make aggressive static scaling wasteful for short inputs.

Summary

NTK-aware scaling provides a principled approach to extending context length in RoPE-based models by recognizing that different frequency bands serve different purposes. The central insight is that uniform scaling, as used in Position Interpolation, treats all frequencies as equal when they are not. High-frequency dimensions carry precise local positional information that the model's training has made sensitive to even small changes, while low-frequency dimensions carry coarse global structure that can absorb more compression. NTK-aware scaling exploits this asymmetry.

  • Key insight: High-frequency components distinguish nearby tokens; low-frequency components capture long-range structure. Uniform scaling damages local precision unnecessarily.

  • The NTK formula: Replace base bb with b′=b⋅sd/(d−2)b' = b \cdot s^{d/(d-2)}, where ss is the context extension factor. This single parameter change propagates through the exponential structure of the frequency formula to produce exactly the dimension-dependent scaling we need.

  • Frequency-dependent scaling: The effective scaling factor increases from 1 (no change) for the highest frequencies to ss (full scaling) for the lowest frequencies. Intermediate dimensions receive intermediate scaling, with the gradient following an exponential curve that mirrors the original frequency spacing.

  • Dynamic variant: Adapt the base in real-time based on current sequence length: b′(L)=b⋅(L/Ltrain)d/(d−2)b'(L) = b \cdot (L/L_{\text{train}})^{d/(d-2)}. This ensures the model applies no unnecessary scaling for short sequences and provides the right amount of extension for long ones.

  • Practical benefits: Better preservation of local attention patterns compared to Position Interpolation, especially for tasks requiring precise nearby-token relationships. Tasks like question answering, code understanding, and span extraction benefit most.

  • Limitations: NTK-aware scaling still requires fine-tuning for optimal performance at large extension factors. It is specific to RoPE, introduces KV-cache incompatibility issues in the dynamic variant, and achieves diminishing returns for very large scale factors beyond approximately 8x to 16x.

NTK-aware scaling improved context extension, but researchers continued seeking even better solutions. In the next chapter, we'll explore YaRN (Yet another RoPE extension method), which builds on NTK-aware principles while adding attention scaling and temperature adjustments for further improvements. YaRN also introduces a more sophisticated frequency treatment that uses a ramp function rather than a smooth exponential, giving practitioners more direct control over which frequency bands are scaled and which are preserved entirely.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about NTK-aware scaling and context extension in transformer models.

NTK-aware Scaling

Question 1 of 80 of 8 completed
What is the main problem with Position Interpolation when extending context length?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025ntkaware, author = {Michael Brenndoerfer}, title = {NTK-aware Scaling: Extending Context Length in LLMs}, year = {2025}, url = {https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). NTK-aware Scaling: Extending Context Length in LLMs. Retrieved from https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension
MLAAcademic
Michael Brenndoerfer. "NTK-aware Scaling: Extending Context Length in LLMs." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension>.
CHICAGOAcademic
Michael Brenndoerfer. "NTK-aware Scaling: Extending Context Length in LLMs." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension.
HARVARDAcademic
Michael Brenndoerfer (2025) 'NTK-aware Scaling: Extending Context Length in LLMs'. Available at: https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). NTK-aware Scaling: Extending Context Length in LLMs. https://mbrenndoerfer.com/writing/ntk-aware-scaling-context-extension

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.