Part of History of Language AI
Transformer-XL extends context with segment recurrence and relative positions. Covers its architecture, efficiency, training method, and influence.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2019: Transformer-XL
Transformer language models were commonly trained on fixed-length pieces of text. A piece started without the hidden states computed for the preceding piece, even when both came from the same document. This created context fragmentation: a dependency crossing the boundary disappeared from the model's input.
Transformer-XL, published in 2019 by researchers at Carnegie Mellon University and Google Brain, connected those pieces. Each layer cached activations from earlier segments and exposed them as memory to the next segment. A new relative-position formulation made those reused activations positionally consistent with the current query.
The architecture did not give attention unlimited or linear-cost memory. It processed one segment at a time against a bounded cache. This reused computation and extended the effective context while retaining standard attention within that working set.
The Problem
Full self-attention over tokens forms an score matrix, so its time and memory costs grow as . Training on an entire long document was therefore expensive. Splitting the document into length- segments controlled the cost but restricted each prediction to its current segment.
This restriction also wasted computation at evaluation time. To predict one segment with a longer history, a vanilla model had to process that history again. Moving the window forward repeated many of the same layer calculations.
Absolute position encodings caused a second problem once hidden states were reused. A cached state already contained the absolute index it had in the old segment. When attached before a new segment, the same state needed to be interpreted at a different distance from each new query. Reusing it without changing the attention formulation produced a positional mismatch.
Transformer-XL addressed these two issues together. Recurrence reused the old activations; relative attention described where a cached key sat with respect to the current query.
The Solution
Let denote the hidden states at layer for segment . When processing the next segment, layer receives a memory made from earlier layer activations and the current segment:
where stops gradients through the cached segment and denotes concatenation. Queries come from the current segment. Keys and values come from the concatenated memory, so a current token can attend to earlier cached states.
Stopping the gradient is an important boundary. Information from the cache affects the forward pass, but backpropagation does not continue through the entire document. Training remains segment-based rather than becoming full-document backpropagation through time.
The memory length is a configuration choice. With current length and memory length , attention for one layer is proportional to . Reusing cached states avoids recomputing them, but attending to a larger memory still costs time and storage. Production implementations normally keep a fixed-size cache and discard older states.
Relative Attention
Transformer-XL decomposes an attention score into content and position terms. The position term depends on the distance between a current query at and a key at . It uses a sinusoidal relative-position vector together with learned projection and bias parameters. It is not a learned lookup embedding for every possible distance.
This formulation solves the cache mismatch because a stored state has no single new absolute index. Its positional contribution is recomputed from its distance to each current query. The model can also evaluate a longer attention length than the segment length used during training, although performance at unseen distances remains an empirical question rather than a guarantee.
Reported Results
The paper reported new results on five language-modeling datasets, including perplexity 18.3 on WikiText-103 and 21.8 on One Billion Word. Its relative effective context length was measured as 450% longer than the vanilla Transformer baseline in the paper's setup.
The reported evaluation speedup of up to 1,800 times compared Transformer-XL's cached generation with a vanilla Transformer that recomputed a long context for each prediction. It was not a claim that Transformer-XL attention was 1,800 times cheaper for every workload.
Applications and Impact
Transformer-XL was introduced and tested as an autoregressive language model. Its most direct application was generation or scoring of a continuous stream while carrying a bounded history between segments. Claims about document classification, coreference, or code generation require separate task-specific evidence and were not results of the original paper.
The released code and checkpoints made the recurrence mechanism reusable. XLNet, published later in 2019 by overlapping authors, adopted Transformer-XL's recurrence and relative attention while changing the pretraining objective.
Transformer-XL also made cached activations a visible design choice for long-context models. Later memory models explored compression or retrieval, while sparse-attention models took a different route by changing which token pairs could interact. These are alternative engineering strategies, not interchangeable forms of recurrence.
Limitations
The cache contains fixed hidden states. A later token can read them, but new evidence does not revise their representations. This asymmetry suits causal language modeling; it is less suitable when a task needs bidirectional reinterpretation of the entire document.
A bounded cache still forgets. Increasing lengthens the available history and raises the attention cost. Segment recurrence reuses old computation, but it does not remove the memory-compute tradeoff.
Gradients stop at the segment boundary. The model can learn to use cached states because they are present in the forward pass, but it cannot assign credit through an arbitrarily long chain of prior segments during one update.
Relative position does not make length extrapolation automatic. Queries at evaluation time may attend over distances that were rare during training. The formulation permits those distances to be represented; model quality still depends on the learned behavior.
Finally, recurrence imposes sequential processing across segments. Tokens within a segment remain parallel, but segment needs the cache produced for segment . That dependency limits document-level training parallelism.
Legacy
Transformer-XL supplied a concrete answer to context fragmentation: retain earlier layer activations and reuse them without backpropagating through the full history. XLNet is the clearest direct continuation of that design.
Its relative-attention formulation also helped establish relative position as an alternative to adding absolute vectors at the input. Later systems used several distinct methods, including relative biases and rotary embeddings. Those methods share a design concern with Transformer-XL but should not all be described as implementations of its exact scheme. GPT-3, for example, used learned absolute position embeddings.
The chapter's main lesson is the tradeoff, not a claim of unlimited context. A fixed cache preserves recent hidden states and avoids recomputation. Extending that cache increases attention cost, and frozen memories cannot be rewritten by later evidence.
Quiz
The following questions review segment recurrence, stopped gradients, relative attention, and cache limits.
Transformer-XL Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!