Part of Language AI Handbook
Examines GPT-2's architecture, model sizes, WebText training, and zero-shot capabilities that changed language modeling through scale.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
GPT-2: Scaling Language Models for Zero-Shot Learning
In February 2019, OpenAI announced GPT-2 with an unusual strategy: they would not release the full model. The 1.5 billion parameter language model was, they claimed, too dangerous. It could generate text so convincingly human-like that they feared misuse, including fake news and spam as well as impersonation. The decision sparked fierce debate about responsible AI release, but it also pointed to a more basic shift. GPT-2 demonstrated that scale, applied to a simple autoregressive language modeling objective, could produce emergent capabilities that no one had explicitly trained for.
Think of GPT-2 as a machine that reads 40 gigabytes of the internet, then attempts to write more of it. It learns not by being told what tasks exist or how to solve them, but by absorbing the patterns of human communication across millions of web pages. When a person posts a question on the internet, someone answers it. When an article headline appears, a body of text follows. When someone writes a passage in French, more French text comes after. GPT-2 learns all of these patterns simultaneously, and the astonishing result is that it can reproduce any of them on demand, even without being explicitly told which pattern is needed.
This chapter examines GPT-2's contributions to language modeling. We'll explore how OpenAI scaled up from GPT-1's 117 million parameters across four model sizes, understand the architectural refinements that enabled stable training at scale, analyze the WebText dataset that taught the model from high-quality internet text, and investigate the zero-shot learning phenomenon that made GPT-2 far more than just a text generator.
The key insight that runs through every section of this chapter is that GPT-2 did not change the basic objective from GPT-1. It still trained by predicting the next token. What changed was everything else: how much data, how large the model, and how carefully the training was engineered to remain stable at scale. These changes in quantity produced changes in quality that surprised even the researchers who built the system.
By the end of this chapter, you'll understand why GPT-2 marked the transition from fine-tuned specialists to general-purpose language models. You'll see the precise architectural choices that made large-scale stable training possible, appreciate the subtle genius of training from naturally occurring task demonstrations embedded in web text, and have a working implementation of the GPT-2 architecture in PyTorch.
GPT-2 was published in February 2019 by Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever at OpenAI. The paper "Language Models are Unsupervised Multitask Learners" challenged the prevailing assumption that language models required task-specific fine-tuning to be useful. OpenAI's staged release strategy, withholding the full 1.5B parameter model for several months, was among the first times an AI lab argued that a language model posed enough misuse risk to warrant controlled release. The full model was eventually released in November 2019 after observing that the feared harms had not materialized from the smaller released versions.
Model Sizes: A Family of Models
GPT-2 wasn't a single model but a family of four. OpenAI trained models ranging from 117 million to 1.5 billion parameters, letting systematic study of how capabilities scale with size. This approach would become standard practice for subsequent large language models.
The decision to train four models of different sizes rather than one final model was itself scientifically significant. By holding the architecture constant and varying only size, the researchers could isolate the effect of scale from the effect of architectural choices. Every observed improvement in the larger models could be attributed directly to having more parameters and more capacity to memorize and generalize. This family-of-models approach became a template: GPT-3 would later present six sizes, and Meta's LLaMA family would publish multiple scales for the same reason.
Think of the four GPT-2 sizes as a staircase, with each step roughly doubling or tripling the parameter count while using the exact same architectural blueprint. GPT-2 Small is essentially GPT-1 replicated. GPT-2 Medium represents a meaningful first expansion. GPT-2 Large approaches what was then considered a very large model. GPT-2 XL, at 1.5 billion parameters, was the final destination that OpenAI initially withheld from public release, citing concerns about misuse potential.
GPT-2 Small matches GPT-1's size at 117M parameters with 12 layers. GPT-2 Medium doubles this to 345M parameters with 24 layers. GPT-2 Large reaches 762M with 36 layers. GPT-2 XL, the full model, contains 1.5B parameters across 48 layers.
Architecture Specifications
The architectural specifications for each variant are:
| Parameter | GPT-2 Small | GPT-2 Medium | GPT-2 Large | GPT-2 XL |
|---|---|---|---|---|
| Layers () | 12 | 24 | 36 | 48 |
| Hidden size () | 768 | 1024 | 1280 | 1600 |
| Attention heads () | 12 | 16 | 20 | 25 |
| Head dimension () | 64 | 64 | 64 | 64 |
| Feed-forward size | 3072 | 4096 | 5120 | 6400 |
| Vocabulary size | 50,257 | 50,257 | 50,257 | 50,257 |
| Context length | 1024 | 1024 | 1024 | 1024 |
| Parameters | ~117M | ~345M | ~762M | ~1.5B |
The key architectural parameters are:
- : number of transformer layers (blocks stacked sequentially)
- : hidden dimension, the size of token representations flowing through the network
- : number of attention heads operating in parallel within each layer
- : dimension per attention head, determining how much information each head can capture
One of the most important patterns in this table is the one that does not change: the head dimension stays fixed at 64 across all four variants. This is a deliberate engineering choice. As the model grows, the researchers did not make individual attention heads larger; they added more heads and more layers instead. Think of it like a company that grows not by making each employee a superhero but by hiring more people with the same specialized role. Each attention head in GPT-2 XL does the same kind of work as each head in GPT-2 Small, just as one of 25 specialists rather than one of 12.
The 4x feed-forward expansion ratio (hidden size multiplied by 4 equals feed-forward size) also remains constant across all variants. Context length doubled from GPT-1's 512 to 1024 tokens, letting the model to condition on longer passages of text before making each prediction. This doubled context is significant: it means the model can attend to roughly two pages of text when generating each word, compared to GPT-1's single page.
Parameter Distribution
Let's compute the parameter distribution to understand where model capacity resides:
def count_gpt2_parameters(
vocab_size: int = 50257,
hidden_size: int = 768,
num_layers: int = 12,
max_position: int = 1024,
) -> dict:
"""Count parameters in each component of GPT-2."""
intermediate_size = hidden_size * 4
params = {}
# Embedding layers (token and position)
params["token_embeddings"] = vocab_size * hidden_size
params["position_embeddings"] = max_position * hidden_size
# Per-layer parameters
# Self-attention: Q, K, V projections + output projection (all combined in c_attn, c_proj)
attention_params = (
3 * (hidden_size * hidden_size) + 3 * hidden_size
) # c_attn
attention_params += hidden_size * hidden_size + hidden_size # c_proj
# Feed-forward: two linear layers (c_fc, c_proj)
ff_params = hidden_size * intermediate_size + intermediate_size # c_fc
ff_params += intermediate_size * hidden_size + hidden_size # c_proj
# Layer norms (2 per layer)
layernorm_params = 4 * hidden_size
params["per_layer"] = attention_params + ff_params + layernorm_params
params["all_layers"] = num_layers * params["per_layer"]
# Final layer norm
params["final_layernorm"] = 2 * hidden_size
# Output projection (tied with token embeddings, so not counted separately)
params["output_projection"] = 0 # Weight tying
# Total
params["embeddings_total"] = (
params["token_embeddings"] + params["position_embeddings"]
)
params["total"] = (
params["embeddings_total"]
+ params["all_layers"]
+ params["final_layernorm"]
)
return paramsGPT-2 Small: Embeddings: 39,383,808 (31.6%) Transformer Layers: 85,054,464 (68.3%) Total: 124,439,808 GPT-2 Medium: Embeddings: 52,511,744 (14.8%) Transformer Layers: 302,309,376 (85.2%) Total: 354,823,168 GPT-2 Large: Embeddings: 65,639,680 (8.5%) Transformer Layers: 708,387,840 (91.5%) Total: 774,030,080 GPT-2 XL: Embeddings: 82,049,600 (5.3%) Transformer Layers: 1,475,558,400 (94.7%) Total: 1,557,611,200
The parameter breakdown reveals where model capacity resides. For GPT-2 Small, embeddings constitute about 31% of parameters, but this fraction shrinks dramatically as models scale. GPT-2 XL dedicates over 90% of its parameters to transformer layers, meaning the large majority of learned knowledge resides in attention and feed-forward weights rather than the vocabulary embeddings.

The embedding layer represents a decreasing fraction of total parameters as models grow. In GPT-2 Small, embeddings account for roughly 31% of parameters, but in GPT-2 XL this drops to about 6%. This pattern reflects a basic principle: vocabulary embeddings scale with vocabulary size (50,257 tokens times hidden dimension), while transformer layers scale quadratically with hidden dimension and linearly with depth. For large models, knowledge increasingly resides in the transformer layers themselves, inside the weights of the attention projections and feed-forward networks.
The key insight here is that knowledge has a different home in small versus large models. In a small model, a disproportionate share of the network is just a lookup table mapping tokens to vectors. In a large model, the transformer machinery dominates, and that machinery is where the model stores relationships between concepts, syntactic patterns, and factual associations.
Architectural Changes from GPT-1
GPT-2 maintained the decoder-only transformer architecture from GPT-1 but introduced several modifications that proved important for stable training at larger scales. These changes became standard practice for subsequent language models. Understanding each change requires understanding the problem it was designed to solve: as networks grow deeper, gradients become harder to propagate reliably during backpropagation, and small numerical instabilities compound across layers.
Think of training a deep network as a game of telephone with 48 players (the layers). Each player passes a message to the next, and each also passes feedback in the opposite direction to help earlier players improve. If any player distorts the feedback signal, every player upstream receives corrupted information. GPT-2's architectural changes are engineering solutions to this telephone problem. This keeps useful gradient information survives the journey from the output layer all the way back to the first embedding layer.
Pre-Normalization
The most significant architectural change was moving layer normalization before each sub-block rather than after. GPT-1 and the original transformer used post-normalization, applying LayerNorm after the residual connection. GPT-2 switched to pre-normalization, applying LayerNorm before the attention and feed-forward operations.
To understand why this matters, consider what happens to activations inside a deep residual network. At each layer, the residual path adds information to the main stream. In post-normalization, unnormalized values from the residual path are added together and then normalized. The problem is that the unnormalized values can have high and unpredictable variance, which makes the normalization layer's job harder and makes gradients flowing backward through the normalization potentially unstable.
With pre-normalization, the LayerNorm is applied to the input before it enters the attention or feed-forward block. The residual connection then adds the block's output to the original, un-normalized input. This means the inputs to each block are always well-behaved, and the gradient signal traveling backward through the residual connection is clean because it bypasses the normalization entirely. For 12-layer networks, both approaches often work. For 48-layer networks, pre-normalization is nearly needed for avoiding training divergence.
class GPT2Block(nn.Module):
"""GPT-2 transformer block with pre-normalization."""
def __init__(
self, hidden_size: int = 768, num_heads: int = 12, dropout: float = 0.1
):
super().__init__()
self.hidden_size = hidden_size
self.num_heads = num_heads
# Pre-normalization: LayerNorm BEFORE attention and FFN
self.ln_1 = nn.LayerNorm(hidden_size, eps=1e-5)
self.ln_2 = nn.LayerNorm(hidden_size, eps=1e-5)
# Attention (simplified for clarity)
self.c_attn = nn.Linear(
hidden_size, 3 * hidden_size
) # Q, K, V combined
self.c_proj = nn.Linear(hidden_size, hidden_size)
# Feed-forward
self.c_fc = nn.Linear(hidden_size, 4 * hidden_size)
self.c_proj_ffn = nn.Linear(4 * hidden_size, hidden_size)
self.dropout = nn.Dropout(dropout)
def forward(
self, x: torch.Tensor, mask: torch.Tensor = None
) -> torch.Tensor:
# Pre-norm attention
residual = x
x = self.ln_1(x) # Normalize BEFORE attention
x = self._attention(x, mask)
x = residual + self.dropout(x) # Residual connection
# Pre-norm feed-forward
residual = x
x = self.ln_2(x) # Normalize BEFORE FFN
x = self.c_fc(x)
x = F.gelu(x) # GELU activation
x = self.c_proj_ffn(x)
x = residual + self.dropout(x) # Residual connection
return x
def _attention(
self, x: torch.Tensor, mask: torch.Tensor = None
) -> torch.Tensor:
B, T, C = x.size()
head_dim = C // self.num_heads
# Compute Q, K, V
qkv = self.c_attn(x)
q, k, v = qkv.split(C, dim=-1)
# Reshape for multi-head attention
q = q.view(B, T, self.num_heads, head_dim).transpose(1, 2)
k = k.view(B, T, self.num_heads, head_dim).transpose(1, 2)
v = v.view(B, T, self.num_heads, head_dim).transpose(1, 2)
# Scaled dot-product attention with causal mask
# Scale by 1/sqrt(d_k) to prevent dot products from growing too large
# which would push softmax into regions with tiny gradients
scale = 1.0 / (head_dim**0.5)
attn = torch.matmul(q, k.transpose(-2, -1)) * scale
# Apply causal mask
causal_mask = torch.triu(
torch.ones(T, T, device=x.device), diagonal=1
).bool()
attn = attn.masked_fill(causal_mask, float("-inf"))
attn = F.softmax(attn, dim=-1)
attn = self.dropout(attn)
# Apply attention to values
out = torch.matmul(attn, v)
out = out.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(out)The causal mask is important for autoregressive generation. It ensures each position can only attend to previous positions (and itself), preventing information leakage from future tokens. Let's visualize what this mask looks like:

The lower triangular pattern means token can only attend to tokens . This is what makes GPT-2 autoregressive: when predicting the next token, the model cannot "peek" at future tokens. During training, this enables efficient parallel computation since the entire sequence can be processed in one forward pass with masking.
Why does this matter? Pre-normalization provides more stable gradients during training. In post-normalization, the residual branch adds unnormalized activations, which can have high variance. The subsequent LayerNorm must handle this variance, but gradients flowing back through the normalization can become unstable at scale. With pre-normalization, the residual connection adds already-normalized values, keeping the residual stream's magnitude more controlled.


GELU Activation
GPT-2 replaced the ReLU activation function with GELU (Gaussian Error Linear Unit). While ReLU simply zeros out negative values, GELU applies a smooth, non-monotonic transformation that weights inputs by their probability under a Gaussian distribution.
The GELU function multiplies each input by its probability of being greater than other inputs from a Gaussian distribution. This creates a smooth gating mechanism where positive values pass through almost unchanged, while negative values are attenuated based on how negative they are.
The exact formulation is:
where:
- : the input value to the activation function
- : the cumulative distribution function (CDF) of the standard normal distribution, representing the probability that a random variable is less than or equal to
- : the error function, a special function that arises in probability and statistics
The intuition is elegant: we multiply each input by the probability that exceeds values drawn from a Gaussian. Large positive values have , so they pass through unchanged. Large negative values have , so they're zeroed out. Values near zero get partially attenuated.
In practice, computing the error function is expensive. GPT-2 and most implementations use this polynomial approximation:
where is the hyperbolic tangent function and the constants and were chosen to minimize approximation error. This formulation is differentiable everywhere and computationally efficient.

Why does this formula make sense? Notice that when is large and positive, approaches 1, so : the activation behaves like the identity. When is large and negative, approaches 0, so : the activation suppresses the signal. But unlike ReLU, the transition is not abrupt. Near , the Gaussian CDF is smooth and differentiable, meaning gradients can flow through even for slightly negative inputs. This smoothness is the key advantage over ReLU.
GELU's smoothness provides better gradient flow compared to ReLU's hard cutoff. To understand why this matters for training, let's compare the derivatives:

The shaded region shows where GELU outperforms ReLU: for slightly negative inputs, ReLU has zero gradient (the "dying ReLU" problem), while GELU allows some gradient to flow. This small difference compounds across billions of training steps and millions of neurons, potentially making the difference between successful and failed optimization at large scales.
Other Modifications
Several additional changes improved GPT-2's training stability. Each reflects a coherent engineering philosophy: keep activations well-behaved throughout the forward pass and keep gradients well-behaved throughout the backward pass.
Increased vocabulary size: GPT-2 uses a vocabulary of 50,257 tokens (versus GPT-1's roughly 40,000), built using byte-level Byte Pair Encoding (BPE). The byte-level encoding is important: it allows the model to represent any text without unknown tokens, because every possible byte can be encoded. This means GPT-2 can handle code, foreign language strings, special characters, and unusual punctuation without discarding information. The odd number 50,257 arose from starting with 256 base bytes, adding 50,000 BPE merge rules, and adding one special end-of-text token.
Extended context length: The context window doubled from 512 to 1024 tokens, letting the model to condition on longer passages. For language modeling, more context is almost always better: a model that can see the last two pages of a document will make better predictions about the next word than one limited to half a page. The doubling here was bounded by memory constraints on the hardware available in 2019.
Modified initialization: Weights in residual layers were scaled by , where is the number of residual layers. Each transformer block has two residual connections (one for attention and one for feed-forward), so for a 12-layer model, , and the scale factor is approximately . Without this scaling, each residual addition increases the variance of activations by a constant amount. After additions, variance grows proportionally to . Scaling by keeps the variance of the residual stream approximately constant regardless of depth, analogous to how the scaling in attention keeps dot-product magnitudes controlled.
Final layer normalization: An additional LayerNorm is applied after the last transformer block, before the output projection. With pre-normalization, the final transformer block's output is added directly to the residual stream without normalization. The final LayerNorm ensures this accumulated residual stream is properly normalized before being projected to logits.
Let's implement a minimal GPT-2 model to see these components together:
class GPT2(nn.Module):
"""Minimal GPT-2 implementation."""
def __init__(
self,
vocab_size: int = 50257,
hidden_size: int = 768,
num_layers: int = 12,
num_heads: int = 12,
max_position: int = 1024,
dropout: float = 0.1,
):
super().__init__()
self.hidden_size = hidden_size
self.num_layers = num_layers
# Token and position embeddings
self.wte = nn.Embedding(vocab_size, hidden_size)
self.wpe = nn.Embedding(max_position, hidden_size)
self.drop = nn.Dropout(dropout)
# Transformer blocks
self.blocks = nn.ModuleList(
[
GPT2Block(hidden_size, num_heads, dropout)
for _ in range(num_layers)
]
)
# Final layer norm (pre-norm style means we need this at the end)
self.ln_f = nn.LayerNorm(hidden_size, eps=1e-5)
# Output projection (tied with token embeddings)
self.lm_head = nn.Linear(hidden_size, vocab_size, bias=False)
self.lm_head.weight = self.wte.weight # Weight tying
# Initialize weights
self.apply(self._init_weights)
# Scale residual projections
for block in self.blocks:
block.c_proj.weight.data *= 1.0 / np.sqrt(2 * num_layers)
block.c_proj_ffn.weight.data *= 1.0 / np.sqrt(2 * num_layers)
def _init_weights(self, module):
if isinstance(module, nn.Linear):
torch.nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
torch.nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
torch.nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, input_ids: torch.Tensor) -> torch.Tensor:
B, T = input_ids.size()
device = input_ids.device
# Get embeddings
position_ids = torch.arange(T, device=device).unsqueeze(0)
token_emb = self.wte(input_ids)
pos_emb = self.wpe(position_ids)
x = self.drop(token_emb + pos_emb)
# Apply transformer blocks
for block in self.blocks:
x = block(x)
# Final layer norm and output projection
x = self.ln_f(x)
logits = self.lm_head(x)
return logitsGPT-2 Small: 124,439,808 parameters
GPT-2 Medium: 354,823,168 parameters
GPT-2 Large: 774,030,080 parameters
GPT-2 XL: 1,557,611,200 parameters
Our implementation produces parameter counts closely matching the official GPT-2 specifications. Small discrepancies arise from minor implementation details like bias terms and layer norm parameters, but the overall scaling matches: roughly 3x increase from Small to Medium, 2.2x from Medium to Large, and 2x from Large to XL. Weight tying between the input embeddings and output projection reduces parameter count while maintaining performance, a technique that became standard in language models. The intuition behind weight tying is elegant: the embedding matrix translates tokens into vectors at the input, and the output projection translates vectors back into token scores. These two operations are inverses of each other, so sharing their weights is parameter-efficient and conceptually coherent.
WebText: Learning from the Internet
GPT-2's training data represented a significant departure from previous approaches. Rather than using carefully curated datasets like BooksCorpus (used for GPT-1) or Wikipedia, OpenAI created WebText by scraping outbound links from Reddit posts with at least 3 upvotes. This might sound like a small methodological detail, but it shaped what the model learned.
The motivation behind this shift was scale and diversity. Books and Wikipedia are high-quality but limited. Books tend toward certain narrative styles and genres. Wikipedia is encyclopedic and factual but covers only a fraction of human knowledge and communication. The internet, for all its chaos, contains text in every genre and style, across registers and domains. By training on web text filtered for some quality signal, the researchers could access training data orders of magnitude larger than curated datasets while still avoiding the worst content.
WebText contains approximately 40 GB of text from 8 million web pages. The filtering heuristic of requiring Reddit upvotes served as a proxy for content quality, as humans had implicitly judged these links worth sharing. Wikipedia was deliberately excluded to enable fair evaluation on Wikipedia-based benchmarks.
Data Collection and Filtering
The WebText construction process involved several key decisions, each which reflects a deliberate choice about the tradeoff between data quantity and quality:
-
Source selection: Links from Reddit with karma (upvotes minus downvotes) were collected. This crowdsourced filtering biased toward content that humans found interesting, informative, or entertaining. A Reddit post linking to random spam is unlikely to accumulate three upvotes; a link to an interesting news article, a thoughtful opinion piece, or a useful tutorial is much more likely to be shared. Reddit thus served as a massive, distributed quality-filtering system for the entire web.
-
Deduplication: Near-duplicate documents were removed to prevent the model from memorizing repeated content. This is important for model quality: a model trained on millions of copies of the same article will reproduce that article well but generalize poorly.
-
Minimal preprocessing: Unlike many previous datasets, WebText preserved formatting and punctuation, along with case. The model learned from text as it appears naturally on the web, including code and tables as well as quoted material. This choice meant the model saw a wider variety of text styles, which likely contributed to its versatility.
-
Exclusion of Wikipedia: Wikipedia was deliberately excluded to enable fair evaluation on Wikipedia-based benchmarks. Including it would have made it impossible to evaluate the model's ability to generalize on any test derived from Wikipedia content.
This data collection strategy reflected a key insight: the internet contains a vast range of text. By learning from this diversity, GPT-2 could potentially generalize to many tasks without task-specific training. The internet also contains implicit task demonstrations: question-and-answer threads, translation examples, text followed by summaries, arguments followed by counterarguments. The model learned these patterns not by being taught them but by absorbing millions of naturally occurring examples.

The WebText approach changed what followed for what the model could do. First, it scaled easily since the internet provides virtually unlimited text. Second, the diversity exposed the model to many natural "tasks" embedded in web text: answering questions, summarizing articles, translating between languages, and explaining concepts. Third, it introduced biases present in Reddit's user base and the content they share. Reddit's demographic skews toward English-speaking, technically educated, younger users, and this bias became part of GPT-2's worldview.
Why Web Text Teaches Tasks Without Supervision
A subtle but important point is that WebText does not just provide language patterns. It provides implicit task demonstrations. Consider what kinds of text appear naturally on the internet: an article followed by a comments section where users summarize or critique it teaches summarization. A Stack Overflow question followed by a correct answer teaches question answering. A news article in English followed by its translation teaches translation. A recipe followed by user reviews teaches evaluation. The model absorbs all of these as instances of next-token prediction, but in doing so it learns the structure of each task type. This is why the zero-shot capabilities described in the next section are not magic: they are the expected consequence of training on text that contains millions of implicit task demonstrations.
Zero-Shot Task Performance
GPT-2's most surprising contribution was showing zero-shot task performance. Without any task-specific fine-tuning, GPT-2 could perform reading comprehension and summarization, as well as translation and question answering, simply by framing these tasks as language modeling. This was not a minor improvement on existing methods; it represented a fundamentally different relationship between a model and the tasks it could perform.
Before GPT-2, the standard workflow in NLP was: pretrain a model on a large corpus, then fine-tune it on labeled examples for each downstream task. Each task required its own fine-tuned model, its own training data, and its own deployment. GPT-2 suggested an alternative: one model, prompted appropriately, could attempt any task. The quality was not always competitive with fine-tuned specialists, but the existence of the capability at all was striking.
The Zero-Shot Paradigm
The key insight was that many NLP tasks can be expressed as conditional text generation. Rather than training separate models for each task, you can prompt a language model with a description of the task:
- Question answering: "Q: What is the capital of France? A:"
- Summarization: "Article: [article text] TL;DR:"
- Translation: "English: Hello, how are you? French:"
The language model, having seen similar patterns in its training data, learns to complete these prompts appropriately.
# Demonstrate zero-shot prompting patterns
zero_shot_examples = {
"reading_comprehension": """
Passage: The Eiffel Tower is a wrought-iron lattice tower on the Champ de Mars
in Paris, France. It is named after the engineer Gustave Eiffel, whose company
designed and built the tower.
Question: Who designed the Eiffel Tower?
Answer:""",
"summarization": """
Article: Scientists have discovered a new species of deep-sea fish in the
Mariana Trench. The fish, nicknamed the "ghost fish" due to its translucent
appearance, was found at depths exceeding 8,000 meters. Researchers believe
the fish has adapted unique biological mechanisms to survive the extreme
pressure at such depths.
TL;DR:""",
"translation": """
English: The weather is beautiful today.
French:""",
"sentiment": """
Review: This movie was absolutely incredible! The acting was superb and the
plot kept me on the edge of my seat the entire time.
Sentiment:""",
}Zero-Shot Prompt Examples: ============================================================ READING COMPREHENSION: ---------------------------------------- Passage: The Eiffel Tower is a wrought-iron lattice tower on the Champ de Mars in Paris, France. It is named after the engineer Gustave Eiffel, whose company designed and built the tower. Question: Who designed the Eiffel Tower? Answer: SUMMARIZATION: ---------------------------------------- Article: Scientists have discovered a new species of deep-sea fish in the Mariana Trench. The fish, nicknamed the "ghost fish" due to its translucent appearance, was found at depths exceeding 8,000 meters. Researchers believe the fish has adapted unique biological mechanisms to survive the extreme pressure at such depths. TL;DR: TRANSLATION: ---------------------------------------- English: The weather is beautiful today. French: SENTIMENT: ---------------------------------------- Review: This movie was absolutely incredible! The acting was superb and the plot kept me on the edge of my seat the entire time. Sentiment:
Each prompt demonstrates a different task expressed as conditional text generation. The reading comprehension prompt provides context then asks a question. The summarization prompt uses "TL;DR:" as a signal the model learned from internet forums. The translation prompt establishes a bilingual pattern the model should continue. The sentiment prompt gives a review and expects the model to complete the word that naturally follows. These formats work because WebText contained millions of similar patterns, each teaching the model what kind of text should complete each kind of prefix.
Benchmark Results
OpenAI evaluated GPT-2 on various benchmarks in zero-shot mode, comparing against supervised baselines that were explicitly trained for each task. Perplexity is the primary metric: it measures how surprised the model is by each word in the test text. Lower perplexity means the model assigns higher probability to the correct next word. A perfect model would have perplexity of 1 (never surprised); a random model over a 50,000-word vocabulary would have perplexity around 50,000.

The relationship between model size and perplexity follows a consistent pattern across benchmarks. Let's examine this scaling behavior more closely:

On a log-log plot, the near-linear relationship between parameters and perplexity suggests power-law scaling. This empirical observation, that performance improves predictably with scale, would later be formalized in the "Scaling Laws for Neural Language Models" paper by Kaplan et al. (2020). The consistency across diverse benchmarks indicates that scaling benefits are not dataset-specific but reflect improvements in language modeling capability. The key insight is that perplexity follows a power law with parameter count, meaning you can predict in advance how much a model will improve if you double its size.
Several patterns emerged from these evaluations:
-
Consistent scaling: Larger models achieved lower perplexity across all benchmarks. GPT-2 XL achieved state-of-the-art on Penn Treebank despite never being explicitly trained on it.
-
Domain transfer: Performance on WikiText (derived from Wikipedia) was strong despite Wikipedia being excluded from training, suggesting the model learned transferable language patterns.
-
Task emergence: Capabilities like reading comprehension and translation emerged without task-specific training, purely from scale and diverse pretraining.
Limitations of Zero-Shot
Zero-shot performance, while impressive, came with clear limitations:
-
Inconsistency: The model sometimes generated plausible-sounding but incorrect answers, particularly for factual questions.
-
Sensitivity to prompting: Small changes in prompt format could significantly affect performance.
-
Task ambiguity: Without examples, the model sometimes misinterpreted what was being asked.
These limitations motivated the subsequent exploration of few-shot learning, where giving a handful of examples dramatically improved task performance. GPT-3 would show that giving just 10 to 20 examples in the prompt could match or exceed the performance of fine-tuned models trained on thousands of examples. The limitations of pure zero-shot learning thus look less like a basic barrier and more like a problem of insufficient context.
Generation Quality
GPT-2's text generation quality captured public attention more than its benchmark scores. The model could produce coherent, contextually appropriate text that often fooled readers into thinking it was human-written. This section examines the technical mechanisms behind that quality and the parameters that control the tradeoff between coherence and creativity.
Coherent Long-Form Generation
Unlike previous language models that quickly devolved into nonsense, GPT-2 maintained coherence over paragraphs and pages. The 1,024-token context window was part of this: the model could remember the topic and characters, along with constraints introduced earlier in a passage. But the transformer's attention mechanism was equally important. Unlike RNN-based models that encoded context into a fixed-size hidden state, the transformer can directly attend to any position in the context window, maintaining specific details without them degrading through recurrent state updates.
Think of the difference this way. An RNN reading a 1,000-word passage must compress everything it reads into a fixed-size vector, the way you might try to remember a long lecture using only five index cards. The transformer, by contrast, keeps the entire lecture transcript available and can look back at any part of it when generating the next word. The 1,024-token context means GPT-2 keeps roughly 750 words of context fully accessible for every generation step.
Let's examine the sampling strategies that enable quality generation:
def top_k_sampling(
logits: torch.Tensor, k: int = 50, temperature: float = 1.0
) -> int:
"""Sample from top-k logits with temperature scaling."""
# Apply temperature
logits = logits / temperature
# Get top-k logits and indices
top_k_logits, top_k_indices = torch.topk(logits, k)
# Convert to probabilities
probs = F.softmax(top_k_logits, dim=-1)
# Sample from the distribution
sampled_idx = torch.multinomial(probs, num_samples=1)
return top_k_indices[sampled_idx].item()
def nucleus_sampling(
logits: torch.Tensor, p: float = 0.9, temperature: float = 1.0
) -> int:
"""Sample from nucleus (top-p) distribution."""
# Apply temperature
logits = logits / temperature
# Sort logits and get cumulative probabilities
sorted_logits, sorted_indices = torch.sort(logits, descending=True)
cumulative_probs = torch.cumsum(F.softmax(sorted_logits, dim=-1), dim=-1)
# Remove tokens outside the nucleus
sorted_indices_to_remove = cumulative_probs > p
# Keep at least one token
sorted_indices_to_remove[0] = False
# Set removed logits to -inf
sorted_logits[sorted_indices_to_remove] = float("-inf")
# Convert to probabilities and sample
probs = F.softmax(sorted_logits, dim=-1)
sampled_idx = torch.multinomial(probs, num_samples=1)
return sorted_indices[sampled_idx].item()The quality of generated text depends heavily on the sampling strategy. Pure random sampling from the full distribution produces incoherent text because low-probability tokens occasionally get selected. Consider a case where the model is 90% sure the next word is "the" but random sampling occasionally selects one of the 49,000 other tokens. Even a 1% chance of selecting a bizarre word means that on average once per hundred tokens, the text swerves in an unexpected direction. Over a 500-word passage, this guarantees several incoherent intrusions. Top-k and nucleus sampling restrict generation to high-probability tokens, dramatically improving coherence.


Temperature and Creativity
The temperature parameter provides a tunable tradeoff between coherence and creativity. Temperature scales the logits before applying softmax, controlling how "peaked" or "flat" the resulting probability distribution becomes.
The temperature-scaled softmax computes the probability of selecting token as:
where:
- : the probability of selecting token
- : the raw logit (unnormalized score) for token from the model's output layer
- : the temperature parameter, a positive scalar
- : the vocabulary size (total number of possible tokens)
- : the exponential function
When , this reduces to standard softmax. As , the distribution becomes increasingly peaked around the maximum logit (approaching argmax). As , the distribution approaches uniform, giving all tokens equal probability regardless of their logits. This happens because dividing by a large compresses all logit differences toward zero.
Why does this formula make sense? Notice that dividing each logit by rescales the differences between logits before softmax is applied. When , a logit difference of 2 becomes a difference of 4, making the leading token much more dominant. When , a logit difference of 2 becomes only 1, compressing the probability gap between tokens. Temperature is thus a direct control on how decisive the model is: low temperature makes it more deterministic and conservative, high temperature makes it more exploratory and potentially unpredictable.

At temperature 0.3, the model almost always selects the highest-probability token, creating repetitive but grammatically correct text. At temperature 1.0 (the default), it balances variety with coherence. At temperature 2.0, even low-probability tokens have significant chance of selection, leading to creative but potentially nonsensical output. For practical applications: use low temperatures (0.3 to 0.7) for factual question answering where accuracy matters, and higher temperatures (0.9 to 1.2) for creative writing or brainstorming where diversity is valuable.
Sample Generation
Let's demonstrate generation using a pretrained GPT-2 model from Hugging Face:
from transformers import GPT2LMHeadModel, GPT2Tokenizer
# Load pretrained GPT-2 (smallest model for efficiency)
tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")
model.eval()
def generate_text(
prompt: str,
max_new_tokens: int = 50,
temperature: float = 0.8,
top_p: float = 0.9,
) -> str:
"""Generate text continuation using GPT-2."""
input_ids = tokenizer.encode(prompt, return_tensors="pt")
with torch.no_grad():
output_ids = model.generate(
input_ids,
max_new_tokens=max_new_tokens,
temperature=temperature,
top_p=top_p,
do_sample=True,
pad_token_id=tokenizer.eos_token_id,
)
return tokenizer.decode(output_ids[0], skip_special_tokens=True)The generated text demonstrates GPT-2's ability to maintain topical coherence and grammatical correctness. It continues prompts in contextually appropriate ways, though close inspection often reveals subtle errors in logic or factual accuracy. The first prompt tends to produce technology commentary that sounds authoritative but may conflate different AI approaches. The creative writing prompt typically produces flowing narrative prose. The science prompt often produces plausible-sounding but fabricated findings. These patterns are characteristic of GPT-2: impressively coherent surface structure, unreliable deep content.
Worked Example: Tracing a Forward Pass
To solidify your understanding of how GPT-2 processes text, let's trace a minimal forward pass step by step with concrete dimensions. We'll use GPT-2 Small as our example, with hidden size , 12 layers, and 12 attention heads.
Suppose the input is the sequence "The cat sat", which tokenizes to 3 tokens with indices, say, [464, 3797, 3332].
Step 1: Token Embedding Lookup. We look up each token index in the token embedding table, which has shape . For 3 tokens, we retrieve a matrix of shape . Each row is the 768-dimensional representation of that token.
Step 2: Position Embedding Lookup. We look up positions 0, 1, 2 in the position embedding table, also of shape . We retrieve another matrix. Adding token embeddings and position embeddings gives the initial hidden state of shape .
Step 3: Pre-Norm Attention in Layer 1. We apply LayerNorm to , creating of shape with mean 0 and variance 1 per position. We then project through the combined projection matrix of shape to get a matrix, which we split into three matrices for , , and .
Step 4: Multi-Head Attention Split. We reshape each of , , and from into by splitting the 768 dimensions across 12 heads of 64 dimensions each. For head , the query matrix is of shape , and similarly for and .
Step 5: Scaled Dot-Product Attention. For each head , we compute attention scores and weighted values:
where:
- has shape : a score matrix where entry measures how much position should attend to position
- : the scaling factor preventing large dot products from saturating the softmax
- The causal mask zeroes out the upper triangle: position 0 only attends to itself, position 1 attends to positions 0 and 1, position 2 attends to all three
The output has shape .
Step 6: Concatenate and Project. We concatenate the 12 head outputs to get a matrix, then project through the output projection of shape , giving a attention output.
Step 7: First Residual Addition. We add the attention output to the original : . Both have shape . The residual connection preserves the original information while the attention output adds what the attention mechanism learned.
Step 8: Pre-Norm Feed-Forward. We apply LayerNorm to , then pass through the feed-forward network. The first linear layer expands from 768 to 3072 dimensions, GELU activation is applied, and the second linear layer compresses back to 768. The result has shape .
Step 9: Second Residual Addition. We add the feed-forward output to : . This completes one transformer block.
Steps 10 through 21: Repeat for layers 2 through 12. The hidden state passes through 11 more identical blocks. Each adds its learned transformations via residual connections, progressively enriching each token's representation with broader context.
Step 22: Final LayerNorm. After all 12 layers, we apply the final layer normalization to the output.
Step 23: Logit Projection. We multiply by the transposed token embedding matrix of shape , creating logits of shape . For next-token prediction, we take the last position's logits (shape ) and sample from this distribution to generate the 4th token.
This entire forward pass computes all three positions simultaneously during training, with the causal mask so each position only uses information from prior positions. The efficiency comes from this parallel computation: rather than running the model three times sequentially to train on positions 1, 2, and 3, we train on all positions in a single forward pass.
Limitations and Impact
GPT-2 introduced capabilities that transformed expectations for language models, but it also revealed basic challenges that persist in larger models today. Understanding these limitations is not just academic: they directly shaped the research agenda that produced GPT-3, InstructGPT, and ChatGPT.
Factual Reliability
GPT-2 generates fluent text that can contain fabricated facts presented with confidence. The model has no mechanism to verify statements against ground truth; it simply produces statistically likely continuations. A question about a historical event might produce a plausible but entirely invented answer, complete with dates and names along with details that fit the expected format of a factual response but bear no relationship to what happened.
This creates a dangerous asymmetry: the text sounds authoritative but may be wrong. The problem is structural, not a matter of training on more data. GPT-2 was never trained to distinguish true from false statements; it was trained to predict the next token. Tokens that follow a confident, authoritative sentence tend to be more confident, authoritative content. The model optimizes for linguistic plausibility, and linguistic plausibility is highly correlated with factual-sounding language but only weakly correlated with actual accuracy. For knowledge-intensive tasks like medical advice or legal guidance, GPT-2's generations are unreliable without external verification. This limitation catalyzed research into retrieval-augmented generation, where the model retrieves relevant documents before generating answers, and fact-checking systems that verify claims post-generation.
Bias and Toxicity
The WebText training data, while filtered for quality through Reddit upvotes, inherited biases from its source. Reddit's user base skews toward younger, English-speaking, male users with particular interests and political views. Content shared heavily on Reddit reflects these demographics. GPT-2 can generate text which reflects stereotypes, political biases, and offensive content present in its training distribution, not because it was designed to but because such language appeared in its training data in statistically meaningful quantities.
OpenAI's decision to stage the model's release was partly motivated by concerns about automated generation of targeted harassment and misinformation. A system that can write convincing fake news articles, generate personalized spam, or produce targeted harassment at scale poses real societal risks. Subsequent work on RLHF (reinforcement learning from human feedback) and constitutional AI directly addresses these safety concerns by training models to refuse harmful requests and align their outputs with human values. GPT-2 made the problem visible; the solutions came later.
Coherence Limits
While GPT-2 produces impressive short-to-medium length text, it often loses coherent structure over very long passages. Characters change names, timelines contradict each other, and factual premises established early in a story are forgotten later. These coherence failures reflect a basic limitation of the causal language modeling objective: the model is trained to predict the next token given the last 1,024 tokens, but it receives no training signal about long-range structural consistency. There is no loss term that penalizes contradicting a premise from 800 tokens ago.
What GPT-2 Enabled
Despite these limitations, GPT-2 fundamentally changed the field in ways that are still felt today:
-
Scale as a strategy: GPT-2 demonstrated that simply making models larger could yield qualitatively new capabilities. This insight drove the push to GPT-3, PaLM, and other massive models. Before GPT-2, the conventional wisdom was that architectural innovations were necessary for capability improvements. After GPT-2, scaling became a primary research direction.
-
Zero-shot learning: The idea that a single model could attempt many tasks without fine-tuning opened research into prompting, in-context learning, and instruction tuning. GPT-2 showed the proof of concept; subsequent work refined and extended it.
-
Generation quality: GPT-2 set a new bar for text generation coherence, making language model outputs useful for drafting and brainstorming, as well as creative applications. The commercial applications that followed, including writing assistants and code completion tools, trace their lineage directly to GPT-2's demonstration that neural text generation could be practically useful.
-
Open research: After staged release, OpenAI eventually published the full GPT-2 model. This enabled extensive research into interpretability and safety, as well as applications. Hundreds of papers studying GPT-2's representations, failure modes, and downstream uses emerged in the following years, creating insights that informed the design of subsequent systems.
The transition from GPT-1 to GPT-2 established the scaling paradigm that would define the next several years of language model development: more parameters, more data, more compute. Each increment revealed new capabilities that smaller models lacked. The path from GPT-2 to GPT-3 (a 100x increase in parameters) and beyond was now clear.
Summary
GPT-2 demonstrated that scaling autoregressive language models yields emergent capabilities beyond next-token prediction. The key contributions include:
-
Model family: Four sizes from 117M to 1.5B parameters, letting systematic study of scaling behavior while keeping architecture constant. This design pattern became the template for virtually all subsequent large language model releases.
-
Architectural refinements: Pre-normalization for stable training at depth, GELU activation for smoother gradients and better gradient flow through slightly negative activations, and modified initialization to control variance growth across deep residual networks.
-
WebText training: 40 GB of text from Reddit-filtered web pages. This provides diverse, naturally-occurring task demonstrations without explicit supervision. The internet's natural embedding of question-answering and summarization, along with translation patterns, was the implicit curriculum that enabled zero-shot transfer.
-
Zero-shot capabilities: Reading comprehension and summarization, along with translation, emerged from pure language modeling, achieved through prompting that aligned test-time queries with training-time text patterns. The model never learned these tasks explicitly; it absorbed them through statistical regularity in the training data.
-
Generation quality: Coherent multi-paragraph text generation that, with appropriate sampling strategies (nucleus sampling, temperature tuning), could pass for human-written. The 1,024-token context window and transformer's direct attention mechanism were the technical enablers of this coherence.
-
Scaling laws foreshadowed: The consistent power-law relationship between model size and perplexity across multiple benchmarks presaged the formal scaling laws that Kaplan et al. would publish in 2020, establishing that capability improvements are predictable functions of compute and data.
GPT-2 marked the transition from fine-tuned specialists to general-purpose language models. Its limitations around factual reliability and bias, along with weak long-range coherence, remain active research areas. But its core insight, that scale enables emergence, fundamentally reshaped expectations for what language models could become. In the next chapter, we'll examine GPT-3, which took this insight to its logical conclusion by scaling to 175 billion parameters and finding that few-shot learning, when the model is large enough, can match or exceed fine-tuned performance on many benchmarks.
Key Parameters
When working with GPT-2 or implementing similar architectures, these parameters most directly affect model behavior:
Architecture Parameters:
-
hidden_size (768 to 1600): The dimensionality of token representations. Larger values increase model capacity but quadratically increase attention computation. GPT-2 variants use 768, 1024, 1280, and 1600.
-
num_layers (12 to 48): The number of stacked transformer blocks. More layers enable deeper reasoning but increase memory linearly. Typical values follow powers of 12.
-
num_heads (12 to 25): The number of parallel attention heads. Must evenly divide hidden_size. More heads allow attending to different relationship types simultaneously.
-
max_position (1024): Maximum sequence length the model can process. Longer contexts enable conditioning on more information but increase memory quadratically with sequence length.
Generation Parameters:
-
temperature (0.0 to 2.0): Controls randomness in sampling. Values below 1.0 sharpen the distribution (more deterministic), values above 1.0 flatten it (more random). Common choices: 0.7 for focused text, 1.0 for balanced, 1.2+ for creative writing.
-
top_k (1 to 100): Restricts sampling to the k highest-probability tokens. Lower values increase coherence but reduce diversity. Common choices: 40-50 for general use, lower for factual tasks.
-
top_p (0.0 to 1.0): Nucleus sampling threshold. Keeps tokens until cumulative probability exceeds p. Adapts to the shape of each distribution. Common choices: 0.9-0.95 for balanced generation.
-
max_new_tokens: Number of tokens to generate. Longer generations risk coherence degradation. Consider the task: summaries need fewer tokens than stories.
Training Parameters:
-
dropout (0.0 to 0.3): Applied to attention weights and feed-forward outputs during training. GPT-2 uses 0.1. Higher values prevent overfitting on small datasets but slow convergence.
-
weight_decay (0.0 to 0.1): L2 regularization strength. GPT-2 uses 0.01. Prevents weights from growing too large, improving generalization.
GPT-2 Knowledge Check
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!