GPT-3: Scale, Few-Shot and In-Context Learning

Michael BrenndoerferUpdated July 21, 202593 min read

Part of Language AI Handbook

Examines GPT-3's 175B parameter architecture, the emergence of few-shot learning, in-context learning mechanisms.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

GPT-3: Scale, Few-Shot Learning, and In-Context Adaptation

GPT-2 demonstrated that language models could perform tasks without task-specific training, achieving modest zero-shot performance on reading comprehension alongside translation and summarization. But this capability was limited: the model often required fine-tuning to match specialized systems, and zero-shot performance lagged well behind purpose-built models trained on labeled data. For most real-world applications, fine-tuning remained the dominant paradigm. You would take a pretrained model, collect thousands of labeled examples for your specific task, and spend GPU time adapting the model's weights to your domain. GPT-3 changed this equation through sheer scale.

With 175 billion parameters, over 100 times larger than GPT-2, GPT-3 revealed that massive language models could learn new tasks from just a few examples provided in the prompt, a capability called few-shot learning. The model did not need gradient updates, additional training, or labeled datasets. You simply described the task and showed a handful of demonstrations inside the prompt, and the model completed the pattern. This was not a minor incremental improvement. It was a qualitative shift in how language models could be used.

This discovery reshaped how the field thinks about language model capabilities. Instead of training separate models for each task, a single pretrained model could adapt on the fly. Prompting replaced fine-tuning for many applications. Practitioners who once spent weeks collecting and labeling data for a new task could instead spend minutes crafting a good prompt. The economic and practical implications were enormous. The implications also extended beyond NLP: GPT-3 demonstrated that scale itself could be a source of emergent capabilities, behaviors that appear suddenly as models grow larger rather than improving smoothly and gradually with size.

The paper introducing GPT-3, "Language Models are Few-Shot Learners" by Brown et al. (2020), became one of the most influential papers in the history of natural language processing. It presented systematic empirical evidence that performance across dozens of diverse NLP benchmarks scaled predictably with model size and the number of in-context examples. It introduced rigorous definitions distinguishing zero-shot, one-shot, and few-shot evaluation. And it raised fundamental questions about what it means for a model to "learn" a task, questions that researchers are still working through today.

In this chapter, we explore GPT-3's architecture and training, the discovery of in-context learning, the difference between zero-shot, one-shot, and few-shot prompting, and what GPT-3 revealed about the relationship between scale and capability. We also examine the limitations that scale did not solve and the lasting impact GPT-3 had on research directions and industry practice.

The Scale of GPT-3

GPT-3's defining characteristic is its size. At 175 billion parameters, it dwarfed all previous language models. To understand what this means in practice, it helps to compare it with its predecessors and to think carefully about what parameters represent.

Each parameter is a single floating-point number, a weight in a matrix that the model adjusts during training to minimize its prediction error. Parameters encode the model's learned knowledge about language: grammar, facts, reasoning patterns, stylistic conventions, and conceptual associations. More parameters mean more representational capacity, which in principle allows the model to store and manipulate more complex patterns. The question GPT-3's creators asked was whether simply scaling up the number of parameters, together with proportionally more data and compute, would yield qualitatively new capabilities rather than just marginal improvements.

The answer, as the table below shows, involved a dramatic escalation across every architectural dimension:

The GPT model progression shows exponential growth in scale. GPT-3 has roughly 117 times more parameters than GPT-2 and 1500 times more than GPT-1.
ModelParametersLayersHidden SizeHeadsContext Length
GPT-1117M1276812512
GPT-21.5B481600251024
GPT-3175B9612288962048

The jump from GPT-2 to GPT-3 represents more than a linear improvement. With 96 transformer layers, a hidden dimension of 12,288, and 96 attention heads, GPT-3 can represent far more complex patterns and relationships. The context length of 2,048 tokens means the model can consider roughly 1,500 words of context when making predictions. Each of these numbers reflects a deliberate design choice about where to invest capacity. The hidden dimension controls how rich the representation of each token can be. The number of layers controls how many successive transformations the model applies before producing its output. The number of attention heads determines how many different "perspectives" the model can use when looking at relationships between tokens. GPT-3 pushed all of these dimensions far beyond what had been tried before.

Think of the hidden dimension as the vocabulary of internal concepts the model works with. GPT-1's 768-dimensional hidden state can represent 768 independent directions of variation in token meaning. GPT-3's 12,288-dimensional space offers 16 times as many dimensions, allowing much finer distinctions. A word like "bank" can be represented differently depending on whether the surrounding context involves rivers, money, or aviation, and a higher-dimensional space gives the model more room to encode these distinctions precisely.

Out[4]:
Visualization
Bar chart comparing GPT-1, GPT-2, and GPT-3 parameter counts on log scale, showing exponential growth.
Parameter count comparison across GPT generations on a logarithmic scale. Each generation represents roughly a 10-100x increase in model size, with GPT-3 reaching 175 billion parameters, over 100 times larger than GPT-2.

Training Data and Compute

GPT-3's training data, while not publicly disclosed in full detail, was substantially larger than GPT-2's WebText. The choice of training data matters as much as its size. Raw internet text contains an enormous quantity of low-quality content: spam, duplicate pages, garbled text, and offensive material. Simply collecting more text does not automatically produce a better model. OpenAI's team invested significant effort in filtering and weighting sources to prioritize quality while retaining diversity.

The training corpus included five sources, each chosen for different reasons:

  • Common Crawl: A filtered version of web pages (410 billion tokens, weighted to contribute 60% of training). This provides the breadth necessary for general language understanding, but required aggressive quality filtering to remove noise.
  • WebText2: An expanded version of GPT-2's Reddit-curated dataset (19 billion tokens). The original WebText dataset used Reddit upvotes as a proxy for quality: only pages linked from Reddit posts with significant engagement were included. WebText2 extends this collection further.
  • Books1 and Books2: Two internet-based book corpora (12 billion and 55 billion tokens). Books provide long-form coherent text with complex narrative structure and reasoning, along with specialized vocabulary that is rare in web text.
  • Wikipedia: The English Wikipedia (3 billion tokens). Despite being small relative to other sources, Wikipedia contributes dense factual content with reliable, encyclopedic coverage of thousands of topics.

The total training corpus contained approximately 499 billion tokens, though the model saw roughly 300 billion tokens during training due to dataset weighting that favored higher-quality sources. Notice that the weighting here is deliberate: Common Crawl contributes 410 billion tokens but is weighted to only 60% of training samples. Meanwhile, Books1 and Books2 together contribute about 67 billion tokens but are sampled at much higher rates than their raw token count suggests. The intuition is straightforward: a well-written book paragraph is more valuable for learning coherent language than a mediocre web page, even if you have far more web pages available. This selective weighting is one of the reasons GPT-3 produces unusually fluent and coherent text compared to what you might expect from internet-scraped data alone.

Out[5]:
Visualization
Horizontal bar chart showing raw token counts for each training data source.
Raw token counts by source. Common Crawl dominates with 410 billion tokens, dwarfing all other sources combined.
Horizontal bar chart showing weighted training contribution percentages.
Weighted contribution to training. Despite its size, Common Crawl was downweighted to 60% while higher-quality book corpora were oversampled relative to their raw token count.
Training Compute

GPT-3 required approximately 3.14×10233.14 \times 10^{23} floating-point operations (FLOPs) to train, roughly 1,000 times more compute than GPT-2. At 2020 cloud computing prices, the estimated training cost exceeded $4 million. This figure does not include the cost of failed runs, hyperparameter searches, or preliminary experiments that informed the final training configuration.

The compute required for training grew even faster than the parameter count. This wasn't just about having more parameters; each parameter needed more training data to be utilized effectively. A rough estimate of training compute for a language model is:

C≈6⋅N⋅DC \approx 6 \cdot N \cdot D

where:

  • CC: total floating-point operations (FLOPs) for training
  • NN: number of model parameters
  • DD: number of training tokens
  • The factor of 6 accounts for the forward and backward passes (approximately 2 FLOPs per parameter per token for the forward pass, and 4 for the backward pass, since the backward pass computes gradients with respect to both activations and weights)

For GPT-3 with N=175×109N = 175 \times 10^9 parameters trained on D≈300×109D \approx 300 \times 10^9 tokens, this yields C≈3.15×1023C \approx 3.15 \times 10^{23} FLOPs, matching the reported training compute. This observation, later formalized in scaling laws, showed that optimal training balances model size with data quantity and training compute. You cannot simply make a model very large and train it briefly; the performance improvements from size only materialize when the model sees enough data to fill its representational capacity.

The GPT-3 Model Family

OpenAI released GPT-3 not as a single model but as a family of eight models spanning four orders of magnitude in size. This range enabled researchers to study how capabilities scale with model size and supplied the evidence for the paper's central claims.

The key insight behind releasing a model family rather than a single large model was methodological: if you claim that some capability emerges at scale, you need to show that smaller models do not have the capability. Without smaller models to compare against, you cannot distinguish between "this requires 175B parameters" and "this requires any reasonable language model." The systematic comparison across sizes transformed anecdotal observations about large models into rigorous empirical findings.

The GPT-3 model family spans from 125 million to 175 billion parameters, enabling systematic study of scaling effects.
Model NameParametersLayersHidden SizeHeads
GPT-3 Small125M1276812
GPT-3 Medium350M24102416
GPT-3 Large760M24153616
GPT-3 XL1.3B24204824
GPT-3 2.7B2.7B32256032
GPT-3 6.7B6.7B32409632
GPT-3 13B13B40512040
GPT-3 175B175B961228896

The smallest GPT-3 model is roughly equivalent to GPT-1 in size, while the largest is over 1,000 times bigger. Each size variant has a distinct architecture with different hidden dimensions and head counts instead of being a truncated version of the 175B model. This design provides cleaner comparisons across compact architectures.

This range allowed researchers to study emergent capabilities: abilities that appear suddenly at certain scales rather than improving gradually. For many tasks, performance on the 125M and 350M models would be near chance level, then improve modestly up to 13B, and then jump dramatically at 175B. This non-linear behavior suggested that in-context learning is not a gradual improvement but a threshold phenomenon: something about the 175B-scale model enables a qualitatively different kind of adaptation.

Emergent Capabilities

An emergent capability is a behavior that appears in larger models but is absent or near-random in smaller ones, even when performance is measured on the same scale. The word "emergent" is somewhat contested: some researchers argue that apparent emergence is an artifact of metric choice (sharp thresholds in accuracy can obscure smooth improvements in log-probability), while others maintain that discrete phase transitions occur. GPT-3 supplied early evidence for emergence and made it a central debate in scaling research.

Out[6]:
Visualization
Horizontal bar chart showing the 8 GPT-3 models from 125M to 175B parameters on a logarithmic scale.
The GPT-3 model family spans nearly four orders of magnitude in parameter count. The exponential spacing between model sizes enabled systematic study of how capabilities scale, revealing emergent behaviors at larger scales where smaller models showed near-chance performance.

Architectural Details

GPT-3 uses the same fundamental decoder-only transformer architecture as GPT-2. The core design remains unchanged: a stack of transformer decoder blocks with masked self-attention, each followed by a feed-forward network. This continuity was a deliberate choice. The GPT-3 authors wanted to isolate the effect of scale, not architecture innovation. By keeping the design consistent with GPT-2, any capability differences observed between the models can be attributed to size rather than architectural changes.

The key architectural differences from GPT-2 include:

  • Alternating dense and sparse attention: In some layers, GPT-3 uses sparse attention patterns that attend to every other token, reducing computational cost for long sequences
  • Learned position embeddings: Like GPT-2, positions are represented by learned embeddings rather than sinusoidal functions
  • Pre-layer normalization: Layer normalization is applied before each sublayer (attention and FFN) rather than after, following the Pre-LN configuration that stabilizes training for deep networks

Pre-layer normalization (Pre-LN) deserves particular attention. In the original transformer design, layer normalization was applied after the residual connection, a pattern called Post-LN. For shallow networks, this works fine. But for very deep networks like GPT-3 with 96 layers, Post-LN creates gradient flow problems: the gradients must pass through each normalization operation during backpropagation, and at great depths this can cause gradients to vanish or explode. Pre-LN addresses this by normalizing before the sublayer operation. This keeps the residual connections carry gradients cleanly from output back to input regardless of depth. This seemingly minor change was essential for training stability at 96 layers.

Parameter Count Calculation

When we say GPT-3 has 175 billion parameters, what exactly are we counting? Understanding where these parameters live reveals the model's structure and helps explain why scale matters so much for capability.

A transformer model stores its learned knowledge in weight matrices. Each weight is a single floating-point number that the model adjusts during training. The total count of these numbers determines the model's capacity to represent patterns in language. To understand GPT-3's scale, we need to trace through each component and count its parameters systematically.

The Building Blocks

A GPT-style model consists of several distinct components, each contributing parameters:

  1. Token embeddings: A lookup table that converts each vocabulary token into a dense vector
  2. Position embeddings: A separate lookup table that encodes where each token appears in the sequence
  3. Transformer layers: The repeated blocks that process and transform token representations
  4. Final layer normalization: A normalization layer before the output projection
  5. Output projection: A matrix that converts hidden states back to vocabulary probabilities

Let's formalize this. For a model with vocabulary size VV, maximum sequence length LL, model dimension dd, number of layers NN, and feed-forward dimension dffd_{ff}, the total parameter count is:

Ptotal=Pembed+Ppos+N⋅Player+Pnorm+PoutputP_{\text{total}} = P_{\text{embed}} + P_{\text{pos}} + N \cdot P_{\text{layer}} + P_{\text{norm}} + P_{\text{output}}

where:

  • Pembed=V⋅dP_{\text{embed}} = V \cdot d: token embedding parameters. Each of the VV vocabulary tokens gets its own dd-dimensional vector, requiring V×dV \times d parameters total.
  • Ppos=L⋅dP_{\text{pos}} = L \cdot d: position embedding parameters. Each of the LL possible positions gets a dd-dimensional vector.
  • PlayerP_{\text{layer}}: parameters per transformer layer, which we'll derive below.
  • Pnorm=2dP_{\text{norm}} = 2d: final layer normalization has a learnable scale and bias vector, each of dimension dd.
  • Poutput=d⋅VP_{\text{output}} = d \cdot V: the output projection maps from hidden dimension dd back to vocabulary size VV. This is often weight-tied with the token embeddings (sharing the same matrix), but GPT-3 uses separate weights.

The formula captures the intuition that larger vocabularies, longer sequences, and higher dimensions all increase parameter count. But the dominant term is N⋅PlayerN \cdot P_{\text{layer}}: the transformer layers, repeated NN times.

Inside a Transformer Layer

Each transformer layer consists of two sublayers, each followed by layer normalization:

  1. Multi-head self-attention: Allows tokens to gather information from other positions
  2. Feed-forward network (FFN): Applies the same transformation to each position independently

The attention sublayer requires four projection matrices: Query (WQ\mathbf{W}_Q), Key (WK\mathbf{W}_K), Value (WV\mathbf{W}_V), and Output (WO\mathbf{W}_O). Each matrix has shape d×dd \times d, contributing d2d^2 parameters. With four matrices, attention adds 4d24d^2 parameters.

The feed-forward network consists of two linear transformations. The first expands from dimension dd to dffd_{ff} (typically 4d4d), and the second contracts back from dffd_{ff} to dd. This requires d×dff+dff×d=2⋅d⋅dffd \times d_{ff} + d_{ff} \times d = 2 \cdot d \cdot d_{ff} parameters.

Each layer also has two layer normalizations (one before attention, one before FFN in the Pre-LN configuration), each with scale and bias vectors of dimension dd. This adds 4d4d parameters per layer.

Putting it together:

Player=Pattn+Pffn+Player_normsP_{\text{layer}} = P_{\text{attn}} + P_{\text{ffn}} + P_{\text{layer\_norms}}

where:

  • Pattn=4d2P_{\text{attn}} = 4d^2: the four attention projection matrices (WQ\mathbf{W}_Q, WK\mathbf{W}_K, WV\mathbf{W}_V, WO\mathbf{W}_O)
  • Pffn=2⋅d⋅dffP_{\text{ffn}} = 2 \cdot d \cdot d_{ff}: the two feed-forward linear layers. With dff=4dd_{ff} = 4d, this becomes 8d28d^2.
  • Player_norms=4dP_{\text{layer\_norms}} = 4d: scale and bias for two layer normalizations

Notice that the per-layer count scales with d2d^2. This quadratic dependence explains why increasing the hidden dimension has such a dramatic effect on model size: doubling dd quadruples the parameters in each layer. It also explains why GPT-3's hidden dimension of 12,288 is so consequential: compared to GPT-2's 1,600, this is roughly an 8x increase in dd, which translates to a 64x increase in parameters per layer from the quadratic terms alone.

A Worked Example: GPT-3 175B

Let's apply these formulas to GPT-3's actual configuration:

  • Vocabulary size: V=50257V = 50257 tokens (the GPT-2 tokenizer vocabulary)
  • Maximum sequence length: L=2048L = 2048 tokens
  • Model dimension: d=12288d = 12288
  • Number of layers: N=96N = 96
  • Feed-forward dimension: dff=4×12288=49152d_{ff} = 4 \times 12288 = 49152

First, we calculate the per-layer parameters:

Pattn=4×122882=4×150,994,944=603,979,776Pffn=2×12288×49152=1,207,959,552Player_norms=4×12288=49,152Player=603,979,776+1,207,959,552+49,152≈1.81 billion\begin{aligned} P_{\text{attn}} &= 4 \times 12288^2 = 4 \times 150{,}994{,}944 = 603{,}979{,}776 \\ P_{\text{ffn}} &= 2 \times 12288 \times 49152 = 1{,}207{,}959{,}552 \\ P_{\text{layer\_norms}} &= 4 \times 12288 = 49{,}152 \\ P_{\text{layer}} &= 603{,}979{,}776 + 1{,}207{,}959{,}552 + 49{,}152 \approx 1.81 \text{ billion} \end{aligned}

Each layer contributes about 1.81 billion parameters. With 96 layers, the transformer stack alone accounts for approximately 174 billion parameters. The remaining 1+ billion come from embeddings and the output projection.

This breakdown reveals an important pattern: the transformer layers dominate the parameter count. Increasing the number of layers or the hidden dimension has a much larger effect than expanding the vocabulary or context length. If you doubled the vocabulary to 100,514 tokens, the additional parameters would be less than 1.5 billion, barely 1% of the total. But if you doubled the hidden dimension to 24,576, you would quadruple the per-layer parameters and nearly quadruple the entire model. In practice, this means that the "knowledge" of a GPT-3-scale model lives overwhelmingly in its attention and FFN matrices, not in the embedding lookup tables.

Implementation

Let's implement this calculation to verify our formulas and explore how parameters distribute across components:

In[7]:
Code
def compute_gpt3_parameters(
    vocab_size=50257,
    n_layers=96,
    d_model=12288,
    n_heads=96,
    d_ff=None,
    max_seq_len=2048,
):
    """
    Calculate parameter count for a GPT-3 style model.

    The feed-forward dimension is 4x the model dimension by default.
    """
    if d_ff is None:
        d_ff = 4 * d_model

    # Token embeddings
    token_embedding_params = vocab_size * d_model

    # Position embeddings
    position_embedding_params = max_seq_len * d_model

    # Per-layer parameters
    # Attention: Q, K, V, O projections
    attention_params_per_layer = 4 * d_model * d_model

    # FFN: two linear layers
    ffn_params_per_layer = 2 * d_model * d_ff

    # Layer norms (2 per layer, each has scale and bias)
    layernorm_params_per_layer = 4 * d_model

    params_per_layer = (
        attention_params_per_layer
        + ffn_params_per_layer
        + layernorm_params_per_layer
    )

    total_layer_params = n_layers * params_per_layer

    # Final layer norm
    final_norm_params = 2 * d_model

    # Output projection (weight-tied with embedding or separate)
    # GPT-3 uses separate output weights
    output_params = d_model * vocab_size

    total = (
        token_embedding_params
        + position_embedding_params
        + total_layer_params
        + final_norm_params
        + output_params
    )

    return {
        "token_embedding": token_embedding_params,
        "position_embedding": position_embedding_params,
        "layers": total_layer_params,
        "final_norm": final_norm_params,
        "output": output_params,
        "total": total,
    }
Out[8]:
Console
GPT-3 175B Parameter Breakdown
=============================================
Token embeddings:        0.62B
Position embeddings:     0.03B
Transformer layers:    173.95B
Final layer norm:        0.02M
Output projection:       0.62B
---------------------------------------------
Total:                 175.21B

The output confirms our manual calculation. The vast majority of parameters, over 173 billion, reside in the 96 transformer layers. This makes sense: each layer contributes about 1.8 billion parameters through its attention projections and feed-forward network, and with 96 layers, these add up quickly.

The embedding layers tell an interesting story. Despite handling a vocabulary of 50,257 tokens, the token embeddings account for only about 0.62 billion parameters, less than 1% of the total. The position embeddings are even smaller at 0.025 billion (25 million) parameters. The output projection mirrors the token embedding in size but is kept separate in GPT-3 rather than being weight-tied.

This distribution has practical implications. If you want a larger model, adding layers or increasing the hidden dimension is far more effective than expanding the vocabulary or context length. Conversely, if you want a smaller model, the layers are where you need to cut.

The visualization below makes this distribution strikingly clear:

Out[9]:
Visualization
Pie chart showing GPT-3 parameter distribution with transformer layers as the largest segment.
Parameter distribution in GPT-3 175B. The transformer layers dominate, containing over 99% of all parameters. Embedding and output projections together account for less than 2%, illustrating how the model's representational capacity lives almost entirely in the processing layers rather than the input/output interfaces.

The pie chart reveals just how lopsided the distribution is: the transformer layers are so dominant that the embeddings and output projection are barely visible. This concentration of parameters in the processing layers, rather than the input/output interfaces, reflects where the model's computational "thinking" happens. The embeddings convert tokens to vectors and back, but the layers are where the model transforms and reasons about those representations. The pattern also has an important implication for interpretability: if you want to understand what a model knows or how it reasons, you need to study the attention and FFN matrices in the processing layers, not the embedding tables.

The Discovery of In-Context Learning

GPT-3's most significant contribution wasn't its architecture but what it revealed about large-scale language models: they can learn new tasks from examples provided in the prompt, without any gradient updates. This capability, called in-context learning (ICL), emerged as a surprise during evaluation.

In-Context Learning

In-context learning is the ability of a language model to perform a task by conditioning on a few examples (demonstrations) in the input prompt, without updating the model's parameters. The model "learns" the task pattern from the examples and applies it to new inputs.

To appreciate how surprising this was, recall the standard pre-GPT-3 paradigm. When you wanted a model to perform sentiment classification, you collected thousands of labeled reviews, fine-tuned a BERT-style model on them, and deployed the resulting specialized model. This process required labeled data (expensive to collect), GPU time (expensive to run), and engineering infrastructure (expensive to maintain). And you had to repeat the entire process for every new task.

In-context learning throws this paradigm out. Instead of updating the model's weights, you modify the input. You write a few examples of the task directly into the prompt, and the model infers what you want from context. The weights never change. No training loop runs. The adaptation happens purely in the model's forward pass, through the same attention mechanism it uses for everything else.

Consider a sentiment classification task. Instead of fine-tuning a model on thousands of labeled examples, you can simply show GPT-3 a few examples in the prompt:

Review: "This movie was absolutely fantastic!" Sentiment: Positive Review: "I wasted two hours of my life on this garbage." Sentiment: Negative Review: "The acting was decent but the plot made no sense." Sentiment:

GPT-3 would then complete this with "Negative" or "Mixed", having inferred the classification pattern from just two examples. The key insight is that the model treats the prompt as part of its input context and continues the pattern it observes. It sees that each review is followed by a sentiment label, and when it encounters a review without a label, it generates one.

This seems almost magical at first. How can the model infer what to do without being trained on it? The answer lies in pretraining. GPT-3 has seen billions of examples of people explaining concepts. This provides examples, and asking others to complete patterns. The format of a few-shot prompt, while novel as a machine learning technique, resembles things GPT-3 has encountered many times: educational texts, quizzes, worked examples, and teaching materials where an instructor shows examples and expects the student to continue the pattern. In-context learning may be less a new capability and more the application of a deeply learned skill, the skill of learning from examples, to the machine learning context.

Zero-Shot, One-Shot, and Few-Shot Learning

OpenAI's GPT-3 paper systematically compared three prompting strategies. Each represents a different level of guidance provided to the model:

  1. Zero-shot: The model receives only a task description, with no examples. It must understand the task from the description alone and generalize immediately.
  2. One-shot: The model sees one example before the test input. This provides a template for the expected input/output format, but offers minimal statistical signal about the task.
  3. Few-shot: The model sees multiple examples (typically 10-100, limited by context length). This provides both format information and statistical patterns about what kinds of inputs map to what kinds of outputs.

The distinction matters because it reveals different aspects of the model's capability. Zero-shot performance measures how well the model can understand natural language descriptions of tasks, essentially testing whether it has built sufficient meta-knowledge to interpret instructions. One-shot performance tests whether a single example is enough to fix the input/output format in the model's representation. Few-shot performance tests whether the model can pick up statistical regularities from a small sample, something closer to what we intuitively mean by "learning."

In practice, GPT-3's zero-shot performance was often impressive but inconsistent. Its one-shot performance was typically better across most tasks. And few-shot performance, with 10-50 examples filling the context window, was usually the strongest. This gradient from zero to few-shot suggested that each type of information, task description, format template, and statistical demonstration, provides complementary signal that the model can use.

In[10]:
Code
def create_prompt(examples, test_input, task_description=None):
    """
    Create a few-shot learning prompt.

    Args:
        examples: List of (input, output) tuples for demonstrations
        test_input: The input to classify/complete
        task_description: Optional zero-shot task description

    Returns:
        Formatted prompt string
    """
    prompt_parts = []

    if task_description:
        prompt_parts.append(task_description + "\n")

    for inp, out in examples:
        prompt_parts.append(f"Input: {inp}")
        prompt_parts.append(f"Output: {out}\n")

    prompt_parts.append(f"Input: {test_input}")
    prompt_parts.append("Output:")

    return "\n".join(prompt_parts)


# Zero-shot example
zero_shot = create_prompt(
    examples=[],
    test_input="The food was cold and the service was slow.",
    task_description="Classify the sentiment of the following review as Positive or Negative.",
)

# One-shot example
one_shot = create_prompt(
    examples=[("Great product, highly recommend!", "Positive")],
    test_input="The food was cold and the service was slow.",
)

# Few-shot example
few_shot = create_prompt(
    examples=[
        ("Great product, highly recommend!", "Positive"),
        ("Terrible experience, never again.", "Negative"),
        ("Best purchase I've ever made!", "Positive"),
        ("Broken on arrival, what a waste.", "Negative"),
    ],
    test_input="The food was cold and the service was slow.",
)
Out[11]:
Console
============================================================
ZERO-SHOT PROMPT
============================================================
Classify the sentiment of the following review as Positive or Negative.

Input: The food was cold and the service was slow.
Output:

============================================================
ONE-SHOT PROMPT
============================================================
Input: Great product, highly recommend!
Output: Positive

Input: The food was cold and the service was slow.
Output:

============================================================
FEW-SHOT PROMPT (4 examples)
============================================================
Input: Great product, highly recommend!
Output: Positive

Input: Terrible experience, never again.
Output: Negative

Input: Best purchase I've ever made!
Output: Positive

Input: Broken on arrival, what a waste.
Output: Negative

Input: The food was cold and the service was slow.
Output:

Look at the difference in what each prompt gives the model. The zero-shot prompt contains the complete task description but leaves the model to infer what the expected output format should be. The one-shot prompt does not explain the task but demonstrates the format. The few-shot prompt provides four demonstrations spanning both classes, giving the model both format information and evidence about what "Positive" and "Negative" mean in this context.

The key insight from GPT-3's evaluation was that performance scaled predictably with both model size and number of examples:

Out[12]:
Visualization
Line plot showing performance curves for different model sizes with increasing few-shot examples.
Illustration of how task performance scales with model size and number of in-context examples. Larger models benefit more from few-shot examples, with the 175B model showing dramatic improvement from zero to few-shot. Smaller models plateau quickly, suggesting in-context learning is an emergent capability of scale.

This scaling pattern revealed something important: in-context learning is not simply pattern matching or memorization. It requires sufficient model capacity to represent the task and generalize from examples. Smaller models plateau quickly, while larger models continue to benefit from additional demonstrations. The 175B model shows particularly steep improvement as examples are added, suggesting it can extract and apply statistical regularities from the demonstrations in a way that smaller models cannot. This is consistent with the hypothesis that in-context learning is an emergent property, one that requires a certain threshold of model capacity before it becomes effective.

Evaluating GPT-3's Capabilities

The GPT-3 paper evaluated the model across dozens of NLP benchmarks, from reading comprehension and question answering to commonsense reasoning and word analogy; it also covered arithmetic, translation, plus coding tasks. This breadth was deliberate: the authors wanted to demonstrate that few-shot learning was a general capability rather than a performance trick on a specific benchmark.

Performance varied dramatically by task type, revealing both the potential and limits of in-context learning. Some tasks saw GPT-3 few-shot performance that matched or exceeded fine-tuned state-of-the-art. Others saw GPT-3 underperform a simple majority-class baseline. Understanding this variance is essential for developing accurate intuitions about what large language models can and cannot do.

Strong Performance: Language Understanding

GPT-3 achieved impressive results on tasks requiring broad language understanding:

GPT-3 few-shot performance compared to fine-tuned state-of-the-art on select benchmarks. Few-shot GPT-3 sometimes matches or exceeds specialized models, particularly on open-ended tasks that benefit from broad world knowledge.
TaskDatasetGPT-3 Few-ShotFine-Tuned SOTA
Reading ComprehensionRACE-h46.8%90.0%
Question AnsweringTriviaQA71.2%75.4%
Commonsense ReasoningPIQA82.8%79.4%
Word ScramblingAnagrams66.9%N/A

On TriviaQA, GPT-3's few-shot performance (71.2%) approached the fine-tuned state-of-the-art (75.4%), despite never being explicitly trained on the task. This is remarkable: models fine-tuned on TriviaQA have access to thousands of question-answer pairs during training, have optimized their weights specifically for this task, and still only barely outperform GPT-3 given a handful of examples. On physical commonsense reasoning (PIQA), GPT-3 exceeded the previous best fine-tuned model, suggesting that some commonsense knowledge is better learned from broad pretraining than from narrow task-specific datasets.

The word scrambling result is particularly striking because it suggests the model has internalized something about the structure of English words. Given a scrambled word like "lykaeh" and asked to find the unscrambled form, GPT-3 could often correctly identify "weakly" or similar targets. This is not a task most humans find trivial, and it was not explicitly part of any training objective. The model appears to have developed implicit knowledge of English phonology and orthography as a byproduct of predicting text.

Emergent Capabilities

Perhaps most striking were capabilities that emerged without explicit training:

  • Arithmetic: GPT-3 could perform multi-digit addition and subtraction with reasonable accuracy, despite never being trained on a math curriculum
  • Code generation: It could write functional code snippets in Python and JavaScript from natural language descriptions
  • Translation: Few-shot GPT-3 matched or exceeded supervised baselines on some language pairs, particularly for translating into English
  • News article generation: Given headlines, it produced articles that human evaluators struggled to distinguish from real news

The arithmetic capability deserves special attention because it cannot be explained by memorization. If GPT-3 had simply memorized a table of addition facts, it would get familiar problems right and novel problems wrong. Instead, performance degraded gradually as problem complexity increased, the hallmark of an imperfect learned procedure rather than a lookup table.

Out[13]:
Visualization
Grouped bar chart showing arithmetic accuracy for 2, 3, and 4 digit operations across addition and subtraction.
GPT-3 arithmetic accuracy by operation type and number of digits. Near-perfect performance on 2-digit problems degrades gracefully as digit count increases, suggesting the model learned genuine algorithmic procedures rather than memorizing specific answers. The degradation pattern mirrors what you would expect from imperfect carry-propagation in mental arithmetic.

The arithmetic results are particularly interesting because they suggest the model learned algorithmic procedures rather than memorizing specific calculations. Performance degrades gracefully with more digits, the pattern you'd expect from imperfect procedure learning rather than lookup table failure. When a lookup table is exhausted, performance drops to zero. When a procedure is imperfect, performance drops gradually as the problem complexity exceeds the model's ability to apply the procedure reliably. GPT-3 shows the latter pattern, implying that whatever computation it performs for arithmetic has the character of a learned algorithm.

Weak Performance: Structured Tasks

GPT-3 struggled on tasks requiring precise, structured outputs or multi-step reasoning:

  • Reading comprehension with multiple choice: On RACE-h, GPT-3 achieved only 46.8% compared to 90% for fine-tuned models. The gap here is enormous and instructive. RACE-h consists of multiple-choice questions from Chinese high school reading comprehension tests. While GPT-3 can understand the English text, the task requires selecting from four options where three are plausible distractors. The model often fails to eliminate plausible-but-wrong options with the precision needed for reliable performance.
  • Natural language inference: Performance on ANLI (adversarial NLI) remained below random chance for some splits. ANLI was specifically designed to be adversarially hard for models that rely on surface-level patterns. GPT-3's failure here reveals that its "reasoning" is often shallow, relying on word overlap and semantic similarity rather than valid logical inference.
  • Math word problems: Multi-step reasoning with intermediate calculations proved challenging. Problems requiring the model to maintain intermediate results across multiple calculation steps fell apart as the number of steps increased, because the model had no mechanism for reliable working memory beyond what fit in its context.
  • Fact verification: Distinguishing true from false claims required external knowledge verification the model could not perform.

These weaknesses pointed to fundamental limitations: GPT-3 excels at pattern completion and surface-level understanding but struggles with tasks requiring careful logical reasoning or precise knowledge retrieval. The model can approximate reasoning in familiar domains where training data provides dense coverage, but it lacks the systematic, step-by-step reasoning that humans use for structured problems.

Historical Context: The Benchmark Race

GPT-3 arrived during a period when NLP benchmarks like GLUE and SuperGLUE were central to the field. These benchmarks tracked progress on a standardized set of tasks, and new models competed to achieve state-of-the-art scores. GPT-3's release complicated this picture: on some benchmarks it was competitive with fine-tuned models despite using no task-specific training, while on others it lagged badly. This inconsistency prompted a broader conversation about whether benchmarks were measuring the right things and whether fine-tuning accuracy was the appropriate metric for evaluating general-purpose language models.

How In-Context Learning Works

The mechanism behind in-context learning remained mysterious. How can a model "learn" a task without updating its weights? Several hypotheses emerged from subsequent research, each offering a different lens on what is happening during the forward pass when a model encounters few-shot demonstrations.

The mechanism matters in practice. If in-context learning works through task recognition, then the examples mainly activate pre-stored knowledge, and their format matters more than their quality. If it works through implicit fine-tuning, then the examples need to provide correct statistical signal about the mapping from inputs to outputs. If it works through Bayesian inference, then more examples are always helpful because they sharpen the posterior. Different mechanisms have different implications for how to engineer effective prompts.

Hypothesis 1: Task Recognition

One theory suggests that pretraining exposes the model to many implicit task formats. When given few-shot examples, GPT-3 recognizes the pattern from pretraining and activates the appropriate "circuit" for that task. The examples don't teach the model something new; they help it identify which of its existing capabilities to apply.

The evidence for this view comes from a striking experimental result: in several studies, randomizing the labels in few-shot examples (assigning "Positive" to negative reviews and vice versa) still improved performance over zero-shot. If the examples were teaching the input-label relation, scrambled labels should hurt performance significantly. That they don't suggests the main benefit of examples is format information, not label information. The model sees the input/output format and activates the corresponding task processing circuit, regardless of whether the specific labels are correct.

In[14]:
Code
def analyze_task_format(examples):
    """
    Analyze the format of few-shot examples to understand task structure.

    This illustrates what patterns the model might recognize.
    """
    analysis = {
        "n_examples": len(examples),
        "avg_input_length": np.mean([len(inp.split()) for inp, _ in examples]),
        "avg_output_length": np.mean([len(out.split()) for _, out in examples]),
        "output_vocabulary": set(out for _, out in examples),
    }

    # Detect if this looks like classification
    if len(analysis["output_vocabulary"]) <= 5:
        analysis["likely_task"] = "classification"
    elif analysis["avg_output_length"] > 10:
        analysis["likely_task"] = "generation"
    else:
        analysis["likely_task"] = "extraction"

    return analysis


# Analyze our sentiment examples
sentiment_examples = [
    ("Great product, highly recommend!", "Positive"),
    ("Terrible experience, never again.", "Negative"),
    ("Best purchase I've ever made!", "Positive"),
    ("Broken on arrival, what a waste.", "Negative"),
]

analysis = analyze_task_format(sentiment_examples)
Out[15]:
Console
Task Format Analysis
=============================================
Number of examples: 4
Average input length: 4.8 words
Average output length: 1.0 words
Unique outputs: {'Positive', 'Negative'}
Likely task type: classification

The analysis reveals structural patterns that a model might use to identify the task. With only two unique outputs (Positive, Negative) and short single-word responses, the format clearly signals a binary classification task. The model may recognize this pattern from similar structures encountered during pretraining, such as review datasets or labeled examples in web text. Think of it like a key fitting a lock: the format of the prompt matches patterns the model has seen before, triggering the relevant internal computation.

The limitation of this hypothesis is that it cannot explain all in-context learning behavior. For tasks absent from the pretraining data, the task recognition account predicts that in-context learning should fail entirely. But GPT-3 shows some ability to adapt to novel input-output formats, suggesting that task recognition alone is insufficient.

Hypothesis 2: Implicit Fine-Tuning

Another perspective views the forward pass through GPT-3's 96 layers as an implicit optimization process. The key insight is that transformer layers can implement gradient-descent-like updates on the representations. Consider a single attention layer operating on demonstrations {(xi,yi)}\{(x_i, y_i)\} followed by a test input xtestx_{\text{test}}. The attention mechanism computes:

htest=htest(0)+∑i=1kαi⋅vih_{\text{test}} = h_{\text{test}}^{(0)} + \sum_{i=1}^{k} \alpha_i \cdot v_i

where:

  • htest(0)h_{\text{test}}^{(0)}: the initial representation of the test input before attention
  • αi\alpha_i: the attention weight from the test input to demonstration ii
  • viv_i: the value vector derived from demonstration ii
  • kk: the number of demonstrations in context

This weighted sum resembles a gradient update: the model adjusts its representation of the test input based on the demonstrations. The attention weights αi\alpha_i function like step sizes, and the value vectors viv_i function like gradient directions. With 96 layers, the model has large "depth" to iteratively refine this adaptation. Research by Akyürek et al. (2022) showed that linear attention transformers can implement gradient descent in-context. This provides a theoretical foundation for this hypothesis. Further work found that in-context learning and explicit fine-tuning produce similar representational changes in the model's hidden states, even though one updates weights and the other doesn't.

The implicit fine-tuning hypothesis has an important implication: the quality of examples matters, not just their format. If the forward pass is implementing something like gradient descent on the demonstrations, then wrong labels should hurt, correct labels should help, and more examples should provide better gradient estimates, all consistent with observed behavior for most tasks.

Hypothesis 3: Bayesian Inference

A third hypothesis frames in-context learning as Bayesian inference over tasks. The model implicitly maintains a prior distribution P(task)P(\text{task}) over possible tasks from pretraining. Each demonstration (xi,yi)(x_i, y_i) updates this posterior according to Bayes' rule:

P(task∣x1,y1,…,xk,yk)∝P(y1,…,yk∣x1,…,xk,task)⋅P(task)P(\text{task} \mid x_1, y_1, \ldots, x_k, y_k) \propto P(y_1, \ldots, y_k \mid x_1, \ldots, x_k, \text{task}) \cdot P(\text{task})

where:

  • P(task)P(\text{task}): the prior probability of each task, learned during pretraining from exposure to diverse text formats
  • P(yi∣xi,task)P(y_i \mid x_i, \text{task}): the likelihood of output yiy_i given input xix_i under a specific task interpretation
  • P(task∣…)P(\text{task} \mid \ldots): the posterior probability of each task after observing demonstrations

By the time the model reaches the test input, it has narrowed down the task distribution sufficiently to make a confident prediction. This view explains why more demonstrations help: each one provides additional evidence that sharpens the posterior. It also explains why demonstrations with correct labels help more than scrambled ones for complex tasks: the likelihood term P(yi∣xi,task)P(y_i \mid x_i, \text{task}) is only informative when the labels are accurate.

The Bayesian view predicts that performance should improve monotonically with more examples, which is roughly consistent with observations. It also predicts that unusual or contradictory demonstrations should hurt performance by introducing conflicting evidence, which is also consistent with empirical findings about prompt quality.

In practice, the true mechanism is likely a blend of all three hypotheses. The model uses format cues to identify the task type (task recognition), adjusts its representations through attention to align with the demonstration distribution (implicit fine-tuning), and these adjustments are consistent with what a Bayesian reasoner would do when updating over task hypotheses (Bayesian inference). The three accounts describe different levels of analysis: the computational level (what is being computed), the algorithmic level (how it is computed), and the representational level (what data structures are used).

Out[16]:
Visualization
Diagram box for task recognition hypothesis.
Task Recognition: Demonstrations activate existing circuits from pretraining.
Diagram box for implicit fine-tuning hypothesis.
Implicit Fine-Tuning: Deep forward passes behave like gradient descent.
Diagram box for Bayesian inference hypothesis.
Bayesian Inference: Demonstrations update posterior over tasks.

Prompt Engineering Principles

The success of in-context learning spawned a new discipline: prompt engineering. Researchers discovered that how you format prompts dramatically affects performance. This was simultaneously exciting and troubling: exciting because it meant that performance could be improved without any model changes, troubling because it made deployment unreliable. A prompt that worked beautifully on test cases might fail unexpectedly on edge cases that differed in subtle ways.

The practical consequence was a new kind of expertise. Organizations deploying GPT-3 needed people who understood how to write effective prompts, test them systematically, and iterate when they failed. Prompt engineering became a recognized skill, and eventually a job title. The discipline drew on insights from cognitive science (how to communicate instructions clearly), information retrieval (how to retrieve relevant context), and software engineering (how to test and validate behavior systematically).

Format Matters

Small changes in prompt format can yield large performance differences:

In[17]:
Code
# Different prompt formats for the same task
formats = {
    "simple": """Review: {review}
Sentiment: """,
    "instructional": """Task: Classify the sentiment of the review as Positive or Negative.

Review: {review}
Sentiment: """,
    "structured": """### Sentiment Classification

**Review:** {review}

**Sentiment:**""",
    "conversational": """You are a sentiment analysis expert. Read the following review and determine if it expresses a Positive or Negative sentiment.

The review is: "{review}"

Your classification:""",
}

test_review = (
    "The battery life is amazing but the camera quality is disappointing."
)
Out[18]:
Console
Prompt Format Variations
============================================================

--- SIMPLE FORMAT ---
Review: The battery life is amazing but the camera quality is disappointing.
Sentiment: 

--- INSTRUCTIONAL FORMAT ---
Task: Classify the sentiment of the review as Positive or Negative.

Review: The battery life is amazing but the camera quality is disappointing.
Sentiment: 

--- STRUCTURED FORMAT ---
### Sentiment Classification

**Review:** The battery life is amazing but the camera quality is disappointing.

**Sentiment:**

--- CONVERSATIONAL FORMAT ---
You are a sentiment analysis expert. Read the following review and determine if it expresses a Positive or Negative sentiment.

The review is: "The battery life is amazing but the camera quality is disappointing."

Your classification:

Research showed that the conversational format often performs best for GPT-3, likely because its training data contains many examples of conversational task-solving. The optimal format varies by model and task, requiring empirical experimentation. The simple format is terse and may not trigger the right pretraining associations. The instructional format is explicit but somewhat mechanical. The structured format uses markdown-style formatting that may or may not match the model's internal representations. The conversational format frames the task in the way you might address a human expert, which often aligns well with the task-solving patterns the model has internalized from instructional web content.

Notice that the conversational format does something the others don't: it establishes a role for the model ("you are a sentiment analysis expert") before asking it to perform. This role assignment can significantly influence the style and accuracy of responses. When the model is told it is an expert, it tends to produce more confident, technically precise outputs. When it is not given a role, outputs can be more hedged and generic. Role assignment is one of the most reliable prompt engineering techniques.

Example Selection and Ordering

The choice and order of few-shot examples strongly affects performance:

  • Diversity: Examples should cover the range of possible inputs and outputs. If you are doing sentiment classification and all your examples are strongly positive or strongly negative, the model may struggle with neutral or mixed inputs.
  • Similarity: Examples semantically similar to the test input often help more. If your test case is a hotel review, hotel review examples will generally outperform product review examples, even for the same classification task.
  • Recency: Examples closer to the end of the prompt have stronger influence due to recency effects in attention. The model attends more strongly to nearby context than to distant context.
  • Balance: For classification, examples should be balanced across classes. An imbalanced prompt where 3 out of 4 examples are Positive will bias the model toward predicting Positive, regardless of the test input.
Out[19]:
Visualization
Horizontal bar chart showing relative influence increasing from the oldest to the newest in-context example.
Relative influence of in-context examples by position in the prompt. Examples at the end (most recent, closest to the test input) exert stronger influence on the model's prediction due to recency effects in causal self-attention. Practical prompt engineering should place the most representative or informative examples closest to the test input.

In practice, this suggests a concrete strategy: if you have one example that is much more similar to your expected test inputs than the others, place it last. The recency effect means the model will weight it most heavily, and the semantic similarity between the last example and the test input will further amplify its influence. Conversely, if you have an outlier example that you want to include for coverage but don't want to dominate predictions, place it early in the prompt.

The Label Matters (Sometimes)

A surprising finding was that the actual labels in few-shot examples don't always matter. In some experiments, randomly assigning labels (positive reviews labeled "Negative" and vice versa) still improved performance over zero-shot, suggesting the model was learning format rather than task semantics.

However, for more complex tasks, correct labels improve performance considerably. The model can use both the format pattern and the input-output mapping. A useful rule of thumb: for simple classification tasks where the labels are self-explanatory (Positive/Negative, True/False), format information dominates. For complex tasks where the input-output mapping is not obvious, correct labels are critical because they provide the only signal about what the task requires.

This finding has a practical implication: you can sometimes bootstrap a few-shot prompt without labeled data, using the correct format with placeholder outputs to establish the input-output structure, then relying on the model's pretraining to supply the semantics. This is particularly useful when labeled data is expensive or unavailable, though it requires careful validation to ensure the model's pretraining assumptions match your intended task.

Limitations and Concerns

GPT-3's release sparked both excitement and concern. While the model demonstrated impressive capabilities, it also exposed fundamental challenges that persist in large language models today. These limitations fall into several categories: reliability of outputs, sensitivity to inputs, societal impacts, and practical constraints. Understanding these limitations is not pessimism; it is necessary context for using large language models responsibly and for understanding what problems remain to be solved.

Factual Errors and Hallucinations

GPT-3 confidently produces plausible-sounding but incorrect information. It might cite nonexistent studies, invent historical events, or provide wrong answers to factual questions. The model has no mechanism to verify claims against external sources or acknowledge uncertainty in a calibrated way.

This limitation is particularly dangerous because GPT-3's outputs are fluent and authoritative. Users may accept incorrect information because it sounds convincing. The model's confidence is orthogonal to its accuracy. This decoupling arises from the training objective: the model is trained to produce text that looks like text a human expert would write. Humans expressing views with confidence write fluently and assertively. The model has learned to mimic this style even when it lacks the knowledge to back it up.

The deeper problem is that GPT-3 has no grounded representation of the world beyond its training data. It has seen statements about the world but has no way to check whether those statements are true. When asked about a topic where its training data contains conflicting or sparse information, it fills in gaps with plausible-sounding constructions that may be entirely fabricated. This phenomenon became known as "hallucination," and it remains one of the central unsolved problems in large language model deployment.

Sensitivity to Prompts

Performance is brittle with respect to prompt phrasing. Changing a single word or reordering examples can shift accuracy by 10-20 percentage points. This makes reliable deployment challenging: a prompt that works well on test examples might fail unexpectedly on edge cases.

The sensitivity to prompt phrasing reveals that the model's "understanding" is often shallow. If the model represented a task consistently, its performance should not depend on whether you say "classify" or "categorize," use bold formatting or plain text, or introduce the task in the first or third person. The effect of these surface variations suggests the model is often matching the prompt's format to patterns in pretraining data rather than constructing a stable representation of the task semantics.

In practice, this means that prompt engineering is never done. Every deployed application needs ongoing monitoring and prompt refinement as the distribution of inputs changes or as new failure modes are discovered. This operational overhead was one of the factors that initially limited GPT-3's adoption in production systems.

Bias and Fairness

GPT-3 amplifies biases present in its training data. It associates certain occupations with genders, exhibits racial stereotypes, and can generate toxic content when prompted. The model learned these patterns from internet text and has no mechanism to distinguish harmful biases from useful patterns.

The scale of GPT-3's training data means it has absorbed biases from an enormous range of sources: historical texts with discriminatory assumptions, contemporary web content with implicit stereotypes, and fringe content with explicit prejudice. These biases manifest in subtle ways: the model may consistently associate "doctor" with male pronouns, describe certain ethnicities with stereotype-laden language, or produce systematically different quality outputs for queries about different demographic groups.

What makes this particularly challenging is that bias is not a simple property that can be measured or fixed by a single intervention. It is distributed across hundreds of billions of parameters, interwoven with useful patterns, and manifests differently across different contexts. Debiasing a language model without degrading its overall capabilities remains an active research area. GPT-3's release accelerated this research by providing a highly capable model whose biases could be systematically studied.

Out[20]:
Visualization
Diagram box for hallucination limitation.
Hallucinations: Generates plausible but factually incorrect content with high confidence.
Diagram box for prompt sensitivity limitation.
Prompt Sensitivity: Small prompt changes cause large performance variations.
Diagram box for bias and fairness limitation.
Bias and Fairness: Reflects and amplifies biases from training data.
Diagram box for reasoning limits.
Reasoning Limits: Struggles with multi-step logical reasoning.

Context Length Constraints

With a 2,048 token context, GPT-3 cannot process long documents or maintain extended conversations. Few-shot learning is limited by how many examples fit in the context window. For tasks requiring synthesis across multiple documents, GPT-3 must work with truncated or summarized inputs.

This constraint is more significant than it might appear. Many real-world tasks require reasoning across documents: legal analysis, scientific literature review, multi-document summarization, and long-form conversation. At 2,048 tokens, GPT-3 can handle about four pages of text. This is sufficient for many question-answering tasks but inadequate for most professional-grade document analysis. The context length limitation was one of the key motivators for subsequent research into efficient attention mechanisms and longer-context models.

Environmental and Access Concerns

Training GPT-3 consumed enormous computational resources with associated carbon emissions. Access was restricted to an API controlled by OpenAI, raising concerns about democratization of AI research. Only well-funded organizations could experiment with the full model, creating imbalances in who could study and critique these systems.

The access restriction had deep consequences for the research community. Traditional NLP research involved open-source models and public datasets that any researcher could use. GPT-3 was a closed system: you could query it through an API, but you could not inspect its weights, run ablation studies, or probe its internal representations. This made GPT-3 difficult to study scientifically and impossible to replicate with limited compute. It also concentrated the ability to identify and publicize safety issues in the hands of those with API access, which was not democratically distributed.

Impact on the Field

GPT-3's impact extended far beyond its benchmark scores. It changed how researchers and practitioners think about NLP, AI development, and the relationship between scale and intelligence. To understand this impact fully, it helps to distinguish between immediate effects on research practice, medium-term effects on industry, and long-term effects on how we think about language models.

The immediate effect was a surge of interest in prompting. Before GPT-3, prompting was a niche technique used primarily to condition generative models. After GPT-3, it became a central paradigm. Hundreds of papers studied how to engineer better prompts, what makes some prompts more effective than others, and how prompting relates to other forms of model adaptation. This research explosion created new subfields: prompt engineering, instruction tuning, and chain-of-thought reasoning all trace their origins to GPT-3's demonstration that prompting could be a general-purpose interface.

Prompting as a New Paradigm

Before GPT-3, the standard approach was to fine-tune pretrained models for specific tasks. GPT-3 demonstrated that sufficiently large models could be steered through prompting alone. This shifted research attention toward:

  • Prompt engineering: Systematic study of how prompt design affects performance. Researchers developed techniques like chain-of-thought prompting (showing the model step-by-step reasoning), zero-shot chain-of-thought (using the phrase "let's think step by step"), and self-consistency (sampling multiple reasoning paths and taking the majority answer).
  • Instruction tuning: Training models to follow natural language instructions. InstructGPT, FLAN, and T0 all explored this direction, training models on diverse collections of instructional tasks to improve their ability to follow natural language descriptions.
  • Chain-of-thought prompting: Encouraging models to show reasoning steps before giving a final answer. This technique, introduced by Wei et al. (2022), showed that providing examples of step-by-step reasoning in the prompt could dramatically improve performance on multi-step reasoning tasks, effectively bootstrapping the kind of explicit intermediate reasoning that GPT-3 struggled to do unaided.

These techniques all share a common insight: the way you communicate with a large language model matters enormously, and there is substantial room to improve performance through better communication without any changes to the model's weights.

Scaling Laws

GPT-3 contributed to the formalization of scaling laws: mathematical relationships between model size, training compute, dataset size, and performance. The key insight is that test loss LL follows a power-law relationship with each scaling factor.

To understand why power laws matter here, consider the alternative. If performance improved logarithmically with scale, you would quickly reach diminishing returns and scaling up would become pointless. If it improved linearly, each doubling of resources would yield the same absolute improvement. Power laws sit between these extremes: performance improves smoothly and predictably, but the improvement per additional resource decreases slowly, making continued scaling worthwhile.

The relationship between model size and test loss takes the form:

L(N)=(NcN)αNL(N) = \left(\frac{N_c}{N}\right)^{\alpha_N}

where:

  • L(N)L(N): the test loss as a function of model parameters
  • NN: the number of model parameters
  • NcN_c: a constant that depends on the dataset and architecture, capturing the baseline difficulty of the prediction task
  • αN≈0.076\alpha_N \approx 0.076: the scaling exponent for parameters (empirically determined from many training runs)

Similar relationships hold for dataset size DD and training compute CC:

L(D)=(DcD)αD,L(C)=(CcC)αCL(D) = \left(\frac{D_c}{D}\right)^{\alpha_D}, \quad L(C) = \left(\frac{C_c}{C}\right)^{\alpha_C}

where αD≈0.095\alpha_D \approx 0.095 for dataset tokens and αC≈0.050\alpha_C \approx 0.050 for compute. These power laws mean that each 10x increase in resources yields a predictable, if diminishing, improvement in loss.

The practical implication is that these power laws let you estimate in advance how much compute and data are needed for a target level of performance. You can also predict the relative value of investing in more parameters versus more data. Kaplan et al. (2020) used scaling laws to argue that models of GPT-3's era were significantly undertrained: given a fixed compute budget, you would do better training a smaller model for longer than training a large model briefly. This observation later motivated the Chinchilla work, which showed that truly compute-optimal training requires much more data per parameter than GPT-3 used.

Out[21]:
Visualization
Log-log plot showing power-law relationship between parameters and test loss.
Parameters scaling: Test loss decreases as a power law with model size. Each order of magnitude in parameters yields a consistent fractional improvement.
Log-log plot showing power-law relationship between data size and test loss.
Data scaling: Test loss decreases with dataset size (tokens). More data helps, but with diminishing returns following the same power-law pattern.
Log-log plot showing power-law relationship between compute and test loss.
Compute scaling: Test loss decreases with training FLOPs. A fixed compute budget can be allocated to more parameters or more training steps, and scaling laws describe the optimal tradeoff.

The scaling law perspective suggested that many apparent limitations might be overcome by simply training larger models on more data. This hypothesis, controversial at the time, proved partially correct with subsequent models. Later work showed a more complicated pattern: some capabilities improve smoothly with scale, while others improve suddenly at specific thresholds. The debate between "smooth scaling" and "emergent capabilities" continues to shape research in this area.

Foundation Models

GPT-3 exemplified the "foundation model" concept: a single large model pretrained on diverse data that can be adapted to many downstream tasks. Rather than training separate models for translation and summarization, plus question answering, a single foundation model handles all tasks through different prompts or lightweight adaptation.

This paradigm consolidated AI development around a few very large models, shifting the field toward centralized training and distributed deployment. The economics of this shift are significant. Training a 175B parameter model requires massive infrastructure and capital, limiting who can train frontier models. But once trained, the model can be deployed cheaply at scale, making its capabilities accessible through an API. This separation of training cost from deployment cost created a new industry structure: a small number of organizations train foundation models, and a much larger number of organizations and developers build applications on top of them.

The foundation model paradigm also changed what "AI research" meant for most practitioners. Instead of training models from scratch, practitioners learned to work with pretrained models, adapting them through prompting, fine-tuning, or other lightweight techniques. The skills required shifted from deep learning architecture design toward prompt engineering and evaluation, along with application development. This shift lowered the barrier to entry for building AI applications. This contributes to the rapid growth of AI-powered products that followed GPT-3's release.

Summary

GPT-3 changed how language models were used. Its scale produced new capabilities: after training 175 billion parameters on hundreds of billions of tokens, the model could adapt to tasks from only a handful of examples in the prompt.

The core technical findings were:

  • In-context learning emerges at scale: Smaller models show minimal benefit from few-shot examples, while GPT-3's performance improved dramatically with each additional demonstration. This suggests a threshold phenomenon where sufficient model capacity is required before in-context adaptation becomes effective.
  • Performance varies by task type: GPT-3 excelled at tasks requiring broad language understanding, commonsense reasoning, and creative generation, while struggling with structured multi-step reasoning, adversarial inference, and tasks requiring precise factual recall.
  • Three hypotheses explain in-context learning: Task recognition (the model identifies the format and activates existing circuits), implicit fine-tuning (the forward pass implements gradient-descent-like updates), and Bayesian inference (demonstrations sharpen a posterior over tasks) each explain different aspects of the observed behavior.
  • Scaling laws predict improvement: Test loss follows power-law relationships with model size, data size, and compute, enabling principled resource allocation for training future models.

The key practical lessons:

  • Scale enables emergence: In-context learning appeared as model size increased, suggesting that some capabilities require a threshold of capacity before manifesting
  • Few-shot learning works: Providing examples in the prompt can match or exceed fine-tuned models for certain tasks, without any parameter updates
  • Prompt design matters: Prompt format and ordering, together with content, strongly affect performance, spawning the field of prompt engineering
  • Limitations persist: Hallucinations, bias, reasoning failures, and context constraints remain challenges that scale alone doesn't solve
  • Paradigm shift: GPT-3 accelerated the move from task-specific fine-tuning toward general-purpose foundation models

The model's release catalyzed an industry-wide race to scale, leading to GPT-4 and Claude, alongside PaLM and other models that pushed beyond GPT-3's capabilities. The research directions it opened, prompt engineering, instruction tuning, chain-of-thought reasoning, and scaling laws, continue to shape the field. But GPT-3's core insight endures: with sufficient scale, language models develop surprising and useful behaviors that weren't explicitly trained, and understanding why requires rethinking what it means for a model to "know" something or "learn" a task.

Key Parameters

When working with GPT-3 and similar large language models, several parameters affect performance and behavior:

  • temperature: Controls randomness in token sampling. Values range from 0.0 (deterministic, always choosing the highest probability token) to 2.0 (highly random). Common values are 0.7-1.0 for creative tasks and 0.0-0.3 for factual tasks. Lower temperatures produce more focused, predictable outputs; higher temperatures increase diversity but may reduce coherence.

  • max_tokens: The maximum number of tokens to generate in the response. Set this based on expected output length to control costs and latency. GPT-3 supports up to 2,048 tokens total (prompt + completion combined), so longer prompts leave less room for generation.

  • top_p (nucleus sampling): An alternative to temperature that samples from the smallest set of tokens whose cumulative probability exceeds the threshold. A value of 0.9 means the model considers tokens until their probabilities sum to 90%. Generally, adjust either temperature or top_p, but not both simultaneously.

  • frequency_penalty: Reduces repetition by penalizing tokens based on how often they appear in the generated text so far. Values range from 0.0 to 2.0. Higher values discourage the model from repeating the same phrases, useful for open-ended generation.

  • presence_penalty: Penalizes tokens based on whether they have appeared at all, regardless of frequency. This encourages the model to introduce new topics. Values range from 0.0 to 2.0. Useful when you want diverse, wide-ranging outputs.

  • n_examples (few-shot count): The number of demonstrations to include in the prompt. More examples generally improve performance but consume context space. Empirically, 4-8 examples often provide good results while leaving room for the test input and response. Beyond 20-30 examples, returns typically diminish.

  • stop sequences: Tokens or strings that signal generation should stop. For few-shot prompts, include delimiters like "\n\n" or "Input:" to prevent the model from generating additional examples. Proper stop sequences ensure clean, usable outputs.

In-Context Learning: A Deeper Look

Now that we have covered the mechanics of how few-shot prompting works, it is worth going deeper into what in-context learning reveals about the nature of large language models. The discovery of in-context learning raised questions that go beyond engineering: what is happening computationally when a model reads a few examples and then generalizes to new inputs? How does this relate to what we call "learning" in more traditional settings? And what does it imply about the internal organization of knowledge in a large language model?

Meta-Learning and the Pretraining Objective

One of the most illuminating frameworks for understanding in-context learning comes from meta-learning, a branch of machine learning concerned with algorithms that can learn to learn. In meta-learning, a model is trained across many different tasks and learns to extract transferable strategies that help it adapt quickly to new tasks. The goal is not to learn any single task well but to develop the capacity for fast task acquisition.

GPT-3 was not explicitly trained with a meta-learning objective. Its training was simply next-token prediction on a diverse corpus of text. Yet the diversity of that corpus may have implicitly induced a form of meta-learning. When GPT-3 trained on a book about cooking, a paper on mathematics, a tutorial on programming, and a forum discussion about personal finance all at once, it was forced to maintain representations that were useful across all of these domains simultaneously. The attention mechanism, which can flexibly attend to any part of the context, is well-suited to this kind of flexible, context-dependent processing.

The key insight from this perspective is that in-context learning is not a special mode that activates when you provide demonstrations. It is the same process the model uses during ordinary language generation, applied to a context that happens to contain input-output examples. When GPT-3 reads the text "Translate English to French: sea otter => loutre de mer, plush girafe => girafe peluche, cheese => fromage", it is processing this as ordinary text, continuing the pattern it has learned to recognize across billions of documents. The few-shot format looks like many things the model has seen during pretraining: flashcard decks, translation dictionaries, worked examples in textbooks, and question-answer pairs. The model recognizes the pattern and applies the corresponding text-generation strategy.

This view makes a testable prediction: the more a task's format resembles text formats common in pretraining, the better in-context learning will work for that task. And this is broadly what researchers found. Translation tasks, which appear extensively in multilingual web text, work well. Mathematical proof tasks, which require formats and reasoning patterns rarely seen in web text, work poorly. Code completion tasks, which appear extensively in GitHub repositories that were likely included in or similar to GPT-3's training data, work better than you might expect for a language model.

What the Demonstrations Provide

A critical experiment in understanding in-context learning was conducted by Min et al. (2022), who systematically studied which information in few-shot demonstrations affects performance. They tested four variants of demonstrations:

  1. Standard demonstrations: Correct input-output pairs in the correct format.
  2. Random labels: Correct inputs paired with randomly assigned labels from the label set.
  3. Random inputs: Inputs replaced with random words or sentences, paired with correct labels.
  4. No demonstrations: Zero-shot baseline.

The surprising finding was that random labels (variant 2) performed nearly as well as standard demonstrations for many tasks, while providing much better performance than zero-shot. This strongly suggested that the labels themselves were not the primary source of benefit. Instead, the main contribution of demonstrations appeared to be: specifying the input distribution (showing what kinds of inputs to expect), specifying the output distribution (showing what kinds of outputs are expected), and specifying the input-output format (showing how inputs and outputs are structured relative to each other).

The semantics of the label mapping (which input class corresponds to which output label) contributed much less than expected. This result was initially controversial because it seemed to undermine the claim that in-context learning was learning the task from the examples. If the labels don't matter, the model isn't learning the input-to-output mapping from the demonstrations; it is doing something else.

Subsequent work refined this picture. While random labels hurt less than expected on simple tasks, they hurt more on complex tasks, especially those involving multiple-step reasoning or subtle semantic distinctions. The conclusion is that demonstrations serve multiple functions simultaneously, and different tasks rely on different aspects. Simple classification tasks benefit primarily from format information. Complex reasoning tasks benefit additionally from the actual input-output mappings in demonstrations.

The Attention Mechanism as the Implementation Vehicle

However in-context learning works at a high level, it must be implemented through the model's actual computational primitives: the attention mechanism and the feed-forward networks. Understanding how attention supports in-context learning gives insight into both the capabilities and the limits of this approach.

Consider what happens when GPT-3 processes a few-shot prompt for sentiment classification. Early layers of the transformer build representations of individual tokens, incorporating local context. Middle layers begin to form representations of longer phrases and sentences. Later layers build higher-level semantic representations that encode the meaning and function of entire passages.

As the model processes the demonstration pairs, the attention mechanism creates associations between input texts and their corresponding labels. When the test input is processed, later-layer attention heads can attend back to demonstration inputs that are semantically similar to the test input. Through the value vectors, the associated labels influence the model's representation of the test input. The more demonstrations, and the more semantically relevant they are to the test input, the stronger these associations become.

Notice that this is not a perfect implementation of the task. The model's ability to find the right demonstrations to attend to depends on the quality of its representations, which depend on pretraining. If the test input is dissimilar to all demonstrations in the model's internal representation space, the attention mechanism cannot bridge the gap. This explains why semantically diverse demonstrations help and why demonstrations similar to the test input are particularly valuable.

The feed-forward networks play a complementary role. Research by Geva et al. (2021) found that feed-forward layers in transformers function somewhat like a key-value memory, where each neuron can be seen as storing a pattern and its associated knowledge. During pretraining, these memories are filled with associations learned from the training corpus. During in-context learning, the representations built by the attention mechanism serve as queries that retrieve relevant memories from the FFN. The final prediction combines the in-context signal from attention with the pretraining knowledge stored in the FFN weights.

This combination is what makes large models so capable in the few-shot setting. A small model can learn in-context signal, but it has limited stored knowledge to combine it with. A large model has both the capacity to learn from in-context demonstrations and a vast store of pretrained knowledge to draw on when interpreting and extending those demonstrations.

The Chinchilla Insight: Was GPT-3 Undertrained?

GPT-3 was released in 2020, and at the time, training a 175B parameter model on 300 billion tokens seemed extraordinarily ambitious. But within two years, research showed that this configuration was suboptimal from a compute-efficiency standpoint. The Chinchilla paper (Hoffmann et al., 2022) presented a systematic study of scaling laws with more careful experimental controls than previous work, and arrived at a striking conclusion: for a given compute budget, you should train a much smaller model on much more data than GPT-3 used.

The key finding was that GPT-3's 175B parameters were paired with only 300 billion training tokens, roughly 1.7 tokens per parameter. The Chinchilla scaling laws suggested that optimal compute-efficient training requires approximately 20 tokens per parameter. Under this prescription, training a 175B parameter model optimally would require about 3.5 trillion tokens, over ten times what GPT-3 used. Equivalently, for the same compute budget that GPT-3 consumed, a model with around 70B parameters trained on 1.4 trillion tokens would achieve better performance than GPT-3 on downstream tasks.

This result had enormous practical implications. It suggested that the scaling race driven by GPT-3's success was partially misguided: organizations racing to train ever-larger models were getting worse performance per FLOP than they could achieve with smaller, better-trained models. The subsequent generation of models, including Meta's LLaMA family and Google's Gemma family, were designed around Chinchilla-optimal or near-optimal training configurations, training smaller models for much longer on more data.

For GPT-3's in-context learning capability specifically, the Chinchilla finding raises an interesting question: would a Chinchilla-optimal model of the same compute budget show stronger or weaker in-context learning? The few-shot learning capability observed in GPT-3 might partly be a function of the model's size rather than its training quality. A smaller model trained on more data might have better world knowledge and more reliable language understanding, but might require a larger model size before in-context learning emerges. The relationship between model size, training data quality, and in-context learning capability is not fully resolved and remains an active area of research.

What the Chinchilla work does establish is that the scaling laws formalized with the help of GPT-3's results are practically useful: they enable rational planning of training runs rather than ad-hoc scaling. The ability to predict in advance what level of performance a given compute budget can achieve transformed how large language model training was approached, turning it from a largely empirical process into something more like applied engineering.

The Role of the Pretraining Corpus in Shaping Capabilities

One of the most important and underappreciated aspects of GPT-3's success is how much the composition of its training corpus shaped its capabilities. The model did not develop its impressive zero-shot and few-shot abilities in a vacuum: it developed them because of exposure to specific kinds of text that implicitly taught the skills needed for in-context learning.

Consider the capability for arithmetic. GPT-3 was not trained on a math curriculum, but the internet contains an enormous amount of arithmetic-adjacent content: programming tutorials that compute intermediate values, science articles that perform calculations, educational websites that work through math problems step by step, and physics discussions that derive numerical answers. A model trained on this content is implicitly exposed to many demonstrations of arithmetic reasoning, encoded in natural language. It learns that numbers relate to each other in specific ways, that certain patterns of text (a series of numbers and operators) predict certain outputs (a result), and that arithmetic errors produce inconsistencies that would not appear in well-written text.

The same logic applies to translation, code generation, and commonsense reasoning. Each of these capabilities correlates with the presence of relevant text in the pretraining corpus. Languages that appear less frequently in the training corpus have weaker translation quality. Programming languages that appear less in training data have weaker code generation. Commonsense situations that are underrepresented in English web text are handled less reliably.

This dependence on corpus composition has a direct practical implication: the capabilities of a large language model are in some sense a mirror of its training data. A model trained predominantly on English text will be better at English tasks than tasks requiring other languages. A model trained predominantly on contemporary web content will reflect contemporary cultural knowledge and biases rather than historical ones. A model trained on filtered, high-quality text may have better factual accuracy but narrower coverage than one trained on raw web crawl data.

Understanding this relationship between training data and capabilities is essential for both using large language models effectively and for understanding their failure modes. When GPT-3 produces confidently wrong information about an obscure topic, the first question to ask is whether that topic is well-represented in its training data. When it handles some linguistic constructions better than others, the explanation often lies in the frequency of those constructions in its pretraining corpus. The model is, in a deep sense, a compression of its training data into a set of weights that can be queried through language.

In Practice: Choosing Between Zero-Shot and Few-Shot

In real deployments, the decision between zero-shot and few-shot prompting involves tradeoffs. Zero-shot is simpler and cheaper (shorter prompts), but few-shot typically performs better. The optimal number of examples varies by task and model, usually somewhere between 4 and 16 for GPT-3. More than 20 examples often provides diminishing returns while consuming context that could be used for longer test inputs. For production systems, it is worth empirically testing 0, 4, 8, and 16 examples on a representative sample of test cases to find the right balance between performance and cost.

GPT-3 and the Problem of Reasoning

One of the clearest limitations revealed by GPT-3's evaluation is the gap between language understanding and systematic reasoning. This distinction is subtle but important, and GPT-3 helped clarify it by providing the first model capable enough to make the distinction crisp.

Language understanding refers to the ability to process and represent the meaning of text: parsing syntactic structure, resolving coreference, understanding word sense, and building representations of discourse. GPT-3 is excellent at this. It can read a complex passage and correctly answer questions about what happened, who did what to whom, and what the author intended. These tasks require understanding, but they can often be answered by pattern matching against familiar textual patterns.

Systematic reasoning refers to the ability to apply logical operations in a step-by-step, reliable way to derive conclusions from premises. This includes deductive reasoning (if P implies Q and P is true, then Q is true), abductive reasoning (finding the most likely explanation for observed evidence), and mathematical reasoning (applying arithmetic and algebraic operations in the correct order). GPT-3 struggles with these tasks not because it cannot understand the problem statement but because applying multi-step logical rules reliably requires a kind of precise, sequential computation that natural language training does not directly optimize for.

The contrast becomes particularly clear on syllogism tasks. Given "All mammals are warm-blooded. Dolphins are mammals. Are dolphins warm-blooded?", GPT-3 consistently answers correctly. The pattern is familiar: the training data is filled with examples of this kind of reasoning. But given "All glorks are flumps. Wumbles are glorks. Are wumbles flumps?", performance drops significantly. The logical structure is identical, but the words are novel, and the model cannot rely on stored associations about glorks, flumps, and wumbles. It must reason from pure logical structure, and here it begins to fail.

This observation motivated chain-of-thought prompting, developed shortly after GPT-3. The key insight is that if you show the model examples of step-by-step reasoning in your few-shot demonstrations, it generates step-by-step reasoning for new problems, and this substantially improves performance on multi-step tasks. The reasoning steps serve as an explicit working memory, externalizing the intermediate computation that the model cannot reliably maintain internally. Rather than trying to "think" through multiple steps in a single forward pass, the model generates text that represents each step, using the generated text as context for the next step.

Chain-of-thought prompting does not fully solve the reasoning problem, but it illustrates an important principle: careful prompt design can use what the model is already good at (generating coherent text) to address what it is bad at (maintaining multi-step computation internally). Many of the most effective prompt engineering techniques follow this pattern: they find ways to reformulate hard tasks as easier language generation tasks, exploiting the model's strengths to work around its limitations.

GPT-3 in the Broader GPT Architecture Evolution

To place GPT-3 in context, it helps to trace the specific architectural and training decisions that evolved across the GPT family and to understand why each change was made.

GPT-1 established the basic paradigm: train a large transformer decoder on next-token prediction, then fine-tune on specific tasks. The model was trained on BookCorpus (800 million words) and showed that pretraining significantly improved downstream performance compared to training task-specific models from scratch. But fine-tuning was still required for good performance; zero-shot generalization was weak.

GPT-2 scaled up both the model (from 117M to 1.5B parameters) and the training data (from BookCorpus to WebText, an 8 million document dataset filtered from Reddit). The key finding was that the model sometimes matched or approached fine-tuned models on certain benchmarks in zero-shot settings, but this capability was inconsistent. The authors famously withheld the full model initially due to concerns about misuse, a decision that generated significant public debate about the dual-use potential of large language models.

GPT-3 continued the scaling trajectory but changed the evaluation methodology. Instead of reporting only fine-tuned performance, the GPT-3 paper systematically evaluated zero-shot, one-shot, and few-shot performance for all benchmarks. This methodological choice was itself a contribution: it established a standard for evaluating in-context learning capabilities that influenced how subsequent models were assessed.

The architectural continuity across GPT-1, GPT-2, and GPT-3 is notable. All three use the same basic decoder-only transformer design. The changes across generations are primarily quantitative (more parameters, more data, more compute) rather than qualitative. This continuity was intentional: it allowed the researchers to attribute capability differences to scale rather than to architectural innovations, making the scaling laws more cleanly interpretable.

Looking forward, the GPT-4 and subsequent Claude models broke from pure architectural continuity, incorporating instruction tuning, reinforcement learning from human feedback (RLHF), and other techniques. These additions addressed some of GPT-3's most significant practical limitations: the tendency to produce harmful content, the inconsistent adherence to instructions, and the lack of context-sensitive helpfulness. But the underlying architecture, the decoder-only transformer trained on next-token prediction, remained the foundation.

Worked Example: Analyzing In-Context Learning Behavior

Let's work through a concrete example that illustrates several key aspects of in-context learning: the effect of example quality, the importance of label balance, and the degradation of performance when demonstrations conflict.

Suppose you want to use GPT-3 for a three-way sentiment classification task (Positive, Neutral, Negative). You have a set of test reviews and you want to construct the best possible few-shot prompt. Here is how different choices affect the setup:

Scenario 1: Unbalanced labels. You select 6 examples, but 4 are Positive, 1 is Neutral, and 1 is Negative. The model will be biased toward predicting Positive because it has seen that label most recently and most frequently. For a well-calibrated classifier, balance matters, and you should aim for roughly equal representation of all classes in your demonstrations.

Scenario 2: Label-text mismatch. You accidentally swap some labels while preparing your examples: a clearly positive review is labeled Neutral, and a neutral review is labeled Positive. For this simple classification task, the model may perform nearly as well as with correct labels (consistent with the Min et al. findings), because it is mainly using format information. But if the task is more complex, the mismatched labels will introduce label noise.

Scenario 3: Semantically distant examples. Your test inputs are about restaurant experiences, but your examples are about software products. The input distribution mismatch means the model has less relevant context for making predictions. Performance may be better than zero-shot (the format is still informative) but worse than if you used restaurant-focused examples.

Scenario 4: Optimal construction. You select 9 examples: 3 Positive, 3 Neutral, 3 Negative, all about restaurants, ordered so that the example most similar to your typical test input is last. This configuration maximizes the format signal, provides balanced label information, minimizes distribution mismatch, and exploits the recency effect. It represents best-practice few-shot construction for this task.

The contrast between these scenarios illustrates that in-context learning, while powerful, requires thoughtful prompt construction. It is not a magic black box where any examples produce good results. Systematic thinking about label balance and semantic relevance, along with ordering and format, pays off in practice.

Understanding GPT-3's Attention Patterns

One way to build intuition about what GPT-3 is doing during in-context learning is to think carefully about how causal self-attention behaves when the context contains few-shot demonstrations followed by a test input.

Recall from our earlier chapters that causal attention allows each token to attend to all previous tokens but not to future ones. In a few-shot prompt, this means the test input token representations can attend to all demonstration tokens, but the demonstration tokens cannot attend to the test input. This asymmetry is important: the model has full access to the demonstrations when processing the test input, but the demonstrations are processed in isolation from each other and from the test input.

Think of the demonstrations as a kind of prefix that conditions the model's predictions. When GPT-3 generates the response to the test input, each generated token can attend to the entire preceding context, including all demonstrations. The attention mechanism thus acts as a soft lookup: the test input representation queries the demonstration representations, and tokens that are semantically similar to the test input receive high attention weights. Their associated value vectors, which encode information about the corresponding outputs, flow into the test input representation and influence the prediction.

The depth of GPT-3's transformer (96 layers) means this conditioning happens iteratively, with each layer refining the representation further. Early layers primarily process syntactic structure; later layers process semantic content and task-relevant patterns. By the final layers, the test input representation has been significantly shaped by the demonstrations, incorporating information about what kind of output is expected. This is the mechanism by which in-context learning integrates the signal from demonstrations without any parameter updates.

A particularly interesting consequence of causal attention is that the order of demonstrations matters in a specific way: later demonstrations have more influence because they can be directly attended to from the test input position with fewer intervening tokens. Earlier demonstrations must propagate their influence through the intermediate positions, which introduces some signal degradation. This is the mechanistic explanation for the recency effect in few-shot prompting.

What GPT-3 Did Not Do

It is worth being clear about what GPT-3 did not achieve, beyond the general limitations discussed earlier. Understanding the specific nature of GPT-3's failures helps clarify what problems remained for future work.

GPT-3 did not develop a persistent world model that it updates as new information arrives. Each query is processed independently; the model has no memory of previous interactions. This means it cannot track changes in the world, maintain consistent beliefs across a conversation, or update its "knowledge" based on corrections you provide in the prompt beyond what fits in the context window.

GPT-3 did not develop reliable calibration. A well-calibrated model should be uncertain about things it does not know and confident about things it does know. GPT-3 often shows the opposite pattern: high confidence about incorrect facts and sometimes hedging unnecessarily about correct ones. The model has no separate "uncertainty module"; its expressed confidence is a product of the token probabilities it assigns, which reflect training data frequency more than epistemic warrant.

GPT-3 did not resolve the grounding problem. Its representations are built entirely from language and have no connection to perceptual experience, physical interaction, or direct observation of the world. All of its "knowledge" is mediated through text, and all of its generation is conditioned on textual patterns. This means it can use words correctly in context without having any model of what those words refer to in the physical world. A model that has never seen a banana but has read millions of sentences about bananas has a different kind of "knowledge" of bananas than a model that has been trained on visual or tactile data.

GPT-3 did not achieve instruction following in a reliable, consistent way. While it could often understand and respond to instructions in prompts, its compliance was uneven and depended heavily on how instructions were phrased. It would sometimes ignore instructions, produce outputs that technically satisfied the letter but not the spirit of instructions, or generate outputs harmful or offensive in ways that violated reasonable expectations. These issues motivated the development of instruction tuning and RLHF, which are central to the design of GPT-4, Claude, and other subsequent models.

These limitations are not failures of GPT-3's design so much as they are clear statements of what next-token prediction on text data alone cannot achieve. They define the research agenda that occupied the field in the years following GPT-3's release and shaped the architectural and training innovations of the next generation of large language models.

The Broader Context: Why 2020 Was the Right Moment

GPT-3 did not arrive in a vacuum. Its emergence was enabled by a convergence of factors that made 2020 the right time for this kind of model to exist and have the impact it did. Understanding this context helps explain why GPT-3 was so influential rather than just one more large language model in a long progression.

Available hardware had shifted decisively in favor of transformer training. NVIDIA's A100 GPU, released in 2020, offered roughly 20 petaFLOPS of bfloat16 performance per chip, dramatically faster than the V100s used for earlier large-scale training. Mixed-precision training techniques had matured, allowing models to train in 16-bit floating point with careful loss scaling while maintaining stability equivalent to 32-bit training. The combination meant that 175B parameter training was practically feasible for the first time, with OpenAI reportedly using a cluster of several thousand A100s for GPT-3's training run.

The software ecosystem had also matured. PyTorch and TensorFlow had both developed stable distributed training frameworks, and OpenAI had developed its own infrastructure for training models at this scale. Techniques like gradient checkpointing (trading compute for memory by not storing all activations during the forward pass) and pipeline parallelism (splitting the model across multiple machines along the depth dimension) made it possible to fit and train models that no single machine could hold in memory.

The dataset situation had evolved significantly from GPT-2's era. Common Crawl had been accumulating web content for years and provided a massive corpus of diverse text. The filtering techniques applied to Common Crawl, including deduplication, quality filtering using a trained classifier, and weighting by content quality metrics, represented a substantial advance over the simpler text collection methods of earlier models. The combination of raw scale and careful filtering produced training data that was both large enough and high-quality enough to train GPT-3 effectively.

Most importantly, the research community had developed the conceptual framework necessary to recognize and articulate what GPT-3's capabilities meant. The vocabulary of "in-context learning," "few-shot prompting," "emergent capabilities," and "scaling laws" did not exist before the GPT-3 paper. The paper itself introduced and defined these concepts. This provides a language for discussing phenomena that had been observed in less striking forms in earlier models but had not been systematically studied. The conceptual framing was as important as the empirical results in making GPT-3 so influential.

Comparing Few-Shot Learning to Fine-Tuning

A natural question when encountering in-context learning for the first time is: when should you use few-shot prompting instead of fine-tuning? GPT-3's paper made the case for prompting as a general alternative to fine-tuning, but the tradeoff is more complicated. Understanding it helps you make the right choice for specific applications.

Fine-tuning updates the model's weights based on task-specific examples. This process changes what the model knows about the task, burning in the learned patterns in a way that persists across all subsequent queries. Fine-tuning is more sample-efficient than in-context learning in the sense that it can achieve good performance on a task with fewer total examples (the updates accumulate, rather than being re-provided with every query). It is also more computationally efficient at inference time: fine-tuned models do not require long prompts containing demonstrations, which reduces inference latency and cost.

In-context learning provides no persistent learning. Every query must include the demonstrations in the prompt, costing context window space and inference compute. Performance is typically lower than a well-fine-tuned model on the same task, especially for complex or specialized tasks where fine-tuning provides substantial signal. In-context learning also requires careful prompt engineering to achieve good performance, whereas fine-tuning involves a more systematic optimization process.

The advantages of in-context learning lie in its flexibility and zero-data-collection requirements. With fine-tuning, you need to collect and label a substantial dataset before you can achieve good performance. You need to run training compute (non-trivial for large models). You need to store and serve a different model checkpoint for each task. With in-context learning, you can try a new task immediately by writing a few examples in a prompt. You can switch between tasks in a single inference pass by changing the prompt. You can experiment with different framings of the same task without retraining.

This makes in-context learning particularly valuable in several scenarios: rapid prototyping where you want to test whether a task is feasible before investing in a fine-tuning pipeline, low-data settings where you cannot collect enough labeled examples for effective fine-tuning, dynamic tasks where the definition of the task changes frequently, and multi-task systems where a single deployment needs to handle many different task types flexibly.

For production systems where a specific task is well-defined, data collection is feasible, and performance needs to be maximized, fine-tuning generally wins over in-context learning. For exploration and flexible zero-data-collection scenarios, in-context learning is the right choice. GPT-3 was the first model capable enough that this tradeoff was interesting rather than trivially resolved in favor of fine-tuning for every serious application.

Tokenization and Its Consequences for GPT-3

GPT-3 uses the same byte-pair encoding (BPE) tokenizer as GPT-2, with a vocabulary of 50,257 tokens. This tokenizer was trained on a corpus similar to WebText and represents a reasonable tradeoff between vocabulary size and coverage. Understanding the tokenizer's properties is important for understanding both GPT-3's capabilities and some of its failure modes.

BPE tokenization splits text into a sequence of subword units. Common words like "the", "is", and "language" are typically single tokens. Less common words are split into component subwords: "tokenization" might become ["token", "ization"]. Rare words or proper nouns may be split into individual characters or byte representations. The exact splits depend on the frequency statistics of the training corpus.

This tokenization has several consequences for GPT-3's behavior. First, the model processes text at the token level, not the character level. This means tasks that require character-level reasoning, such as counting letters, reversing strings, or identifying if a word is a palindrome, are difficult because GPT-3 does not have direct access to character-level information. A word like "level" is represented as a single token; to reverse it, the model would need to reason about the characters that compose that token, which is not a natural operation for a model trained on token sequences.

Second, the tokenization introduces a specific granularity for arithmetic. Numbers are tokenized in ways that do not correspond to their place-value structure. The number "1234" might be tokenized as a single token, or as "123" and "4", or as "12" and "34", depending on what subword units were learned. This inconsistent granularity helps explain GPT-3's arithmetic performance: the model can learn arithmetic operations on token sequences, but the fact that the same number can be tokenized differently in different contexts makes reliable computation difficult. A model with a character-level tokenizer would have more consistent access to the digit structure needed for reliable arithmetic.

Third, the tokenizer handles different languages with varying efficiency. Languages that appear frequently in the training corpus have more specialized tokens, requiring fewer tokens to represent the same amount of text. Languages that appear rarely are encoded more inefficiently, requiring more tokens per word. This efficiency difference translates to a context length disadvantage: GPT-3 can fit more English text in its 2,048 token context than equivalent text in a low-resource language, which disadvantages non-English speakers when using the model.

These tokenization effects are not unique to GPT-3, they affect all BPE-based models, but GPT-3's widespread deployment made them practically significant. Understanding that "the model struggles with character-level tasks" is not a vague capability limitation but a specific consequence of how text is represented provides a much clearer basis for predicting where the model will fail and for designing prompts that work around the limitation.

The Legacy of GPT-3

Four years after its release, GPT-3's specific capabilities have been vastly surpassed by subsequent models. GPT-4, Claude 3, Gemini, and Llama 3 all substantially outperform GPT-3 on every benchmark. The models that matter for production use in 2024 were not available in 2020. In this sense, GPT-3 is historical rather than current.

But GPT-3's legacy is not primarily about its performance numbers. It is about what it demonstrated, made people believe was possible, and what research directions it catalyzed. The specific ideas it contributed continue to shape the field.

GPT-3 brought systematic attention to in-context learning, which had appeared in weaker form in GPT-2. The prompting approach that now is common in how practitioners interact with large language models traces directly to GPT-3's demonstration that prompting was a viable alternative to fine-tuning. The scaling laws literature, which provides the theory for rational investment in large-scale training, was materially advanced by the systematic multi-scale comparisons in GPT-3's evaluation. GPT-3 changed the public conversation about AI capabilities and risks, along with governance, in ways that continue to shape policy discussions.

Most concretely, GPT-3 changed what developers believed was possible and started building toward. Its API demonstrated a new business model: training a single large model and providing API access, rather than selling software that runs locally. It demonstrated that a single general-purpose model could provide value across dozens of different tasks, enabling product development that would have required multiple specialized models before. It showed that language generation could be fluent enough and controllable enough to be the output interface for a wide range of applications.

The follow-on work that GPT-3 directly motivated includes instruction tuning (improving model adherence to natural language instructions), RLHF and AI alignment research (addressing the value alignment and safety issues that GPT-3's deployment revealed), efficient adaptation methods like LoRA and prefix tuning (making it practical to fine-tune large models with limited compute), retrieval-augmented generation (addressing the knowledge cutoff and hallucination limitations), and constitutional AI (building in behavioral constraints at the training level rather than relying on prompt engineering). Each of these areas can trace a direct line back to limitations that GPT-3's deployment surfaced.

In this sense, GPT-3's most important contribution may have been its failures. By being capable enough to deploy and useful enough for people to try, it revealed why large language models were not yet dependable or safe enough to help users consistently. These failures defined the work of the half-decade that followed.

The next generation of models did not simply scale GPT-3 up. They refined the training recipe based on the lessons GPT-3 taught: better data filtering, compute-optimal training ratios, instruction following, and safety alignment. These refinements are the subject of later chapters. But the foundation, the decoder-only transformer trained on next-token prediction as a path to general language understanding, is the framework that GPT-3 most definitively established and that continues to underpin the state of the art today.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about GPT-3, few-shot learning, and in-context learning.

GPT-3 and In-Context Learning Quiz

Question 1 of 100 of 10 completed
How many parameters does GPT-3's largest model have?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025gpt3-2, author = {Michael Brenndoerfer}, title = {GPT-3: Scale, Few-Shot and In-Context Learning}, year = {2025}, url = {https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). GPT-3: Scale, Few-Shot and In-Context Learning. Retrieved from https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning
MLAAcademic
Michael Brenndoerfer. "GPT-3: Scale, Few-Shot and In-Context Learning." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning>.
CHICAGOAcademic
Michael Brenndoerfer. "GPT-3: Scale, Few-Shot and In-Context Learning." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning.
HARVARDAcademic
Michael Brenndoerfer (2025) 'GPT-3: Scale, Few-Shot and In-Context Learning'. Available at: https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). GPT-3: Scale, Few-Shot and In-Context Learning. https://mbrenndoerfer.com/writing/gpt-3-scale-few-shot-in-context-learning

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.