Attention Visualization: Extracting and Interpreting Weights

Michael BrenndoerferFebruary 11, 202651 min read

Part of Language AI Handbook

Extract attention weights from transformer models, visualize head patterns, measure head entropy, and understand the key caveats.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Attention Visualization

Transformer models route information through attention mechanisms at every layer. When a model processes the sentence "The bank can guarantee deposits will eventually cover future tuition costs," the word "bank" needs to be disambiguated. Does it mean a financial institution or a river bank? The model resolves this by attending to surrounding context, and that process leaves a trace: the attention weights that governed information flow. Attention visualization is the practice of extracting and displaying those weights to understand, at least partially, what the model is "looking at" when it processes text.

Interest in attention visualization grew rapidly after the transformer architecture was introduced in 2017. Researchers and practitioners wanted to know whether the internal computations of these black-box models could be made legible. If attention weights reveal which words influence which other words, then perhaps they also reveal something about the model's "reasoning." This hope turned out to be both partly justified and partly misleading, a tension you will explore throughout this chapter and in the next.

The technique has a particular appeal because attention weights look interpretable by construction. They are non-negative, they sum to one across each row, and they are indexed by the tokens in the input sequence. You can point to an entry in the attention matrix and say "this token attended 0.4 to that token," and the statement feels meaningful. The challenge is that the relationship between this appearance of interpretability and the actual computational role of attention weights is more complicated than it first seems. This chapter develops both the practical skill of producing attention visualizations and the critical judgment to interpret them appropriately.

The goal here is practical: you will learn how to extract attention weights from transformer models, visualize them in interpretable ways, understand what individual attention heads tend to specialize in, and appreciate the significant caveats that govern what conclusions you can and cannot draw from this type of analysis. By the end, you will be equipped to use attention visualization as one component in a broader interpretability toolkit, positioned alongside gradient-based methods and probing classifiers that are covered in subsequent chapters.

Attention Mechanisms: A Brief Recap

Building on the transformer architecture covered in the earlier parts of this book, recall that attention computes a weighted combination of values, where the weights reflect how relevant each position in the input is to the current query. This chapter assumes you are familiar with that mechanism. The key point for visualization purposes is that attention is not a single scalar per layer but a matrix per head per layer. This dimensionality is both a strength and a challenge: a model like BERT produces so many attention matrices that there is too much to examine by eye, which is why systematic, entropy-guided analysis is necessary rather than optional.

For a sequence of nn tokens, each attention head produces an n×nn \times n weight matrix. Entry (i,j)(i, j) in this matrix represents the degree to which position ii attends to position jj when the model processes that head. With a model that has LL layers and HH heads per layer, a single forward pass generates L×HL \times H such matrices. For BERT-base, which has 12 layers and 12 heads per layer, that is 144 attention matrices for every sentence. Each of those 144 matrices encodes a different view of the same sequence, learned independently through gradient descent. This redundancy is a feature: it allows the model to simultaneously reason about multiple types of relationships within the same input.

Attention Head

An attention head is one of the parallel attention mechanisms within a transformer layer. Each head applies its own learned query, key, and value projections to the input, producing an independent attention pattern. Multiple heads allow the model to simultaneously attend to different aspects of the input. The outputs of all heads are concatenated and linearly projected to produce the layer's final output.

The attention weight matrix for a single head is computed as:

A=softmax(QKTdk)A = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)

where:

  • QQ: the query matrix, shape (n×dk)(n \times d_k), obtained by projecting the input sequence through a learned weight matrix WQW^Q
  • KK: the key matrix, shape (n×dk)(n \times d_k), obtained through a learned projection WKW^K
  • dkd_k: the dimension of the key vectors, used to scale the dot products and prevent gradient saturation from large values
  • softmax\text{softmax}: applied row-wise, so each row of AA sums to 1 and can be interpreted as a probability distribution over positions

The resulting matrix AA has shape (n×n)(n \times n), and row ii is the attention distribution for token ii: it tells you how much weight token ii places on each token in the sequence when computing its updated representation.

Why does this formulation matter for visualization? The softmax normalization means that every row of AA is a proper probability distribution. You can interpret each row as a question: "Given that I am at position ii and I want to update my representation, which other positions should I consult, and by how much?" The answer varies by head and by layer, and that variation is exactly what attention visualization tries to expose.

Why Attention Weights Are Accessible

Attention weights occupy a privileged position among neural network internals: they are explicitly computed as an intermediate variable during the forward pass, before being used to weight the value vectors. Unlike the raw numerical weights of a dense layer, which transform inputs in high-dimensional space and resist direct interpretation, attention weights sit between zero and one and sum to one across each row. That probabilistic structure makes them look interpretable even when they may not fully reflect causal reasoning.

This accessibility explains why attention visualization emerged so early in transformer research. By 2018, papers like "Visualizing Attention in Transformer-Based Language Representation Models" and the original transformer paper's supplementary material were already displaying attention heatmaps. The technique requires minimal tooling, produces easy-to-read outputs, and maps onto intuitive concepts like "which words are important?" Whether the visual appeal tracks the underlying computation is the deeper question this chapter addresses.

Extracting Attention Weights

Attention weights are not outputs you receive by default when you call a language model. You need to explicitly request them and hook into the model's internals. Modern transformer libraries make this straightforward, but it is worth understanding what you are requesting and how the data is structured before you begin visualizing it.

Using Hugging Face Transformers

The Hugging Face transformers library provides a simple mechanism for extracting attention weights. When you call a model with output_attentions=True, the return value includes all attention matrices from all layers.

The attention output is a tuple of tensors, one per layer. Each tensor has shape (batch_size, num_heads, sequence_length, sequence_length). The entry at [batch, head, query_position, key_position] is the attention weight that the token at query_position places on the token at key_position, for that head.

In[3]:
Code
import torch
from transformers import AutoModel, AutoTokenizer

# Load model and tokenizer
model_name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name, output_attentions=True)
model.eval()

# Example sentence
text = "The bank can guarantee deposits will eventually cover future tuition costs."
inputs = tokenizer(text, return_tensors="pt")

# Forward pass with attention output
with torch.no_grad():
    outputs = model(**inputs, output_attentions=True)

# Extract attention weights: tuple of (num_layers,) tensors
# Each tensor: (batch=1, num_heads, seq_len, seq_len)
attention_weights = outputs.attentions

# Get tokens for display
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
num_layers = len(attention_weights)
num_heads = attention_weights[0].shape[1]
seq_len = attention_weights[0].shape[2]
Out[4]:
Console
Number of layers:   12
Attention heads per layer: 12
Sequence length (with [CLS]/[SEP]): 14
Tokens: ['[CLS]', 'the', 'bank', 'can', 'guarantee', 'deposits', 'will', 'eventually', 'cover', 'future', 'tuition', 'costs', '.', '[SEP]']

BERT-base produces 12 layers, 12 heads each, and the sequence includes the special [CLS] token at position 0 and [SEP] at the end. These special tokens participate in attention like any other token, which is worth remembering when interpreting visualizations.

The Role of Special Tokens

The [CLS] and [SEP] tokens are not accidental participants in the attention computation. BERT was trained with [CLS] serving as the aggregate sentence representation: the vector at [CLS] after all 12 layers accumulates information from the entire sentence and is fed into the classification head for downstream tasks. This means that [CLS] must actively gather information, and many attention heads in later layers show substantial weight directed toward or from [CLS].

Similarly, [SEP] marks sentence boundaries. Research has shown that some attention heads treat [SEP] as a "null" or "no-op" target: when no other position in the sequence is highly relevant for a given query, the head routes attention to [SEP] as a way of adding minimal information. This is a learned behavior, not an architectural constraint, and it appears with striking consistency across heads.

Understanding the special token dynamics is important when reading attention heatmaps. A head that appears to strongly "attend to punctuation" may be attending to [SEP], which is functionally different. Similarly, a head whose attention is concentrated on [CLS] may be performing sentence-level aggregation rather than identifying an important word.

Accessing Individual Heads

Once you have the attention tuple, you can extract any specific layer and head combination. The convention is zero-indexed: layer 0 is the first transformer block, head 0 is the first attention head within that block.

In[5]:
Code
# Extract layer 5, head 0 (zero-indexed)
layer_idx = 5
head_idx = 0

# Shape: (seq_len, seq_len)
attn_matrix = attention_weights[layer_idx][0, head_idx].numpy()
Out[6]:
Console
Attention matrix shape: (14, 14)
First row (attention FROM token 'the'):
  [CLS]           0.0357
  the             0.0444
  bank            0.2899
  can             0.2238
  guarantee       0.1598
  deposits        0.0650
  will            0.0054
  eventually      0.0009
  cover           0.0005
  future          0.0001
  tuition         0.0004
  costs           0.0005
  .               0.0017
  [SEP]           0.1720

The first token in the sequence is always [CLS]. The output above shows which tokens the second token attends to when processing this sentence. Each row sums to 1.0 because of the softmax normalization.

The fact that every row sums to 1 has a subtle implication for interpretation. If a sequence has 14 tokens, the "average" attention weight for any given row is 1/14≈0.0711/14 \approx 0.071. A weight of 0.3 on a particular token is therefore roughly four times the average, which is meaningful. But if a sequence has 100 tokens, a weight of 0.3 on any single token stands out dramatically from the 0.010.01 average. Interpreting attention weights without accounting for sequence length is a common mistake that can lead to comparing patterns across sentences of different lengths without adjusting for the different baselines.

Visualizing Attention Patterns

Raw numbers are difficult to interpret. Heatmap visualizations of the attention matrix make patterns visible at a glance, and the choice of visualization format significantly affects what patterns you notice and what you miss.

Single-Head Heatmap

The most common visualization maps the attention matrix to a color grid where rows correspond to query tokens and columns to key tokens. Cell (i,j)(i, j) is colored proportionally to the attention weight that token ii places on token jj. Reading across a row tells you the full attention distribution for a single query token: where does this token "look" when updating its representation? Reading down a column tells you which queries attend heavily to this key token: which tokens find this position useful?

Out[7]:
Visualization
Heatmap of attention weights where rows are query tokens and columns are key tokens with varying shades
Attention weight heatmap for BERT layer 6, head 1 (zero-indexed layer 5, head 0), processing an ambiguous sentence about 'bank'. Each row shows the attention distribution for one token across all sequence positions. Darker cells indicate stronger attention, with most rows concentrated on a small subset of positions rather than spread evenly across all tokens.

Patterns often visible in these heatmaps include diagonal bands (each token attending mostly to itself), strong attention toward the [CLS] token or punctuation tokens, and occasional long-range connections between semantically related tokens. The fact that certain patterns recur across different sentences is evidence that a head has learned a general function rather than fitting a single example.

Choosing Color Scales

The color scale in an attention heatmap matters more than it might seem. By default, matplotlib normalizes the colormap to span the range of values present in the matrix. If one cell has an attention weight of 0.9 and all others are below 0.1, the single dominant cell will appear very dark and everything else will appear nearly identical shades of light blue. This can make diffuse patterns look more uniform than they are.

Two useful alternatives are: (1) fixing the color scale from 0 to 1 (vmin=0, vmax=1), which puts all heads on a common scale and makes sparse versus dense patterns visually distinct, and (2) using a logarithmic color scale, which amplifies small differences in low-weight cells and can reveal secondary patterns that would otherwise be washed out. Each choice tells a different story, so it is worth producing both when exploring a new head.

Comparing Multiple Heads

Because each head learns an independent set of query, key, and value projections, different heads in the same layer often develop distinct attention patterns. Visualizing multiple heads side by side reveals this specialization and demonstrates that the layer is computing multiple things simultaneously.

In[8]:
Code
# Extract four heads from layer 6 (zero-indexed layer 5) for comparison
layer_idx = 5
heads_to_show = 4  # Show first 4 heads

layer_attns = [
    attention_weights[layer_idx][0, h].numpy() for h in range(heads_to_show)
]
Out[9]:
Visualization
Heatmap of attention weights for head 1 of BERT layer 6, with a partial diagonal and strong final SEP column
Layer 6, Head 1. Attention follows a partial diagonal across neighboring content tokens and also concentrates on the final `[SEP]` position.
Heatmap of attention weights for head 2 of BERT layer 6, dominated by attention to SEP
Layer 6, Head 2. Most query positions direct their strongest attention to `[SEP]`, with comparatively little weight elsewhere.
Out[10]:
Visualization
Heatmap of attention weights for head 3 of BERT layer 6, with diffuse content attention and a strong SEP column
Layer 6, Head 3. Attention is more diffuse across content tokens while retaining a strong column at `[SEP]`.
Heatmap of attention weights for head 4 of BERT layer 6, combining SEP attention with localized token connections
Layer 6, Head 4. Attention combines a prominent `[SEP]` column with several localized content-token connections.

Looking across heads, you can see that attention is not monolithic. Some heads produce sharp, focused patterns while others spread attention diffusely. This diversity is intentional: the multi-head mechanism exists precisely to allow different heads to capture different types of relationships simultaneously. If all heads learned the same pattern, most of the model's parameter budget for attention would be wasted. The fact that heads specialize is evidence that the training objective is pushing each head toward a distinct functional niche, even though no explicit diversity loss is imposed during training.

Aggregating Attention Across Heads

A tempting operation is to average all attention matrices in a layer into a single summary heatmap. This produces a view of "what the layer attends to on average." However, averaged attention is almost always more diffuse than any individual head and often has lower information content than the most specialized heads. The averaging smears together patterns that operate at very different scales and for different purposes.

A better aggregation strategy for identifying important tokens is to take the maximum attention weight received by each key token across all heads. This shows you which tokens any head found relevant, rather than which tokens all heads found moderately relevant. For understanding model structure, examining individual heads and comparing them directly produces deeper insight than any aggregate.

Attention Head Specialization

One of the most striking findings from early attention analysis is that individual heads often exhibit recognizable functional patterns. Researchers at Hugging Face and elsewhere systematically analyzed attention heads in BERT and found that certain heads reliably perform specific linguistic roles.

Common Head Patterns

Through careful examination of many sentences, several recurring head types have been documented. These patterns are not accidental: they reflect how the model solved the masked language modeling objective, which requires understanding contextual meaning, syntactic agreement, and semantic relationships across a sentence.

Positional heads attend predominantly to the previous or next token, implementing a form of local context window. These heads appear consistently and are easy to spot: the heatmap shows a band just above or below the diagonal. The previous-token head is particularly common in early layers, where it serves a function analogous to a left-context n-gram window.

Syntactic heads track grammatical relationships. For example, some heads connect verbs to their subjects or objects, or link determiners to the nouns they modify. When you visualize these heads on sentences with clear syntactic structure, the attention matrix looks almost like a dependency parse tree: non-zero cells correspond to syntactic arcs. Clark et al. (2019) identified specific BERT heads that align well with dependency parse relations at rates significantly above chance, particularly for relations like nominal subjects, direct objects, and adjectival modifiers.

Coreference heads link pronouns to their antecedents. In a sentence like "The lawyer questioned the witness. She doubted his testimony," a coreference head in a later layer might show strong attention from "She" toward "The lawyer" and from "his" toward "the witness." These heads are particularly interesting because coreference is a long-range phenomenon that requires integrating information across sentence boundaries. The fact that some heads specialize in this suggests that even within-sentence attention can partially solve cross-sentence reference tasks when the context fits within the model's maximum sequence length.

Separator-attending heads show strong attention toward the [SEP] token, particularly in later layers. These heads are thought to implement a form of "no-op" or soft reset, attending to a semantically neutral token when no other position is relevant. The [SEP] head is a kind of "I have nothing useful to look at" signal that the model can route to when the query position has already gathered the context it needs. This pattern is specific to BERT-style models trained with sentence-pair tasks; GPT-style models do not have [SEP] tokens and therefore cannot develop this behavior.

Self-attending heads have a strong diagonal, meaning each token attends mostly to itself. These appear to aggregate local information without changing the representation significantly. They are most prevalent in very early layers and may serve an initialization function before the model builds richer cross-token representations.

The existence of these patterns provided early hope that transformers were learning interpretable linguistic structure rather than performing opaque statistical computation. However, the relationship between attention patterns and linguistic roles is correlational rather than causal. A head that looks like it tracks subjects need not be computing subject identification in any meaningful sense. The weights might have converged to this pattern for other reasons that happen to produce similar-looking attention matrices.

The Specialization Discovery Process

How did researchers discover these patterns in the first place? The methodology matters, because it affects how confident you should be in the findings.

The most common approach is qualitative: visualize many attention heads on many sentences and look for patterns that recur. When head X in layer Y consistently shows strong attention from pronouns to nearby nouns, you hypothesize that it tracks coreference. This is hypothesis generation, not hypothesis testing.

A more rigorous approach is correlation analysis: probe whether the attention pattern aligns quantitatively with a linguistic annotation. For a syntactic dependency head, you would take a parsed corpus, extract the dependency arcs, and measure what fraction of the model's attention edges match gold parse arcs. If the match rate is significantly above the random baseline, you have quantitative evidence for the hypothesis. Clark et al. (2019) used exactly this approach and found several BERT heads with dependency alignment scores above 0.5, meaning they matched the gold parse arc more than half the time across diverse sentences.

The limitation of correlation analysis is that it tells you whether the head's pattern resembles a linguistic structure, not whether the model is using that structure for its predictions. A head can produce syntactically organized attention without the downstream layers caring about those attention weights. This distinction motivates more interventional tools like probing classifiers and activation patching, covered in later chapters.

Attention Entropy as a Specialization Measure

One quantitative way to characterize heads is through the entropy of their attention distributions. A head with low entropy has sharp, focused attention concentrated on a few tokens. A head with high entropy distributes attention broadly. Plotting entropy across layers reveals structural patterns in how information flows through the model.

Entropy for a single attention row is:

H(Ai)=−∑j=1nAijlog⁡AijH(A_i) = -\sum_{j=1}^{n} A_{ij} \log A_{ij}

where:

  • AiA_i: the ii-th row of the attention matrix, representing the attention distribution for token ii
  • AijA_{ij}: the attention weight from token ii to token jj
  • The sum runs over all nn positions in the sequence

Higher entropy means more diffuse attention; lower entropy means more focused attention on specific positions.

To interpret this concretely: if a sequence has 14 tokens and a head distributes attention perfectly uniformly, each weight is 1/141/14 and the entropy is log⁡(14)≈2.64\log(14) \approx 2.64 nats. A head that concentrates all attention on a single token has entropy 00. Real heads fall between these extremes, but syntactic and coreference heads typically have much lower entropy than sentence-level aggregation heads.

The entropy measure is useful precisely because it is aggregatable. You can compute mean entropy across all query positions in a sentence, then average across sentences, to get a single number per head that characterizes how selective that head is across the whole dataset. This makes it possible to rank all 144 BERT-base heads from most focused to most diffuse and focus your visual inspection on the most interesting candidates.

In[11]:
Code
import numpy as np


def attention_entropy(attn_matrix):
    """Compute mean entropy across all query positions for one attention head."""
    # Clip to avoid log(0)
    attn_clipped = np.clip(attn_matrix, 1e-9, 1.0)
    entropy_per_row = -np.sum(attn_clipped * np.log(attn_clipped), axis=1)
    return float(np.mean(entropy_per_row))


# Compute entropy for all heads across all layers
entropy_matrix = np.zeros((num_layers, num_heads))
for layer in range(num_layers):
    for head in range(num_heads):
        attn = attention_weights[layer][0, head].numpy()
        entropy_matrix[layer, head] = attention_entropy(attn)
Out[12]:
Visualization
Heatmap of attention entropy values across 12 layers and 12 heads of BERT-base
Mean attention entropy per head across all 12 layers of BERT-base, with each cell representing one attention head. Lower entropy (darker blue) indicates focused, selective attention concentrated on few tokens, while higher entropy (lighter blue) corresponds to diffuse attention spread broadly. Early layers tend to show more variable entropy, while later layers often develop more specialized low-entropy heads alongside high-entropy aggregate heads.
Out[13]:
Console
Most focused head: Layer 3, Head 1 (entropy = 0.024)
Most diffuse head: Layer 1, Head 1 (entropy = 2.540)

The entropy profile often shows that early layers have relatively high-entropy heads (diffuse attention gathering broad context), while later layers develop more specialized, low-entropy heads. However, this pattern varies substantially across sentences, models, and fine-tuning tasks.

Attention Interpretation Caveats

Attention visualization is useful for forming hypotheses and building intuitions. It is also easy to misuse. Before acting on what you see in an attention heatmap, you should understand several well-documented limitations that have emerged from careful empirical work since 2019.

Attention Is Not Explanation

The most important caveat: high attention weight does not mean causal influence on the model's output. Attention weights tell you how information was routed, not why the model produced the output it did. A head might place high attention weight on a token for reasons entirely unrelated to the downstream prediction.

Researchers have demonstrated this directly. Jain and Wallace (2019) published "Attention is not Explanation," a study showing that in text classification tasks, alternative attention distributions that are very different from the learned distribution can produce the same predictions. If swapping the attention weights produces identical outputs, then the specific attention pattern cannot be the explanation for the output. Their experiments covered multiple NLP tasks and model types, and the finding was robust: attention weights and gradient-based feature importance measures frequently diverged.

Wiegreffe and Pinter (2019) responded with "Attention is not not Explanation," arguing that the question requires a more precise distinction. Under a different definition of "explanation" (one focused on faithfulness rather than counterfactual equivalence), attention weights can serve explanatory purposes in some settings. This debate has not been fully resolved, and it has productively pushed the field toward more careful definitions of what an "explanation" should provide.

The current consensus is pragmatic: treat attention as a hypothesis-generating tool, not a ground-truth explanation. The next chapter, Attention Analysis Limitations, covers this issue in depth with specific experiments.

Gradient and Attention Divergence

Studies using gradient-based attribution methods frequently disagree with attention-based attributions. You can find sentences where attention strongly highlights certain tokens but gradients show those same tokens have minimal effect on the output probability. You can equally find tokens that gradients flag as highly influential but that receive low attention weights.

This divergence is informative rather than disqualifying. It suggests that attention and gradients are measuring different things. Attention measures information routing: which positions' representations were blended into other positions' representations. Gradients measure sensitivity: how much the output would change if the input representation at a given position were perturbed. Both views capture measurable properties of the computation, but neither alone is the complete story.

Attention Heads May Not Be Modular

The observation that some heads look like they perform specific linguistic functions does not mean those heads are the only or even primary mechanism for those functions. When researchers ablate attention heads by setting their weights to uniform distributions, models often maintain performance far better than expected. The functions performed by the ablated heads appear to be distributed across other components or are simply not needed for the tasks tested.

This distributional redundancy means that even a correct identification of "head X tracks subject-verb relations" may not translate to "removing head X would impair subject-verb agreement." The model may just use other heads or MLP layers to compensate. Michel et al. (2019) showed that BERT models can be pruned to 50% of their heads with minimal performance loss on many tasks, and some heads could be pruned individually with zero impact. This resilience makes functional interpretation harder: if a function is not concentrated in one head, you cannot identify it by examining heads in isolation.

Visualization Tools May Mislead

BertViz, the most widely used attention visualization library, displays multiple heads simultaneously by overlaying or averaging them. When you average attention across all heads, the result is often much more diffuse and less informative than any individual head. Conversely, cherry-picking the head that produces the most interpretable pattern may give a misleading impression of how the model works overall.

Averaging attention across layers is even more problematic because the semantics of attention weights change across layers. Layer 1 attention operates on raw token embeddings; layer 12 attention operates on highly transformed representations. The numerical value 0.3 at layer 1 and layer 12 cannot be compared directly. When aggregating across layers, the resulting "average attention" conflates computations that are fundamentally incomparable.

Some researchers proposed "attention rollout" (Abnar and Zuidema, 2020) as a more principled aggregation method. Attention rollout computes how information flows from the input through all layers by multiplying attention matrices across layers and accounting for residual connections. The result is an attention flow matrix that attempts to quantify how much each input token contributes to each output token's representation through the full stack of layers. Attention rollout is more informative than naive layer averaging, but it still does not account for the value projections, which shape what information is passed forward.

Context Dependency

Attention patterns are sensitive to the specific input. The same head that tracks subject-verb relations on one sentence may show a completely different pattern on another. When you visualize attention on a single example, you are seeing one instantiation of a learned function, not a stable, general-purpose representation.

To characterize a head reliably, you need to examine it across hundreds or thousands of diverse examples and aggregate the patterns. Single-example visualization is useful for exploration but insufficient for claims about model behavior. A head that tracks coreference on three example sentences might track something entirely different on a fourth. The only way to know is to test systematically.

Attention Visualization Tools

Several libraries have been developed to make attention visualization accessible without writing custom matplotlib code. Each makes different design choices that affect what aspects of attention are easy to see.

BertViz

BertViz is the most widely used library for transformer attention visualization. It provides two main views that serve different analytical purposes.

The head view shows attention as arcs connecting query tokens to key tokens. The thickness of each arc represents the attention weight. You can show a specific token to see all attention flowing to and from it, or view all tokens simultaneously. The arc representation makes it easy to follow individual attention paths but can become cluttered when many tokens have non-trivial weights.

The model view provides a compact overview of all attention patterns across all layers, allowing you to compare layers at a glance and spot which layers contain the most interesting patterns. Each layer's attention pattern is shown as a small thumbnail that you can click to expand. This overview is particularly useful for identifying which layers are worth examining closely.

In[14]:
Code
# BertViz is designed for Jupyter notebooks
# Installation: uv pip install bertviz

# from bertviz import head_view, model_view
# head_view(attention_weights, tokens)   # Interactive head visualization
# model_view(attention_weights, tokens)  # Overview of all layers

BertViz outputs interactive HTML, which works in Jupyter but not in static documents. For paper figures or non-interactive reports, matplotlib heatmaps like those shown earlier are more suitable.

One limitation of BertViz's head view is that the arc representation is optimized for short sequences. With sequences of 30 or more tokens, the arcs overlap heavily and the visualization becomes difficult to read. For long sequences, a matrix heatmap scales much better visually.

Transformers Interpret

The transformers-interpret library provides integrated gradients alongside attention, allowing you to compare attention-based and gradient-based attributions for the same prediction. This comparison is valuable precisely because of the caveats discussed above: seeing where attention and gradients agree (and disagree) helps calibrate how much trust to place in attention alone.

When the two methods agree on a set of important tokens, that agreement is informative: both the routing mechanism (attention) and the sensitivity measure (gradients) point to the same tokens as relevant. When they disagree, the disagreement is a signal to investigate further. It may indicate that attention is routing information through an indirect path, that a token's influence operates through its effect on other tokens rather than directly, or that the gradient signal is dominated by a different computation than what attention shows.

Attention Rollout and Beyond

Beyond BertViz and transformers-interpret, a growing set of tools implements more principled attention-based interpretability methods.

Attention rollout (Abnar and Zuidema, 2020) propagates attention weights backward through all layers, accounting for residual connections. Instead of looking at a single layer's attention, you get an estimate of the full information path from each input token to each output token. The formula multiplies the augmented attention matrices (attention matrix plus identity, to represent the residual connection) layer by layer:

A~(l)=0.5⋅A(l)+0.5⋅I\tilde{A}^{(l)} = 0.5 \cdot A^{(l)} + 0.5 \cdot I R=A~(1)⋅A~(2)⋯A~(L)R = \tilde{A}^{(1)} \cdot \tilde{A}^{(2)} \cdots \tilde{A}^{(L)}

where:

  • A(l)A^{(l)}: the attention matrix at layer ll, averaged across heads
  • II: the identity matrix, representing the residual path that skips the attention computation
  • RR: the final rollout matrix, whose entry (i,j)(i, j) estimates how much input token jj influences output position ii through all layers
  • The 0.50.5 weight splits information equally between the attended path and the residual path

Attention rollout is still an approximation (it ignores value projections and nonlinearities), but it is more principled than naive layer averaging and has been shown to correlate better with gradient-based saliency maps on several benchmarks.

Manual Extraction vs. Library Wrappers

For research purposes, direct extraction using the Hugging Face output_attentions=True flag, as shown in this chapter, gives you maximum flexibility. Library wrappers like BertViz are useful for interactive exploration but can obscure the underlying data structure. When you need to aggregate patterns across a dataset, run quantitative analyses, or feed attention data into downstream analyses, working directly with the raw tensors is preferable.

The raw tensor approach also makes it easy to integrate attention analysis with other interpretability methods. You can compute attention entropy for the same sentences where you also compute gradient norms, compare the two, and feed both into a downstream analysis without being constrained by what a library exposes through its API.

Worked Example: Tracing "Bank" Disambiguation

Before scaling to systematic analysis, it helps to walk through one concrete example that illustrates both what attention visualization can reveal and its limits. The ambiguous word "bank" is a classic test case for disambiguation because the same surface form maps to two completely different meanings depending on context.

Consider processing "The bank can guarantee deposits will eventually cover future tuition costs" through BERT. You would like to know how the model resolves the meaning of "bank." The hypothesis: if BERT correctly understands that this is a financial bank, some attention head in a later layer should show "bank" attending to contextually disambiguating tokens like "deposits," "guarantee," or "tuition."

You can test this by extracting attention weights and checking whether there is a head in layers 5-12 where the token "bank" places substantial weight on "deposits." If you find such a head and confirm the pattern holds for other financial sentences but not for river-related sentences (e.g., "She walked along the bank of the river"), you have preliminary evidence of a disambiguation-related head.

But here is the problem: you can almost always find such a head if you look through all 144 attention matrices. With 144 matrices of different patterns for every sentence, some head will happen to show the pattern you are looking for. This is the multiple comparisons problem applied to attention analysis. The appropriate response is to define your hypothesis before examining the attention matrices, then test it on held-out sentences rather than confirming it on the same sentence you used to generate the hypothesis.

This worked example illustrates the general shape of attention visualization's epistemic situation: the visualizations are informative and hypothesis-generating, but confirming those hypotheses requires the kind of controlled analysis covered in subsequent chapters. The visual output is the beginning of the investigation, not its conclusion.

Code Implementation: Systematic Head Analysis

Rather than examining heads one at a time, a practical workflow involves computing summary statistics across the full dataset to find which heads are most distinctive. Let us walk through an example that identifies heads with the most focused attention and examines their patterns.

We first need a small dataset of sentences to analyze across. Using multiple sentences provides more reliable estimates of head behavior than any single example. The eight sentences below cover a range of syntactic structures and semantic domains, making sure that heads which show consistent patterns are general rather than fitting a specific input.

In[15]:
Code
sentences = [
    "The bank can guarantee deposits will eventually cover future tuition costs.",
    "The cat sat on the mat and looked out the window.",
    "Scientists discovered a new species of frog in the Amazon rainforest.",
    "She opened the door and walked into the brightly lit room.",
    "The algorithm processes each token in the sequence one step at a time.",
    "His mother gave him a book about ancient Roman history for his birthday.",
    "The company announced record profits despite difficult market conditions.",
    "Researchers found that sleep deprivation impairs decision-making ability.",
]

# Process all sentences and collect entropy values
all_entropies = []

for sent in sentences:
    enc = tokenizer(sent, return_tensors="pt")
    with torch.no_grad():
        out = model(**enc, output_attentions=True)
    attn_weights_sent = out.attentions

    sent_entropy = np.zeros((num_layers, num_heads))
    for layer in range(num_layers):
        for head in range(num_heads):
            attn = attn_weights_sent[layer][0, head].numpy()
            sent_entropy[layer, head] = attention_entropy(attn)
    all_entropies.append(sent_entropy)

# Average entropy across sentences
mean_entropy = np.mean(all_entropies, axis=0)  # shape: (num_layers, num_heads)
Out[16]:
Console
5 most focused heads (lowest mean entropy):
  Layer  3, Head  1: entropy = 0.021
  Layer  3, Head 10: entropy = 0.068
  Layer  2, Head  7: entropy = 0.311
  Layer  6, Head  2: entropy = 0.499
  Layer  5, Head  4: entropy = 0.501

5 most diffuse heads (highest mean entropy):
  Layer  1, Head  1: entropy = 2.547
  Layer  1, Head  7: entropy = 2.356
  Layer  1, Head  5: entropy = 2.351
  Layer  1, Head  9: entropy = 2.319
  Layer  2, Head  9: entropy = 2.195

The heads with lowest entropy are the most selective: they consistently attend to a small number of positions across diverse sentences. These are the heads most worth examining visually, as they likely have the clearest interpretable function.

Identifying the Most Focused Head

Once you have ranked heads by entropy, you can retrieve and display the most focused one to examine what it is attending to. The minimum-entropy head across your test corpus is a strong candidate for having a clear, consistent function.

In[17]:
Code
# Find the most focused head across the dataset
best_layer, best_head = np.unravel_index(
    mean_entropy.argmin(), mean_entropy.shape
)

# Re-run the first sentence through this specific head for display
test_inputs = tokenizer(sentences[0], return_tensors="pt")
with torch.no_grad():
    test_outputs = model(**test_inputs, output_attentions=True)

test_tokens = tokenizer.convert_ids_to_tokens(test_inputs["input_ids"][0])
focused_attn = test_outputs.attentions[best_layer][0, best_head].numpy()
Out[18]:
Visualization
Heatmap of the most focused BERT-base attention head identified by minimum entropy
Attention pattern for the most focused BERT-base head identified by minimum mean entropy across eight sentences. This head consistently concentrates attention on a small number of positions, making its pattern easier to interpret qualitatively. Examining this head across multiple sentences reveals whether the focused attention tracks a stable linguistic property or varies with input content.
Out[19]:
Visualization
Heatmap of mean attention entropy across 12 layers and 12 heads aggregated over multiple sentences
Mean attention entropy aggregated across eight diverse sentences, showing which BERT-base heads are most focused versus most diffuse. The darkest cells identify the most selective heads, while lighter cells show heads whose attention remains more broadly distributed.

Key Parameters

The key parameters in this analysis are:

  • layer_idx: Which transformer layer to examine (0-indexed). Earlier layers capture surface-level patterns; later layers capture more abstract semantic information. For hypothesis generation about syntactic structure, start with layers 4-7. For semantic roles and coreference, start with layers 8-11.
  • head_idx: Which attention head within the layer to examine. Heads are independently parameterized and may develop completely different attention strategies. There is no global convention for what head 0 or head 7 does across models; the identities of specialized heads vary with the random seed, model size, and pre-training data.
  • entropy threshold: A quantitative cutoff for identifying focused heads. Heads below a threshold entropy value are candidates for closer qualitative inspection. A reasonable starting point is to examine heads below the 20th percentile of entropy across your dataset. This typically yields 20-30 candidates in a model like BERT-base.
  • aggregation method: Whether to examine individual sentences or average across a corpus. Single examples reveal per-sentence behavior; corpus averages reveal stable head functions. For claims about model behavior in general, you need corpus-level aggregation. For debugging a specific prediction, single-example visualization is appropriate and useful.

The interaction between these parameters matters. A head may appear unfocused (high entropy) on a general corpus but become sharply focused on domain-specific text. A head that looks syntactic on standard English sentences may lose its syntactic character on code-switching text or non-standard dialects. Attention behavior is a function of both the model and the input distribution.

Attention Across Layers

Patterns in attention change substantially across the depth of the model. Understanding this layer-wise progression is important for knowing where to look when investigating a specific linguistic phenomenon. Early, middle, and late layers tend to serve different purposes, and that structure affects both what you can find and what you should expect to see.

Early Layers: Surface Patterns

Early layers (layers 1-4 in BERT-base) tend to show simpler patterns: strong self-attention, positional attention, or attention toward punctuation tokens. These early layers are building basic representations of word meaning and local context. At this stage, the model is essentially performing a form of contextualized smoothing: blending nearby token representations to create richer local context before attempting to resolve longer-range dependencies.

The positional heads that appear in early layers are not learning syntax. They are learning that adjacent words tend to be relevant to each other, a statistical regularity that holds across almost all natural language. A token like "quickly" is more likely to be relevant to the verb it modifies than to a noun three positions away, and early positional heads capture this general proximity bias before the model has built enough contextual information to make more specific decisions.

Middle Layers: Syntactic Structure

Middle layers (layers 5-8) show more syntactically-organized patterns. It is in these layers that the heads most resembling dependency parsers or part-of-speech taggers tend to appear. The representations at this depth encode sentence structure as well as word meaning. By layer 5, the model has built representations that distinguish "bank" in a financial context from "bank" in a geographical one, and the attention patterns reflect this richer understanding.

The syntactic specialization of middle layers makes intuitive sense given how BERT was trained. The masked language modeling objective forces the model to predict masked words in context. To predict "deposits" correctly when "bank" is visible, the model benefits from understanding the subject-verb-object structure of the sentence. Middle layers appear to be where the model encodes that structural information most explicitly, because it is at this depth that the representations are abstract enough to capture syntax but not yet dominated by task-specific semantics.

Late Layers: Semantic Aggregation

Late layers (layers 9-12) attend more selectively to semantically relevant tokens. In a sentence completion or classification task, these layers are focused on gathering the specific information needed for the output. The [CLS] token, which carries the aggregate sentence representation used for classification, accumulates information aggressively in these layers.

In later layers, you often see lower average entropy because the model is making more decisive choices about which information to aggregate. A [CLS] token in layer 11 that is computing a sentiment representation may strongly attend to sentiment-bearing words like "excellent" or "disappointing," producing a focused, interpretable pattern. But this focus is task-driven: a model fine-tuned for sentiment analysis will develop different late-layer patterns than a model fine-tuned for question answering.

This layered progression is not universal. It was documented primarily in BERT trained with masked language modeling. Decoder-only models like GPT, which always have causal (lower-triangular) attention masks, show different progressions because each token can only attend to preceding tokens. The absence of bidirectional context changes what functions early versus late layers can perform.

We can visualize this layer progression directly by plotting the mean entropy per layer (averaged across all heads and sentences). Layers where entropy is lower contain more selective heads on average; layers where entropy is higher contain broader, more aggregating heads.

Out[20]:
Visualization
Line plot showing mean attention entropy per BERT-base layer, with a shaded band spanning the minimum and maximum head entropy
Mean attention entropy per layer in BERT-base, averaged across all 12 heads and 8 sentences. The entropy profile shows how focused versus diffuse attention is at each depth. Lower entropy layers contain heads that selectively attend to a small number of positions, while higher entropy layers distribute attention more broadly across the sequence.

The shaded band shows the range from the most focused head to the most diffuse head within each layer. Wide bands indicate high diversity among heads in that layer; narrow bands indicate that heads in that layer behave more uniformly.

Causal vs. Bidirectional Attention

A central visualization difference exists between encoder-only models (BERT) and decoder-only models (GPT). BERT uses bidirectional attention: every token can attend to every other token. The attention matrix is fully populated. This means every entry in the attention matrix carries useful information, and you can read patterns in any direction.

GPT-style models use causal masking: token ii can only attend to tokens j≤ij \leq i. The attention matrix is lower-triangular, with the upper triangle set to −∞-\infty before softmax (which maps to exactly zero after softmax). When visualizing GPT attention, you will see this triangular structure, and any interpretation must account for the fact that "future" tokens are invisible to earlier positions. This asymmetry fundamentally changes what patterns are possible: you cannot have a GPT attention head that connects a pronoun back to a later antecedent, because earlier tokens simply cannot attend to later ones.

In[21]:
Code
# Demonstrate causal attention pattern with GPT-2
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer as GPT2Tokenizer

gpt_tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
gpt_model = AutoModelForCausalLM.from_pretrained("gpt2", output_attentions=True)
gpt_model.eval()

gpt_inputs = gpt_tokenizer(
    "The bank can guarantee deposits will", return_tensors="pt"
)
with torch.no_grad():
    gpt_outputs = gpt_model(**gpt_inputs, output_attentions=True)

gpt_tokens = gpt_tokenizer.convert_ids_to_tokens(gpt_inputs["input_ids"][0])
gpt_attn = gpt_outputs.attentions

# Extract one head to show the causal (triangular) mask
gpt_head_attn = gpt_attn[5][0, 0].numpy()
Out[22]:
Visualization
Lower-triangular heatmap showing GPT-2 causal attention where the upper triangle is zero; raw token labels use a G-dot prefix to mark preceding whitespace
Causal attention pattern from GPT-2 layer 6, head 1, illustrating the lower-triangular structure enforced by the causal mask. No token can attend to future tokens, so the upper triangle is always zero. The `Ġ` prefix in GPT-2's raw tokenizer output marks a token that begins after whitespace.

The upper triangle is always zero in GPT-style models. This has a practical implication for visualization: entropy calculations, aggregation, and interpretation must account for the fact that later tokens have access to more context than earlier tokens. The attention distribution for the last token in a long sequence has the richest context; the distribution for the first token attends only to itself.

This positional asymmetry also means that entropy analysis of GPT-style attention requires different baselines than BERT analysis. A token at position 1 can attend to at most 2 positions (itself and position 0), while a token at the last position can attend to all positions. The maximum possible entropy scales with position, so a naive entropy comparison across positions within the same GPT attention matrix is not meaningful without position-dependent normalization.

Attention Flow and Residual Connections

An important subtlety in attention visualization is that attention weights describe how information is routed within a single layer's attention sub-module, but the final output of that layer also includes the residual connection. The residual connection passes the layer's input directly to the layer's output and adds it to the attention sub-module's output. This means the information at any position after a layer is a mixture of what attention routed there and what was already present before the layer.

As a consequence, even if a token places zero attention weight on another token in a given layer, the second token's information is not necessarily absent from the first token's representation after that layer. Information that was gathered by attention in earlier layers is already embedded in the representation coming into the current layer via the residual stream. Attention visualization at a single layer sees only one piece of the information routing story.

This is one reason attention rollout, which propagates information backwards through all layers while accounting for residual connections, provides a more complete picture than single-layer attention heatmaps. The residual connection also partially explains why models are robust to attention head ablation: even if you zero out an entire head's attention, the residual path continues to carry information from earlier layers.

Understanding the residual stream's role leads to a more careful interpretation of attention heatmaps. When you see a token attend weakly to distant positions in a middle layer, do not conclude that the model has not gathered that information. It may have gathered it in earlier layers and carried it forward through the residual stream. What you are seeing in the middle layer is only the incremental routing decision at that depth.

Practical Workflow for Attention Analysis

When you want to use attention visualization to understand a model's behavior on a specific task, a structured workflow helps you avoid the most common pitfalls. The goal is to move systematically from broad exploration to specific, testable hypotheses rather than cherry-picking heatmaps that look interesting.

Start with the head entropy overview. Before examining individual heads, compute entropy across all heads and layers for your dataset. This tells you which heads are the most focused and therefore the most likely to have interpretable, stable patterns. Examining random heads wastes time and invites pareidolia: the tendency to see patterns in noise. Entropy-ranked candidates have earned their place in your analysis.

Examine multiple examples for each head of interest. A single example can be misleading. Collect 20-50 diverse examples and visualize the same head across all of them. If the pattern changes dramatically with different inputs, the head does not have a stable function that attention visualization can reveal. Stability across diverse inputs is a prerequisite for meaningful interpretation.

Form specific hypotheses and test them. Rather than staring at heatmaps looking for inspiration, start with a linguistic question: "Does this head track subject-verb agreement?" Then construct minimal pairs where the subject-verb relation changes, and check whether the head's attention pattern changes accordingly. If the head tracks subjects, the attention from the verb should shift when you swap the subject. If it does not, the head is not tracking subject-verb relations in the way you hypothesized.

Cross-validate with gradient-based methods. If you find an interesting attention pattern, run integrated gradients or attention rollout on the same examples. Attention and gradient methods often agree on the most important tokens, but when they disagree, that disagreement is informative. Divergence suggests that attention is doing something different from what determines the output.

Report aggregated patterns, not cherry-picked examples. When communicating findings, show patterns that hold across many examples. A single well-chosen heatmap can illustrate a concept, but quantitative support requires dataset-level analysis. If your claim is that "head X tracks subject-verb relations," your evidence should include a corpus-level precision and recall analysis against gold dependency parses, not just three examples that support the claim.

Following this workflow turns attention visualization from an exploratory art into a semi-systematic analysis method. You will still discover things by looking at individual heatmaps, but you will have a principled way to evaluate whether those discoveries reflect consistent model behavior or sampling noise.

Limitations and Impact

Attention visualization was the first generation of transformer interpretability tools, and its record is mixed but substantial. Understanding both what it contributed and where it fell short is important for using it appropriately today.

What Attention Visualization Got Right

Attention visualization demonstrated that transformers do not operate as complete black boxes. Their internal computations show structured, sometimes linguistically meaningful patterns that can be examined and categorized. The finding that attention heads specialize was a concrete scientific contribution: it revealed that multi-head attention acts as both a capacity mechanism and a structural mechanism that allows the model to maintain multiple distinct views of the input simultaneously.

The work on identifying syntactic heads contributed to the scientific understanding of what these models learn. Finding heads that align with dependency parses at above-chance rates was evidence that the masked language modeling objective incidentally teaches syntactic structure as a byproduct of learning contextual meaning. This connection between distributional learning and formal syntactic structure was a meaningful result, even if the causal implications were overstated.

Attention visualization also built bridges between NLP interpretability and formal linguistics. Linguists could look at attention patterns and recognize structures they studied. This cross-disciplinary legibility was valuable for building the research community's understanding of what transformers were doing, even if the full picture turned out to be more complex.

Where Attention Visualization Fell Short

The negative side is that attention visualization generated overconfidence. Researchers and practitioners sometimes treated attention weights as reliable explanations, drawing causal conclusions from correlational evidence. "The model made this decision because it attended strongly to that word" became a common but often unjustified claim. This misuse led to direct rebuttals, especially Jain and Wallace's 2019 paper "Attention is not Explanation" and the follow-up debate over whether there are conditions under which attention can serve explanatory purposes.

The field's response has been to develop richer, more causally grounded tools: gradient-based saliency methods, probing classifiers, activation patching, and mechanistic interpretability techniques. These tools are covered in subsequent chapters of this part. Attention visualization remains useful as a quick hypothesis-generation step and for communication with non-expert audiences, but it is now understood as one tool among many rather than a sufficient analysis on its own.

The proliferation of attention analysis papers between 2018 and 2020 also produced a large body of work with inconsistent methodologies. Different papers measured "syntactic alignment" differently, used different evaluation corpora, and drew different conclusions. Reproducing and comparing these results is difficult, and some of the specific claims about which heads perform which functions have not held up to replication across different random seeds or pre-training configurations.

Practical Value Today

From a practical standpoint, attention visualization is most valuable when:

  • Debugging unexpected model behavior (the model gets a specific example wrong; which tokens is it attending to?)
  • Communicating model behavior to stakeholders who need an accessible, visual explanation
  • Forming initial hypotheses that can be tested with more rigorous methods
  • Identifying attention heads that might be candidates for pruning or fine-tuning

It is least reliable when:

  • You want to claim that attention reveals why the model made a specific decision
  • You are comparing attention patterns across models or architectures (attention weights are not on a common scale across models)
  • You are working with very long sequences where the attention matrix becomes too large to interpret visually
  • You need causal rather than correlational evidence about model behavior

The best current practice is to treat attention visualization as a first step in an analysis pipeline, not as a standalone method. Use entropy analysis to identify candidate heads, use heatmaps to form hypotheses, then use probing classifiers or activation patching to test those hypotheses with more controlled experiments.

Summary

Attention visualization extracts the attention weight matrices produced during a forward pass and displays them as heatmaps or arc diagrams. Each transformer layer has multiple heads, each producing an independent n×nn \times n attention matrix for a sequence of nn tokens.

Extracting attention weights requires the output_attentions=True flag in Hugging Face models. The resulting tuple contains one tensor per layer, with shape (batch, heads, seq_len, seq_len). Individual heads can be sliced out and visualized directly.

Common patterns in attention heads include positional attention (attending to adjacent tokens), syntactic heads (tracking grammatical relations), coreference heads (linking pronouns to antecedents), and separator-attending heads. Entropy provides a quantitative measure of how focused a head's attention is, with low-entropy heads tending to perform more interpretable, specialized functions.

Attention weights are not causal explanations. High attention weight does not imply high causal influence on the output. Heads may not be modular, and the same head may show very different patterns on different inputs. Visualization tools like BertViz can mislead when they average across heads or layers without accounting for the fundamental incomparability of attention weights across depths.

The layer-wise progression of attention matters for interpretation. Early layers show positional and local patterns, middle layers encode syntactic structure, and late layers focus on semantically relevant tokens. Encoder-only models like BERT have fully populated attention matrices, while decoder-only models like GPT have lower-triangular matrices due to causal masking.

Residual connections carry information between layers independent of what any single attention head routes, which means single-layer heatmaps capture only part of the information routing story. Attention rollout provides a more complete picture by propagating attention weights backward through all layers while accounting for the residual stream.

These limitations motivate the more rigorous techniques covered in Attention Analysis Limitations and the broader interpretability toolkit developed in subsequent chapters. Attention visualization remains useful as a quick hypothesis-generation step and for communication with non-expert audiences, but it is now understood as one tool among many rather than a sufficient analysis on its own. A structured workflow that combines entropy analysis, multi-example examination, specific hypothesis testing, and cross-validation with gradient-based methods produces more reliable conclusions than isolated heatmap inspection.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about attention visualization.

Attention Visualization Quiz

Question 1 of 80 of 8 completed
For a BERT-base model processing a sentence of 15 tokens, how many attention matrices are produced in a single forward pass?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026attentionvisualization, author = {Michael Brenndoerfer}, title = {Attention Visualization: Extracting and Interpreting Weights}, year = {2026}, url = {https://mbrenndoerfer.com/writing/attention-visualization}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Attention Visualization: Extracting and Interpreting Weights. Retrieved from https://mbrenndoerfer.com/writing/attention-visualization
MLAAcademic
Michael Brenndoerfer. "Attention Visualization: Extracting and Interpreting Weights." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/attention-visualization>.
CHICAGOAcademic
Michael Brenndoerfer. "Attention Visualization: Extracting and Interpreting Weights." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/attention-visualization.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Attention Visualization: Extracting and Interpreting Weights'. Available at: https://mbrenndoerfer.com/writing/attention-visualization (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Attention Visualization: Extracting and Interpreting Weights. https://mbrenndoerfer.com/writing/attention-visualization

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.