BART Architecture: Encoder-Decoder Design for NLP

Michael BrenndoerferJanuary 13, 202557 min read

Part of Language AI Handbook

Covers BART's encoder-decoder architecture combining BERT and GPT designs. Examines attention patterns, model configurations, and implementation details.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

BART Architecture

BART (Bidirectional and Auto-Regressive Transformers) represents one of the most thoughtfully designed models to emerge from the 2019 generation of large pretrained language models. When Facebook AI Research introduced BART, the NLP field was defined by a clear division: encoder-only models like BERT excelled at understanding tasks, applying bidirectional attention to build rich contextual representations, while decoder-only models like GPT dominated text generation, applying causal attention to produce fluent, coherent continuations. Both families had demonstrated remarkable capabilities, but they had also exposed a basic tension. The architectures optimized for understanding were poorly suited to generation, and those optimized for generation could not exploit full bidirectional context. BART asked a simple but powerful question: what if we designed a model that inherently did both?

The answer was to take the encoder-decoder framework, which had been introduced for machine translation in the original "Attention Is All You Need" transformer paper, and adapt it specifically for the pretraining paradigm. Rather than treating the encoder and decoder as novel components designed from scratch, BART's designers made an elegant choice: make the encoder as close to BERT as possible, and make the decoder as close to GPT as possible. Then connect them through cross-attention, and pretrain the resulting model as a denoising autoencoder, teaching it to reconstruct original text from arbitrarily corrupted versions. This combination of established components with a flexible pretraining objective produced a model that achieved state-of-the-art performance on summarization benchmarks and opened new research directions in sequence-to-sequence learning.

What makes BART interesting is its performance and the clarity of its design philosophy. Every architectural choice can be traced back to a principled decision: inherit from BERT for understanding, inherit from GPT for generation, and let the pretraining objective unify them. This transparency makes BART an excellent model for understanding encoder-decoder architectures more broadly. By studying what BART does and why, you gain insight into the design space that all sequence-to-sequence models inhabit. The choices BART makes, which positions it adopts from its predecessors and which it modifies, illuminate the tradeoffs that any practitioner working with language models must eventually consider.

Where T5, which we covered in previous chapters, takes an encoder-decoder approach with a specific span corruption objective, BART uses a more flexible and philosophically distinct approach. T5 reformulates every NLP task as text-to-text mapping and uses a uniform span corruption objective across all tasks. BART, by contrast, directly inherits BERT's bidirectional encoder and GPT's autoregressive decoder, and explores a much richer space of pretraining noise functions, including token masking, deletion, infilling, sentence permutation, and document rotation. These differences matter in practice, leading to distinct strengths for different applications, and understanding them builds your ability to choose the right architecture for a given problem.

This chapter examines BART's architecture in detail: how its encoder and decoder are structured, how the three attention mechanisms work and why each is configured as it is, how BART compares to T5 in design and performance, and what its limitations are. The next chapter covers BART's pretraining objectives in depth. Together, these two chapters complete your understanding of one of the foundational encoder-decoder models in modern NLP.

Historical Context

BART was introduced in October 2019 by Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer at Facebook AI Research, in a paper titled "BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension." The paper appeared at ACL 2020. It was published almost simultaneously with T5 from Google Research, and the two models became the dominant encoder-decoder architectures of the era. BART's design directly reflected the state of the field at the time: BERT had been published roughly a year earlier, GPT-2 had demonstrated the power of large-scale autoregressive generation, and the challenge was to combine their strengths rather than choose between them. BART's contribution was showing that a simple combination of these two architectures, coupled with flexible denoising pretraining, could outperform more complex designs on generation-heavy tasks.

The BART Encoder-Decoder Design

BART follows the encoder-decoder framework we discussed in Part XII and Part XVI, but its design philosophy draws explicitly from BERT and GPT. To appreciate this design choice, consider the basic trade-off in language modeling: understanding requires seeing the full context (including what comes before and after a word), while generation must proceed sequentially since you cannot use words you have not yet produced. These two requirements seem contradictory, yet both are needed for tasks like summarization where you must deeply understand a document before creating a coherent condensed version. A model that processes input bidirectionally can build richer representations, but it cannot generate text one token at a time, because generation requires a strictly left-to-right discipline. A model that processes input left-to-right can generate text naturally, but it builds weaker representations because each token can only incorporate information from tokens to its left.

BART resolves this tension through architectural separation. The encoder is essentially a BERT-style transformer, using bidirectional self-attention over the input sequence and letting each token to attend to all other tokens. This bidirectional view means that when the encoder processes the word "bank" in a sentence, it can simultaneously consider both the preceding context ("walked along the") and the following context ("of the river") to determine that we are discussing a riverbank rather than a financial institution. The decoder, in contrast, is essentially a GPT-style transformer, using causal (left-to-right) self-attention to ensure the model can only use previously generated tokens when predicting the next one. This constraint is not a limitation but a necessity: during generation, future tokens simply do not exist yet, and a model trained to exploit them would fail at inference time.

Think of BART's architecture as two specialist consultants working in sequence. The encoder is a research analyst who reads the entire input document with complete access to all its parts, building a complete, context-aware understanding of the content. The decoder is a writer who can consult the analyst's notes (through cross-attention) but must produce words one at a time without knowing what will come next in the output. The writer's sequential discipline ensures that what gets produced is valid text, and the analyst's complete view ensures that what the writer produces is grounded in the input. The key insight is that separating these two roles into distinct architectural components allows each to be optimized for its particular function, avoiding the compromises that would arise from forcing a single architecture to serve both purposes simultaneously.

The architecture can be summarized as:

BART=BERT Encoder+GPT Decoder\text{BART} = \text{BERT Encoder} + \text{GPT Decoder}

This equation describes the architecture directly. The BART authors explicitly designed the encoder to match BERT's architecture and the decoder to match GPT's, then connected them with cross-attention. The "+" here represents architectural composition: the encoder and decoder are separate components linked through a cross-attention mechanism that allows the decoder to query the encoder's representations as it generates each token.

This design means BART inherits the bidirectional contextual understanding that made BERT successful for classification and extraction tasks, while also gaining the autoregressive generation capabilities that made GPT successful for text generation. The result is a model that can first build a rich, context-aware representation of the input, then use that representation to generate fluent, coherent output. This combination is precisely what tasks such as summarization and translation require. It also supports question-answering.

Encoder Structure

The BART encoder processes the input sequence using standard transformer encoder blocks, transforming raw token embeddings into contextualized representations. Think of each encoder layer as a refinement pass: the first layer turns bare token embeddings into slightly contextualized representations, the second layer enriches those with broader context, and by the final layer each token's representation encodes a sophisticated understanding of its role within the full input sequence. The encoder performs this transformation with complete access to all input positions at every layer, a capability that distinguishes it from the decoder and that makes it capable of resolving complex ambiguities that require long-range context.

These representations capture the meaning of each token within its full surrounding context. Each block contains:

  1. Multi-head self-attention with bidirectional (non-causal) masking. This is the mechanism that allows each token to "see" every other token in the input, gathering information from both directions to build context-aware representations. With multiple attention heads, the model can simultaneously attend to different aspects of the context: one head might focus on syntactic dependencies, another on semantic similarity, and another on positional proximity.

  2. Feed-forward network with GeLU activation. After attention aggregates information across positions, this two-layer neural network transforms each position's representation independently, letting the model to compute complex non-linear functions of the attended information. The feed-forward network is where much of the model's factual knowledge is believed to be stored, based on analysis of what information can be retrieved from its weights.

  3. Residual connections around both sublayers. These skip connections add the input of each sublayer to its output, creating direct gradient pathways that facilitate training deep networks and letting the model to learn incremental refinements rather than complete transformations. Without residual connections, gradients in deep networks tend to vanish or explode, making training unstable.

  4. Layer normalization applied after each sublayer (post-norm). This normalizes the activations to have zero mean and unit variance, stabilizing training by preventing the hidden representations from growing too large or too small as they pass through many layers.

The encoder produces a sequence of hidden states, one for each input token. These states capture rich bidirectional context, as each encoder representation can incorporate information from the entire input sequence. When you pass a sentence through BART's encoder, the representation for each word includes its meaning within the sentence context. This is the foundation upon which the decoder builds its generation: rather than generating from a blank slate, the decoder starts with a complete, contextually-rich picture of the input.

Decoder Structure

The decoder generates output tokens autoregressively, one at a time, using the encoder's representations to understand the input. This process mirrors how a human might write a summary: first, you read and understand the source document (the encoder), then you compose the summary word by word while referring back to the original (the decoder with cross-attention). Each decoder block contains:

  1. Causal self-attention over previously generated tokens. This allows each position to attend only to earlier positions in the output sequence, building up a representation of what has been generated so far. The causal constraint ensures the model cannot "cheat" by looking at tokens it has not yet produced. During training, the model processes the entire target sequence at once, but the mask enforces the discipline of sequential generation: each position can only see positions to its left.

  2. Cross-attention over encoder hidden states. This is the bridge between understanding and generation. It allows the decoder to query the encoder's representations, focusing on different parts of the input as needed for generating each output token. Without cross-attention, the decoder would be a purely unconditional language model with no grounding in the input. With cross-attention, the decoder can dynamically focus on whatever parts of the source are most relevant to the word it is currently generating.

  3. Feed-forward network with GeLU activation. Just as in the encoder, this transforms the combined self-attention and cross-attention information through a non-linear function, computing complex features of the attended representations.

  4. Residual connections around all three sublayers. With three sublayers instead of two, the decoder has even more opportunity to benefit from these direct pathways that preserve information and facilitate gradient flow.

  5. Layer normalization after each sublayer (post-norm). This maintains training stability across the decoder's deeper structure.

The causal masking in self-attention prevents the decoder from "cheating" by looking at future tokens during training. When training on a target sequence like "The study found significant results," the decoder predicting "found" can see "The" and "study" but not "significant" or "results." This ensures that the learned conditional distribution P(tokent∣token1,…,tokent−1)P(\text{token}_t \mid \text{token}_1, \ldots, \text{token}_{t-1}) is trained with exactly the same information available during inference. Cross-attention allows each decoder position to attend to all encoder positions, letting the decoder to ground its generation in the input context. When generating "found," the decoder can look back at the relevant parts of the source document to decide what verb is appropriate.

Key Architectural Choices

BART makes several design decisions that distinguish it from other encoder-decoder models. These choices reflect its heritage from BERT and GPT and affect training dynamics and model behavior in ways that practitioners should understand.

GeLU activation is the first key choice. Unlike T5, which uses ReLU in its original form, BART follows BERT and GPT-2 in using the GeLU (Gaussian Error Linear Unit) activation function in feed-forward layers. The GeLU function is defined as GeLU(x)=x⋅Φ(x)\text{GeLU}(x) = x \cdot \Phi(x), where Φ(x)\Phi(x) is the standard Gaussian cumulative distribution function. In practice, a common approximation is used:

GeLU(x)≈0.5⋅x⋅(1+tanh⁡(2π⋅(x+0.044715x3)))\text{GeLU}(x) \approx 0.5 \cdot x \cdot \left(1 + \tanh\left(\sqrt{\frac{2}{\pi}} \cdot (x + 0.044715 x^3)\right)\right)

where:

  • xx: the pre-activation value at a given neuron
  • Φ(x)\Phi(x): the Gaussian CDF, which determines what fraction of a Gaussian distribution lies below xx
  • The product x⋅Φ(x)x \cdot \Phi(x): scales the input by its probability under a standard normal, smoothly transitioning between near-zero for negative inputs and near-linear for positive inputs

GeLU provides smoother gradients than ReLU because it does not have a sharp transition at zero. Instead, it smoothly interpolates between passing and blocking signals. This smoothness can lead to more stable optimization and has become the preferred activation for many modern language models. Why does this formula make sense? Notice that for very negative xx, Φ(x)≈0\Phi(x) \approx 0, so GeLU outputs near zero, similar to ReLU. For very positive xx, Φ(x)≈1\Phi(x) \approx 1, so GeLU is approximately linear, also like ReLU. The difference is in the transition region near x=0x = 0, where GeLU changes smoothly rather than abruptly.

Post-layer normalization is the second key choice. BART applies layer normalization after residual connections (known as post-norm), following the original transformer design. T5 uses pre-norm, applying layer normalization before each sublayer. The placement of normalization affects gradient flow: pre-norm tends to produce more stable gradients at initialization, which makes it easier to train very deep models. Post-norm can achieve slightly better final performance when training succeeds, but requires more careful tuning of learning rate schedules. This trade-off explains why both conventions persist in practice.

Learned positional embeddings are the third key choice. BART uses learned absolute position embeddings, similar to BERT and GPT-2, rather than the relative position encodings used in T5. With learned embeddings, the model maintains a separate embedding vector for each position (1, 2, 3, and so on up to some maximum), and these vectors are learned during training just like token embeddings. This approach is simple and effective but creates a hard limit on sequence length: the model has no embedding for position 1025 if it was trained with a maximum of 1024.

No parameter sharing is the fourth key choice. Unlike T5, BART does not tie encoder and decoder embeddings by default. The encoder and decoder maintain separate embedding matrices. This increases parameter count but allows the encoder and decoder to learn specialized representations suited to their different roles. The encoder's embeddings are tuned for building contextual representations of input text, while the decoder's embeddings are tuned for the language modeling objective of predicting the next token. These roles are related but distinct, and separate embedding matrices give the model the freedom to specialize each.

Attention Configuration

BART's attention patterns show how information flows through the model. Attention is the mechanism by which transformers route information between positions, and the pattern of allowed attention, meaning which positions can attend to which other positions, fundamentally shapes what the model can learn and compute. There are three distinct attention mechanisms in BART, each with its own configuration and purpose: bidirectional self-attention in the encoder, causal self-attention in the decoder, and cross-attention connecting the decoder to the encoder. Understanding these three patterns and why each is configured the way it is gives you a deep understanding of how information moves through the model.

Think of the attention mechanisms as a communication network within the model. Bidirectional encoder attention is like a town hall meeting where everyone can hear everyone else simultaneously. This allows rich information exchange. Causal decoder attention is like a relay race where each runner can only receive information from those who ran before. This keeps valid sequential generation. Cross-attention is like a telephone line from the decoder to the encoder. This allows the decoder to query the encoder's representations at any moment during generation. Together, these three communication patterns implement the full pipeline from input understanding to output generation.

Encoder Self-Attention

In the encoder, every token can attend to every other token. This bidirectional attention gives the encoder the power to build contextual representations, because a word's meaning often depends on context that appears both before and after it. Consider disambiguating "The bank was eroding" versus "The bank was closing": you need both the subject and the verb to understand which sense of "bank" is intended. With bidirectional attention, the token "bank" can simultaneously look at "eroding" or "closing" to resolve this ambiguity. In a unidirectional model, "bank" would have to represent its meaning before seeing the disambiguating word that follows it.

To compute bidirectional self-attention, we construct a mask that permits all token pairs. For an input sequence of length nn, the attention mask is an n×nn \times n matrix of ones:

Mencoder=[11⋯111⋯1⋮⋮⋱⋮11⋯1]\mathbf{M}_{\text{encoder}} = \begin{bmatrix} 1 & 1 & \cdots & 1 \\ 1 & 1 & \cdots & 1 \\ \vdots & \vdots & \ddots & \vdots \\ 1 & 1 & \cdots & 1 \end{bmatrix}

where:

  • Mencoder\mathbf{M}_{\text{encoder}}: the attention mask matrix of shape n×nn \times n
  • nn: the length of the input sequence
  • Each entry of 1 indicates that attention is permitted between that query-key pair; a 0 would block that pair

Reading this matrix, row ii describes which positions token ii can attend to: a 1 in column jj means token ii can attend to token jj. Since every entry is 1, every token can attend to every other token, including itself. This complete connectivity allows information to flow freely through the sequence, letting the encoder to build representations that incorporate arbitrarily distant context.

Why does this formula make sense? Notice that the matrix is symmetric: if token ii can attend to token jj, then token jj can also attend to token ii. This symmetry reflects the bidirectional nature of the encoder: attention flows in both directions simultaneously. The uniformity of the matrix, with every entry equal to 1, reflects the fact that there are no structural constraints on encoder attention. Any token might be relevant to any other token, and the model learns through training which connections are useful.

This bidirectional attention is what gives encoder-only models like BERT their power for understanding tasks. Each token's representation is informed by the complete context, not just preceding tokens. A word at the beginning of a sentence can be influenced by words at the end, and vice versa, letting the rich contextual representations that make BERT effective for tasks like sentiment analysis and named entity recognition. BART inherits this power in its encoder, giving it the same contextual understanding capability that made BERT so successful.

Decoder Causal Self-Attention

The decoder uses causal (or autoregressive) masking in its self-attention layers. The word "causal" refers to the structure of language generation: each token is caused by (depends on) the tokens that came before it, not those that come after. This constraint is not artificial. During generation, future tokens do not yet exist. A model that attended to future tokens during training would be learning a task that is fundamentally different from the task it faces at inference time, where future tokens are simply unavailable.

To enforce this constraint, we use a lower-triangular mask. For a sequence of length mm, the attention mask is:

Mdecoder=[100⋯0110⋯0111⋯0⋮⋮⋮⋱⋮111⋯1]\mathbf{M}_{\text{decoder}} = \begin{bmatrix} 1 & 0 & 0 & \cdots & 0 \\ 1 & 1 & 0 & \cdots & 0 \\ 1 & 1 & 1 & \cdots & 0 \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ 1 & 1 & 1 & \cdots & 1 \end{bmatrix}

where:

  • Mdecoder\mathbf{M}_{\text{decoder}}: the causal attention mask matrix of shape m×mm \times m
  • mm: the length of the decoder sequence (the number of tokens generated so far)
  • Entry (i,j)=1(i, j) = 1 if j≤ij \leq i, meaning position ii can attend to position jj
  • Entry (i,j)=0(i, j) = 0 if j>ij > i, meaning position ii cannot attend to the future position jj

The lower-triangular structure emerges directly from the causality constraint. Position 1 can only attend to position 1 (itself), so the first row has a single 1. Position 2 can attend to positions 1 and 2, giving two 1s. Position ii can attend to all positions from 1 through ii, creating the triangular pattern. The zeros above the diagonal represent the blocked attention to future positions, which is information that the model must not use because it will not be available at generation time.

Why does this formula make sense? Consider the diagonal of the mask: every position can attend to itself, which makes sense because a token's representation should at minimum encode information about the token itself. The lower triangle means that as we move to later positions in the sequence, more context becomes available: position 1 has only itself, position 2 has itself and position 1, and so on. This is precisely the structure of autoregressive generation, where each new token is produced with access to all previously generated tokens.

This ensures that when generating token tt, the model can only attend to tokens at positions 1,2,…,t1, 2, \ldots, t. This causal structure makes autoregressive generation possible. During training, we can compute attention for all positions in parallel (the mask handles the constraints), but the model learns to predict each token using only its predecessors, exactly as it will during generation. This training efficiency, combined with the discipline of causal masking, is one of the key innovations of transformer-based language models over earlier recurrent architectures.

Cross-Attention

Cross-attention connects the decoder to the encoder, bridging the understanding and generation phases. Without cross-attention, the decoder would not know what input to generate output for. It would become an unconditional language model, generating plausible text without grounding in any specific input. With cross-attention, the decoder can dynamically consult the encoder's representations, focusing on whatever parts of the input are most relevant for generating the current output token.

The mechanism of cross-attention is identical to standard scaled dot-product attention, but with a important difference: the queries come from the decoder, while the keys and values come from the encoder. This asymmetry is meaningful. The decoder is asking questions (represented as queries) about the input, and the encoder provides the answers (as values) along with ways to determine relevance (through keys). For each decoder position, queries come from the decoder hidden states while keys and values come from the encoder hidden states. There is no masking in cross-attention: every decoder position can attend to every encoder position.

We compute cross-attention as:

CrossAttn(Qdec,Kenc,Venc)=softmax(QdecKencTdk)Venc\text{CrossAttn}(\mathbf{Q}_{\text{dec}}, \mathbf{K}_{\text{enc}}, \mathbf{V}_{\text{enc}}) = \text{softmax}\left(\frac{\mathbf{Q}_{\text{dec}} \mathbf{K}_{\text{enc}}^T}{\sqrt{d_k}}\right) \mathbf{V}_{\text{enc}}

where:

  • Qdec\mathbf{Q}_{\text{dec}}: query vectors derived from decoder hidden states, shape (m,dk)(m, d_k). These represent "what the decoder is looking for" at each position.
  • Kenc\mathbf{K}_{\text{enc}}: key vectors derived from encoder hidden states, shape (n,dk)(n, d_k). These represent "what information each encoder position offers" and are compared against queries.
  • Venc\mathbf{V}_{\text{enc}}: value vectors derived from encoder hidden states, shape (n,dv)(n, d_v). These contain the actual information that will be retrieved and combined.
  • dkd_k: the dimension of query and key vectors, used for scaling
  • dvd_v: the dimension of value vectors
  • dk\sqrt{d_k}: a scaling factor that prevents dot products from growing too large, which would push softmax into regions with vanishing gradients

The computation proceeds in three stages. First, the dot product QdecKencT\mathbf{Q}_{\text{dec}} \mathbf{K}_{\text{enc}}^T computes a compatibility score between each decoder position and each encoder position, resulting in an m×nm \times n matrix of raw attention scores. Second, the softmax operation converts these scaled scores into attention weights that sum to 1 across encoder positions, creating a proper probability distribution over the input that indicates how much attention each output position pays to each input position. Third, these weights are used to compute a weighted combination of encoder values, where positions with higher attention weights contribute more to the final representation.

Why does this formula make sense? The dot product QdecKencT\mathbf{Q}_{\text{dec}} \mathbf{K}_{\text{enc}}^T measures alignment between what the decoder is looking for (queries) and what each encoder position offers (keys). High scores mean the decoder's needs align well with what that encoder position provides. After scaling and softmax, these alignment scores become a probability distribution that the decoder uses to blend together the encoder's value vectors. The result is a context vector: a weighted mixture of all encoder outputs, where the weights reflect relevance to the current decoding step.

This lets the decoder focus on different parts of the input as it generates each output token. When summarizing a document, the decoder might focus on the introduction when generating the first sentence of the summary, then shift attention to specific details when describing particular findings, and attend to the conclusion when wrapping up. This dynamic alignment is what makes encoder-decoder models so effective for conditional generation tasks, a mechanism we explored thoroughly in our discussion of Bahdanau and Luong attention in Part XII.

Attention Flow Visualization

To understand how information flows through BART, consider processing an input sequence and generating an output. The flow follows a clear two-phase structure that separates understanding from generation while connecting them through cross-attention. The encoder and decoder do not operate completely independently: the encoder runs first, creating a complete set of hidden states, and the decoder then uses those hidden states at every layer as it generates output tokens one by one.

The encoder phase processes the entire input once. Input tokens are transformed into embeddings, positional embeddings are added, and the combined representations are passed through LL encoder layers. At each layer, bidirectional self-attention allows information to flow freely between all positions. By the end of the encoder, each token's representation includes context from the entire input. This phase processes the entire input once, creating a fixed set of representations that the decoder will query during generation. The computational cost of this phase is O(n2d)O(n^2 d) where nn is the sequence length and dd is the hidden dimension, dominated by the self-attention operation.

The decoder phase generates one token at a time. For each output token, the decoder performs a sequence of operations that integrate three sources of information:

  • Causal self-attention integrates information from previously generated tokens, building a representation of the output generated so far. This allows the decoder to maintain coherence across the output sequence: later tokens can attend to earlier ones. This keeps the generated text flows logically.
  • Cross-attention retrieves relevant information from the encoder, grounding the generation in the input. At each step, the decoder issues queries against all encoder hidden states and receives a weighted combination of their values as a context signal.
  • The feed-forward network transforms the combined representation, computing complex functions of the attended information.
  • The output projection produces a probability distribution over the vocabulary, from which the next token is selected.

This two-phase structure is efficient for tasks with long inputs and shorter outputs. The expensive encoder computation happens once, and the decoder reuses those representations for every generated token. If generating a 100-word summary from a 1000-word document, the encoder runs once over 1000 tokens, and the decoder runs 100 times, each time performing cross-attention over the 1000 encoder states. The encoder's computation is amortized across all generated tokens.

Out[3]:
Visualization
Heatmap showing encoder bidirectional attention pattern with all ones.
Encoder bidirectional attention mask where every token attends to all others, letting full contextual representation building.
Heatmap showing decoder causal attention with lower-triangular pattern.
Decoder causal self-attention mask with lower-triangular structure, enforcing the left-to-right discipline required for autoregressive generation.
Heatmap showing cross-attention with full connectivity.
Cross-attention mask showing full connectivity from decoder to encoder, letting the decoder to query any encoder position at each generation step.

BART vs T5 Comparison

BART and T5 were developed around the same time and share the encoder-decoder architecture, but they differ in key ways that reflect fundamentally different design philosophies. T5 was designed around the principle of maximum unification: by reformulating every NLP task as text-to-text mapping and using a single uniform pretraining objective (span corruption), the T5 authors aimed to create a general-purpose framework applicable to any task without architectural modifications. BART was designed around a different principle: maximum fidelity to proven components. By directly inheriting BERT's encoder and GPT's decoder, the BART authors accepted greater architectural specificity in exchange for the ability to use the extensive knowledge already accumulated about those architectures' behavior and optimization. Understanding these differences helps you choose the right model and reveals the design space of encoder-decoder models.

The differences between BART and T5 are not arbitrary: each choice in BART can be traced to a corresponding choice in BERT or GPT, and each choice in T5 can be traced to the text-to-text unification principle. When you see that BART uses post-norm while T5 uses pre-norm, you are seeing the difference between BERT's original design and T5's architectural innovation for stability. When you see that BART uses BPE tokenization while T5 uses SentencePiece, you are seeing BART inherit GPT-2's tokenizer versus T5 developing its own. These differences accumulate into models with distinct strengths and weaknesses in practice.

Architectural Differences

The table below summarizes the key architectural distinctions between BART and T5:

Architectural comparison between BART and T5.
ComponentBARTT5
Activation functionGeLUReLU (original) / GeGLU (v1.1)
NormalizationPost-normPre-norm
Position encodingLearned absoluteRelative (bucketed)
Embedding sharingSeparateTied (encoder-decoder)
VocabularyBPE (GPT-2 tokenizer)SentencePiece

Pre-norm versus post-norm affects training stability in important ways. Pre-norm normalizes the input to each sublayer before the sublayer's computation. This keeps the residual connection adds a well-scaled update to a normalized representation. Post-norm normalizes after adding the residual connection, which means the residual pathway can accumulate unnormalized updates across layers. Pre-norm tends to produce more stable gradients at initialization. This enables training of very deep models with fewer numerical difficulties. Post-norm models often achieve slightly better final performance when training is successful, because the normalization at the output of each sublayer constrains representations more tightly. The practical effect is that most modern large language models, from T5 onward, have adopted pre-norm because it makes training more predictable, especially at large scale.

BART's use of learned absolute position embeddings means it has a fixed maximum sequence length, typically 1024 tokens, while T5's relative position encoding theoretically generalizes to longer sequences. With learned absolute embeddings, each position from 1 to 1024 has a dedicated embedding vector that the model learns during training. A sequence of length 1025 would require a position-1025 embedding that does not exist. T5's relative position encoding computes compatibility between positions based on their relative distance rather than their absolute positions, and uses a small number of learned bias terms corresponding to distance buckets. This approach can extrapolate to distances not seen during training, though with some performance degradation. In practice, both models require additional techniques for handling very long contexts, as we covered in Part XVIII.

The vocabulary difference is also practically significant. BART inherits GPT-2's byte-pair encoding (BPE) tokenizer with a vocabulary of approximately 50,000 tokens, while T5 uses SentencePiece with a vocabulary of 32,000 tokens. BPE operates on bytes and can represent any Unicode text without unknown tokens, which is particularly valuable for handling code and other unusual inputs. SentencePiece provides language-agnostic subword tokenization and is somewhat more compact. These tokenization differences affect how each model handles out-of-vocabulary words, multilingual inputs, and special character sequences.

Input-Output Formatting

The most noticeable practical difference between BART and T5 is how input and output are structured during fine-tuning. This reflects the two models' different design philosophies about task generalization.

T5 uses a text-to-text format where every task is framed as mapping an input text to an output text. Tasks are specified through prefixes that tell the model what operation to perform:

summarize: The researchers conducted experiments... → The study found... translate English to German: Hello world → Hallo Welt

This uniformity is elegant: the model architecture and loss function are identical for every task, and tasks can even be mixed during fine-tuning. The task prefix teaches the model to condition its behavior on the instruction, prefiguring the instruction-following capability that later became central to large language model development.

BART treats tasks more naturally by using its encoder-decoder separation. For sequence-to-sequence tasks like summarization, the input goes to the encoder and the output comes from the decoder. No task prefix is needed because the architecture itself implies the operation: encode the source, decode the target. For classification tasks, a special token's representation from the decoder is used for prediction, which is closer to how BERT handles classification with its [CLS] token. This means BART can handle classification without reformulating it as text generation, an efficiency advantage for discriminative tasks.

The practical implication is that T5's text-to-text format is more flexible for multi-task learning and instruction following, while BART's format is more natural for generation tasks and slightly more efficient for classification. Neither is universally superior: the choice depends on the application, available computational resources, and the specific mix of tasks you need to support.

Pre-training Objectives

The pretraining objectives of BART and T5 differ in both philosophy and practical effect. While we will cover BART's pretraining in detail in the next chapter, the high-level contrast is important for understanding why these models behave differently in practice.

T5 uses span corruption, replacing random spans of tokens with sentinel tokens (like <extra_id_0>, <extra_id_1>, etc.) and training the model to generate the missing spans in sequence. The key properties of span corruption are its computational efficiency: since the target sequence contains only the masked spans (not the entire input), each training example requires generating far fewer tokens than the input length. For a 512-token input where 15% of tokens are masked in spans of average length 3, the target is roughly 25 tokens, making each training step roughly 20 times more efficient in terms of decoding computation.

BART explores multiple noise functions and trains the model to reconstruct the full original document. The target sequence is the complete uncorrupted text, which is the same length as the input. This means each BART pretraining step requires generating significantly more tokens than the corresponding T5 step. However, this more demanding objective provides a stronger training signal for generation tasks: the model must fill in masked regions and reconstruct the complete coherent text. The noise functions BART explores, including token deletion (which forces the model to determine how many tokens are missing), text infilling (which forces it to generate multiple tokens in place of a single mask), and sentence permutation (which forces it to reorder scrambled sentences), each train different aspects of text reconstruction ability.

T5's span corruption is more computationally efficient because the target sequence is much shorter than the input. BART's document reconstruction means the target is the same length as the uncorrupted input, requiring more computation per training step but giving a richer learning signal for generation-heavy downstream tasks. For applications where generation quality is important, such as abstractive summarization or dialogue generation, BART's more demanding pretraining objective tends to pay off at fine-tuning time.

Performance Trade-offs

In benchmarks, both models show strong performance with different strengths that reflect their pretraining objectives and architectural choices:

  • Summarization: BART tends to perform better, particularly on abstractive summarization benchmarks like CNN/DailyMail and XSum. The reason is that BART's pretraining requires generating coherent, full-length text rather than filling in short spans, which directly exercises the skills needed for summarization.
  • Translation: Performance is similar, with both models achieving strong results when fine-tuned on parallel data. For low-resource translation where fine-tuning data is scarce, the pretraining differences become more pronounced.
  • Question answering: Both perform well, with specific results depending on the dataset and fine-tuning setup. For extractive QA, BERT-style encoder models remain competitive. For abstractive or open-domain QA, encoder-decoder models have an advantage.
  • Classification: BART can perform classification by using the decoder's representation of a special token, though encoder-only models remain competitive for pure classification tasks where generation is not required.

The practical difference often comes down to implementation details and fine-tuning approach rather than architecture alone. For most practitioners, the availability of pretrained model weights, the quality of existing fine-tuned checkpoints for the target task, and compatibility with the application's requirements will matter more than the raw architectural differences.

BART Model Sizes

BART was released in two primary sizes, following the naming convention established by BERT. This two-tier approach was standard for 2019-era models: a "base" configuration suitable for research and resource-constrained applications, and a "large" configuration that maximizes performance at the cost of greater computational requirements. Understanding the differences between these sizes and how they compare to contemporary models helps you calibrate the computational demands of working with BART.

The design principle behind BART's sizing is that the base model should be large enough to demonstrate the architecture's capabilities convincingly, while remaining tractable for research experiments on single GPUs. The large model should push capabilities further while remaining feasible for fine-tuning on multi-GPU research servers. Both sizes use the same architectural blueprint, differing only in depth (number of layers) and width (hidden dimension size), a standard practice for scaling transformer models.

BART-base

The base model provides a good balance between capability and computational requirements. Its configuration reflects the original BERT-base design, with 6 encoder layers and 6 decoder layers instead of BERT's 12 layers. This reduction is necessary because BART's decoder adds approximately the same number of parameters as the encoder, so a 12+12 layer BART would be much larger than BERT-base. By halving the depth to 6+6, BART-base achieves a total parameter count similar to BERT-base and GPT-2 Small, which makes it comparable in cost to those models.

  • Encoder layers: 6
  • Decoder layers: 6
  • Hidden dimension: 768
  • Attention heads: 12
  • Feed-forward dimension: 3072
  • Parameters: approximately 140 million

BART-base works well when computational resources are limited or fast inference is required. It can run on consumer GPUs and provides strong performance on tasks such as summarization and question answering. For development and experimentation, starting with BART-base is advisable: it converges faster during fine-tuning and allows you to iterate on hyperparameters before committing to the more expensive BART-large.

BART-large

The large model increases capacity significantly by doubling both depth and width. Each of the 12 encoder layers and 12 decoder layers contains wider attention heads and larger feed-forward networks, compounding the capacity increase. The result is a model with nearly three times the parameters of BART-base.

  • Encoder layers: 12
  • Decoder layers: 12
  • Hidden dimension: 1024
  • Attention heads: 16
  • Feed-forward dimension: 4096
  • Parameters: approximately 400 million

BART-large performs better on most benchmarks, particularly for complex generation tasks requiring fine-grained understanding of long documents. It requires more memory and computation, typically requiring a GPU with at least 16GB of memory for fine-tuning. For production summarization systems where quality is important and computational cost is acceptable, BART-large is generally the preferred choice.

The hidden dimension increase from 768 to 1024 has a cascading effect on all components: the attention key, query, and value dimensions scale up, the feed-forward dimensions scale proportionally, and the position embeddings become larger. The number of attention heads increases from 12 to 16, and since the hidden dimension also grows, each individual attention head becomes slightly wider as well (64 dimensions per head in both base and large, since 1024/16=641024 / 16 = 64). This consistency in per-head dimension is intentional: keeping the per-head dimension constant ensures that the scaled dot-product attention mechanism has consistent properties across model sizes.

Comparison with Other Models

The following table contextualizes BART's sizes relative to related models:

Parameter counts and layer configurations for BART and related models.
ModelParametersEncoder LayersDecoder Layers
BERT-base110M12-
BERT-large340M24-
GPT-2 Small124M-12
GPT-2 Medium355M-24
BART-base140M66
BART-large400M1212
T5-base220M1212
T5-large770M2424

BART's parameter count is somewhat lower than T5 for the same size label because T5 uses more layers. BART-large with 12+12 layers is closer to T5-base with 12+12 layers in terms of depth, though T5 uses parameter-efficient relative position encodings while BART has separate learned embeddings for each position. The naming conventions across model families are inconsistent: a "base" BART is comparable in depth to a "base" T5, but BART-large is not comparable to T5-large in terms of parameter count. When comparing models across families, examining layer counts and hidden dimensions directly is more informative than relying on size labels.

Out[4]:
Visualization
Horizontal bar chart comparing parameter counts of BERT, GPT-2, BART, and T5 models.
Parameter counts for encoder-decoder and related models sorted by total parameters. BART-base and BART-large sit between BERT and T5 in terms of total parameters, which reflects their different layer configurations and the overhead of maintaining both encoder and decoder components.

Worked Example: Tracing a Forward Pass

To make the architecture concrete, let's trace a short input through BART's encoder and decoder step by step. This numerical walkthrough illustrates exactly how each component turns representations and where the three attention mechanisms come into play.

Suppose we want to summarize the sentence "Scientists discovered a new planet." We want BART to generate the summary "New planet found." Let's trace the processing at a conceptual level, using simplified numbers to illustrate the key transformations.

Step 1: Tokenization. The tokenizer first splits the input into BPE subword tokens. For our sentence, the result might be: <s>, Scientists, discovered, a, new, planet, ., </s>. Call this sequence of 8 tokens x=[x1,x2,x3,x4,x5,x6,x7,x8]\mathbf{x} = [x_1, x_2, x_3, x_4, x_5, x_6, x_7, x_8]. Similarly, the target output tokenizes as: <s>, New, planet, found, ., </s>, giving 6 tokens y=[y1,y2,y3,y4,y5,y6]\mathbf{y} = [y_1, y_2, y_3, y_4, y_5, y_6].

Step 2: Input embeddings. Each input token xix_i is looked up in BART's embedding table to produce an embedding vector ei∈R768\mathbf{e}_i \in \mathbb{R}^{768} (for BART-base). Then the corresponding positional embedding pi∈R768\mathbf{p}_i \in \mathbb{R}^{768} for position ii is added, giving the initial encoder input:

hi(0)=ei+pi\mathbf{h}_i^{(0)} = \mathbf{e}_i + \mathbf{p}_i

where:

  • hi(0)\mathbf{h}_i^{(0)}: the initial representation for token ii, combining token identity and position
  • ei\mathbf{e}_i: the token embedding for xix_i, encoding the token's meaning
  • pi\mathbf{p}_i: the positional embedding for position ii, encoding the token's location in the sequence

After this step, we have 8 vectors, each of dimension 768, representing the combined token-and-position information for each input token. These are the inputs to the first encoder layer.

Step 3: Encoder processing. The 8 vectors pass through 6 encoder layers. In each layer, multi-head self-attention allows all 8 tokens to interact with each other simultaneously. Consider what happens to the token "planet" at position 6. In the first encoder layer, "planet" can already attend to "Scientists" and "discovered," picking up contextual signals that this is a scientific discovery context. In later layers, its representation becomes increasingly contextualized, incorporating information about the full sentence structure. After 6 layers, the encoder produces 8 hidden state vectors hi(L)∈R768\mathbf{h}_i^{(L)} \in \mathbb{R}^{768}, one for each input token, where each vector encodes that token's meaning in the full sentence context.

Step 4: Decoder initialization. The decoder starts with the beginning-of-sequence token <s> as its first input. This is mapped to an embedding and positional embedding just as in the encoder, giving an initial decoder hidden state d1(0)∈R768\mathbf{d}_1^{(0)} \in \mathbb{R}^{768}.

Step 5: Generating the first output token. The decoder processes d1(0)\mathbf{d}_1^{(0)} through 6 decoder layers. In each layer, three operations occur. First, causal self-attention: since we only have one decoder token so far, this simply returns the token's own representation (trivially causal with one token). Second, cross-attention: the decoder issues a query against all 8 encoder hidden states, computing compatibility scores for each. The decoder's query asks, in effect, "given that I'm about to generate the first word of the summary, which encoder states are most relevant?" The cross-attention mechanism computes a weighted combination of encoder values, likely focusing heavily on the "new planet" part of the input since that is the most content-bearing phrase. Third, the feed-forward network processes the combined representation. After all 6 layers, the decoder projects the final hidden state onto the vocabulary dimension (50,265 dimensions for BART's BPE vocabulary) and applies softmax to get a probability distribution. The highest-probability token is selected (assuming greedy decoding): "New."

Step 6: Generating subsequent tokens. The decoder now has two tokens: <s> and "New." It processes both through the decoder layers, with causal self-attention letting "New" to attend to <s> but not vice versa. Cross-attention again queries the encoder, but now from the perspective of generating the second word given that "New" was already generated. The model selects "planet" as the next token. This process continues, generating "found," ".", and finally </s>, at which point generation stops.

The key insight from this trace is that the encoder runs once and the decoder runs once per output token. For our 6-token output, the cross-attention mechanism queries the same 8 encoder hidden states 6 times, each time from a different perspective as the output sequence grows. The encoder's computation is reused across all generation steps, which is the basic efficiency advantage of the encoder-decoder architecture over pure decoder architectures for conditional generation.

Code Implementation

Let's explore BART's architecture using the Hugging Face Transformers library. We'll examine the model structure, inspect attention patterns, and see how the encoder and decoder interact. The goal is to verify the architectural details we've discussed and develop intuition for how BART behaves in practice.

In[8]:
Code
from transformers import BartConfig, BartModel, BartTokenizer

# Load BART-base model and tokenizer with attention output enabled in config
tokenizer = BartTokenizer.from_pretrained("facebook/bart-base")
config = BartConfig.from_pretrained(
    "facebook/bart-base", output_attentions=True
)
model = BartModel.from_pretrained("facebook/bart-base", config=config)
model.eval()

First, let's examine the model configuration to confirm the architectural details we discussed:

In[10]:
Code
# Store configuration for inspection
model_config = model.config

The configuration confirms our discussion: 6 encoder and decoder layers, 768-dimensional hidden states, 12 attention heads, and GeLU activation. The maximum position embeddings of 1024 reflects BART's use of learned absolute position encoding, matching the architectural specifications from the BART paper.

Now let's pass an example through the model and examine the outputs:

In[13]:
Code
import torch

# Prepare input
input_text = "BART is a denoising autoencoder for pretraining sequence-to-sequence models."
inputs = tokenizer(input_text, return_tensors="pt")

# Create decoder input (shifted right, starting with BOS token)
decoder_input_ids = torch.tensor([[tokenizer.bos_token_id]])
In[14]:
Code
# Forward pass through the full model
with torch.no_grad():
    outputs = model(
        input_ids=inputs["input_ids"],
        attention_mask=inputs["attention_mask"],
        decoder_input_ids=decoder_input_ids,
        output_attentions=True,
    )

The output shapes reveal the flow of information. The encoder produces hidden states for each input token, while the decoder produces hidden states for each position in the output sequence (currently just one, for the BOS token). The number of attention outputs equals the number of layers (6 for BART-base), one set per layer.

Let's visualize the attention patterns from the last layer of each component:

In[17]:
Code
# Extract attention weights (last layer, first head)
encoder_attn = (
    outputs.encoder_attentions[-1][0, 0].detach().numpy()
)  # [seq_len, seq_len]
decoder_self_attn = (
    outputs.decoder_attentions[-1][0, 0].detach().numpy()
)  # [1, 1]
cross_attn = outputs.cross_attentions[-1][0, 0].detach().numpy()  # [1, seq_len]

# Get tokens for labeling
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
Out[5]:
Visualization
Heatmap of encoder self-attention weights showing bidirectional attention pattern.
Encoder self-attention weights from the last layer of BART-base (simulated). Each cell shows how much attention token (row) pays to token (column). The diffuse, bidirectional pattern reflects the encoder's ability to integrate context from all positions simultaneously, with slight diagonal emphasis showing each token also attends to itself.

The encoder attention pattern shows the bidirectional nature of BART's encoder. Each token can attend to all other tokens, with the model learning to focus on contextually relevant positions. Notice that attention is not uniform: the model concentrates attention on positions that provide useful contextual information, such as subword tokens that together form a single word attending strongly to each other.

Now let's examine cross-attention by generating a longer output sequence:

In[20]:
Code
import torch
from transformers import BartConfig, BartForConditionalGeneration

# Load model with config that enables attention outputs
config_gen = BartConfig.from_pretrained(
    "facebook/bart-base", output_attentions=True
)
model_gen = BartForConditionalGeneration.from_pretrained(
    "facebook/bart-base", config=config_gen
)
model_gen.eval()

# We'll manually decode a few steps to capture attention
decoder_ids = [tokenizer.bos_token_id]
cross_attentions_per_step = []

with torch.no_grad():
    for step in range(5):
        decoder_input = torch.tensor([decoder_ids])
        gen_outputs = model_gen(
            input_ids=inputs["input_ids"],
            attention_mask=inputs["attention_mask"],
            decoder_input_ids=decoder_input,
            output_attentions=True,
        )

        # Get cross-attention from last layer, first head
        cross_attn_step = (
            gen_outputs.cross_attentions[-1][0, 0, -1, :].detach().numpy()
        )
        cross_attentions_per_step.append(cross_attn_step)

        # Greedy decoding for next token
        next_token = gen_outputs.logits[0, -1, :].argmax().item()
        decoder_ids.append(next_token)

# Stack cross-attentions
import numpy as np

cross_attn_matrix = np.stack(cross_attentions_per_step)
generated_tokens = tokenizer.convert_ids_to_tokens(decoder_ids[1:])  # Skip BOS
Out[6]:
Visualization
Heatmap of cross-attention from decoder to encoder positions.
Cross-attention weights showing how each generated token attends to different encoder positions (simulated). Each row corresponds to a generated token, and each column corresponds to an encoder input position. The peaked attention pattern shows the model focusing on semantically corresponding input spans when generating each output token.

The cross-attention visualization reveals how the decoder grounds its generation in the input. Each row shows which encoder positions a generated token attended to. This shows the dynamic alignment between input and output. Tokens like "BART" generated first focus attention on the corresponding "BART" in the input, while later tokens like "sequence" attend to the "sequence" span in the encoder input. This alignment is not hardcoded: it emerges from the model learning through training which input positions are relevant for generating each output position.

Let's also count parameters to verify the model sizes we discussed:

In[23]:
Code
def count_parameters(model):
    """Count trainable parameters in model components."""
    encoder_params = sum(p.numel() for p in model.model.encoder.parameters())
    decoder_params = sum(p.numel() for p in model.model.decoder.parameters())
    embed_params = sum(p.numel() for p in model.model.shared.parameters())
    lm_head_params = sum(p.numel() for p in model.lm_head.parameters())
    total_params = sum(p.numel() for p in model.parameters())
    return (
        encoder_params,
        decoder_params,
        embed_params,
        lm_head_params,
        total_params,
    )

The parameter count confirms that BART-base has approximately 140 million parameters, split roughly evenly between the encoder and decoder with significant additional parameters in the embedding matrices. This aligns with our earlier discussion of BART model sizes and demonstrates how the encoder-decoder architecture distributes capacity across both components.

Out[7]:
Visualization
Pie chart showing BART-base parameter distribution across components.
Distribution of parameters across BART-base components. The encoder and decoder contain nearly identical parameter counts (each about 60.6M), which reflects the symmetric depth of 6 layers each. Shared embeddings account for roughly 14% of total parameters, a consequence of the large BPE vocabulary size.

Key Parameters

Understanding BART's configuration parameters is important for adapting the model to new tasks and for interpreting the behavior of different model variants. Each parameter controls a specific aspect of the architecture's capacity and behavior.

The key parameters for BART's architecture are:

  • d_model: The hidden dimension size, 768 for base and 1024 for large. This is the basic width of the model: it determines the dimensionality of all token representations throughout the network. Larger values increase the model's representational capacity but proportionally increase computation and memory requirements across all layers.
  • encoder_layers / decoder_layers: The number of transformer blocks in each component. For BART-base, both are 6; for BART-large, both are 12. More layers increase model capacity by letting more refinement passes over the representations, but also increase training time and inference latency proportionally to depth.
  • encoder_attention_heads: The number of attention heads, 12 for base and 16 for large. Multiple heads allow the model to attend to different aspects of the input simultaneously. Each head operates on a slice of the hidden dimension (dk=dmodel/headsd_k = d_{\text{model}} / \text{heads}, so 64 dimensions per head in both base and large), letting the model to learn multiple distinct attention patterns.
  • encoder_ffn_dim: The dimension of the feed-forward network's hidden layer, typically 4 times the hidden dimension (3072 for base, 4096 for large). This controls the capacity of the non-linear transformations within each block. A larger FFN dimension allows the model to represent more complex functions of the attended information.
  • max_position_embeddings: The maximum sequence length the model can process, 1024 for BART. This is a hard architectural limit imposed by the learned absolute position embeddings. Attempting to process longer sequences requires extrapolation beyond the trained embedding vectors, which typically degrades performance.
  • activation_function: The non-linearity used in feed-forward layers, GeLU for BART. This affects the local curvature of the loss function and the speed of training convergence.

These parameters interact in important ways. Doubling d_model roughly quadruples the computation in attention layers (due to the O(n2d)O(n^2 d) complexity) and the computation in feed-forward layers (due to the O(nd2)O(n d^2) complexity of the two linear projections). Doubling encoder_layers doubles the depth and thus the computation linearly. When scaling BART to larger sizes, both dimensions are typically increased together, leading to the super-linear scaling behavior that characterizes large language model training costs.

Limitations and Impact

BART's architecture has several important limitations that practitioners should understand before deploying it or comparing it to more recent alternatives. These limitations are not defects so much as consequences of the design choices that gave BART its strengths in 2019, but which have been superseded by architectural advances in the years since.

The post-norm design that BART inherited from the original transformer makes training less stable at large scales compared to pre-norm architectures. When training BART-like models at the scale of hundreds of billions of parameters, the post-norm placement requires extremely careful tuning of learning rate warmup schedules and initialization strategies to prevent training instability. Pre-norm architectures, by contrast, are more forgiving: the normalization at the input of each sublayer ensures that residual connections add well-scaled increments regardless of depth. This practical advantage explains why T5 and virtually all subsequent large language models have adopted pre-norm. For BART at its original scales of 140M and 400M parameters, post-norm works adequately, but scaling further requires addressing this limitation.

The use of learned absolute position embeddings creates a hard context length limit of 1024 tokens. This was a reasonable constraint in 2019, when available compute made processing longer sequences prohibitive and most benchmark tasks used shorter inputs. But it has become a significant practical limitation as use cases requiring long-document understanding have grown. Summarizing a 10-page research paper, answering questions about a book chapter, or processing lengthy code files all require context lengths well beyond 1024 tokens. Extending BART to longer contexts requires either interpolating or extrapolating the position embeddings, which degrades performance in ways that are difficult to fully recover through fine-tuning. More modern architectures using rotary position embeddings (RoPE) or ALiBi can extrapolate to longer sequences with much less degradation, as we covered in Part XIV.

Computationally, BART's pretraining objective requires reconstructing the entire input document, making pretraining more expensive than T5's span corruption approach. For downstream applications, however, this difference disappears since both models fine-tune and generate output tokens autoregressively at the same cost. The pretraining cost difference matters primarily when training new BART-style models from scratch, a scenario that only large research labs or organizations with substantial compute budgets will encounter. For the large majority of practitioners, who use pretrained model weights, this limitation is invisible.

BART also lacks mechanisms for efficient long-sequence processing. Its quadratic attention complexity means processing 4,096 tokens requires 16 times more attention computation than processing 1,024 tokens. Modern efficient attention mechanisms such as FlashAttention, sparse attention, and linear attention variants have significantly reduced this computational bottleneck, but the original BART architecture predates these developments. Fine-tuned BART models used in production today often incorporate these efficiency improvements, but they were not part of the original design.

Despite these limitations, BART's impact on NLP research and practice has been substantial. It demonstrated that combining BERT-style encoding with GPT-style decoding produces a model that handles both understanding and generation more effectively than either family alone. The denoising pretraining framework established that text reconstruction under arbitrary corruption is a powerful general-purpose objective, opening research directions into what kinds of corruption are most beneficial and why. BART also provided the first clear evidence that encoder-decoder architectures could match or exceed encoder-only models on understanding-heavy tasks while substantially outperforming them on generation tasks, a finding that influenced the subsequent development of models like mBART, BART-large fine-tuned on summarization (Pegasus, ProphetNet), and the mT5 family.

BART also showed that encoder-decoder architectures work well for conditional generation tasks. While decoder-only models like GPT have since dominated many applications due to their simplicity and scalability, encoder-decoder models like BART remain competitive for tasks where the input and output differ structurally, such as document summarization, data-to-text generation, and machine translation. The key advantage of encoder-decoder models for these tasks is that the encoder can build a complete, contextualized representation of the entire input before generation begins, rather than treating the input tokens as just the first part of a unified sequence. For tasks where the input is semantically distinct from the output, this separation of understanding and generation continues to provide meaningful benefits.

Summary

BART combines a BERT-style bidirectional encoder with a GPT-style autoregressive decoder, creating a model designed for both understanding and generation. The design philosophy is explicit: inherit proven components rather than innovating architecturally, and let the pretraining objective (denoising reconstruction) unify the two inherited components into a coherent whole.

The three attention mechanisms in BART each serve a distinct role. Encoder bidirectional self-attention, represented by an all-ones mask matrix, allows every input token to incorporate context from every other token, building rich contextual representations. Decoder causal self-attention, represented by a lower-triangular mask, enforces the sequential discipline required for autoregressive generation. Cross-attention, with queries from the decoder and keys and values from the encoder, connects the understanding and generation phases, letting the decoder to dynamically focus on relevant parts of the input as it generates each output token.

Compared to T5, BART makes different architectural choices that reflect its design philosophy. GeLU activation instead of ReLU, post-norm instead of pre-norm, learned absolute positions instead of relative positions, and separate embedding matrices for encoder and decoder: these differences reflect BART's decision to directly combine BERT and GPT rather than developing novel architectural components. Each choice has practical implications for training stability, sequence length handling, and task-specific performance.

BART comes in base (140M parameters, 6+6 layers) and large (400M parameters, 12+12 layers) configurations. Both use the same architectural blueprint, with the large model giving greater representational capacity at higher computational cost. Compared to contemporary models, BART-large sits between BERT-large and T5-base in parameter count. This reflects the overhead of maintaining both a full encoder and a full decoder.

The limitations of BART, including post-norm instability at scale, fixed context length, and expensive pretraining, reflect the state of the field in 2019 and have been addressed in subsequent work. Its impact, establishing denoising encoder-decoder pretraining as a powerful framework for generation tasks, remains evident in the architectures and training objectives of models developed in its wake.

The next chapter explores BART's pretraining in detail, examining the various noising functions that teach the model to reconstruct corrupted text and how these objectives shape the model's capabilities for downstream tasks.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about BART's encoder-decoder architecture.

BART Architecture

Question 1 of 80 of 8 completed
What is the fundamental architectural design of BART?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025bartarchitecture, author = {Michael Brenndoerfer}, title = {BART Architecture: Encoder-Decoder Design for NLP}, year = {2025}, url = {https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). BART Architecture: Encoder-Decoder Design for NLP. Retrieved from https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers
MLAAcademic
Michael Brenndoerfer. "BART Architecture: Encoder-Decoder Design for NLP." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers>.
CHICAGOAcademic
Michael Brenndoerfer. "BART Architecture: Encoder-Decoder Design for NLP." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers.
HARVARDAcademic
Michael Brenndoerfer (2025) 'BART Architecture: Encoder-Decoder Design for NLP'. Available at: https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). BART Architecture: Encoder-Decoder Design for NLP. https://mbrenndoerfer.com/writing/bart-architecture-encoder-decoder-transformers

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.