Part of History of Language AI
Bahdanau attention replaced fixed context vectors with dynamic source alignment, improving neural machine translation and shaping later transformer models.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2015: Attention Mechanism
In work published in 2015, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio introduced an attention mechanism for neural machine translation. It allowed the decoder to compute a different weighted combination of source representations for each generated word. The method addressed a limitation of early encoder-decoder systems, which compressed the source sentence into one fixed-size vector.
Statistical machine translation systems used pipelines with separate stages for phrase extraction, translation, and reordering. Sequence-to-sequence architectures offered an alternative in 2014: an encoder and decoder could be trained together to map source text to target text. Their reliance on a single fixed-size source representation, however, created an information bottleneck.
Sequence-to-sequence models worked by having an encoder RNN process the source sentence and compress it into a single fixed-size vector, which a decoder RNN then used to generate the target sentence. While this approach showed promise for short sentences, it struggled with longer sequences. The encoder's final hidden state had to capture all information about the source sentence in a single vector, creating an information bottleneck that made it difficult to handle sentences longer than about 15 words. Longer source sentences led to degraded translation quality as important details were lost in the compression process.
Bahdanau and his colleagues removed that constraint by giving the decoder access to every encoder hidden state. When translating "the cat sat on the mat" to French, the model could assign high weight to "cat" while generating "chat," then change the weights for "s'est assis" or "tapis." The learned weights improved translation and produced useful source-target alignment visualizations.
For a detailed derivation and implementation of the additive alignment model, continue with Bahdanau Attention: Dynamic Context for NMT in the Language AI Handbook.
The Problem
In the 2014 sequence-to-sequence architecture, an encoder RNN processed the source sentence and passed its final hidden state to the decoder. That one fixed-size vector was the decoder's only representation of the source, creating what became known as the bottleneck problem.
For short sentences, this approach worked reasonably well. A sentence like "Hello, how are you?" contains relatively little information, and a typical hidden state size of 256 or 512 dimensions could capture the essential meaning. However, as source sentences grew longer, the fixed-size vector became increasingly inadequate. Longer sentences contained more information: additional noun phrases, modifiers, subordinate clauses, and complex grammatical structures. The encoder's final hidden state had to somehow compress all of this information while the decoder had to reconstruct the entire meaning from this single representation.
This compression problem manifested in several concrete ways. First, translation quality degraded noticeably for sentences longer than about 15 words. The model would lose track of information from earlier parts of the sentence, leading to translations that omitted important details or produced incorrect interpretations. Second, the model struggled with sentences that required maintaining long-range dependencies. In a sentence like "The keys that the man who visited yesterday left are on the table", the relationship between "keys" and "left" spans multiple words, and the final hidden state often failed to preserve this connection.
Third, the model had difficulty handling sentences with multiple independent pieces of information. A sentence like "John loves Mary, and she loves him too" contains two distinct relationships that both needed to be preserved. When compressed into a single vector, these relationships could interfere with each other or one might be lost entirely. Finally, the approach made it impossible to align specific source words with specific target words, which meant the model couldn't provide interpretable information about which source words influenced which target words.
The bottleneck problem became more severe when dealing with languages that had different word orders than the target language. In translating from English to Japanese, where the verb typically appears at the end, the encoder would process the entire English sentence before the decoder began generating Japanese. Information about the English verb, processed early by the encoder, had to be preserved through the entire encoding process and then accessed correctly by the decoder much later. The fixed-size bottleneck made this particularly challenging, as the verb information had to compete with all other sentence information for representation in the final hidden state.
The Solution
Bahdanau and his colleagues let the decoder access all encoder hidden states instead of only the final one. At each generation step, the model computed weights over those states and formed a step-specific context vector.
Different target words can depend on different source positions. The attention mechanism computed an alignment score between the decoder state and each encoder state, normalized the scores, and used them to form a weighted combination of encoder states.
Attention Computation
The attention mechanism worked by computing alignment scores between the decoder's current hidden state and each encoder hidden state. For each position in the source sentence and the current decoder step , the model computed an alignment score that measured how well the source word at position aligned with the target word being generated at step . These scores were computed using a small neural network, often called an alignment model, that took the decoder hidden state and encoder hidden state as inputs.
The alignment scores were then normalized using a softmax function to create attention weights. This normalization ensured that the weights summed to one and could be interpreted as a probability distribution over source positions. For decoder step , the attention weight for source position was computed as:
where is the length of the source sentence. These weights determined how much each encoder hidden state contributed to the context vector used for generating the current target word.
The context vector for decoder step was computed as a weighted sum of all encoder hidden states, where the weights came from the attention mechanism:
where represents the encoder hidden state at position . This context vector contained information from all source positions, weighted by their relevance to the current decoding step, and was combined with the decoder's hidden state to generate the next target word.
Alignment Model Variants
Bahdanau's original paper proposed an additive attention mechanism, where the alignment score was computed using a feedforward network with a single hidden layer. The score was calculated as:
where is the decoder hidden state at the previous step, is the encoder hidden state at position , and are weight matrices, is a learned vector, and is the activation function. This additive approach allowed the model to learn complex relationships between decoder and encoder states.
Later work, particularly by Minh-Thang Luong and colleagues, introduced a simpler dot-product attention that computed alignment scores directly from the inner product between transformed decoder and encoder states. This variant reduced computational complexity while maintaining similar performance:
where and are learned transformation matrices. The dot-product approach was simpler and faster to compute, making it attractive for practical applications while preserving the core attention mechanism's ability to learn dynamic alignments.
Integration with Decoder
At each decoding step, the model computed weights over the encoder positions and formed a weighted context vector. It combined that context with the decoder hidden state to predict the next target word.
The weights could be plotted as an alignment matrix between source and target positions. In the running example, high values might connect "chat" with "cat," "s'est assis" with "sat," and "tapis" with "mat." These alignments were learned from the translation objective rather than direct alignment labels.
Applications and Impact
Attention improved neural machine translation, especially on longer sentences that exposed the fixed-vector bottleneck. The decoder no longer had to recover every source detail from the encoder's final state alone.
Attention mechanisms also improved handling of long-range dependencies, which had been a persistent challenge for RNN-based models. In translating complex sentences with multiple clauses or embedded structures, attention allowed the decoder to directly access encoder states from much earlier in the sequence. A sentence like "The book that the professor who taught the advanced course recommended is excellent" contains nested dependencies spanning many words. Attention mechanisms could learn to focus on "professor" when generating the relevant target word, even if it appeared much earlier in the source sentence.
Attention matrices also provided alignment diagnostics. Researchers could inspect whether source and target positions received plausible correspondences, including cases where word order differed between languages. A diffuse or unexpected pattern could reveal a translation error, although the weights were not a complete explanation of the model's decision.
Researchers used these plots to debug alignments and compare architectures. The visualization exposed one internal weighting pattern, but did not make the rest of the network transparent.
Attention soon became a standard component of neural machine translation. Direct access to source positions helped with reordering and with copying or translating rare words and proper nouns.
Limitations
Encoder-decoder attention computes a score for each source-target position pair. With source length and target length , this requires alignment scores. The cost is quadratic when both lengths grow together and becomes expensive for long sequences.
The score calculations for a given decoder step can be parallelized across source positions. The surrounding RNN decoder, however, still generated target states sequentially, and the growing score matrix increased computation and memory use. Later self-attention architectures removed recurrence but retained the all-pairs cost.
Attention mechanisms also struggled with certain types of linguistic phenomena. While they excelled at word-level alignments, they had difficulty handling phrase-level or syntactic-level correspondences. A complex source phrase might need to be translated as a single target word, or vice versa, and attention mechanisms sometimes failed to capture these multi-word correspondences effectively. The mechanism worked best when alignments were roughly one-to-one or one-to-many, but struggled with many-to-one or many-to-many alignments that required more complex coordination.
Another limitation was the lack of explicit modeling of attention history. The standard attention mechanism computed weights independently for each decoding step, without explicitly tracking which source positions had already been attended to. This could lead to problems like repetition, where the model would attend to the same source words multiple times, or omission, where important source words were never attended to. While the decoder's hidden state implicitly tracked some of this information, explicit coverage mechanisms would later be developed to address these issues.
The attention mechanism also required storing all encoder hidden states in memory throughout the decoding process. For long sequences, this memory requirement could become prohibitive, especially when processing batches of sequences in parallel. Unlike simpler encoder-decoder models that only needed the final encoder hidden state, attention-based models needed to maintain all intermediate states, increasing memory requirements linearly with sequence length.
Finally, while attention provided interpretable alignments, these alignments were not always linguistically meaningful. The model learned attention patterns that improved translation quality, but these patterns did not necessarily correspond to semantic or syntactic relationships in ways that linguists would recognize. Attention weights could be noisy or spread across multiple source positions when a single focused alignment would be more appropriate. This reflected the model's optimization for translation accuracy rather than linguistic interpretability.
Legacy and Looking Forward
Bahdanau attention replaced one fixed source encoding with a context vector that changed at every decoding step. Its use in translation also helped establish learned, content-dependent weighting as a reusable component in neural sequence models.
The transformer architecture, introduced in 2017, extended attention to relationships within a sequence. Self-attention lets each position combine information from other positions, and transformer blocks use it in place of recurrent sequence processing. Models such as GPT and BERT are built from these transformer blocks.
Attention visualizations became a common diagnostic for inspecting alignments and comparing model behavior. Later work also showed why attention weights should not be treated as a definitive explanation: other internal states and parameters contribute to each prediction.
Multimodal systems also use attention to connect representations of text, images, and audio. Cross-attention can associate a generated caption with image regions or relate a question's tokens to visual features.
The translation objective trained the alignment model jointly with the encoder and decoder. This replaced a separately engineered alignment stage with weights learned end to end, though the learned alignments were optimized for prediction rather than linguistic analysis.
Transformer language models continue to use attention, with variants that change the sparsity pattern, approximate the computation, or implement it more efficiently. Across these variants, learned weights determine how representations at different positions are combined.
Quiz
The following questions review the encoder-decoder bottleneck, alignment weights, computational cost, and the connection to self-attention.
Attention Mechanism Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!