Part of Language AI Handbook
Explains how feed-forward networks provide nonlinearity in transformers, with 2-layer architecture, 4x dimension expansion, parameter analysis.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Feed-Forward Networks
Self-attention lets tokens gather contextual information from across the sequence. But attention alone is limited: it computes weighted averages of value vectors, a fundamentally linear operation. After softmax normalization, each output position is a convex combination of value vectors. No matter how many attention heads you stack, no matter how carefully you tune the query and key projections, the mapping from input to output remains linear in the values. To learn complex functions of language, transformers need something more.
The feed-forward network (FFN) provides exactly this missing ingredient. Sitting in every transformer block alongside the self-attention sublayer, it applies a nonlinear transformation to each token's representation independently. Think of it as the transformer's "thinking" layer: after attention has routed information between positions and assembled contextually-aware representations, the FFN processes each of those representations through a small neural network that can learn curved, high-dimensional decision boundaries. This is where the actual computational work of language understanding happens.
Every transformer block contains two main components: a self-attention sublayer and a feed-forward sublayer. While attention handles inter-token communication, the FFN handles per-token computation. The division of labor is elegant and deliberate. Attention routes information: it reads from the full sequence and assembles each position's context. The FFN transforms information: it takes each position's assembled context and applies a learned nonlinear function to produce a richer representation. Together, they enable transformers to learn the rich, hierarchical, compositional representations that power modern language AI.
Understanding the FFN matters beyond academic interest. In most transformer architectures, the FFN contains significantly more parameters than the attention mechanism, often two-thirds or more of each layer's total parameters. It also dominates computational cost for short sequences. This combination, large parameter count and heavy compute, makes FFN design and optimization one of the most important engineering problems in modern LLM deployment. Techniques like mixture of experts, activation sparsity, and quantization all target the FFN specifically.
The position-wise feed-forward network was introduced in the original transformer paper "Attention Is All You Need" by Vaswani et al. (2017). The authors used a two-layer architecture with ReLU activation and a 4x hidden dimension expansion factor, choices that have become de facto standards across the field. What started as a practical design decision, pairing a nonlinear sublayer with the attention sublayer, turned out to encode deep inductive biases about how information should be processed in language models. Subsequent research has refined the activation function (ReLU to GELU to SiLU/Swish), explored gated variants (GLU, SwiGLU), and investigated the FFN's role as associative memory, but the basic two-layer expand-and-contract structure has remained remarkably stable across nearly a decade of architectural innovation.
This chapter examines the FFN in detail. You'll learn its two-layer architecture and understand why the hidden dimension is expanded. You'll see how position independence enables massive parallelization and how the FFN acts as an associative memory that stores factual knowledge in its weights. You'll work through a numerical example step-by-step, trace through a complete implementation, and calculate the substantial parameter count that makes FFNs the largest component of most transformer models. By the end, you'll understand what the FFN computes, why it's designed the way it is, and what would be lost without it.
The Position-Wise Feed-Forward Network
To understand why transformers need the feed-forward network, consider what attention alone provides. Self-attention computes weighted averages of value vectors, where the weights come from query-key similarities. This is powerful for gathering contextual information: a token can attend to any position in the sequence and assemble a representation that reflects its surrounding context. But no matter how sophisticated the attention patterns, the operation remains a convex combination of value vectors. Each output is literally a weighted sum of inputs. This is a linear function.
Linear functions have severe limitations for language understanding. They can only rotate and scale the input space, then translate it. They cannot learn the curved decision boundaries that distinguish "bank" (financial institution) from "bank" (riverbank). They cannot capture the complex feature interactions that tell you whether "not bad" is negative or positive in context. They cannot represent the conditional relationships that make "the cat sat on the mat" grammatically fine but "the cat sat on the cat" slightly odd. For a model to approximate arbitrary functions of language, it needs the ability to carve up representation space in nonlinear ways.
The key insight is that any sufficiently wide network with a nonlinear activation function can approximate arbitrary continuous functions. This is the universal approximation theorem, a foundational result in neural network theory. The FFN uses this result by introducing a nonlinear activation between two linear layers. The first layer projects the input into a higher-dimensional space. The activation function applies a nonlinearity. The second layer projects back. The result is a function that can approximate far more complex relationships than any linear operation could.
There is a critical design choice embedded in how the FFN is applied: instead of mixing information across positions like attention does, the FFN processes each position independently. If your sequence has 100 tokens, the FFN applies the same transformation to each of those 100 representations separately, using identical weights for all positions. This is called a position-wise or position-independent transformation, and it fundamentally shapes the FFN's role in the transformer.
A position-wise operation applies the same function to each position in a sequence independently. The function's parameters are shared across positions, but the inputs and outputs at each position don't interact with each other.
This division of labor is elegant. Attention handles inter-position communication: it routes information between tokens, allowing the model to build representations that depend on context. The FFN handles intra-position computation: it transforms each position's representation using a learned nonlinear function. By separating these concerns, the transformer achieves both contextual awareness (from attention) and expressive power (from the FFN).
Why share the same network across all positions? Two reasons justify this choice. First, language exhibits translation invariance in a deep sense: the grammatical patterns that help understand "the cat" are equally useful whether those words appear at the beginning, middle, or end of a sentence. A transformation that converts a noun phrase's representation into something more useful should work wherever that noun phrase appears. Second, parameter sharing dramatically reduces model size. Instead of learning separate networks for each of the potentially thousands of positions in a sequence, we learn one network that generalizes across all positions. This is a major efficiency win.
Think of the position-wise FFN as a shared "dictionary of transformations." After attention has looked up what each position means in context, the FFN looks up how to transform that meaning into a richer representation, using the same dictionary regardless of where the position appears in the sequence.
The Two-Layer Architecture
With this motivation in place, let's examine what the FFN computes. The architecture is surprisingly simple: two linear transformations with a nonlinear activation function sandwiched between them. But this simplicity is deceptive. The expand-transform-contract structure enables remarkably expressive transformations, and the specific choice of expansion factor shapes the model's capacity in ways that affect both learning and efficiency.
Consider a single token's representation: a vector with dimensions. This vector encodes everything the model currently knows about that token in its context, having passed through however many transformer layers preceded this one. The FFN transforms this vector in three stages:
- Expand: Project into a higher-dimensional space
- Transform: Apply nonlinearity to enable learning of curved boundaries
- Contract: Project back to the original dimension
The complete formula captures all three stages in one expression. Given an input vector , the FFN computes:
where:
- : the input vector for a single position (the token's representation from the previous layer)
- : the first projection matrix, which expands the dimension from to
- : the first bias vector, added after the first linear transformation
- : a nonlinear activation function (e.g., ReLU, GELU) applied element-wise to the hidden representation
- : the second projection matrix, which contracts the dimension back from to
- : the second bias vector, added to produce the final output
- : the hidden (or intermediate) dimension, typically set to
To understand this formula, let's read it from the inside out, following the order of operations.
Stage 1: The Expansion ()
The input vector is multiplied by a weight matrix and shifted by a bias vector . This is a standard linear transformation, but with a twist: the output dimension is larger than the input dimension. If has 512 dimensions, the result might have 2048 dimensions. We're projecting into a higher-dimensional space where the data is easier to manipulate. In this expanded space, the model has more "room" to organize features before applying nonlinearity.
Why does projecting to higher dimensions help? Think of it this way: in the original -dimensional space, the token's representation may conflate many different features in a compressed form. By expanding to dimensions, the first layer can disentangle these features, spreading them across a larger space where they are more separable. The nonlinearity can then act on these individual features independently, creating the curved boundaries that linear operations cannot.
Stage 2: The Nonlinearity ()
The activation function is applied element-wise to the expanded representation. Common choices include ReLU (which zeros out negative values) and GELU (a smoother alternative that we'll examine in detail in the next chapter). This is where the magic happens: the nonlinearity allows the network to learn curved decision boundaries that would be impossible with linear transformations alone.
The element-wise nature of the activation matters. Because it operates on each hidden dimension independently, it creates a kind of gating: some hidden dimensions activate strongly and contribute to the output, while others are suppressed to near zero. Different input patterns activate different subsets of hidden dimensions, which is exactly how the network implements different transformations for different types of inputs.
Stage 3: The Contraction ()
Finally, the transformed representation is projected back to the original dimension via weight matrix and bias . The output has the same dimensionality as the input, which is essential for the residual connection that adds the FFN output back to its input. The second linear layer combines the activated hidden dimensions into a compressed -dimensional output, essentially summarizing the computation done in the expanded space.
Why does this formula make sense? Notice that without the nonlinearity , the composition of two linear maps is itself a linear map: is just another linear function of . The nonlinearity is the only thing that prevents the two layers from collapsing into a single linear layer. The biases also allow the network to shift the activation function's operating range, making it easier to learn the right threshold for each hidden dimension.
The expansion ratio is a key hyperparameter. The original transformer used a 4x expansion: for , the hidden dimension was . This ratio has become a de facto standard, though modern architectures sometimes use different values, especially when combined with gated variants, as we'll explore later.
Implementation
With the formula understood, translating it to code reveals its simplicity. The entire FFN is just two matrix multiplications with a nonlinearity in between:
import numpy as np
def relu(x):
"""ReLU activation function."""
return np.maximum(0, x)
def ffn(x, W1, b1, W2, b2, activation=relu):
"""
Position-wise feed-forward network.
Args:
x: Input tensor, shape (n, d_model) or (d_model,)
W1: First layer weights, shape (d_model, d_ff)
b1: First layer bias, shape (d_ff,)
W2: Second layer weights, shape (d_ff, d_model)
b2: Second layer bias, shape (d_model,)
activation: Nonlinear activation function
Returns:
Output tensor, same shape as x
"""
hidden = activation(x @ W1 + b1)
output = hidden @ W2 + b2
return output
# Example dimensions (from original transformer)
d_model = 512 # Model dimension
d_ff = 2048 # Hidden dimension (4x expansion)
# Initialize weights with Xavier/Glorot initialization
W1 = np.random.randn(d_model, d_ff) * np.sqrt(2.0 / (d_model + d_ff))
b1 = np.zeros(d_ff)
W2 = np.random.randn(d_ff, d_model) * np.sqrt(2.0 / (d_ff + d_model))
b2 = np.zeros(d_model)
# Process a single position
x_single = np.random.randn(d_model)
y_single = ffn(x_single, W1, b1, W2, b2)Single position FFN: Input shape: (512,) Output shape: (512,) Input norm: 22.2545 Output norm: 13.1094
The FFN preserves the dimensionality of its input: a 512-dimensional vector goes in, and a 512-dimensional vector comes out. This is essential for the residual connection that adds the FFN output back to its input. The residual connection (which adds the original input to the FFN output) requires both tensors to have the same shape, so the FFN's dimension-preserving property is architecturally necessary rather than merely convenient.
Notice that we initialize the weights using Xavier/Glorot initialization, scaling by . This initialization keeps the variance of activations stable across layers, preventing the exploding or vanishing gradients that plagued deep networks before careful initialization became standard practice. The biases are initialized to zero, which is standard; the weight matrices carry all the initial structure.
The Hidden Dimension Expansion
The most striking aspect of the FFN architecture is the dimension expansion. The first linear layer projects from to , typically with . For a model with , this means expanding to 2048 dimensions before projecting back down. For GPT-3 with , the hidden dimension reaches 49152, nearly fifty thousand neurons per FFN layer.
This expansion might seem wasteful. Why create a 2048-dimensional intermediate representation just to immediately compress it back to 512 dimensions? The answer lies in the expressiveness of the network, and it connects to some deep results in neural network theory.
A fundamental result in approximation theory shows that wider hidden layers can approximate more complex functions. Consider what happens without expansion: if , the hidden layer has the same dimensionality as the input. While the formula still applies, the network has limited capacity to decompose and recombine features. Each hidden neuron is a linear combination of the input features, gated by nonlinearity. With only neurons, the network can detect different patterns. With neurons, it can detect four times as many patterns and combine them in far richer ways.
Think of the expansion as temporarily working in a higher-dimensional space where the data is easier to manipulate. In the expanded 2048-dimensional space, the activation function can create very specific, highly discriminating patterns. One hidden neuron might activate only when the input represents a verb in past tense. Another might activate for proper nouns that appear after articles. These fine-grained detectors would be impossible to implement in 512 dimensions because there simply isn't enough room to disentangle all the features. The expansion provides the working space to learn them, and the contraction summarizes the results back into the original dimensionality for the residual connection.
The 4x expansion factor was established in the original "Attention Is All You Need" paper and has become a standard choice. It represents a pragmatic balance between expressiveness and efficiency: large enough to provide substantial nonlinear capacity, small enough to remain computationally tractable. Modern models like LLaMA use slightly different ratios (around 2.7x) when using gated linear units (GLUs), which effectively increase the hidden dimension through gating, but the underlying logic of expanding before contracting remains the same.
# Visualize the dimension flow through the FFN
def analyze_ffn_dimensions(d_model, expansion_factor):
"""Analyze dimensions through FFN layers."""
d_ff = d_model * expansion_factor
return {
"input": d_model,
"after_W1": d_ff,
"expansion_ratio": d_ff / d_model,
"after_W2": d_model,
}
# Common configurations
configs = [
("GPT-2 Small", 768, 4),
("GPT-2 Medium", 1024, 4),
("GPT-2 Large", 1280, 4),
("GPT-3 (175B)", 12288, 4),
("LLaMA-7B", 4096, 2.6875), # Uses 11008 hidden dim
]FFN dimension expansion across models: Model d_model d_ff Ratio ---------------------------------------------- GPT-2 Small 768 3072 4.00x GPT-2 Medium 1024 4096 4.00x GPT-2 Large 1280 5120 4.00x GPT-3 (175B) 12288 49152 4.00x LLaMA-7B 4096 11008 2.69x
The pattern is consistent: every model expands significantly in the hidden layer. LLaMA's slightly lower ratio reflects the use of the SwiGLU activation, a gated variant that includes an additional linear projection. Even with a nominally lower ratio, SwiGLU-based FFNs have similar or greater effective capacity because the gate mechanism provides additional expressive power.
Let's visualize how the actual activation values change through each stage of the FFN, to build intuition about what the expand-transform-contract process looks like in practice:



The three histograms tell an important story. The input distribution is centered at zero and roughly Gaussian, which is typical after layer normalization. After ReLU, the distribution becomes one-sided: roughly half the hidden units collapse to exactly zero (shown by the red line), while the other half pass through positive values unchanged. This sparsity is not a bug but a feature: different inputs will activate different subsets of hidden dimensions, creating a sparse code that allows the network to implement many different transformations. The output distribution returns to a balanced form around zero after the second linear layer combines the sparse hidden activations.
Now let's visualize the dimension sizes at each stage with a schematic:

Position Independence
A important property of the FFN is that it processes each position independently. Unlike attention, where every position can influence every other position through the query-key-value mechanism, the FFN applies an identical transformation to each position in isolation. This has practical consequences for both computation and interpretation, and it reflects a deep design decision about the division of responsibilities in the transformer architecture.
When we say the FFN is position-independent, we mean that the output at position depends only on the input at position . The FFN does not know or care about adjacent tokens. It does not look at the whole sequence. It simply applies the same transformation to whatever vector it receives, and that vector could have come from any position in the sequence.
This might seem like a limitation. After all, language is deeply contextual. The meaning of "bank" depends on whether "river" or "money" appeared nearby. But here is the key insight: by the time the FFN sees the representation at position , that representation has already been processed by the self-attention sublayer in the same block. Attention has already incorporated contextual information. The FFN is operating on a context-enriched representation, not on a raw token embedding. Its job is to transform that already-contextualized representation further, not to gather more context.
Think of the transformer block as a two-phase processor: attention gathers context (inter-position), and the FFN processes the gathered context (intra-position). The FFN can focus entirely on the "what to do with this information" question because attention has already handled the "what context is relevant" question.
Let's verify this independence empirically:
# Demonstrate position independence
seq_len = 5
X = np.random.randn(seq_len, d_model)
# Process all positions at once (batch processing)
Y_batch = ffn(X, W1, b1, W2, b2)
# Process each position individually
Y_individual = np.zeros_like(X)
for i in range(seq_len):
Y_individual[i] = ffn(X[i], W1, b1, W2, b2)
# Check they're identical
difference = np.abs(Y_batch - Y_individual).max()Position independence verification: Maximum difference between batch and individual processing: 3.77e-15 Outputs are identical: True
The batch and individual processing produce identical results, up to floating-point precision. This confirms that positions don't interact within the FFN. Changing the token at position 3 has absolutely no effect on the FFN output at position 1, even if they share the same FFN weights. The information barrier between positions is absolute.
Position independence means the FFN can be computed in parallel across all positions. On a GPU, this is extremely efficient: instead of processing tokens sequentially, we process the entire sequence simultaneously as a batch matrix multiplication. The FFN is embarrassingly parallel, meaning no position needs to wait for any other position's result. For a sequence of 1024 tokens, we can compute all 1024 FFN outputs simultaneously, limited only by the size of the weight matrices and the amount of memory available.
Position independence also clarifies the division of labor in a transformer block. Attention handles inter-position communication: it routes information between tokens, allowing the model to build representations that depend on context. The FFN handles intra-position transformation: it transforms each position's representation using the same learned function, adding nonlinearity and processing capacity to what would otherwise be a purely linear attention mechanism. This division of labor is one of the transformer's most elegant architectural properties.

Interpreting the FFN as Key-Value Memory
One of the most illuminating ways to understand the FFN is to see it as a neural network layer and as an associative memory: a structure that stores key-value pairs in its weights and retrieves the appropriate values when presented with matching inputs. This interpretation, developed by researchers at Tel Aviv University and other institutions, helps explain how transformers store and retrieve factual knowledge, and why simply scaling up the FFN increases a model's capacity to "know" things.
The idea is elegant. Consider the first layer's weight matrix . Each column of can be thought of as a "key" that matches certain input patterns. When we compute , we are essentially computing the dot product between the input and each of the keys. Each element of the resulting vector measures how well the input matches the corresponding key.
To make this precise: the pre-activation value for hidden dimension measures the alignment between the current input and the -th key pattern. Given an input and the -th column of (call it ), this alignment is:
where:
- : the pre-activation value for hidden dimension , measuring how well the input matches key
- : the input vector (-dimensional), representing the current token's contextual state
- : the -th column of , encoding the pattern that key is tuned to detect
- : the -th element of the bias vector , which acts as a threshold, biasing the key toward activation or suppression
When the input aligns well with key (high dot product), the corresponding hidden dimension activates strongly. After ReLU, only positive activations survive, so keys that match the current input "fire" while keys that don't match are silenced. This is exactly how an associative memory works: present a query, retrieve only the memories that match.
The second layer's weight matrix contains the "values" associated with each key. Each row of (call it for the -th row) is the value vector that gets added to the output when hidden dimension is active. The final FFN output is a weighted sum of these values, where the weights are the hidden activations:
where:
- : the post-activation strength of the -th key match
- : the -th row of , encoding the value retrieved when key fires
- The sum runs over all key-value pairs, but in practice only the active (nonzero) values contribute
The feed-forward layer can be interpreted as an associative memory where columns are keys, hidden activations are match scores, and rows are values. Input patterns that match certain keys retrieve their associated values. With key-value pairs per layer and dozens of layers per model, the total storage capacity of the FFN weights is enormous.
This memory interpretation explains several important empirical observations about transformers. Researchers have found that specific neurons in FFN layers activate for particular concepts: individual neurons that fire consistently for "the Eiffel Tower," for "programming languages," or for "past tense verbs." These neurons are not hand-designed. They emerge from training, as the model discovers that certain hidden dimensions reliably encode certain patterns. The FFN learns to store factual associations in its weights, and the attention mechanism retrieves them by routing relevant input patterns through the FFN.
This interpretation also explains a peculiar phenomenon called "knowledge editing" in language models. When researchers want to update a model's stored facts (for instance, changing "The president of France is X" to a newer value), they can sometimes accomplish this by modifying specific rows of FFN weight matrices. The change propagates naturally because those rows encode the associated value that gets retrieved when the "president of France" key fires.
The memory interpretation also clarifies why larger FFN dimensions store more knowledge. More hidden dimensions means more key-value pairs, which means more distinct patterns the model can recognize and more distinct responses it can retrieve. This is why scaling the FFN dimension is one of the primary ways to increase a model's factual capacity.
Let's visualize this interpretation with a concrete example:
# Interpret FFN as key-value memory
d_model_small = 8
d_ff_small = 16
W1_small = np.random.randn(d_model_small, d_ff_small) * 0.5
b1_small = np.zeros(d_ff_small)
W2_small = np.random.randn(d_ff_small, d_model_small) * 0.5
b2_small = np.zeros(d_model_small)
# Create an input that strongly activates certain hidden dimensions
x_test = np.random.randn(d_model_small)
# Compute hidden activations (before and after ReLU)
pre_activation = x_test @ W1_small + b1_small
hidden_activations = relu(pre_activation)
# See which "keys" were matched (high activation)
active_dims = hidden_activations > 0.5FFN as Key-Value Memory: Input dimension: 8 Number of 'keys' (hidden dim): 16 Pre-activation (dot product with keys): [ 2.89 -0.3 -1.39 1.7 0.28 0.9 1.25 -0.09 -0.54 1.3 -2.01 -0.44 -0.15 3.36 0.76 -1.39] After ReLU (matched keys): [2.89 0. 0. 1.7 0.28 0.9 1.25 0. 0. 1.3 0. 0. 0. 3.36 0.76 0. ] Strongly activated dimensions: [ 0 3 5 6 9 13 14]
Let's visualize this sparsity pattern across multiple inputs to see how different inputs activate different subsets of hidden dimensions:

The heatmap reveals the key-value memory character directly. Each column represents one hidden dimension (one key-value pair). The white cells are zero, meaning that key did not fire for that input. The blue cells indicate firing keys, with darker blue indicating stronger activation. Notice that no two input samples have exactly the same activation pattern: different inputs retrieve different combinations of stored associations. This input-dependent sparse access is what makes the FFN such a powerful memory structure.
This sparsity also has practical implications for efficiency. In a model with large , only a fraction of hidden dimensions are nonzero for any given input. Methods that exploit this sparsity can skip computation for inactive dimensions, reducing the compute needed well below the theoretical maximum. Recent work on ReLU-based sparse FFNs has shown that sparsity rates of 90% or higher can be achieved with minimal quality loss, meaning only 10% of hidden dimensions are needed for most tokens.
Parameter Count Analysis
Feed-forward networks are the largest component of transformer models by parameter count. In a standard transformer, the FFN contains significantly more parameters than the attention mechanism. This matters for model scaling and efficiency optimization, and it helps explain why so many recent research directions (mixture of experts, weight sharing, factorization) target the FFN specifically.
For a single FFN layer, we count the total number of learnable parameters by summing the sizes of all weight matrices and bias vectors. The calculation is straightforward. Matrix has shape , contributing parameters. Matrix has shape , contributing an equal parameters. The two bias vectors contribute and parameters respectively. Adding these together:
where:
- : the number of elements in weight matrix (rows columns)
- : the number of elements in weight matrix (rows columns, equal to the count in )
- : the number of elements in bias vector
- : the number of elements in bias vector
- The factor of 2 in front of accounts for both weight matrices and having the same total element count
With the standard expansion factor , we can simplify this expression. Substituting :
For large models where is in the hundreds or thousands, the quadratic term dominates and the linear bias terms become negligible, giving approximately parameters per FFN layer.
Why does this formula make sense? Notice that the parameter count grows quadratically with : if you double the model dimension, you quadruple the FFN parameter count. This is the same quadratic scaling that makes attention expensive with respect to sequence length, but here the quadratic factor is in the model dimension rather than sequence length. Doubling from 512 to 1024 increases FFN parameters from roughly 2 million to 8 million per layer, a 4x increase for a 2x dimension increase.
def count_ffn_params(d_model, d_ff, include_bias=True):
"""Count parameters in a feed-forward network."""
weight_params = 2 * d_model * d_ff # W1 and W2
bias_params = d_ff + d_model if include_bias else 0
return weight_params + bias_params
def count_attention_params(d_model, num_heads, include_bias=True):
"""Count parameters in multi-head attention."""
# Q, K, V projections and output projection
weight_params = 4 * d_model * d_model # W_Q, W_K, W_V, W_O
bias_params = 4 * d_model if include_bias else 0
return weight_params + bias_params
# Compare for different model sizes
model_configs = [
("GPT-2 Small", 768, 3072, 12),
("GPT-2 Medium", 1024, 4096, 16),
("BERT-Base", 768, 3072, 12),
("GPT-3 (175B)", 12288, 49152, 96),
]Parameter count comparison: FFN vs Attention (per layer) Model d_model d_ff FFN Params Attn Params FFN/Attn -------------------------------------------------------------------------------- GPT-2 Small 768 3072 4,718,592 2,359,296 2.0x GPT-2 Medium 1024 4096 8,388,608 4,194,304 2.0x BERT-Base 768 3072 4,718,592 2,359,296 2.0x GPT-3 (175B) 12288 49152 1,207,959,552 603,979,776 2.0x
The FFN consistently contains about twice as many parameters as the attention mechanism per layer. This 2:1 ratio emerges from a straightforward calculation: attention uses four weight matrices of size for QKV and output projections, totaling parameters. The FFN uses two matrices with total size parameters. The ratio is exactly , so the FFN always has twice the weight parameters of attention when the 4x expansion factor is used.
For large models like GPT-3, each transformer layer has over 2.4 billion parameters in the FFN alone. Multiplied across 96 layers, the FFN accounts for the vast majority of the model's 175 billion total parameters. This makes FFN optimization critical for model efficiency, both in terms of memory and computation.
Let's visualize how FFN parameters scale with model dimension:

Let's also visualize the parameter distribution in a transformer block to see the FFN's dominance at a glance:

Computational Cost
Beyond parameter count, we need to consider computational cost, measured in floating-point operations (FLOPs). Understanding FFN compute requirements helps explain why these layers dominate inference time for short sequences, and why long-context models face a different bottleneck.
The FFN's computational cost per token follows directly from its structure. For each token, the FFN performs two matrix-vector multiplications. Each multiply-add operation consists of one multiplication and one addition, counting as 2 floating-point operations. Working through the arithmetic:
The first layer computes where is a -dimensional vector and is . This requires multiply-adds, so FLOPs. The second layer computes where is -dimensional and is , requiring another FLOPs. The total FLOPs per token for the FFN is therefore:
where:
- : the input and output dimension of the FFN
- : the hidden dimension (typically )
- The factor of 4 comes from: 2 layers 2 FLOPs per multiply-add operation
For a sequence of tokens, each token is processed independently, so the total FFN cost scales linearly with sequence length. Because the FFN has no interaction between positions, the computation at each position is identical and independent:
where is the number of tokens in the sequence. This is scaling: double the sequence length, double the FLOPs.
This linear scaling contrasts sharply with attention, which has quadratic complexity due to the attention matrix computation. For every pair of positions, attention must compute a query-key dot product. With positions, there are such pairs. As sequences grow longer, attention's quadratic term eventually dominates.
The crossover point depends on model dimensions. For a model with and , let's compute when attention FLOPs equal FFN FLOPs. FFN FLOPs scale as . Attention's quadratic term (for the attention matrix alone) scales as . Setting these equal: , giving tokens. At sequence lengths below 3000, FFN dominates. Above 3000, attention takes over.
def ffn_flops(n, d_model, d_ff):
"""Calculate FFN FLOPs for sequence length n."""
return 4 * n * d_model * d_ff
def attention_flops(n, d_model):
"""Calculate attention FLOPs (simplified)."""
# Q, K, V projections: 3 * 2 * n * d_model^2
# QK^T: 2 * n^2 * d_model
# Attention @ V: 2 * n^2 * d_model
# Output projection: 2 * n * d_model^2
projection_flops = 4 * 2 * n * d_model * d_model
attention_matrix_flops = 4 * n * n * d_model
return projection_flops + attention_matrix_flops
# Compare across sequence lengths
sequence_lengths = [128, 512, 1024, 2048, 4096, 8192]
d_model_test = 768
d_ff_test = 3072FLOPs comparison: FFN vs Attention (GPT-2 Small)
Seq Length FFN FLOPs Attn FLOPs FFN/Attn
------------------------------------------------------------
128 1,207,959,552 654,311,424 1.85x
512 4,831,838,208 3,221,225,472 1.50x
1024 9,663,676,416 8,053,063,680 1.20x
2048 19,327,352,832 22,548,578,304 0.86x
4096 38,654,705,664 70,866,960,384 0.55x
8192 77,309,411,328 244,813,135,872 0.32xFor short sequences (128-512 tokens), the FFN contributes over twice the FLOPs of attention. As sequences grow to 4096 or 8192 tokens, attention's quadratic cost catches up and eventually dominates. The crossover point around 1024-2048 tokens is where attention and FFN have comparable costs for this model configuration.

This crossover point has important practical implications. Short-context applications like question answering and classification, as well as code completion, typically operate with sequences well below the crossover point, meaning the FFN is the primary compute bottleneck. Research into FFN efficiency (quantization, pruning, low-rank approximation) offers the greatest gains here. Long-context applications like document summarization, long-form generation, and retrieval-augmented generation operate above the crossover, making attention efficiency (sparse attention, linear attention, sliding window attention) the priority.
A Complete Worked Example
The formula is compact, but its compactness can obscure what happens at each step. To understand the FFN, let's trace through a concrete computation with actual numbers. We'll use intentionally small dimensions ( and ) so you can follow every multiplication and addition by hand if you like, verifying each step against the formulas.
The goal of this example is threefold: to see exactly how the input vector is transformed at each stage, to observe how ReLU zeros out negative activations to introduce nonlinearity, and to verify that the output dimension matches the input dimension, ready for the residual connection. We'll use simple, hand-readable weight values to keep the arithmetic transparent.
# Small example for hand-traceable computation
d_model_tiny = 3
d_ff_tiny = 4
# Initialize weights with simple values
W1_tiny = np.array(
[
[0.5, -0.3, 0.8, 0.2],
[-0.2, 0.6, 0.1, -0.4],
[0.3, 0.1, -0.5, 0.7],
]
) # Shape: (3, 4)
b1_tiny = np.array([0.1, -0.1, 0.2, 0.0]) # Shape: (4,)
W2_tiny = np.array(
[
[0.4, -0.2, 0.3],
[0.1, 0.5, -0.1],
[-0.3, 0.2, 0.4],
[0.2, -0.4, 0.1],
]
) # Shape: (4, 3)
b2_tiny = np.array([0.05, -0.05, 0.1]) # Shape: (3,)
# Input vector
x_tiny = np.array([1.0, -0.5, 0.8])We have chosen an input with mixed positive and negative values (), which is representative of what you'd see in practice after layer normalization. The weight matrices use a mix of positive and negative values. This ensures some hidden dimensions will have positive pre-activations (and survive ReLU) while others will be zeroed out.
Stage 1: Expansion via the First Linear Layer
The first operation computes . Our 3-dimensional input vector gets multiplied by a weight matrix, producing a 4-dimensional hidden representation. Each element of this output is a weighted sum of the input elements plus a bias term. Specifically, the -th element of the pre-activation is : a dot product between the input and the -th column of .
Step 1: First linear transformation (x @ W1 + b1) Input x: [ 1. -0.5 0.8] W1: [[ 0.5 -0.3 0.8 0.2] [-0.2 0.6 0.1 -0.4] [ 0.3 0.1 -0.5 0.7]] b1: [ 0.1 -0.1 0.2 0. ] x @ W1 + b1 = [ 0.94 -0.62 0.55 0.96]
Notice the dimension change: a 3-dimensional vector goes in, and a 4-dimensional vector comes out. This expansion is the "higher-dimensional working space" we discussed earlier, where the network has more room to manipulate the representation before applying nonlinearity. The result is called the pre-activation because we haven't applied the nonlinearity yet. Some of these pre-activation values are positive and some are negative, which sets up the selective filtering that ReLU will perform in the next step.
Stage 2: Nonlinearity via ReLU
Here is where the FFN gains its expressive power. The ReLU activation function applies a simple rule: keep positive values unchanged, but set negative values to zero. Mathematically, . This is one of the simplest possible nonlinear functions, yet it is sufficient to make the FFN a universal approximator when the hidden dimension is large enough.
Step 2: Apply ReLU activation Pre-activation: [ 0.94 -0.62 0.55 0.96] After ReLU: [0.94 0. 0.55 0.96] Negative values become zero, positive values pass through unchanged.
By zeroing out some dimensions, ReLU creates sparse hidden representations: only a subset of hidden dimensions are active for any given input. Different inputs activate different subsets, allowing the network to learn piece-wise linear functions that approximate arbitrary curves. Each hidden dimension can be thought of as a detector for a particular pattern in the input. When the detector fires (positive pre-activation), it contributes to the output. When it doesn't fire (negative pre-activation, zeroed by ReLU), it contributes nothing. The network learns during training which patterns each detector should respond to, and what contribution to make when it does respond.
The key insight here is that ReLU's behavior changes based on which side of zero the pre-activation falls. For pre-activations far above zero, the gradient of ReLU is 1 and learning proceeds normally. For pre-activations below zero, the gradient is 0 and the network cannot update that hidden dimension for that particular input. This creates a form of implicit regularization: hidden dimensions that consistently fail to fire for a category of inputs effectively "opt out" of processing that category, allowing specialization.


Stage 3: Contraction via the Second Linear Layer
Finally, we project back to the original dimension. The 4-dimensional hidden vector is multiplied by a weight matrix, and a 3-dimensional bias is added. This contraction reduces the dimension and combines the contributions of all active hidden dimensions into an output vector. Each output dimension receives a weighted sum of the active hidden dimensions, where the weights are the corresponding elements of .
Think of the second linear layer as the "summarization" step. The first layer asked "which features of this input are relevant?" The nonlinearity filtered out irrelevant dimensions. The second layer asks "given these relevant features, what transformation should we apply to the representation?"
Step 3: Second linear transformation (hidden @ W2 + b2) Hidden: [0.94 0. 0.55 0.96] W2: [[ 0.4 -0.2 0.3] [ 0.1 0.5 -0.1] [-0.3 0.2 0.4] [ 0.2 -0.4 0.1]] b2: [ 0.05 -0.05 0.1 ] hidden @ W2 + b2 = [ 0.453 -0.512 0.698]
The output has the same dimension as the input (3 elements), which is essential. In the full transformer, this output will be added to the original input via a residual connection, and both must have the same shape. The FFN has transformed the representation while preserving its dimensionality, enabling the residual addition that stabilizes training in deep networks.
Verification and Summary
Let's verify that our step-by-step calculation matches the complete FFN function applied directly:
Verification using ffn() function: Output: [ 0.453 -0.512 0.698] Match: True Summary: Input: [ 1. -0.5 0.8] (dimension 3) Output: [ 0.453 -0.512 0.698] (dimension 3)
The step-by-step and function-based computations match exactly. This worked example demonstrates the complete journey: a 3-dimensional input expands to 4 dimensions, passes through nonlinearity (with some dimensions zeroed out), and contracts back to 3 dimensions. The transformation is nonlinear (different inputs will activate different subsets of hidden dimensions, producing qualitatively different transformations), position-independent (the same computation would apply to any token's representation), and dimension-preserving (the output shape matches the input shape for the residual connection).
Implementation: A Complete FFN Module
Having traced through the mathematics by hand, we can now build a reusable FFN module that encapsulates everything we've learned. This implementation follows patterns used in production transformer libraries: it initializes weights using proper scaling (Xavier/Glorot initialization), supports optional biases, and handles both single vectors and batched sequences without any special casing.
naturally the class structure maps onto the mathematical formula. The __init__ method creates the four learnable parameter tensors (, , , ). The __call__ method implements the formula in three readable lines. The num_parameters method computes the formula we derived earlier. The entire forward pass is just three matrix multiplications (including the bias additions). This reflects the fundamental simplicity of the FFN despite its expressive power.
class FeedForwardNetwork:
"""
Position-wise feed-forward network for transformer blocks.
Implements: FFN(x) = activation(x @ W1 + b1) @ W2 + b2
"""
def __init__(self, d_model, d_ff, activation=relu, use_bias=True):
"""
Initialize the feed-forward network.
Args:
d_model: Input and output dimension
d_ff: Hidden dimension (typically 4 * d_model)
activation: Nonlinear activation function
use_bias: Whether to include bias terms
"""
self.d_model = d_model
self.d_ff = d_ff
self.activation = activation
self.use_bias = use_bias
# Xavier/Glorot initialization
self.W1 = np.random.randn(d_model, d_ff) * np.sqrt(
2.0 / (d_model + d_ff)
)
self.W2 = np.random.randn(d_ff, d_model) * np.sqrt(
2.0 / (d_ff + d_model)
)
if use_bias:
self.b1 = np.zeros(d_ff)
self.b2 = np.zeros(d_model)
else:
self.b1 = None
self.b2 = None
def __call__(self, x):
"""
Apply the feed-forward transformation.
Args:
x: Input tensor of shape (..., d_model)
Returns:
Output tensor of shape (..., d_model)
"""
# First linear layer
hidden = x @ self.W1
if self.use_bias:
hidden = hidden + self.b1
# Activation
hidden = self.activation(hidden)
# Second linear layer
output = hidden @ self.W2
if self.use_bias:
output = output + self.b2
return output
def num_parameters(self):
"""Return total parameter count."""
params = self.d_model * self.d_ff + self.d_ff * self.d_model
if self.use_bias:
params += self.d_ff + self.d_model
return params# Test the module
ffn_module = FeedForwardNetwork(d_model=512, d_ff=2048)
# Single vector
x_single_test = np.random.randn(512)
y_single_test = ffn_module(x_single_test)
# Batch of vectors (sequence)
x_batch_test = np.random.randn(16, 512) # 16 tokens
y_batch_test = ffn_module(x_batch_test)FeedForwardNetwork module test: Configuration: d_model: 512 d_ff: 2048 Parameters: 2,099,712 Single vector: Input shape: (512,) Output shape: (512,) Batch (sequence): Input shape: (16, 512) Output shape: (16, 512)
The module correctly handles both single vectors and batched sequences. When a 2D input with shape is provided (a batch of token representations), the matrix multiplication automatically broadcasts across all positions simultaneously. This is the source of the FFN's computational efficiency: instead of calling the function once per token in a loop, the entire sequence is processed in a single pair of batch matrix multiplications.
The parameter count matches our theoretical formula: , close to 2.1 million parameters just for this single FFN layer. A 12-layer GPT-2 Small model has 12 such layers, plus attention layers, adding up to the 117 million total parameters commonly cited.
Limitations and Impact
The feed-forward network is conceptually simple: two linear layers with a nonlinearity between them. Yet this simplicity masks significant computational cost, and understanding those costs is essential for anyone working with large language models in practice.
The FFN's most significant limitation is its memory footprint. Because FFN weights are static after training, they must reside in memory at inference time. For a model like LLaMA-70B with a hidden dimension of 8192 and an FFN dimension of 28672 (approximately 3.5x), each FFN layer holds about 2 8192 28672 470 million parameters. Across 80 layers, the FFN alone stores roughly 37 billion parameters, each requiring at least 2 bytes in FP16 format, meaning over 74 GB of memory just for FFN weights. This is why running large models requires high-memory GPUs or specialized hardware.
Weight quantization is the most common approach to this memory problem. By representing weights in lower precision (INT8, INT4, or even INT2), the memory footprint shrinks proportionally. A model quantized to INT4 requires one-quarter the memory of its FP16 counterpart, enabling deployment on consumer hardware. The challenge is maintaining model quality: aggressive quantization can degrade generation quality, particularly for reasoning tasks that require precise numerical computations in the FFN's hidden layers. Research into quantization-aware training and calibration has made significant progress, but the tradeoff between compression and quality remains an active area of work.
Sparsity offers another path forward. The ReLU activation naturally creates sparse hidden representations, as negative pre-activations become zero. Researchers have exploited this by identifying which hidden dimensions will be active for a given input and computing only those, skipping computation for dimensions that would be zeroed anyway. This requires predicting active neurons before the full computation, which introduces overhead, but the net savings can be substantial for highly sparse models. The GELU and SiLU activations that have replaced ReLU in many modern models are smoother alternatives that do not produce exact zeros, trading the sparsity benefit for better gradient flow during training.
The Mixture of Experts (MoE) architecture takes the sparsity idea further. Instead of a single FFN per layer, an MoE layer has multiple "expert" FFNs and a router that directs each token to only a small subset (typically 2 of 8, or 2 of 64). The model scales parameters without proportionally scaling compute, because only a fraction of experts process each token. Models like Mixtral-8x7B and GPT-4 (reportedly) use MoE to achieve parameter counts that would otherwise be computationally prohibitive. The key challenge is load balancing: ensuring that all experts are used approximately equally across the training data, preventing some experts from specializing in everything while others specialize in nothing.
Despite its computational weight, the FFN's role is needed and irreplaceable. It provides the nonlinearity that maps attention's linear weighted averages into expressive function approximation. It is the model's "memory," storing factual associations in its weights that attention retrieves based on context. Its position-wise nature enables massive parallelization that makes transformer training tractable at the scale of billions of parameters and trillions of training tokens.
The interplay between attention and FFN defines transformer expressiveness. Attention routes information between positions, creating context-dependent representations. The FFN transforms those representations position-by-position, adding computational depth and nonlinear capacity. Together, they enable the hierarchical, compositional language understanding that powers modern NLP. Neither component alone would be sufficient: attention without the FFN is a purely linear contextual aggregation; the FFN without attention is a powerful token-level classifier with no awareness of surrounding context.
Summary
The feed-forward network is the workhorse of transformer computation, applying identical transformations to each position independently. This chapter covered its architecture and efficiency, then examined its interpretation within the broader transformer block.
Key takeaways:
-
Two-layer architecture: The FFN formula consists of an expansion layer (), a nonlinear activation (ReLU, GELU, etc.), and a contraction layer (). The standard expansion factor is , establishing a working space four times larger than the model dimension where nonlinear transformations are easier to learn.
-
Position independence: Each position is processed separately with shared weights. This enables parallel computation and cleanly separates the FFN's role (per-position nonlinear transformation) from attention's role (inter-position information routing). The position independence is not a limitation but a design choice: by the time the FFN processes each token, attention has already incorporated contextual information into that token's representation.
-
Key-value memory interpretation: The FFN can be viewed as an associative memory where columns are keys, hidden activations are match scores, and rows are values. Different inputs activate different subsets of hidden dimensions, creating sparse, input-dependent retrievals. This explains how transformers store factual knowledge and why larger FFN dimensions increase a model's capacity to "know" things.
-
Dominant parameter count: The FFN contains approximately two-thirds of a transformer block's parameters. With 4x expansion, the total is approximately parameters per layer (from the formula ). This quadratic scaling with makes FFN optimization critical for model efficiency.
-
Linear computational scaling: FFN compute cost is FLOPs, scaling linearly with sequence length . This contrasts with attention's quadratic scaling. For short sequences, FFN dominates compute; for long sequences, attention becomes the bottleneck.
-
Crossover point: Around 1000-2000 tokens, FFN and attention have comparable computational costs for typical model sizes. This crossover point shapes optimization strategy: short-context systems prioritize FFN compression, while long-context systems prioritize attention efficiency.
The next chapter examines activation functions used in FFNs in depth, comparing ReLU, GELU, SiLU/Swish, and their gated variants (GLU, SwiGLU), and exploring why modern models have moved beyond the original ReLU choice toward smoother, more expressive alternatives.
Key Parameters
When implementing or configuring feed-forward networks in transformers, these parameters control capacity and efficiency:
-
d_model (input/output dimension): The embedding dimension that the FFN preserves. Typical values range from 256 (small models) to 12288 (GPT-3 scale). This dimension must match the attention layer output and determines the FFN's interface with the rest of the transformer.
-
d_ff (hidden dimension): The expanded dimension of the intermediate representation. The standard choice is , though modern architectures like LLaMA use ratios around 2.7x when combined with gated linear units. Larger values increase expressiveness but proportionally increase parameters and compute.
-
activation: The nonlinear function applied element-wise after the first linear layer. ReLU was used in the original transformer, but GELU has become the standard for encoder models (BERT, RoBERTa) and SiLU/Swish for decoder models (LLaMA, GPT-NeoX). The choice affects gradient flow, sparsity patterns, and training stability.
-
use_bias: Whether to include bias terms and . Some modern architectures (LLaMA, PaLM) omit biases entirely to reduce parameters and simplify quantization. The impact on model quality is typically minimal.
-
dropout (not shown in our implementation): Dropout rate applied to the hidden representation after activation. Values of 0.1-0.2 are common during training to prevent overfitting. Set to 0 during inference.
-
initialization: Weight initialization scale affects training stability. Xavier/Glorot initialization (scaling by ) is standard. Some architectures use scaled initialization for residual paths to maintain signal magnitude through deep networks.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about feed-forward networks in transformers.
Feed-Forward Networks Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!