Feed-Forward Networks in Transformers: Architecture

Michael BrenndoerferUpdated June 10, 202556 min read

Part of Language AI Handbook

Explains how feed-forward networks provide nonlinearity in transformers, with 2-layer architecture, 4x dimension expansion, parameter analysis.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Feed-Forward Networks

Self-attention lets tokens gather contextual information from across the sequence. But attention alone is limited: it computes weighted averages of value vectors, a fundamentally linear operation. After softmax normalization, each output position is a convex combination of value vectors. No matter how many attention heads you stack, no matter how carefully you tune the query and key projections, the mapping from input to output remains linear in the values. To learn complex functions of language, transformers need something more.

The feed-forward network (FFN) provides exactly this missing ingredient. Sitting in every transformer block alongside the self-attention sublayer, it applies a nonlinear transformation to each token's representation independently. Think of it as the transformer's "thinking" layer: after attention has routed information between positions and assembled contextually-aware representations, the FFN processes each of those representations through a small neural network that can learn curved, high-dimensional decision boundaries. This is where the actual computational work of language understanding happens.

Every transformer block contains two main components: a self-attention sublayer and a feed-forward sublayer. While attention handles inter-token communication, the FFN handles per-token computation. The division of labor is elegant and deliberate. Attention routes information: it reads from the full sequence and assembles each position's context. The FFN transforms information: it takes each position's assembled context and applies a learned nonlinear function to produce a richer representation. Together, they enable transformers to learn the rich, hierarchical, compositional representations that power modern language AI.

Understanding the FFN matters beyond academic interest. In most transformer architectures, the FFN contains significantly more parameters than the attention mechanism, often two-thirds or more of each layer's total parameters. It also dominates computational cost for short sequences. This combination, large parameter count and heavy compute, makes FFN design and optimization one of the most important engineering problems in modern LLM deployment. Techniques like mixture of experts, activation sparsity, and quantization all target the FFN specifically.

Historical Context

The position-wise feed-forward network was introduced in the original transformer paper "Attention Is All You Need" by Vaswani et al. (2017). The authors used a two-layer architecture with ReLU activation and a 4x hidden dimension expansion factor, choices that have become de facto standards across the field. What started as a practical design decision, pairing a nonlinear sublayer with the attention sublayer, turned out to encode deep inductive biases about how information should be processed in language models. Subsequent research has refined the activation function (ReLU to GELU to SiLU/Swish), explored gated variants (GLU, SwiGLU), and investigated the FFN's role as associative memory, but the basic two-layer expand-and-contract structure has remained remarkably stable across nearly a decade of architectural innovation.

This chapter examines the FFN in detail. You'll learn its two-layer architecture and understand why the hidden dimension is expanded. You'll see how position independence enables massive parallelization and how the FFN acts as an associative memory that stores factual knowledge in its weights. You'll work through a numerical example step-by-step, trace through a complete implementation, and calculate the substantial parameter count that makes FFNs the largest component of most transformer models. By the end, you'll understand what the FFN computes, why it's designed the way it is, and what would be lost without it.

The Position-Wise Feed-Forward Network

To understand why transformers need the feed-forward network, consider what attention alone provides. Self-attention computes weighted averages of value vectors, where the weights come from query-key similarities. This is powerful for gathering contextual information: a token can attend to any position in the sequence and assemble a representation that reflects its surrounding context. But no matter how sophisticated the attention patterns, the operation remains a convex combination of value vectors. Each output is literally a weighted sum of inputs. This is a linear function.

Linear functions have severe limitations for language understanding. They can only rotate and scale the input space, then translate it. They cannot learn the curved decision boundaries that distinguish "bank" (financial institution) from "bank" (riverbank). They cannot capture the complex feature interactions that tell you whether "not bad" is negative or positive in context. They cannot represent the conditional relationships that make "the cat sat on the mat" grammatically fine but "the cat sat on the cat" slightly odd. For a model to approximate arbitrary functions of language, it needs the ability to carve up representation space in nonlinear ways.

The key insight is that any sufficiently wide network with a nonlinear activation function can approximate arbitrary continuous functions. This is the universal approximation theorem, a foundational result in neural network theory. The FFN uses this result by introducing a nonlinear activation between two linear layers. The first layer projects the input into a higher-dimensional space. The activation function applies a nonlinearity. The second layer projects back. The result is a function that can approximate far more complex relationships than any linear operation could.

There is a critical design choice embedded in how the FFN is applied: instead of mixing information across positions like attention does, the FFN processes each position independently. If your sequence has 100 tokens, the FFN applies the same transformation to each of those 100 representations separately, using identical weights for all positions. This is called a position-wise or position-independent transformation, and it fundamentally shapes the FFN's role in the transformer.

Position-Wise Transformation

A position-wise operation applies the same function to each position in a sequence independently. The function's parameters are shared across positions, but the inputs and outputs at each position don't interact with each other.

This division of labor is elegant. Attention handles inter-position communication: it routes information between tokens, allowing the model to build representations that depend on context. The FFN handles intra-position computation: it transforms each position's representation using a learned nonlinear function. By separating these concerns, the transformer achieves both contextual awareness (from attention) and expressive power (from the FFN).

Why share the same network across all positions? Two reasons justify this choice. First, language exhibits translation invariance in a deep sense: the grammatical patterns that help understand "the cat" are equally useful whether those words appear at the beginning, middle, or end of a sentence. A transformation that converts a noun phrase's representation into something more useful should work wherever that noun phrase appears. Second, parameter sharing dramatically reduces model size. Instead of learning separate networks for each of the potentially thousands of positions in a sequence, we learn one network that generalizes across all positions. This is a major efficiency win.

Think of the position-wise FFN as a shared "dictionary of transformations." After attention has looked up what each position means in context, the FFN looks up how to transform that meaning into a richer representation, using the same dictionary regardless of where the position appears in the sequence.

The Two-Layer Architecture

With this motivation in place, let's examine what the FFN computes. The architecture is surprisingly simple: two linear transformations with a nonlinear activation function sandwiched between them. But this simplicity is deceptive. The expand-transform-contract structure enables remarkably expressive transformations, and the specific choice of expansion factor shapes the model's capacity in ways that affect both learning and efficiency.

Consider a single token's representation: a vector xx with dmodeld_{\text{model}} dimensions. This vector encodes everything the model currently knows about that token in its context, having passed through however many transformer layers preceded this one. The FFN transforms this vector in three stages:

  1. Expand: Project xx into a higher-dimensional space
  2. Transform: Apply nonlinearity to enable learning of curved boundaries
  3. Contract: Project back to the original dimension

The complete formula captures all three stages in one expression. Given an input vector x∈Rdmodelx \in \mathbb{R}^{d_{\text{model}}}, the FFN computes:

FFN(x)=σ(xW1+b1)W2+b2\text{FFN}(x) = \sigma(xW_1 + b_1)W_2 + b_2

where:

  • x∈Rdmodelx \in \mathbb{R}^{d_{\text{model}}}: the input vector for a single position (the token's representation from the previous layer)
  • W1∈Rdmodel×dffW_1 \in \mathbb{R}^{d_{\text{model}} \times d_{ff}}: the first projection matrix, which expands the dimension from dmodeld_{\text{model}} to dffd_{ff}
  • b1∈Rdffb_1 \in \mathbb{R}^{d_{ff}}: the first bias vector, added after the first linear transformation
  • σ\sigma: a nonlinear activation function (e.g., ReLU, GELU) applied element-wise to the hidden representation
  • W2∈Rdff×dmodelW_2 \in \mathbb{R}^{d_{ff} \times d_{\text{model}}}: the second projection matrix, which contracts the dimension back from dffd_{ff} to dmodeld_{\text{model}}
  • b2∈Rdmodelb_2 \in \mathbb{R}^{d_{\text{model}}}: the second bias vector, added to produce the final output
  • dffd_{ff}: the hidden (or intermediate) dimension, typically set to 4×dmodel4 \times d_{\text{model}}

To understand this formula, let's read it from the inside out, following the order of operations.

Stage 1: The Expansion (xW1+b1xW_1 + b_1)

The input vector xx is multiplied by a weight matrix W1W_1 and shifted by a bias vector b1b_1. This is a standard linear transformation, but with a twist: the output dimension is larger than the input dimension. If xx has 512 dimensions, the result might have 2048 dimensions. We're projecting into a higher-dimensional space where the data is easier to manipulate. In this expanded space, the model has more "room" to organize features before applying nonlinearity.

Why does projecting to higher dimensions help? Think of it this way: in the original dmodeld_{\text{model}}-dimensional space, the token's representation may conflate many different features in a compressed form. By expanding to dffd_{ff} dimensions, the first layer can disentangle these features, spreading them across a larger space where they are more separable. The nonlinearity can then act on these individual features independently, creating the curved boundaries that linear operations cannot.

Stage 2: The Nonlinearity (σ(⋅)\sigma(\cdot))

The activation function σ\sigma is applied element-wise to the expanded representation. Common choices include ReLU (which zeros out negative values) and GELU (a smoother alternative that we'll examine in detail in the next chapter). This is where the magic happens: the nonlinearity allows the network to learn curved decision boundaries that would be impossible with linear transformations alone.

The element-wise nature of the activation matters. Because it operates on each hidden dimension independently, it creates a kind of gating: some hidden dimensions activate strongly and contribute to the output, while others are suppressed to near zero. Different input patterns activate different subsets of hidden dimensions, which is exactly how the network implements different transformations for different types of inputs.

Stage 3: The Contraction ((⋅)W2+b2(\cdot)W_2 + b_2)

Finally, the transformed representation is projected back to the original dimension via weight matrix W2W_2 and bias b2b_2. The output has the same dimensionality as the input, which is essential for the residual connection that adds the FFN output back to its input. The second linear layer combines the activated hidden dimensions into a compressed dmodeld_{\text{model}}-dimensional output, essentially summarizing the computation done in the expanded space.

Why does this formula make sense? Notice that without the nonlinearity σ\sigma, the composition of two linear maps is itself a linear map: xW1W2+constx W_1 W_2 + \text{const} is just another linear function of xx. The nonlinearity is the only thing that prevents the two layers from collapsing into a single linear layer. The biases also allow the network to shift the activation function's operating range, making it easier to learn the right threshold for each hidden dimension.

The expansion ratio dff/dmodeld_{ff} / d_{\text{model}} is a key hyperparameter. The original transformer used a 4x expansion: for dmodel=512d_{\text{model}} = 512, the hidden dimension was dff=2048d_{ff} = 2048. This ratio has become a de facto standard, though modern architectures sometimes use different values, especially when combined with gated variants, as we'll explore later.

Implementation

With the formula understood, translating it to code reveals its simplicity. The entire FFN is just two matrix multiplications with a nonlinearity in between:

In[3]:
Code
import numpy as np


def relu(x):
    """ReLU activation function."""
    return np.maximum(0, x)


def ffn(x, W1, b1, W2, b2, activation=relu):
    """
    Position-wise feed-forward network.

    Args:
        x: Input tensor, shape (n, d_model) or (d_model,)
        W1: First layer weights, shape (d_model, d_ff)
        b1: First layer bias, shape (d_ff,)
        W2: Second layer weights, shape (d_ff, d_model)
        b2: Second layer bias, shape (d_model,)
        activation: Nonlinear activation function

    Returns:
        Output tensor, same shape as x
    """
    hidden = activation(x @ W1 + b1)
    output = hidden @ W2 + b2
    return output


# Example dimensions (from original transformer)
d_model = 512  # Model dimension
d_ff = 2048  # Hidden dimension (4x expansion)

# Initialize weights with Xavier/Glorot initialization
W1 = np.random.randn(d_model, d_ff) * np.sqrt(2.0 / (d_model + d_ff))
b1 = np.zeros(d_ff)
W2 = np.random.randn(d_ff, d_model) * np.sqrt(2.0 / (d_ff + d_model))
b2 = np.zeros(d_model)

# Process a single position
x_single = np.random.randn(d_model)
y_single = ffn(x_single, W1, b1, W2, b2)
Out[4]:
Console
Single position FFN:
  Input shape:  (512,)
  Output shape: (512,)
  Input norm:   22.2545
  Output norm:  13.1094

The FFN preserves the dimensionality of its input: a 512-dimensional vector goes in, and a 512-dimensional vector comes out. This is essential for the residual connection that adds the FFN output back to its input. The residual connection (which adds the original input xx to the FFN output) requires both tensors to have the same shape, so the FFN's dimension-preserving property is architecturally necessary rather than merely convenient.

Notice that we initialize the weights using Xavier/Glorot initialization, scaling by 2/(din+dout)\sqrt{2 / (d_{in} + d_{out})}. This initialization keeps the variance of activations stable across layers, preventing the exploding or vanishing gradients that plagued deep networks before careful initialization became standard practice. The biases are initialized to zero, which is standard; the weight matrices carry all the initial structure.

The Hidden Dimension Expansion

The most striking aspect of the FFN architecture is the dimension expansion. The first linear layer projects from dmodeld_{\text{model}} to dffd_{ff}, typically with dff=4×dmodeld_{ff} = 4 \times d_{\text{model}}. For a model with dmodel=512d_{\text{model}} = 512, this means expanding to 2048 dimensions before projecting back down. For GPT-3 with dmodel=12288d_{\text{model}} = 12288, the hidden dimension reaches 49152, nearly fifty thousand neurons per FFN layer.

This expansion might seem wasteful. Why create a 2048-dimensional intermediate representation just to immediately compress it back to 512 dimensions? The answer lies in the expressiveness of the network, and it connects to some deep results in neural network theory.

A fundamental result in approximation theory shows that wider hidden layers can approximate more complex functions. Consider what happens without expansion: if dff=dmodeld_{ff} = d_{\text{model}}, the hidden layer has the same dimensionality as the input. While the formula still applies, the network has limited capacity to decompose and recombine features. Each hidden neuron is a linear combination of the input features, gated by nonlinearity. With only dmodeld_{\text{model}} neurons, the network can detect dmodeld_{\text{model}} different patterns. With 4×dmodel4 \times d_{\text{model}} neurons, it can detect four times as many patterns and combine them in far richer ways.

Think of the expansion as temporarily working in a higher-dimensional space where the data is easier to manipulate. In the expanded 2048-dimensional space, the activation function can create very specific, highly discriminating patterns. One hidden neuron might activate only when the input represents a verb in past tense. Another might activate for proper nouns that appear after articles. These fine-grained detectors would be impossible to implement in 512 dimensions because there simply isn't enough room to disentangle all the features. The expansion provides the working space to learn them, and the contraction summarizes the results back into the original dimensionality for the residual connection.

The 4x expansion factor was established in the original "Attention Is All You Need" paper and has become a standard choice. It represents a pragmatic balance between expressiveness and efficiency: large enough to provide substantial nonlinear capacity, small enough to remain computationally tractable. Modern models like LLaMA use slightly different ratios (around 2.7x) when using gated linear units (GLUs), which effectively increase the hidden dimension through gating, but the underlying logic of expanding before contracting remains the same.

In[5]:
Code
# Visualize the dimension flow through the FFN
def analyze_ffn_dimensions(d_model, expansion_factor):
    """Analyze dimensions through FFN layers."""
    d_ff = d_model * expansion_factor

    return {
        "input": d_model,
        "after_W1": d_ff,
        "expansion_ratio": d_ff / d_model,
        "after_W2": d_model,
    }


# Common configurations
configs = [
    ("GPT-2 Small", 768, 4),
    ("GPT-2 Medium", 1024, 4),
    ("GPT-2 Large", 1280, 4),
    ("GPT-3 (175B)", 12288, 4),
    ("LLaMA-7B", 4096, 2.6875),  # Uses 11008 hidden dim
]
Out[6]:
Console
FFN dimension expansion across models:

Model               d_model     d_ff    Ratio
----------------------------------------------
GPT-2 Small             768     3072     4.00x
GPT-2 Medium           1024     4096     4.00x
GPT-2 Large            1280     5120     4.00x
GPT-3 (175B)          12288    49152     4.00x
LLaMA-7B               4096    11008     2.69x

The pattern is consistent: every model expands significantly in the hidden layer. LLaMA's slightly lower ratio reflects the use of the SwiGLU activation, a gated variant that includes an additional linear projection. Even with a nominally lower ratio, SwiGLU-based FFNs have similar or greater effective capacity because the gate mechanism provides additional expressive power.

Let's visualize how the actual activation values change through each stage of the FFN, to build intuition about what the expand-transform-contract process looks like in practice:

Out[7]:
Visualization
Histogram of input values showing normal distribution centered at zero.
Input values follow a normal distribution centered at zero. This reflects the typical distribution of token representations entering the FFN after layer normalization.
Histogram of hidden activations after ReLU showing only positive values and many zeros.
After ReLU, the distribution becomes one-sided: negative values collapse to exactly zero (red line) while positive values pass through unchanged, creating sparsity in the hidden layer.
Histogram of output values showing distribution around zero.
Output values return to a distribution around zero after the second linear projection contracts the sparse hidden representation back to d_model dimensions.

The three histograms tell an important story. The input distribution is centered at zero and roughly Gaussian, which is typical after layer normalization. After ReLU, the distribution becomes one-sided: roughly half the hidden units collapse to exactly zero (shown by the red line), while the other half pass through positive values unchanged. This sparsity is not a bug but a feature: different inputs will activate different subsets of hidden dimensions, creating a sparse code that allows the network to implement many different transformations. The output distribution returns to a balanced form around zero after the second linear layer combines the sparse hidden activations.

Now let's visualize the dimension sizes at each stage with a schematic:

Out[8]:
Visualization
Diagram showing dimension sizes at each FFN layer, with expansion from 512 to 2048 then contraction back to 512.
Dimension flow through the FFN. The input at d_model (512) expands through the first weight matrix W1 to the hidden layer at d_ff (2048, shown as a taller rectangle), passes through the nonlinear activation, then contracts back to d_model (512) via W2. The dramatic size difference illustrates why the 4x expansion dominates the FFN's parameter count.

Position Independence

A important property of the FFN is that it processes each position independently. Unlike attention, where every position can influence every other position through the query-key-value mechanism, the FFN applies an identical transformation to each position in isolation. This has practical consequences for both computation and interpretation, and it reflects a deep design decision about the division of responsibilities in the transformer architecture.

When we say the FFN is position-independent, we mean that the output at position ii depends only on the input at position ii. The FFN does not know or care about adjacent tokens. It does not look at the whole sequence. It simply applies the same transformation FFN(x)=σ(xW1+b1)W2+b2\text{FFN}(x) = \sigma(xW_1 + b_1)W_2 + b_2 to whatever vector it receives, and that vector could have come from any position in the sequence.

This might seem like a limitation. After all, language is deeply contextual. The meaning of "bank" depends on whether "river" or "money" appeared nearby. But here is the key insight: by the time the FFN sees the representation at position ii, that representation has already been processed by the self-attention sublayer in the same block. Attention has already incorporated contextual information. The FFN is operating on a context-enriched representation, not on a raw token embedding. Its job is to transform that already-contextualized representation further, not to gather more context.

Think of the transformer block as a two-phase processor: attention gathers context (inter-position), and the FFN processes the gathered context (intra-position). The FFN can focus entirely on the "what to do with this information" question because attention has already handled the "what context is relevant" question.

Let's verify this independence empirically:

In[9]:
Code
# Demonstrate position independence
seq_len = 5
X = np.random.randn(seq_len, d_model)

# Process all positions at once (batch processing)
Y_batch = ffn(X, W1, b1, W2, b2)

# Process each position individually
Y_individual = np.zeros_like(X)
for i in range(seq_len):
    Y_individual[i] = ffn(X[i], W1, b1, W2, b2)

# Check they're identical
difference = np.abs(Y_batch - Y_individual).max()
Out[10]:
Console
Position independence verification:
  Maximum difference between batch and individual processing: 3.77e-15
  Outputs are identical: True

The batch and individual processing produce identical results, up to floating-point precision. This confirms that positions don't interact within the FFN. Changing the token at position 3 has absolutely no effect on the FFN output at position 1, even if they share the same FFN weights. The information barrier between positions is absolute.

Position independence means the FFN can be computed in parallel across all positions. On a GPU, this is extremely efficient: instead of processing tokens sequentially, we process the entire sequence simultaneously as a batch matrix multiplication. The FFN is embarrassingly parallel, meaning no position needs to wait for any other position's result. For a sequence of 1024 tokens, we can compute all 1024 FFN outputs simultaneously, limited only by the size of the weight matrices and the amount of memory available.

Position independence also clarifies the division of labor in a transformer block. Attention handles inter-position communication: it routes information between tokens, allowing the model to build representations that depend on context. The FFN handles intra-position transformation: it transforms each position's representation using the same learned function, adding nonlinearity and processing capacity to what would otherwise be a purely linear attention mechanism. This division of labor is one of the transformer's most elegant architectural properties.

Out[11]:
Visualization
Diagram showing three input positions being processed independently through the same FFN to produce three output positions.
Position independence in the FFN illustrated for three sequence positions. Each input position flows vertically through the shared FFN (orange box), producing its output without any horizontal communication between positions. The label 'No interaction' highlights the absence of cross-position information flow, contrasting with the self-attention layer where every position can attend to every other position.

Interpreting the FFN as Key-Value Memory

One of the most illuminating ways to understand the FFN is to see it as a neural network layer and as an associative memory: a structure that stores key-value pairs in its weights and retrieves the appropriate values when presented with matching inputs. This interpretation, developed by researchers at Tel Aviv University and other institutions, helps explain how transformers store and retrieve factual knowledge, and why simply scaling up the FFN increases a model's capacity to "know" things.

The idea is elegant. Consider the first layer's weight matrix W1W_1. Each column of W1W_1 can be thought of as a "key" that matches certain input patterns. When we compute xW1+b1xW_1 + b_1, we are essentially computing the dot product between the input xx and each of the dffd_{ff} keys. Each element of the resulting vector measures how well the input matches the corresponding key.

To make this precise: the pre-activation value for hidden dimension ii measures the alignment between the current input and the ii-th key pattern. Given an input xx and the ii-th column of W1W_1 (call it ki\mathbf{k}_i), this alignment is:

hi=x⋅ki+b1,ih_i = x \cdot \mathbf{k}_i + b_{1,i}

where:

  • hih_i: the pre-activation value for hidden dimension ii, measuring how well the input matches key ii
  • xx: the input vector (dmodeld_{\text{model}}-dimensional), representing the current token's contextual state
  • ki\mathbf{k}_i: the ii-th column of W1W_1, encoding the pattern that key ii is tuned to detect
  • b1,ib_{1,i}: the ii-th element of the bias vector b1b_1, which acts as a threshold, biasing the key toward activation or suppression

When the input xx aligns well with key ki\mathbf{k}_i (high dot product), the corresponding hidden dimension activates strongly. After ReLU, only positive activations survive, so keys that match the current input "fire" while keys that don't match are silenced. This is exactly how an associative memory works: present a query, retrieve only the memories that match.

The second layer's weight matrix W2W_2 contains the "values" associated with each key. Each row of W2W_2 (call it vi\mathbf{v}_i for the ii-th row) is the value vector that gets added to the output when hidden dimension ii is active. The final FFN output is a weighted sum of these values, where the weights are the hidden activations:

FFN(x)=∑i=1dffhi+⋅vi\text{FFN}(x) = \sum_{i=1}^{d_{ff}} h_i^+ \cdot \mathbf{v}_i

where:

  • hi+=ReLU(hi)=max⁡(0,hi)h_i^+ = \text{ReLU}(h_i) = \max(0, h_i): the post-activation strength of the ii-th key match
  • vi\mathbf{v}_i: the ii-th row of W2W_2, encoding the value retrieved when key ii fires
  • The sum runs over all dffd_{ff} key-value pairs, but in practice only the active (nonzero) hi+h_i^+ values contribute
FFN as Key-Value Memory

The feed-forward layer can be interpreted as an associative memory where W1W_1 columns are keys, hidden activations are match scores, and W2W_2 rows are values. Input patterns that match certain keys retrieve their associated values. With dffd_{ff} key-value pairs per layer and dozens of layers per model, the total storage capacity of the FFN weights is enormous.

This memory interpretation explains several important empirical observations about transformers. Researchers have found that specific neurons in FFN layers activate for particular concepts: individual neurons that fire consistently for "the Eiffel Tower," for "programming languages," or for "past tense verbs." These neurons are not hand-designed. They emerge from training, as the model discovers that certain hidden dimensions reliably encode certain patterns. The FFN learns to store factual associations in its weights, and the attention mechanism retrieves them by routing relevant input patterns through the FFN.

This interpretation also explains a peculiar phenomenon called "knowledge editing" in language models. When researchers want to update a model's stored facts (for instance, changing "The president of France is X" to a newer value), they can sometimes accomplish this by modifying specific rows of FFN weight matrices. The change propagates naturally because those rows encode the associated value that gets retrieved when the "president of France" key fires.

The memory interpretation also clarifies why larger FFN dimensions store more knowledge. More hidden dimensions means more key-value pairs, which means more distinct patterns the model can recognize and more distinct responses it can retrieve. This is why scaling the FFN dimension is one of the primary ways to increase a model's factual capacity.

Let's visualize this interpretation with a concrete example:

In[12]:
Code
# Interpret FFN as key-value memory
d_model_small = 8
d_ff_small = 16

W1_small = np.random.randn(d_model_small, d_ff_small) * 0.5
b1_small = np.zeros(d_ff_small)
W2_small = np.random.randn(d_ff_small, d_model_small) * 0.5
b2_small = np.zeros(d_model_small)

# Create an input that strongly activates certain hidden dimensions
x_test = np.random.randn(d_model_small)

# Compute hidden activations (before and after ReLU)
pre_activation = x_test @ W1_small + b1_small
hidden_activations = relu(pre_activation)

# See which "keys" were matched (high activation)
active_dims = hidden_activations > 0.5
Out[13]:
Console
FFN as Key-Value Memory:
  Input dimension: 8
  Number of 'keys' (hidden dim): 16

Pre-activation (dot product with keys):
  [ 2.89 -0.3  -1.39  1.7   0.28  0.9   1.25 -0.09 -0.54  1.3  -2.01 -0.44
 -0.15  3.36  0.76 -1.39]

After ReLU (matched keys):
  [2.89 0.   0.   1.7  0.28 0.9  1.25 0.   0.   1.3  0.   0.   0.   3.36
 0.76 0.  ]

Strongly activated dimensions: [ 0  3  5  6  9 13 14]

Let's visualize this sparsity pattern across multiple inputs to see how different inputs activate different subsets of hidden dimensions:

Out[14]:
Visualization
Heatmap showing hidden activations across inputs and dimensions, with many zero values demonstrating sparsity.
Hidden activation patterns for 10 different random inputs passing through a small FFN with 16 hidden dimensions. Each row shows the activation values after ReLU for one input. White cells indicate zero activation (the feature did not match), while blue cells indicate positive activations (the key fired). Different inputs activate different subsets of dimensions. This shows that the FFN implements input-dependent sparse retrieval, similar to a content-addressable memory.

The heatmap reveals the key-value memory character directly. Each column represents one hidden dimension (one key-value pair). The white cells are zero, meaning that key did not fire for that input. The blue cells indicate firing keys, with darker blue indicating stronger activation. Notice that no two input samples have exactly the same activation pattern: different inputs retrieve different combinations of stored associations. This input-dependent sparse access is what makes the FFN such a powerful memory structure.

This sparsity also has practical implications for efficiency. In a model with large dffd_{ff}, only a fraction of hidden dimensions are nonzero for any given input. Methods that exploit this sparsity can skip computation for inactive dimensions, reducing the compute needed well below the theoretical maximum. Recent work on ReLU-based sparse FFNs has shown that sparsity rates of 90% or higher can be achieved with minimal quality loss, meaning only 10% of hidden dimensions are needed for most tokens.

Parameter Count Analysis

Feed-forward networks are the largest component of transformer models by parameter count. In a standard transformer, the FFN contains significantly more parameters than the attention mechanism. This matters for model scaling and efficiency optimization, and it helps explain why so many recent research directions (mixture of experts, weight sharing, factorization) target the FFN specifically.

For a single FFN layer, we count the total number of learnable parameters by summing the sizes of all weight matrices and bias vectors. The calculation is straightforward. Matrix W1W_1 has shape dmodel×dffd_{\text{model}} \times d_{ff}, contributing dmodel×dffd_{\text{model}} \times d_{ff} parameters. Matrix W2W_2 has shape dff×dmodeld_{ff} \times d_{\text{model}}, contributing an equal dff×dmodeld_{ff} \times d_{\text{model}} parameters. The two bias vectors contribute dffd_{ff} and dmodeld_{\text{model}} parameters respectively. Adding these together:

FFN parameters=2×dmodel×dff+dff+dmodel\text{FFN parameters} = 2 \times d_{\text{model}} \times d_{ff} + d_{ff} + d_{\text{model}}

where:

  • dmodel×dffd_{\text{model}} \times d_{ff}: the number of elements in weight matrix W1W_1 (rows ×\times columns)
  • dff×dmodeld_{ff} \times d_{\text{model}}: the number of elements in weight matrix W2W_2 (rows ×\times columns, equal to the count in W1W_1)
  • dffd_{ff}: the number of elements in bias vector b1b_1
  • dmodeld_{\text{model}}: the number of elements in bias vector b2b_2
  • The factor of 2 in front of dmodel×dffd_{\text{model}} \times d_{ff} accounts for both weight matrices W1W_1 and W2W_2 having the same total element count

With the standard expansion factor dff=4×dmodeld_{ff} = 4 \times d_{\text{model}}, we can simplify this expression. Substituting dff=4dmodeld_{ff} = 4 d_{\text{model}}:

FFN parameters=2×dmodel×(4×dmodel)+(4×dmodel)+dmodel=8×dmodel2+5×dmodel\begin{aligned} \text{FFN parameters} &= 2 \times d_{\text{model}} \times (4 \times d_{\text{model}}) + (4 \times d_{\text{model}}) + d_{\text{model}} \\ &= 8 \times d_{\text{model}}^2 + 5 \times d_{\text{model}} \end{aligned}

For large models where dmodeld_{\text{model}} is in the hundreds or thousands, the quadratic term dominates and the linear bias terms become negligible, giving approximately 8×dmodel28 \times d_{\text{model}}^2 parameters per FFN layer.

Why does this formula make sense? Notice that the parameter count grows quadratically with dmodeld_{\text{model}}: if you double the model dimension, you quadruple the FFN parameter count. This is the same quadratic scaling that makes attention expensive with respect to sequence length, but here the quadratic factor is in the model dimension rather than sequence length. Doubling dmodeld_{\text{model}} from 512 to 1024 increases FFN parameters from roughly 2 million to 8 million per layer, a 4x increase for a 2x dimension increase.

In[15]:
Code
def count_ffn_params(d_model, d_ff, include_bias=True):
    """Count parameters in a feed-forward network."""
    weight_params = 2 * d_model * d_ff  # W1 and W2
    bias_params = d_ff + d_model if include_bias else 0
    return weight_params + bias_params


def count_attention_params(d_model, num_heads, include_bias=True):
    """Count parameters in multi-head attention."""
    # Q, K, V projections and output projection
    weight_params = 4 * d_model * d_model  # W_Q, W_K, W_V, W_O
    bias_params = 4 * d_model if include_bias else 0
    return weight_params + bias_params


# Compare for different model sizes
model_configs = [
    ("GPT-2 Small", 768, 3072, 12),
    ("GPT-2 Medium", 1024, 4096, 16),
    ("BERT-Base", 768, 3072, 12),
    ("GPT-3 (175B)", 12288, 49152, 96),
]
Out[16]:
Console
Parameter count comparison: FFN vs Attention (per layer)

Model            d_model     d_ff     FFN Params    Attn Params   FFN/Attn
--------------------------------------------------------------------------------
GPT-2 Small          768     3072      4,718,592      2,359,296        2.0x
GPT-2 Medium        1024     4096      8,388,608      4,194,304        2.0x
BERT-Base            768     3072      4,718,592      2,359,296        2.0x
GPT-3 (175B)       12288    49152  1,207,959,552    603,979,776        2.0x

The FFN consistently contains about twice as many parameters as the attention mechanism per layer. This 2:1 ratio emerges from a straightforward calculation: attention uses four weight matrices of size dmodel×dmodeld_{\text{model}} \times d_{\text{model}} for QKV and output projections, totaling 4dmodel24 d_{\text{model}}^2 parameters. The FFN uses two matrices with total size 2dmodel×dff=2dmodel×4dmodel=8dmodel22 d_{\text{model}} \times d_{ff} = 2 d_{\text{model}} \times 4 d_{\text{model}} = 8 d_{\text{model}}^2 parameters. The ratio is exactly 8/4=28/4 = 2, so the FFN always has twice the weight parameters of attention when the 4x expansion factor is used.

For large models like GPT-3, each transformer layer has over 2.4 billion parameters in the FFN alone. Multiplied across 96 layers, the FFN accounts for the vast majority of the model's 175 billion total parameters. This makes FFN optimization critical for model efficiency, both in terms of memory and computation.

Let's visualize how FFN parameters scale with model dimension:

Out[17]:
Visualization
Line plot showing FFN parameters increasing quadratically from millions to billions as d_model increases from 256 to 12288.
FFN parameter count (blue circles) scales quadratically with model dimension, shown here on a log-log scale. The dashed red line shows the approximation 8 times d_model squared, which matches the exact count closely for large d_model values where bias terms are negligible. Doubling d_model quadruples the parameter count, which is why large language models devote such significant engineering effort to FFN compression.

Let's also visualize the parameter distribution in a transformer block to see the FFN's dominance at a glance:

Out[18]:
Visualization
Pie chart showing FFN with ~67% of parameters, attention with ~33%, and layer norm with <1%.
Parameter distribution across components in a single GPT-2 Small transformer block (d_model=768). The feed-forward network accounts for roughly two-thirds of all learnable parameters, with the attention mechanism holding most of the remainder. Layer normalization parameters are negligible in comparison. This distribution motivates the heavy focus on FFN optimization in efficient transformer research.

Computational Cost

Beyond parameter count, we need to consider computational cost, measured in floating-point operations (FLOPs). Understanding FFN compute requirements helps explain why these layers dominate inference time for short sequences, and why long-context models face a different bottleneck.

The FFN's computational cost per token follows directly from its structure. For each token, the FFN performs two matrix-vector multiplications. Each multiply-add operation consists of one multiplication and one addition, counting as 2 floating-point operations. Working through the arithmetic:

The first layer computes xW1xW_1 where xx is a dmodeld_{\text{model}}-dimensional vector and W1W_1 is dmodel×dffd_{\text{model}} \times d_{ff}. This requires dmodel×dffd_{\text{model}} \times d_{ff} multiply-adds, so 2×dmodel×dff2 \times d_{\text{model}} \times d_{ff} FLOPs. The second layer computes hW2hW_2 where hh is dffd_{ff}-dimensional and W2W_2 is dff×dmodeld_{ff} \times d_{\text{model}}, requiring another 2×dff×dmodel2 \times d_{ff} \times d_{\text{model}} FLOPs. The total FLOPs per token for the FFN is therefore:

FLOPsFFN per token=2×(dmodel×dff)+2×(dff×dmodel)=4×dmodel×dff\text{FLOPs}_{\text{FFN per token}} = 2 \times (d_{\text{model}} \times d_{ff}) + 2 \times (d_{ff} \times d_{\text{model}}) = 4 \times d_{\text{model}} \times d_{ff}

where:

  • dmodeld_{\text{model}}: the input and output dimension of the FFN
  • dffd_{ff}: the hidden dimension (typically 4×dmodel4 \times d_{\text{model}})
  • The factor of 4 comes from: 2 layers ×\times 2 FLOPs per multiply-add operation

For a sequence of nn tokens, each token is processed independently, so the total FFN cost scales linearly with sequence length. Because the FFN has no interaction between positions, the computation at each position is identical and independent:

Total FLOPsFFN=4×n×dmodel×dff\text{Total FLOPs}_{\text{FFN}} = 4 \times n \times d_{\text{model}} \times d_{ff}

where nn is the number of tokens in the sequence. This is O(n)O(n) scaling: double the sequence length, double the FLOPs.

This linear scaling contrasts sharply with attention, which has quadratic complexity O(n2⋅dk)O(n^2 \cdot d_k) due to the n×nn \times n attention matrix computation. For every pair of positions, attention must compute a query-key dot product. With nn positions, there are n2n^2 such pairs. As sequences grow longer, attention's quadratic term eventually dominates.

The crossover point depends on model dimensions. For a model with dmodel=768d_{\text{model}} = 768 and dff=3072d_{ff} = 3072, let's compute when attention FLOPs equal FFN FLOPs. FFN FLOPs scale as 4n×768×3072≈9.4×106×n4n \times 768 \times 3072 \approx 9.4 \times 10^6 \times n. Attention's quadratic term (for the attention matrix alone) scales as 4n2×768≈3×103×n24n^2 \times 768 \approx 3 \times 10^3 \times n^2. Setting these equal: 9.4×106×n=3×103×n29.4 \times 10^6 \times n = 3 \times 10^3 \times n^2, giving n≈3000n \approx 3000 tokens. At sequence lengths below 3000, FFN dominates. Above 3000, attention takes over.

In[19]:
Code
def ffn_flops(n, d_model, d_ff):
    """Calculate FFN FLOPs for sequence length n."""
    return 4 * n * d_model * d_ff


def attention_flops(n, d_model):
    """Calculate attention FLOPs (simplified)."""
    # Q, K, V projections: 3 * 2 * n * d_model^2
    # QK^T: 2 * n^2 * d_model
    # Attention @ V: 2 * n^2 * d_model
    # Output projection: 2 * n * d_model^2
    projection_flops = 4 * 2 * n * d_model * d_model
    attention_matrix_flops = 4 * n * n * d_model
    return projection_flops + attention_matrix_flops


# Compare across sequence lengths
sequence_lengths = [128, 512, 1024, 2048, 4096, 8192]
d_model_test = 768
d_ff_test = 3072
Out[20]:
Console
FLOPs comparison: FFN vs Attention (GPT-2 Small)

  Seq Length       FFN FLOPs      Attn FLOPs     FFN/Attn
------------------------------------------------------------
         128   1,207,959,552     654,311,424         1.85x
         512   4,831,838,208   3,221,225,472         1.50x
        1024   9,663,676,416   8,053,063,680         1.20x
        2048  19,327,352,832  22,548,578,304         0.86x
        4096  38,654,705,664  70,866,960,384         0.55x
        8192  77,309,411,328 244,813,135,872         0.32x

For short sequences (128-512 tokens), the FFN contributes over twice the FLOPs of attention. As sequences grow to 4096 or 8192 tokens, attention's quadratic cost catches up and eventually dominates. The crossover point around 1024-2048 tokens is where attention and FFN have comparable costs for this model configuration.

Out[21]:
Visualization
Line plot showing FFN FLOPs as a straight line and attention FLOPs as a curve that crosses over and grows faster at longer sequences.
Computational cost in FLOPs for FFN (red) and attention (blue) as sequence length increases, shown on a log-log scale. FFN cost grows linearly while attention cost grows quadratically. The vertical dotted line marks the crossover point where attention surpasses the FFN in compute. Below this crossover, optimizing the FFN yields greater efficiency gains; above it, attention efficiency becomes the priority.

This crossover point has important practical implications. Short-context applications like question answering and classification, as well as code completion, typically operate with sequences well below the crossover point, meaning the FFN is the primary compute bottleneck. Research into FFN efficiency (quantization, pruning, low-rank approximation) offers the greatest gains here. Long-context applications like document summarization, long-form generation, and retrieval-augmented generation operate above the crossover, making attention efficiency (sparse attention, linear attention, sliding window attention) the priority.

A Complete Worked Example

The formula FFN(x)=σ(xW1+b1)W2+b2\text{FFN}(x) = \sigma(xW_1 + b_1)W_2 + b_2 is compact, but its compactness can obscure what happens at each step. To understand the FFN, let's trace through a concrete computation with actual numbers. We'll use intentionally small dimensions (dmodel=3d_{\text{model}} = 3 and dff=4d_{ff} = 4) so you can follow every multiplication and addition by hand if you like, verifying each step against the formulas.

The goal of this example is threefold: to see exactly how the input vector is transformed at each stage, to observe how ReLU zeros out negative activations to introduce nonlinearity, and to verify that the output dimension matches the input dimension, ready for the residual connection. We'll use simple, hand-readable weight values to keep the arithmetic transparent.

In[22]:
Code
# Small example for hand-traceable computation
d_model_tiny = 3
d_ff_tiny = 4

# Initialize weights with simple values
W1_tiny = np.array(
    [
        [0.5, -0.3, 0.8, 0.2],
        [-0.2, 0.6, 0.1, -0.4],
        [0.3, 0.1, -0.5, 0.7],
    ]
)  # Shape: (3, 4)

b1_tiny = np.array([0.1, -0.1, 0.2, 0.0])  # Shape: (4,)

W2_tiny = np.array(
    [
        [0.4, -0.2, 0.3],
        [0.1, 0.5, -0.1],
        [-0.3, 0.2, 0.4],
        [0.2, -0.4, 0.1],
    ]
)  # Shape: (4, 3)

b2_tiny = np.array([0.05, -0.05, 0.1])  # Shape: (3,)

# Input vector
x_tiny = np.array([1.0, -0.5, 0.8])

We have chosen an input with mixed positive and negative values (x=[1.0,−0.5,0.8]x = [1.0, -0.5, 0.8]), which is representative of what you'd see in practice after layer normalization. The weight matrices use a mix of positive and negative values. This ensures some hidden dimensions will have positive pre-activations (and survive ReLU) while others will be zeroed out.

Stage 1: Expansion via the First Linear Layer

The first operation computes xW1+b1xW_1 + b_1. Our 3-dimensional input vector gets multiplied by a 3×43 \times 4 weight matrix, producing a 4-dimensional hidden representation. Each element of this output is a weighted sum of the input elements plus a bias term. Specifically, the jj-th element of the pre-activation is x[0]⋅W1[0,j]+x[1]⋅W1[1,j]+x[2]⋅W1[2,j]+b1[j]x[0] \cdot W_1[0,j] + x[1] \cdot W_1[1,j] + x[2] \cdot W_1[2,j] + b_1[j]: a dot product between the input and the jj-th column of W1W_1.

Out[23]:
Console
Step 1: First linear transformation (x @ W1 + b1)

Input x: [ 1.  -0.5  0.8]

W1:
[[ 0.5 -0.3  0.8  0.2]
 [-0.2  0.6  0.1 -0.4]
 [ 0.3  0.1 -0.5  0.7]]

b1: [ 0.1 -0.1  0.2  0. ]

x @ W1 + b1 = [ 0.94 -0.62  0.55  0.96]

Notice the dimension change: a 3-dimensional vector goes in, and a 4-dimensional vector comes out. This expansion is the "higher-dimensional working space" we discussed earlier, where the network has more room to manipulate the representation before applying nonlinearity. The result is called the pre-activation because we haven't applied the nonlinearity yet. Some of these pre-activation values are positive and some are negative, which sets up the selective filtering that ReLU will perform in the next step.

Stage 2: Nonlinearity via ReLU

Here is where the FFN gains its expressive power. The ReLU activation function applies a simple rule: keep positive values unchanged, but set negative values to zero. Mathematically, ReLU(z)=max⁡(0,z)\text{ReLU}(z) = \max(0, z). This is one of the simplest possible nonlinear functions, yet it is sufficient to make the FFN a universal approximator when the hidden dimension is large enough.

Out[24]:
Console

Step 2: Apply ReLU activation

Pre-activation: [ 0.94 -0.62  0.55  0.96]
After ReLU:     [0.94 0.   0.55 0.96]

Negative values become zero, positive values pass through unchanged.

By zeroing out some dimensions, ReLU creates sparse hidden representations: only a subset of hidden dimensions are active for any given input. Different inputs activate different subsets, allowing the network to learn piece-wise linear functions that approximate arbitrary curves. Each hidden dimension can be thought of as a detector for a particular pattern in the input. When the detector fires (positive pre-activation), it contributes to the output. When it doesn't fire (negative pre-activation, zeroed by ReLU), it contributes nothing. The network learns during training which patterns each detector should respond to, and what contribution to make when it does respond.

The key insight here is that ReLU's behavior changes based on which side of zero the pre-activation falls. For pre-activations far above zero, the gradient of ReLU is 1 and learning proceeds normally. For pre-activations below zero, the gradient is 0 and the network cannot update that hidden dimension for that particular input. This creates a form of implicit regularization: hidden dimensions that consistently fail to fire for a category of inputs effectively "opt out" of processing that category, allowing specialization.

Out[25]:
Visualization
Bar chart showing pre-activation values with some negative (red) and some positive (blue).
Pre-activation values before ReLU. Negative values (red bars) indicate hidden dimensions where the input pattern did not match the corresponding key. These values will be zeroed out by ReLU.
Bar chart showing values after ReLU with negative values now at zero.
After ReLU: the two negative pre-activations (shown as red bars at zero) have been suppressed. Only the positive activations (blue bars) contribute to the final output, creating the sparse hidden representation characteristic of ReLU-based FFNs.

Stage 3: Contraction via the Second Linear Layer

Finally, we project back to the original dimension. The 4-dimensional hidden vector is multiplied by a 4×34 \times 3 weight matrix, and a 3-dimensional bias is added. This contraction reduces the dimension and combines the contributions of all active hidden dimensions into an output vector. Each output dimension receives a weighted sum of the active hidden dimensions, where the weights are the corresponding elements of W2W_2.

Think of the second linear layer as the "summarization" step. The first layer asked "which features of this input are relevant?" The nonlinearity filtered out irrelevant dimensions. The second layer asks "given these relevant features, what transformation should we apply to the representation?"

Out[26]:
Console

Step 3: Second linear transformation (hidden @ W2 + b2)

Hidden: [0.94 0.   0.55 0.96]

W2:
[[ 0.4 -0.2  0.3]
 [ 0.1  0.5 -0.1]
 [-0.3  0.2  0.4]
 [ 0.2 -0.4  0.1]]

b2: [ 0.05 -0.05  0.1 ]

hidden @ W2 + b2 = [ 0.453 -0.512  0.698]

The output has the same dimension as the input (3 elements), which is essential. In the full transformer, this output will be added to the original input via a residual connection, and both must have the same shape. The FFN has transformed the representation while preserving its dimensionality, enabling the residual addition that stabilizes training in deep networks.

Verification and Summary

Let's verify that our step-by-step calculation matches the complete FFN function applied directly:

Out[27]:
Console

Verification using ffn() function:
  Output: [ 0.453 -0.512  0.698]
  Match: True

Summary:
  Input:  [ 1.  -0.5  0.8] (dimension 3)
  Output: [ 0.453 -0.512  0.698] (dimension 3)

The step-by-step and function-based computations match exactly. This worked example demonstrates the complete journey: a 3-dimensional input expands to 4 dimensions, passes through nonlinearity (with some dimensions zeroed out), and contracts back to 3 dimensions. The transformation is nonlinear (different inputs will activate different subsets of hidden dimensions, producing qualitatively different transformations), position-independent (the same computation would apply to any token's representation), and dimension-preserving (the output shape matches the input shape for the residual connection).

Implementation: A Complete FFN Module

Having traced through the mathematics by hand, we can now build a reusable FFN module that encapsulates everything we've learned. This implementation follows patterns used in production transformer libraries: it initializes weights using proper scaling (Xavier/Glorot initialization), supports optional biases, and handles both single vectors and batched sequences without any special casing.

naturally the class structure maps onto the mathematical formula. The __init__ method creates the four learnable parameter tensors (W1W_1, b1b_1, W2W_2, b2b_2). The __call__ method implements the formula σ(xW1+b1)W2+b2\sigma(xW_1 + b_1)W_2 + b_2 in three readable lines. The num_parameters method computes the formula we derived earlier. The entire forward pass is just three matrix multiplications (including the bias additions). This reflects the fundamental simplicity of the FFN despite its expressive power.

In[28]:
Code
class FeedForwardNetwork:
    """
    Position-wise feed-forward network for transformer blocks.

    Implements: FFN(x) = activation(x @ W1 + b1) @ W2 + b2
    """

    def __init__(self, d_model, d_ff, activation=relu, use_bias=True):
        """
        Initialize the feed-forward network.

        Args:
            d_model: Input and output dimension
            d_ff: Hidden dimension (typically 4 * d_model)
            activation: Nonlinear activation function
            use_bias: Whether to include bias terms
        """
        self.d_model = d_model
        self.d_ff = d_ff
        self.activation = activation
        self.use_bias = use_bias

        # Xavier/Glorot initialization
        self.W1 = np.random.randn(d_model, d_ff) * np.sqrt(
            2.0 / (d_model + d_ff)
        )
        self.W2 = np.random.randn(d_ff, d_model) * np.sqrt(
            2.0 / (d_ff + d_model)
        )

        if use_bias:
            self.b1 = np.zeros(d_ff)
            self.b2 = np.zeros(d_model)
        else:
            self.b1 = None
            self.b2 = None

    def __call__(self, x):
        """
        Apply the feed-forward transformation.

        Args:
            x: Input tensor of shape (..., d_model)

        Returns:
            Output tensor of shape (..., d_model)
        """
        # First linear layer
        hidden = x @ self.W1
        if self.use_bias:
            hidden = hidden + self.b1

        # Activation
        hidden = self.activation(hidden)

        # Second linear layer
        output = hidden @ self.W2
        if self.use_bias:
            output = output + self.b2

        return output

    def num_parameters(self):
        """Return total parameter count."""
        params = self.d_model * self.d_ff + self.d_ff * self.d_model
        if self.use_bias:
            params += self.d_ff + self.d_model
        return params
In[29]:
Code
# Test the module
ffn_module = FeedForwardNetwork(d_model=512, d_ff=2048)

# Single vector
x_single_test = np.random.randn(512)
y_single_test = ffn_module(x_single_test)

# Batch of vectors (sequence)
x_batch_test = np.random.randn(16, 512)  # 16 tokens
y_batch_test = ffn_module(x_batch_test)
Out[30]:
Console
FeedForwardNetwork module test:

Configuration:
  d_model: 512
  d_ff: 2048
  Parameters: 2,099,712

Single vector:
  Input shape: (512,)
  Output shape: (512,)

Batch (sequence):
  Input shape: (16, 512)
  Output shape: (16, 512)

The module correctly handles both single vectors and batched sequences. When a 2D input with shape (n,dmodel)(n, d_{\text{model}}) is provided (a batch of nn token representations), the matrix multiplication automatically broadcasts across all nn positions simultaneously. This is the source of the FFN's computational efficiency: instead of calling the function once per token in a loop, the entire sequence is processed in a single pair of batch matrix multiplications.

The parameter count matches our theoretical formula: 2×512×2048+2048+512=2,099,200+2,560=2,101,7602 \times 512 \times 2048 + 2048 + 512 = 2{,}099{,}200 + 2{,}560 = 2{,}101{,}760, close to 2.1 million parameters just for this single FFN layer. A 12-layer GPT-2 Small model has 12 such layers, plus attention layers, adding up to the 117 million total parameters commonly cited.

Limitations and Impact

The feed-forward network is conceptually simple: two linear layers with a nonlinearity between them. Yet this simplicity masks significant computational cost, and understanding those costs is essential for anyone working with large language models in practice.

The FFN's most significant limitation is its memory footprint. Because FFN weights are static after training, they must reside in memory at inference time. For a model like LLaMA-70B with a hidden dimension of 8192 and an FFN dimension of 28672 (approximately 3.5x), each FFN layer holds about 2 ×\times 8192 ×\times 28672 ≈\approx 470 million parameters. Across 80 layers, the FFN alone stores roughly 37 billion parameters, each requiring at least 2 bytes in FP16 format, meaning over 74 GB of memory just for FFN weights. This is why running large models requires high-memory GPUs or specialized hardware.

Weight quantization is the most common approach to this memory problem. By representing weights in lower precision (INT8, INT4, or even INT2), the memory footprint shrinks proportionally. A model quantized to INT4 requires one-quarter the memory of its FP16 counterpart, enabling deployment on consumer hardware. The challenge is maintaining model quality: aggressive quantization can degrade generation quality, particularly for reasoning tasks that require precise numerical computations in the FFN's hidden layers. Research into quantization-aware training and calibration has made significant progress, but the tradeoff between compression and quality remains an active area of work.

Sparsity offers another path forward. The ReLU activation naturally creates sparse hidden representations, as negative pre-activations become zero. Researchers have exploited this by identifying which hidden dimensions will be active for a given input and computing only those, skipping computation for dimensions that would be zeroed anyway. This requires predicting active neurons before the full computation, which introduces overhead, but the net savings can be substantial for highly sparse models. The GELU and SiLU activations that have replaced ReLU in many modern models are smoother alternatives that do not produce exact zeros, trading the sparsity benefit for better gradient flow during training.

The Mixture of Experts (MoE) architecture takes the sparsity idea further. Instead of a single FFN per layer, an MoE layer has multiple "expert" FFNs and a router that directs each token to only a small subset (typically 2 of 8, or 2 of 64). The model scales parameters without proportionally scaling compute, because only a fraction of experts process each token. Models like Mixtral-8x7B and GPT-4 (reportedly) use MoE to achieve parameter counts that would otherwise be computationally prohibitive. The key challenge is load balancing: ensuring that all experts are used approximately equally across the training data, preventing some experts from specializing in everything while others specialize in nothing.

Despite its computational weight, the FFN's role is needed and irreplaceable. It provides the nonlinearity that maps attention's linear weighted averages into expressive function approximation. It is the model's "memory," storing factual associations in its weights that attention retrieves based on context. Its position-wise nature enables massive parallelization that makes transformer training tractable at the scale of billions of parameters and trillions of training tokens.

The interplay between attention and FFN defines transformer expressiveness. Attention routes information between positions, creating context-dependent representations. The FFN transforms those representations position-by-position, adding computational depth and nonlinear capacity. Together, they enable the hierarchical, compositional language understanding that powers modern NLP. Neither component alone would be sufficient: attention without the FFN is a purely linear contextual aggregation; the FFN without attention is a powerful token-level classifier with no awareness of surrounding context.

Summary

The feed-forward network is the workhorse of transformer computation, applying identical transformations to each position independently. This chapter covered its architecture and efficiency, then examined its interpretation within the broader transformer block.

Key takeaways:

  • Two-layer architecture: The FFN formula FFN(x)=σ(xW1+b1)W2+b2\text{FFN}(x) = \sigma(xW_1 + b_1)W_2 + b_2 consists of an expansion layer (dmodel→dffd_{\text{model}} \to d_{ff}), a nonlinear activation σ\sigma (ReLU, GELU, etc.), and a contraction layer (dff→dmodeld_{ff} \to d_{\text{model}}). The standard expansion factor is dff=4×dmodeld_{ff} = 4 \times d_{\text{model}}, establishing a working space four times larger than the model dimension where nonlinear transformations are easier to learn.

  • Position independence: Each position is processed separately with shared weights. This enables parallel computation and cleanly separates the FFN's role (per-position nonlinear transformation) from attention's role (inter-position information routing). The position independence is not a limitation but a design choice: by the time the FFN processes each token, attention has already incorporated contextual information into that token's representation.

  • Key-value memory interpretation: The FFN can be viewed as an associative memory where W1W_1 columns are keys, hidden activations are match scores, and W2W_2 rows are values. Different inputs activate different subsets of hidden dimensions, creating sparse, input-dependent retrievals. This explains how transformers store factual knowledge and why larger FFN dimensions increase a model's capacity to "know" things.

  • Dominant parameter count: The FFN contains approximately two-thirds of a transformer block's parameters. With 4x expansion, the total is approximately 8×dmodel28 \times d_{\text{model}}^2 parameters per layer (from the formula 2×dmodel×dff+dff+dmodel2 \times d_{\text{model}} \times d_{ff} + d_{ff} + d_{\text{model}}). This quadratic scaling with dmodeld_{\text{model}} makes FFN optimization critical for model efficiency.

  • Linear computational scaling: FFN compute cost is 4×n×dmodel×dff4 \times n \times d_{\text{model}} \times d_{ff} FLOPs, scaling linearly with sequence length nn. This contrasts with attention's quadratic O(n2)O(n^2) scaling. For short sequences, FFN dominates compute; for long sequences, attention becomes the bottleneck.

  • Crossover point: Around 1000-2000 tokens, FFN and attention have comparable computational costs for typical model sizes. This crossover point shapes optimization strategy: short-context systems prioritize FFN compression, while long-context systems prioritize attention efficiency.

The next chapter examines activation functions used in FFNs in depth, comparing ReLU, GELU, SiLU/Swish, and their gated variants (GLU, SwiGLU), and exploring why modern models have moved beyond the original ReLU choice toward smoother, more expressive alternatives.

Key Parameters

When implementing or configuring feed-forward networks in transformers, these parameters control capacity and efficiency:

  • d_model (input/output dimension): The embedding dimension that the FFN preserves. Typical values range from 256 (small models) to 12288 (GPT-3 scale). This dimension must match the attention layer output and determines the FFN's interface with the rest of the transformer.

  • d_ff (hidden dimension): The expanded dimension of the intermediate representation. The standard choice is dff=4×dmodeld_{ff} = 4 \times d_{\text{model}}, though modern architectures like LLaMA use ratios around 2.7x when combined with gated linear units. Larger values increase expressiveness but proportionally increase parameters and compute.

  • activation: The nonlinear function applied element-wise after the first linear layer. ReLU was used in the original transformer, but GELU has become the standard for encoder models (BERT, RoBERTa) and SiLU/Swish for decoder models (LLaMA, GPT-NeoX). The choice affects gradient flow, sparsity patterns, and training stability.

  • use_bias: Whether to include bias terms b1b_1 and b2b_2. Some modern architectures (LLaMA, PaLM) omit biases entirely to reduce parameters and simplify quantization. The impact on model quality is typically minimal.

  • dropout (not shown in our implementation): Dropout rate applied to the hidden representation after activation. Values of 0.1-0.2 are common during training to prevent overfitting. Set to 0 during inference.

  • initialization: Weight initialization scale affects training stability. Xavier/Glorot initialization (scaling by 2/(din+dout)\sqrt{2/(d_{in} + d_{out})}) is standard. Some architectures use scaled initialization for residual paths to maintain signal magnitude through deep networks.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about feed-forward networks in transformers.

Feed-Forward Networks Quiz

Question 1 of 80 of 8 completed
What is the primary purpose of the feed-forward network (FFN) in a transformer block?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025feedforward, author = {Michael Brenndoerfer}, title = {Feed-Forward Networks in Transformers: Architecture}, year = {2025}, url = {https://mbrenndoerfer.com/writing/transformer-feed-forward-networks}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Feed-Forward Networks in Transformers: Architecture. Retrieved from https://mbrenndoerfer.com/writing/transformer-feed-forward-networks
MLAAcademic
Michael Brenndoerfer. "Feed-Forward Networks in Transformers: Architecture." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/transformer-feed-forward-networks>.
CHICAGOAcademic
Michael Brenndoerfer. "Feed-Forward Networks in Transformers: Architecture." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/transformer-feed-forward-networks.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Feed-Forward Networks in Transformers: Architecture'. Available at: https://mbrenndoerfer.com/writing/transformer-feed-forward-networks (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Feed-Forward Networks in Transformers: Architecture. https://mbrenndoerfer.com/writing/transformer-feed-forward-networks

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.