The MLP block
Per-token nonlinear processing
Sources: GLU Variants Improve Transformer — Shazeer, 2020; The Llama 3 Herd of Models — Grattafiori et al., 2024
In a standard dense transformer block, attention lets tokens look at each other. The other main sublayer works on each token in isolation. That sublayer is called the MLPMLPMulti-Layer Perceptron — a stack of dense (matrix-multiply + nonlinearity) layers applied per-token. The transformer’s feed-forward block.See in glossary → block, short for Multi-Layer Perceptron, a network that transforms inputs through layers of weighted sums and nonlinear functions.
If attention is the social half of a transformer layer, the MLP is the private half: every token gets pulled aside, run through the same little network, and handed back with its representation refined.
What’s actually inside
The calculation has four steps:
- Take the token’s vector (size , say 4,096).
- Project it up to a much wider vector, the expanded layer (size , usually 3–4× wider).
- Apply a nonlinearity to each entry separately that decides which entries of that wider vector “fire.”
- Project the result back down to .
That’s it. Two matrix multiplications with a nonlinearity in between. The MLP output is added back to the input through the residual connectionresidual connectionoutput = x + f(x). Lets gradients flow through deep stacks and means each block adds a refinement rather than rewriting.See in glossary →, which adds a block’s output to its input, and we move on.
A useful way to think about the middle “expanded” layer: it’s a collection of feature detectors. Each entry asks something like “does this token look like a verb of motion?” or “does the residual stream look like it’s building up to a comma?”. The nonlinearity decides whether each detector fires; the down-projection blends the firing detectors back into a new vector, which the rest of the model reads.
The detectors aren’t designed. They’re learned. Probing studies can sometimes find individual MLP neurons with surprisingly clean correlations, such as indented Python code or quoted strings. But representations are often distributed across many features, so a single neuron should not be treated as a complete, human-readable concept. The wide middle layer holds a large share of the model’s parameters and is an important site of learned computation.
Why there has to be a nonlinearity
Stack two matrix multiplies with nothing in between and they collapse into a single matrix multiply. A matrix multiply is a linear transformation, so without the nonlinearity the MLP could not represent conditional, nonlinear behavior such as “if A and B, then turn on C.” The nonlinearity is the bend that lets the network do more than a linear transformation.
The choice of nonlinearity and gating does matter, though the basic role is the same. The original Transformer used ReLU. GPT-2 used GELUGELUGaussian Error Linear Unit — a smooth nonlinearity used inside the MLP. SiLU/SwiGLU are common modern variants.See in glossary →. Modern Llama-class models use a gated variant called SwiGLUSwiGLUA gated MLP variant (Llama, PaLM): output = SiLU(xW₁) ⊙ (xW₂), then projected. Outperforms plain MLPs at the same param count.See in glossary →, which adds a second up-projection that modulates the first. Under matched model budgets, gated variants often outperform simpler ungated activations, but the effect depends on the architecture and training setup.
The MLP holds most of the weights
The parameter counts show why MLPs matter for inference speed. For a Llama-3-8B layer:
- Attention projections: roughly 42 million parameters per layer.
- MLP: roughly 176 million parameters per layer, over 4× the attention projections.
Multiplied by 32 layers, that’s about 5.6 billion of the model’s 8 billion parameters living in MLP blocks. Most of a modern LLM, by parameter count, is feed-forward networks. Moving those weights takes a substantial share of inference time.
This is one reason small-batch dense decoding is memory-bandwidth-boundmemory-boundLimited mainly by the speed of moving data to the processor, rather than by the speed of performing calculations.See in glossary →: limited by how quickly weights can be read from memory. Every new token requires reading most of those MLP weights from HBMHBMHigh-Bandwidth Memory — the DRAM stack soldered next to the GPU die. H100 SXM has 80 GB at ~3.35 TB/s.See in glossary → (the GPU’s high-bandwidth main memory) for only one token’s worth of math. Compared with the arithmetic speed of tensor coresTensor CoreA specialized unit in an NVIDIA GPU that performs blocks of matrix arithmetic at high speed.See in glossary →, the GPU units specialized for matrix multiplication, moving those bytes is the bottleneck. The chapters on prefill and decode and GPU memory examine the resulting limits.
Position-wise, not sequence-wise
The MLP runs the same way at every position, independently. There is no mixing between tokens inside this sublayer. In a standard transformer block, attention is the cross-token mixing operation and the MLP is the main per-token transformation; residual connections and normalization, rescaling intermediate values to control their size, support both. A transformer layer alternates these components.
A complete transformer block combines attention and the MLP with shortcuts that preserve the input and normalization that controls the scale of intermediate values.