How a model learns
Gradient descent and backpropagation
We have an objective (minimize the cross-entropy loss) and a model with billions of parametersparametersThe numbers (weights) inside a model that get adjusted during training. A “7B model” has 7 billion of them.See in glossary → to adjust. The update method is gradient descentgradient descentThe core training algorithm: repeatedly nudge each parameter a small step in the direction that lowers the loss, as told by the gradient.See in glossary →, powered by backpropagationbackpropagationThe algorithm that computes the loss gradient for every parameter efficiently by applying the chain rule backward through the network, reusing intermediate results from the forward pass.See in glossary →. Repeated over many examples, these updates turn initially random parameters into a useful predictor.
The loss is a landscape; training walks downhill
Fix the training data and think of the loss as a function of the parameters alone. With billions of parameters this loss landscapeloss landscapeThe (extremely high-dimensional) surface of loss as a function of the parameters. Training is a walk downhill on this surface toward a low-loss region.See in glossary → lives in billions of dimensions, but the three-dimensional picture carries over: two parameters spread across the ground, height is the loss, and we want to reach a low valley.
The tool for going downhill is the gradientgradientThe vector of partial derivatives of the loss with respect to every parameter — it points in the direction of steepest loss increase, so we step the opposite way to reduce the loss.See in glossary →, (the symbol is read “nabla” or “del”): a vectorvectorAn ordered list of numbers, such as the measurements describing one example or the learned values representing a token.See in glossary →, or ordered list of numbers, containing one partial derivativepartial derivativeThe rate at which a function changes when one input changes while all its other inputs are held fixed.See in glossary → per parameter. Each derivative measures how the loss changes when that parameter changes a little while the others stay fixed. It points in the direction of steepest increase, so to decrease the loss we step the opposite way:
The scalar (the Greek letter eta) is the learning ratelearning rateThe size of each parameter step. Too high and training can diverge; too low and it crawls.See in glossary →: how big a step to take. That single update rule, repeated, is gradient descent.
Play with the learning rate above and the central tension of all training appears immediately. Too small and progress is glacial. Too large and the ball overshoots the valley, zig-zags, or flies off entirely. And because the surface is steeper in one direction than the other, no single learning rate is ideal for both axes at once. A real network has billions of axes with wildly different curvatures. An optimizer, the rule that converts gradients into parameter updates, can adjust its steps to those differences.
Stochastic gradient descent: don’t read the whole library each step
The true loss is an average over the entire corpus. Computing its exact gradient would mean a forward pass over trillions of tokens for a single update. Absurd. Instead we estimate the gradient from a mini-batchmini-batchThe chunk of training examples processed together in one step. Gradients are averaged over the mini-batch, trading off gradient noise against memory and compute.See in glossary →, a small group of sequences sampled from the data. The estimate is noisy, but it is unbiased, meaning correct on average, and millions of times cheaper, and the noise even helps escape bad regions. This is stochastic gradient descentSGDStochastic Gradient Descent — gradient descent using a noisy gradient estimated from one mini-batch at a time rather than the whole dataset.See in glossary → (stochastic means involving randomness), and one such update is a training steptraining stepOne iteration of the loop: forward pass on a batch, backward pass to get gradients, optimizer update. A large model is trained for hundreds of thousands of steps.See in glossary →.
A frontier run is hundreds of thousands to millions of these steps. Teams often avoid repeatedly sampling the same general-web data when fresh data is available, but curated sources can be deliberately upsampled and data can recur across training stages. The key idea is that, when plentiful fresh data exists, repeated exposure usually has diminishing returns. Scaling laws describe how these returns change with data and model size.
Backpropagation: getting a billion derivatives for the price of one pass
The catch is step 3. We need for every parameter: billions of derivatives. Computing each one independently would be hopeless. BackpropagationbackpropagationThe algorithm that computes the loss gradient for every parameter efficiently by applying the chain rule backward through the network, reusing intermediate results from the forward pass.See in glossary → computes them all in a single backward sweep, and it is the algorithm that makes deep learning feasible.
The chain rulechain ruleThe calculus rule for differentiating composed functions. Backpropagation is just the chain rule applied layer by layer, from the loss back to the inputs.See in glossary → calculates the sensitivity of a chain of operations by combining the sensitivities of its individual steps. A neural network is a long composition of simple operations (multiplication of matricesmatrixA rectangular array of numbers arranged in rows and columns. Matrix multiplication combines these arrays through weighted sums.See in glossary →, rectangular arrays of numbers; normalizationnormalizationRescaling values to control their size or spread. In neural networks, this helps keep intermediate calculations at useful numerical scales.See in glossary →, which rescales values to control their size; and nonlinear functionsnonlinear functionA function whose output isn't just a scaled, shifted copy of its input — e.g. ReLU, GELU, sigmoid. Stacking nonlinearities between matrix multiplies is what lets a neural net represent anything more interesting than scaling and rotation.See in glossary →, which do more than scale and add inputs). In the forward passforward passRunning inputs through the network to produce outputs (logits) and the loss, caching intermediate activations that backpropagation will need.See in glossary → we run inputs through to produce the loss, caching the intermediate activationsactivationsThe values produced by a layer after applying its activation function. During training, intermediate activations are often kept for the backward pass.See in glossary →, the values produced by the model’s operations. In the backward passbackward passThe second half of a training step: backpropagation walks from the loss back through the network, computing each parameter's gradient.See in glossary → we walk the operations in reverse, and at each one multiply the gradient flowing back by that operation’s local derivative. Each layer needs only its cached inputs and the gradient arriving from the layer above, so the whole gradient costs about the same as two forward passes, regardless of how many parameters there are.
Two consequences of backprop shape everything downstream:
- Memory. The backward pass needs the forward activations, so they must be kept in memory until used. At long context lengthcontext lengthThe maximum number of tokens the model can attend to at once (also called the context window or sequence length). Pre-training picks a context length; later stages often extend it.See in glossary → these activations balloon, which is why gradient checkpointinggradient checkpointingActivation recomputation — saving memory by discarding most activations in the forward pass and recomputing them during the backward pass, trading extra compute for far less memory.See in glossary → (recomputing activations instead of storing them) becomes essential at scale.
- Numerics. Gradients are computed through many multiplications and can become very small or very large. Keeping them representable is the job of the precision machinery: rescaling small gradients, choosing suitable number formats, and controlling the size of intermediate values.
You rarely write backprop by hand; frameworks like PyTorch build a computational graph during the forward pass and differentiate it automatically. But knowing that the gradient is cheap to compute but expensive to store explains an enormous amount about how large models are actually engineered.
With gradients in hand, the remaining question is how to use them well, which is the optimizer.