Batches, epochs & noise
Why we estimate the gradient from samples
An optimizer needs a gradientgradientThe vector of partial derivatives of the loss with respect to every parameter — it points in the direction of steepest loss increase, so we step the opposite way to reduce the loss.See in glossary → to choose its update. The negative gradient gives the steepest downhill direction at the current parameters. But the lossloss functionA single number measuring how wrong the model's predictions are on a batch of data. Training works by adjusting the model to make this number smaller.See in glossary → we actually care about is an average over the whole dataset, and computing its exact gradient means running every example through the model before taking a single step. With a huge dataset, that would make each step painfully expensive. Estimating it from a sample makes each update much cheaper.
Estimate the gradient from a sample
The trick is a staple of statistics: to estimate an average, you don’t need the whole population. A random sample will do. So each step we draw a small random mini-batchmini-batchThe chunk of training examples processed together in one step. Gradients are averaged over the mini-batch, trading off gradient noise against memory and compute.See in glossary → of examples, compute the gradient on just those, and step. If the batch is sampled fairly, its gradient is an unbiased estimate of the true one (correct on average) but any single batch is a bit off, so the estimate is noisy.
That’s the whole idea of stochastic gradient descentSGDStochastic Gradient Descent — gradient descent using a noisy gradient estimated from one mini-batch at a time rather than the whole dataset.See in glossary →: trade an exact, unaffordable gradient for a cheap, slightly-wrong one, and take many more steps to make up for it. In practice it’s not even close: the cheap version wins by orders of magnitude.
Three words worth pinning down
- A batchbatchThe group of training examples used for one gradient estimate. Bigger batches reduce gradient noise but use more memory and compute per step.See in glossary → is the group of examples used for one gradient estimate; the batch size is how many.
- A steptraining stepOne iteration of the loop: forward pass on a batch, backward pass to get gradients, optimizer update. A large model is trained for hundreds of thousands of steps.See in glossary → is one update: one batch in, one nudge to the parameters out.
- An epochepochOne full pass over the training dataset. Frontier LLMs are often trained for roughly a single epoch over a deduplicated corpus, so each token is seen about once.See in glossary → is one full pass through the entire dataset: dataset size ÷ batch size steps.
Training is many steps grouped into epochs; the loss is expected to fall across steps, with each epoch usually revisiting the whole dataset in a fresh random order.
The noise is a feature and a bug
Turn the batch size down and watch the descent path jitter; turn it up and the path smooths toward the clean, exact-gradient curve. The learning rate is held fixed here, so the only thing changing is how noisy each step’s gradient is.
Small batches wander, but that wandering isn’t purely a defect. A little randomness can rattle the marble out of a shallow dip or off a plateau that a perfectly smooth descent would have settled into. This is why moderate batch noise can help generalizationgeneralizationHow well a model performs on data it never saw during training. The whole point of pre-training is to generalize, not to memorize the corpus.See in glossary →: how well the model performs on examples it did not train on.
We can now train efficiently: estimate the gradient from a batch, step with an optimizer, repeat for many epochs. A linear model still has a limit, however: changing its weights cannot make its boundary curve. A neural network overcomes that limit by combining weighted sums with functions that bend their outputs.