From this loop to an LLM
The same loop, where the prediction is the next token
The training loop is input → model → prediction → loss → gradient → update, looped until the model generalizes. A large language model uses that same basic machine. The pieces are chosen so that the “prediction” is what comes next in human text. The inputs, predictions, and dataset change; the method for adjusting parameters remains the same.
The training loop applied to text
| In this explainer | In an LLM |
|---|---|
| The input | A stretch of text: the tokens seen so far |
| The prediction | A probability for every possible next tokentokenThe atomic unit of text the model sees. Roughly a word-fragment — “tokenization” is a piece of text → list of token IDs.See in glossary → |
| The model | A transformertransformerA neural-network architecture introduced in "Attention Is All You Need" (2017), built from stacked self-attention and feed-forward layers.See in glossary →: deep stacks of the same linear-plus-nonlinear layers |
| The loss | Cross-entropycross-entropy lossA classification loss that penalizes the model according to the negative log-probability it assigned to the correct answer.See in glossary →: the surprise at the true next token |
| The training loop | Forward, loss, backprop, update: unchanged |
| The dataset | Trillions of tokens of text from books, web, and code |
The prediction is the next token
During its initial text training, an LLM’s main job is next-token predictionnext-token predictionThe pre-training objective for GPT-style models: given the tokens so far, predict a probability distribution over the next token. Also called causal or autoregressive language modeling.See in glossary →: given the text so far, output a probability distribution over the entire vocabularyvocabularyThe fixed set of tokens a model knows about. Modern LLMs have ~32k–200k entries.See in glossary →, the fixed list of tokens the model can choose from. That’s a classification problem (“which of these ~100,000 options comes next?”) scored with the same loss as other classification tasks. The model is scored by cross-entropy: the negative log of the probability it placed on the token that actually came next. Confidently predict the right next token and the loss is near zero; be blindsided and it’s large. It is the same measure of surprise used for other classification tasks.
The model is bigger, but it’s the same kind of thing
The transformertransformerA neural-network architecture introduced in "Attention Is All You Need" (2017), built from stacked self-attention and feed-forward layers.See in glossary → looks intimidating from the outside, but structurally it’s what you’ve built: layers of weighted sums with nonlinearities between them, trained by gradient descent via backpropagation. Its special ingredient — attention — is a particular layer that lets each position gather information from earlier ones, so context flows through the network. Attention is trained along with the rest of the network. The knobs are still knobs; there are just hundreds of billions of them, and backprop still hands each one its gradient in a single backward pass.
Training at scale
The LLM Pre-training explainer examines the engineering needed for scale: how the same basic loop is run on tens of thousands of GPUs, over trillions of tokens, with the numerical and engineering tricks that make a run that big finish at all. At its center is the same prediction, loss, gradient, and update cycle. The surrounding machinery gets more complex; the core idea stays recognizable.