What is pre-training?
Learning from raw text, no labels required
When people say a model “knows” things (that Paris is in France, that Python numbers list positions starting at zero, that a sonnet has fourteen lines), much of that broad knowledge was first acquired during one process: pre-trainingpre-trainingThe first phase of building a language model: training on an enormous corpus of raw text to predict the next token, learning general-purpose language ability before any task-specific tuning.See in glossary →. It is the initial, broad training of a large language modelLLMLarge Language Model — a neural network trained on huge text corpora to predict the next token given previous tokens.See in glossary →, a model that processes and generates text.
The training task is simple to state. Take a large collection of text (large, curated slices of web text alongside books, code, and other sources) and train a neural networkneural networkA function built by stacking many simple operations — mostly matrix multiplies with nonlinearities between them — whose behavior is shaped by tuning billions of internal numbers (its parameters) from data.See in glossary →, a mathematical function whose adjustable numbers are learned from examples, to do one thing: predict the next word. Do that at enough scale, for long enough, and the network is forced to learn grammar, facts, reasoning patterns, translation, arithmetic, and code, because all of those are useful for guessing what comes next. No one ever tells it the rules of French or the syntax of Python. It infers them, because inferring them lowers its prediction error.
Three phases: pre-training, post-training, inference
Building and using a language model involves three broad phases.
- Pre-training learns a foundation modelfoundation modelA large model pre-trained on broad data that can be adapted to many downstream tasks. The pre-trained LLM is the foundation; fine-tuning specializes it.See in glossary →, a general-purpose model that can be adapted to many tasks, from raw text. This is where the parametersparametersThe numbers (weights) inside a model that get adjusted during training. A “7B model” has 7 billion of them.See in glossary → (the billions of numbers that are the model) acquire most of the broad capabilities later stages build on. It can take months on tens of thousands of graphics processing units (GPUs)GPUGraphics Processing Unit — a processor designed to perform many calculations in parallel, widely used for training and running neural networks.See in glossary →, chips that perform many arithmetic operations in parallel.
- Post-training then continues updating that foundation through fine-tuningfine-tuningContinuing to train a pre-trained model on a smaller, task- or behavior-specific dataset. This explainer is about pre-training; fine-tuning and other post-training steps are out of scope.See in glossary →, further training on more targeted examples or feedback, to shape behavior and strengthen targeted capabilities. It typically uses far less diverse data and compute than pre-training, but can still teach the model to follow instructions, reason in particular ways, or use tools.
- InferenceinferenceRunning a trained model to produce outputs. Training learns the weights once; inference uses them many times.See in glossary → is running the finished model to answer your prompts.
The split matters because the phases do fundamentally different jobs. Pre-training is hugely expensive (a frontier run can cost tens of millions of dollars) and establishes most of the model’s broad knowledge and general-purpose competence. Post-training shapes how that knowledge is expressed and can strengthen targeted skills, such as following instructions, refusing harmful requests, or reasoning step by step. That is why a new “frontier model” is news: somebody ran pre-training again, bigger or better, and the foundation moved.
Learning from the text itself
In supervised learningsupervised learningTraining on examples where each input is paired with a human-provided correct answer (a label). Regression and classification both live here; this explainer focuses on it.See in glossary →, you need labeled examples, photos tagged “cat” or “dog,” sentences tagged with their sentiment. Labels are made by humans, so they are scarce and expensive. You will never hand-label a trillion examples.
Next-token prediction sidesteps this entirely. Given the text “the cat sat on the”, the “label” for what comes next is just the actual next word in the document: “mat”. The data labels itself. This is self-supervised learningself-supervised learningTraining where the labels come from the data itself — e.g. hide part of an example and ask the model to predict it. No human annotation needed.See in glossary →, and it is the reason pre-training can consume trillions of tokenstokenThe atomic unit of text the model sees. Roughly a word-fragment — “tokenization” is a piece of text → list of token IDs.See in glossary →, the pieces of text, often words or word fragments, that a language model processes: every sentence ever written is already a pile of free training examples, one per position.
Why scale works at all
The bet underneath modern AI is that prediction is understanding in disguise. To predict the next token of a murder mystery’s final page, it helps to have tracked the plot. To predict the next line of a proof, it helps to have learned the math. To complete a function, it helps to know the application programming interface (API)APIApplication Programming Interface — a defined way for one piece of software to request data or actions from another.See in glossary →, the rules for calling a piece of software. The pressure to predict well, applied across a broad enough corpuscorpusThe body of text a model is trained on. Modern pre-training corpora are measured in trillions of tokens drawn from web crawls, books, code, and more.See in glossary →, or collection of texts, pushes the network to build internal machinery that looks a lot like knowledge and reasoning.
For a long time it was not obvious this would keep paying off. It did, and remarkably smoothly. Bigger models trained on more data get predictably better, following clean mathematical curves described by the scaling-laws chapters. That predictability is what justified spending ever-larger sums: you could forecast the payoff before you paid.
A pre-trained model is valued precisely because it generalizesgeneralizationHow well a model performs on data it never saw during training. The whole point of pre-training is to generalize, not to memorize the corpus.See in glossary →: it performs well on downstream tasksdownstream taskAny specific job (translation, question answering, coding) a pre-trained model is later applied to. Pre-training is deliberately task-agnostic so it transfers to many downstream tasks.See in glossary →, the applications it is used for after its initial training, and text it never saw. It can still memorize some training examples: deduplication and limiting repeated exposure reduce redundant training and memorization risk, but do not eliminate either.
The science and engineering of pre-training
Five parts connect the basic training task to large-scale model development:
- Foundations: how prediction errors are measured, how parameters are adjusted, how numbers are stored, how work is distributed across GPUs, and how training data is prepared. This is the machinery every model shares.
- The transformer and the paradigm: the papers that invented the architecture and the pre-train-then-adapt recipe. Attention Is All You Need, GPT-1, BERT, GPT-2, T5.
- Scaling laws: how the field learned to predict and budget training. Kaplan, GPT-3, Chinchilla.
- The modern era: what today’s open models actually do differently. Llama 3, DeepSeek-V3, Qwen2.5, the Gemmas, synthetic data.
- The 2026 frontier: the latest reports, with their changes to training methods and systems.
A training method has to work mathematically and fit the available hardware. The prediction error links the two: it defines what the model should improve and what the training system must repeatedly calculate.