Section 01

What is post-training?

Turning a base model into an assistant

Sources: Training language models to follow instructions with human feedback — Ouyang et al., 2022; Direct Preference Optimization — Rafailov et al., 2023; DeepSeek-R1 — DeepSeek-AI, 2025

Ask a freshly pre-trained language model “What is the capital of France?” and you might get this back:

What is the capital of France? What is the largest city in France? What is the official language of France? List three famous French painters.

It didn’t answer. It did something stranger and, on reflection, completely logical: it continued the document. Somewhere in its training data lived a quiz, a worksheet, a list of trivia questions, and the single most likely thing to follow one question is another question. The model did exactly what it was built to do. It is a brilliant autocomplete. It is also a terrible assistant.

A chat product combines a post-trained model with product-level components such as tool runtimes and safety systems. Post-trainingpost-trainingEverything done to a model after pre-training to turn a raw next-token predictor into a useful assistant: supervised fine-tuning, RLHF, and RL from verifiable rewards.See in glossary → is the phase that teaches the model to answer questions, follow instructions, and prefer particular response styles rather than merely continue text. It changes the model’s training signal to favor useful responses.

Two models, one set of weights

Pre-training, the subject of the sibling explainer, produces a base modelbase modelA model straight out of pre-training — a powerful text continuator that has not yet been taught to follow instructions, hold a conversation, or refuse harmful requests.See in glossary →: a neural networkneural networkA function built by stacking many simple operations — mostly matrix multiplies with nonlinearities between them — whose behavior is shaped by tuning billions of internal numbers (its parameters) from data.See in glossary →, a mathematical function whose adjustable numbers are learned from examples, trained on large curated collections of text and other data to do one thing, predict the next tokennext-token predictionThe pre-training objective for GPT-style models: given the tokens so far, predict a probability distribution over the next token. Also called causal or autoregressive language modeling.See in glossary →, a piece of text such as a word or word fragment. That objective encourages it to absorb grammar, facts, code, and reasoning patterns, because all of those help it guess what comes next. The result is a model with broad learned capabilities but no training signal that specifically says how to respond to a user.

The reason is subtle but important. The base model’s training distribution is the broad mix pre-training draws on: web pages, books, code, and other primary sources. Very little of that text is a helpful assistant responding to a user. It is articles, forum flame wars, half-finished code, published prose, and, yes, trivia worksheets. When you prompt the base model, it doesn’t ask “how would a helpful assistant respond?” It asks “what is the most likely continuation of this text?” And that continuation is often unhelpful, repetitive, or actively wrong, because plenty of unhelpful, repetitive, wrong text exists.

Post-training often reshapes how existing capabilities are expressed, though it can also teach new facts, skills, and tool-use conventions from its own data. The capital of France may already be represented in a base model; what can be missing is the disposition to answer directly, in a helpful tone, while declining requests it should not honor. Full fine-tuningfine-tuningContinuing to train a pre-trained model on a smaller, task- or behavior-specific dataset. This explainer is about pre-training; fine-tuning and other post-training steps are out of scope.See in glossary →, further training of an existing model, updates the same parametersparametersThe numbers (weights) inside a model that get adjusted during training. A “7B model” has 7 billion of them.See in glossary →, the model’s adjustable numbers, set during pre-training, while parameter-efficient methods can train adaptersadapterA small set of extra parameters trained on top of a frozen base model, so fine-tuning updates only the adapter rather than the full network. Adapters cut memory and storage cost and can be swapped in and out per task.See in glossary → instead: small added sets of parameters that modify the model’s behavior. Both usually use far less data than pre-training, but their cost and ease of iteration depend on the goals and method.

The boundary between the two phases isn’t perfectly sharp. Many modern pipelines insert a mid-trainingmid-trainingA phase between the main pre-training run and post-training, used to inject specialized data or capabilities (e.g. long context, code-from-execution) while still training the base model on a next-token-style objective.See in glossary → phase in between: continued next-token training on a mixture selected to strengthen particular abilities (math, code, reasoning traces, long-context data) that still looks like pre-training in mechanism, covered in the pre-training explainer. Post-training then shapes how the resulting base model responds to users.

A note on cost

The relative cost of pre-training and post-training depends on the methods and goals. Pre-training is a massive one-time compute bill: a frontier run can cost tens of millions of dollars in GPUGPUGraphics Processing Unit — a processor designed to perform many calculations in parallel, widely used for training and running neural networks.See in glossary → time. Graphics processing units are chips that perform many arithmetic operations in parallel. But modern post-training is far from trivial: it involves collecting human preference labels at scale, training auxiliary reward models, and running reinforcement-learning loops that repeatedly sample from the model. Reinforcement learning (RL)reinforcement learningLearning from trial and error: an agent takes actions and receives a single-number reward signal, with no labeled "right answer" for each step.See in glossary → improves behavior using numerical rewards for sampled actions. Reasoning-focused RL can spend substantial computation generating and grading rolloutsrolloutA complete generated sample from the policy — for an LLM, one full response to a prompt. RL collects rollouts, scores them, and updates the policy.See in glossary →, complete attempts at a task.

The phases spend their budgets differently. Pre-training uses large amounts of arithmetic to learn patterns from data. Post-training spends a mix of human labor, careful data curation, and a different shape of compute to build behavior. Neither is categorically “the expensive one.” It depends on the model, the goals, and the year.

The post-training stack

Here is the whole arc at a glance. A base model goes through some subset of these stages, in roughly this order:

  1. Supervised fine-tuning (SFT) / instruction tuning. Show the model high-quality examples of instructions paired with good responses, and continue next-token training on those examples, usually computing loss on the assistant-response tokens. This teaches the model the format of being an assistant: that a user turn should be followed by a helpful answer, not another question. This is instruction tuning.

  2. Preference optimization. SFT learns from demonstrations, while preference data provides a relative signal about which of two generated sample responses is better. This is the heart of RLHFRLHFReinforcement Learning from Human Feedback — train a reward model on human preference comparisons, then optimize the policy against that reward with RL (typically PPO), with a KL leash to a reference.See in glossary → (reinforcement learning from human feedback), where a learned reward modelreward model (RM)A model trained from human preference data to output a scalar score for how good a response is. Stands in for a human judge so RL can query reward millions of times.See in glossary → uses human comparisons to score further responses, reducing the need to label every new example.

  3. Reinforcement learning from verifiable rewards (RLVR) / reasoning RL. For tasks where correctness can be checked automatically (math with a known answer or code that must pass tests), we can use a verifier rather than a learned reward model. Optimizing such signals with specialized reinforcement-learning algorithms has been an important ingredient in recent reasoning-model pipelines, alongside data, prompting, and other training choices.

The map below makes this navigable. Each node is a stage; click through to see how they connect.

The post-training stack
How a raw pretrained model becomes an aligned reasoning assistant — click any stage to see what it does and jump to its chapter.
Base modelpretrained

The raw pretrained language model. It has absorbed broad world knowledge from next-token prediction over a huge corpus, but it only continues text — it has not yet been taught to follow instructions, hold a conversation, or behave like a helpful assistant.

The modern post-training stack — click any stage to jump to its chapter.

Three eras, briefly

It helps to see how this stack assembled itself historically, because each layer was a response to the limits of the last.

  • 2021–2022: instruction tuning and RLHF. FLAN and T0 showed that fine-tuning on instructions phrased in natural language can improve zero-shotzero-shotPerforming a task from instructions alone, with no examples given. GPT-2 showed a pre-trained LM can do many tasks zero-shot, just by being prompted.See in glossary → performance: handling a new task from its instruction without example answers in the prompt. InstructGPT established the influential three-step recipe (SFT, then a reward model, then reinforcement-learning optimization against it) that later informed many chat-model pipelines. Alignment became a training problem, not just a prompting trick.

  • 2023: offline and direct methods. RLHF’s RL loop can be finicky and expensive to run. A direct, offline method showed that, under a particular preference-modeling formulation, a loss on preference pairs can change the model’s response probabilities without a separate reward model or new responses sampled during each update. A wave of variants followed, broadening the use of offline preference optimization.

  • 2024–2026: verifiable rewards and reasoning. OpenAI’s o1 highlighted the impact of reinforcement learning on reasoning, though its full training recipe is not public. DeepSeek-R1-Zero demonstrated reasoning behavior emerging from pure RL, while DeepSeek-R1 combined cold-start data and multi-stage RL. Verifiable rewards are now an important ingredient in reasoning and agentic RL research.

Measuring and improving behavior

Post-training changes the probabilities of the model’s responses. Its objectives measure which responses become more likely, how far the model changes, and how varied its outputs remain. Choosing what to reward also requires defining the desired behavior: the alignment problem.

The main methods build on demonstrations, comparisons, or checked outcomes. Instruction tuning learns from examples; preference methods learn which answers people favor; verifiable rewards support training on tasks such as mathematics and programming. Longer tasks extend the same problem to sequences of tool calls and actions.

Each method needs both a suitable objective and a workable training process. The quality of the data, the stability of the updates, and failures in the feedback signal determine how well the mathematical goal translates into useful behavior.