Section 02

Kinds of learning

Supervised, unsupervised, self-supervised, reinforcement

Learning from examples with known answers is one common approach to machine learning. Four common approaches differ largely in one way: where does the signal for learning come from? These categories are useful rather than perfectly separate: real systems often combine them.

Supervised: learning from an answer key

In supervised learningsupervised learningTraining on examples where each input is paired with a human-provided correct answer (a label). Regression and classification both live here; this explainer focuses on it.See in glossary →, every training example arrives with the correct answer attached. You show the model a house’s size and tell it the price; a photo and tell it “cat.” The model guesses, compares against the known answer, and adjusts. That paired data (input plus label) is the “supervision.”

Two shapes of answer live here. When the answer is a quantity (a price, a temperature) it’s regressionregressionA prediction task whose answer is a number (a price, a temperature), as opposed to a category. Usually scored with squared error.See in glossary →. When the answer is a category (spam or not, cat or dog) it’s classificationclassificationA prediction task whose answer is one of a fixed set of categories (spam / not spam, or which of 100,000 tokens comes next). The model outputs a probability per option and is scored with cross-entropy.See in glossary →. Both are supervised; they differ only in what the label looks like. A house-price predictor and a spam classifier make useful starting examples because their answers are easy to specify.

Unsupervised: finding structure with no answers

Sometimes there are no labels at all: just a pile of data. Unsupervised learningunsupervised learningFinding structure in unlabeled data — grouping, compressing, or otherwise organizing it — with no answer key to imitate.See in glossary → asks the model to find structure on its own. Two classic jobs: clusteringclusteringAn unsupervised task that groups examples resembling one another, without being told what the groups are.See in glossary →, which groups examples that resemble each other (say, sorting shoppers into natural segments no one defined in advance), and dimensionality reductiondimensionality reductionCompressing many correlated features into a few underlying dimensions that capture most of the variation.See in glossary →, which replaces a long list of measurements with a shorter one that retains useful patterns, especially when the original measurements tend to vary together. Nobody hands over a “right answer”; the signal is the shape of the data itself.

Self-supervised: labels for free

Large language models can get their answer key from the raw data. In self-supervised learningself-supervised learningTraining where the labels come from the data itself — e.g. hide part of an example and ask the model to predict it. No human annotation needed.See in glossary →, you hide part of the input and ask the model to predict the hidden part. Cover part of a sentence: the missing piece is right there in the text, so it doubles as the label. No human labeling required, yet every guess can still be scored against a known truth.

This is the core trick behind training many LLMs: turn raw text into billions of “predict what comes next” puzzles, each one supervised by the text that actually followed. It’s technically supervised learning, but the supervision comes from the data itself and can be produced at enormous scale. The same prediction-and-adjustment loop can therefore learn from text without a person supplying each answer.

Reinforcement: learning from reward

The fourth approach drops the per-example answer key. In reinforcement learningreinforcement learningLearning from trial and error: an agent takes actions and receives a single-number reward signal, with no labeled "right answer" for each step.See in glossary →, an agentagentA system that chooses actions, observes their effects in an environment, and uses those observations to decide what to do next.See in glossary →, the system choosing what to do, takes actions in an environmentenvironmentThe world or software system an agent interacts with, which provides observations and responds to its actions.See in glossary →, the setting it interacts with. It tries to collect as much rewardrewardA single-number feedback signal in reinforcement learning that says how good an outcome was. The agent tries to collect as much reward as possible over time.See in glossary → as possible over time: a numerical score for the outcomes of its actions. Actions affect what happens (and therefore what the agent sees and can do) next. Rewards may arrive at every step or only occasionally. Chess illustrates the difficult sparse, delayed case: a bot isn’t told the best move for each board; it may learn from only the win or loss at the end, leaving the algorithm to work out which of forty moves deserves credit. Game-playing agents and, increasingly, fine-tuningfine-tuningContinuing to train a pre-trained model on a smaller, task- or behavior-specific dataset. This explainer is about pre-training; fine-tuning and other post-training steps are out of scope.See in glossary →, or further training, of LLMs live here.

Which kind of learning is this?
Pick a task. The families differ mostly in one thing: where the training signal comes from, and whether a human had to label anything.
SupervisedLearn from input → known answer pairs a human provided.
what plays the "answer"
The correct tag "cat" or "dog" attached to each training photo.
where labels come from
A human labeled every photo by hand — expensive but explicit.
Different as they look, the same pattern keeps returning: produce something, score it against some signal, and change the parameters to do better. The source of that signal changes — a human label, the data’s own structure, a hidden piece of the input, or a reward.

Supervised learning makes the training loop especially easy to examine: there is an input, a known answer, and a measurable error. Before a model can use any of them, sizes, colors, and categories must become numbers it can calculate with.