Kinds of learning
Supervised, unsupervised, self-supervised, reinforcement
Learning from examples with known answers is one common approach to machine learning. Four common approaches differ largely in one way: where does the signal for learning come from? These categories are useful rather than perfectly separate: real systems often combine them.
Supervised: learning from an answer key
In supervised learningsupervised learningTraining on examples where each input is paired with a human-provided correct answer (a label). Regression and classification both live here; this explainer focuses on it.See in glossary →, every training example arrives with the correct answer attached. You show the model a house’s size and tell it the price; a photo and tell it “cat.” The model guesses, compares against the known answer, and adjusts. That paired data (input plus label) is the “supervision.”
Two shapes of answer live here. When the answer is a quantity (a price, a temperature) it’s regressionregressionA prediction task whose answer is a number (a price, a temperature), as opposed to a category. Usually scored with squared error.See in glossary →. When the answer is a category (spam or not, cat or dog) it’s classificationclassificationA prediction task whose answer is one of a fixed set of categories (spam / not spam, or which of 100,000 tokens comes next). The model outputs a probability per option and is scored with cross-entropy.See in glossary →. Both are supervised; they differ only in what the label looks like. A house-price predictor and a spam classifier make useful starting examples because their answers are easy to specify.
Unsupervised: finding structure with no answers
Sometimes there are no labels at all: just a pile of data. Unsupervised learningunsupervised learningFinding structure in unlabeled data — grouping, compressing, or otherwise organizing it — with no answer key to imitate.See in glossary → asks the model to find structure on its own. Two classic jobs: clusteringclusteringAn unsupervised task that groups examples resembling one another, without being told what the groups are.See in glossary →, which groups examples that resemble each other (say, sorting shoppers into natural segments no one defined in advance), and dimensionality reductiondimensionality reductionCompressing many correlated features into a few underlying dimensions that capture most of the variation.See in glossary →, which replaces a long list of measurements with a shorter one that retains useful patterns, especially when the original measurements tend to vary together. Nobody hands over a “right answer”; the signal is the shape of the data itself.
Self-supervised: labels for free
Large language models can get their answer key from the raw data. In self-supervised learningself-supervised learningTraining where the labels come from the data itself — e.g. hide part of an example and ask the model to predict it. No human annotation needed.See in glossary →, you hide part of the input and ask the model to predict the hidden part. Cover part of a sentence: the missing piece is right there in the text, so it doubles as the label. No human labeling required, yet every guess can still be scored against a known truth.
This is the core trick behind training many LLMs: turn raw text into billions of “predict what comes next” puzzles, each one supervised by the text that actually followed. It’s technically supervised learning, but the supervision comes from the data itself and can be produced at enormous scale. The same prediction-and-adjustment loop can therefore learn from text without a person supplying each answer.
Reinforcement: learning from reward
The fourth approach drops the per-example answer key. In reinforcement learningreinforcement learningLearning from trial and error: an agent takes actions and receives a single-number reward signal, with no labeled "right answer" for each step.See in glossary →, an agentagentA system that chooses actions, observes their effects in an environment, and uses those observations to decide what to do next.See in glossary →, the system choosing what to do, takes actions in an environmentenvironmentThe world or software system an agent interacts with, which provides observations and responds to its actions.See in glossary →, the setting it interacts with. It tries to collect as much rewardrewardA single-number feedback signal in reinforcement learning that says how good an outcome was. The agent tries to collect as much reward as possible over time.See in glossary → as possible over time: a numerical score for the outcomes of its actions. Actions affect what happens (and therefore what the agent sees and can do) next. Rewards may arrive at every step or only occasionally. Chess illustrates the difficult sparse, delayed case: a bot isn’t told the best move for each board; it may learn from only the win or loss at the end, leaving the algorithm to work out which of forty moves deserves credit. Game-playing agents and, increasingly, fine-tuningfine-tuningContinuing to train a pre-trained model on a smaller, task- or behavior-specific dataset. This explainer is about pre-training; fine-tuning and other post-training steps are out of scope.See in glossary →, or further training, of LLMs live here.
Supervised learning makes the training loop especially easy to examine: there is an input, a known answer, and a measurable error. Before a model can use any of them, sizes, colors, and categories must become numbers it can calculate with.