Measuring prediction error
Loss functions and the cost of a mistake
The loss is the number the whole machine is trying to make small, so it’s worth getting right. A useful loss has to reward predictions we actually consider good and give the training method a downhill direction to follow at almost every point. Some useful losses have sharp corners rather than perfectly smooth curves; training software can still handle those. How we build the loss depends on what kind of thing we’re predicting.
Predicting a number: squared error
When the answer is a quantity (a price, a temperature, a length) we already have our loss from the last section, the mean squared errormean squared errorA common loss for predicting numbers: average the squared gap between each prediction and its true value. Squaring makes all errors positive and gives especially large misses extra weight. L = (1/n) Σ (ŷ − y)².See in glossary →:
Each prediction’s gap from the truth is squared and averaged. Zero means every prediction is spot on. This is a common loss for regressionregressionA prediction task whose answer is a number (a price, a temperature), as opposed to a category. Usually scored with squared error.See in glossary →: it is simple and smooth, and it emphasizes large misses. That last property can be undesirable when a few examples are extremely unusual, so other losses sometimes count each miss in more direct proportion to its size.
Predicting a category: the surprise view
Many problems ask “which one?” rather than “what number?” Is this email spam or not? Which item in a huge list comes next? This is classificationclassificationA prediction task whose answer is one of a fixed set of categories (spam / not spam, or which of 100,000 tokens comes next). The model outputs a probability per option and is scored with cross-entropy.See in glossary →, where the loss usually measures how much probability the model gave the correct category, rather than the distance between numeric labels. A classifier outputs probabilities: for a yes/no problem it might give the probability of “yes,” and for a many-way problem it gives a probability for each option. A probability expresses how likely an outcome is, on a scale from 0 (impossible) to 1 (certain).
How wrong is a probability? The trick is to score it by surprise. If the true answer was “spam” and the model said 90%, it was barely surprised: small loss. If it had confidently said 1%, reality just blindsided it: huge loss. A logarithmlogarithmThe exponent needed to obtain a number from a chosen base. The natural logarithm uses the base e, approximately 2.718; logarithms turn products into sums.See in glossary → reverses exponentiation: if , then , where is a mathematical constant approximately equal to 2.718. A useful measure of surprise is the negative logarithm of the probability the model assigned to the correct answer:
This is cross-entropy losscross-entropy lossA classification loss that penalizes the model according to the negative log-probability it assigned to the correct answer.See in glossary →. It is the standard loss for modern classifiers, and it will show up again when we get to language models. Say the right answer with probability 1 and the loss is : no surprise, perfect. Say it with probability near 0 and shoots toward infinity: total shock, enormous penalty.
Feel it
Drag the probability the model assigns to the correct class and watch the loss respond. The relationship is not a straight line: that curve, gentle near “confident and right” and cliff-like near “confident and wrong,” is what makes cross-entropy penalize confident mistakes so strongly.
One number, whatever the task
Regression or classification, the shape is the same: predictions in, one loss number out, low when the model is good. That single number is what makes automatic learning possible, because now “improve the model” has a precise meaning: reduce this number. But we glossed over something: cross-entropy needs a probability for each class, and a plain weighted sum of inputs can land anywhere from minus to plus infinity. How does a raw score become a probability, and how does that carve the world into categories? A classifier needs a function that converts its raw scores into probabilities before this loss can be calculated.