Drawing a boundary
Turning scores into yes/no decisions
Cross-entropy needs a probability for each class, but a plain weighted sum of inputs (the only thing our model knows how to compute) can land anywhere from minus infinity to plus infinity. A score of −7.2 or +340 isn’t a probability. A function called the sigmoid converts such a score into a probability for a yes/no classificationclassificationA prediction task whose answer is one of a fixed set of categories (spam / not spam, or which of 100,000 tokens comes next). The model outputs a probability per option and is scored with cross-entropy.See in glossary → problem.
From a number to a probability
Start with the raw output of a linear model, the score (often called the logitlogitsThe raw, pre-softmax scores the model produces — one per vocabulary token, per position. Bigger logit = the model finds that token more likely; the actual value can be any real number, positive or negative. Applying softmax across the vocabulary turns logits into a probability distribution that sums to 1. Sampling then picks one token from that distribution.See in glossary →): . It’s an unbounded real number: big and positive when the model leans toward “yes,” big and negative for “no,” near zero when it’s unsure. We want to squash that whole range into the interval so it reads as a probability.
The sigmoid function (also called the logistic function) does exactly that:
Feed it a huge positive score and it returns something just under 1; a huge negative score gives something just above 0; a score of 0 gives exactly 0.5. Its S-shape is gentle in the middle and flattens at the ends: confident scores map to confident probabilities, and once you’re already sure, being more sure barely moves the output. That output is precisely the probability cross-entropy was waiting for.
The score also draws a line
Here’s the geometric payoff. In a space of two features, the set of points where the score is exactly zero (where , so the probability is exactly 0.5) is a straight line. On one side the score is positive and the model predicts “yes”; on the other it’s negative and predicts “no.” That line is the decision boundary, and the weights set its angle while the bias sets its offset.
A linear classifier therefore separates its inputs with a straight boundary. The sigmoid didn’t add any bending power — the boundary is still a line — it only converted the linear score into a probability-shaped number so we can measure and penalize confidence.
Notice the threshold. We called the boundary “probability = 0.5,” but nothing forces 0.5. Slide the threshold and the line slides with it: demand 0.9 confidence before you cry “fraud” and you’ll raise fewer false alarms but miss more real cases; drop it to 0.2 and you catch more fraud at the cost of crying wolf. The weights fix the direction of the cut; the threshold picks where along it you’re willing to commit.
More than two classes: softmax
Sigmoid handles a yes/no question. For “which of classes?” (which item in a long list comes next) we need scores turned into nonnegative numbers that sum to 1. That’s the softmaxsoftmaxFunction that turns any vector into a probability distribution (positive, sums to 1) by exponentiating and normalizing.See in glossary → function: raise to the power of each score to make it positive, then divide by the total so the whole set can be treated as one probability distribution. A two-class softmax and a sigmoid are closely related ways to express the same yes/no choice.
One caution: softmax gives numbers shaped like probabilities, but that does not automatically mean a score of 0.9 is right 90% of the time. When predicted probabilities match real-world frequencies that way, the model is called calibrated. Calibration is useful, but it is a separate property from simply summing to 1.
We can now score, squash, and classify: turn raw inputs into a probability and a boundary, and measure how wrong that probability was. What we still can’t do is improve it automatically. We have a loss; which way do we turn each knob to make it smaller? Answering that needs one idea from calculus (the slope of the loss with respect to a knob) which tells us how a small change to a parameter changes the loss.