Which way is downhill?
Slopes and just enough calculus
We have a model with knobs and a loss that scores them. To learn automatically, the machine needs to answer one question, over and over: if I nudge this knob a tiny bit, does the loss go up or down? That question has a name in mathematics — the derivative — a measure of how fast one quantity changes as another changes. It is an idea from calculuscalculusThe math of how things change. The one piece we need is the slope: at any point on a curve, how steeply it is rising or falling, and in which direction. That single idea is what lets us tell which way to turn each dial to lower the loss — no heavy math required.See in glossary →, the branch of mathematics that studies change and accumulation.
The slope is the answer
Fix everything except one knob, w, and plot the loss as w varies. You get a curve. At any point on that curve, the derivativederivativeThe slope of a function at a point: how fast the output changes as you nudge the input, and in which direction. In training, the derivative of the loss with respect to a parameter tells you which way to turn that knob.See in glossary → is simply its slope there: how steeply the loss rises or falls as you increase w, and in which direction.
- Slope positive: increasing
wincreases the loss. So to lower the loss, decreasew. Move left. - Slope negative: increasing
wdecreases the loss. Move right. - Slope zero: the curve is flat right here. You might be at a bottom, a top, or a flat spot. With many knobs, you could also be at a saddle pointsaddle pointA flat-looking point that is downhill in some directions but uphill in others — like the center of a horse saddle. A zero gradient can mark a saddle rather than a minimum.See in glossary →: downhill in some directions but uphill in others. The slope alone cannot tell you which.
That’s it. The sign of the slope tells you which way is downhill; the size of the slope tells you how steep the ground is. You never need the model to “understand” anything. It just needs to feel the slope under each knob.
Drag the point
Move the dot along the loss curve. The straight line touching it is the tangent, and its steepness is the slope at that spot. Watch the sign flip as you cross the bottom of the valley, and watch the arrow point the way the loss falls.
The rule the widget keeps stating is the one to memorize:
Read it plainly: new w = old w, minus a small step in the direction of the slope. The minus is the important part. It’s what makes you go against the slope, i.e. downhill. The little (the Greek letter “eta”) controls how big a step to take; it’s called the learning ratelearning rateThe size of each parameter step. Too high and training can diverge; too low and it crawls.See in glossary →, and it determines whether the updates make steady progress or overshoot.
From one knob to many
A real model has millions of knobs, not one. The idea generalizes without any new concepts: for each knob, ask “what’s the slope of the loss along this knob, holding the others fixed?” That per-knob slope is called a partial derivativepartial derivativeThe rate at which a function changes when one input changes while all its other inputs are held fixed.See in glossary →, and the full list of them (one number per parameter) is the gradient. It’s just “the slope” pointing in every knob’s direction at once.
The gradient gives a downhill direction at the current settings. Repeatedly stepping in that direction is gradient descentgradient descentThe core training algorithm: repeatedly nudge each parameter a small step in the direction that lowers the loss, as told by the gradient.See in glossary →. Its success also depends on the size of each step: too small is slow, while too large can send the loss upward.