Scaling laws
Loss as a power law in size, data, and compute
Paper: Scaling Laws for Neural Language Models — Kaplan et al., 2020
GPT-2 showed that scaling works. Kaplan et al.’s 2020 Scaling Laws for Neural Language Models showed that scaling is predictable, so precisely that you can forecast a model’s loss before you train it. Its fitted relationships let researchers estimate the return from a larger training budget.
Loss is a power law in scale
The experiments found a consistent relationship across scales. Train transformers of many sizes on many amounts of data, and the test losscross-entropy lossA classification loss that penalizes the model according to the negative log-probability it assigned to the correct answer.See in glossary → follows a smooth power lawpower lawA relationship of the form y = a·x^(−b): on log-log axes it's a straight line. Pre-training loss follows a power law in scale, so each 10× of compute buys a roughly constant drop in loss.See in glossary → in each of the three scale factors (model size , dataset size , and compute ) over many orders of magnitude:
with (non-embedding parameters) and tokens. There’s a matching law for compute with exponent . A power lawpower lawA relationship of the form y = a·x^(−b): on log-log axes it's a straight line. Pre-training loss follows a power law in scale, so each 10× of compute buys a roughly constant drop in loss.See in glossary → plots as a straight line on log-log axes: each multiplicative step in scale reduces the scale-dependent loss by a fixed proportion, equivalently, by a fixed additive amount in log-loss.

Figure 1 from Scaling Laws for Neural Language Models (Kaplan et al., 2020), arXiv 2001.08361.
These are scaling lawsscaling lawsEmpirical formulas showing that test loss falls as a smooth power law in model size, dataset size, and compute. They let you predict a large model's performance from small experiments.See in glossary →
A few properties made the result so influential:
- Smoothness. No bumps, no plateaus across the studied range: just clean power laws.
- Universality. The shape barely depends on architectural details (depth vs. width, etc.) within reason; it’s dominated by scale.
- Sample efficiency of large models. Bigger models reach any given loss using fewer tokens. A large model “learns more per example.”
Kaplan’s recipe, and the catch
From these laws, Kaplan derived how to spend a fixed compute budget6ND ruleA rule of thumb: training a dense model with N parameters on D tokens costs about 6ND floating-point operations (≈2ND forward + ≈4ND backward).See in glossary → () optimally. Their answer: pour most of the increase into model size, with data growing only slowly. They estimated , i.e. as you scale compute, grow the model fast and the dataset gently, training very large models and stopping well before convergence.
This recipe shaped a generation of models, including GPT-3: make them enormous, don’t worry too much about training on proportionally more tokens. It was also, in one important respect, wrong. A couple of years later, Chinchilla showed Kaplan had badly under-weighted data, largely because of a subtle flaw in how the learning rate was scheduled across the experiments. The Chinchilla study revisited the balance between model size and data.
GPT-3 put much of its training budget into model size. Its results also revealed an ability that the loss curves alone did not describe: learning a task from examples placed in the prompt.