Section 32

RL scaling laws

How far does RL compute take reasoning?

Papers: ProRL: Prolonged RL Expands Reasoning Boundaries — Liu et al. (NVIDIA), 2025 · The Art of Scaling RL Compute for LLMs — Khatri et al., 2025 · RL vs. Distillation — Kim et al., 2025

Pre-training has scaling laws: spend more compute, in the right proportions, and loss falls along a predictable curve. That predictability is what lets labs commit hundreds of millions of dollars to a single run. For reinforcement learning, the relationship between compute and performance has been less predictable. Two strands of 2025 work examine whether longer training continues to improve results, and whether those improvements represent new abilities or greater reliability.

Does RL expand the boundary, or only sharpen it?

Several studies suggest that RL from verifiable rewardsRLVRReinforcement Learning from Verifiable Rewards — use an automatic checker (unit tests, an answer key, a math grader) as the reward instead of a learned reward model. No reward hacking of a neural proxy.See in glossary → mainly makes existing abilities more reliable. Measure a model with pass@k: the chance that at least one of kk sampled answers is correct. RL reliably lifts pass@1 (the model’s single best guess gets better), but several studies found it often doesn’t lift pass@k at large kk. The interpretation: RL isn’t teaching new solutions, it’s just concentrating probability onto solutions the base model could already stumble onto occasionally. Under this interpretation, the reasoning was latent in the base model all along, and RL only made it reliable.

The counter-evidence: prolonged RL

Training duration may help explain these results. ProRL (NVIDIA, 2025) argues that with enough training — thousands of RL steps, plus stability machinery to keep the policy from collapsing (entropy control, periodic reference resets, other stability measures) — RL does expand the boundary. Their prolonged-RL models beat the base model across a wide range of pass@k, including on tasks where the base model scored essentially zero no matter how many samples you drew. If the base model can never produce a correct answer, RL can’t merely be reweighting existing samples: it has found something genuinely new.

These findings allow for both effects: short RL runs mostly sharpen, and the “RL only reweights” result is real for that regime; but sustained, stable RL on the right problems can push past the base model’s reach. How long and how stably you train turns out to matter as much as the algorithm.

RL, made predictable

Other work in 2025 studied how RL performance scales with compute. The Art of Scaling RL Compute (Khatri et al., 2025) ran a systematic ~400,000-GPU-hour study and found that RL performance follows a sigmoidal, or S-shaped, curve in compute: improvement starts slowly, accelerates, then levels off toward a limiting value. The curve can be extrapolated—fitted to smaller runs to predict the results of larger ones—and the design choices (the loss form, advantage normalization, batch construction) mostly move the efficiency and the ceiling rather than the shape. They distill the best choices into a recipe, ScaleRL, whose scaling matches the predictability long taken for granted in pre-training.

RL scaling lawsRL scaling lawsEmpirical curves predicting how RL post-training performance grows with compute — the RL analogue of pre-training scaling laws. Recent work fits a sigmoidal curve that can be extrapolated from small runs.See in glossary → are young and far less settled than their pre-training cousins: the curves depend heavily on the task, the verifier, and stability tricks, and the field is actively arguing about how far the boundary really moves. These studies offer a basis for planning RL compute, while leaving open how well their curves transfer to other tasks and training recipes.