Section 17

TRPO to PPO

Trust regions and the clipped surrogate

An advantage estimate A^t\hat{A}_t supplies a learning signal with less noise than raw rewards. The remaining danger is the step size. Policy gradients tell you a direction, not a safe distance, and in RL a single over-eager step can be catastrophic in a way it never is in supervised learning. PPO addresses this with a clipped surrogate objective: a substitute training objective that limits the reward for moving response probabilities too far.

Why one big step can destroy a policy

In supervised learning, the data is fixed. If you overshoot, the loss goes up, the next gradient points back, and you recover. In RL the data is generated by the policy itself. Take too large a step and you change the very distribution that produces your training data. The policy might lurch into a region where it generates garbage, every rollout from that broken policy scores near zero, the advantages collapse to noise, and there is no useful gradient left to climb back out. The feedback loop that was helping you is now actively hurting you. RL training can fall off a cliff and never return.

So we want to improve the policy, but only cautiously: never letting the new policy stray too far from the old one in a single update. We need to measure that “distance,” and the natural ruler is the KL divergenceKL divergenceKullback–Leibler divergence — a measure of how far one probability distribution is from another. Used in post-training as a "leash" that keeps a model close to a reference policy.See in glossary → between the old and new policy distributions.

TRPO: improve, but stay in a trust region

This is the idea behind Trust Region Policy Optimization (TRPOTRPOTrust Region Policy Optimization (Schulman, 2015) — take the largest policy-gradient step that stays within a trust region (a KL bound), guaranteeing stable improvement. PPO’s parent.See in glossary →, Schulman 2015). Define a trust regiontrust regionA bound on how far the policy may move in one update (measured in KL divergence), so the update stays in the region where the local approximation is trustworthy.See in glossary →: a neighborhood around the current policy within which we trust our local estimate of “improvement” to be reliable. TRPO maximizes the expected advantage subject to a hard KL constraint:

max⁡θ  E ⁣[ πθ(a∣s)πθold(a∣s) A^ ]subject toE[ KL(πθold ∥ πθ) ]≤δ\max_\theta \; \mathbb{E}\!\left[\, \frac{\pi_\theta(a \mid s)}{\pi_{\theta_{\text{old}}}(a \mid s)} \, \hat{A} \,\right] \quad \text{subject to} \quad \mathbb{E}\big[\, \mathrm{KL}(\pi_{\theta_{\text{old}}} \,\|\, \pi_\theta) \,\big] \le \delta

Two pieces to unpack. The constraint says: take the biggest improving step you can, but don’t let the new policy diverge from the old one by more than δ\delta in KL. That is the guard rail against the cliff. The objective contains a new and important ratio.

The importance-sampling ratio

That fraction in the objective,

rt(θ)=πθ(at∣st)πθold(at∣st),r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)},

is the importance-samplingimportance samplingReweighting samples from one distribution to estimate expectations under another, via the probability ratio π_new/π_old. The ratio PPO clips comes from here.See in glossary → ratio. It exists to solve the off-policyoff-policyRL that learns from data generated by a different (older or separate) policy. DPO and rejection-sampling methods are off-policy / offline.See in glossary → problem from chapter 13. We generated our rollouts with the old policy πθold\pi_{\theta_{\text{old}}}, but we want to optimize the new policy πθ\pi_\theta. Importance sampling corrects for the mismatch: it reweights each old sample by how much more (or less) likely the new policy is to have produced it. If the new policy now favors a good action, rt>1r_t > 1 and its advantage counts for more; if it has turned away from it, rt<1r_t < 1.

This ratio is what lets us take several gradient steps on one batch of rollouts instead of throwing the data away after a single update, a major efficiency win. But it’s also dangerous: if rtr_t runs away from 1, we’re extrapolating wildly from samples the new policy would rarely generate, and the estimate becomes meaningless. That’s precisely what the trust region must contain.

TRPO approximately solves this constrained update with second-order Fisher-matrix machinery and a line search. It is more complicated and expensive to implement at scale. The natural question is whether a first-order method can obtain useful stability without solving that constraint directly.

PPO: clipping the probability ratio

The answer is Proximal Policy Optimization (PPOPPOProximal Policy Optimization (Schulman, 2017) — nudges the policy toward higher reward in small, clipped steps with a KL leash to a reference model. The RLHF workhorse: stable, simple, widely used.See in glossary →, Schulman 2017), which uses a simpler objective than TRPO’s constrained optimization. Instead of a hard KL constraint, PPO bakes the “stay close” pressure directly into the objective by clipping the ratio. The clipped surrogateclipped surrogate objectivePPO’s loss: maximize the probability-ratio-weighted advantage, but clip the ratio to [1−ε, 1+ε] so a single update can’t move the policy too far.See in glossary → objective is:

LCLIP(θ)=Et[ min⁡( rt(θ) A^t,    clip(rt(θ), 1−ϵ, 1+ϵ) A^t )]L^{\text{CLIP}}(\theta) = \mathbb{E}_t\Big[\, \min\big(\, r_t(\theta)\, \hat{A}_t, \;\; \mathrm{clip}\big(r_t(\theta),\, 1 - \epsilon,\, 1 + \epsilon\big)\, \hat{A}_t \,\big) \Big]

Here ϵ\epsilon is a small constant (typically 0.10.1 or 0.20.2). The clip\mathrm{clip} function pins rtr_t to the interval [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon]. We compute the surrogate two ways (once with the true ratio, once with the clipped ratio) and take the minimum. That minimum is what makes the whole thing work, and it’s worth walking through both cases carefully.

The effect is a cheap, first-order surrogate inspired by trust-region methods. There is no KL constraint to solve and no second-order matrixmatrixA rectangular array of numbers arranged in rows and columns. Matrix multiplication combines these arrays through weighted sums.See in glossary →, a table of numbers describing how the local slopes change: just a clamp and a min. But clipping is not a hard trust region or a monotonic-improvement guarantee: shared parameters, multiple minibatch epochs, and unsampled actions can still produce a large KL change. Implementations therefore monitor KL and may early-stop or add a KL penalty. PPO’s practical balance of simplicity and empirical stability, not an exact TRPO guarantee, is why it became widely used.

Try it

The plot below shows the clipped surrogate as a function of the probability ratio rtr_t. Flip the sign of the advantage and slide ϵ\epsilon. Watch the objective rise linearly with rtr_t and then go flat at the clip boundary. That flat region is the trust region in disguise. Notice how the flat side switches depending on whether the advantage is positive or negative, and how a wider ϵ\epsilon permits bigger steps before the brakes engage.

PPO clipped surrogate objective
Plotted over the probability ratio r = π_new/π_old: the raw term r·A vs PPO's clipped objective. The flat region is the cap.
00.511.522.5r1−ε1+ε
r·A (unclipped) PPO objective current r
r·A (unclipped)
1.00
PPO objective
1.00
Region
active
PPO maximizes min( r·A, clip(r, 1−ε, 1+ε)·A ). When A > 0 (the action was good), the objective stops rising once r > 1+ε — so a single good sample can't shove the policy arbitrarily far. When A < 0 (the action was bad), the cap is symmetric on the r < 1−ε side. In the flat clipped zone the gradient is zero, which is how PPO approximates a trust region without the expensive second-order math of TRPO.

The clipped objective limits the incentive for large probability changes while preserving a gradient that can correct movement in the wrong direction. An RLHF pipeline combines this update with a learned reward, a value estimate, and a KL penalty to a frozen reference model.