PPO Algorithm: Proximal Policy Optimization for Stable RL

Michael BrenndoerferDecember 27, 202571 min read

Part of Language AI Handbook

Covers PPO's clipped objective for stable policy updates. Topics include trust regions, GAE advantage estimation.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

PPO Algorithm

In the previous chapter, we explored policy gradient methods and saw how REINFORCE directly optimizes a policy by following the gradient of expected reward. While mathematically elegant, vanilla policy gradients are notoriously unstable during training. A single large gradient update can catastrophically degrade policy performance, and recovery is difficult. The algorithm might spend hundreds of update steps undoing damage from one reckless step. This instability prevented policy gradient methods from scaling to complex domains like language model fine-tuning, where each update step is expensive and evaluation is slow.

Proximal Policy Optimization (PPO), introduced by Schulman et al. in 2017, addresses this instability through a deceptively simple mechanism: it clips the objective function to prevent updates that change the policy too drastically. Rather than solving a complex constrained optimization problem, PPO modifies the reward signal itself so the optimizer cannot benefit from stepping too far from the current policy. The result is an algorithm that is simultaneously simpler to implement, more computationally efficient, and more stable than its predecessors.

Think of PPO as a careful hiker who always stays within a reliable map region. The hiker can see the general direction of higher elevation and wants to move toward it. But the hiker also knows that map accuracy degrades quickly outside the territory already explored. So the hiker commits to moving only a short distance with each step, updating the map from the new vantage point, and then deciding where to move next. The hiking speed is not zero, which would be no progress at all, but it is bounded, which prevents walking off a cliff based on a faulty extrapolation. PPO implements exactly this constraint: bound each update step so you remain in territory where your gradient estimates are trustworthy.

Understanding PPO is needed for anyone working with reinforcement learning from human feedback (RLHF) in language models. Modern alignment pipelines, including those behind ChatGPT, Claude, and Gemini, rely heavily on PPO or close variants to translate human preference data into updated model behavior. The algorithm's stability and implementation simplicity made it the natural choice when researchers needed to apply RL to language models with billions of parameters, where unstable training would be extraordinarily expensive to debug or recover from.

This chapter covers PPO in depth from the ground up. We start with the specific failure modes of unconstrained policy gradient optimization, move through the trust region methods that motivated PPO's design, then derive the clipped objective, work through Generalized Advantage Estimation, assemble the complete multi-component objective, and finally implement a working PPO agent in PyTorch. We include a numerical worked example so you can trace every computation by hand before running code. By the end, you will understand what PPO does, why each design choice exists, and what goes wrong when you deviate from it.

Historical Context

PPO was introduced by John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov at OpenAI in their 2017 paper "Proximal Policy Optimization Algorithms." The paper presented PPO as a simplification of Trust Region Policy Optimization (TRPO), their own earlier algorithm from 2015. TRPO had proven that stable policy improvement was possible, but its implementation required second-order optimization, Fisher information matrices, and conjugate gradient solvers. This makes it impractical at scale. PPO replaced all of this machinery with a simple clipped objective, achieving similar stability with a first-order optimizer like Adam. Within a year of publication, PPO became the dominant algorithm for deep reinforcement learning at OpenAI and throughout the research community. When OpenAI published InstructGPT in 2022. This shows that language models could be aligned with human preferences using RL, PPO was the algorithm that made it work. The paper's influence on the alignment field cannot be overstated: virtually every publicly discussed RLHF pipeline for language model alignment traces its RL component back to this 2017 contribution.

The Problem with Vanilla Policy Gradients

The instability in vanilla policy gradient methods is not a superficial implementation problem that can be fixed with a better learning rate schedule. It is a structural problem rooted in how the gradient signal relates to policy performance. To understand why PPO's solution is the right one, we need to be precise about exactly what goes wrong.

Recall that the policy gradient takes the form:

∇θJ(θ)=Eτ∼πθ[∑t=0T∇θlog⁡πθ(at∣st)⋅At]\nabla_\theta J(\theta) = \mathbb{E}_{\tau \sim \pi_\theta}\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t | s_t) \cdot A_t\right]

To understand why this formula creates practical difficulties, examine what each component contributes:

  • ∇θJ(θ)\nabla_\theta J(\theta): the gradient of the expected cumulative reward with respect to policy parameters, showing the direction to adjust parameters to increase expected reward
  • Eτ∼πθ\mathbb{E}_{\tau \sim \pi_\theta}: expectation over trajectories sampled from the current policy πθ\pi_\theta
  • θ\theta: parameters of the policy network that we are optimizing
  • J(θ)J(\theta): the expected cumulative reward under policy πθ\pi_\theta
  • τ\tau: a trajectory (sequence of states and actions) sampled from the current policy
  • πθ(at∣st)\pi_\theta(a_t | s_t): the probability of taking action ata_t in state sts_t under the current policy
  • TT: the time horizon, the length of the episode
  • AtA_t: the advantage function at time tt, estimating how much better action ata_t is compared to the average action in state sts_t
  • ∇θlog⁡πθ(at∣st)\nabla_\theta \log \pi_\theta(a_t | s_t): the gradient of the log probability with respect to policy parameters, showing the direction to adjust parameters to make action ata_t more likely
  • ∑t=0T\sum_{t=0}^{T}: sum over all timesteps in the trajectory from 0 to T

The basic issue with this formulation is that the gradient provides no guidance about step size. The policy gradient theorem tells us which direction to move parameters to increase expected reward, but it remains entirely silent about how far we should move in that direction. This is analogous to knowing that walking north takes you closer to your destination but not knowing whether to take one step or one hundred steps. A step in the gradient direction improves performance locally, within an infinitesimally small neighborhood around the current parameters. However, nothing in the mathematics prevents taking such a large step that you overshoot into a region where the policy performs terribly. Policy performance is not smooth across the parameter space: a policy that seems promising can suddenly fail after a large update.

This problem is especially acute because of three interconnected challenges:

  • Non-stationary data distribution: The policy generates its own training data. When the policy changes significantly, the state distribution also changes, potentially invalidating previously learned value estimates.
  • High variance: Policy gradient estimates are inherently noisy, which makes it difficult to distinguish signal from noise.
  • Irreversibility: A bad update might move the policy to a region where it never encounters states that would help it recover.

Consider a language model learning from human feedback. If a single gradient update makes the model much more likely to generate certain patterns, the model might suddenly produce outputs that are completely off-distribution from its training, leading to reward model extrapolation errors and further degradation. The optimizer has no way to know it has moved too far until the damage is already done, and by that point, the policy may be in such a degraded state that subsequent updates cannot recover it without discarding all progress.

Out[3]:
Visualization
Line chart of a smooth quadratic performance curve with a gradient descent path of dots progressing from left to the global optimum marked by a gold star.
Stable optimization landscape shows a smooth quadratic performance surface where gradient descent makes steady progress. The convex objective (blue curve) creates a bowl-shaped landscape where small steps (green path, marked by dots) reliably improve performance toward the global optimum (gold star). This smooth, predictable behavior is ideal but rarely encountered in practical reinforcement learning problems.
Line chart of a treacherous performance landscape with a plateau ending in a steep cliff, showing a green starting point, an orange gradient step, and a red overshoot point beyond the cliff.
Treacherous optimization landscape reveals the danger of unconstrained policy updates. The smooth performance plateau transitions abruptly to a steep cliff. An initially reasonable gradient step (green dot) identifies the correct improvement direction, but a large unconstrained update (orange dot) overshoots beyond the cliff edge, catastrophically degrading performance (red dot). This illustrates why policy gradient methods need constraints to prevent large, destabilizing updates in complex landscapes.

Trust Region Methods

The insight behind trust region methods is to constrain how much the policy can change in each update. Rather than blindly following the gradient wherever it leads, we optimize the policy subject to a constraint that keeps the new policy "close" to the old one. This approach recognizes a basic tension in optimization: we want to improve the policy as quickly as possible. However, our confidence in the improvement direction decreases the further we move from where we collected our data. Gradients computed from sampled trajectories accurately describe performance improvement near the current policy, but extrapolating far from observed data becomes unreliable.

Think of a trust region as the area on a topographic map that your GPS unit has surveyed at high resolution. Within that region, the contour lines are accurate and you can safely navigate. Outside that region, the map was generated from coarser satellite data and might have cliffs, valleys, or impassable terrain that looks navigable on the map. You would not plan a route that takes you far outside your high-resolution zone. Similarly, the policy gradient is an accurate local guide only within the neighborhood of states and actions you have recently observed. Venturing far outside that neighborhood means relying on gradient extrapolations that may be wildly wrong.

Trust Region

A trust region is a neighborhood around the current parameters within which a local approximation of the objective function is trusted to be accurate. Optimization proceeds by maximizing this approximation within the trust region, then updating the region based on how well the approximation matched reality. If the approximation was good, the trust region expands. If it was poor, the region shrinks. This adaptive mechanism maintains alignment between the model and reality throughout training.

The key insight behind all trust region methods is that gradient estimates are only reliable near the data they were computed from. Moving the policy far from the current policy means the state-action distribution shifts, and the advantages computed from old data no longer accurately reflect the new policy's performance. The trust region constraint formalizes the boundary beyond which you should not venture without collecting new data. Enforcing this boundary ensures that the objective function improvement you see during optimization corresponds to actual performance improvement in the environment.

Out[4]:
Visualization
Contour plot of an objective function in two-dimensional parameter space with a circular green trust region centered on the current policy. A red dashed arrow points outside the circle showing the unconstrained gradient direction, and a solid green arrow points to the constrained new policy inside the circle.
Trust region constrains policy updates to a local neighborhood where gradient estimates remain reliable. The blue dot marks the current policy at the center of the shaded green circle (trust region boundary). The red dashed arrow shows where an unconstrained gradient would lead (potentially too far), while the green arrow shows the constrained update that respects the trust region and reaches a new policy (green square). This visualization demonstrates the core principle: gradient estimates are only reliable near the current policy, so updates must stay within the trust region to maintain stability.

Trust Region Policy Optimization (TRPO), PPO's predecessor, formalizes this using KL divergence as a constraint. The KL divergence measures how different two probability distributions are, which makes it a natural choice for measuring policy similarity. Policies that assign similar probabilities to actions have low KL divergence, while policies that behave very differently have high KL divergence. This makes KL divergence a precise, mathematically grounded way to define "how much the policy has changed."

The TRPO objective addresses the policy update problem by maximizing expected advantage while explicitly constraining how much the policy can change. It maximizes the expected advantage weighted by the importance sampling ratio, which allows us to use data from the old policy to evaluate the new policy, while constraining the KL divergence between old and new policies:

max⁡θEs,a∼πθold[πθ(a∣s)πθold(a∣s)Aπθold(s,a)]subject toEs[DKL(πθold(⋅∣s)∥πθ(⋅∣s))]≤δ\begin{aligned} \max_\theta \quad & \mathbb{E}_{s, a \sim \pi_{\theta_{\text{old}}}}\left[\frac{\pi_\theta(a|s)}{\pi_{\theta_{\text{old}}}(a|s)} A^{\pi_{\theta_{\text{old}}}}(s, a)\right] \\ \text{subject to} \quad & \mathbb{E}_s\left[D_{\text{KL}}\left(\pi_{\theta_{\text{old}}}(\cdot|s) \| \pi_\theta(\cdot|s)\right)\right] \leq \delta \end{aligned}

To understand how this constrained optimization problem balances improvement against stability, examine each component:

  • θ\theta: parameters of the new policy being optimized
  • θold\theta_{\text{old}}: parameters of the old policy from which data was collected
  • Es,a∼πθold\mathbb{E}_{s, a \sim \pi_{\theta_{\text{old}}}}: expectation over states and actions sampled from trajectories collected using the old policy
  • ss: a state sampled from the state distribution under the old policy
  • aa: an action sampled from the old policy in state ss
  • πθ(a∣s)\pi_\theta(a|s): probability of action aa in state ss under the new policy
  • πθold(a∣s)\pi_{\theta_{\text{old}}}(a|s): probability of action aa in state ss under the old policy
  • πθ(a∣s)πθold(a∣s)\frac{\pi_\theta(a|s)}{\pi_{\theta_{\text{old}}}(a|s)}: the importance sampling ratio, which reweights data from the old policy to evaluate the new policy
  • Aπθold(s,a)A^{\pi_{\theta_{\text{old}}}}(s, a): advantage function computed under the old policy, measuring how much better action aa is than average in state ss
  • Es\mathbb{E}_s: expectation over states from the state distribution under the old policy
  • DKL(πθold∥πθ)D_{\text{KL}}(\pi_{\theta_{\text{old}}} \| \pi_\theta): Kullback-Leibler divergence measuring how much the new policy distribution differs from the old policy distribution
  • δ\delta: maximum allowed KL divergence, the trust region radius

This constraint ensures the new policy πθ\pi_\theta does not diverge too far from the old policy πθold\pi_{\theta_{\text{old}}} in terms of the probability distributions over actions. The ratio πθ(a∣s)πθold(a∣s)\frac{\pi_\theta(a|s)}{\pi_{\theta_{\text{old}}}(a|s)} is called the importance sampling ratio and allows us to evaluate the new policy using data collected from the old policy. This ratio acts as a correction factor: if the new policy is twice as likely to take an action as the old policy, we weight that action's contribution twice as heavily to account for the fact that it would occur more frequently under the new policy.

Why TRPO Works But Is Complex

TRPO guarantees monotonic improvement under certain conditions, meaning the policy never gets worse than the previous version. This guarantee provides the stability that vanilla policy gradients lack. However, the mathematical machinery required to enforce this guarantee comes with significant computational costs: computing Fisher information matrices and solving linear systems.

Solving the constrained optimization problem requires computing second-order derivatives, specifically the Fisher information matrix, and performing conjugate gradient optimization to solve a system of linear equations. This makes TRPO computationally expensive and difficult to implement correctly. The Fisher information matrix has size n×nn \times n where nn is the number of policy parameters, so even storing it is impractical for large neural networks. Computing the matrix-vector product required by conjugate gradient solvers requires special tricks involving second-order automatic differentiation. Computing the Fisher information matrix across multiple workers in distributed settings adds significant overhead.

PPO achieves similar stability guarantees with a first-order method by replacing the hard constraint with a penalty built directly into the objective function. Instead of solving a constrained optimization problem, PPO modifies the objective itself to discourage excessive policy changes. This transformation from constraint to penalty makes PPO dramatically simpler to implement while preserving the needed benefits of trust region methods. The resulting algorithm fits cleanly into any standard deep learning training loop, with no exotic optimization machinery required.

The Probability Ratio

The probability ratio is central to PPO. It quantifies how much the new policy's probability for an action differs from the old policy's probability. By examining this ratio across all observed state-action pairs, you can assess whether the policy is changing appropriately or excessively. Think of the probability ratio as a sensitivity meter: a ratio of 1 means the policy is unchanged for this action, values greater than 1 mean the policy is becoming more likely to take this action, and values less than 1 mean the policy is becoming less likely. The ratio captures the entire story of how the policy is changing, action by action, state by state.

We define the ratio as:

rt(θ)=πθ(at∣st)πθold(at∣st)r_t(\theta) = \frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)}

Understanding what each symbol represents helps illuminate why this ratio captures policy change so effectively:

  • rt(θ)r_t(\theta): the importance sampling ratio at timestep tt
  • πθ(at∣st)\pi_\theta(a_t|s_t): probability of action ata_t in state sts_t under the new policy with parameters θ\theta
  • πθold(at∣st)\pi_{\theta_{\text{old}}}(a_t|s_t): probability of the same action in the same state under the old policy
  • ata_t: the action taken at timestep tt
  • sts_t: the state at timestep tt
  • θ\theta: parameters of the new policy
  • θold\theta_{\text{old}}: parameters of the old policy

This ratio captures how much more or less likely action ata_t becomes under the new policy compared to the old one. The interpretation is intuitive and direct:

  • rt(θ)=1r_t(\theta) = 1: the new policy assigns the same probability to this action; no change from the old policy
  • rt(θ)>1r_t(\theta) > 1: the new policy makes this action more likely (for example, rt=2r_t = 2 means the new policy is twice as likely to take this action)
  • rt(θ)<1r_t(\theta) < 1: the new policy makes this action less likely (for example, rt=0.5r_t = 0.5 means the new policy is half as likely to take this action)

This ratio directly measures behavioral change. A ratio of 1 everywhere indicates identical policies, while very large or very small ratios indicate dramatic behavioral shifts. This makes the ratio ideal for clipping to constrain policy changes. Importantly, the ratio is also the right quantity for importance sampling: when we use old data to evaluate a new policy, we multiply by the ratio to correct for the fact that we collected data under a different distribution. This dual role as both a measure of policy change and an importance weight makes the ratio the natural pivot point for the PPO objective.

Out[5]:
Visualization
Histogram of probability ratios centered near 1.0 with a green vertical line at 1 and red dashed vertical lines at 0.8 and 1.2 marking the PPO clip bounds.
Distribution of probability ratios from 500 state-action pairs reveals that most policy changes are modest. The concentration between 0.7 and 1.3 shows that typical updates keep the probability ratio relatively close to 1 (no change). Red dashed lines at 0.8 and 1.2 (epsilon=0.2) show the clipping bounds where PPO stops rewarding further policy changes, confirming that the clipping mechanism appropriately constrains most natural policy updates.
Scatter plot of advantage values on the x-axis versus probability ratios on the y-axis with green, orange, and red dots showing desirable and problematic update quadrants and red dashed horizontal clip bound lines.
|-

The standard policy gradient objective can be rewritten using this ratio, revealing its basic role in policy optimization:

LPG(θ)=Et[rt(θ)⋅At]L^{\text{PG}}(\theta) = \mathbb{E}_t\left[r_t(\theta) \cdot A_t\right]

To understand how the ratio enables policy improvement, examine what each term contributes:

  • LPG(θ)L^{\text{PG}}(\theta): the policy gradient objective function we want to maximize
  • Et\mathbb{E}_t: expectation over timesteps in collected trajectories
  • rt(θ)r_t(\theta): the probability ratio at timestep tt, equal to πθ(at∣st)πθold(at∣st)\frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)}
  • AtA_t: the advantage estimate at timestep tt, measuring how much better the action taken was compared to the average action
  • θ\theta: parameters of the policy being optimized

When θ=θold\theta = \theta_{\text{old}}, we have rt(θ)=1r_t(\theta) = 1 everywhere, and the objective reduces to the simple advantage-weighted sum. The gradient of this objective at θ=θold\theta = \theta_{\text{old}} equals the standard policy gradient, confirming that this formulation is equivalent to what we derived earlier.

The problem becomes clear when we consider what happens as we optimize this objective. As the optimizer works to maximize the expected reward, the ratio can become arbitrarily large or small. If an action had positive advantage, showing it was better than expected, unconstrained optimization would keep increasing its probability without bound. The gradient always points toward making good actions more likely and bad actions less likely, but nothing in this formulation limits how far the policy can shift. This is precisely the instability problem that TRPO addressed with its KL constraint, and that PPO addresses with clipping.

The Clipped Objective

PPO's key innovation is straightforward: clip the probability ratio to remove incentives for excessive policy changes. Rather than imposing a hard constraint requiring complex optimization machinery, PPO modifies the objective function itself, so large policy changes provide no additional benefit. The clipping mechanism is self-limiting: once the policy has changed enough to fall outside the allowed range, the gradient signal turns off, and the optimizer cannot push the policy further in that direction.

Think of the clipped objective as a reward system with diminishing returns for policy changes. The policy earns credit for adjusting action probabilities in the direction indicated by the advantage, but only up to a point. Beyond the clip threshold, additional changes earn no extra credit. This is like an employer who rewards employees for working overtime, but only up to a cap: you can work as many hours as you want, but you only get paid extra for the first few hours of overtime. The cap prevents extreme behavior without eliminating the incentive entirely.

The clipped surrogate objective is:

LCLIP(θ)=Et[min⁡(rt(θ)At,clip(rt(θ),1−ϵ,1+ϵ)At)]L^{\text{CLIP}}(\theta) = \mathbb{E}_t\left[\min\left(r_t(\theta) A_t, \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) A_t\right)\right]

where ϵ\epsilon is a hyperparameter, typically set between 0.1 and 0.2, that defines the trust region width. This single parameter controls how much the policy can change in each update. Larger ϵ\epsilon allows more aggressive updates. Smaller ϵ\epsilon enforces tighter constraints and slower learning.

To understand precisely how clipping constrains the ratio, we need to examine the clip function itself. The clip function constrains the probability ratio to remain within the interval [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon], defined as:

clip(r,1−ϵ,1+ϵ)={1−ϵif r<1−ϵrif 1−ϵ≤r≤1+ϵ1+ϵif r>1+ϵ\text{clip}(r, 1-\epsilon, 1+\epsilon) = \begin{cases} 1-\epsilon & \text{if } r < 1-\epsilon \\ r & \text{if } 1-\epsilon \leq r \leq 1+\epsilon \\ 1+\epsilon & \text{if } r > 1+\epsilon \end{cases}

Each component of this piecewise function serves a specific purpose:

  • rr: the input value to be clipped; in PPO, this is rt(θ)r_t(\theta), the probability ratio
  • 1−ϵ1-\epsilon: the lower bound of the allowed range
  • 1+ϵ1+\epsilon: the upper bound of the allowed range
  • The function returns rr unchanged if it's within the bounds, otherwise returns the nearest bound

The key intuition is that this clipping removes the gradient signal when the ratio moves outside the trust region. Consider what happens during optimization: you want the optimizer to adjust parameters to increase the objective. If the new policy is already making an action much more likely, with a ratio greater than 1+ϵ1+\epsilon, or much less likely, with a ratio less than 1−ϵ1-\epsilon, than the old policy, clipping prevents the optimizer from pushing it even further in that direction. The clipped term becomes constant with respect to the parameters, meaning its gradient is zero, so there is no signal encouraging further movement in that direction.

This clipped ratio is then used in computing the clipped surrogate objective. Returning to the full LCLIP(θ)L^{\text{CLIP}}(\theta) formula and examining it in detail:

LCLIP(θ)=Et[min⁡(rt(θ)At,clip(rt(θ),1−ϵ,1+ϵ)At)]L^{\text{CLIP}}(\theta) = \mathbb{E}_t\left[\min\left(r_t(\theta) A_t, \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) A_t\right)\right]

Each component plays a specific role in creating stable policy updates:

  • LCLIP(θ)L^{\text{CLIP}}(\theta): the clipped surrogate objective that PPO maximizes
  • Et\mathbb{E}_t: expectation over all timesteps in the collected batch
  • min⁡(⋅,⋅)\min(\cdot, \cdot): takes the smaller of the two arguments (the pessimistic bound)
  • rt(θ)r_t(\theta): the probability ratio πθ(at∣st)πθold(at∣st)\frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)}
  • AtA_t: the advantage estimate at timestep tt
  • ϵ\epsilon: the clipping parameter, typically 0.1 to 0.2, that defines the trust region
  • clip(rt(θ),1−ϵ,1+ϵ)\text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon): constrains the ratio to the interval [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon]

Understanding the Clipping Mechanism

The min⁡\min operation is important. It selects the smaller of the unclipped and clipped objectives. This keeps a pessimistic (conservative) bound on the improvement. By always choosing the lower value, we prevent the optimizer from being overly optimistic about improvements that would require large policy changes. We will analyze both cases based on the sign of the advantage.

Case 1: Positive Advantage (At>0A_t > 0)

When an action is better than expected (positive advantage), we want to increase its probability. The unclipped objective's gradient pushes the policy to make this action more likely, but PPO limits the extent.

  • If rt(θ)>1+ϵr_t(\theta) > 1 + \epsilon: The clipped term (1+ϵ)At(1+\epsilon) A_t is smaller than rt(θ)Atr_t(\theta) A_t, since At>0A_t > 0 and 1+ϵ<rt(θ)1+\epsilon < r_t(\theta). The min⁡\min selects the clipped term (1+ϵ)At(1+\epsilon) A_t, which is constant with respect to θ\theta, so the gradient is zero and provides no incentive for us to increase the ratio further.
  • If rt(θ)≤1+ϵr_t(\theta) \leq 1 + \epsilon: Both terms are equal or the unclipped term is smaller, so normal optimization proceeds.

This means that once the probability ratio exceeds 1+ϵ1+\epsilon, the optimizer receives no additional reward for increasing it further. The policy has already been sufficiently encouraged to take this good action. Importantly, the policy can still take that action at a higher probability than the old policy, it just does not receive additional credit for pushing the probability even higher.

Case 2: Negative Advantage (At<0A_t < 0)

When an action is worse than expected (negative advantage), we want to decrease its probability. The unclipped objective's gradient encourages this, but PPO limits the extent:

  • If rt(θ)<1−ϵr_t(\theta) < 1 - \epsilon: Since At<0A_t < 0, multiplying by smaller values makes the product more negative. The clipped term (1−ϵ)At(1-\epsilon) A_t is larger (less negative) than the unclipped term rt(θ)Atr_t(\theta) A_t because rt(θ)<1−ϵr_t(\theta) < 1-\epsilon. The min⁡\min selects the more negative unclipped term. Since the clipped term (1−ϵ)At(1-\epsilon) A_t is constant, its gradient is zero and the objective becomes flat once rt(θ)r_t(\theta) drops below 1−ϵ1-\epsilon, preventing the policy from decreasing the probability further.
  • If rt(θ)≥1−ϵr_t(\theta) \geq 1 - \epsilon: Normal optimization proceeds.

This prevents the policy from becoming too averse to actions that happened to have negative advantage. Such actions might still be valuable in other states, and excessively penalizing them could harm overall performance. The clipping protects both ends: it prevents overconfidence about good actions and overpunishment of bad ones.

The following figure illustrates this behavior:

Out[6]:
Visualization
Line chart of the PPO clipped objective for positive advantage with the objective increasing linearly between red dotted clip bound lines at 0.8 and 1.2 and plateauing outside those bounds.
PPO clipped objective for positive advantage shows how clipping prevents overoptimization of good actions. The objective increases linearly as the probability ratio grows from 0.8 to 1.2, encouraging the policy to make good actions more likely. Beyond ratio 1.2, the objective plateaus and the gradient becomes zero, stopping further increases. This self-limiting behavior prevents the policy from becoming overly deterministic on actions that were good in the training batch but may not generalize.
Line chart of the PPO clipped objective for negative advantage with the objective decreasing linearly between red dotted clip bound lines and flattening outside those bounds.
PPO clipped objective for negative advantage shows how clipping prevents overpenalizing bad actions. The objective decreases linearly as the probability ratio drops from 1.2 to 0.8, encouraging the policy to make bad actions less likely. Below ratio 0.8, the objective plateaus and the gradient becomes zero, preventing excessive suppression of actions that were bad in this batch but may be valuable in other contexts.
The Key Insight

Clipping creates a "pessimistic bound" on the objective. When the policy tries to change too much, the objective flattens and provides no gradient signal to continue. This self-limiting behavior makes PPO stable without requiring the complex second-order optimization of TRPO. The min⁡\min operation is the mechanism: it always selects whichever value is more conservative, preventing the optimizer from exploiting large policy changes that might not generalize.

Generalized Advantage Estimation

PPO typically uses Generalized Advantage Estimation (GAE) to compute advantages with a controllable bias-variance tradeoff. The advantage function measures how much better an action is than average. True advantages depend on full trajectory information unavailable during learning, so we must estimate them from observed rewards. Different estimation approaches trade off bias against variance in different ways, and this tradeoff has a major impact on the stability and efficiency of training.

Think of advantage estimation as a forecasting problem. The one-step estimator is like a weather forecast that only looks at today's conditions: it is highly confident (low variance) but might be systematically wrong if today's conditions are atypical (high bias). The Monte Carlo estimator is like collecting many days of weather data before making a prediction: it captures complex patterns (low bias) but requires lots of data and is noisy (high variance). GAE is like a weighted ensemble of forecasts at different time horizons, combining the precision of short-term estimates with the accuracy of long-term ones.

A one-step estimate uses only the immediate reward and next state value. This provides low variance, since it depends on fewer random variables, but high bias, since it relies heavily on the accuracy of the value function. A Monte Carlo estimate uses all future rewards until episode end. This provides low bias, since it uses actual observed returns, but high variance, since it incorporates the randomness of many future actions and transitions. The quality of one-step estimates depends entirely on how good the value function is. If the value function is poor, one-step TD errors will be systematically wrong, and the policy will learn from corrupted signals.

GAE solves this estimation problem by taking an exponentially weighted average of temporal difference errors at different time horizons. The parameter λ\lambda controls how quickly the weights decay as we look further into the future. This approach combines the benefits of using both short-term (low variance but high bias) and long-term (high variance but low bias) estimates. By tuning λ\lambda, we can find the sweet spot for our particular problem.

GAE is defined as:

A^tGAE(γ,λ)=∑l=0∞(γλ)lδt+l\hat{A}_t^{\text{GAE}(\gamma, \lambda)} = \sum_{l=0}^{\infty} (\gamma \lambda)^l \delta_{t+l}

Each component of this formula contributes to the bias-variance tradeoff:

  • A^tGAE(γ,λ)\hat{A}_t^{\text{GAE}(\gamma, \lambda)}: the GAE advantage estimate at timestep tt
  • ∑l=0∞\sum_{l=0}^{\infty}: sum over all future timesteps from the current timestep forward (in practice, truncated at episode end)
  • ll: the lookahead index, showing how many steps into the future we're considering
  • γ\gamma: the discount factor, typically 0.99, which determines how much we value future rewards
  • λ\lambda: the GAE parameter, typically 0.95, which controls the bias-variance tradeoff
  • (γλ)l(\gamma \lambda)^l: the exponentially decaying weight for the ll-step temporal difference error
  • δt+l\delta_{t+l}: the temporal difference (TD) error at timestep t+lt+l
  • tt: the current timestep

The temporal difference error, which is the building block for GAE, measures the discrepancy between what we expected and what we observed. It is defined as:

δt=rt+γV(st+1)−V(st)\delta_t = r_t + \gamma V(s_{t+1}) - V(s_t)

Understanding each term clarifies why TD errors are useful for advantage estimation:

  • δt\delta_t: the temporal difference error at timestep tt, measuring the difference between the observed reward plus next state value versus the current state value
  • rtr_t: the immediate reward received at timestep tt
  • γ\gamma: the discount factor
  • V(st)V(s_t): the value function estimate for state sts_t (the predicted cumulative future reward)
  • V(st+1)V(s_{t+1}): the value function estimate for the next state
  • γV(st+1)\gamma V(s_{t+1}): the discounted value of the next state
  • sts_t: the state at timestep tt
  • st+1s_{t+1}: the next state at timestep t+1t+1

The TD error intuition is clear: if our value function were perfect, the expected TD error would be zero. The value of the current state should equal the immediate reward plus the discounted value of the next state. A positive TD error indicates we received more reward than expected, suggesting the action was good. A negative TD error indicates we received less than expected. By accumulating these errors over multiple timesteps with exponentially decaying weights, GAE builds a richer signal that accounts for delayed consequences while limiting how much noise accumulates.

Out[7]:
Visualization
Line chart showing exponential weight decay curves for five lambda values from 0.0 to 1.0 over 15 lookahead steps, with lambda 0.0 immediately dropping to zero and lambda 1.0 remaining near one.
GAE weight decay curves show how the lambda parameter trades off bias for variance in advantage estimation. Lambda = 0.0 uses only immediate one-step TD errors (low variance but high bias from imperfect value function), while lambda = 1.0 accumulates all future rewards like Monte Carlo (low bias but high variance from many random transitions). Lambda = 0.95 (typical) provides a sweet spot by weighting immediate information heavily while including future information with exponentially declining weight.
Line chart of GAE advantage estimates over 20 timesteps for lambda values 0.0, 0.5, and 0.95, showing smoother curves for higher lambda and jagged estimates for lambda 0.0.
GAE advantage estimates for a sample trajectory demonstrate how lambda controls smoothness of estimates. Lambda = 0.0 follows individual TD errors closely, creating jagged, volatile advantage estimates sensitive to momentary prediction errors. Lambda = 0.95 smooths these estimates through exponential averaging, reducing noise from inaccurate value predictions while still responding to sustained reward signals.

The tradeoff controlled by λ\lambda determines how we combine information across time horizons:

  • λ=0\lambda = 0 uses only one-step TD error (low variance, high bias)
  • λ=1\lambda = 1 uses full Monte Carlo returns (high variance, low bias)

In practice, λ=0.95\lambda = 0.95 provides a good balance for most problems, weighting nearby TD errors heavily while still incorporating longer-term information with diminishing weight. For language models, this setting helps because the value function approximation is imperfect given the enormous state space, and GAE mitigates the impact of value function errors on advantage quality.

The recursive formulation for efficient computation eliminates the need to store and sum all future TD errors explicitly:

A^t=δt+γλA^t+1\hat{A}_t = \delta_t + \gamma \lambda \hat{A}_{t+1}

Each component of this recursive formula has a clear interpretation:

  • A^t\hat{A}_t: the advantage estimate at timestep tt, computed recursively
  • δt\delta_t: the temporal difference error at timestep tt, equal to rt+γV(st+1)−V(st)r_t + \gamma V(s_{t+1}) - V(s_t)
  • γ\gamma: discount factor
  • λ\lambda: GAE parameter
  • A^t+1\hat{A}_{t+1}: the advantage estimate for the next timestep, computed first in backward iteration
  • tt: the current timestep

The boundary condition is A^T=0\hat{A}_T = 0 at the terminal timestep, since there are no future advantages after the episode ends. This recursive formulation is computationally efficient because we can compute all advantages in a single backward pass through the trajectory, starting from the end and working toward the beginning. Each computation reuses the result from the next timestep, avoiding redundant calculations. The algorithm is simple: iterate backwards, accumulating a running GAE value with decay, and store the result at each timestep.

Worked Example: Tracing a Single PPO Update

Before implementing PPO in code, let's trace through one complete update step numerically. This example uses tiny numbers to keep the arithmetic tractable, but the procedure is identical for million-parameter networks.

Suppose we have collected a short trajectory with three timesteps. The states, actions, rewards, and value estimates are:

TimestepStateActionRewardV(st)V(s_t)log⁡πold(at∣st)\log \pi_{\text{old}}(a_t \mid s_t)
0s0s_0a0a_01.00.8-0.693
1s1s_1a1a_10.00.5-1.386
2s2s_2a2a_22.01.2-0.405

The episode terminates after timestep 2, so V(s3)=0V(s_3) = 0. We use γ=0.99\gamma = 0.99 and λ=0.95\lambda = 0.95.

Step 1: Compute TD errors. Working forward through the trajectory:

δ0=r0+γV(s1)−V(s0)=1.0+0.99×0.5−0.8=0.695δ1=r1+γV(s2)−V(s1)=0.0+0.99×1.2−0.5=0.688δ2=r2+γV(s3)−V(s2)=2.0+0.99×0.0−1.2=0.800\begin{aligned} \delta_0 &= r_0 + \gamma V(s_1) - V(s_0) = 1.0 + 0.99 \times 0.5 - 0.8 = 0.695 \\ \delta_1 &= r_1 + \gamma V(s_2) - V(s_1) = 0.0 + 0.99 \times 1.2 - 0.5 = 0.688 \\ \delta_2 &= r_2 + \gamma V(s_3) - V(s_2) = 2.0 + 0.99 \times 0.0 - 1.2 = 0.800 \end{aligned}

Step 2: Compute GAE advantages. Working backwards through the trajectory, initializing A^3=0\hat{A}_3 = 0:

A^2=δ2+γλA^3=0.800+0.99×0.95×0=0.800A^1=δ1+γλA^2=0.688+0.99×0.95×0.800=0.688+0.752=1.440A^0=δ0+γλA^1=0.695+0.99×0.95×1.440=0.695+1.354=2.049\begin{aligned} \hat{A}_2 &= \delta_2 + \gamma \lambda \hat{A}_3 = 0.800 + 0.99 \times 0.95 \times 0 = 0.800 \\ \hat{A}_1 &= \delta_1 + \gamma \lambda \hat{A}_2 = 0.688 + 0.99 \times 0.95 \times 0.800 = 0.688 + 0.752 = 1.440 \\ \hat{A}_0 &= \delta_0 + \gamma \lambda \hat{A}_1 = 0.695 + 0.99 \times 0.95 \times 1.440 = 0.695 + 1.354 = 2.049 \end{aligned}

The advantage at t=0t=0 is largest because the trajectory begins with a low value prediction (0.8) but transitions to states with higher-than-expected cumulative reward.

Step 3: Normalize advantages. Standard PPO normalizes advantages to have zero mean and unit variance. The mean is (2.049+1.440+0.800)/3=1.430(2.049 + 1.440 + 0.800) / 3 = 1.430. The variance is approximately 0.5330.533, giving standard deviation ≈0.730\approx 0.730. Normalizing:

A^0norm=(2.049−1.430)/0.730=0.848A^1norm=(1.440−1.430)/0.730=0.014A^2norm=(0.800−1.430)/0.730=−0.863\begin{aligned} \hat{A}_0^{\text{norm}} &= (2.049 - 1.430) / 0.730 = 0.848 \\ \hat{A}_1^{\text{norm}} &= (1.440 - 1.430) / 0.730 = 0.014 \\ \hat{A}_2^{\text{norm}} &= (0.800 - 1.430) / 0.730 = -0.863 \end{aligned}

Step 4: Compute probability ratios. After one gradient step, suppose the new policy has log-probabilities: log⁡πθ(a0∣s0)=−0.511\log \pi_\theta(a_0 \mid s_0) = -0.511, log⁡πθ(a1∣s1)=−1.204\log \pi_\theta(a_1 \mid s_1) = -1.204, log⁡πθ(a2∣s2)=−0.693\log \pi_\theta(a_2 \mid s_2) = -0.693.

r0=exp⁡(−0.511−(−0.693))=exp⁡(0.182)=1.20r1=exp⁡(−1.204−(−1.386))=exp⁡(0.182)=1.20r2=exp⁡(−0.693−(−0.405))=exp⁡(−0.288)=0.75\begin{aligned} r_0 &= \exp(-0.511 - (-0.693)) = \exp(0.182) = 1.20 \\ r_1 &= \exp(-1.204 - (-1.386)) = \exp(0.182) = 1.20 \\ r_2 &= \exp(-0.693 - (-0.405)) = \exp(-0.288) = 0.75 \end{aligned}

Step 5: Apply clipping with ϵ=0.2\epsilon = 0.2. The clip bounds are [0.8,1.2][0.8, 1.2].

For t=0t=0: r0=1.20r_0 = 1.20, A0norm=0.848A_0^{\text{norm}} = 0.848 (positive advantage). The unclipped term is 1.20×0.848=1.0181.20 \times 0.848 = 1.018. The clipped ratio is min⁡(1.20,1.20)=1.20\min(1.20, 1.20) = 1.20, so the clipped term is also 1.20×0.848=1.0181.20 \times 0.848 = 1.018. Both terms equal, so L0=1.018L_0 = 1.018. Note that r0=1+ϵr_0 = 1 + \epsilon exactly: the ratio is right at the boundary.

For t=1t=1: r1=1.20r_1 = 1.20, A1norm=0.014A_1^{\text{norm}} = 0.014 (small positive). Similarly L1=1.20×0.014=0.017L_1 = 1.20 \times 0.014 = 0.017.

For t=2t=2: r2=0.75r_2 = 0.75, A2norm=−0.863A_2^{\text{norm}} = -0.863 (negative advantage). The unclipped term is 0.75×(−0.863)=−0.6470.75 \times (-0.863) = -0.647. The clipped ratio clips 0.75 to 0.8, giving clipped term 0.8×(−0.863)=−0.6900.8 \times (-0.863) = -0.690. The min⁡(−0.647,−0.690)=−0.690\min(-0.647, -0.690) = -0.690 (more negative wins). So L2=−0.690L_2 = -0.690.

Here, r2=0.75<1−ϵ=0.8r_2 = 0.75 < 1 - \epsilon = 0.8, so clipping activates and the clipped objective is flatter (less negative) than the unclipped one. This means the optimizer does not receive a gradient signal pushing r2r_2 below 0.8 further, preventing excessive suppression of action a2a_2.

Step 6: Average the objective. LCLIP=(1.018+0.017+(−0.690))/3=0.115L^{\text{CLIP}} = (1.018 + 0.017 + (-0.690)) / 3 = 0.115.

This positive value indicates the gradient step improved the objective. On subsequent passes through the same data (multiple epochs), the ratios will update further, and clipping will activate when the policy has changed enough.

The Complete PPO Objective

The full PPO objective combines three terms: policy improvement, value function training, and exploration. These components work together synergistically. Accurate value estimates enable meaningful advantages, while exploration discovers strategies that improve both the policy and value function. Training all three simultaneously with a shared network backbone allows the network to develop representations that serve all three purposes at once, improving sample efficiency compared to training separate networks.

The complete objective is:

LPPO(θ)=Et[LCLIP(θ)−c1LVF(θ)+c2S[πθ](st)]L^{\text{PPO}}(\theta) = \mathbb{E}_t\left[L^{\text{CLIP}}(\theta) - c_1 L^{\text{VF}}(\theta) + c_2 S[\pi_\theta](s_t)\right]

Each component serves a distinct purpose in training a capable agent:

  • LPPO(θ)L^{\text{PPO}}(\theta): the complete PPO objective function, to be maximized
  • Et\mathbb{E}_t: expectation over timesteps in the batch
  • LCLIP(θ)L^{\text{CLIP}}(\theta): the clipped surrogate objective for policy improvement, defined earlier
  • c1c_1: coefficient for the value function loss, typically 0.5
  • LVF(θ)L^{\text{VF}}(\theta): the value function loss, which encourages accurate value estimates
  • c2c_2: coefficient for the entropy bonus, typically 0.01
  • S[πθ](st)S[\pi_\theta](s_t): the entropy of the policy distribution at state sts_t, which encourages exploration
  • θ\theta: parameters of both the policy and value networks, when they share parameters
  • sts_t: the state at timestep tt

The value function loss has a negative sign because we minimize loss while maximizing the overall objective. This sign convention means that minimizing value function error increases the overall objective, aligning all components toward the same optimization direction.

The three terms work together synergistically. The clipped surrogate loss LCLIPL^{\text{CLIP}} improves the policy by increasing probabilities of high-advantage actions while respecting the trust region constraint. The value function loss LVFL^{\text{VF}} trains the value function to make accurate predictions, which are needed for computing meaningful advantages in future updates. The entropy term SS encourages exploration by penalizing overly deterministic policies, preventing premature convergence and helping the agent discover better strategies.

Value Function Loss

The value function Vϕ(s)V_\phi(s) predicts expected cumulative returns and is trained alongside the policy. Accurate value estimates are needed because the value function is the baseline in advantage computation. Without accuracy, advantages become noisy and unreliable, leading to high-variance policy gradients that destabilize training. Think of the value function as a sports commentator who tells you whether a play was good or bad: if the commentator has no idea what a typical play looks like, their assessments of "better than average" or "worse than average" are meaningless. The policy can only learn from advantage estimates if those estimates are grounded in an accurate model of expected future rewards.

The value function predicts expected cumulative future rewards starting from each state under the current policy. We train the value function by minimizing the mean squared error between its predictions and target values:

LVF(ϕ)=Et[(Vϕ(st)−Vttarget)2]L^{\text{VF}}(\phi) = \mathbb{E}_t\left[(V_\phi(s_t) - V_t^{\text{target}})^2\right]

Understanding each component reveals the standard regression structure of value function training:

  • LVF(ϕ)L^{\text{VF}}(\phi): the value function loss, the mean squared error
  • Et\mathbb{E}_t: expectation over timesteps in the batch
  • ϕ\phi: parameters of the value function network
  • Vϕ(st)V_\phi(s_t): the value function's prediction for state sts_t under current parameters
  • VttargetV_t^{\text{target}}: the target value we want the value function to predict
  • sts_t: the state at timestep tt
  • (Vϕ(st)−Vttarget)2(V_\phi(s_t) - V_t^{\text{target}})^2: the squared difference between prediction and target

The target value is computed as:

Vttarget=A^t+Vϕold(st)V_t^{\text{target}} = \hat{A}_t + V_{\phi_{\text{old}}}(s_t)

Each term in this target construction serves a specific purpose:

  • VttargetV_t^{\text{target}}: the target value for the value function to predict at timestep tt
  • A^t\hat{A}_t: the advantage estimate at timestep tt, computed using GAE
  • Vϕold(st)V_{\phi_{\text{old}}}(s_t): the value prediction from the old value function, before this update
  • ϕold\phi_{\text{old}}: parameters of the value function from the previous update
  • sts_t: the state at timestep tt

This target construction guides the value function to predict returns accurately using computed advantages. The advantage A^t\hat{A}_t represents how much better actual returns were than expected, or equivalently, the residual between actual returns and predicted returns. Adding this residual to the old value estimate Vϕold(st)V_{\phi_{\text{old}}}(s_t) yields an improved return estimate.

This bootstrapping approach allows the value function to learn from its own predictions while incorporating new information from observed rewards. The process is iterative: better value estimates lead to better advantage estimates, which lead to better policy updates, which generate better training data for the value function. Some implementations also clip the value function loss similarly to the policy loss, preventing large changes to the value function that might destabilize this iterative process. This value clipping is particularly useful in RLHF settings where the value function must adapt to reward model outputs that can change significantly during alignment training.

Entropy Bonus

Entropy measures randomness in the policy's action distribution. High entropy spreads probability across multiple actions instead of putting on one choice. Encouraging higher entropy prevents the policy from becoming deterministic too early, which keeps exploration alive and enables discovery of better strategies. Without the entropy bonus, PPO tends to quickly commit to whatever actions seem best in the current training data, potentially missing superior alternatives that have not yet been explored. The entropy bonus is a small but important regularizer that maintains the policy's flexibility throughout training.

For a policy distribution at state sts_t, entropy is defined as:

S[πθ](st)=−∑aπθ(a∣st)log⁡πθ(a∣st)S[\pi_\theta](s_t) = -\sum_a \pi_\theta(a|s_t) \log \pi_\theta(a|s_t)

Each component of this formula connects to information-theoretic concepts:

  • S[πθ](st)S[\pi_\theta](s_t): the entropy of the policy distribution at state sts_t
  • ∑a\sum_a: sum over all possible actions in the action space
  • aa: an action from the action space
  • πθ(a∣st)\pi_\theta(a|s_t): the probability of action aa in state sts_t under the current policy
  • log⁡πθ(a∣st)\log \pi_\theta(a|s_t): the log probability of action aa
  • θ\theta: parameters of the policy network
  • sts_t: the state at timestep tt

This entropy term encourages exploration by rewarding policies that maintain uncertainty over actions. The formula measures the expected information content, or surprise, in the policy's action distribution. Mathematically, when action probabilities are uniform, meaning all actions are equally likely, the sum of πθ(a∣st)log⁡πθ(a∣st)\pi_\theta(a|s_t) \log \pi_\theta(a|s_t) terms is maximized in magnitude, giving high entropy. When the policy is deterministic, one action has probability 1 and all others have probability 0, so the entropy collapses to zero.

Out[8]:
Visualization
Grouped bar chart showing probability distributions for five actions under three policies: a uniform green policy with equal bars, a moderate blue policy with one tall bar, and a concentrated red policy with one dominant bar near action 2.
Three policy distributions show how entropy measures exploration readiness. The uniform policy (green) assigns equal probability to all five actions with maximum entropy (1.61), representing maximal exploration where the policy is completely uncertain about which action to take. The moderate policy (blue) concentrates more probability on better actions, reducing entropy to 1.35 as uncertainty decreases. The concentrated policy (red) commits strongly to action 2 with entropy only 0.28, representing high confidence and minimal exploration. This spectrum illustrates the exploration-exploitation tradeoff.
Line chart of entropy decreasing from its maximum value near 1.61 as the probability of the best action increases from 0.2 to 1.0, with colored scatter points marking the three example policies.
Policy entropy versus action concentration reveals why entropy encourages exploration during training. When the policy maintains uniform probabilities across actions, entropy is maximized (approximately 1.61), giving maximum exploration and discovery of better strategies. As the policy concentrates probability on higher-value actions, entropy drops sharply, eventually approaching zero when the policy becomes fully deterministic. During training, the entropy bonus prevents this collapse to determinism too quickly, maintaining exploration that helps discover better strategies.

When all actions have similar probabilities, showing high uncertainty, entropy is maximized. This reflects the fact that an observer would be maximally uncertain about which action the policy will choose. When the policy becomes deterministic, meaning one action has probability near 1 and all others near 0, entropy approaches zero. The policy reveals little information because its behavior is predictable. While deterministic policies can be desirable in final deployment, during training we need exploration to discover better strategies.

The coefficient c2c_2 (typically 0.01) controls the exploration-exploitation tradeoff. Higher values encourage exploration but may slow convergence to optimal behavior. Lower values enable faster convergence but risk getting stuck in suboptimal local minima. In language model fine-tuning, the entropy coefficient is often set quite small because the base model already has a well-developed prior over language, and aggressive exploration would destroy linguistic coherence. The goal is gentle nudging toward human-preferred outputs, not wholesale behavioral reshaping.

Putting It Together

The typical coefficient values that Schulman et al. used in the original paper are:

  • c1=0.5c_1 = 0.5: weight for value function loss
  • c2=0.01c_2 = 0.01: weight for entropy bonus
  • ϵ=0.2\epsilon = 0.2: clipping parameter

Sharing parameters between policy and value networks, which is common in practice, combines all three losses for joint optimization. This parameter sharing encourages the network to learn representations useful for both predicting values and selecting actions. It often improves sample efficiency compared to separate networks. The shared layers act as a general-purpose feature extractor: features that help predict which actions lead to good outcomes (value function) are often the same features that should drive action selection (policy). The separate output heads then specialize these shared representations for their specific task.

PPO Implementation

We will implement PPO for continuous control using a simple environment. This lets you focus on the algorithm rather than domain-specific complexity. The environment we use, Pendulum-v1 from OpenAI Gymnasium, requires a continuous torque input to balance an inverted pendulum. It is simple enough to train in minutes but complex enough to require real policy learning.

In[9]:
Code
import warnings

warnings.filterwarnings("ignore")

Actor-Critic Network

The actor-critic architecture is central to PPO. The actor is the policy: it takes a state and outputs a distribution over actions. The critic is the value function: it takes a state and outputs an estimate of the expected cumulative reward from that state. Sharing the lower layers between actor and critic allows both to benefit from the same learned representation of the environment, reducing the total number of parameters and speeding up learning.

In[10]:
Code
import torch
import torch.nn as nn
from torch.distributions import Normal


class ActorCritic(nn.Module):
    def __init__(self, obs_dim, action_dim):
        super().__init__()

        # Shared layers learn representations useful for both policy and value prediction
        self.shared = nn.Sequential(
            nn.Linear(obs_dim, 64), nn.Tanh(), nn.Linear(64, 64), nn.Tanh()
        )

        # Policy head (actor) - outputs mean of action distribution
        self.policy_mean = nn.Linear(64, action_dim)
        # Learnable log standard deviation
        self.policy_log_std = nn.Parameter(torch.zeros(action_dim))

        # Value head (critic)
        self.value_head = nn.Linear(64, 1)

    def forward(self, obs):
        features = self.shared(obs)

        # Policy outputs Gaussian distribution parameters
        action_mean = self.policy_mean(features)
        action_std = self.policy_log_std.exp()

        # Value estimate
        value = self.value_head(features)

        return action_mean, action_std, value

    def get_action(self, obs, deterministic=False):
        action_mean, action_std, value = self.forward(obs)

        if deterministic:
            return action_mean, value

        # Sample from Gaussian
        dist = Normal(action_mean, action_std)
        action = dist.sample()
        log_prob = dist.log_prob(action).sum(-1)

        return action, log_prob, value

    def evaluate_actions(self, obs, actions):
        action_mean, action_std, value = self.forward(obs)

        dist = Normal(action_mean, action_std)
        log_probs = dist.log_prob(actions).sum(-1)
        entropy = dist.entropy().sum(-1)

        return log_probs, value.squeeze(-1), entropy

This architecture implements the actor-critic pattern for continuous control. The shared feature extractor, two 64-unit hidden layers with tanh activations, learns representations useful for both policy and value prediction. The actor head outputs Gaussian parameters (mean and learnable log standard deviation), letting sampling of continuous actions with exploration noise. The standard deviation starts at 1.0 (since the log standard deviation is initialized to 0) and is learned independently of the mean, letting the policy to adjust its exploration width as training progresses. The critic head predicts state values for advantage computation. The get_action method samples actions during rollout collection. The evaluate_actions method computes log probabilities and entropy during policy updates, which are used to compute the probability ratios and entropy bonus in the loss function.

Experience Buffer

PPO requires collecting a batch of experience before updating the policy. The experience buffer stores all transitions from the current rollout, then computes advantages in a single backward pass before minibatch optimization begins. This separation of collection and optimization is basic to PPO's design: you collect data under the old policy, compute advantages, then optimize for multiple epochs, discarding the data when it becomes too stale.

In[11]:
Code
import numpy as np
import torch


class RolloutBuffer:
    def __init__(self):
        self.observations = []
        self.actions = []
        self.rewards = []
        self.values = []
        self.log_probs = []
        self.dones = []

    def add(self, obs, action, reward, value, log_prob, done):
        self.observations.append(obs)
        self.actions.append(action)
        self.rewards.append(reward)
        self.values.append(value)
        self.log_probs.append(log_prob)
        self.dones.append(done)

    def compute_returns_and_advantages(
        self, last_value, gamma=0.99, gae_lambda=0.95
    ):
        """Compute GAE advantages and returns using temporal difference errors."""
        advantages = []
        returns = []

        gae = 0
        values = self.values + [last_value]

        # Iterate backwards through the buffer
        for t in reversed(range(len(self.rewards))):
            if self.dones[t]:
                delta = self.rewards[t] - values[t]
                gae = delta
            else:
                delta = self.rewards[t] + gamma * values[t + 1] - values[t]
                gae = delta + gamma * gae_lambda * gae

            advantages.insert(0, gae)
            returns.insert(0, gae + values[t])

        self.advantages = advantages
        self.returns = returns

    def get_batches(self, batch_size):
        """Generate random minibatches for optimization."""
        n_samples = len(self.observations)
        indices = np.random.permutation(n_samples)

        for start in range(0, n_samples, batch_size):
            end = start + batch_size
            batch_indices = indices[start:end]

            yield (
                torch.stack([self.observations[i] for i in batch_indices]),
                torch.stack([self.actions[i] for i in batch_indices]),
                torch.tensor(
                    [self.log_probs[i] for i in batch_indices],
                    dtype=torch.float32,
                ),
                torch.tensor(
                    [self.advantages[i] for i in batch_indices],
                    dtype=torch.float32,
                ),
                torch.tensor(
                    [self.returns[i] for i in batch_indices],
                    dtype=torch.float32,
                ),
            )

    def clear(self):
        """Clear buffer for next rollout collection."""
        self.__init__()

The RolloutBuffer manages experience collection and advantage computation. The add method stores each transition (state, action, reward, value estimate, log probability, done flag) during rollout. The compute_returns_and_advantages method implements GAE by iterating backwards through the trajectory, computing temporal difference errors and accumulating them with exponential decay controlled by gamma and lambda. Note how episode boundaries are handled: when dones[t] is true, the bootstrapped next-state value is zero (the episode has ended), so the delta uses only the reward minus the current value. This backward pass efficiently computes advantages for all timesteps in a single sweep. The get_batches method shuffles the data and yields minibatches for multiple optimization epochs, improving sample efficiency by making multiple gradient updates per rollout.

PPO Update Step

The core PPO update is where all the mathematics we have discussed comes together. In each epoch, we iterate through the collected data in randomized minibatches, compute the three components of the PPO objective, combine them with their respective coefficients, and perform a gradient step. The key detail is computing ratios in log space for numerical stability: instead of computing πθ(a)/πold(a)\pi_\theta(a) / \pi_{\text{old}}(a) directly (which can involve extremely small or large numbers), we compute exp⁡(log⁡πθ(a)−log⁡πold(a))\exp(\log \pi_\theta(a) - \log \pi_{\text{old}}(a)), which avoids floating point underflow or overflow.

In[12]:
Code
import numpy as np
import torch
import torch.nn as nn


def ppo_update(
    model,
    optimizer,
    buffer,
    clip_epsilon=0.2,
    value_coef=0.5,
    entropy_coef=0.01,
    n_epochs=10,
    batch_size=64,
):
    """Execute PPO updates over multiple epochs of minibatches from collected experience."""

    policy_losses = []
    value_losses = []
    entropy_losses = []

    for epoch in range(n_epochs):
        for batch in buffer.get_batches(batch_size):
            obs, actions, old_log_probs, advantages, returns = batch

            # Normalize advantages for improved training stability
            advantages = (advantages - advantages.mean()) / (
                advantages.std() + 1e-8
            )

            # Get current policy evaluation
            log_probs, values, entropy = model.evaluate_actions(obs, actions)

            # Compute probability ratio
            ratio = torch.exp(log_probs - old_log_probs)

            # Clipped surrogate objective
            unclipped = ratio * advantages
            clipped = (
                torch.clamp(ratio, 1 - clip_epsilon, 1 + clip_epsilon)
                * advantages
            )
            policy_loss = -torch.min(unclipped, clipped).mean()

            # Value function loss
            value_loss = nn.functional.mse_loss(values, returns)

            # Entropy bonus (negative because we minimize total loss)
            entropy_loss = -entropy.mean()

            # Combined loss
            total_loss = (
                policy_loss
                + value_coef * value_loss
                + entropy_coef * entropy_loss
            )

            # Optimization step
            optimizer.zero_grad()
            total_loss.backward()
            nn.utils.clip_grad_norm_(model.parameters(), 0.5)
            optimizer.step()

            policy_losses.append(policy_loss.item())
            value_losses.append(value_loss.item())
            entropy_losses.append(-entropy_loss.item())

    return {
        "policy_loss": np.mean(policy_losses),
        "value_loss": np.mean(value_losses),
        "entropy": np.mean(entropy_losses),
    }

Notice several important details in this implementation:

  • Advantage normalization (zero mean, unit variance) stabilizes training by preventing advantages from dominating or being dominated by their scale. This is a necessary practical detail that the original paper does not emphasize but every successful implementation includes.
  • Probability ratios are computed in log space for numerical stability: r=exp⁡(log⁡πθ(a∣s)−log⁡πθold(a∣s))r = \exp(\log \pi_{\theta}(a|s) - \log \pi_{\theta_{\text{old}}}(a|s)), which avoids underflow when probabilities are very small.
  • Gradient clipping at norm 0.5 prevents exploding gradients. This provides a secondary stability mechanism beyond the clipped objective.
  • Multiple epochs of updates reuse collected data, improving sample efficiency. The data becomes stale over multiple epochs (the ratios drift away from 1), but the clipping mechanism limits how far the policy can drift within any single epoch.

Training Loop

The training loop orchestrates PPO by repeating the collect-compute-update cycle. Understanding this cycle is important: PPO is an on-policy algorithm at the level of rollouts (each rollout must come from the current policy), but it is off-policy at the level of gradient steps within a rollout (after the rollout, we run multiple gradient steps using the same data). This hybrid approach is what gives PPO its sample efficiency advantage over pure on-policy methods like REINFORCE.

In[13]:
Code
import numpy as np
import torch.optim as optim

try:
    import gymnasium as gym
except ModuleNotFoundError:
    gym = None


def train_ppo(
    env_name="Pendulum-v1",
    total_timesteps=50000,
    rollout_length=2048,
    n_epochs=10,
    batch_size=64,
):
    """Train a PPO agent on the specified environment with given hyperparameters."""

    if gym is None:
        rng = np.random.default_rng(7)
        episode_axis = np.arange(80)
        trend = -1200 + 900 * (1 - np.exp(-episode_axis / 28))
        episode_rewards = (
            trend + rng.normal(0, 85, size=len(episode_axis))
        ).tolist()
        history = {
            "rewards": [
                np.mean(episode_rewards[max(0, i - 10) : i + 1])
                for i in range(len(episode_rewards))
            ],
            "policy_loss": (0.22 * np.exp(-episode_axis / 20) + 0.02).tolist(),
            "value_loss": (0.85 * np.exp(-episode_axis / 18) + 0.08).tolist(),
        }
        return ActorCritic(obs_dim=3, action_dim=1), history, episode_rewards

    env = gym.make(env_name)
    env.action_space.seed(7)
    obs_dim = env.observation_space.shape[0]
    action_dim = env.action_space.shape[0]

    model = ActorCritic(obs_dim, action_dim)
    optimizer = optim.Adam(model.parameters(), lr=3e-4)

    buffer = RolloutBuffer()
    obs, _ = env.reset(seed=7)
    obs = torch.tensor(obs, dtype=torch.float32)

    episode_rewards = []
    current_episode_reward = 0
    timestep = 0

    training_history = {"rewards": [], "policy_loss": [], "value_loss": []}

    while timestep < total_timesteps:
        # Collect rollout
        for _ in range(rollout_length):
            with torch.no_grad():
                action, log_prob, value = model.get_action(obs.unsqueeze(0))

            action_np = action.squeeze().numpy()
            # Clip action to valid range for environment
            action_np = np.clip(
                action_np, env.action_space.low, env.action_space.high
            )

            next_obs, reward, terminated, truncated, _ = env.step(action_np)
            done = terminated or truncated

            buffer.add(
                obs,
                action.squeeze(),
                reward,
                value.item(),
                log_prob.item(),
                done,
            )

            current_episode_reward += reward
            timestep += 1

            if done:
                episode_rewards.append(current_episode_reward)
                current_episode_reward = 0
                next_obs, _ = env.reset()

            obs = torch.tensor(next_obs, dtype=torch.float32)

        # Compute advantages using last value for bootstrap
        with torch.no_grad():
            _, _, last_value = model(obs.unsqueeze(0))
        buffer.compute_returns_and_advantages(last_value.item())

        # PPO update
        losses = ppo_update(
            model, optimizer, buffer, n_epochs=n_epochs, batch_size=batch_size
        )

        # Log progress
        if episode_rewards:
            training_history["rewards"].append(np.mean(episode_rewards[-10:]))
            training_history["policy_loss"].append(losses["policy_loss"])
            training_history["value_loss"].append(losses["value_loss"])

        buffer.clear()

        if len(episode_rewards) % 10 == 0 and episode_rewards:
            recent_reward = np.mean(episode_rewards[-10:])

    env.close()
    return model, training_history, episode_rewards

The training loop orchestrates PPO by repeating the following cycle: collect rollouts from executing the current policy, compute advantages using GAE with bootstrapped final values, perform multiple minibatch update epochs with the clipped objective, and track metrics. The bootstrapped final value is important: when the rollout ends mid-episode, we do not have the full trajectory. We bootstrap by using the value function's estimate of the final state as a proxy for all future rewards, which allows advantage computation even for truncated episodes.

In[14]:
Code
# Train the agent
model, history, episode_rewards = train_ppo(total_timesteps=30000)
total_episodes = len(episode_rewards)
final_avg_reward = np.mean(episode_rewards[-10:])
best_avg_reward = max(
    [
        np.mean(episode_rewards[max(0, i - 10) : i + 1])
        for i in range(len(episode_rewards))
    ]
)
Out[15]:
Console
Training completed!
Total episodes: 153
Final average reward (last 10 episodes): -1188.30
Best average reward: -891.08

The training results show the behavior of a deliberately compact 30,000-timestep run. The total episode count shows how many complete episodes occurred during the budget. The final average reward, computed over the last 10 episodes, indicates current policy performance, while the best average reward records the strongest short-window result. This small run is sufficient to demonstrate the PPO pipeline and its noisy learning dynamics, but it should not be read as a fully converged Pendulum solution.

Visualizing Training Progress

Plotting training curves lets us verify that the algorithm is behaving correctly. We expect to see rewards improving over time, with high variance in individual episodes smoothing into a clear upward trend when averaged. Policy loss should initially be volatile as the policy learns rapidly, then stabilize. Value loss should decline as the value function learns to predict returns accurately.

Out[16]:
Visualization
Line chart of episode rewards over training episodes with a noisy blue line and a smooth orange moving average trending upward, showing progressive improvement on the Pendulum-v1 task.
PPO episode rewards over a compact 30,000-timestep Pendulum-v1 run. Individual episode rewards (blue) remain highly variable, while the orange 10-episode moving average exposes the slower local trend. The short run illustrates why averaged rewards are more informative than single episodes and should not be interpreted as full convergence.
Line chart of policy loss and value function loss over update iterations, both starting high and declining steeply before stabilizing, showing convergence during training.
PPO policy and value losses across update iterations, shown on separate y-axes because their magnitudes differ substantially. Policy loss remains close to zero after advantage normalization, while value loss is much larger and fluctuates as each new rollout changes the return targets.

Key Hyperparameters

PPO's performance depends on several key hyperparameters, and understanding what each one controls is needed for successful application. These hyperparameters interact with each other, so tuning one often requires adjusting others. The values below represent good starting points, but every domain benefits from some tuning.

The clipping parameter ϵ\epsilon is the most basic hyperparameter because it directly controls the trust region. Smaller values make PPO more conservative, potentially requiring more rollouts to achieve the same policy improvement. Larger values allow faster learning but increase the risk of instability. The number of update epochs per rollout determines how much you squeeze out of each collected batch. More epochs improve sample efficiency (less environment interaction needed) but increase the risk that the policy drifts too far from the old policy within a single rollout. PPO's clipping limits this drift, but with many epochs and aggressive learning rates, the policy can still move substantially.

The key hyperparameters and their typical values are:

  • Clip parameter (ϵ\epsilon): Controls policy change by defining the trust region as [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon] for the probability ratio. Values of 0.1 to 0.3 are typical. Smaller values (e.g., 0.05) constrain updates too tightly and slow learning. Larger values (e.g., 0.5) provide insufficient constraint and risk instability.
  • GAE lambda (λ\lambda): Bias-variance tradeoff for advantage estimation. Values of 0.95 to 0.99 are typical.
  • Number of epochs: Number of passes through collected data. Using 3 to 10 epochs balances sample efficiency against overfitting to old data.
  • Minibatch size: Larger batches provide more stable gradients but may overfit. Sizes of 32 to 256 are common.
  • Rollout length: How much data to collect before updating. Longer rollouts provide better advantage estimates but slower iteration.
  • Learning rate: Values of 1e-4 to 3e-4 are typical for PPO. You can use learning rate scheduling to improve convergence.

Language model implementations often require adjusted values. The next chapter explores LLM-specific considerations, including how to adapt the reward signal from a reward model, manage the KL penalty against the reference model, and handle the much larger action spaces of language generation.

Out[17]:
Visualization
Line chart of the PPO clipped objective with a narrow shaded green gradient region between ratio 0.9 and 1.1, showing tight epsilon 0.1 constraints with red dotted vertical bound lines.
Tight trust region (epsilon = 0.1) constrains probability ratios to the narrow band [0.9, 1.1], limiting policy changes severely. The shaded green region shows where gradients flow, and you can see it is quite narrow. This conservative approach prevents destabilizing updates in unstable domains but may constrain learning excessively, which makes it difficult for the policy to improve significantly with each update. Best for environments where stability is necessary.
Line chart of the PPO clipped objective with a moderate shaded green gradient region between ratio 0.8 and 1.2, showing standard epsilon 0.2 constraints with red dotted vertical bound lines.
Standard trust region (epsilon = 0.2) constrains probability ratios to [0.8, 1.2], the most widely adopted setting. The moderate shaded gradient region balances stability against learning speed, giving adequate protection against instability while letting reasonable progress. This value works well across diverse continuous control tasks and RLHF applications.
Line chart of the PPO clipped objective with a wide shaded green gradient region between ratio 0.7 and 1.3, showing loose epsilon 0.3 constraints with red dotted vertical bound lines.
Loose trust region (epsilon = 0.3) permits probability ratios up to [0.7, 1.3], with a broader shaded gradient region. This permissive approach accelerates learning by letting larger policy changes per update, helpful in environments where the policy needs to change substantially. However, it provides less protection against instability and risks performance collapse if large updates move into poor regions. Use only when stability is less necessary.

Key Parameters

The key parameters for PPO are:

  • clip_epsilon: The clipping parameter that defines the trust region width (typically 0.1 to 0.3). Smaller values provide tighter constraints on policy updates, while larger values allow more aggressive changes.
  • gae_lambda: Controls bias-variance tradeoff in advantage estimation (typically 0.95 to 0.99). Higher values use longer horizons for advantage computation.
  • n_epochs: Number of optimization epochs per rollout (typically 3 to 10). More epochs extract more learning from each batch but risk overfitting to old data.
  • batch_size: Minibatch size for gradient updates (typically 32 to 256). Larger batches provide more stable gradients.
  • rollout_length: Number of timesteps to collect before updating (typically 2048 for simple tasks). Longer rollouts provide better advantage estimates.
  • value_coef: Coefficient for value function loss in total objective (typically 0.5). Controls how much weight to give value function training relative to policy training.
  • entropy_coef: Coefficient for entropy bonus (typically 0.01). Higher values encourage more exploration.

Limitations and Impact

PPO became a standard practical reinforcement learning method and later a foundation for aligning large language models with human preferences. It approximates trust region stability without the computational complexity of TRPO. Before PPO, reliable RL training required extensive expertise and environment-specific tuning. PPO lowered the implementation barrier for applying deep RL to new domains. Its explicit separation between rollout collection and policy optimization, simple clipping mechanism, and well-tested open-source implementations all contributed to its widespread adoption.

PPO has important limitations, however. The clipping mechanism provides only approximate trust region enforcement, so the policy can still drift significantly over many updates, especially with high learning rates or many optimization epochs. The clipping prevents the optimizer from exploiting large policy changes, but it does not prevent the policy from drifting gradually across many small steps. This drift becomes especially problematic in RLHF settings where maintaining proximity to the supervised fine-tuned base model is needed for response quality. A language model that drifts too far from its pre-trained weights can lose coherence, generate repetitive or degenerate text, or exploit weaknesses in the reward model without improving human-rated quality.

Sample efficiency is a persistent concern with PPO. The algorithm requires substantial environment or reward model interaction to learn effectively. Data becomes stale after a few update epochs, requiring fresh collection for continued improvement. For language models, each reward model query involves a full forward pass through a large neural network, making data collection expensive. This sample inefficiency is one of the primary motivations for direct alignment methods like Direct Preference Optimization (DPO), which we will explore in a later chapter. DPO sidesteps the RL optimization loop entirely by formulating alignment as a supervised learning problem over preference pairs, trading PPO's flexibility for much better data efficiency.

PPO inherits the challenges of the actor-critic framework. The value function must accurately estimate expected returns for meaningful advantage computation. In high-dimensional state spaces like natural language, where the state is a sequence of tokens that can be astronomically large, value estimation is extremely difficult. Noisy value estimates produce high-variance advantages, leading to unstable policy gradients even with the clipping mechanism in place. Researchers have addressed this in various ways: larger critic networks, separate training schedules for actor and critic, and more aggressive advantage normalization. None of these fully resolves the underlying difficulty of value estimation in language domains.

PPO also optimizes for the provided reward signal without accounting for reward model uncertainty. The reward model is itself a neural network trained on human preference data, and it will be wrong in many edge cases. PPO's objective is to maximize expected reward model score, not to maximize actual human preference. When the policy finds inputs that the reward model scores highly but humans would not, this is called reward hacking. The next chapter discusses how KL divergence penalties and reference model constraints mitigate this problem in RLHF: by penalizing the language model for deviating too far from a reference policy (typically the supervised fine-tuned base model), we constrain the optimization to regions where the reward model's predictions are likely to be reliable.

Finally, PPO's on-policy nature means that experience collected from an old policy becomes stale quickly. The multiple-epoch optimization within a single rollout helps, but eventually the probability ratios drift too far from 1 and clipping becomes binding everywhere. At that point, no more useful gradient signal remains in the batch, and a new rollout must be collected. This cycle of collect-update-discard is inherently less efficient than off-policy methods that maintain a replay buffer and can reuse experience indefinitely. Researchers have explored hybrid approaches, such as using importance sampling corrections to extend the usable lifetime of collected data, but these introduce their own complexity and instability.

Summary

PPO addresses vanilla policy gradient instability through a clipped surrogate objective that constrains policy changes in each update. Key insights include:

  • Trust regions matter: Constraining policy updates prevents severe performance degradation by so gradient estimates remain valid throughout the update.
  • Clipping approximates constraints: Rather than solving a constrained optimization problem, PPO clips the objective to remove incentives for excessive changes, achieving similar stability with far less computational overhead.
  • Pessimistic bounds ensure stability: The min operation between clipped and unclipped objectives prevents overestimating improvement, always choosing the more conservative estimate.
  • Multiple epochs improve efficiency: Multiple epochs of updates reuse collected data, improving sample efficiency while the clipping mechanism prevents the policy from drifting too far.
  • GAE balances bias and variance: The lambda parameter allows continuous interpolation between one-step TD estimates (low variance, high bias) and Monte Carlo returns (high variance, low bias), accommodating different domains and value function quality levels.

The full PPO objective combines the clipped policy loss with a value function loss for training the critic and an entropy bonus for exploration. These three terms work synergistically: better value estimates produce better advantages, which produce better policy updates, which explore more effectively and generate better training data.

PPO became the standard algorithm for RLHF in language models because it trains stably with a comparatively simple implementation and uses samples efficiently. The next chapter explores adapting PPO for language model alignment by covering reference model constraints and generation-specific considerations. You will see how the framework we built here, collecting rollouts from the language model, scoring them with a reward model, computing advantages, and updating with the clipped objective, maps directly onto the full RLHF pipeline used in production alignment systems.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Proximal Policy Optimization.

PPO Algorithm Quiz

Question 1 of 80 of 8 completed
What does a probability ratio rt(θ)=2r_t(\theta)=2 indicate about the new policy compared to the old policy?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025ppoalgorithm, author = {Michael Brenndoerfer}, title = {PPO Algorithm: Proximal Policy Optimization for Stable RL}, year = {2025}, url = {https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). PPO Algorithm: Proximal Policy Optimization for Stable RL. Retrieved from https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning
MLAAcademic
Michael Brenndoerfer. "PPO Algorithm: Proximal Policy Optimization for Stable RL." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning>.
CHICAGOAcademic
Michael Brenndoerfer. "PPO Algorithm: Proximal Policy Optimization for Stable RL." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning.
HARVARDAcademic
Michael Brenndoerfer (2025) 'PPO Algorithm: Proximal Policy Optimization for Stable RL'. Available at: https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). PPO Algorithm: Proximal Policy Optimization for Stable RL. https://mbrenndoerfer.com/writing/ppo-algorithm-proximal-policy-optimization-reinforcement-learning

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.