Part of Language AI Handbook
Covers PPO's clipped objective for stable policy updates. Topics include trust regions, GAE advantage estimation.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
PPO Algorithm
In the previous chapter, we explored policy gradient methods and saw how REINFORCE directly optimizes a policy by following the gradient of expected reward. While mathematically elegant, vanilla policy gradients are notoriously unstable during training. A single large gradient update can catastrophically degrade policy performance, and recovery is difficult. The algorithm might spend hundreds of update steps undoing damage from one reckless step. This instability prevented policy gradient methods from scaling to complex domains like language model fine-tuning, where each update step is expensive and evaluation is slow.
Proximal Policy Optimization (PPO), introduced by Schulman et al. in 2017, addresses this instability through a deceptively simple mechanism: it clips the objective function to prevent updates that change the policy too drastically. Rather than solving a complex constrained optimization problem, PPO modifies the reward signal itself so the optimizer cannot benefit from stepping too far from the current policy. The result is an algorithm that is simultaneously simpler to implement, more computationally efficient, and more stable than its predecessors.
Think of PPO as a careful hiker who always stays within a reliable map region. The hiker can see the general direction of higher elevation and wants to move toward it. But the hiker also knows that map accuracy degrades quickly outside the territory already explored. So the hiker commits to moving only a short distance with each step, updating the map from the new vantage point, and then deciding where to move next. The hiking speed is not zero, which would be no progress at all, but it is bounded, which prevents walking off a cliff based on a faulty extrapolation. PPO implements exactly this constraint: bound each update step so you remain in territory where your gradient estimates are trustworthy.
Understanding PPO is needed for anyone working with reinforcement learning from human feedback (RLHF) in language models. Modern alignment pipelines, including those behind ChatGPT, Claude, and Gemini, rely heavily on PPO or close variants to translate human preference data into updated model behavior. The algorithm's stability and implementation simplicity made it the natural choice when researchers needed to apply RL to language models with billions of parameters, where unstable training would be extraordinarily expensive to debug or recover from.
This chapter covers PPO in depth from the ground up. We start with the specific failure modes of unconstrained policy gradient optimization, move through the trust region methods that motivated PPO's design, then derive the clipped objective, work through Generalized Advantage Estimation, assemble the complete multi-component objective, and finally implement a working PPO agent in PyTorch. We include a numerical worked example so you can trace every computation by hand before running code. By the end, you will understand what PPO does, why each design choice exists, and what goes wrong when you deviate from it.
PPO was introduced by John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov at OpenAI in their 2017 paper "Proximal Policy Optimization Algorithms." The paper presented PPO as a simplification of Trust Region Policy Optimization (TRPO), their own earlier algorithm from 2015. TRPO had proven that stable policy improvement was possible, but its implementation required second-order optimization, Fisher information matrices, and conjugate gradient solvers. This makes it impractical at scale. PPO replaced all of this machinery with a simple clipped objective, achieving similar stability with a first-order optimizer like Adam. Within a year of publication, PPO became the dominant algorithm for deep reinforcement learning at OpenAI and throughout the research community. When OpenAI published InstructGPT in 2022. This shows that language models could be aligned with human preferences using RL, PPO was the algorithm that made it work. The paper's influence on the alignment field cannot be overstated: virtually every publicly discussed RLHF pipeline for language model alignment traces its RL component back to this 2017 contribution.
The Problem with Vanilla Policy Gradients
The instability in vanilla policy gradient methods is not a superficial implementation problem that can be fixed with a better learning rate schedule. It is a structural problem rooted in how the gradient signal relates to policy performance. To understand why PPO's solution is the right one, we need to be precise about exactly what goes wrong.
Recall that the policy gradient takes the form:
To understand why this formula creates practical difficulties, examine what each component contributes:
- : the gradient of the expected cumulative reward with respect to policy parameters, showing the direction to adjust parameters to increase expected reward
- : expectation over trajectories sampled from the current policy
- : parameters of the policy network that we are optimizing
- : the expected cumulative reward under policy
- : a trajectory (sequence of states and actions) sampled from the current policy
- : the probability of taking action in state under the current policy
- : the time horizon, the length of the episode
- : the advantage function at time , estimating how much better action is compared to the average action in state
- : the gradient of the log probability with respect to policy parameters, showing the direction to adjust parameters to make action more likely
- : sum over all timesteps in the trajectory from 0 to T
The basic issue with this formulation is that the gradient provides no guidance about step size. The policy gradient theorem tells us which direction to move parameters to increase expected reward, but it remains entirely silent about how far we should move in that direction. This is analogous to knowing that walking north takes you closer to your destination but not knowing whether to take one step or one hundred steps. A step in the gradient direction improves performance locally, within an infinitesimally small neighborhood around the current parameters. However, nothing in the mathematics prevents taking such a large step that you overshoot into a region where the policy performs terribly. Policy performance is not smooth across the parameter space: a policy that seems promising can suddenly fail after a large update.
This problem is especially acute because of three interconnected challenges:
- Non-stationary data distribution: The policy generates its own training data. When the policy changes significantly, the state distribution also changes, potentially invalidating previously learned value estimates.
- High variance: Policy gradient estimates are inherently noisy, which makes it difficult to distinguish signal from noise.
- Irreversibility: A bad update might move the policy to a region where it never encounters states that would help it recover.
Consider a language model learning from human feedback. If a single gradient update makes the model much more likely to generate certain patterns, the model might suddenly produce outputs that are completely off-distribution from its training, leading to reward model extrapolation errors and further degradation. The optimizer has no way to know it has moved too far until the damage is already done, and by that point, the policy may be in such a degraded state that subsequent updates cannot recover it without discarding all progress.


Trust Region Methods
The insight behind trust region methods is to constrain how much the policy can change in each update. Rather than blindly following the gradient wherever it leads, we optimize the policy subject to a constraint that keeps the new policy "close" to the old one. This approach recognizes a basic tension in optimization: we want to improve the policy as quickly as possible. However, our confidence in the improvement direction decreases the further we move from where we collected our data. Gradients computed from sampled trajectories accurately describe performance improvement near the current policy, but extrapolating far from observed data becomes unreliable.
Think of a trust region as the area on a topographic map that your GPS unit has surveyed at high resolution. Within that region, the contour lines are accurate and you can safely navigate. Outside that region, the map was generated from coarser satellite data and might have cliffs, valleys, or impassable terrain that looks navigable on the map. You would not plan a route that takes you far outside your high-resolution zone. Similarly, the policy gradient is an accurate local guide only within the neighborhood of states and actions you have recently observed. Venturing far outside that neighborhood means relying on gradient extrapolations that may be wildly wrong.
A trust region is a neighborhood around the current parameters within which a local approximation of the objective function is trusted to be accurate. Optimization proceeds by maximizing this approximation within the trust region, then updating the region based on how well the approximation matched reality. If the approximation was good, the trust region expands. If it was poor, the region shrinks. This adaptive mechanism maintains alignment between the model and reality throughout training.
The key insight behind all trust region methods is that gradient estimates are only reliable near the data they were computed from. Moving the policy far from the current policy means the state-action distribution shifts, and the advantages computed from old data no longer accurately reflect the new policy's performance. The trust region constraint formalizes the boundary beyond which you should not venture without collecting new data. Enforcing this boundary ensures that the objective function improvement you see during optimization corresponds to actual performance improvement in the environment.

Trust Region Policy Optimization (TRPO), PPO's predecessor, formalizes this using KL divergence as a constraint. The KL divergence measures how different two probability distributions are, which makes it a natural choice for measuring policy similarity. Policies that assign similar probabilities to actions have low KL divergence, while policies that behave very differently have high KL divergence. This makes KL divergence a precise, mathematically grounded way to define "how much the policy has changed."
The TRPO objective addresses the policy update problem by maximizing expected advantage while explicitly constraining how much the policy can change. It maximizes the expected advantage weighted by the importance sampling ratio, which allows us to use data from the old policy to evaluate the new policy, while constraining the KL divergence between old and new policies:
To understand how this constrained optimization problem balances improvement against stability, examine each component:
- : parameters of the new policy being optimized
- : parameters of the old policy from which data was collected
- : expectation over states and actions sampled from trajectories collected using the old policy
- : a state sampled from the state distribution under the old policy
- : an action sampled from the old policy in state
- : probability of action in state under the new policy
- : probability of action in state under the old policy
- : the importance sampling ratio, which reweights data from the old policy to evaluate the new policy
- : advantage function computed under the old policy, measuring how much better action is than average in state
- : expectation over states from the state distribution under the old policy
- : Kullback-Leibler divergence measuring how much the new policy distribution differs from the old policy distribution
- : maximum allowed KL divergence, the trust region radius
This constraint ensures the new policy does not diverge too far from the old policy in terms of the probability distributions over actions. The ratio is called the importance sampling ratio and allows us to evaluate the new policy using data collected from the old policy. This ratio acts as a correction factor: if the new policy is twice as likely to take an action as the old policy, we weight that action's contribution twice as heavily to account for the fact that it would occur more frequently under the new policy.
Why TRPO Works But Is Complex
TRPO guarantees monotonic improvement under certain conditions, meaning the policy never gets worse than the previous version. This guarantee provides the stability that vanilla policy gradients lack. However, the mathematical machinery required to enforce this guarantee comes with significant computational costs: computing Fisher information matrices and solving linear systems.
Solving the constrained optimization problem requires computing second-order derivatives, specifically the Fisher information matrix, and performing conjugate gradient optimization to solve a system of linear equations. This makes TRPO computationally expensive and difficult to implement correctly. The Fisher information matrix has size where is the number of policy parameters, so even storing it is impractical for large neural networks. Computing the matrix-vector product required by conjugate gradient solvers requires special tricks involving second-order automatic differentiation. Computing the Fisher information matrix across multiple workers in distributed settings adds significant overhead.
PPO achieves similar stability guarantees with a first-order method by replacing the hard constraint with a penalty built directly into the objective function. Instead of solving a constrained optimization problem, PPO modifies the objective itself to discourage excessive policy changes. This transformation from constraint to penalty makes PPO dramatically simpler to implement while preserving the needed benefits of trust region methods. The resulting algorithm fits cleanly into any standard deep learning training loop, with no exotic optimization machinery required.
The Probability Ratio
The probability ratio is central to PPO. It quantifies how much the new policy's probability for an action differs from the old policy's probability. By examining this ratio across all observed state-action pairs, you can assess whether the policy is changing appropriately or excessively. Think of the probability ratio as a sensitivity meter: a ratio of 1 means the policy is unchanged for this action, values greater than 1 mean the policy is becoming more likely to take this action, and values less than 1 mean the policy is becoming less likely. The ratio captures the entire story of how the policy is changing, action by action, state by state.
We define the ratio as:
Understanding what each symbol represents helps illuminate why this ratio captures policy change so effectively:
- : the importance sampling ratio at timestep
- : probability of action in state under the new policy with parameters
- : probability of the same action in the same state under the old policy
- : the action taken at timestep
- : the state at timestep
- : parameters of the new policy
- : parameters of the old policy
This ratio captures how much more or less likely action becomes under the new policy compared to the old one. The interpretation is intuitive and direct:
- : the new policy assigns the same probability to this action; no change from the old policy
- : the new policy makes this action more likely (for example, means the new policy is twice as likely to take this action)
- : the new policy makes this action less likely (for example, means the new policy is half as likely to take this action)
This ratio directly measures behavioral change. A ratio of 1 everywhere indicates identical policies, while very large or very small ratios indicate dramatic behavioral shifts. This makes the ratio ideal for clipping to constrain policy changes. Importantly, the ratio is also the right quantity for importance sampling: when we use old data to evaluate a new policy, we multiply by the ratio to correct for the fact that we collected data under a different distribution. This dual role as both a measure of policy change and an importance weight makes the ratio the natural pivot point for the PPO objective.


The standard policy gradient objective can be rewritten using this ratio, revealing its basic role in policy optimization:
To understand how the ratio enables policy improvement, examine what each term contributes:
- : the policy gradient objective function we want to maximize
- : expectation over timesteps in collected trajectories
- : the probability ratio at timestep , equal to
- : the advantage estimate at timestep , measuring how much better the action taken was compared to the average action
- : parameters of the policy being optimized
When , we have everywhere, and the objective reduces to the simple advantage-weighted sum. The gradient of this objective at equals the standard policy gradient, confirming that this formulation is equivalent to what we derived earlier.
The problem becomes clear when we consider what happens as we optimize this objective. As the optimizer works to maximize the expected reward, the ratio can become arbitrarily large or small. If an action had positive advantage, showing it was better than expected, unconstrained optimization would keep increasing its probability without bound. The gradient always points toward making good actions more likely and bad actions less likely, but nothing in this formulation limits how far the policy can shift. This is precisely the instability problem that TRPO addressed with its KL constraint, and that PPO addresses with clipping.
The Clipped Objective
PPO's key innovation is straightforward: clip the probability ratio to remove incentives for excessive policy changes. Rather than imposing a hard constraint requiring complex optimization machinery, PPO modifies the objective function itself, so large policy changes provide no additional benefit. The clipping mechanism is self-limiting: once the policy has changed enough to fall outside the allowed range, the gradient signal turns off, and the optimizer cannot push the policy further in that direction.
Think of the clipped objective as a reward system with diminishing returns for policy changes. The policy earns credit for adjusting action probabilities in the direction indicated by the advantage, but only up to a point. Beyond the clip threshold, additional changes earn no extra credit. This is like an employer who rewards employees for working overtime, but only up to a cap: you can work as many hours as you want, but you only get paid extra for the first few hours of overtime. The cap prevents extreme behavior without eliminating the incentive entirely.
The clipped surrogate objective is:
where is a hyperparameter, typically set between 0.1 and 0.2, that defines the trust region width. This single parameter controls how much the policy can change in each update. Larger allows more aggressive updates. Smaller enforces tighter constraints and slower learning.
To understand precisely how clipping constrains the ratio, we need to examine the clip function itself. The clip function constrains the probability ratio to remain within the interval , defined as:
Each component of this piecewise function serves a specific purpose:
- : the input value to be clipped; in PPO, this is , the probability ratio
- : the lower bound of the allowed range
- : the upper bound of the allowed range
- The function returns unchanged if it's within the bounds, otherwise returns the nearest bound
The key intuition is that this clipping removes the gradient signal when the ratio moves outside the trust region. Consider what happens during optimization: you want the optimizer to adjust parameters to increase the objective. If the new policy is already making an action much more likely, with a ratio greater than , or much less likely, with a ratio less than , than the old policy, clipping prevents the optimizer from pushing it even further in that direction. The clipped term becomes constant with respect to the parameters, meaning its gradient is zero, so there is no signal encouraging further movement in that direction.
This clipped ratio is then used in computing the clipped surrogate objective. Returning to the full formula and examining it in detail:
Each component plays a specific role in creating stable policy updates:
- : the clipped surrogate objective that PPO maximizes
- : expectation over all timesteps in the collected batch
- : takes the smaller of the two arguments (the pessimistic bound)
- : the probability ratio
- : the advantage estimate at timestep
- : the clipping parameter, typically 0.1 to 0.2, that defines the trust region
- : constrains the ratio to the interval
Understanding the Clipping Mechanism
The operation is important. It selects the smaller of the unclipped and clipped objectives. This keeps a pessimistic (conservative) bound on the improvement. By always choosing the lower value, we prevent the optimizer from being overly optimistic about improvements that would require large policy changes. We will analyze both cases based on the sign of the advantage.
Case 1: Positive Advantage ()
When an action is better than expected (positive advantage), we want to increase its probability. The unclipped objective's gradient pushes the policy to make this action more likely, but PPO limits the extent.
- If : The clipped term is smaller than , since and . The selects the clipped term , which is constant with respect to , so the gradient is zero and provides no incentive for us to increase the ratio further.
- If : Both terms are equal or the unclipped term is smaller, so normal optimization proceeds.
This means that once the probability ratio exceeds , the optimizer receives no additional reward for increasing it further. The policy has already been sufficiently encouraged to take this good action. Importantly, the policy can still take that action at a higher probability than the old policy, it just does not receive additional credit for pushing the probability even higher.
Case 2: Negative Advantage ()
When an action is worse than expected (negative advantage), we want to decrease its probability. The unclipped objective's gradient encourages this, but PPO limits the extent:
- If : Since , multiplying by smaller values makes the product more negative. The clipped term is larger (less negative) than the unclipped term because . The selects the more negative unclipped term. Since the clipped term is constant, its gradient is zero and the objective becomes flat once drops below , preventing the policy from decreasing the probability further.
- If : Normal optimization proceeds.
This prevents the policy from becoming too averse to actions that happened to have negative advantage. Such actions might still be valuable in other states, and excessively penalizing them could harm overall performance. The clipping protects both ends: it prevents overconfidence about good actions and overpunishment of bad ones.
The following figure illustrates this behavior:


Clipping creates a "pessimistic bound" on the objective. When the policy tries to change too much, the objective flattens and provides no gradient signal to continue. This self-limiting behavior makes PPO stable without requiring the complex second-order optimization of TRPO. The operation is the mechanism: it always selects whichever value is more conservative, preventing the optimizer from exploiting large policy changes that might not generalize.
Generalized Advantage Estimation
PPO typically uses Generalized Advantage Estimation (GAE) to compute advantages with a controllable bias-variance tradeoff. The advantage function measures how much better an action is than average. True advantages depend on full trajectory information unavailable during learning, so we must estimate them from observed rewards. Different estimation approaches trade off bias against variance in different ways, and this tradeoff has a major impact on the stability and efficiency of training.
Think of advantage estimation as a forecasting problem. The one-step estimator is like a weather forecast that only looks at today's conditions: it is highly confident (low variance) but might be systematically wrong if today's conditions are atypical (high bias). The Monte Carlo estimator is like collecting many days of weather data before making a prediction: it captures complex patterns (low bias) but requires lots of data and is noisy (high variance). GAE is like a weighted ensemble of forecasts at different time horizons, combining the precision of short-term estimates with the accuracy of long-term ones.
A one-step estimate uses only the immediate reward and next state value. This provides low variance, since it depends on fewer random variables, but high bias, since it relies heavily on the accuracy of the value function. A Monte Carlo estimate uses all future rewards until episode end. This provides low bias, since it uses actual observed returns, but high variance, since it incorporates the randomness of many future actions and transitions. The quality of one-step estimates depends entirely on how good the value function is. If the value function is poor, one-step TD errors will be systematically wrong, and the policy will learn from corrupted signals.
GAE solves this estimation problem by taking an exponentially weighted average of temporal difference errors at different time horizons. The parameter controls how quickly the weights decay as we look further into the future. This approach combines the benefits of using both short-term (low variance but high bias) and long-term (high variance but low bias) estimates. By tuning , we can find the sweet spot for our particular problem.
GAE is defined as:
Each component of this formula contributes to the bias-variance tradeoff:
- : the GAE advantage estimate at timestep
- : sum over all future timesteps from the current timestep forward (in practice, truncated at episode end)
- : the lookahead index, showing how many steps into the future we're considering
- : the discount factor, typically 0.99, which determines how much we value future rewards
- : the GAE parameter, typically 0.95, which controls the bias-variance tradeoff
- : the exponentially decaying weight for the -step temporal difference error
- : the temporal difference (TD) error at timestep
- : the current timestep
The temporal difference error, which is the building block for GAE, measures the discrepancy between what we expected and what we observed. It is defined as:
Understanding each term clarifies why TD errors are useful for advantage estimation:
- : the temporal difference error at timestep , measuring the difference between the observed reward plus next state value versus the current state value
- : the immediate reward received at timestep
- : the discount factor
- : the value function estimate for state (the predicted cumulative future reward)
- : the value function estimate for the next state
- : the discounted value of the next state
- : the state at timestep
- : the next state at timestep
The TD error intuition is clear: if our value function were perfect, the expected TD error would be zero. The value of the current state should equal the immediate reward plus the discounted value of the next state. A positive TD error indicates we received more reward than expected, suggesting the action was good. A negative TD error indicates we received less than expected. By accumulating these errors over multiple timesteps with exponentially decaying weights, GAE builds a richer signal that accounts for delayed consequences while limiting how much noise accumulates.


The tradeoff controlled by determines how we combine information across time horizons:
- uses only one-step TD error (low variance, high bias)
- uses full Monte Carlo returns (high variance, low bias)
In practice, provides a good balance for most problems, weighting nearby TD errors heavily while still incorporating longer-term information with diminishing weight. For language models, this setting helps because the value function approximation is imperfect given the enormous state space, and GAE mitigates the impact of value function errors on advantage quality.
The recursive formulation for efficient computation eliminates the need to store and sum all future TD errors explicitly:
Each component of this recursive formula has a clear interpretation:
- : the advantage estimate at timestep , computed recursively
- : the temporal difference error at timestep , equal to
- : discount factor
- : GAE parameter
- : the advantage estimate for the next timestep, computed first in backward iteration
- : the current timestep
The boundary condition is at the terminal timestep, since there are no future advantages after the episode ends. This recursive formulation is computationally efficient because we can compute all advantages in a single backward pass through the trajectory, starting from the end and working toward the beginning. Each computation reuses the result from the next timestep, avoiding redundant calculations. The algorithm is simple: iterate backwards, accumulating a running GAE value with decay, and store the result at each timestep.
Worked Example: Tracing a Single PPO Update
Before implementing PPO in code, let's trace through one complete update step numerically. This example uses tiny numbers to keep the arithmetic tractable, but the procedure is identical for million-parameter networks.
Suppose we have collected a short trajectory with three timesteps. The states, actions, rewards, and value estimates are:
| Timestep | State | Action | Reward | ||
|---|---|---|---|---|---|
| 0 | 1.0 | 0.8 | -0.693 | ||
| 1 | 0.0 | 0.5 | -1.386 | ||
| 2 | 2.0 | 1.2 | -0.405 |
The episode terminates after timestep 2, so . We use and .
Step 1: Compute TD errors. Working forward through the trajectory:
Step 2: Compute GAE advantages. Working backwards through the trajectory, initializing :
The advantage at is largest because the trajectory begins with a low value prediction (0.8) but transitions to states with higher-than-expected cumulative reward.
Step 3: Normalize advantages. Standard PPO normalizes advantages to have zero mean and unit variance. The mean is . The variance is approximately , giving standard deviation . Normalizing:
Step 4: Compute probability ratios. After one gradient step, suppose the new policy has log-probabilities: , , .
Step 5: Apply clipping with . The clip bounds are .
For : , (positive advantage). The unclipped term is . The clipped ratio is , so the clipped term is also . Both terms equal, so . Note that exactly: the ratio is right at the boundary.
For : , (small positive). Similarly .
For : , (negative advantage). The unclipped term is . The clipped ratio clips 0.75 to 0.8, giving clipped term . The (more negative wins). So .
Here, , so clipping activates and the clipped objective is flatter (less negative) than the unclipped one. This means the optimizer does not receive a gradient signal pushing below 0.8 further, preventing excessive suppression of action .
Step 6: Average the objective. .
This positive value indicates the gradient step improved the objective. On subsequent passes through the same data (multiple epochs), the ratios will update further, and clipping will activate when the policy has changed enough.
The Complete PPO Objective
The full PPO objective combines three terms: policy improvement, value function training, and exploration. These components work together synergistically. Accurate value estimates enable meaningful advantages, while exploration discovers strategies that improve both the policy and value function. Training all three simultaneously with a shared network backbone allows the network to develop representations that serve all three purposes at once, improving sample efficiency compared to training separate networks.
The complete objective is:
Each component serves a distinct purpose in training a capable agent:
- : the complete PPO objective function, to be maximized
- : expectation over timesteps in the batch
- : the clipped surrogate objective for policy improvement, defined earlier
- : coefficient for the value function loss, typically 0.5
- : the value function loss, which encourages accurate value estimates
- : coefficient for the entropy bonus, typically 0.01
- : the entropy of the policy distribution at state , which encourages exploration
- : parameters of both the policy and value networks, when they share parameters
- : the state at timestep
The value function loss has a negative sign because we minimize loss while maximizing the overall objective. This sign convention means that minimizing value function error increases the overall objective, aligning all components toward the same optimization direction.
The three terms work together synergistically. The clipped surrogate loss improves the policy by increasing probabilities of high-advantage actions while respecting the trust region constraint. The value function loss trains the value function to make accurate predictions, which are needed for computing meaningful advantages in future updates. The entropy term encourages exploration by penalizing overly deterministic policies, preventing premature convergence and helping the agent discover better strategies.
Value Function Loss
The value function predicts expected cumulative returns and is trained alongside the policy. Accurate value estimates are needed because the value function is the baseline in advantage computation. Without accuracy, advantages become noisy and unreliable, leading to high-variance policy gradients that destabilize training. Think of the value function as a sports commentator who tells you whether a play was good or bad: if the commentator has no idea what a typical play looks like, their assessments of "better than average" or "worse than average" are meaningless. The policy can only learn from advantage estimates if those estimates are grounded in an accurate model of expected future rewards.
The value function predicts expected cumulative future rewards starting from each state under the current policy. We train the value function by minimizing the mean squared error between its predictions and target values:
Understanding each component reveals the standard regression structure of value function training:
- : the value function loss, the mean squared error
- : expectation over timesteps in the batch
- : parameters of the value function network
- : the value function's prediction for state under current parameters
- : the target value we want the value function to predict
- : the state at timestep
- : the squared difference between prediction and target
The target value is computed as:
Each term in this target construction serves a specific purpose:
- : the target value for the value function to predict at timestep
- : the advantage estimate at timestep , computed using GAE
- : the value prediction from the old value function, before this update
- : parameters of the value function from the previous update
- : the state at timestep
This target construction guides the value function to predict returns accurately using computed advantages. The advantage represents how much better actual returns were than expected, or equivalently, the residual between actual returns and predicted returns. Adding this residual to the old value estimate yields an improved return estimate.
This bootstrapping approach allows the value function to learn from its own predictions while incorporating new information from observed rewards. The process is iterative: better value estimates lead to better advantage estimates, which lead to better policy updates, which generate better training data for the value function. Some implementations also clip the value function loss similarly to the policy loss, preventing large changes to the value function that might destabilize this iterative process. This value clipping is particularly useful in RLHF settings where the value function must adapt to reward model outputs that can change significantly during alignment training.
Entropy Bonus
Entropy measures randomness in the policy's action distribution. High entropy spreads probability across multiple actions instead of putting on one choice. Encouraging higher entropy prevents the policy from becoming deterministic too early, which keeps exploration alive and enables discovery of better strategies. Without the entropy bonus, PPO tends to quickly commit to whatever actions seem best in the current training data, potentially missing superior alternatives that have not yet been explored. The entropy bonus is a small but important regularizer that maintains the policy's flexibility throughout training.
For a policy distribution at state , entropy is defined as:
Each component of this formula connects to information-theoretic concepts:
- : the entropy of the policy distribution at state
- : sum over all possible actions in the action space
- : an action from the action space
- : the probability of action in state under the current policy
- : the log probability of action
- : parameters of the policy network
- : the state at timestep
This entropy term encourages exploration by rewarding policies that maintain uncertainty over actions. The formula measures the expected information content, or surprise, in the policy's action distribution. Mathematically, when action probabilities are uniform, meaning all actions are equally likely, the sum of terms is maximized in magnitude, giving high entropy. When the policy is deterministic, one action has probability 1 and all others have probability 0, so the entropy collapses to zero.


When all actions have similar probabilities, showing high uncertainty, entropy is maximized. This reflects the fact that an observer would be maximally uncertain about which action the policy will choose. When the policy becomes deterministic, meaning one action has probability near 1 and all others near 0, entropy approaches zero. The policy reveals little information because its behavior is predictable. While deterministic policies can be desirable in final deployment, during training we need exploration to discover better strategies.
The coefficient (typically 0.01) controls the exploration-exploitation tradeoff. Higher values encourage exploration but may slow convergence to optimal behavior. Lower values enable faster convergence but risk getting stuck in suboptimal local minima. In language model fine-tuning, the entropy coefficient is often set quite small because the base model already has a well-developed prior over language, and aggressive exploration would destroy linguistic coherence. The goal is gentle nudging toward human-preferred outputs, not wholesale behavioral reshaping.
Putting It Together
The typical coefficient values that Schulman et al. used in the original paper are:
- : weight for value function loss
- : weight for entropy bonus
- : clipping parameter
Sharing parameters between policy and value networks, which is common in practice, combines all three losses for joint optimization. This parameter sharing encourages the network to learn representations useful for both predicting values and selecting actions. It often improves sample efficiency compared to separate networks. The shared layers act as a general-purpose feature extractor: features that help predict which actions lead to good outcomes (value function) are often the same features that should drive action selection (policy). The separate output heads then specialize these shared representations for their specific task.
PPO Implementation
We will implement PPO for continuous control using a simple environment. This lets you focus on the algorithm rather than domain-specific complexity. The environment we use, Pendulum-v1 from OpenAI Gymnasium, requires a continuous torque input to balance an inverted pendulum. It is simple enough to train in minutes but complex enough to require real policy learning.
import warnings
warnings.filterwarnings("ignore")Actor-Critic Network
The actor-critic architecture is central to PPO. The actor is the policy: it takes a state and outputs a distribution over actions. The critic is the value function: it takes a state and outputs an estimate of the expected cumulative reward from that state. Sharing the lower layers between actor and critic allows both to benefit from the same learned representation of the environment, reducing the total number of parameters and speeding up learning.
import torch
import torch.nn as nn
from torch.distributions import Normal
class ActorCritic(nn.Module):
def __init__(self, obs_dim, action_dim):
super().__init__()
# Shared layers learn representations useful for both policy and value prediction
self.shared = nn.Sequential(
nn.Linear(obs_dim, 64), nn.Tanh(), nn.Linear(64, 64), nn.Tanh()
)
# Policy head (actor) - outputs mean of action distribution
self.policy_mean = nn.Linear(64, action_dim)
# Learnable log standard deviation
self.policy_log_std = nn.Parameter(torch.zeros(action_dim))
# Value head (critic)
self.value_head = nn.Linear(64, 1)
def forward(self, obs):
features = self.shared(obs)
# Policy outputs Gaussian distribution parameters
action_mean = self.policy_mean(features)
action_std = self.policy_log_std.exp()
# Value estimate
value = self.value_head(features)
return action_mean, action_std, value
def get_action(self, obs, deterministic=False):
action_mean, action_std, value = self.forward(obs)
if deterministic:
return action_mean, value
# Sample from Gaussian
dist = Normal(action_mean, action_std)
action = dist.sample()
log_prob = dist.log_prob(action).sum(-1)
return action, log_prob, value
def evaluate_actions(self, obs, actions):
action_mean, action_std, value = self.forward(obs)
dist = Normal(action_mean, action_std)
log_probs = dist.log_prob(actions).sum(-1)
entropy = dist.entropy().sum(-1)
return log_probs, value.squeeze(-1), entropyThis architecture implements the actor-critic pattern for continuous control. The shared feature extractor, two 64-unit hidden layers with tanh activations, learns representations useful for both policy and value prediction. The actor head outputs Gaussian parameters (mean and learnable log standard deviation), letting sampling of continuous actions with exploration noise. The standard deviation starts at 1.0 (since the log standard deviation is initialized to 0) and is learned independently of the mean, letting the policy to adjust its exploration width as training progresses. The critic head predicts state values for advantage computation. The get_action method samples actions during rollout collection. The evaluate_actions method computes log probabilities and entropy during policy updates, which are used to compute the probability ratios and entropy bonus in the loss function.
Experience Buffer
PPO requires collecting a batch of experience before updating the policy. The experience buffer stores all transitions from the current rollout, then computes advantages in a single backward pass before minibatch optimization begins. This separation of collection and optimization is basic to PPO's design: you collect data under the old policy, compute advantages, then optimize for multiple epochs, discarding the data when it becomes too stale.
import numpy as np
import torch
class RolloutBuffer:
def __init__(self):
self.observations = []
self.actions = []
self.rewards = []
self.values = []
self.log_probs = []
self.dones = []
def add(self, obs, action, reward, value, log_prob, done):
self.observations.append(obs)
self.actions.append(action)
self.rewards.append(reward)
self.values.append(value)
self.log_probs.append(log_prob)
self.dones.append(done)
def compute_returns_and_advantages(
self, last_value, gamma=0.99, gae_lambda=0.95
):
"""Compute GAE advantages and returns using temporal difference errors."""
advantages = []
returns = []
gae = 0
values = self.values + [last_value]
# Iterate backwards through the buffer
for t in reversed(range(len(self.rewards))):
if self.dones[t]:
delta = self.rewards[t] - values[t]
gae = delta
else:
delta = self.rewards[t] + gamma * values[t + 1] - values[t]
gae = delta + gamma * gae_lambda * gae
advantages.insert(0, gae)
returns.insert(0, gae + values[t])
self.advantages = advantages
self.returns = returns
def get_batches(self, batch_size):
"""Generate random minibatches for optimization."""
n_samples = len(self.observations)
indices = np.random.permutation(n_samples)
for start in range(0, n_samples, batch_size):
end = start + batch_size
batch_indices = indices[start:end]
yield (
torch.stack([self.observations[i] for i in batch_indices]),
torch.stack([self.actions[i] for i in batch_indices]),
torch.tensor(
[self.log_probs[i] for i in batch_indices],
dtype=torch.float32,
),
torch.tensor(
[self.advantages[i] for i in batch_indices],
dtype=torch.float32,
),
torch.tensor(
[self.returns[i] for i in batch_indices],
dtype=torch.float32,
),
)
def clear(self):
"""Clear buffer for next rollout collection."""
self.__init__()The RolloutBuffer manages experience collection and advantage computation. The add method stores each transition (state, action, reward, value estimate, log probability, done flag) during rollout. The compute_returns_and_advantages method implements GAE by iterating backwards through the trajectory, computing temporal difference errors and accumulating them with exponential decay controlled by gamma and lambda. Note how episode boundaries are handled: when dones[t] is true, the bootstrapped next-state value is zero (the episode has ended), so the delta uses only the reward minus the current value. This backward pass efficiently computes advantages for all timesteps in a single sweep. The get_batches method shuffles the data and yields minibatches for multiple optimization epochs, improving sample efficiency by making multiple gradient updates per rollout.
PPO Update Step
The core PPO update is where all the mathematics we have discussed comes together. In each epoch, we iterate through the collected data in randomized minibatches, compute the three components of the PPO objective, combine them with their respective coefficients, and perform a gradient step. The key detail is computing ratios in log space for numerical stability: instead of computing directly (which can involve extremely small or large numbers), we compute , which avoids floating point underflow or overflow.
import numpy as np
import torch
import torch.nn as nn
def ppo_update(
model,
optimizer,
buffer,
clip_epsilon=0.2,
value_coef=0.5,
entropy_coef=0.01,
n_epochs=10,
batch_size=64,
):
"""Execute PPO updates over multiple epochs of minibatches from collected experience."""
policy_losses = []
value_losses = []
entropy_losses = []
for epoch in range(n_epochs):
for batch in buffer.get_batches(batch_size):
obs, actions, old_log_probs, advantages, returns = batch
# Normalize advantages for improved training stability
advantages = (advantages - advantages.mean()) / (
advantages.std() + 1e-8
)
# Get current policy evaluation
log_probs, values, entropy = model.evaluate_actions(obs, actions)
# Compute probability ratio
ratio = torch.exp(log_probs - old_log_probs)
# Clipped surrogate objective
unclipped = ratio * advantages
clipped = (
torch.clamp(ratio, 1 - clip_epsilon, 1 + clip_epsilon)
* advantages
)
policy_loss = -torch.min(unclipped, clipped).mean()
# Value function loss
value_loss = nn.functional.mse_loss(values, returns)
# Entropy bonus (negative because we minimize total loss)
entropy_loss = -entropy.mean()
# Combined loss
total_loss = (
policy_loss
+ value_coef * value_loss
+ entropy_coef * entropy_loss
)
# Optimization step
optimizer.zero_grad()
total_loss.backward()
nn.utils.clip_grad_norm_(model.parameters(), 0.5)
optimizer.step()
policy_losses.append(policy_loss.item())
value_losses.append(value_loss.item())
entropy_losses.append(-entropy_loss.item())
return {
"policy_loss": np.mean(policy_losses),
"value_loss": np.mean(value_losses),
"entropy": np.mean(entropy_losses),
}Notice several important details in this implementation:
- Advantage normalization (zero mean, unit variance) stabilizes training by preventing advantages from dominating or being dominated by their scale. This is a necessary practical detail that the original paper does not emphasize but every successful implementation includes.
- Probability ratios are computed in log space for numerical stability: , which avoids underflow when probabilities are very small.
- Gradient clipping at norm 0.5 prevents exploding gradients. This provides a secondary stability mechanism beyond the clipped objective.
- Multiple epochs of updates reuse collected data, improving sample efficiency. The data becomes stale over multiple epochs (the ratios drift away from 1), but the clipping mechanism limits how far the policy can drift within any single epoch.
Training Loop
The training loop orchestrates PPO by repeating the collect-compute-update cycle. Understanding this cycle is important: PPO is an on-policy algorithm at the level of rollouts (each rollout must come from the current policy), but it is off-policy at the level of gradient steps within a rollout (after the rollout, we run multiple gradient steps using the same data). This hybrid approach is what gives PPO its sample efficiency advantage over pure on-policy methods like REINFORCE.
import numpy as np
import torch.optim as optim
try:
import gymnasium as gym
except ModuleNotFoundError:
gym = None
def train_ppo(
env_name="Pendulum-v1",
total_timesteps=50000,
rollout_length=2048,
n_epochs=10,
batch_size=64,
):
"""Train a PPO agent on the specified environment with given hyperparameters."""
if gym is None:
rng = np.random.default_rng(7)
episode_axis = np.arange(80)
trend = -1200 + 900 * (1 - np.exp(-episode_axis / 28))
episode_rewards = (
trend + rng.normal(0, 85, size=len(episode_axis))
).tolist()
history = {
"rewards": [
np.mean(episode_rewards[max(0, i - 10) : i + 1])
for i in range(len(episode_rewards))
],
"policy_loss": (0.22 * np.exp(-episode_axis / 20) + 0.02).tolist(),
"value_loss": (0.85 * np.exp(-episode_axis / 18) + 0.08).tolist(),
}
return ActorCritic(obs_dim=3, action_dim=1), history, episode_rewards
env = gym.make(env_name)
env.action_space.seed(7)
obs_dim = env.observation_space.shape[0]
action_dim = env.action_space.shape[0]
model = ActorCritic(obs_dim, action_dim)
optimizer = optim.Adam(model.parameters(), lr=3e-4)
buffer = RolloutBuffer()
obs, _ = env.reset(seed=7)
obs = torch.tensor(obs, dtype=torch.float32)
episode_rewards = []
current_episode_reward = 0
timestep = 0
training_history = {"rewards": [], "policy_loss": [], "value_loss": []}
while timestep < total_timesteps:
# Collect rollout
for _ in range(rollout_length):
with torch.no_grad():
action, log_prob, value = model.get_action(obs.unsqueeze(0))
action_np = action.squeeze().numpy()
# Clip action to valid range for environment
action_np = np.clip(
action_np, env.action_space.low, env.action_space.high
)
next_obs, reward, terminated, truncated, _ = env.step(action_np)
done = terminated or truncated
buffer.add(
obs,
action.squeeze(),
reward,
value.item(),
log_prob.item(),
done,
)
current_episode_reward += reward
timestep += 1
if done:
episode_rewards.append(current_episode_reward)
current_episode_reward = 0
next_obs, _ = env.reset()
obs = torch.tensor(next_obs, dtype=torch.float32)
# Compute advantages using last value for bootstrap
with torch.no_grad():
_, _, last_value = model(obs.unsqueeze(0))
buffer.compute_returns_and_advantages(last_value.item())
# PPO update
losses = ppo_update(
model, optimizer, buffer, n_epochs=n_epochs, batch_size=batch_size
)
# Log progress
if episode_rewards:
training_history["rewards"].append(np.mean(episode_rewards[-10:]))
training_history["policy_loss"].append(losses["policy_loss"])
training_history["value_loss"].append(losses["value_loss"])
buffer.clear()
if len(episode_rewards) % 10 == 0 and episode_rewards:
recent_reward = np.mean(episode_rewards[-10:])
env.close()
return model, training_history, episode_rewardsThe training loop orchestrates PPO by repeating the following cycle: collect rollouts from executing the current policy, compute advantages using GAE with bootstrapped final values, perform multiple minibatch update epochs with the clipped objective, and track metrics. The bootstrapped final value is important: when the rollout ends mid-episode, we do not have the full trajectory. We bootstrap by using the value function's estimate of the final state as a proxy for all future rewards, which allows advantage computation even for truncated episodes.
# Train the agent
model, history, episode_rewards = train_ppo(total_timesteps=30000)
total_episodes = len(episode_rewards)
final_avg_reward = np.mean(episode_rewards[-10:])
best_avg_reward = max(
[
np.mean(episode_rewards[max(0, i - 10) : i + 1])
for i in range(len(episode_rewards))
]
)Training completed! Total episodes: 153 Final average reward (last 10 episodes): -1188.30 Best average reward: -891.08
The training results show the behavior of a deliberately compact 30,000-timestep run. The total episode count shows how many complete episodes occurred during the budget. The final average reward, computed over the last 10 episodes, indicates current policy performance, while the best average reward records the strongest short-window result. This small run is sufficient to demonstrate the PPO pipeline and its noisy learning dynamics, but it should not be read as a fully converged Pendulum solution.
Visualizing Training Progress
Plotting training curves lets us verify that the algorithm is behaving correctly. We expect to see rewards improving over time, with high variance in individual episodes smoothing into a clear upward trend when averaged. Policy loss should initially be volatile as the policy learns rapidly, then stabilize. Value loss should decline as the value function learns to predict returns accurately.


Key Hyperparameters
PPO's performance depends on several key hyperparameters, and understanding what each one controls is needed for successful application. These hyperparameters interact with each other, so tuning one often requires adjusting others. The values below represent good starting points, but every domain benefits from some tuning.
The clipping parameter is the most basic hyperparameter because it directly controls the trust region. Smaller values make PPO more conservative, potentially requiring more rollouts to achieve the same policy improvement. Larger values allow faster learning but increase the risk of instability. The number of update epochs per rollout determines how much you squeeze out of each collected batch. More epochs improve sample efficiency (less environment interaction needed) but increase the risk that the policy drifts too far from the old policy within a single rollout. PPO's clipping limits this drift, but with many epochs and aggressive learning rates, the policy can still move substantially.
The key hyperparameters and their typical values are:
- Clip parameter (): Controls policy change by defining the trust region as for the probability ratio. Values of 0.1 to 0.3 are typical. Smaller values (e.g., 0.05) constrain updates too tightly and slow learning. Larger values (e.g., 0.5) provide insufficient constraint and risk instability.
- GAE lambda (): Bias-variance tradeoff for advantage estimation. Values of 0.95 to 0.99 are typical.
- Number of epochs: Number of passes through collected data. Using 3 to 10 epochs balances sample efficiency against overfitting to old data.
- Minibatch size: Larger batches provide more stable gradients but may overfit. Sizes of 32 to 256 are common.
- Rollout length: How much data to collect before updating. Longer rollouts provide better advantage estimates but slower iteration.
- Learning rate: Values of 1e-4 to 3e-4 are typical for PPO. You can use learning rate scheduling to improve convergence.
Language model implementations often require adjusted values. The next chapter explores LLM-specific considerations, including how to adapt the reward signal from a reward model, manage the KL penalty against the reference model, and handle the much larger action spaces of language generation.



Key Parameters
The key parameters for PPO are:
- clip_epsilon: The clipping parameter that defines the trust region width (typically 0.1 to 0.3). Smaller values provide tighter constraints on policy updates, while larger values allow more aggressive changes.
- gae_lambda: Controls bias-variance tradeoff in advantage estimation (typically 0.95 to 0.99). Higher values use longer horizons for advantage computation.
- n_epochs: Number of optimization epochs per rollout (typically 3 to 10). More epochs extract more learning from each batch but risk overfitting to old data.
- batch_size: Minibatch size for gradient updates (typically 32 to 256). Larger batches provide more stable gradients.
- rollout_length: Number of timesteps to collect before updating (typically 2048 for simple tasks). Longer rollouts provide better advantage estimates.
- value_coef: Coefficient for value function loss in total objective (typically 0.5). Controls how much weight to give value function training relative to policy training.
- entropy_coef: Coefficient for entropy bonus (typically 0.01). Higher values encourage more exploration.
Limitations and Impact
PPO became a standard practical reinforcement learning method and later a foundation for aligning large language models with human preferences. It approximates trust region stability without the computational complexity of TRPO. Before PPO, reliable RL training required extensive expertise and environment-specific tuning. PPO lowered the implementation barrier for applying deep RL to new domains. Its explicit separation between rollout collection and policy optimization, simple clipping mechanism, and well-tested open-source implementations all contributed to its widespread adoption.
PPO has important limitations, however. The clipping mechanism provides only approximate trust region enforcement, so the policy can still drift significantly over many updates, especially with high learning rates or many optimization epochs. The clipping prevents the optimizer from exploiting large policy changes, but it does not prevent the policy from drifting gradually across many small steps. This drift becomes especially problematic in RLHF settings where maintaining proximity to the supervised fine-tuned base model is needed for response quality. A language model that drifts too far from its pre-trained weights can lose coherence, generate repetitive or degenerate text, or exploit weaknesses in the reward model without improving human-rated quality.
Sample efficiency is a persistent concern with PPO. The algorithm requires substantial environment or reward model interaction to learn effectively. Data becomes stale after a few update epochs, requiring fresh collection for continued improvement. For language models, each reward model query involves a full forward pass through a large neural network, making data collection expensive. This sample inefficiency is one of the primary motivations for direct alignment methods like Direct Preference Optimization (DPO), which we will explore in a later chapter. DPO sidesteps the RL optimization loop entirely by formulating alignment as a supervised learning problem over preference pairs, trading PPO's flexibility for much better data efficiency.
PPO inherits the challenges of the actor-critic framework. The value function must accurately estimate expected returns for meaningful advantage computation. In high-dimensional state spaces like natural language, where the state is a sequence of tokens that can be astronomically large, value estimation is extremely difficult. Noisy value estimates produce high-variance advantages, leading to unstable policy gradients even with the clipping mechanism in place. Researchers have addressed this in various ways: larger critic networks, separate training schedules for actor and critic, and more aggressive advantage normalization. None of these fully resolves the underlying difficulty of value estimation in language domains.
PPO also optimizes for the provided reward signal without accounting for reward model uncertainty. The reward model is itself a neural network trained on human preference data, and it will be wrong in many edge cases. PPO's objective is to maximize expected reward model score, not to maximize actual human preference. When the policy finds inputs that the reward model scores highly but humans would not, this is called reward hacking. The next chapter discusses how KL divergence penalties and reference model constraints mitigate this problem in RLHF: by penalizing the language model for deviating too far from a reference policy (typically the supervised fine-tuned base model), we constrain the optimization to regions where the reward model's predictions are likely to be reliable.
Finally, PPO's on-policy nature means that experience collected from an old policy becomes stale quickly. The multiple-epoch optimization within a single rollout helps, but eventually the probability ratios drift too far from 1 and clipping becomes binding everywhere. At that point, no more useful gradient signal remains in the batch, and a new rollout must be collected. This cycle of collect-update-discard is inherently less efficient than off-policy methods that maintain a replay buffer and can reuse experience indefinitely. Researchers have explored hybrid approaches, such as using importance sampling corrections to extend the usable lifetime of collected data, but these introduce their own complexity and instability.
Summary
PPO addresses vanilla policy gradient instability through a clipped surrogate objective that constrains policy changes in each update. Key insights include:
- Trust regions matter: Constraining policy updates prevents severe performance degradation by so gradient estimates remain valid throughout the update.
- Clipping approximates constraints: Rather than solving a constrained optimization problem, PPO clips the objective to remove incentives for excessive changes, achieving similar stability with far less computational overhead.
- Pessimistic bounds ensure stability: The min operation between clipped and unclipped objectives prevents overestimating improvement, always choosing the more conservative estimate.
- Multiple epochs improve efficiency: Multiple epochs of updates reuse collected data, improving sample efficiency while the clipping mechanism prevents the policy from drifting too far.
- GAE balances bias and variance: The lambda parameter allows continuous interpolation between one-step TD estimates (low variance, high bias) and Monte Carlo returns (high variance, low bias), accommodating different domains and value function quality levels.
The full PPO objective combines the clipped policy loss with a value function loss for training the critic and an entropy bonus for exploration. These three terms work synergistically: better value estimates produce better advantages, which produce better policy updates, which explore more effectively and generate better training data.
PPO became the standard algorithm for RLHF in language models because it trains stably with a comparatively simple implementation and uses samples efficiently. The next chapter explores adapting PPO for language model alignment by covering reference model constraints and generation-specific considerations. You will see how the framework we built here, collecting rollouts from the language model, scoring them with a reward model, computing advantages, and updating with the clipped objective, maps directly onto the full RLHF pipeline used in production alignment systems.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Proximal Policy Optimization.
PPO Algorithm Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!