Part of World Models Handbook
Compare PILCO, PETS, and MBPO for decision-centric model-based reinforcement learning, covering deep ensembles, uncertainty, planning, and policy optimization.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
PILCO, PETS, Ensembles, and MBPO
Introduction
Imagine you are handed a physical system you have never touched before: a remote-controlled car on a smooth floor, a robot arm with unfamiliar payload, or a chemical reactor with unknown gains. You are allowed a small budget of trials, perhaps fifty or a few hundred interactions, and then judged on how well you can steer the system to a target. When that budget is small relative to the task and failures cost time, hardware, or safety, direct trial-and-error learning may be impractical. One strategy is to learn a model of how the system responds, then use that model to reason about the outcome of candidate actions before committing to any of them.
This chapter compares three decision-centric model-based reinforcement learning algorithms: PILCO, PETS, and MBPO. It also covers deep ensembles, a general predictive-uncertainty method adopted by PETS and related systems. They form a useful conceptual progression for studying model error, but not a literal inheritance chain. Deep ensembles predate PETS, PETS combines them with neural dynamics and sampled planning, and MBPO reuses a related ensemble model for policy learning rather than PETS-style online planning.
The central tension is straightforward to state and difficult to resolve. A model learned from limited data can be wrong in ways that matter. If a policy or action sequence is optimized aggressively against that model, the optimizer can select behavior that benefits from its errors, especially outside the observed data distribution. This is called model exploitation, and it explains why low average one-step prediction error need not produce good control. The methods in this chapter offer different, imperfect ways to limit that failure.
- PILCO uses Gaussian-process dynamics, analytic moment matching with repeated Gaussian approximation, and analytic policy gradients to optimize a parameterized feedback policy under a predictive state distribution.
- PETS combines probabilistic neural ensembles, particle-based uncertainty propagation, CEM, and receding-horizon control.
- Deep ensembles provide a scalable empirical proxy for epistemic uncertainty; they are useful in PETS but are a separate method with their own history and limitations.
- MBPO uses short model-generated branches from states in a real replay buffer to supply training data for an off-policy policy learner.
Before diving in, recall the frame from Part II: Inference, Decisions, and Control: a world model predicts consequences, while a decision procedure consumes those predictions. This chapter is about the interface between the two. We will distinguish what each method predicts, what assumptions it makes, and whether the model is used for online planning or for policy training. We will also lean on the treatment of uncertainty from Part IV: Space, Physics, and Causality and on model predictive control from Part VII: Planning and Agency.
Three notions of model quality recur throughout the chapter. Predictive fidelity concerns the accuracy of predictive distributions on a relevant evaluation or deployment distribution. Uncertainty quality concerns calibration, sharpness, and behavior under relevant distribution shift. Decision usefulness concerns whether the composed model-and-decision system produces good actions. A fourth notion, representation quality, asks whether the model retains the variables needed for control rather than incidental detail. This chapter defines that fourth lens but treats it only briefly; later latent-model chapters develop it directly. The properties are not interchangeable: good average predictions need not be calibrated off distribution, and good uncertainty estimates cannot rescue a defective objective or optimizer.
There is a useful continuity with system identification, filtering, and predictive control, but the correspondence is an analogy rather than a sharp historical divide. Classical identification and MPC already address model mismatch, disturbances, uncertain parameters, validation, and both hard and soft constraints. System identification is itself model learning. PETS uses sampled MPC with a learned probabilistic dynamics model. MBPO instead generates model transitions to train a policy, typically with SAC, rather than selecting actions with online MPC. These choices affect approximation and data reuse without removing the older concerns. Even when the physical plant is stationary, the estimated model can change as evidence accumulates. Likewise, an uncertainty penalty is one optional learned-control design choice, not a defining replacement for hard constraints.
Probabilistic dynamics and data efficiency
The motivating problem is data efficiency. On some benchmark tasks, model-free policy-gradient and actor-critic baselines require millions of environment steps, as illustrated by the MBPO benchmark comparison. A model-based method adds a supervised learning problem: each observed transition trains a predictor of local dynamics. Under an appropriate smoothness prior or other inductive bias, that observation can inform predictions at nearby state-action pairs. Once the model is useful, planning or policy learning can consume synthetic transitions without further environment interaction. PILCO was an influential early demonstration that this strategy could achieve exceptional data efficiency on continuous-control benchmarks, as a later probabilistic-MPC comparison also notes.
The contrast is not that model-free methods reduce every transition to a few bits or ignore its next state. Value and policy algorithms may use states, actions, rewards, next states, temporal-difference errors, advantages, or distributional targets. The distinction is that a dynamics learner uses the transition as direct supervision for a reusable transition model. Generalization from one transition to a neighborhood comes from the model class and its fitted inductive bias, not from the tuple alone. Queries to the learned model then produce synthetic transitions, not additional real experience. Suitable priors or smoothness assumptions may improve data efficiency by helping a dynamics model generalize from limited data. Propagating uncertainty and choosing an appropriate objective can limit exposure to model bias when their assumptions hold. None of these choices guarantees safe extrapolation.
State uncertainty and dynamics uncertainty are also distinct rather than mutually exclusive. A state estimator may maintain a belief over the current hidden state while using known, estimated, or jointly inferred dynamics. PILCO learns uncertain dynamics and propagates uncertain state distributions through them. In both cases dispersion records uncertainty conditional on the model assumptions, and new observations update the relevant belief. A fuller belief-space controller can maintain uncertainty over state and model parameters or functions at the same time.
Why a point estimate is not enough
Suppose you fit a deterministic model and optimize a controller against its predicted trajectories. Away from observed data, the predictions are weakly constrained by evidence and are determined largely by architecture, regularization, and optimization. A sufficiently aggressive controller search can then select actions that benefit from those off-support errors. This is a risk of optimizing one fitted hypothesis, not proof that every point-estimate controller must fail.
The failure can arise structurally when three conditions coincide: the model is inaccurate in reachable regions, the optimizer can search those regions, and the objective rewards the resulting prediction errors. Better coverage, regularization, constraints, robust objectives, or benign task structure can prevent it. Representing model uncertainty can help, but only if the decision rule uses that information appropriately; posterior averaging alone is not a universal safety mechanism.
Formally, let the true dynamics be the conditional distribution . A probabilistic model assigns a distribution parameterized by . In the Bayesian view, itself is uncertain, giving a posterior after seeing data . To predict the next state, we average over all models weighted by how plausible each is under the data, which means marginalizing over the uncertain model parameters:
where:
- : the state at time
- : the action taken at time
- : the parameters of the probabilistic dynamics model
- : the model's predicted distribution over the next state given the current state and action
- : the posterior distribution over model parameters after observing data
- : the dataset of observed transitions
The integral averages predictions from all parameter settings according to their posterior plausibility. Under an identifiable, well-specified model with an appropriate prior and likelihood, informative data can concentrate the posterior and reduce latent predictive uncertainty. In data-sparse regions, uncertainty may remain larger. Neither behavior is automatic under misspecification, and integrating over parameters is not itself a risk penalty: it defines a posterior-predictive distribution.
Consider a dynamics model with one unknown gain. Data near a well-explored operating point may narrow the posterior over that gain, while predictions elsewhere may remain sensitive to the prior and model assumptions. Bayesian averaging prevents the planner from evaluating only the single most flattering gain, but a risk-neutral expected return can still favor an uncertain action if its posterior mean payoff is high. A controller that should avoid optimistic tails needs a suitable nonlinear utility, a lower-confidence or worst-case objective, CVaR, or another explicit risk criterion.
Two kinds of uncertainty can contribute to the posterior-predictive distribution . At fixed , can represent aleatoric outcome variability; uncertainty over adds an epistemic component after marginalization. Aleatoric uncertainty is outcome variability that remains after the modeled state, action, observation process, and environment regime are fixed. Epistemic uncertainty concerns the unknown dynamics parameters or function and can shrink with informative data under suitable model and training assumptions. Epistemic uncertainty is the component most directly associated with ignorance in unexplored regions, but aleatoric risk and model misspecification can also dominate a decision.
An operational test requires controlled conditions. Repeating the same fully specified physical transition in a stationary regime can reveal intrinsic process variation, although unmodeled hidden state can confound the comparison. If the prediction target is the latent physical transition, stochastic sensor noise can look like process noise; for a measured-outcome target, residual stochastic sensor noise is an aleatoric component, while systematic sensor bias is not automatically one. Disagreement among models retrained with different samples or optimization randomness is a practical epistemic proxy, not a pure measurement: it contains sampling and training variability. Additional informative local data may narrow disagreement when the model class and training procedure turn those observations into more similar local fits, but the change is not guaranteed or monotone. The two uncertainties call for different responses. Irreducible variability may require expected-outcome or robust planning; reducible ignorance may justify targeted data collection. Models that expose only a collapsed total predictive variance make that distinction harder to use.
Gaussian processes and the analytic route
Let and , and set with . Let denote one scalar component of a latent dynamics function. Assume and observations , where the are mutually independent and independent of . A positive-semidefinite kernel defines the covariance of any finite collection of function values. Assume , so is invertible even if training inputs coincide. Given training inputs and scalar-output targets , the posterior distribution of the latent function value at is Gaussian with
Here has entries , contains the covariances between the query and all training inputs, is the identity, and . The displayed is latent-function variance. The predictive variance of a new noisy observation is under this homoscedastic likelihood.
For fixed inputs and fixed kernel and noise hyperparameters, the posterior mean is linear in the targets, and its effective weights are the entries of . They are obtained by solving the full kernel system, generally depend jointly on the training inputs, and can even be negative. A training point's query-to-point similarity alone therefore does not determine its weight; the full Gram matrix, query covariance vector, and noise variance matter. With a local squared-exponential kernel, nearby isolated observations often exert more influence, but that is an intuition rather than a coefficient-by-coefficient guarantee.
The latent variance depends on the input locations and fitted kernel and noise parameters, not on the observed target values. It generally decreases near informative observations, but with noisy data it need not reach zero even at a training input. For a stationary kernel whose covariance decays with distance, it approaches the prior variance when all query-to-training covariances vanish. These statements remain conditional on the kernel, likelihood, hyperparameters, and model specification: a misspecified GP can still be confidently wrong. With PILCO's squared-exponential kernel and differentiable moment, policy, and cost expressions, the resulting approximate predictive moments are differentiable in the quantities used for policy optimization.
The kernel encodes assumptions about smoothness and characteristic variation scales. PILCO uses a squared-exponential kernel,
where and concatenate state and action, is the signal variance, and
is a matrix containing the squared characteristic length scales. A shorter permits covariance to decay more rapidly along input dimension ; a longer one favors slower variation. PILCO fits these and other GP hyperparameters with gradient-based log-marginal-likelihood optimization. Its released GP trainer also adds numerical-stability penalties for extreme length scales and signal-to-noise ratios.
The squared-exponential prior favors extremely smooth functions: its sample paths are infinitely differentiable under the usual construction, and nearby inputs receive high prior covariance. The length scales control typical variation, but they do not impose a deterministic bound on a function's derivative; derivative distributions still have unbounded Gaussian support. This strong smoothness prior can work well for smoothly varying mechanics, but stationary squared-exponential dynamics can fit contact, stiction, and discrete switching poorly because those regimes conflict with the prior's smoothness and stationarity assumptions.
Closed-form moment matching
Suppose the state at time is approximated by and the deterministic feedback policy is . A generic additive-noise transition can be written as
PILCO specifically fits independent GPs to state increments and reconstructs the next state by adding the predicted increment to . The fitted target noise and the GP's latent-function uncertainty must therefore be distinguished from any physical process-noise interpretation.
Passing an uncertain Gaussian state through a nonlinear policy and GP predictive model generally produces a non-Gaussian distribution. Non-Gaussianity alone does not make moments intractable, but the required integrals are not available in closed form for arbitrary kernels and policies. PILCO exploits compatible squared-exponential kernels and policy forms to compute the needed first and second moments analytically, then projects the result back to a Gaussian. Thus the integrals used for each update are analytic while the repeated Gaussian moment-matching inference is approximate. The PILCO authors also note that their rollout does not retain temporal correlation in model uncertainty: it treats that uncertainty similarly to temporally uncorrelated noise, which can underestimate it. This caveat does not erase the benefit of uncertainty-aware evaluation, but it rules out reading the rollout as an exact joint distribution over a fixed uncertain dynamics function.
The resulting and are differentiable functions of , , and , so gradients of predicted cost can flow back to the policy parameters. Many alternative kernels or policy nonlinearities would require new moment derivations or numerical approximation; the limitation is tractability for this pipeline, not a claim that every other kernel has intractable moments.
With the moment-matching machinery in place, we can state the control objective: choose policy parameters that minimize the total cost the agent expects to incur along a finite horizon. Let be a cost function (PILCO uses a saturating cost; more on that shortly). The expected cost over a finite horizon of length is
where:
- : the expected cost over the horizon, as a function of the policy parameters
- : the finite planning horizon
- : the per-state cost function
- : the Gaussian predictive distribution over states at time , obtained by moment matching
- : the mean and covariance of that predictive distribution
Because the matched moments are differentiable in , PILCO computes without Monte Carlo trajectory gradients. The expected cost incorporates the GP's predictive uncertainty and can mitigate model exploitation, but approximate inference and model misspecification prevent an unconditional guarantee.
The objective is a finite, undiscounted cost sum. PILCO does not optimize one fixed open-loop action sequence: it propagates the state distribution induced by the same parameterized feedback policy over the horizon. Each step depends on the previous predictive moments and on , so the gradient follows the same dependency pattern as backpropagation through time. The recurrent operation is an analytic moment update under an approximate Gaussian state distribution rather than a deterministic update of one state.
The saturating cost
One more design choice deserves attention. A positive-definite quadratic distance cost grows without bound as the state moves away from the target. Under a Gaussian predictive state distribution, its expectation contains both squared mean error and a term proportional to predictive covariance. Large uncertainty can therefore strongly influence the objective, although unstable gradients or literal tail dominance do not follow automatically. For normalized or otherwise commensurate state coordinates, an isotropic PILCO-style bounded saturating cost is
where sets the distance scale. The exponential term is largest at the target and decays with distance, so the cost moves smoothly from low cost near the target to a high but bounded cost far away. Specifically, at the target and approaches as the distance grows.
The bounded cost limits how much very distant predictive states can contribute. Once the distance is large relative to , moving still farther changes the cost only slightly. Compared with an unbounded quadratic, this reduces the influence of remote uncertain states and emphasizes distinctions near the task-relevant target scale. It does not prove that those remote states are impossible or irrelevant; it is an objective-design choice that controls their influence on policy optimization.
What data efficiency bought, and what it cost
The PILCO experiments reported control learning in a small number of trials on several continuous-control systems. The result made the method an influential example of data-efficient policy search. Its restrictions matter just as much:
- Computational cost scales poorly. Standard exact GP training with a dense, unstructured covariance matrix costs time per Cholesky factorization and memory, and moment propagation is expensive. The workable point count is implementation- and problem-dependent; hundreds to low thousands is a rough historical scale, not a hard limit.
- The prior and analytic forms are restrictive. Stationary squared-exponential kernels are poorly matched to discontinuities and switching, while analytic moment calculations require compatible kernels and policy forms.
- Repeated Gaussian projection can degrade. Long-horizon propagation repeatedly compresses a potentially non-Gaussian distribution into its first two moments.
PETS directly addresses GP scalability by using probabilistic neural ensembles and Monte Carlo propagation. MBPO addresses a different problem: bias from using learned dynamics to generate policy-training data over long horizons. It would be misleading to present MBPO's short branches as a direct remedy for PILCO's particular approximation errors.
Trajectory sampling with ensembles
PETS, short for Probabilistic Ensembles with Trajectory Sampling, estimates action-sequence values with particles propagated through a probabilistic neural ensemble. The ensemble members are plausible fitted functions and their disagreement is an empirical epistemic proxy; ordinary deep-ensemble members are not formal draws from a Bayesian posterior.
PILCO's one-step GP posterior is analytic under its model assumptions, but its long-term state prediction uses approximate Gaussian moment matching. Its data efficiency cannot be attributed to end-to-end exactness. PETS instead estimates action-sequence returns by finite-sample particle rollout through its ensemble. Increasing particle count can reduce Monte Carlo error in a candidate's value estimate, while model bias, ensemble approximation error, and action-search error remain separate limitations.
The probabilistic neural network model
At the model level, PETS represents the conditional next-state distribution as a diagonal Gaussian. In the released implementation, the network may receive a preprocessed observation with the action and predicts Gaussian parameters for a task-specific target, with variances parameterized to be strictly positive: the Cartpole configuration uses state increments, while HalfCheetah uses an absolute first next-state coordinate and increments for the others. Deterministic postprocessing reconstructs the next state. The following and denote moments of that induced next-state distribution, not necessarily raw network outputs:
where:
- : the model's predicted distribution over the next state
- : the predicted mean next-state vector after target postprocessing
- : the predicted diagonal next-state covariance after target postprocessing
- : the network parameters
The diagonal covariance treats modeled output dimensions as conditionally independent. The released target maps above are coordinatewise affine with unit Jacobian, so the Gaussian data-fit term can also be written as a negative log-likelihood of the induced next states. It penalizes both an incorrect mean and an overconfident variance: an incorrect mean increases squared error relative to predicted variance, and a variance that is too small magnifies that error. The unregularized data-fit term is
where:
- : the Gaussian negative log-likelihood data-fit term as a function of the network parameters
- : the dataset of observed transitions
- : the observed next state
- : the likelihood the model assigns to the observed transition
The released PETS training code adds weight decay and penalties on learned log-variance bounds to this Gaussian data-fit term; the full optimization is regularized, not pure maximum likelihood. Training an input-dependent variance output with log-likelihood, rather than fitting only a mean by squared error, permits a heteroscedastic model. With adequate data and a well-fitted conditional model, it can predict higher variance where transitions have greater irreducible variability and lower variance where they are consistent. The learned is then an aleatoric proxy, not a clean measurement: mean-model error can also inflate it, and the training loss does not identify its value in unseen regions. A single network can therefore output a confident, low-variance prediction where it has never seen data. Ensemble disagreement supplies a separate empirical proxy for epistemic uncertainty.
A single probabilistic network with a likelihood data-fit term can learn input-dependent aleatoric variance on observed data, but that term does not directly constrain epistemic uncertainty at unobserved inputs. Extra capacity can produce confident off-support behavior, but it does not necessarily make that behavior worse; the outcome depends on architecture, optimization, regularization, and data.
Why ensembles provide an epistemic proxy
The deep-ensemble recipe trains networks with independent randomization and, in PETS, bootstrap resampling. For vector-valued predicted means, define
For a scalar output, the covariance expression reduces to the familiar average squared deviation. For a state-vector output, and the covariance is a matrix. The between-member covariance is used as an epistemic proxy. With probabilistic members, a useful approximate total-covariance decomposition is
which keeps average within-member predictive variance separate from between-member disagreement.
In observed regions, independently trained members can converge toward similar functions; away from the data, different random initializations, training paths, and optional bootstrap resampling can lead to different extrapolations. This behavior is common enough to be useful, but it is not guaranteed or monotone in data density. Correlated members may agree while sharing the same error, and optimization variability can create disagreement even in densely sampled regions.
Ordinary deep-ensemble members are not posterior samples. Bootstrap resampling approximates a frequentist sampling distribution under its own conditions, not generally the Bayesian function posterior conditional on one observed dataset. The method is therefore best described as an empirical ensemble used as an epistemic proxy. A related Randomized Prior Functions construction adds a fixed, randomly drawn prior function to each trainable member's output; it does not merely randomize the final layer and train the remainder. These methods can work well empirically without turning their fitted-member distribution into an analytic or simulation-exact posterior.
Trajectory sampling with ancestral sampling
PETS evaluates each candidate action sequence with multiple Monte Carlo state particles. Candidates, particles, and ensemble members are different objects. Let index a CEM candidate, index one of its particles, and identify the ensemble member assigned to that particle at simulated time . Starting every particle from the measured state, PETS propagates
The candidate score is the particle average of cumulative reward,
Here is an abstract task-supplied transition score, which may depend on the sampled next state. In the released PETS controller, each predicted next observation receives a task-supplied observation cost and its action receives a task-supplied action cost; maximizing corresponds to minimizing their sum. The dynamics network does not learn those costs through the Gaussian transition likelihood. This differs from original MBPO's learned dynamics-and-reward ensemble.
Thus CEM's candidate population controls which action sequences are searched, while controls Monte Carlo uncertainty propagation and the precision of each candidate's return estimate. PETS does not require one trajectory per ensemble member, nor does a particle denote an action-sequence candidate.
The original PETS study defines two bootstrap-assignment semantics:
- TS1 resamples at every simulated step. It represents a transition process that marginalizes over or resamples the ensemble hypothesis through time.
- TS samples one member for each particle and holds it fixed across the horizon. It represents a time-invariant sampled dynamics hypothesis and preserves trajectory-level effects of that member.
Both are valid PETS variants. A TS1 trajectory need not coincide with the rollout of any one fixed member, but it is not therefore a statistical fiction. TS1 and TS make different assumptions about temporal dependence in model uncertainty. There is no general identity saying that TS1 must average away long-horizon uncertainty or that TS is always more conservative.
CEM proposes candidate action sequences. PETS then propagates multiple state particles for each candidate. Each particle samples aleatoric transitions from an assigned probabilistic ensemble member, with that assignment updated according to TS1 or TS. The candidate is ranked by its average particle return.
Cross-entropy method for action selection
PETS uses the cross-entropy method (CEM) to search over action sequences:
- Initialize a proposal distribution over sequences.
- Sample a population of candidate sequences.
- Estimate each candidate's return with its own state particles.
- Select a high-return elite set.
- Refit the proposal to the elites.
- Repeat the sampling and refitting updates until a stopping criterion is met or the maximum number of iterations is reached. Return a sequence derived from the final proposal; the released PETS optimizer returns its smoothed mean. The surrounding receding-horizon controller executes only the first action.
Candidate and particle evaluations can be batched across neural-network calls. CEM is derivative-free, but it is not immune to search failure. Too few candidates can miss narrow high-value regions; too little proposal variance can collapse the search prematurely; too much variance can spend computation on poor regions; and an extreme elite fraction can make updates either brittle or weak. The appropriate population depends on action dimension, horizon, proposal family, and compute budget rather than a field-wide typical range.
Sampled elite return is one optimization diagnostic, not proof that concentration is beneficial. Proposal variance and returns reevaluated with fresh or additional particles are useful diagnostics for sampling noise and concentration. Distinguishing premature collapse from convergence to a strong solution also requires broader-search or restart evidence.
The receding-horizon loop
PETS plans over a finite horizon, executes only the first action, observes the resulting state, and replans. Receding-horizon control therefore limits each open-loop commitment to one action and each uncorrected forecast to the configured planning horizon. It does not numerically bound model error or guarantee global competence. Success can still depend on horizon length, terminal design, feasibility, exploration, and accuracy in regions the task must traverse.
The system loop is:
- Collect an initial transition dataset.
- Train the probabilistic ensemble.
- Use CEM and particle trajectory sampling to rank candidate action sequences.
- Execute the first action, observe the transition, and append it to the dataset.
- Retrain the ensemble according to a configured cadence.
The published PETS score is the arithmetic mean of particle returns. That is a risk-neutral Monte Carlo estimate of expected return under the fitted ensemble predictive distribution, not an automatic uncertainty hedge. One low return lowers the mean arithmetically, but two return distributions with the same mean receive the same score even if their spreads differ.
Retraining too rarely leaves the ensemble stale, while retraining too often increases computation. A usable trigger should be observable: for example, retrain after a fixed number of transitions and monitor held-out likelihood, calibration, or prediction drift in recently visited regions. A claim that the dataset has changed the model needs one of those quantities to make it operational.
Uncertainty-aware model predictive control
"Uncertainty-aware" can describe several different mechanisms. Keeping them separate prevents a risk-neutral estimator from being mistaken for a safety objective.
Three roles for uncertainty
The first role is uncertainty propagation under an expected-return objective. PETS samples predictive trajectories and averages their returns. This estimates expected return under the fitted ensemble predictive distribution. It accounts for how uncertainty passes through nonlinear dynamics and rewards, but it imposes no generic penalty on return variance or ensemble disagreement.
The second role is explicit risk sensitivity or an uncertainty penalty. One possible objective is
where might measure return spread, a lower-confidence adjustment, or another task-defined uncertainty statistic. The coefficient makes the risk tradeoff explicit and must put the uncertainty term on a scale commensurate with return. Worst-case aggregation and CVaR are other possibilities. Baseline PETS does not include such a term.
The third role is uncertainty as an information signal. A controller may choose actions partly because the resulting observations are expected to improve its model. That dual-control objective values future information rather than merely charging for present uncertainty.
A method can combine these roles, but its behavior should be attributed to the actual objective. PETS performs uncertainty propagation with risk-neutral averaging. An added penalty would be a separate modifier, and an information bonus would add an exploration motive.
What receding-horizon control changes
For PETS, MPC executes one action from each optimized sequence and then replans from a new observation. This limits the uncorrected forecast span used by one plan and limits commitment to its first action. It does not place a numerical bound on error or make global task success depend only on local accuracy. Reported PETS horizons are on the order of tens of steps for its tasks, but the required horizon is task-specific; delayed rewards, terminal conditions, constraints, and necessary travel through poorly modeled regions can all require additional design.
Ensemble disagreement in practice
Suppose ensemble members disagree near a joint limit. If that disagreement affects reward-relevant predictions within the planning horizon, particle sampling can produce a wider or differently shaped return distribution for action sequences entering that region. Under arithmetic averaging, disagreement alone has no direction: pessimistic and optimistic deviations may cancel, reward curvature may lower or raise the mean, or the mean may remain unchanged. CEM ranks the resulting means unless an explicit risk criterion is added.
Consequently, baseline PETS does not automatically avoid every blind spot, and ensemble size or regularization is not a principled caution knob. If a controller with an explicit uncertainty penalty avoids a necessary region, reducing that penalty is one possible response. If risk-neutral PETS avoids the region because its estimated expected return is poor, the relevant remedies are better targeted data, an information objective, a changed task objective, or a deliberately risk-sensitive decision rule. There is no penalty to relax in baseline PETS.
Separating uncertainty from risk attitude
Aleatoric and epistemic uncertainty can both alter expected return through nonlinear dynamics and rewards; neither has a universal "variance only" or "downward bias" effect. Their practical distinction is instead about what can be reduced. Within-member predictive variance is an aleatoric proxy under a well-fitted member model, while between-member mean disagreement is an epistemic proxy. Retaining both terms preserves a useful operational decomposition; collapsing them into total variance loses that distinction. Even when retained, the terms do not uniquely identify pure aleatoric and epistemic uncertainty: model misspecification can affect within-member variance, and training variability or shared error can affect between-member disagreement.
Irreducible variability may call for expected-outcome, robust, or risk-sensitive planning. Reducible disagreement may motivate data collection, but only after checking that the ensemble spread is calibrated enough to serve as a local signal. Readers interested in the deeper treatment of stochasticity should revisit Part IV, Ch 5: Stochasticity and Uncertainty.
Choosing horizon, ensemble size, particles, and candidates
The main compute allocations have different jobs:
- Ensemble size . More members can reduce finite-ensemble estimation noise and may improve functional diversity, but correlated or systematically biased members need not become a better posterior proxy. Training and inference cost grow approximately linearly before batching effects.
- Particles per candidate . More particles improve predictive-distribution coverage and reduce Monte Carlo noise in each candidate's expected-return estimate. They do not increase the number of action sequences searched.
- CEM candidate population . More candidates increase action-sequence search coverage within an iteration.
- Horizon . A longer horizon exposes delayed consequences but uses more modeled steps, which can increase sensitivity to model error.
- CEM updates and elite fraction. More updates can refine the proposal. Too small an elite fraction can collapse it prematurely, while too large a fraction weakens each update.
These parameters share a compute budget and should be tuned jointly, but one does not mechanically justify increasing another. A larger ensemble does not prove that a longer horizon is safe, and more particles do not by themselves justify more CEM iterations.
Failure signatures should follow the same distinctions. Too few candidates underpower action search. Too few particles produce noisy or poorly resolved candidate-value estimates. Too few or highly correlated ensemble members make disagreement unreliable. An excessive horizon exposes the score to longer model rollouts. Aggressive elite selection can lock onto sampling noise. CEM can fail abruptly through premature collapse or failure to sample a narrow optimum, so graceful degradation should not be assumed.
Short model rollouts for policy optimization
We now shift from online planning to policy learning. MBPO uses its learned model to generate transitions for an off-policy learner during environment training; it does not use the model as an MPC planner at action time.
Long free-running model trajectories can expose policy learning to distribution shift. Multi-step deviation need not increase monotonically: stable errors can saturate or cancel, and an exact model does not drift. Even so, the opportunity for accumulated error and off-support states generally increases with horizon. Calling arbitrary long model-generated policy data "model-based value expansion" is also too broad: Model-Based Value Expansion is a specific short-rollout value-target method.
MBPO instead starts branches from states sampled from the real-environment replay buffer and rolls the current policy through the model for up to steps. A short branch limits consecutive model use and repeatedly refreshes branch starts from real data. This does not guarantee that every generated transition is accurate, but it reduces exposure relative to a full free-running trajectory.
The useful horizon is task- and training-stage-dependent. The original paper reports single-step rollouts as a strong baseline and a Hopper configuration whose rollout length is scheduled from 1 to 15; that is not a universal 1–15 rule. The claim that most generative-model value always lies in the first few steps should likewise be read as an empirical result on the reported benchmarks, not a general theorem.
Before the algorithmic details, consider a deliberately hand-designed illustration. It defines a saturating benefit and a linear cost as functions of rollout length. Their difference has an interior maximum by construction; this is not an empirical MBPO result or a derivation of its theorem.

The three ingredients of MBPO
Original MBPO combines:
- A probabilistic dynamics-and-reward ensemble. In the original MBPO paper, each member predicts next state and reward conditioned on current state and action. At each simulated transition, the algorithm samples a model from the ensemble; disagreement is not used as an explicit stopping rule or caution penalty.
- A short-branch generator. Branch starts are sampled from the real-environment replay buffer. The current policy supplies actions for model branches up to a globally configured rollout cap , and generated transitions enter a model replay buffer.
- An off-policy learner. In the released implementation, SAC updates its actor and Q-functions from minibatches containing both real and model-generated replay data.
The released MBPO model is trained to predict reward and a state increment. Before sampling, the released simulator adds the current state to each ensemble member's predicted increment mean; the sampled vector is then split into reward and next state. A task-specific termination function supplies the done flag. The resulting synthetic replay entry contains ; termination is not predicted by an extra model head.
Sampling across ensemble members can diversify generated transitions when their predictive distributions differ, but baseline MBPO does not infer a locally safe horizon from ensemble spread, reject uncertain transitions, or make SAC intrinsically cautious. Those would be additional algorithmic mechanisms.
Reading MBPO's rollout bounds
The paper's analysis compares true return with return under -step branches. Its first worst-case bound uses model error measured under the data-collecting policy together with divergence between that policy and the updated policy. The paper explicitly describes this version as overly pessimistic: taken literally, it favors real off-policy data over model-generated data and does not justify a positive rollout length.
The paper's refined analysis aims to estimate model error on the updated policy's state-action distribution. The expression printed in Theorem 4.3 has two qualitative dependencies on :
- policy-divergence contributions decrease geometrically with , where is the discount factor;
- the current-policy model-error contribution grows linearly with .
All terms are scaled by reward and discount-dependent factors. For this printed penalty, sufficiently small can make a positive finite preferable to ; that algebra does not show that actual policy return must have an interior optimum. Lemma B.4 in the full paper with appendices defines pre- and post-branch errors but applies them to reversed time ranges in its proof, a mismatch also identified by an independent reproduction. The main Theorem 4.3 statement also describes updated-policy model error while indexing the displayed condition by the data-collecting policy; Appendix A.2 uses the updated policy. We therefore treat the printed tradeoff as the paper's motivation for short model use, not an independently verified performance guarantee. Model error on the current policy distribution still has to be estimated, and rollout schedules require empirical validation.
Original rollout schedule and data mixture
Original MBPO uses a rollout length scheduled globally by training epoch, not a state-dependent horizon chosen from local disagreement. Every branch generated at a given training stage uses that configured as a maximum; terminal branches can end sooner. Ensemble spread remains a heuristic disagreement statistic; using it for adaptive local horizons would require a separately specified and validated method.
The algorithm maintains a real replay buffer and a model replay buffer. In the released implementation, policy minibatches mix both sources using a configured real-data fraction. The paper's Algorithm 2 pseudocode, by contrast, writes policy updates on alone. For example, a released HalfCheetah configuration fixes the real-data fraction at 0.05 within a run, while the MBPO class separately defaults to 0.1 if no override is supplied; model-generated samples form most of the minibatch in either case. Meanwhile, the rollout length may increase on a global schedule as training progresses. These are implementation-specific choices, not a recommendation that every task use the same fraction, and original MBPO does not prescribe shrinking the synthetic share as real data accumulates.
Worked example: fixed-controller branch evaluation under model bias
We will use a one-dimensional system to isolate multi-step prediction error. The true dynamics and biased model receive the same initial states, action sequences, and process-noise draws. These common random numbers remove sampling differences from the model-versus-truth comparison; the remaining discrepancy comes from the action-gain bias.
import numpy as np
A_true, B_true = 0.9, 0.5
A_hat, B_hat = 0.9, 0.6
horizon = 12
n_trials = 200
rng_rmse = np.random.default_rng(0)
initial_states = rng_rmse.normal(size=n_trials)
actions = rng_rmse.normal(size=(n_trials, horizon))
process_noise = 0.05 * rng_rmse.normal(size=(n_trials, horizon))
s_true = initial_states.copy()
s_biased = initial_states.copy()
s_exact = initial_states.copy()
biased_sq_error = np.zeros(horizon)
exact_sq_error = np.zeros(horizon)
for t in range(horizon):
eps_t = process_noise[:, t]
a_t = actions[:, t]
s_true = A_true * s_true + B_true * a_t + eps_t
s_biased = A_hat * s_biased + B_hat * a_t + eps_t
s_exact = A_true * s_exact + B_true * a_t + eps_t
biased_sq_error[t] = np.mean((s_biased - s_true) ** 2)
exact_sq_error[t] = np.mean((s_exact - s_true) ** 2)
rmse = np.sqrt(biased_sq_error)
rmse_exact = np.sqrt(exact_sq_error)The exact-model control line is zero because it receives the same disturbances as the true system. This is a diagnostic use of common random numbers, not a claim that a predictive distribution has zero aleatoric variance.
step 1: biased-model RMSE = 0.1043 step 2: biased-model RMSE = 0.1308 step 3: biased-model RMSE = 0.1524 step 4: biased-model RMSE = 0.1720 step 5: biased-model RMSE = 0.1906 step 6: biased-model RMSE = 0.1969 step 7: biased-model RMSE = 0.2177 step 8: biased-model RMSE = 0.2192 step 9: biased-model RMSE = 0.2189 step 10: biased-model RMSE = 0.2201 step 11: biased-model RMSE = 0.2241 step 12: biased-model RMSE = 0.2226 Action-gain bias: B_hat - B_true = 0.10 Final-step RMSE (12 steps): 0.2226

The biased curve need not increase at every sampled step. In this stable system, the error obeys a stable recurrence driven by the random term , so its population variance approaches a finite limit. Different systems can amplify, cancel, oscillate, or saturate model error. Dimensionality alone does not determine error propagation: stability, coupling, error direction, and controller sensitivity matter. The task objective can change how consequential an error is and, when it changes action selection, which trajectories are encountered.
Coherent fixed-controller branch evaluation
The next toy is deliberately narrower than policy training. No policy is learned, no replay mixture is optimized, and this is not a literal MBPO implementation. We evaluate a grid of fixed proportional controllers, , under the true and biased simulators. Each branch starts from the same bank of states, runs coherently for consecutive steps, and receives the same process-noise draws in both simulators and for every candidate gain.
rng_branches = np.random.default_rng(43)
n_branches = 4096
max_branch = 12
branch_lengths = np.arange(1, max_branch + 1)
Ks = np.linspace(0.0, 2.0, 41)
branch_starts = rng_branches.normal(size=n_branches)
branch_noise = 0.05 * rng_branches.normal(size=(max_branch, n_branches))
def cumulative_branch_scores(A, B):
"""Score fixed feedback gains on coherent branches with shared randomness."""
scores = np.empty((max_branch, len(Ks)))
for j, K in enumerate(Ks):
s = branch_starts.copy()
cumulative_cost = np.zeros(n_branches)
for t in range(max_branch):
a = -K * s
s = A * s + B * a + branch_noise[t]
cumulative_cost += s**2 + 0.05 * a**2
scores[t, j] = -np.mean(cumulative_cost)
return scores
true_branch_scores = cumulative_branch_scores(A_true, B_true)
model_branch_scores = cumulative_branch_scores(A_hat, B_hat)
score_rmse = np.sqrt(
np.mean((model_branch_scores - true_branch_scores) ** 2, axis=1)
)
true_best_gain = Ks[np.argmax(true_branch_scores, axis=1)]
model_best_gain = Ks[np.argmax(model_branch_scores, axis=1)]k= 1: true-best K=1.50, model-best K=1.30, score RMSE=0.0516 k= 3: true-best K=1.55, model-best K=1.35, score RMSE=0.0915 k=12: true-best K=1.55, model-best K=1.35, score RMSE=0.1234

Every model branch advances from its own current model state for exactly steps; there is no temporally lagged "real" state and no periodic reset inside a branch. Common random numbers make comparisons insensitive to candidate-evaluation order and reduce Monte Carlo noise in model-versus-truth differences. In this seeded toy, the plotted score discrepancy grows over the displayed horizons and then begins to level as the stable closed-loop dynamics contract.
The printed gains are evaluation results for fixed controllers, not evidence about policy training. The example supports only the local conclusion that longer coherent branches expose cumulative controller scores to more consequences of this particular model bias. It does not establish that three steps are negligible, that a short branch always selects a controller closer to the true optimum, or that changing one bias parameter moves an optimal monotonically.
For integer branch lengths, a candidate should be compared with neighboring values and validated on the downstream task. Equality of marginal benefit and marginal cost is only an intuition for a differentiable continuous relaxation; boundary optima and non-monotone finite differences are possible.
Limitations and impact
These methods are best read as a comparison of design choices rather than a strict sequence in which each solves its predecessor's central problem.
PILCO demonstrated striking data efficiency with GP dynamics, analytic moment matching with repeated Gaussian approximation, and analytic policy gradients, but cubic GP scaling and restrictive analytic forms limit its practical range. Its expected-cost formulation and bounded objective remain useful examples, although direct historical inheritance by every later method should not be assumed.
PETS showed sample-efficient learned control on challenging simulated continuous-control benchmarks, including a contact-rich simulated 7-DOF Pusher task and a separate 7-DOF Reacher task. The paper did not report those experiments on physical robots. PETS pays for online planning with ensemble inference, particle rollouts, and CEM search. Particle count controls return-estimation precision, while action dimension and horizon primarily make action-sequence optimization harder; those are related compute costs, not the same phenomenon.
MBPO's published analysis proposes a lower-bound tradeoff between policy divergence and current-policy model error under truncated branches. Because a later reproduction disputes pre/post-branch indexing in the proof, its printed expression is a motivation for positive model use, not a settled performance certificate. The analysis does not establish that earlier work was generally misguided or prescribe a universally optimal . Its practical design introduces rollout-schedule and replay-mixture choices that require empirical validation.
Several recurring failure modes remain useful as a checklist:
- Distribution shift. Optimization can seek states poorly represented by model-training data. Coverage, regularization, constraints, explicit risk objectives, and real-state branch starts can reduce this risk without eliminating it.
- Accumulated rollout error. Imperfect open-loop learned-model predictions can diverge, saturate, oscillate, or cancel. Shorter horizons and new observations reduce consecutive reliance on them; an exact model would not drift.
- Uncertainty mis-estimation. Ensemble disagreement is an empirical proxy and can be miscalibrated. A collapsed total variance discards the within-/between-member distinction; retaining both terms permits an operational decomposition, but neither is a pure uncertainty category under misspecification.
- Objective and usage misspecification. A saturating cost and an explicit uncertainty penalty change the objective, while a short rollout limits model use. These are different mechanisms and can fail in different ways.
This list identifies common risks, not a complete safety certificate or a prediction that every controller exhibits all four. The methods advanced learned model-based RL on benchmark control problems and supplied mechanisms that recur in later work. Later latent-control methods may reuse short planning horizons, terminal values, ensembles, or uncertainty proxies in different combinations. For example, TD-MPC and Control-Centric Representations uses short latent trajectory optimization with a terminal value; that is not the same as MBPO's branched synthetic policy data. The following chapter, World Models, SimPLe, and PlaNet, compares distinct modeling choices, including compact latent dynamics and action-conditioned video prediction.
Summary
This chapter compared three decision-centric methods (PILCO, PETS, and MBPO) and deep ensembles, an uncertainty-estimation method used by PETS.
- PILCO. Gaussian-process dynamics, analytic moment integrals, repeated Gaussian approximation, and analytic gradients make data-efficient feedback-policy search possible. Modeling uncertainty can mitigate exploitation, but neither the GP assumptions nor expected-cost optimization guarantees it away.
- Deep ensembles and PETS. Ensemble disagreement is an empirical epistemic proxy. PETS evaluates each CEM candidate by averaging returns over state particles, with TS1 or TS controlling bootstrap assignments. That arithmetic mean is risk-neutral; explicit pessimism, variance penalties, lower-confidence scores, or CVaR are separate choices.
- Receding-horizon control. PETS executes one planned action and replans from the next observation. This limits commitment and the uncorrected forecast span, but it does not place a numerical bound on model error or guarantee global competence.
- MBPO. Short branches from real replay states generate data for an off-policy learner. The expression printed in Theorem 4.3 favors positive truncated rollouts when current-policy model error is sufficiently low, but a later reproduction disputes pre/post-branch indexing in its proof. Treat that expression as the paper's motivation, not a verified performance guarantee; actual return and the practical schedule still require empirical validation.
Uncertainty representations and trust controls should not be conflated. A GP posterior and an ensemble predictive distribution represent uncertainty under different assumptions. A short rollout does not represent uncertainty; it limits consecutive reliance on the model. Predictive fidelity, uncertainty quality, and decision usefulness are the main lenses developed here. Representation quality is introduced but deferred to the later latent-model chapters. These methods supplied influential mechanisms that recur in learned control, but they are not a single lineage or the foundation of every later world-model system.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about model-based reinforcement learning with PILCO, PETS, and MBPO.
PILCO, PETS, Ensembles, and MBPO
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!