World Models, SimPLe, and PlaNet

Michael BrenndoerferJuly 14, 202652 min read

Part of World Models Handbook

Compare World Models, SimPLe, and PlaNet: latent dynamics, video prediction, and CEM planning for model-based reinforcement learning.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

World Models, SimPLe, and PlaNet

Three other model-based reinforcement learning (MBRL) methods provide a useful contrast: they fit dynamics in observation space or in low-dimensional hand-designed state spaces, then used those dynamics for planning or for generating synthetic rollouts. PILCO learned probabilistic dynamics over compact, low-dimensional continuous state vectors. PETS used probabilistic neural-network ensembles to plan action sequences. MBPO used model ensembles for short synthetic rollouts that augment a model-free learner. Those methods share a common feature: their learned models predict task-defined state-vector transitions used for planning, model-based policy optimization, or generating policy-training data, rather than dynamics in an image-compression latent state. When a low-dimensional state vector is sufficiently observed and predictive for the task, direct state-space modeling and planning can be practical. When the observation is a 210-by-160-pixel Atari screen or a first-person 3D view, per-pixel modeling and planning become computationally demanding and can scale poorly.

This chapter compares three responses, each with a different place for prediction. World Models compresses images and predicts dynamics in that compressed representation. SimPLe instead rolls an observation-space video model forward; its discrete stochastic latent is an internal source of variation, not the compact state consumed by the policy. PlaNet learns a compact recurrent state-space model and plans through it online. Treating all three as latent-dynamics methods would hide the most useful contrast.

First posted to arXiv within roughly one year and addressing overlapping image-based model-learning problems, the papers are historically adjacent but conceptually distinct. All three use a learned predictive model to support decisions. They differ in what is predicted, what the decision mechanism consumes, and whether a controller is trained or actions are planned online. Reading them together shows why the model's coupling to a policy, planner, or value function matters: that coupling helps identify which predictions matter, alongside the task and the states, actions, and horizons the decision procedure queries. Low average pixel error does not by itself guarantee a useful controller, and relatively high pixel error does not by itself rule one out, because control depends on errors in the decision-relevant quantities and trajectories queried by the particular decision mechanism. The Dreamer family, which we examine next, pushes this insight further. Here we build the foundation.

To keep the concepts clean as we proceed, it helps to hold four separate notions of "model quality" in mind, because the rest of this chapter repeatedly depends on the distinction. Predictive fidelity asks how accurately the model forecasts its declared prediction targets, such as future observations, latent transitions, rewards, or termination and discount signals. Representation quality asks how well the latent organizes the world into useful, disentangled, or stable factors. Uncertainty quality asks whether predictive uncertainty is calibrated to the conditional variation remaining given the model's information (from stochastic transitions, observation noise, or unresolved hidden state), and whether epistemic uncertainty reflects gaps in the model's knowledge. Those are related but distinct questions. Decision usefulness asks whether the model supports good actions once it is plugged into a planner, a policy, or a value function. These four can be related, but they need not move together. Optimizing one does not ensure improvement in another. The three papers emphasize different predictive targets, representation choices, and decision couplings; none of them, by itself, establishes calibrated predictive or epistemic uncertainty.

Our plan: first, we develop variational latent dynamics in general and then the recurrent state-space model (RSSM) used by PlaNet, which gives us a concrete reference point for later latent models. Next, we look at controllers learned in dreams, contrasting the World Models training objective with the SimPLe procedure. Then we examine video prediction for Atari as a modeling problem, since SimPLe's learned environment is a video-prediction model and many of its model-related difficulties arise there. Finally, we look at latent planning, in particular PlaNet's CEM-based planner over the RSSM latent state, and explain why planning in latent space is not the same as planning in the true state.

Variational latent dynamics

A latent dynamics model introduces an unobserved or inferred state and models its action-conditioned evolution. World Models and PlaNet aim to make their learned state more compact than pixels or easier to use for prediction and control, but neither property is guaranteed merely by introducing a latent variable. Formally, we imagine a latent state ztz_t, an observation xtx_t, an action ata_t, and a reward rtr_t, and we model:

zt∼pθ(zt∣zt−1,at−1),xt∼pθ(xt∣zt),rt∼pθ(rt∣zt).z_{t} \sim p_\theta(z_t \mid z_{t-1}, a_{t-1}), \qquad x_t \sim p_\theta(x_t \mid z_t), \qquad r_t \sim p_\theta(r_t \mid z_t).

where:

  • ztz_t: the latent state at time tt, evolved by the transition model and, in planning-based agents, rolled forward to evaluate candidate actions
  • at−1a_{t-1}: the action taken at the previous step, which conditions how the latent changes
  • pθ(zt∣zt−1,at−1)p_\theta(z_t \mid z_{t-1}, a_{t-1}): the transition (prior) model, parameterized by θ\theta, giving the distribution over the next latent given the current latent and action
  • xtx_t: the observation at time tt (for example, an Atari frame)
  • pθ(xt∣zt)p_\theta(x_t \mid z_t): the observation (decoder) model, rendering an observation from the latent
  • rtr_t: the reward at time tt
  • pθ(rt∣zt)p_\theta(r_t \mid z_t): the reward model, predicting reward from the latent

The transition lets latent states evolve, the observation model renders observations, and the reward model predicts per-step rewards that accumulate into the return used for control. If we could reliably infer ztz_t from the observation history, predict its transitions over the planning horizon, and estimate rewards accurately enough to compare candidate returns, we could plan entirely in latent space: roll the latent forward under candidate actions, evaluate predicted rewards, and pick the best action sequence. The obstacle is that we never observe ztz_t. We need an inference procedure. In the end-to-end variational formulation developed here, the observation model and inference network are optimized jointly; staged or pretrained alternatives are also possible.

It is worth pausing on why this factorization can be a reasonable modeling assumption. In image-based control, many pixel configurations can share a control-relevant situation: appearance and background variation may add visual complexity without changing the underlying dynamics. A latent can compress some of that redundancy and expose dynamics that are easier to model. In an Atari game, visible sprite positions and velocities are useful factors, but a sufficient emulator state may also contain timers, mode flags, lives, off-screen objects, and pseudorandom state. World Models and PlaNet aim to retain variables needed for prediction and control in a representation more compact than the image; this does not imply that every image-based method here, including SimPLe, feeds a compact latent to its decision mechanism or that the hidden state is always a handful of coordinates.

This is the setting of a variational autoencoder extended over time. We introduced latent-variable representations and temporal state in Part III. Here we reconnect those ideas to dynamics and control. For nonlinear transition and decoder networks, the exact sequence posterior is generally intractable. Let TT be the sequence length, let hth_t be the deterministic state computed from the sampled latent and action history before observing xtx_t, and let qϕq_\phi be the filtering variational posterior with inference-network parameters ϕ\phi. The approximate sequence posterior then factorizes as

qϕ(z1:T∣x1:T,a1:T−1)=∏t=1Tqϕ(zt∣ht,xt).q_\phi(z_{1:T} \mid x_{1:T}, a_{1:T-1}) = \prod_{t=1}^{T} q_\phi(z_t \mid h_t, x_t).

Assume the generative joint factorizes as ∏t=1Tpθ(zt∣ht)pθ(xt∣ht,zt)pθ(rt∣ht,zt)\prod_{t=1}^{T}p_\theta(z_t\mid h_t)p_\theta(x_t\mid h_t,z_t)p_\theta(r_t\mid h_t,z_t), so observation and reward are conditionally independent given (ht,zt)(h_t,z_t), with the deterministic recursion ht=fθ(ht−1,zt−1,at−1)h_t=f_\theta(h_{t-1},z_{t-1},a_{t-1}) and fixed initial conditions. Under these assumptions, L(θ,ϕ)\mathcal{L}(\theta,\phi) below is a sequential ELBO: it lower-bounds the conditional log evidence log⁡pθ(x1:T,r1:T∣a1:T−1)\log p_\theta(x_{1:T},r_{1:T}\mid a_{1:T-1}).

L(θ,ϕ)=Eqϕ(z1:T∣x1:T,a1:T−1)[∑t=1T(log⁡pθ(xt∣ht,zt)+log⁡pθ(rt∣ht,zt)+log⁡pθ(zt∣ht)−log⁡qϕ(zt∣ht,xt))].\mathcal{L}(\theta, \phi) = \mathbb{E}_{q_\phi(z_{1:T} \mid x_{1:T}, a_{1:T-1})} \left[ \sum_{t=1}^{T} \left( \log p_\theta(x_t \mid h_t,z_t) + \log p_\theta(r_t \mid h_t,z_t) + \log p_\theta(z_t \mid h_t) - \log q_\phi(z_t \mid h_t,x_t) \right) \right].

The outer expectation is over the whole sequential latent trajectory, so it also averages over the random histories that determine later priors. The two likelihood terms encourage the inferred state to retain information useful for predicting observations and rewards. They do not guarantee interpretable factors or prevent collapse. The log-prior minus log-posterior terms are the negative KL contributions. They encourage posterior mass, on average, to remain in regions assigned substantial density by the transition prior. Without that agreement, prior outputs can drift away from the states seen by the decoder and reward head, making multi-step prior rollouts progressively unreliable.

There is a well-known tradeoff hiding in the KL term. If we weight reconstruction heavily, the model may allocate capacity according to pixel likelihood rather than task relevance, for example to background animation or texture. If we weight the prior too heavily, the latent may carry too little information and both observation and reward prediction can degrade. The β-VAE objective also changes the KL coefficient, and Burgess et al. analyze its capacity–reconstruction tradeoff from a rate–distortion perspective; the same coefficient-level reasoning applies to this sequential objective. PlaNet proposed latent overshooting as an additional multi-step regularizer: from an earlier posterior state, the transition is rolled forward for several steps and its multi-step predictive prior is matched by KL to later posterior beliefs. The released implementation treats those posterior targets as stopped-gradient references by default. Those extra terms encourage consistency at several prediction distances. The paper's reported final RSSM used three free nats and did not require overshooting; Appendix D reports a slight RSSM performance decrease when overshooting was added. The released configuration likewise defaults to three free nats and an overshooting loss scale of zero. These two settings agree between the final paper agent and released defaults, even though the paper's proposed overshooting objective was evaluated separately.

To make the tradeoff concrete, think of the KL term as a communication-channel budget. The posterior can carry information about the current observation beyond what the prior predicts, though it may also ignore that observation. Averaged over data, its KL to the transition prior upper-bounds this conditional information and also penalizes mismatch between the aggregate posterior and the prior. With natural logarithms, the KL is an expected log-density excess in nats; under a common discretization or an ideal coding construction, this corresponds to expected excess code length. Divide by ln⁡2\ln 2 to express nats in bits. Too little regularization can let the posterior encode detail that prior-only rollouts cannot reproduce. Too much can leave the state too vague for observation or reward prediction. The useful operating point depends on the task and on the relative loss weights.

A second design decision is how hth_t, the history summary, is formed. A feedforward latent derived from one frame is Markov only if that frame is predictively sufficient for action-conditioned future observations and rewards. This can fail when histories with the same current frame imply different futures—for example, if persistent velocity or hidden object state can be inferred from earlier frames. Chronologically, World Models came first, PlaNet followed, and SimPLe came later. Architecturally, they offered different answers: World Models encodes each frame and adds separate recurrent memory; PlaNet integrates deterministic recurrence with stochastic state in an RSSM; SimPLe predicts in observation space from four-frame inputs, while its selected video model also maintains an internal recurrent spatial state and a separate latent-bit LSTM (released configuration; video-model implementation). The SimPLe policy consumes frame stacks, not those internal model states. These papers belong to overlapping observation-space and latent-state research lines rather than a single progression. RSSM-style recurrent latent models became one influential modern lineage, alongside transformer-based observation-token models and other alternatives.

The single-frame Markov issue deserves emphasis because it can silently merge histories that share a frame but imply different future distributions. Consider a game in which two visually identical situations have different futures because one has an enemy moving toward the agent and the other has an enemy moving away. A model that sees only the current frame cannot distinguish them. If both histories have positive conditional probability under that frame and the same action, the true next-step predictive distribution mixes their possible outcomes; a misspecified model need not cover both. Recurrence can mitigate the ambiguity when the history contains enough identifying information: the recurrent state can accumulate evidence about which situation the agent is in, and the transition can condition on that summary. It does not solve partial observability by itself. Finite capacity, aliased histories, and imperfect optimization can still leave different situations indistinguishable. This resembles belief-state methods in partially observable control; PlaNet's RSSM supplies a learned approximate state summary rather than an exact belief.

Latent dynamics are a learned-model component of model-based RL

A latent dynamics model is a predictive component, not itself a planner or policy. An agent is conventionally model-based when model-predicted consequences score candidate action sequences or generate policy or value updates. PlaNet's CEM planner and World Models' dream-trained VizDoom controller meet this operational criterion. The World Models CarRacing controller was instead optimized on real-environment returns while consuming a reconstruction-trained VAE code and the hidden state of a separately trained predictive MDN-RNN; classifying that entire procedure solely by its feature source is taxonomy-dependent. The model predicts action consequences through an inferred representation rather than directly in observed state; World Models and PlaNet design that representation to compress images. The representation and dynamics may be trained jointly, as in PlaNet, or in stages, as in World Models. PlaNet's complete agent used fewer episodes than selected model-free baselines on its six evaluated control tasks; that result is not a guaranteed benefit of compression or either training schedule. Learned-model agents remain vulnerable to model exploitation, and learned representations and inference add potential failure modes. A policy or planner can favor actions whose predicted returns are spuriously high in unsupported regions.

From VAE to recurrent state-space model

PlaNet's RSSM is a variational latent dynamics architecture that Dreamer later reused, so we describe it carefully. The RSSM splits the latent state into two parts. First, the deterministic recurrent state hth_t summarizes the past. It follows the recurrence

ht=fθ(ht−1,zt−1,at−1),h_t = f_\theta(h_{t-1}, z_{t-1}, a_{t-1}),

where:

  • hth_t: the deterministic recurrent state at time tt, carrying memory of the past
  • fθf_\theta: the recurrent network (for example, a GRU) parameterized by θ\theta
  • ht−1h_{t-1}: the previous recurrent state
  • zt−1z_{t-1}: the previous stochastic latent state
  • at−1a_{t-1}: the action taken at the previous step

This component is deterministic and carries memory. Second, the stochastic state ztz_t is sampled from a diagonal Gaussian. It can represent per-step conditional variation, but the architecture does not require it to contain only information absent from hth_t.

Why split the latent at all, rather than using a single stochastic latent? The two paths provide different inductive biases. The finite recurrent state is a fixed-size learned summary of history that can retain useful temporal information without resampling every component at every step. It may discard task-relevant information when capacity or training is inadequate; for some tasks, a finite state can instead be sufficient. The stochastic state represents conditional variation through a distribution. A diagonal Gaussian supplies a mean and variance, but its presence does not by itself imply calibrated uncertainty, and a single conditional Gaussian latent density is unimodal. If calibrated predictive uncertainty is claimed, the induced predictive distribution should be evaluated against held-out outcomes under a stated calibration target and metric. Richer one-step multimodality in the latent transition requires a richer latent distribution; a nonlinear decoder can nevertheless induce a multimodal observation distribution, and multi-step propagation can produce complicated trajectory distributions.

The transition prior, conditioned on the history, is:

pθ(zt∣ht)=N ⁣(μθp(ht), diag⁡(σθp(ht)2)),p_\theta(z_t \mid h_t) = \mathcal{N}\!\big(\mu_\theta^p(h_t),\, \operatorname{diag}(\sigma_\theta^p(h_t)^2)\big),

where:

  • hth_t: the deterministic recurrent state summarizing the preceding inferred history and action at−1a_{t-1}, before the current observation xtx_t is incorporated
  • μθp(ht)\mu_\theta^p(h_t): the prior mean, a function of the history produced by the transition network
  • σθp(ht)2\sigma_\theta^p(h_t)^2: the vector of coordinatewise prior variances, also a function of the history
  • N(⋅,⋅)\mathcal{N}(\cdot, \cdot): a multivariate Gaussian specified by its mean and covariance matrix, diagonal here with coordinatewise variances

The approximate posterior, used during training and whenever we want to incorporate a new observation, is:

qϕ(zt∣ht,xt)=N ⁣(μϕq(ht,xt), diag⁡(σϕq(ht,xt)2)),q_\phi(z_t \mid h_t, x_t) = \mathcal{N}\!\big(\mu_\phi^q(h_t, x_t),\, \operatorname{diag}(\sigma_\phi^q(h_t, x_t)^2)\big),

where:

  • xtx_t: the current observation, the extra input the posterior receives
  • μϕq(ht,xt)\mu_\phi^q(h_t, x_t): the posterior mean, a function of both the history and the current observation
  • σϕq(ht,xt)2\sigma_\phi^q(h_t, x_t)^2: the vector of coordinatewise posterior variances

The posterior sees the current observation, the prior does not. During imagination, we sample ztz_t from the prior; during inference from real observations we may sample from the posterior to stay grounded. The KL term now compares two Gaussians over the stochastic state, and because both are diagonal, it has a closed form:

DKL ⁣(q ∥ p)=12∑k=1K(σq,k2+(μq,k−μp,k)2σp,k2−1+log⁡σp,k2−log⁡σq,k2)\begin{aligned} D_{\mathrm{KL}}\!\big( q \,\|\, p \big) &= \frac{1}{2} \sum_{k=1}^{K} \Big( \frac{\sigma_{q,k}^2 + (\mu_{q,k} - \mu_{p,k})^2}{\sigma_{p,k}^2} - 1 + \log \sigma_{p,k}^2 - \log \sigma_{q,k}^2 \Big) \end{aligned}

where:

  • qq: the approximate posterior Gaussian over the stochastic state (mean μq\mu_q, variance σq2\sigma_q^2)
  • pp: the transition prior Gaussian over the stochastic state (mean μp\mu_p, variance σp2\sigma_p^2)
  • KK: the dimensionality of the stochastic latent ztz_t
  • μq,k,σq,k2\mu_{q,k}, \sigma_{q,k}^2: the posterior mean and variance of the kk-th latent dimension
  • μp,k,σp,k2\mu_{p,k}, \sigma_{p,k}^2: the prior mean and variance of the kk-th latent dimension

This expression measures, dimension by dimension, how much the posterior differs from the prior. Holding the positive prior variance σp,k2\sigma_{p,k}^2 fixed, the mean-mismatch contribution grows quadratically with ∣μq,k−μp,k∣|\mu_{q,k}-\mu_{p,k}|. The variance ratio and logarithm must be read together: their contribution is minimized when the posterior and prior variances match and increases when the posterior variance is either larger or smaller. Because both distributions are diagonal, the total separates into per-dimension contributions.

The closed form matters practically as well as conceptually. The analytic Gaussian KL is differentiable and cheap to compute, but that does not imply a gradient everywhere in a thresholded training objective. The released PlaNet objective subtracts three free nats and clips the divergence penalty at zero. Where the standard one-step penalty has a nonzero gradient, it trains the prior toward the posterior and regularizes the posterior toward the prior; below the free-nats threshold, it supplies no such gradient. Optional multi-step overshooting instead uses stopped-gradient posterior targets and is disabled by default in the released configuration. Active KL pressure can encourage prior-only rollouts (which planning uses) to remain closer to posterior-inferred latents seen during training, but it does not prevent multi-step divergence.

Two structural advantages follow. First, because hth_t is deterministic and recurrent, the model need not be Markov in the current frame alone: it can retain evidence about velocity or partially observed variables when the history and learned capacity are sufficient. Second, because ztz_t is stochastic, the model can represent conditional variation rather than being restricted to one deterministic next state. That is not the same as representing arbitrary multimodality or calibrated epistemic uncertainty. The combination of a deterministic memory path and a stochastic path recurs in later RSSM-based world models, including the Dreamer family.

In PlaNet, the observation and reward heads consume the combined state [ht,zt][h_t,z_t]. The paper evaluated six pixel-based continuous-control tasks from the DeepMind Control Suite, using 64×6464\times64 RGB observations and scalar rewards (Hafner et al., 2019). The observation decoder is not needed inside candidate rollouts at control time: CEM advances the prior and queries predicted rewards without rendering future frames. The encoder and posterior incorporate the retained real observation at each decision step, after the task's configured action-repeat interval.

It is worth stating explicitly what these components do. The recurrent state hth_t summarizes the preceding inferred history through zt−1z_{t-1} and at−1a_{t-1}, and parameterizes the time-tt prior pθ(zt∣ht)p_\theta(z_t\mid h_t) before xtx_t is incorporated. The stochastic state ztz_t supplies the sampled per-step component. PlaNet's observation and reward heads use their combination; CEM consumes the returns produced by rolling that combined state forward rather than passing ztz_t through a learned controller. The decoded observation supplies a training signal but is not rendered during planning.

Controller learning in dreams

Once latent dynamics are combined with models of the task's reward and termination or discount, they can provide a learned simulator for control. The next question is how to couple that simulator to a decision mechanism. The chapter contrasts three answers: World Models and SimPLe train controllers or policies, while PlaNet uses online planning instead of training a separate policy network.

The first answer, used by World Models, is to train a compact controller separately from the vision and memory modules. A VAE compresses each 64×6464\times64 frame into ztz_t: the reported latent dimension was 32 for CarRacing and 64 for the VizDoom dream-training experiment. Given (zt,at,ht)(z_t,a_t,h_t), an MDN-RNN predicts a distribution over zt+1z_{t+1} and advances its recurrent state to ht+1h_{t+1}. The paper's main controller equation is an affine map of [zt,ht][z_t,h_t], while Appendix A.3's Doom-specific controller description specifies [zt,ct,ht][z_t,c_t,h_t], including the LSTM cell state. The released VizDoom code uses those inputs through a weight matrix without an explicit learned bias, followed by tanh and a task-specific action interpretation (state construction; controller implementation). Concentrating most complexity in the vision and memory modules keeps the controller's CMA-ES search space small.

CMA-ES treats return as a black box: it probes controller parameters with a population of perturbations and updates a search distribution from the better candidates. That made it practical for the small controller, but it did not guarantee protection from model exploitation. The paper documents controllers exploiting imperfections in the learned VizDoom dream and tests increased sampling temperature in the MDN-RNN as a mitigation. Controller size and gradient-free optimization are capacity and search choices, not demonstrated safeguards.

The second answer, used by SimPLe, is to train a policy inside a learned observation-space simulator. SimPLe stands for "Simulated Policy Learning." Its pipeline alternates between collecting real Atari interaction, fitting the video-and-reward model, and updating PPO mostly from simulated interactions. Real frames train and refresh the model and provide restart states. Policy learning is dominated by simulated rollouts, although the reported implementation also applied PPO directly to the much smaller real-data stream.

SimPLe uses a video model whose simulated environment emits visual observations to PPO; unlike World Models' controller inputs and PlaNet's planner state, its policy does not consume the model's internal latent state. Because the model predicts pixels directly, the policy can be trained on the same kinds of inputs a real agent would see, so the transfer from imagination to reality is at least in principle direct. Predicting a high-dimensional per-pixel distribution is computationally expensive; when the training objective aggregates similarly weighted per-pixel losses, large visual regions can dominate the gradient. Its clipped visual loss suppresses gradients from already-confident pixels, while a discrete stochastic latent can represent sharp conditional alternatives instead of averaging them into blur; neither removes the cost of pixel-space prediction. Multi-step rollouts can still accumulate visual artifacts and drift, so the training loop keeps imagined horizons short and re-trains the world model in each of its 15 data-collection iterations, after each 6,400-interaction batch.

PlaNet takes a different route: it trains no separate policy network and plans online in latent space with model-predictive control. Dreamer, in the next chapter, returns to policy learning and trains actor and value networks in imagination.

What unifies the three is that decisions are mediated by a learned predictive model, so decision quality can be affected by errors in the predictions or learned features that the policy or planner actually uses. SimPLe trains its policy in the learned simulator. World Models did so for VizDoom, whereas its CarRacing controller was optimized through real-environment rollouts while consuming learned VAE and RNN features. PlaNet trains no separate policy network and instead searches candidate actions online. Shorter rollouts and repeated grounding in real observations can limit open-loop drift, but they do not make transition or reward rankings correct. World Models additionally tested model temperature against dream exploitation. None of these devices is a guarantee.

These choices have different benefits and limitations; they need not be opposed. SimPLe, for example, combined short simulated rollouts with restarts from real buffered states. Shorter horizons can make the learning or planning target myopic. Restricting the controller to a smaller function class can reduce expressiveness, although the smaller search space may be easier to optimize. Restarting finite-horizon simulated rollouts from real buffered states limits each trajectory to the chosen number of model steps from a data state; it does not restrict the model itself. The three papers choose different points in this design space.

Why learned dreams differ from real interaction

Imagined rollouts differ from real rollout data because their transitions come from a fitted model rather than the environment. Errors can be correlated along the rollout: if the model under-renders a moving sprite, that error can affect later inputs and predictions. SimPLe used a default simulated-rollout horizon of 50 model steps, restarting from four-frame observation stacks sampled from its real-data buffer; it ablated horizons of 25, 50, and 100 steps and ran the outer training loop for 15 iterations (Kaiser et al., 2019). A model can therefore fail locally in a region the improving policy increasingly visits even when its average prediction error looks acceptable.

This policy–model feedback loop is one mechanism behind failures of learning in dreams. As a policy changes, its state-occupancy distribution can also change. It may remain in well-modeled regions, generalize into sparsely covered regions, or seek states where prediction errors make return look spuriously high. More frequent visits to affected states increase their weight in on-policy model-error averages, all else equal, and create more opportunities for those errors to influence decisions; the resulting control harm need not grow monotonically. If the model is not refreshed, the policy may continue exploiting the same unsupported region while the error there remains untested. Periodic refresh with recent real data can interrupt or reduce this feedback loop by exposing the model to states the current policy visits, but finite data, capacity, and optimization mean that accurate coverage is not guaranteed.

Learning in dreams

Learning in dreams means updating a policy or value function from model-generated trajectories rather than environment transitions. Online planning through a model also uses imagined trajectories, but the planner is searching rather than being trained on them. Model exploitation and rollout drift remain possible. Short rollouts and real-state restarts limit the number of consecutive model-generated steps from a buffered real state and can reduce compounding error; they do not guarantee geometric or distributional proximity to observed data. Targeted regularization may discourage particular exploitative behavior, and model refresh can improve coverage where new real data are collected. None guarantees accurate return rankings or prevents exploitation.

Video prediction for Atari

SimPLe's world model is a video-prediction model, and to understand why it is hard and what it teaches, we should look at the objective directly. Given past frames x1:tx_{1:t}, past actions, and candidate future actions, an action-conditioned predictor models

p(xt+1:T∣x1:t,a1:T−1).p(x_{t+1:T} \mid x_{1:t}, a_{1:T-1}).

The future action sequence at:T−1a_{t:T-1} matters because later frames depend on actions that have not yet occurred at time tt. Accurate pixel prediction requires accounting for the pixel-level effects of motion, appearance, interactions, occlusions, and actions that affect the target frames; useful control may depend on only a subset of those effects. Some factors are visually dominant but irrelevant to the current decision; others are small but useful. An on-screen score can be a useful visual cue, for example, but the environment still supplies reward as a separate variable. The pixels and reward should not be conflated.

The mismatch between what pixel loss rewards and what control needs is a central tension in video-model world models. Pixel loss is an imperfect proxy for decision-relevant modeling: optimizing it can favor visually dominant but control-irrelevant features. A model that reproduces the background perfectly but loses the agent, hazards, and their motion can achieve low pixel error while being useless for control. Conversely, a model with crude rendering that preserves the controllable objects and their interactions can be more useful for control than its pixel error suggests. SimPLe's primary policy evaluation reported game score and sample efficiency (Kaiser et al., 2019). Those test the learned policy's behavior, whereas pixel reconstruction error tests a different property of the model.

SimPLe's selected stochastic-discrete predictor was conditioned on four stacked frames and an action. It combined a convolutional video model with recurrent internal state across predicted steps (released configuration; video-model implementation). A separate auxiliary autoregressive LSTM generated its discrete binary latent bits within each step. The model downscaled Atari frames to 105×80105\times80 and used a 256-way categorical distribution for each output pixel channel. That per-pixel categorical likelihood is separate from the stochastic latent bits. It changes pixel regression into classification. Independent per-pixel categoricals can be multimodal at individual pixels, but their factorized joint cannot capture cross-pixel correlations needed to coordinate coherent alternative frames; the shared stochastic latent provides a mechanism for that coordination (Kaiser et al., 2019).

There are design decisions worth naming because they recur in later video-model world models. The first is input conditioning. An action-conditioned predictor needs the candidate action for each modeled transition; past actions can also help infer hidden state when the retained observations are insufficient. SimPLe feeds each next-frame prediction the action applied at that step, so its autoregressive rollout is conditioned step by step on the chosen action sequence. The second is stochasticity handling. The Atari emulator's core transition is deterministic given its complete state and applied action (Machado et al., 2018). ALE's optional sticky actions add stochasticity relative to the commanded action. Four observed frames need not reveal the complete emulator state. When the retained frame and action history aliases complete emulator states (for example, because relevant variables remain hidden, occluded, or affected by Atari's flicker), different futures may remain possible under the same candidate action, leaving conditional uncertainty for the model. SimPLe includes a discrete stochastic latent to represent conditional alternatives. The third is frame stacking and resolution. SimPLe worked at reduced resolution and stacked frames so the model could see motion (Kaiser et al., 2019). The fourth is evaluation for decisions versus evaluation for pixels. Pixel-averaged reconstruction error can look favorable when a model reproduces a large static background but misses a small control-critical object. The paper also reported model and reward-prediction diagnostics. Neither those diagnostics nor reconstruction quality alone establishes control utility.

A useful way to reason about video prediction is to keep three quantities separate. Let ct=(x1:t,a1:t)c_t=(x_{1:t},a_{1:t}) denote the available context for a one-step prediction.

  • Irreducible conditional uncertainty under log loss is the entropy H[p(⋅∣ct)]=Ep[−log⁡p(xt+1∣ct)]H[p(\cdot\mid c_t)]=\mathbb{E}_{p}[-\log p(x_{t+1}\mid c_t)].
  • Fitted-model discrepancy can be defined as DKL(p(⋅∣ct) ∥ p^θ(⋅∣ct))D_{\mathrm{KL}}(p(\cdot\mid c_t)\,\|\,\hat p_\theta(\cdot\mid c_t)), averaged over contexts. The model's cross-entropy is their sum: H(p,p^θ)=H(p)+DKL(p ∥ p^θ)H(p,\hat p_\theta)=H(p)+D_{\mathrm{KL}}(p\,\|\,\hat p_\theta). Representative data, adequate capacity, and successful optimization can reduce the second term; more data alone does not guarantee that it shrinks.
  • Cumulative rollout error is a chosen trajectory metric, for example ET=∑j=1TE[∥xt+j−x^t+j∥]E_T=\sum_{j=1}^{T}\mathbb{E}[\|x_{t+j}-\hat x_{t+j}\|] under a specified coupling of true and model trajectories. Under a fixed action sequence, if the transition is uniformly LL-contractive with L<1L<1 in the chosen metric, true and model randomness are coupled consistently, and one-step discrepancy is at most ϵ\epsilon, then the per-step error is bounded by Lje0+ϵ(1−Lj)/(1−L)L^j e_0+\epsilon(1-L^j)/(1-L); this prevents exponential amplification and gives a uniform per-step bound. The partial sums ETE_T remain uniformly bounded as T→∞T\to\infty only when the nonnegative per-step errors are summable; persistent error can instead make ETE_T grow linearly. There is no universal theorem that it must grow superlinearly.

Separating these quantities points to different interventions. A richer stochastic predictor can represent conditional variation, but the distribution still has to be checked for calibration. Better coverage and modeling can reduce fitted-model discrepancy on the contexts represented in the data. Shorter rollouts and restarts from real buffered observation stacks reduce how long prediction errors feed back as inputs. SimPLe uses all three ideas, but the existence of a stochastic latent does not prove that its uncertainty is calibrated.

SimPLe limits compounding error by restarting simulated trajectories from real buffered frames after a finite horizon. MBPO starts short model rollouts from real replay states. Dreamer starts imagined trajectories from posterior states inferred from experience. These are related ways to limit reliance on unsupported long rollouts, not evidence of a direct lineage from SimPLe.

SimPLe's procedure alternates real-data collection, model fitting, and policy updates dominated by experience from the learned simulator; it also applies a much smaller set of PPO updates from collected real trajectories. This shared loop does not imply a one-way architectural lineage to the later methods discussed above. Its evidence remains specifically about data-efficient learning on Atari with an observation-space video model.

Latent planning with learned state-space models

Now we turn to PlaNet's online-planning approach, where its RSSM supports fast latent-space rollouts. Given a learned latent dynamics model, an inferred current latent state, and a reward predictor, we can evaluate candidate action sequences by imagining their latent trajectories and summing predicted rewards. We then execute the first planned action for the task's configured action-repeat interval, observe the resulting real frame, infer a new latent, and repeat. This is model-predictive control (MPC) in latent space, and it is the direct latent-space counterpart of the sampling-based MPC we discussed in Part VII, Chapter 1.

PlaNet's planner uses the cross-entropy method (CEM), a sampling-and-elite-refitting optimization procedure (Rubinstein, 1997). The following operational summary combines PlaNet's Algorithm 1 control loop and supplementary Algorithm 2 CEM planner with the released implementation's action clipping. At each time step:

  1. Infer the current latent state from the observation history using the posterior.
  2. Sample NN candidate action sequences of horizon HH from an initial Gaussian proposal, then clip each sampled action to the environment's valid range before evaluating it.
  3. For each candidate, roll the latent forward using the prior transition (no observations available for future steps), accumulating predicted reward from the reward head.
  4. Keep the best KK elites, refit the Gaussian proposal to their bounded action sequences, and repeat for a few iterations.
  5. Select the first action vector of the final proposal mean; during data collection add the configured exploration noise, then apply the bounded action for the task's action-repeat interval, observe, and re-infer the latent.

Each of these steps has a rationale worth understanding. Posterior inference incorporates the latest real observation before planning. Sampling lets the planner search a non-convex return surface, while prior rollouts are necessary because future observations are unavailable. The elite refit concentrates the proposal on higher-scoring action sequences. Re-observing and replanning after each decision-level action, after its configured action-repeat interval, limits open-loop error accumulation, but it does not repair a transition or reward model that ranks the local candidates incorrectly.

One conditional planning objective, after fixing a sample ztz_t from the current posterior, is:

arg⁡max⁡at:t+H−1  Epθ(zt+1:t+H∣ht,zt,at:t+H−1)[∑j=1Hrθ(ht+j,zt+j)],\arg\max_{a_{t:t+H-1}} \; \mathbb{E}_{p_\theta(z_{t+1:t+H}\mid h_t,z_t,a_{t:t+H-1})} \left[ \sum_{j=1}^{H} r_\theta(h_{t+j},z_{t+j}) \right],

where:

  • at:t+H−1a_{t:t+H-1}: the candidate action sequence of horizon HH that the planner is choosing
  • HH: the planning horizon, the number of future steps the planner looks ahead
  • rθ(ht+j,zt+j)r_\theta(h_{t+j},z_{t+j}): the predicted reward after action at+j−1a_{t+j-1} has advanced the model state
  • ht+j=fθ(ht+j−1,zt+j−1,at+j−1)h_{t+j}=f_\theta(h_{t+j-1},z_{t+j-1},a_{t+j-1}): the deterministic recurrent update
  • pθ(zt+1:t+H∣ht,zt,at:t+H−1)p_\theta(z_{t+1:t+H}\mid h_t,z_t,a_{t:t+H-1}): the joint action-conditioned rollout distribution, equal to the product of the sequential priors pθ(zt+j∣ht+j)p_\theta(z_{t+j}\mid h_{t+j})
  • E[⋅]\mathbb{E}[\cdot]: expected return over that joint stochastic trajectory, conditional on the sampled current state and the full candidate action sequence

The objective aligns action and reward timing: every candidate action affects one included next-state reward, including the final action. Monte Carlo rollouts can estimate the expectation. PlaNet's published Algorithm 2 samples a current posterior state for each candidate trajectory, so its candidate-trajectory distribution includes current-state uncertainty as well as future transition uncertainty. It uses one sampled trajectory per candidate action sequence rather than averaging repeated trajectories for the same candidate; the CEM population supplies the broader sample set. The released agent instead obtains one posterior state per real decision, and the planner implementation tiles that state across its CEM candidates. Its future transition samples can still differ across candidates. This paper-versus-code distinction does not change the receding-horizon control loop, but it matters for which uncertainty is sampled within one planning call.

Two points matter in latent planning. The first is a risk shared by any planner using approximate dynamics; the second is specific to learned latent representations and inference.

First, the planner's return rankings depend in part on how the model extrapolates along candidate action sequences. If an action sequence drives the latent into a region with no training support and receives a spuriously high predicted return relative to alternatives, the planner may prefer it over better real-world sequences. This is model exploitation: short planning horizons and re-planning after every step limit exposure to extrapolation regions, though they cannot eliminate it.

Second, PlaNet plans in a learned latent trained jointly for observation reconstruction, reward prediction, and dynamics. This training may preserve decision-relevant variables, but it does not guarantee that they are emphasized over visual nuisance information. PlaNet does not render candidate future frames for planning. For accurate reward and transition predictions, the complete recurrent state [ht,zt][h_t,z_t] must retain the information those models require, though it may also retain nuisance information. If that complete inferred state omits a task-relevant variable, more search over it cannot recover the missing information.

The contrast with MPC over an explicitly specified state-space model is instructive. Classical MPC can also face model mismatch, disturbances, state-estimation error, constraints, and optimization error (Rawlings et al., 2017). Learned latent MPC additionally introduces potential representation error, while approximation in its learned encoder or posterior is a source of state-estimation error. Replanning after each observation incorporates new evidence and limits commitment to an open-loop prediction; it does not guarantee recovery of the true state under partial observability or encoder error (PlaNet).

Latent MPC versus direct policy learning

Latent MPC plans online with a model and does not require a separately trained policy network. Direct policy learning amortizes decision computation into a mapping from method-specific inputs to actions: latent features for World Models and Dreamer, frame stacks for SimPLe. PlaNet's CEM-based MPC pays online candidate-search cost and can re-anchor its search after each observation. In the PlaNet–Dreamer-style comparison, an amortized policy replaces candidate search with one policy evaluation and is typically cheaper online, though relative cost depends on model and policy size, planning budget, and hardware. Horizon sensitivity and long-horizon performance vary by task: Dreamer's evaluation found that its value model reduced sensitivity to imagination horizon, while its PlaNet-planner and reward-only-actor comparisons were shortsighted on acrobot and hopper but succeeded on more reactive walker tasks. This does not establish a general robustness ordering. PlaNet sits on the planning end, while Dreamer sits on the policy-learning end.

Comparing the three

Putting the three side by side clarifies what each contributed and what each does at test time:

  • World Models: A VAE encodes each frame, an MDN-RNN predicts compressed dynamics and updates memory, and a tiny controller acts on latent and recurrent features. The paper's main equation uses [zt,ht][z_t,h_t] plus a bias; Appendix A.3's Doom-specific controller description and the released VizDoom state construction use [zt,ct,ht][z_t,c_t,h_t]. The released controller has no explicit learned bias. CarRacing controller optimization used the real environment; the VizDoom controller was trained in the learned dream. At deployment, the VAE and MDN-RNN remain active because they supply the controller inputs.
  • SimPLe: An observation-space video model predicts future frames and rewards from recent frames and actions, using a discrete stochastic latent internally. The policy is trained with PPO primarily on rollouts from the simulated environment; collected real-environment trajectories also supply a much smaller set of PPO updates. At execution, the policy consumes real frame observations and the video model is not needed.
  • PlaNet: An RSSM with deterministic and stochastic state components, plus observation and reward heads. Online CEM searches action sequences by rolling the latent transition and reward models forward. At each real decision step, the encoder and posterior incorporate the new observation before replanning. Each chosen action is then applied for the task's configured action-repeat interval. The observation decoder is not used inside planning, and no separate policy or value network is trained.

These are three different couplings: learned compressed features and memory feeding an evolved controller; a refreshed observation-space simulator supplying PPO experience; and online action-sequence search scored by a latent model. Each coupling helps identify which model errors affect decisions on the states, actions, and horizons actually queried.

World Models uses its memory in two roles: to generate VizDoom dream transitions during controller training and to provide recurrent features to the deployed controller. The paper's general controller equation uses hth_t, but Appendix A.3's Doom-specific controller description already includes both ctc_t and hth_t; the released VizDoom state construction follows that task-specific input choice. SimPLe needs video rollouts accurate enough that a policy trained in the simulator transfers to Atari. PlaNet needs sufficiently accurate multi-step return rankings across its planning horizon. Latent overshooting directly encourages multi-step latent-prior consistency against later posteriors and can support return ranking only indirectly. The released configuration defaulted its scale to zero. Receding-horizon execution limits how long PlaNet commits to a candidate sequence, but choosing the first action still depends on all rewards predicted across that sequence.

Worked example

To make the ideas concrete without needing an Atari emulator, let us build a small variational latent dynamics model and compare two decision couplings: a tiny controller optimized through imagined rollouts and a CEM planner that searches the same learned prior online. SimPLe's alternating observation-space training loop is not reproduced by this example. The synthetic environment exposes its full two-dimensional state as the observation, so this is a dynamics example rather than a demonstration of partial observability. It is illustrative, not a benchmark.

We start with the shared plotting setup, then define the synthetic world.

The synthetic world is a controlled damped linear system. Its fully observed state contains position and velocity, and reward peaks only when both are near the origin. The system is already passively stable, so a no-action baseline is essential: a learned controller should be credited only for improvement beyond that contraction.

In[3]:
Code
import numpy as np
import torch
import torch.nn.functional as F
from torch import nn

torch.manual_seed(7)
rng = np.random.default_rng(7)

## Fully observed state s = [position, velocity]; action a is a 1-D acceleration.
A = np.array([[0.95, 0.10], [0.00, 0.90]])
B = np.array([[0.0], [0.10]])
ACTION_LIMIT = 1.0


def step_world(s, a):
    """Advance and return the next fully observed state."""
    s_next = A @ s + B.flatten() * a
    return s_next


def reward_of(s):
    """A smooth reward peaking at the origin; asks the controller to stabilize."""
    return float(np.exp(-0.5 * np.sum(s**2)))

We now generate a training batch of short trajectories with exploratory actions and record transitions of the form (observation, action, next observation, reward). Exploration draws from an unbounded Gaussian. ACTION_LIMIT does not clip these training actions. It bounds only the tiny controller and CEM candidates later. The learned model therefore sees some action magnitudes outside their control range.

In[4]:
Code
def collect_data(n_episodes=200, horizon=40):
    obs, act, next_obs, rew = [], [], [], []
    for _ in range(n_episodes):
        s = rng.normal(scale=0.6, size=2)
        for _ in range(horizon):
            a = float(rng.normal(scale=0.8))
            s_next = step_world(s, a)
            obs.append(s.copy())
            act.append([a])
            next_obs.append(s_next.copy())
            rew.append([reward_of(s_next)])
            s = s_next
    return (
        np.asarray(obs, dtype=np.float32),
        np.asarray(act, dtype=np.float32),
        np.asarray(next_obs, dtype=np.float32),
        np.asarray(rew, dtype=np.float32),
    )


obs, act, next_obs, rew = collect_data()
print(
    f"Collected {obs.shape[0]} transitions, observation dimension {obs.shape[1]}"
)

We can inspect the collected transitions before training. Because reward depends jointly on next position and next velocity, color is a more faithful encoding than a curve obtained by sorting one coordinate.

Out[6]:
Visualization
Dense scatter of next-state position against velocity, colored from low reward away from the origin to high reward near the origin.
State coverage in the exploratory dataset. Each point is an observed next state and its color is the aligned reward, which is largest when both position and velocity are near zero.

With a two-dimensional observed state, we do not need a VAE for compression, but the same probabilistic interfaces make the temporal targets explicit. The encoder maps each state to a Gaussian latent. The prior maps (zt,at)(z_t,a_t) to a distribution over zt+1z_{t+1}. The decoder maps latents back to states. During training, the reward head receives samples from the observation-conditioned next-state posterior. During held-out one-step evaluation, imagined-controller optimization, and CEM planning, each reward-head query receives a transitioned prior mean instead. The deployed tiny controller selects actions from the encoded current state and does not itself query the reward head. The KL term encourages the next-posterior and transition-prior distributions to agree, but it does not make posterior samples equal to the prior mean used by these evaluation and planning paths. This toy setup therefore has a train-use shift in reward prediction. For image observations, SimPLe and PlaNet use convolutional encoders or decoders to exploit spatial structure. When the current observation omits decision-relevant information that can be recovered from history, a temporal state can help retain it. The small fully connected networks below are only for this toy system.

In[7]:
Code
class Encoder(nn.Module):
    def __init__(self, obs_dim, latent_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(obs_dim, 32), nn.ReLU(), nn.Linear(32, 2 * latent_dim)
        )

    def forward(self, x):
        mu, log_var = self.net(x).chunk(2, dim=-1)
        return mu, log_var.clamp(-6, 2)


class Prior(nn.Module):
    def __init__(self, latent_dim, act_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(latent_dim + act_dim, 32),
            nn.ReLU(),
            nn.Linear(32, 2 * latent_dim),
        )

    def forward(self, z, a):
        mu, log_var = self.net(torch.cat([z, a], dim=-1)).chunk(2, dim=-1)
        return mu, log_var.clamp(-6, 2)


class Decoder(nn.Module):
    def __init__(self, latent_dim, obs_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(latent_dim, 32), nn.ReLU(), nn.Linear(32, obs_dim)
        )

    def forward(self, z):
        return self.net(z)


class RewardHead(nn.Module):
    def __init__(self, latent_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(latent_dim, 32), nn.ReLU(), nn.Linear(32, 1)
        )

    def forward(self, z):
        return self.net(z)


def reparameterize(mu, log_var):
    sigma = torch.exp(0.5 * log_var)
    eps = torch.randn_like(sigma)
    return mu + sigma * eps


def gaussian_kl(mu_q, log_var_q, mu_p, log_var_p):
    var_q, var_p = torch.exp(log_var_q), torch.exp(log_var_p)
    return 0.5 * torch.sum(
        (var_q + (mu_q - mu_p) ** 2) / var_p - 1.0 + log_var_p - log_var_q,
        dim=-1,
    )


latent_dim = 3
encoder = Encoder(obs.shape[1], latent_dim)
prior = Prior(latent_dim, act.shape[1])
decoder = Decoder(latent_dim, obs.shape[1])
reward_head = RewardHead(latent_dim)

params = (
    list(encoder.parameters())
    + list(prior.parameters())
    + list(decoder.parameters())
    + list(reward_head.parameters())
)
optimizer = torch.optim.Adam(params, lr=3e-3)

obs_t = torch.from_numpy(obs)
act_t = torch.from_numpy(act)
next_obs_t = torch.from_numpy(next_obs)
rew_t = torch.from_numpy(rew)

## Held-out split for honest evaluation.
n_train = int(0.9 * obs_t.shape[0])
train_idx = torch.arange(n_train)
test_idx = torch.arange(n_train, obs_t.shape[0])

Each transition supplies both current and next posterior targets. We reconstruct xtx_t from q(zt∣xt)q(z_t\mid x_t) and xt+1x_{t+1} from q(zt+1∣xt+1)q(z_{t+1}\mid x_{t+1}), match the transition prior p(zt+1∣zt,at)p(z_{t+1}\mid z_t,a_t) to the next posterior, and predict rt+1r_{t+1} from zt+1z_{t+1}. An additional prior-decoding loss compares the transitioned prior mean with xt+1x_{t+1}; this is an auxiliary one-step prediction term rather than a separate ELBO likelihood.

In[8]:
Code
def train_step(batch_size=256):
    idx = train_idx[torch.randint(len(train_idx), (batch_size,))]
    x = obs_t[idx]
    a = act_t[idx]
    x_next = next_obs_t[idx]
    r_next = rew_t[idx]

    mu_q, log_var_q = encoder(x)
    mu_q_next, log_var_q_next = encoder(x_next)
    z = reparameterize(mu_q, log_var_q)
    z_next = reparameterize(mu_q_next, log_var_q_next)
    mu_p_next, log_var_p_next = prior(z, a)

    current_recon = F.mse_loss(decoder(z), x)
    next_recon = F.mse_loss(decoder(z_next), x_next)
    prior_next_recon = F.mse_loss(decoder(mu_p_next), x_next)
    reward_loss = F.mse_loss(reward_head(z_next), r_next)
    kl = gaussian_kl(
        mu_q_next, log_var_q_next, mu_p_next, log_var_p_next
    ).mean()

    loss = (
        current_recon + next_recon + prior_next_recon + reward_loss + 0.1 * kl
    )
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    return (
        float(loss.detach()),
        float(current_recon.detach()),
        float(prior_next_recon.detach()),
        float(reward_loss.detach()),
        float(kl.detach()),
    )


history = []
for step in range(1500):
    loss, recon, pred, rloss, kl = train_step()
    if step % 250 == 0:
        history.append((step, loss, recon, pred, rloss, kl))

We can also evaluate the prior's one-step prediction error on held-out transitions, which is the first hint at whether planning in this model could work at all.

Out[9]:
Console
Training trace (step, loss, current recon, prior next, reward, kl):
  step    0: loss=1.9210 recon=0.1665 prior_next=0.1251 reward=1.4458 kl=0.2322
  step  250: loss=0.1106 recon=0.0254 prior_next=0.0215 reward=0.0033 kl=0.3922
  step  500: loss=0.0700 recon=0.0021 prior_next=0.0021 reward=0.0014 kl=0.6216
  step  750: loss=0.0652 recon=0.0012 prior_next=0.0010 reward=0.0006 kl=0.6131
  step 1000: loss=0.0731 recon=0.0011 prior_next=0.0010 reward=0.0003 kl=0.6970
  step 1250: loss=0.0655 recon=0.0009 prior_next=0.0007 reward=0.0003 kl=0.6263
Held-out prior next-state MSE:  0.0003
Held-out prior next-reward MSE: 0.0005

After 1,500 optimizer updates, the held-out metrics above test one-step prior predictions rather than posterior reconstructions: the prior is invoked with (zt,at)(z_t,a_t), its next latent is decoded against xt+1x_{t+1}, and the same next latent predicts rt+1r_{t+1}. These checks do not establish long-horizon accuracy, but they provide a basic sanity check before control; multi-step and closed-loop evaluation are still needed to assess decision usefulness.

First, we optimize a tiny tanh-squashed affine controller through imagined rollouts generated by the trained prior. We backpropagate through deterministic prior-mean transitions and the reward head; no stochastic latent trajectory is sampled during this optimization. This is not the CMA-ES procedure from World Models. The analogy is limited to a small controller acting on learned latent features.

In[10]:
Code
controller = nn.Linear(latent_dim, 1)

## Seed imagined rollouts from real encoded latents.
with torch.no_grad():
    seed_idx = torch.randint(len(train_idx), (64,))
    seed_mu, _ = encoder(obs_t[train_idx[seed_idx]])
    seed_z = seed_mu


def imagine_return(z0, horizon=15, gamma=0.95):
    """Roll prior means forward and sum the predicted next-state rewards."""
    z = z0
    total = torch.zeros(z0.shape[0])
    disc = 1.0
    for _ in range(horizon):
        a = ACTION_LIMIT * torch.tanh(controller(z))
        mu_p, _ = prior(z, a)
        z = mu_p
        total = total + disc * reward_head(z).squeeze(-1)
        disc *= gamma
    return total


def train_controller(steps=400):
    # Freeze the fitted world model; retain gradients through it to controller actions.
    for module in (encoder, prior, decoder, reward_head):
        module.requires_grad_(False)
    opt = torch.optim.Adam(controller.parameters(), lr=1e-2)
    for _ in range(steps):
        returns = imagine_return(seed_z)
        loss = -returns.mean()
        opt.zero_grad()
        loss.backward()
        opt.step()
    with torch.no_grad():
        return float(imagine_return(seed_z).mean())


imagined_return = train_controller()
print(f"Imagined return of the trained tiny controller: {imagined_return:.4f}")

Second, at each real step we infer the current latent and run a CEM-style search over action sequences using the same prior means and reward head. Prior means remove latent-transition sampling; CEM still samples action sequences. The fixed NumPy and PyTorch seeds support repeatable runs in the same software, platform, and device environment, but do not guarantee identical numbers across releases or hardware. PlaNet itself sampled one stochastic latent trajectory per candidate. As in the historical planner's bounded-action step, each iteration here samples an unbounded Gaussian parameterized in action coordinates, clips each sampled action once to the same [−1,1][-1,1] range as the tiny controller, and refits the Gaussian moments from the clipped elites. The returned action is the first component of the final proposal mean, clamped to that range.

In[11]:
Code
def cem_plan(z0, horizon=8, n_candidates=128, n_elites=16, n_iter=3):
    """Plan a single action for a single latent state using CEM."""
    action_mean = torch.zeros(horizon, 1)
    action_std = torch.full((horizon, 1), 0.7)

    with torch.no_grad():
        for _ in range(n_iter):
            samples = action_mean + action_std * torch.randn(
                n_candidates, horizon, 1
            )
            samples = samples.clamp(-ACTION_LIMIT, ACTION_LIMIT)
            scores = torch.zeros(n_candidates)
            z = z0.unsqueeze(0).repeat(n_candidates, 1)
            disc = 1.0
            for t in range(horizon):
                mu_p, _ = prior(z, samples[:, t])
                z = mu_p
                scores += disc * reward_head(z).squeeze(-1)
                disc *= 0.95
            elite_idx = torch.topk(scores, n_elites).indices
            elites = samples[elite_idx]
            action_mean = elites.mean(dim=0)
            action_std = elites.std(dim=0, correction=0).clamp_min(1e-2)
    return action_mean[0].clamp(-ACTION_LIMIT, ACTION_LIMIT)


def evaluate(policy_fn, initial_states, horizon=25):
    returns = []
    for initial_state in initial_states:
        s = initial_state.copy()
        total = 0.0
        for _ in range(horizon):
            a = policy_fn(s)
            s = step_world(s, a)
            total += reward_of(s)
        returns.append(total)
    return np.asarray(returns)


def tiny_controller_policy(s):
    with torch.no_grad():
        mu_q, _ = encoder(torch.from_numpy(s).unsqueeze(0).float())
        a = ACTION_LIMIT * torch.tanh(controller(mu_q)).item()
    return a


def cem_policy(s):
    with torch.no_grad():
        mu_q, _ = encoder(torch.from_numpy(s).unsqueeze(0).float())
        a = cem_plan(mu_q[0])
    return float(a)


def no_action_policy(s):
    return 0.0


evaluation_starts = rng.normal(scale=0.6, size=(24, 2)).astype(np.float32)
returns_no_action = evaluate(no_action_policy, evaluation_starts)
returns_tiny = evaluate(tiny_controller_policy, evaluation_starts)
returns_cem = evaluate(cem_policy, evaluation_starts)

for name, values in (
    ("No-action baseline", returns_no_action),
    ("Tiny imagined controller", returns_tiny),
    ("Latent CEM planner", returns_cem),
):
    print(
        f"{name:27s} mean 25-step return: {values.mean():.4f} "
        f"(SE across starts {values.std(ddof=1) / np.sqrt(len(values)):.4f})"
    )
Out[11]:
Console
No-action baseline          mean 25-step return: 22.3583 (SE across starts 0.4295)
Tiny imagined controller    mean 25-step return: 23.7577 (SE across starts 0.2144)
Latent CEM planner          mean 25-step return: 23.6934 (SE across starts 0.2070)

The printed results use the same 24 initial states for all three methods and include the passive baseline. They are a seeded illustration, not evidence that one coupling is generally better. After the common encoder pass, the tiny controller needs one controller pass per action, whereas CEM repeatedly evaluates its prior and reward head over a population of sequences online; the code therefore demonstrates a computation tradeoff, not a robustness guarantee.

Let's visualize the real-state trajectories of the passive baseline, tiny controller, and CEM planner, plus one imagined reward trajectory for the tiny controller.

Out[13]:
Visualization
Position-velocity phase portrait with three paths from the same start: a dashed no-action baseline, a tiny-controller path, and a CEM-planner path. Circles mark final states and a star marks the origin.
Paired 25-step state trajectories from one shared initial state. The tiny controller and latent CEM planner constrain actions to [-1,1], while the passive baseline always selects zero; circles mark final states and the star marks the reward peak.
Out[14]:
Visualization
Line chart of predicted reward over 30 imagined steps, declining gradually from just under 1.0 to a slightly lower value.
Predicted next-state reward along a 30-step prior-mean rollout of the tiny controller from one encoded training observation. Training instead maximized mean 15-step discounted return over 64 fixed encoded states, so this longer curve is descriptive; its per-step reward declines slightly.

The phase portrait shows three descriptive trajectories, one per method, from a shared start; the paired return summary above is the broader comparison. In the seeded prior-mean rollout, predicted reward declines slightly rather than rising monotonically. That curve does not, by itself, demonstrate model exploitation: doing so would require comparing imagined and true outcomes for the same action sequence and showing that the controller selected a systematic model error.

Limitations and impact

The three papers taught different lessons, and each carries a limitation. World Models showed that a compact controller trained entirely in a learned VizDoom dream could transfer to the real VizDoom environment; its CarRacing controller was instead optimized in the real environment. The separately trained VAE and memory need not retain every task-relevant feature, the no-hidden-layer controller limits expressiveness, and the paper directly observed dream exploitation. SimPLe showed that an observation-space video model could support data-efficient Atari policy learning, but its pixel objective is not directly decision-aware, and errors can feed back even through its deliberately truncated multi-step rollouts. PlaNet used an RSSM for online latent planning on six image-based DeepMind Control Suite tasks. It also proposed latent overshooting as a training regularizer, though its reported final RSSM did not require it. Its CEM search adds execution cost and can still prefer action sequences whose predicted returns are wrong.

These limitations arise from the model objectives, the controller or planner design, and the way the two are coupled.

  • World Models: The VAE and MDN-RNN remain active during deployment to supply ztz_t and hth_t; the paper's Doom-specific controller description and the released VizDoom controller also consume the LSTM cell state ctc_t. Their separate training objectives may omit task-relevant information, while a controller trained in the fixed dream can exploit its errors.
  • SimPLe: The observation-space simulator must be accurate enough for policy transfer. Its pixel likelihood is not directly optimized for control. The clipped visual loss stops gradients from already-confident pixels; the authors conjectured that this reduces background dominance. A separate reward head adds a task-specific signal. Remaining observation and reward errors can still bias the learned policy.
  • PlaNet: At each real decision step, the model is queried over candidate action sequences of its configured fixed decision-step horizon; the selected first action is repeated for a task-specific number of environment transitions (Algorithms 1–2). Re-observation limits open-loop commitment, but useful multi-step transition predictions and return rankings remain necessary.

Taken together, the three establish a spectrum of couplings between the world model and the controller, and they make one tension in model-based RL explicit. For these return-maximizing methods, expected return is the control objective; global predictive fidelity alone is not. Low average next-frame error does not by itself guarantee a good policy, and relatively high pixel error does not by itself rule one out, because control depends on errors in the decision-relevant quantities and trajectories queried by the particular controller. For the model's contribution to success, the relevant question is whether its predictions are decision-faithful on the states, action sequences, and horizons the controller queries, including the return rankings or value updates needed to choose good actions. Planner or policy quality, optimization, exploration, and data collection remain separate requirements. This framing connects several mechanisms discussed here: short grounded rollouts, KL-regularized latents, constrained policies, and model refresh.

The impact is clearest in the design questions that remained active. Dreamer retained an RSSM-like imagination model but learned actor and value networks instead of planning with CEM. It started imagined trajectories from posterior model states inferred from stored experience. DreamerV2 retained that starting point while replacing continuous Gaussian stochastic states with categorical ones. IRIS took another route, encoding frames as discrete observation tokens and predicting token dynamics with an autoregressive Transformer. These are distinct representation and control choices, not one direct lineage from SimPLe. Together, World Models, PlaNet, and SimPLe show observation-space simulators and learned latent-state models coexisting, with prediction coupled to recurrent controller features, simulator-trained policies, or online latent planning.

Summary

  • World Models predicts compressed latent dynamics, SimPLe predicts observation-space video with an internal recurrent state and a discrete stochastic latent, and PlaNet plans through a compact recurrent latent state-space model.
  • A sequential ELBO averages over the latent trajectory: observation and reward likelihoods encourage useful information, while KL regularization encourages agreement between each observation-conditioned posterior and its transition prior. None of these outcomes is guaranteed by a loss term alone.
  • The RSSM combines a finite deterministic history summary hth_t with a stochastic state ztz_t. Stochasticity permits conditional variation but does not by itself establish calibration or arbitrary one-step multimodality. PlaNet proposed latent overshooting to train multi-step prior consistency, but its reported final RSSM did not require it and the released configuration disabled that loss by default.
  • World Models deploys its VAE, MDN-RNN, and no-hidden-layer controller together; SimPLe alternates Atari model fitting with PPO updates driven primarily by experience from an observation-space simulator, plus a much smaller direct update from real trajectories; PlaNet trains no separate policy network and runs CEM online.
  • Under log loss, fitted predictive cross-entropy separates into irreducible conditional entropy plus a KL model-discrepancy term. Autoregressive rollout error depends on the chosen metric and dynamics stability, so its growth rate is not universal.
  • On the model side, exploitation risk and decision-relevant fidelity are two organizing concerns: a useful world model preserves the predictions, return rankings, or value updates its controller needs on queried states and horizons, rather than merely reconstructing observations well. Control performance also depends on the planner or policy and on data collection. Predictive fidelity, representation structure, uncertainty calibration, and decision usefulness remain distinct evaluation axes; improvement on one does not guarantee improvement on the others.
  • The useful comparison is between compressed features and memory feeding a controller, an observation-space simulator training a policy, and action sequences searched online through a latent model.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about World Models, SimPLe, and PlaNet.

World Models, SimPLe, and PlaNet

Question 1 of 80 of 8 completed
What key representation contrast does the chapter draw among World Models, SimPLe, and PlaNet?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026worldmodels, author = {Michael Brenndoerfer}, title = {World Models, SimPLe, and PlaNet}, year = {2026}, url = {https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). World Models, SimPLe, and PlaNet. Retrieved from https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison
MLAAcademic
Michael Brenndoerfer. "World Models, SimPLe, and PlaNet." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison>.
CHICAGOAcademic
Michael Brenndoerfer. "World Models, SimPLe, and PlaNet." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison.
HARVARDAcademic
Michael Brenndoerfer (2026) 'World Models, SimPLe, and PlaNet'. Available at: https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). World Models, SimPLe, and PlaNet. https://mbrenndoerfer.com/writing/world-models-simple-planet-model-based-rl-comparison

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.