Part of World Models Handbook
Compare Atari, board games, and open worlds as learned game engines. Measure one-step prediction, recursive rollout drift, and model-predictive control.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Games and Learned Game Engines
Press right in a grid world, and the agent should move one cell unless a boundary stops it. A conventional engine implements that rule directly. A learned engine predicts the next observation from examples of actions and outcomes; it may produce the right picture while getting the rule wrong. That difference matters both to a player who expects a coherent game and to an agent planning through predicted transitions. This chapter uses a small learned grid-world engine to measure the gap between one-step prediction, longer rollouts, and decisions made with the model.
Games are useful testbeds because an evaluator can often reset them, choose actions, and observe rewards. But the game families below pose different problems. A single Atari frame can omit emulator state; a board position in Go exposes the pieces and legal moves; a generated first-person scene must remain responsive and recognizable across actions. No one architecture or metric solves all three.
We will compare Atari, board games, and more open-ended interactive worlds; distinguish models used for agent decisions from worlds rendered for people; and test prediction, rollout error, and model-predictive control in one executable example. A compact evaluation checklist below separates controllability, visual quality, rule fidelity, and decision usefulness; it is not a new benchmark. The chapter builds on Part V: World-Model Architectures and Part VIII: Decision-Centric Research Lineages, before the application domains that follow.
Atari, Board Games, and Open Worlds
The word "game" hides wildly different problems. A world model that works on Atari will not automatically work on Go, and both will struggle in Minecraft for reasons that have nothing to do with one another. Part of the confusion is that "game" is defined by its user interface and cultural role, not by its mathematical structure. Two games that look similar on a screen can live in completely different regimes when characterized as dynamical systems. What separates them is a small set of structural properties: state dimension, observability, branching factor, horizon, determinism, and the amount of hidden mechanics. Understanding those properties is what turns "we trained a model that plays games" into a statement we can reason about. It is also what tells you, in advance, which techniques are likely to help and which are likely to stall.
A learned game engine is a model that, given the current state of a game and an action, produces a plausible next state (and often a reward, termination flag, or auxiliary quantities). The engine may be probabilistic, may be imperfect, and may violate the rules of the game it is supposed to be emulating. You can use it either as a training environment for agents or as a playable artifact for humans.
The definition is deliberately loose on the word "plausible." It might mean a close match to the next frame, or only enough accuracy to support a specified decision. The relevant standard depends on the consumer: an agent needs decision-relevant quantities, while a player also needs a responsive, intelligible view.
Atari: High-Dimensional Observations, Low-Dimensional Truth
The Arcade Learning Environment (ALE), introduced by Bellemare and colleagues, provides a common interface to hundreds of Atari 2600 environments; the familiar 57-game evaluation set is a benchmark selection, not the interface's limit. Standard RGB observations are typically 210 × 160 pixels, though game-specific observation spaces should be checked; the emulator exposes discrete joystick/button combinations, scores, and configurable evaluation choices. Those choices matter. Frame skipping, episode starts, and stochasticity can change the task being measured. In particular, the later ALE protocol study introduced sticky actions to make action execution stochastic and reduce exploitation of deterministic trajectories. A score is comparable across papers only after the game set and protocol have been specified.
A typical 210 × 160 ALE RGB frame contains 210 × 160 × 3 = 100,800 channel values, but image size is not the same as state complexity. The emulator can carry variables not recoverable from a single image, including timing and off-screen objects. Some Atari tasks therefore require a history of observations or another memory mechanism to predict what follows. A compact learned state should retain what matters for the intended task, rather than reconstructing every pixel; that is the sufficient-state question from Part III: Representing Agents and Worlds. Neither "all Atari is deterministic" nor "all Atari state is tiny" follows from the image dimensions.
The 2015 DQN result showed that a convolutional value function could learn control from processed game frames across multiple Atari titles, using the same algorithmic setup but separate training for each game. Its short frame history helps expose motion. Sparse-reward games such as Montezuma's Revenge remained difficult, however; that observation alone does not isolate representation, memory, or exploration as the cause. DQN is not itself a learned transition model, so a stack of frames should not be called a world model.
Board Games: Perfect Information, Enormous Search
In chess, Go, and shogi, both players see the board and the legal moves are defined by known rules. The challenge is less hidden state than the number of possible continuations. A Go position can offer hundreds of legal moves, and search depth multiplies those choices. AlphaZero used known game rules with learned policy and value predictions to guide Monte Carlo tree search. MuZero kept search but learned the dynamics used inside it. Those are importantly different uses of a model: AlphaZero learns how to evaluate and prioritize positions, whereas MuZero learns a compact transition representation as well.
The value-equivalence lesson belongs to MuZero's learned search model, not to AlphaZero's known-rule transition function. MuZero trains its representation, dynamics, and prediction components for reward, policy, and value targets without requiring an image decoder. A latent model can be useful for search without faithfully reconstructing the visible board, but success on those targets does not mean it has learned every rule or is valid under every action sequence. We return to this distinction in Part VIII's planned chapter on MuZero and value-equivalent models.
Open Worlds: Composition, Hierarchy, and Memory
Minecraft illustrates open-ended, compositional tasks: an agent may need to remember locations and combine objects, resources, and intermediate actions over a long horizon. DOOM is a different, level-based first-person example used in learned video generation; it should not be treated as an open-world benchmark simply because it has a rendered 3D view. Both can expose only part of the underlying state through the current image. A predictor trained on familiar scenes also needs to be tested on changed arrangements, as discussed in Compositionality and Systematic Generalization.
The difficulty of a Minecraft objective such as collecting diamonds includes sequential dependencies: gathering resources and crafting tools precede mining. A model or policy can fail at any dependency, so final reward by itself does not diagnose the failure. DreamerV3 reported collecting diamonds from scratch, but that result does not make every generated-world system an open-world agent. GameNGen modeled DOOM, DIAMOND trained an Atari agent with a diffusion world model and demonstrated a separate Counter-Strike: Global Offensive generator, and the Genie family pursued interactive generation. These are different tasks, data sources, and evaluation protocols.
What the Three Regimes Actually Stress
It is worth tabulating the structural differences because they determine which techniques help and which do not. Two regimes that look superficially alike, both rendered as pixel grids, for instance, can differ in ways that force completely different architecture choices.
- Observation. An Atari image and a Minecraft view may omit state; a board position in chess or Go exposes the game pieces, with any clock modeled separately.
- Branching and horizon. Legal move counts, episode lengths, and useful planning horizons depend on the game and evaluation protocol. Go search and a multi-stage Minecraft task stress different kinds of depth.
- Stochasticity. Board-game rules are usually deterministic. ALE can be evaluated with stochastic sticky actions, while some 3D game environments, including procedurally generated Minecraft worlds, include random elements. A stochastic predictive distribution is useful when uncertainty affects decisions, not automatically mandatory for every open world.
- Rules. Chess rules are directly specified; a pixel-only learner must infer relevant game mechanics from experience. Rule consistency is a testable property only after the rule and observation conditions are stated.
A method demonstrated in one regime should therefore be evaluated again in the next, rather than assumed to fail or transfer automatically. The relevant tests change with the task: hidden-state inference for a pixel-only Atari agent, search quality for board play, and persistence and compositional behavior for a longer generated world.
Learned Simulators and Playable Generation
A learned simulator replaces some or all hand-coded transitions with a model trained on interaction data. A frame predictor outputs a next image from recent images and actions; an interactive generator must also keep responding over many steps. Both are world models, but the right architecture and tests depend on whether the output is used for planning, training, or play.
The purpose matters, because it changes which errors are consequential. Two uses are useful to distinguish.
- Simulation for agents. The model generates transitions used for planning or imagined policy updates. Reward, value, termination, and action consequences must be accurate enough for the chosen decision process; photorealistic reconstruction is optional. Dreamer and MuZero are examples with learned latent dynamics.
- Playable generation for humans. The system renders an action-responsive view under a latency budget. Visual continuity, controllability, and rule behavior matter together. GameNGen and the Genie systems are examples, but their tasks and hardware differ.
The representation follows that objective. A compact latent can reduce the cost of simulating candidate futures, but its efficiency depends on model size, rollout length, hardware, and any decoder or search. A renderable system may still predict compressed visual tokens rather than raw pixels. Ha and Schmidhuber's World Models is already a hybrid: it learns visual compression and latent dynamics, then reconstructs frames for inspection. The practical comparison is task- and implementation-specific, not an overnight-versus-month rule.
The Latent Lineage: Dreamer, PlaNet, MuZero
Building on Recurrent State-Space Models, the latent lineage encodes observations and predicts future latent states under actions. PlaNet combines a deterministic recurrent path with stochastic latent variables. Dreamer uses imagined latent trajectories to improve its actor and critic, while collecting additional environment interactions to train and update the world model. DreamerV2 and V3 extended that line. A stochastic latent can represent multiple possible futures, though deterministic models remain useful when their assumptions fit the task. With denoting reward after action and denoting whether the episode continues after that action, one generic transition description is
where:
- : a latent representation at time that may summarize observation history
- : the action taken at time
- : the learned transition distribution over the next latent given the current latent and action
- : the reward after action , whose value the model predicts during imagination
- : the continuation indicator after action (one if the episode continues, zero if it terminates)
- : a generic reward-and-continuation prediction; individual implementations may condition these heads on only a subset of the listed variables
The sampled transition can represent more than one possible next latent for the same inputs. Whether the latent should be interpreted as a belief over hidden states depends on the inference model and observation history; stochasticity alone does not establish that interpretation. The indexing above keeps the after-action reward aligned with the worked example below.
MuZero trains representation, dynamics, and prediction networks for search targets without an observation-reconstruction objective. Its learned latent need not depict the board or Atari frame. That can avoid spending model capacity on visual details irrelevant to its search targets; it does not prove a general compute advantage over every reconstructive model. Search cost still depends on model architecture and the number of nodes expanded.
The relevant intuition is decision sufficiency: a model need not reconstruct everything to support a particular search objective. If the task only requires choosing an action in two visually different states, preserving the action ranking can suffice. Equal action-conditional rewards and values are a stronger sufficient condition, not a necessary one. A new task or action set may require distinctions the representation discarded.
The Pixel Lineage: GameNGen, DIAMOND, Genie
Ha and Schmidhuber's World Models (2018) combined a VAE for image compression, an MDN-RNN for latent dynamics, and a small controller. The VizDoom controller was trained in a generated dream environment; the CarRacing experiment used a different training arrangement. The result is an influential early hybrid of learned compression and dynamics.
The visual-generation projects differ in both training data and output. DIAMOND (2024) used a diffusion world model to train an Atari agent and separately demonstrated an interactive model trained on static Counter-Strike: Global Offensive gameplay. GameNGen (2024) first recorded DOOM play from an RL agent, then trained an action-conditioned diffusion frame predictor; its reported interactive rate was about 20 frames per second on a single TPU. Genie (2024) learned latent actions from unlabeled video for controllable generation. That training description should not be copied verbatim onto every successor. Genie 2 generated action-controlled 3D scenes for up to a minute in reported examples, most of which lasted 10–20 seconds. More recently, Genie 3 reported real-time 720p interaction at 24 frames per second with consistency over a few minutes. These are reported demonstrations under different settings, not matched benchmark scores.
The differences between these methods come down to a small set of choices.
- Observation representation. Pixels, latent vectors, discrete tokens. Each has different compression efficiency and different reconstruction quality.
- Dynamics model. Recurrent latent prediction (Dreamer), token prediction (for example, IRIS), or diffusion-based visual prediction (DIAMOND and GameNGen).
- Conditioning data. Recorded actions are available for GameNGen's gameplay data; Genie learns a latent action representation from video without action labels.
- Temporal context. A longer input history can help with hidden state or visual continuity, but increases memory and compute in architecture-dependent ways.
These choices interact, but none fixes long-horizon behavior automatically. The right design depends on whether the output is used for decisions, rendered to a player, or evaluated as a research model.
What Makes a Simulator "Playable"
Playable generation places stricter human-facing requirements on three properties that can also matter to an agent-facing simulator.
- Controllability. Actions must causally change the generated world, with feedback fast enough for the intended interaction. A game that ignores a button press is not playable, regardless of visual quality.
- Interactive throughput. Latency must be low enough for the intended control loop. GameNGen reports about 20 frames per second on a single TPU; that is a measured configuration, not a universal threshold.
- Persistence. Revisiting a location should reveal a coherent continuation of the earlier scene when the task requires it. A fixed context window may make that difficult, but its limit must be measured for the particular model.
One-step frame quality alone does not establish playability. A practical evaluation should record the model, hardware, frame rate and input-to-output latency; test whether changing an action changes the next outcome; compare multi-step trajectories and revisited locations against a reference when one exists; count specified rule violations; and, if agent training is the goal, evaluate the resulting policy in the reference environment. Short attractive clips and downstream return answer different questions. None of these measurements by itself certifies a complete game engine.
We will return to memory and consistency in the next section. For now, the takeaway is that learned simulators are not one thing. They are a family of compromises between fidelity, speed, controllability, and abstraction, and the right point on that surface depends on whether the consumer is an agent computing gradients or a human holding a controller. Choosing the wrong point can make an otherwise excellent model useless for your purpose, which is why the "what is it for" question has to be answered first.
A Worked Example: Learning a Grid-World Engine
We will now build a grid-world engine with a movable agent and fixed goal. Its two-channel observation marks the agent and goal, and its after-action reward is negative Manhattan distance to the goal. A small network predicts the next observation and reward. We will test how its predictions change under recursion, then use it to score candidate action sequences while executing and evaluating the resulting choices in the reference environment.
The observation contains 50 binary entries: two 5 × 5 channels, one for the agent and one for the goal. That makes mistakes easy to inspect. The experiment is a teaching example, not a timing or scaling result for pixel-based engines.
The Environment and Its Data
We import the scientific stack and set NumPy and PyTorch seeds so this notebook run can be reproduced in a compatible environment. A seed does not remove sampling uncertainty, implementation differences, or training-data effects; later curves describe one seeded run.
import matplotlib.pyplot as plt
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
torch.manual_seed(0)
np.random.seed(0)Actions are the four cardinal moves, clipped at the boundary. The reward is the negative Manhattan distance from the agent to the goal after the action; reaching the goal gives zero, the maximum reward. An action that leaves distance unchanged has unchanged reward, not an improvement signal. Episodes end at the goal or after twenty steps. The budget bounds trajectory length but does not make state coverage uniform.
class GridWorld:
"""A tiny 5x5 grid world with a movable agent and a random goal."""
ACTIONS = [(-1, 0), (1, 0), (0, -1), (0, 1)] # up, down, left, right
def __init__(self, size=5, max_steps=20, seed=0):
self.size = size
self.max_steps = max_steps
self.rng = np.random.default_rng(seed)
self.agent = None
self.goal = None
self.t = 0
def reset(self):
self.agent = self.rng.integers(0, self.size, size=2)
self.goal = self.rng.integers(0, self.size, size=2)
while np.array_equal(self.goal, self.agent):
self.goal = self.rng.integers(0, self.size, size=2)
self.t = 0
return self._obs()
def _obs(self):
obs = np.zeros((2, self.size, self.size), dtype=np.float32)
obs[0, self.agent[0], self.agent[1]] = 1.0
obs[1, self.goal[0], self.goal[1]] = 1.0
return obs
def step(self, a):
dy, dx = self.ACTIONS[a]
self.agent = np.array(
[
int(np.clip(self.agent[0] + dy, 0, self.size - 1)),
int(np.clip(self.agent[1] + dx, 0, self.size - 1)),
]
)
self.t += 1
dist = float(np.abs(self.agent - self.goal).sum())
reward = -dist
done = np.array_equal(self.agent, self.goal) or self.t >= self.max_steps
return self._obs(), reward, doneTo generate data, we interact with the environment using uniformly sampled actions. This is a simple behavior policy, not passive observation: the collector chooses actions and receives their consequences. Nor does uniform action sampling imply uniform state coverage; boundaries, early termination, and the reset distribution shape which transitions appear. Interaction Data and Passive Observation explains why the behavior policy and coverage matter.
def collect_random_transitions(n_episodes=2000, seed=1):
env = GridWorld(seed=seed)
obs_list, act_list, next_obs_list, rew_list = [], [], [], []
for _ in range(n_episodes):
obs = env.reset()
done = False
while not done:
a = int(env.rng.integers(0, 4))
next_obs, r, done = env.step(a)
obs_list.append(obs)
act_list.append(a)
next_obs_list.append(next_obs)
rew_list.append(r)
obs = next_obs
return (
np.array(obs_list, dtype=np.float32),
np.array(act_list, dtype=np.int64),
np.array(next_obs_list, dtype=np.float32),
np.array(rew_list, dtype=np.float32),
)
obs_arr, act_arr, next_obs_arr, rew_arr = collect_random_transitions()Collected 32158 transitions Obs shape: (32158, 2, 5, 5) Mean reward: -3.488 Reward range: -8.0 to -0.0
Each transition is , where is the reward after action and . The model maps a flattened observation and one-hot action to a predicted next observation and scalar reward. Unlike Dreamer, it has no recurrent latent state. It is deliberately small enough to inspect, and it can still expose the same gap between one-step prediction and recursive use.
The Model
The architecture is minimal. A two-hidden-layer MLP takes the flattened observation and one-hot action; separate linear heads predict the next observation and reward. The heads allow distinct losses or weights, although this example simply sums two mean-squared errors. Other world models may use different state representations, objectives, and rollout tests; this MLP is not a template for all of them.
OBS_DIM = 2 * 5 * 5 # two channels (agent, goal), each a 5x5 grid
N_ACTIONS = 4
class WorldModel(nn.Module):
def __init__(self, obs_dim=OBS_DIM, n_actions=N_ACTIONS, hidden=128):
super().__init__()
self.n_actions = n_actions
self.trunk = nn.Sequential(
nn.Linear(obs_dim + n_actions, hidden),
nn.ReLU(),
nn.Linear(hidden, hidden),
nn.ReLU(),
)
self.next_obs_head = nn.Linear(hidden, obs_dim)
self.reward_head = nn.Linear(hidden, 1)
def forward(self, obs_flat, a):
a_oh = F.one_hot(a, self.n_actions).float()
h = self.trunk(torch.cat([obs_flat, a_oh], dim=-1))
return self.next_obs_head(h), self.reward_head(h).squeeze(-1)W_model = WorldModel()
optimizer = torch.optim.Adam(W_model.parameters(), lr=1e-3)
obs_t = torch.from_numpy(obs_arr.reshape(len(obs_arr), -1))
act_t = torch.from_numpy(act_arr)
nobs_t = torch.from_numpy(next_obs_arr.reshape(len(next_obs_arr), -1))
rew_t = torch.from_numpy(rew_arr)
n_val = 5000
train_obs, train_act, train_nobs, train_rew = (
obs_t[:-n_val],
act_t[:-n_val],
nobs_t[:-n_val],
rew_t[:-n_val],
)
val_obs, val_act, val_nobs, val_rew = (
obs_t[-n_val:],
act_t[-n_val:],
nobs_t[-n_val:],
rew_t[-n_val:],
)Training sums mean-squared error for next-observation entries and reward. Equal coefficients are a simple demonstration choice, not evidence that the two error scales are comparable or appropriately balanced. Their values should be inspected separately when that balance matters; Stochasticity and Uncertainty discusses predictive uncertainty rather than prescribing this weighting.
loss_history = []
for epoch in range(15):
W_model.train()
perm = torch.randperm(len(train_obs))
epoch_loss = 0.0
for i in range(0, len(train_obs), 256):
idx = perm[i : i + 256]
pred_no, pred_r = W_model(train_obs[idx], train_act[idx])
target_no = train_nobs[idx] + 0.02 * torch.randn_like(train_nobs[idx])
target_r = train_rew[idx] + 0.10 * torch.randn_like(train_rew[idx])
loss = F.mse_loss(pred_no, target_no) + F.mse_loss(pred_r, target_r)
optimizer.zero_grad()
loss.backward()
optimizer.step()
epoch_loss += loss.item() * len(idx)
loss_history.append(epoch_loss / len(train_obs))Final training loss: 0.0344

The training objective falls, but it cannot by itself tell us how well the model predicts held-out transitions. The environment is deterministic, while the training targets include small random perturbations. The next tests separate fitting this objective from one-step prediction and recursive rollout fidelity.
One-Step Accuracy Is Not Enough
We start with 5,000 validation transitions set aside from the end of the collected array. This is a transition-level split: the boundary need not be episode-disjoint, and the small grid permits repeated states. We compute the mean Euclidean (L2) distance between flattened predicted and true next observations:
over that held-out set of transitions. Here:
- : the number of held-out transitions evaluated
- : the model's predicted next observation for the -th held-out transition, flattened to a vector
- : the true next observation for the -th held-out transition, flattened to a vector
- : the Euclidean (L2) norm, so each example contributes the overall pixel-space distance between prediction and truth
with torch.no_grad():
pred_next, pred_r = W_model(val_obs, val_act)
per_example_err = ((pred_next - val_nobs) ** 2).sum(dim=-1).sqrt().numpy()
reward_err = (pred_r - val_rew).abs().numpy()Mean 1-step L2 error: 0.9972 Mean reward absolute error: 0.0492

The errors cluster around an L2 distance of one on the flattened two-channel grid, not around zero. That is a meaningful one-step mismatch in this small environment: the clustered error should not be mistaken for near-exact predictions. One-step validation remains useful, but it does not reveal how errors change when predictions become inputs to later predictions. We therefore measure recursive rollout error separately.
Long-Horizon Consistency and Rules
One-step accuracy is not enough when a model is fed its own predictions. A small error may change the next input, which can create further error; whether the deviation grows, shrinks, or changes decisions depends on the learned transition and the states visited. The bound below shows one possible relationship between one-step and recursive errors under an explicit smoothness assumption. It is not a universal forecast that every video model fails after a fixed number of frames.
Compounding Error: The Math
Suppose the true dynamics are and a model rollout is under the same action sequence. Let and let be the one-step model error evaluated at the true state. If is -Lipschitz in its state input along both compared trajectories, the triangle inequality gives
where:
- : the error between the predicted and true state at rollout step
- : the Lipschitz constant of with respect to its state input, i.e., the worst-case factor by which a small state perturbation is amplified in one model step
- : the one-step model error at step , i.e., the error introduced when the model predicts the next state from the true current state
Unrolling the recursion gives the error at horizon :
Here is the initial error. If every , then for the bound approaches as grows while the initial-error term decays; for it is ; and for it may grow exponentially. This is an upper bound, not an equality or a measured growth law. Its premises require the same actions and a Lipschitz bound covering the compared states, including the model-generated ones; a bound measured only on training states would not suffice for an out-of-distribution rollout.
The three regimes can be compared directly by iterating the recursion with a constant per-step error.
# Theoretical rollout-error envelopes from e_{t+1} = L * e_t + eps_t,
# with a constant per-step error eps = 0.05 and initial error e_0 = eps.
# L < 1 damps prior error toward a finite bound, L = 1 accumulates
# linearly, and L > 1 permits exponential growth of the bound. These
# curves iterate the bound above; they are not measurements, so no seed is needed.
lipschitz_values = [0.8, 1.0, 1.2]
per_step_error = 0.05
theoretical_horizon = np.arange(0, 21)
theoretical_error_curves = {}
for L_val in lipschitz_values:
errs = np.empty(theoretical_horizon.size)
errs[0] = per_step_error
for t in range(1, theoretical_horizon.size):
errs[t] = L_val * errs[t - 1] + per_step_error
theoretical_error_curves[L_val] = errs
The analytical curves illustrate possible amplification. They do not estimate this fitted grid-world model's Lipschitz constant. Shortening a rollout limits the number of recursive predictions, and re-anchoring replaces a predicted state with an observation; whether either improves decisions still requires evaluation.
Two Failure Modes: Drift and Collapse
Two useful descriptions of long-horizon failure are gradual drift and abrupt rule failure. They are not exhaustive or tied to a universal time scale.
- Drift. A predicted trajectory can gradually diverge in position, appearance, or timing while each frame still looks plausible. Measure it against a reference trajectory under a specified action sequence; do not infer a fixed failure frame count from a single demonstration.
- Rule failure. A model may generate a transition that violates a specified mechanic, such as a character jumping again in mid-air when the game permits only one jump. The error may appear abruptly or after a long rollout; it need not be caused only by an out-of-distribution input.
Different measurements reveal different problems. A frame-distance curve may reveal accumulating prediction error, while a targeted rule test can expose a failure that looks visually minor. Noise augmentation, broader training coverage, longer context, and explicit constraints are possible interventions, but their effects depend on the model and rule being tested.
Diffusion does not make temporal errors disappear: an autoregressive diffusion world model still conditions future frames on generated history. DIAMOND's project report shows a concrete rule violation—repeated jumps not allowed by the reference game. GameNGen used conditioning augmentations for stable autoregressive generation, and its authors evaluated clips after five minutes of generation. Neither example supports a blanket claim that visual generators fail after 30–60 frames or that one model family is intrinsically immune to collapse.
Rule Consistency: The Underappreciated Signal
Games have mechanics that can be stated and tested. A learned model that lets an agent pass through an impassable wall may support a policy that fails in the reference game. But scores can legitimately decrease in some games, and a visible inconsistency does not guarantee that a trained agent will exploit it. To assess rule consistency, specify the rule, action distribution, and observation conditions, then count violations over evaluated transitions.
This idea is worth stating as a criterion.
A learned game engine is rule-consistent on a distribution of trajectories if the fraction of transitions that violate a specified rule of the underlying game is small. Rule consistency is a separate requirement from perceptual similarity: a model can look right while breaking a rule that the eye does not check.
Downstream agent return in the reference game is a separate, indirect test of decision usefulness: poor transfer could stem from rule errors, reward errors, policy optimization, or another mismatch. It cannot by itself identify which mechanic failed. Direct tests can replay a chosen action sequence and score specified mechanics. Part XI's planned chapter on state, physics, causality, and memory evaluation develops that distinction further.
Techniques for Long-Horizon Consistency
Several techniques can be tested for longer rollouts; none is universal.
- Exposure to predicted states. Training on some model-generated histories can reduce the mismatch between training and rollout inputs, if those histories are representative. Do not assume every cited system uses the same procedure.
- Conditioning augmentation. Corrupting conditioning frames during training may improve robustness to prediction noise; GameNGen reports this intervention. It can also change one-step accuracy, so both should be measured.
- Temporal context or external memory. Access to earlier observations can help recover hidden or revisited details, at a model-dependent compute cost. A demonstrated one-minute video is not evidence of a one-minute input context window. See Temporal State and Memory.
- Latent prediction. A compact latent can omit pixel detail irrelevant to a decision objective; it does not make state or reward prediction error vanish.
- Re-anchoring. Replacing a predicted state with a new observation truncates the executed model rollout. It cannot retroactively remove errors in the candidate futures that selected the preceding action.
In the mathematical bound, a shorter horizon reduces the number of propagated terms. The other techniques may change the local errors or dynamics but do not guarantee a smaller or . They need before-and-after evaluation under the same protocol.
Measuring Drift in Our Grid-World Engine
We now collect fifty new random-policy episodes with a separate seed and compare model-generated observations with the reference environment under the same recorded actions, for at most twenty steps. Episodes terminate at different times, so the number contributing to each horizon can change. This is a fresh-episode test, unlike the transition-level validation split above.
def recursive_rollout(model, obs0_flat, actions):
"""Roll the model forward autoregressively. Returns predictions for t=0..T."""
preds = [obs0_flat.detach().clone()]
cur = obs0_flat.unsqueeze(0)
with torch.no_grad():
for a in actions:
a_t = torch.tensor([int(a)])
next_obs, _ = model(cur, a_t)
preds.append(next_obs.squeeze(0))
cur = next_obs
return torch.stack(preds)
def collect_held_out_episodes(n_episodes=50, seed=999):
env = GridWorld(seed=seed)
episodes = []
for _ in range(n_episodes):
obs = env.reset()
ep_obs, ep_act = [obs], []
done = False
while not done:
a = int(env.rng.integers(0, 4))
next_obs, _, done = env.step(a)
ep_act.append(a)
ep_obs.append(next_obs)
obs = next_obs
episodes.append(
(
np.array(ep_obs, dtype=np.float32),
np.array(ep_act, dtype=np.int64),
)
)
return episodes
held_out = collect_held_out_episodes()
T = 20
horizon_err_sum = np.zeros(T + 1)
horizon_counts = np.zeros(T + 1)
with torch.no_grad():
for ep_obs, ep_act in held_out:
obs0 = torch.from_numpy(ep_obs[0].reshape(-1))
preds = recursive_rollout(W_model, obs0, ep_act[:T])
L = len(preds)
true_obs = torch.from_numpy(ep_obs[:L].reshape(L, -1))
errs = ((preds - true_obs) ** 2).sum(dim=-1).sqrt().numpy()
horizon_err_sum[:L] += errs
horizon_counts[:L] += 1
horizon_errors = horizon_err_sum / np.maximum(horizon_counts, 1)Rollout L2 error by horizon: [0. 1.001 1.113 1.182 1.226 1.296 1.334 1.401 1.444 1.469 1.527 1.58 1.607 1.669 1.746 1.815 1.929 2.065 2.201 2.41 2.585]

The mean L2 distance rises from around one after the first predicted transition to about 2.6 at step twenty in this seeded run. At each horizon the mean uses episodes still available there, so the changing set is another reason not to read the curve as a fitted growth law. It neither identifies the model's Lipschitz constant nor proves that a trajectory left the training-state distribution.
The aggregate curve hides the appearance of an individual trajectory. For one selected held-out episode, we decode the highest-scoring agent cell in each predicted frame and compare it with the true path. This illustration does not estimate how frequently such divergences occur.
# Decode the agent cell from channel 0 of each predicted observation and
# compare it with the true agent cell from the held-out episode.
# Selection rule: keep episodes with at least 15 actions; choose the episode
# whose true path visits the most distinct cells (first index breaks ties).
# Episodes come from collect_held_out_episodes(seed=999) collected above.
trajectory_rollout_horizon = 15
candidate_episodes = [
idx for idx, (ep_obs, ep_act) in enumerate(held_out) if len(ep_act) >= 15
]
# Selection uses only the true path, not model error; the figure is illustrative.
candidate_episodes = sorted(
candidate_episodes,
key=lambda i: len(
np.unique(
[
tuple(np.unravel_index(int(c.argmax()), (5, 5)))
for c in held_out[i][0].reshape(-1, 2, 5, 5)[:, 0]
],
axis=0,
)
),
reverse=True,
)
selected_episodes = candidate_episodes[:1]
trajectory_episodes = []
with torch.no_grad():
for idx in selected_episodes:
ep_obs, ep_act = held_out[idx]
obs0 = torch.from_numpy(ep_obs[0].reshape(-1))
preds = recursive_rollout(
W_model, obs0, ep_act[:trajectory_rollout_horizon]
)
n_steps = len(preds)
true_agent = ep_obs[:n_steps].reshape(n_steps, 2, 5, 5)[:, 0]
pred_agent = preds.reshape(n_steps, 2, 5, 5)[:, 0]
true_cells = np.array(
[np.unravel_index(int(p.argmax()), (5, 5)) for p in true_agent]
)
pred_cells = np.array(
[np.unravel_index(int(p.argmax()), (5, 5)) for p in pred_agent]
)
goal_cell = np.unravel_index(int(ep_obs[0, 1].argmax()), (5, 5))
trajectory_episodes.append(
{
"episode": idx,
"goal": goal_cell,
"true": true_cells,
"pred": pred_cells,
}
)
The natural questions are: can an agent still use this model to make decisions, and how careful must it be about the drift? The next section addresses that.
Agents Trained Inside Generated Games
A learned model can supply imagined trajectories for policy updates or candidate-action search, but the model itself first needs data. Dreamer continues to collect environment experience while learning from imagined trajectories; MuZero learns dynamics useful for search; Ha and Schmidhuber tested a dream-trained controller in VizDoom. GameNGen's DOOM generation is primarily a playable-model demonstration, not evidence that an agent trained in it transferred to the reference game. Computational and data advantages depend on the environment, hardware, model, and evaluation protocol.
A model used for decisions needs at least three properties under the intended task distribution.
- Decision sufficiency. It should retain information needed to rank actions for the objective, even when visual reconstruction is not required.
- Rollout stability over the planning horizon. The drift we just measured must be small enough that short plans remain accurate.
- Reward realism. The predicted reward must be aligned with the true reward on states the planner tends to visit.
These conditions differ from visual playability, rather than forming a guaranteed sufficient test. A model can be useful for one planner while looking poor, and a rendered world can look persuasive while giving wrong rewards. Even when the three conditions appear to hold on a test set, distribution shift under optimization can expose further errors.
Two Ways to Train an Agent with a Model
The literature separates two styles of use.
- Policy learning in imagination. The agent updates an actor and critic from modeled trajectories, as Dreamer does. The learned policy can exploit errors in the world model; actual training also needs environment data for that model.
- Planning in the model. A planner evaluates candidate futures and chooses an action, as in MuZero's MCTS or the simple MPC below. Replanning from a new observation limits the length of the executed open-loop plan but does not make the first decision immune to errors in simulated futures.
Both styles can exploit errors. An imagined policy is trained against model predictions; a planner can choose a bad first action because an inaccurate later reward made its candidate sequence look attractive. Receding-horizon search then observes the new real state before choosing again, while an actor does not necessarily search anew at each decision. Search also has a per-decision cost that depends on its budget. The example below demonstrates random-shooting MPC, not a general superiority claim for planning.
Model Exploitation Is the Central Risk
Optimization can favor model-predicted outcomes that do not occur in the reference environment. For the grid world, separate the true reward rule from the network's reward-head prediction. The environment returns
Here is the true post-action agent coordinate, is the true goal coordinate, and is Manhattan distance. The learned model instead predicts
where is the separate scalar reward head after the MLP trunk. It is not computed by decoding the predicted next observation or measuring its distance to a decoded goal. In particular, a one-cell observation-prediction error does not imply a one-unit error in . The planner below sums these direct reward-head outputs. Model exploitation remains possible if a candidate action sequence receives optimistic reward predictions, but this code does not demonstrate that it occurs.
The distinction between the two heads matters diagnostically: observation drift and reward error need separate measurements. A visually wrong rollout might still rank actions well, while a visually plausible one might misrank them. Repeated optimization can search for optimistic predictions, but whether that happens here must be measured against true returns.
Common responses to this risk change the model, planner, or objective:
- Ensembles and uncertainty penalties. Model disagreement can signal poorly supported actions, although disagreement is not a perfect error estimate. PILCO, PETS, Ensembles, and MBPO compares related approaches.
- Short imagination horizons. Dreamer imagines finite rollouts and uses a value model to estimate returns beyond them. Shorter rollouts limit how far the learned dynamics are applied directly, although accumulated model error is not the paper's stated reason for this design.
- Real-data correction. New reference-environment transitions can reveal and repair model errors in visited regions, at an interaction cost.
- Conservatism. MOPO subtracts a model-uncertainty penalty from imagined rewards; MOReL routes unsupported state-actions to a pessimistic unknown state. Part VIII's planned chapter on offline and conservative model-based RL covers this family.
Each mitigation changes what the agent is allowed to optimize or what evidence it receives. None guarantees transfer for an arbitrary learned model. Evaluate the chosen policy in the reference environment, not solely against its training model.
Demonstrating Planning Inside the Learned Engine
At each environment step, random-shooting MPC samples action sequences of length , rolls them out inside the learned model, and selects the sequence with the highest sum of predicted reward-head outputs. It executes only that sequence's first action, observes the real next state, and plans again. Re-anchoring prevents errors in an unexecuted trajectory from becoming the next observed state, but errors at any simulated depth can still affect which first action wins the search.
def mpc_plan(model, obs_flat, horizon=5, n_samples=64, rng=None):
"""Random-shooting MPC. Returns the best first action."""
if rng is None:
rng = np.random.default_rng()
best_total = -np.inf
best_a0 = 0
with torch.no_grad():
for _ in range(n_samples):
seq = rng.integers(0, N_ACTIONS, size=horizon)
cur = obs_flat.unsqueeze(0)
total = 0.0
for a in seq:
a_t = torch.tensor([int(a)])
next_obs, r = model(cur, a_t)
total += float(r.item())
cur = next_obs
if total > best_total:
best_total = total
best_a0 = int(seq[0])
return best_a0We evaluate the planner in the reference grid world with a seed separate from data collection. This closed-loop test checks whether actions selected using the model reach the real goal; it complements, rather than replaces, transition and reward-error measurements.
def evaluate_mpc(model, horizon=5, n_episodes=30, seed=1234, n_samples=64):
env = GridWorld(seed=seed)
rng = np.random.default_rng(seed)
wins = 0
for _ in range(n_episodes):
obs = env.reset()
done = False
while not done:
obs_flat = torch.from_numpy(obs.reshape(-1))
a = mpc_plan(
model, obs_flat, horizon=horizon, n_samples=n_samples, rng=rng
)
obs, _, done = env.step(a)
if np.array_equal(env.agent, env.goal):
wins += 1
return wins / n_episodes
results = {}
for h in [1, 2, 5, 10]:
results[h] = evaluate_mpc(W_model, horizon=h)Horizon 1: success rate 1.00 Horizon 2: success rate 1.00 Horizon 5: success rate 1.00 Horizon 10: success rate 1.00

All four tested horizons reach the goal on every one of the thirty seeded evaluation episodes for that horizon. Each evaluation call starts its environment and planner generators from the same seed; different horizon lengths group and consume the planner draws differently, so this is not a paired comparison of identical candidate sequences. These saturated results show that the learned model supports this planner on these small test cases despite observation rollout error. They do not show a benefit from increasing the horizon, a drift-induced decline, an optimal horizon, or guaranteed improvement from more samples. On a harder task, those effects would need a broader paired evaluation.
Pixel-level rollout error and decision usefulness answer different questions. Here the planner reaches the goal on the tested episodes despite substantial observation error; thirty trials per horizon are too narrow to establish reliability elsewhere. Evaluate both model predictions and closed-loop outcomes on held-out starts and, where possible, new tasks, reporting uncertainty when comparing methods.
Limitations & Impact
The cited systems establish different accomplishments under different tests. DreamerV3 reports results across more than 150 tasks, not blanket superhuman performance on all 57 Atari games. AlphaZero used known rules for board-game search, and MuZero learned search dynamics. GameNGen demonstrated action-responsive DOOM generation on specified hardware; Genie 2 and Genie 3 reported increasingly long interactive visual worlds. Those claims should be read with their game sets, protocols, hardware, and time horizons attached. They do not imply that one learned engine handles every game or that model-generated experience removes the need for reference-environment tests.
The limitations are equally real, and it is worth being honest about them. Four stand out.
First, evaluation conditions vary even within games. A deterministic board game, Atari with sticky actions, and Minecraft do not share one observability or reward structure. Success in one setting does not establish transfer to physical manipulation or driving, where sensing, action, contact, safety, and data collection impose different constraints. The remaining Part X: Application Domains chapters treat those settings separately.
Second, consistency depends on what is being preserved and for how long. GameNGen reported multi-minute DOOM trajectories, Genie 2 reported worlds up to a minute with most demonstrations shorter, and Genie 3 reported real-time consistency for a few minutes. These are important but not equivalent to arbitrary-length persistence of every object, rule, and reward. Conditioning augmentation, longer context, memory, and re-anchoring incur different costs; no universal doubling or quadrupling factor follows from their names.
Third, optimizing against an imperfect model can favor prediction errors. Uncertainty penalties and short rollouts can reduce particular risks, but a policy still needs evaluation outside the model. The toy MPC result above is a useful reminder: high closed-loop success on a small seeded set does not measure exploitation pressure across other starts, goals, or planners.
Fourth, no single metric covers all intended uses. PSNR compares images, distributional video metrics compare sampled videos, targeted tests count rule violations, and closed-loop return tests whether decisions transfer under a specified protocol. The evaluation checklist earlier keeps these questions separate. Part XI: Evaluation and Understanding develops the broader evaluation framework.
The useful conclusion is narrower than a single progress ladder: learned dynamics can support search, imagined policy updates, or interactive rendering, but each success is conditional on its task and measurement. Before transferring a method to another domain, identify which state, action, reward, latency, and persistence assumptions change.
Summary
Atari, board games, and open-ended interactive environments stress different properties: hidden state in pixel observations, branching search under known rules, and composition or persistence over longer tasks.
- Atari challenges models to retain action-relevant state from high-dimensional, partially observable observations.
- Board games distinguish AlphaZero's search with known rules from MuZero's learned, decision-oriented dynamics.
- Long-horizon open-world tasks often stress compositionality and memory, but the requirements depend on the task.
Latent dynamics are central to PlaNet, Dreamer, and MuZero; visual generation is central to DIAMOND, GameNGen, and Genie. World Models combined visual compression with latent prediction. The design trade-offs include decision targets, renderability, action conditioning, context, and measured compute rather than a universal latent-versus-pixel speed ranking.
Longer recursive use requires measuring both prediction error and specified rule behavior. Under the displayed same-action and Lipschitz assumptions, the error bound is linear for and may grow exponentially for ; it does not fit a particular rollout. Conditioning augmentation, temporal context, memory, and re-anchoring are candidates for testing, not guaranteed fixes.
An imperfect model can still support useful decisions. Dreamer-style imagination updates and model-based search are distinct uses, and both can be misled by reward or transition errors. The grid-world planner succeeds on its seeded test set even though its observation predictions drift; that does not establish transfer beyond this example.
The next chapter turns to robotic manipulation, where contact, sensing, and data collection change what a useful model must predict. We should carry the evaluation habit forward, not assume that a game result transfers unchanged.
Key Parameters
The key parameters for the grid-world learned engine are:
- size: Grid dimension. Controls the state space and the number of possible agent and goal positions.
- max_steps: Reference-environment episode budget. It caps collected episodes and recorded-action rollout evaluation; MPC candidate length is set separately by horizon.
- hidden: Width of the MLP trunk. Increasing it raises parameter count and compute per model step; effects on held-out accuracy must be measured.
- lr: Adam learning rate for the world model. An excessively high value can destabilize training; a very low one can slow optimization.
- n_val: Number of held-out transitions. Sets the size of the evaluation split.
- horizon: Random-shooting MPC planning horizon. Longer horizons evaluate more predicted steps and can expose planning to more model error; the amount and direction must be measured on the planner's candidate distribution.
- n_samples: Number of random action sequences evaluated per MPC step. More samples increase search work and may find a higher predicted score; true performance need not improve if reward predictions are wrong.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about games and learned game engines.
Games and Learned Game Engines Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!