Video Generators as Candidate World Models

Michael BrenndoerferJuly 21, 202647 min read

Part of World Models Handbook

Can video generators serve as world models? Tests cover action-conditioned prediction, a qualification ladder, occlusion memory, and closed-loop planning.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Video Generators as Candidate World Models

Picture a small object just left of an occluder on a flat table. An agent can push left, push right, or do nothing. Suppose a generator makes a convincing clip for each choice. The shadows and lighting look plausible, yet the left-push clip shows the object drifting right; in another clip the object disappears behind the occluder and never returns. The question is whether the generated difference follows the chosen action and the scene's dynamics, not merely whether each clip looks natural.

Every frame can be crisp and the motion smooth while the predicted consequence is wrong for the action. Plausible continuation asks what a future might look like under the training distribution. A controller asks what would happen after its chosen intervention. The two conditionals can agree, but only when the data and causal assumptions justify that agreement.

A plausible clip does not by itself establish a useful world model. For the control tasks in this chapter, the candidate must predict consequences of feasible actions from an adequately specified state or history. Its video output makes certain failures inspectable, but pixels are not required of world models in general. Action response, persistence through occlusion, contact behavior, and uncertainty are separate properties to test. Which ones matter depends on the task and observation regime. A model may pass a short-horizon action test yet drift over longer rollouts, or predict a visible trajectory well while remaining overconfident about hidden outcomes.

We build on Part IX, Chapter 2, Pretrained Visual Models and DINO-WM. Its dynamics model predicts frozen visual features for planning without a pixel-reconstruction term; a separate diagnostic decoder can visualize features. Here, the candidate generates future frames directly. Those frames help reveal visible failures, but perceptual quality does not establish action-dependent dynamics or decision quality.

The comparison to classical methods is useful when their roles stay distinct: system identification estimates dynamics, state estimation infers hidden state from observations, and model predictive control chooses actions using predictions. A video generator may contribute to one or more of those roles, but its rendered output alone does not satisfy all three.

We first define the conditioning interface, then separate short-horizon prediction, long-horizon rollouts, physical constraints, and action response. A qualification ladder makes the evaluation questions explicit. A small corridor experiment shows how to measure rollout and action-response errors; its visibility-stratified error alone will not certify hidden-state memory.

Formal Setting: Frames, Actions, and Conditioned Futures

A video generator used as a world model samples futures conditioned on a short history and, when available, a sequence of actions. Write a frame as

ot∈RH×W×C,o_t \in \mathbb{R}^{H \times W \times C},

where HH and WW are pixel dimensions and CC is the channel count (three for RGB). This chapter examines a camera-observation setting: the physical state is hidden from the generator, while frames and any recorded controls are available. Other world models may also receive proprioception, maps, or privileged simulator state.

A history at time tt is a window of recent frames and recorded actions:

ht=(ot−L+1:t,  at−L+1:t−1),h_t = (o_{t-L+1:t},\; a_{t-L+1:t-1}),

Here LL is the context length, and the displayed full window assumes zero-indexed observations with t≥L−1t\geq L-1; for earlier steps, use the available shorter history. The action window ends at t−1t-1 because those controls have already been applied. A finite history can omit an earlier event needed to predict the future; a persistent memory state or another sensor may carry information beyond this window. We keep any other condition cc separate from hth_t. It may be text, a goal, or a time-varying prescribed camera trajectory; a camera command changes the observation process and is not automatically a physical action on the scene.

A generator samples a block of future frames

o~t+1:t+K∼pθ(o~t+1:t+K∣ht,  at:t+K−1,  c),\tilde{o}_{t+1:t+K} \sim p_\theta(\tilde{o}_{t+1:t+K} \mid h_t,\; a_{t:t+K-1},\; c),

Here KK is the prediction horizon, at:t+K−1a_{t:t+K-1} is the planned sequence of world-affecting controls, and cc holds other conditioning. The action space A\mathcal{A} may be discrete (left/right/stay) or continuous (an end-effector command). Timing must be specified somewhere in the conditioning for a timed-response test; a time-ordered action sequence is one explicit way to do so. Permuting the same actions can change intermediate states and sometimes the final state. A model given only an unordered set cannot distinguish the two sequences when their consequences differ.

Prompt-only versus action-conditioned generation

A prompt-only generator conditions on cc (usually text) but has no structured action sequence in this interface. It samples pθ(o~t+1:t+K∣ht,c)p_\theta(\tilde{o}_{t+1:t+K} \mid h_t, c). Such a model may learn temporal patterns and can be useful for passive prediction. A free-form prompt, however, is not a validated, time-aligned substitute for the specific action and no-op commands a controller must compare.

Even when a generator does accept actions, the conditioning distribution matters. Suppose a model was trained on a fixed behavioral policy πβ(at∣ot)\pi_\beta(a_t \mid o_t). The conditional distribution it learns is

p(ot+1∣ht,at)=p(ot+1,at∣ht)p(at∣ht)p(o_{t+1} \mid h_t, a_t) = \frac{p(o_{t+1}, a_t \mid h_t)}{p(a_t \mid h_t)}

where p(at∣ht)>0p(a_t\mid h_t)>0. This identity defines an observational conditional; a fitted pθp_\theta approximates it only to the extent supported by its data and model fit. If the policy tends to push only when the object is visible, a predictor may use visibility as a shortcut and underuse the action input. For control we want the consequence of setting the action:

p(ot+1∣do⁡(at),ht),p\big(o_{t+1} \mid \operatorname{do}(a_t), h_t\big),

This is the next-observation distribution when the agent sets ata_t. The observational and interventional conditionals coincide at a given history and action under consistency, conditional exchangeability (no unmeasured common cause of action and outcome after conditioning on hth_t), and positivity for that action. Randomized action assignment can help establish exchangeability; dense observed action support alone cannot. At zero support, p(ot+1∣ht,at)p(o_{t+1}\mid h_t,a_t) for that action cannot be estimated from those data—it is not thereby proven unequal to the interventional quantity. Hernán and Robins, Causal Inference: What If develops these identification conditions.

Consider a hidden property UU that affects both a policy's choice and the object's motion. Even if both actions occur at every observed history, conditioning on the chosen action can reveal information about UU; setting that action need not have the same effect. Positivity lets us compare actions within supported histories, while exchangeability rules out that hidden-confounding explanation. A nearly deterministic policy makes off-policy effects hard to estimate and may permit a good in-policy predictor to ignore actions. Neither excellent in-policy fit nor inevitable off-policy failure follows without testing the model.

This distinction echoes Part VI, Chapter 3, where we distinguished interaction data from passive observation. It will also frame our counterfactual probes below.

What Video Generation Adds (and What It Does Not)

Rendering the future as video has a genuine advantage and a genuine hazard.

The advantage is inspectability. Part IX, Chapter 2's feature-space rollouts require a probe or decoder for visual inspection. Generated frames expose visible events such as an object leaving a target region or an arm intersecting a clearly visible wall. A frame-based check still needs an appropriate viewpoint and, for automated scoring, a validated perception procedure. Hidden state, forces, and ambiguous contact cannot be read reliably from the clip alone.

The hazard is that visual plausibility and predictive correctness are different tests. A good perceptual score can coexist with a failed action-response or dynamics probe. In the occluded-corridor example, a smooth clip could reintroduce the object at a position inconsistent with the specified actions. Modern video generators may model whole clips and long-range temporal structure—not just isolated frames—yet their training objectives do not automatically enforce the action-conditioned physical constraints of a particular task. The Sora technical report discusses both emerging simulation-like behavior and explicit failure modes.

Perceptual coherence

The property that a clip resembles a plausible continuous recording: coherent motion, lighting, textures, and object appearance. Paired-frame scores such as PSNR, SSIM, and LPIPS and video-distribution scores such as FVD each probe different aspects of visual quality; none alone isolates temporal coherence or action fidelity.

State sufficiency (for a task)

The property that the model's internal state at time tt, together with the planned action sequence, is enough to predict task-relevant futures correctly. Sufficiency is always relative to a task and a policy class. A latent state that suffices for "where is the object?" may not suffice for "will the robot's gripper collide with the wall?"

A useful way to organize the analysis is to ask three questions of any video generator:

  1. Does it look right? Perceptual coherence on held-out clips.
  2. Does it behave right? Contact geometry, inertia, occlusion memory, object identity, and count preservation.
  3. Does it help decide right? Uncertainty-aware planning quality in closed loop.

The first question can use perceptual metrics and human inspection, but not one universal score. The second needs probes that isolate a physical or causal property. The third tests the intended decision use in a control loop against baselines. A model can look good but fail a dynamics probe; it can also predict well yet fail in closed loop because of planner search, task-cost design, latency, or model exploitation. Diagnose those components before blaming the predictor.

The gap between visual quality and control utility is an empirical question for each model and task, not a prevalence claim about the field.

Generator Families in One Paragraph Each

Video generators as world models come in a few architectural flavors. We do not re-derive the algorithms here; earlier chapters cover the mechanics, and only the differences that matter for dynamics matter here.

Autoregressive token video models. A tokenizer maps video into discrete codes, and a dynamics model predicts subsequent codes from prior context. Actions, when available, may enter as tokens, embeddings, or another conditioning signal. Recursive sampling permits extended generation but can compound error. Genie (Bruce et al., 2024) combines a spatiotemporal video tokenizer, an autoregressive dynamics model, and a latent-action model. Its interactive interface belongs to Chapter 52.

Diffusion and flow video models. A model denoises or transports a video representation, sometimes generating a whole clip or block jointly rather than one frame at a time. Sampling cost depends on the number of network evaluations and the video extent processed per evaluation; it is not universally KNK N separate frame passes. The Sora report (2024) and UniSim paper (2023) provide examples with different conditioning and evidence. GameNGen (2024) uses diffusion for next-frame game generation and rolls it forward interactively.

Hybrid and interaction-oriented models. These mix representation, dynamics, and rendering choices rather than forming a single latent-state family. GAIA-1 (Hu et al., 2023) predicts discrete video-related tokens autoregressively and uses a diffusion decoder for rendering. IRASim (2024) conditions robot-scene video on actions. DINO-WM predicts visual features for planning, with a separate optional diagnostic decoder; it is not itself a video generator.

The useful comparison is what each system conditions on, what it predicts, how it renders observations, and whether it can roll out under a specified action sequence. Error type cannot be read off the model-family label alone. A tokenized dynamics model may drift in object state or appearance; a diffusion model may also lose identity or miss an action response. Measure these failures directly.

Part IX, Chapter 4, Genie and Interactive Generated Environments, examines latent-action inference and interactive generation in detail. Here we focus on evidence needed before trusting generated video for a specified predictive or control task.

Prediction Horizons: Short, Long, and Mode Collapse

Longer horizons expose more opportunities for model error, although an exact model or a contracting system need not show increasing error at every step. A single average can hide where a rollout fails, so report performance by horizon and condition.

Short-horizon prediction (a handful of frames). Slow motion and stable backgrounds can make frame similarity an easy baseline. Check whether changing a feasible action changes the near future in the expected direction. Prompt sensitivity alone does not demonstrate response to a timed action.

Mid-horizon prediction (hundreds of milliseconds to seconds in a chosen task). A contact event can reveal a failure that frame-similarity scores obscure: bodies interpenetrate, or a grasp does not change the object's motion. The relevant time span depends on the task and frame rate, not on a universal model-failure threshold.

Long-horizon prediction (many transitions at the task's sampling rate). Recursive prediction can compound bias, stochastic error, or off-distribution feedback. Probe position, object identity and count, and camera pose separately: a model might maintain one while losing another.

Rollout drift

In a recursive generator, sampled outputs or an updated internal state condition later predictions. Under additive persistent bias and nonexpanding dynamics, an error term can accumulate approximately linearly in horizon; expanding dynamics can amplify perturbations geometrically. Independent zero-mean errors may instead accumulate in root-mean-square scale near kϵ\sqrt{k}\epsilon under suitable assumptions, while contraction can limit them. One-step error alone therefore does not determine a rollout-error curve.

Two models with similar one-step error can have different long-horizon behavior because their errors differ in bias, correlation, state sensitivity, or exposure to self-generated inputs. Bias is one mechanism, not the only one. The corridor experiment therefore reports recursive RMSE against horizon rather than inferring rollout quality from teacher-forced error.

A second issue is multimodal futures. A state near an unstable choice, or an observation that hides relevant state, can support several plausible outcomes. A stochastic generator can represent multiple samples, but a diffusion, flow, or autoregressive architecture does not guarantee calibrated mode probabilities. A deterministic predictor may return a mean trajectory that corresponds to none of the outcomes. Calibration must be measured against repeated outcomes or a trusted reference distribution.

Occlusion memory. When a task requires tracking a hidden object, its reappearance should be consistent with prior observations and intervening actions, within uncertainty. A generator that always restores the last visible position can fail this test even if the visible background stays coherent. IntPhys (Riochet et al., 2018) evaluates physical plausibility and object-permanence expectations in videos; this chapter's controlled reappearance test asks a narrower, action-conditioned question.

An occlusion probe is useful because the current pixels do not reveal the object's location. A frame generator may still use temporal context or learned regularities; it is not forced to interpolate or freeze. Conversely, an explicit latent state may still track the wrong trajectory. Compare reappearance positions after controlled hidden intervals and actions, and include a simple velocity-extrapolation baseline. The aggregate visibility split in the toy below is only a diagnostic, not an isolated memory test.

Controllability: Prompts, Cameras, Actions

For this chapter, controllability means that a specified input changes the generated future in a predictable, task-relevant way. Three conditioning channels need different tests.

Text prompts. A phrase such as "a red ball rolls right" can specify semantic content and sometimes temporal order. Unless the system validates a mapping from prompt to action timing and units, prompt sensitivity is not evidence that it can execute a controller's timed command.

Camera and viewpoint controls. Specified camera pose, intrinsics, or motion can change the rendered observation while leaving world state fixed. In an embodied system, moving a camera can also move hardware or alter future information; state explicitly which mechanism is under test. Do not score camera-induced image motion as an object-action response.

Time-indexed action sequences. For a planner issuing physical controls, each ata_t has a defined time interval and meaning. Evaluate whether output changes with that action at the appropriate time and magnitude. A free-form text description does not establish an equivalent interface without calibration and validation.

All three can change pixels, but an output change alone does not identify its cause. Use controls that hold the scene and camera fixed while varying a feasible world action, and a separate camera intervention with the scene fixed. If the generator responds to these in indistinguishable ways, its outputs are unsuitable for a planner that needs their effects separated.

Matched-initial-state counterfactual test

From a resettable scene with the same initial state sts_t and observation history hth_t, execute two feasible actions at(0)≠at(1)a^{(0)}_t \ne a^{(1)}_t in the reference environment. Condition the generator on each action while holding other inputs fixed. Where possible, pair random seeds to reduce sampling variance, then repeat across seeds. Compare the distribution and timing of model branch differences with the reference branches; a zero difference is not an error when the two actions truly have the same task-relevant effect.

Matched branches expose weak action sensitivity when the reference actions have distinguishable effects. The toy corridor shows an action-blind predictor returning identical branches for left and right pushes, while the reference branches separate. This controlled comparison is stronger than comparing two attractive clips from unrelated initial states, but it still tests only the chosen actions, histories, and outcome measure.

Training-policy coverage is a separate axis from causal validity. A rarely chosen action has little direct evidence in that context, and a smooth generated response may still be extrapolation. A correct response on supported actions shows local predictive competence for the tested distribution; it does not by itself prove the dataset is adequate or that interventions outside support will generalize. Report action frequencies and evaluate supported and shifted action regimes separately.

Physical Plausibility versus Visual Realism

The following probes separate visual appearance from properties a particular task may require. None is a universal pass/fail rule for every scene.

  • Contact and intersection. Under a rigid-body assumption, rendered solids should not pass through one another when the relevant surfaces are visible.
  • Motion under forces. Changes in velocity should agree with the modeled forces, contacts, friction, and camera motion; deceleration can have an ordinary cause.
  • Occlusion and identity. Track an object through hidden intervals, allowing for uncertainty, and test whether reappearance is consistent with the earlier evidence.
  • Object count. Check counts only in a controlled field of view where entry, exit, creation, and destruction rules are specified.
  • Uncertainty calibration. Compare predictive samples with repeated reference outcomes, including both stochastic and near-deterministic transitions.
  • Decision usefulness. Compare closed-loop task outcomes against appropriate policies and planners under the same observation, action, and compute budgets.

Some properties can be judged from a rendered view; others require an external tracker, simulator state, repeated trials, or a downstream control task. Contact is geometric, while motion tests require a stated dynamics model and forces. An attractive frame cannot substitute for those references.

These benchmarks ask different questions. IntPhys (Riochet et al., 2018) and Physion (Bear et al., 2021) test visual physical prediction or reasoning. Physics-IQ (2025) and VideoPhy (Bansal et al., 2024) directly evaluate physical behavior or commonsense in generated video and report shortcomings in the systems they tested. Their tasks, sampled models, and scoring procedures differ; neither establishes a universal ranking of video generators for control.

On the perceptual side, FVD (Unterthiner et al., 2018) compares feature statistics of sets of videos; frame FID, PSNR/SSIM, and LPIPS use different image-level comparisons. Their scores can be useful, but none isolates whether a generated object's motion follows the commanded action. Do not assume FVD is invariant to a displacement; test whether the chosen metric actually detects the task-critical change in the evaluation set.

Task-relevant physical constraints

A task-relevant constraint is a predicate on a specified physical or estimated state. Examples include a cube ending inside a target region or a gripper avoiding an obstacle. Whether the predicate can be judged from generated frames depends on visibility and the accuracy of any perception head; manual inspection of a few frames is not ground truth for hidden states.

For action-conditioned planning, include at least one task-relevant action-response test and, where the task demands it, a physical constraint test. Perceptual metrics supplement rather than replace those tests. A model that ignores actions can satisfy a static appearance constraint, so vary actions while holding the initial scene fixed.

A Qualification Ladder for Video World Models

We propose a seven-rung ladder. Each rung has a ground truth, a baseline, a mode of false positive, and a characteristic way to interpret failure. Passing a higher rung provides stronger evidence for the capability tested on that task; passing only the early prediction rungs does not establish action-sensitive planning utility. None of these results, on its own, certifies or rules out the model's usefulness for every other task.

The ladder orders tests by what they ask of a model, from local prediction to decision-making. The middle rungs probe hidden-state tracking and action response; the later rungs probe uncertainty and closed-loop use. These are task-relative diagnostics, not a certification scheme: success on one rung neither guarantees the next nor makes every other rung necessary for a useful system. Report failures and evaluation conditions alongside successes.

Rung 1: Held-Out One-Step Prediction

What it tests. Given a held-out frame pair and the intervening action, does the model predict the next frame? This is the gentlest test.

  • Ground truth. The recorded next frame from a held-out trajectory.
  • Baseline. The trivial predictor that outputs the current frame (no motion). Also useful: the previous frame, or the mean frame over the dataset.
  • False positive. A model that has memorized the training set will do unusually well on scenes that resemble training, without generalizing. Fix this by holding out scenes, not just timesteps, and by reporting per-scene performance.
  • Failure interpretation. Beating persistence on a suitable held-out metric supports local predictive ability. Failing it calls for diagnosis of the data, metric, and task before making broader claims.

The one-step rung is easy to over-read. Much of a natural scene is predictable from the current frame alone, so a model can score well without tracking dynamics. Conversely, persistence is unusually strong when changes are small or a pixel metric penalizes valid stochastic alternatives. Treat this as a cheap diagnostic, then test the property the intended task needs.

Rung 2: Open-Loop Long-Horizon Rollouts

What it tests. Condition on a real history and held-out action sequence, generate a KK-frame continuation, and compare with reference outcomes. For recursive models this exposes compounding error; block or joint-clip models can fail differently.

  • Ground truth. A held-out KK-frame clip and its action trace.
  • Baseline. A task-relevant latent dynamics model, such as those discussed in Part VIII, Chapter 2, and naive frame repetition.
  • False positive. Models that generate plausible-looking but dynamics-wrong clips. Watch for slow motion, slowed time, or "frozen" backgrounds that hide drift in the salient object.
  • Failure interpretation. Plot task-relevant state or event error against horizon, per scene, where reference state is available. Pixel error and perceptual scores are supplementary for stochastic futures. A steep curve limits the tested planning horizon; it does not settle whether the model has any world-model utility.

The key discipline on this rung is to report error as a function of horizon rather than as a single averaged value. A model whose error grows slowly with horizon is qualitatively different from one whose error saturates at a fixed offset, even if their average errors coincide. For a planner that scores trajectories out to horizon KK, inspect error near KK as well as the average: late errors can change which action sequence looks best.

Rung 3: Occlusion-Memory Probes

What it tests. Does the model maintain a belief about objects that are hidden but still exist?

  • Ground truth. The true state of the occluded object while hidden, recorded in the simulator or estimated by external tracking from a second view.
  • Baseline. A predictor that freezes the object's position at the moment it disappeared, and a predictor that uses the prior velocity to extrapolate.
  • False positive. A model that appears to remember the object because in the held-out data the object always disappears while stationary. Fix by including occluded episodes with sustained motion.
  • Failure interpretation. If an accessible state estimate or probe is less accurate than velocity extrapolation, that probe gives weak evidence of useful hidden-state tracking in this setting. Output-only generators may need a reappearance test instead. The corridor below illustrates visibility-stratified rollout error, not an isolated proof of memory.

Freezing the object at disappearance is a weak baseline. Velocity extrapolation is stronger, but beating it is not a definition of persistent state: the object's motion may itself be close to constant velocity, or the model's state may encode useful information that this particular probe misses. Use several motion regimes, evaluate after reappearance, and separate state-estimation error from rendering quality.

Rung 4: Matched Counterfactual Action Branches

What it tests. Does the model's action channel actually control its predictions? See the earlier counterfactual note.

  • Ground truth. Two real branches from a reset (a simulator episode with the same initial state and two different action traces).
  • Baseline. A model that ignores the action channel; its two conditionings should produce statistically indistinguishable samples.
  • False positive. A generator that is conditioned on the action in a correlated way because the training policy always applied the same action in that context. To expose this, evaluate on matched pairs where the two actions are both plausible under the training policy.
  • Failure interpretation. Compare the predicted difference between branches with the reference difference. A near-zero predicted response is evidence of an action-channel failure only where the reference actions produce measurably different outcomes.

Matched initial states remove a major source of variation. For stochastic generators, compare branch distributions or couple samples with the same random seed; otherwise sampling noise can masquerade as an action effect. Correct branch differences support action sensitivity under the tested interventions, provided the reference reset and intervention are valid. They do not by themselves establish causal generalization to every policy.

Rung 5: Interventions on Object and Action

What it tests. Can the model respond appropriately to specified scene or action changes, including shifts beyond well-covered training conditions? Recolor an object, move a distractor, change a container's shape, or apply an infrequent action where a reference response can be measured.

  • Ground truth. The real simulator response to the intervention.
  • Baseline. The base model without intervention on the same scene.
  • False positive. Models that pass this test because the intervention is trivially local (recoloring the background does not change dynamics). The interesting interventions perturb dynamics-relevant attributes: mass, friction, contact shape.
  • Failure interpretation. Compare the direction and magnitude of the predicted response with the reference. An unchanged prediction is a failure only if the intervention should change the measured outcome. This connects to the compositionality discussion in Part IV, Chapter 6.

This is a transfer test, but the input must contain enough information to identify the changed dynamics. Recoloring need not affect motion; a visible shape change might. Mass or friction cannot always be inferred from appearance alone. A model's response to a well-specified intervention is more informative than whether its output merely looks different.

Rung 6: Uncertainty and Multimodal Calibration

What it tests. When the true future is ambiguous, does the model's sample distribution reflect that ambiguity, and when it is deterministic, does it concentrate?

  • Ground truth. The empirical distribution of the next observation under the true stochastic dynamics. Often approximated by running the simulator many times.
  • Baseline. A deterministic model that outputs the conditional mean, which will have under-dispersed samples.
  • False positive. A model that is over-dispersed everywhere, including on deterministic transitions; this is easy to produce by cranking sampling temperature and looks like "handling uncertainty" without actually tracking it.
  • Failure interpretation. Use reliability checks for defined events and proper scores where applicable: CRPS for scalar targets, or negative log-likelihood if the model exposes a tractable density. Report sharpness and predictive error separately. Part IV, Chapter 5 discusses broader uncertainty diagnostics.

Calibration asks whether stated probabilities agree with frequencies under a specified evaluation distribution; accuracy asks whether useful outcomes receive high probability or low error. A conditional-mean predictor can have low squared error yet miss a bimodal future. Broad samples may improve coverage while worsening sharpness. Neither property replaces the other, and empirical calibration requires repeated comparable conditions or carefully chosen event bins.

Rung 7: Closed-Loop Replanning

What it tests. Does using the model as part of a planner actually solve tasks? This is the only rung where the criterion is decision quality, not prediction quality.

  • Ground truth. Task success rate and cost under a trusted reference on the same physical or simulated environment.
  • Baseline. A strong model-free policy, a planning baseline using trusted dynamics where available, and a simple policy or random sequence, all under comparable observation and compute budgets.
  • False positive. The model can pass this rung accidentally if the task is easy enough that a strong model-free baseline already succeeds, or if the evaluation environment leaks true state to the policy. The closed-loop eval must be a genuine planning loop with the model in the inner simulation step.
  • Failure interpretation. A perceptually strong generator may still misrank actions. Inspect whether the planner exploits model error, as discussed in Part VI, Chapter 1, and retest the resulting actions in the reference environment. Report the action-distribution shift instead of assuming every closed-loop failure has this cause.

Closed-loop evaluation is decisive when the intended use is planning; other world models may be built for prediction or scientific analysis. A planner and model can jointly exploit errors in the learned simulator, so test chosen actions in the reference environment. When trusted dynamics are available, compare against an otherwise matched planner using them. Report uncertainty and policy shift rather than treating a single success rate as a complete diagnosis.

FVD is supplementary

FVD estimates a distance between finite sets of clips in a chosen feature embedding. A low value may indicate resemblance under that representation, but it does not certify action response, persistent state, causal transfer, or planning utility. Use it as supporting evidence only.

A Planning Interface Over Generated Futures

Let us now make the connection to decisions explicit. Suppose a generator produces action-conditioned videos. A receding-horizon planner over generated futures solves, at each step tt,

At∗∈arg⁡min⁡A∈AH  Eo~t+1:t+H∼pθ(⋅∣ht,A,c) ⁣[∑j=1Hγj−1J(o~t+j,at+j−1)]  +  λ U(A∣ht,c),A_t^* \in \arg\min_{A \in \mathcal{A}^H} \;\mathbb{E}_{\tilde{o}_{t+1:t+H} \sim p_\theta(\cdot\mid h_t,A,c)}\!\left[\sum_{j=1}^{H} \gamma^{j-1} J(\tilde{o}_{t+j}, a_{t+j-1})\right] \;+\; \lambda \, \mathcal{U}(A\mid h_t,c),

where A=(at,…,at+H−1)A=(a_t,\ldots,a_{t+H-1}), H≥1H\geq1 is the horizon, γ∈(0,1]\gamma\in(0,1] discounts later costs, and JJ scores a predicted next observation and its preceding action. The weight λ≥0\lambda\geq0 multiplies a nonnegative estimated uncertainty penalty U(A∣ht,c)\mathcal{U}(A\mid h_t,c). Define g(o,a)≤0g(o,a)\leq0 as satisfying a specified safety rule. A safety filter might reject AA when Pr⁡pθ(∃j∈{1,…,H}:g(o~t+j,at+j−1)>0∣ht,A,c)>δ\Pr_{p_\theta}(\exists j\in\{1,\ldots,H\}:g(\tilde{o}_{t+j},a_{t+j-1})>0\mid h_t,A,c)>\delta for a chosen δ∈[0,1]\delta\in[0,1]. This is a model-estimated chance constraint, not a guarantee of physical safety. Only the first action at∗a_t^* of the chosen sequence is executed; the planner then observes the true next frame and replans.

The summation scores a candidate action sequence by expected discounted cost. The uncertainty term can discourage actions with poorly supported predictions, but only if that estimate is informative about model error. Sample-based risk filters are similarly limited by generator calibration and sample count. Multimodal futures matter because a single plausible sample may conceal rare but costly outcomes; an expectation alone can also conceal tail risk.

This is a model-predictive-control-style interface, with a generator supplying the candidate futures. Part VII, Chapter 1 develops the broader planning setup.

The compute cost is where video-space planning differs from latent-space planning.

  • Video-space rollouts generate HH frames or an equivalent clip for each of MM candidate action sequences. A frame-by-frame diffusion implementation with NN denoising evaluations per frame may use roughly MHNM H N network evaluations; a joint-clip implementation instead denoises a spatiotemporal representation per candidate. Autoregressive token models may need about MHPM H P sequential token predictions for PP tokens per frame, subject to batching, caching, and model design. Runtime and memory depend on token count, representation, network size, sampler, hardware, and horizon—not pixel resolution alone.
  • Latent-space rollouts, including the feature-space prediction approach of DINO-WM in Part IX, Chapter 2, evolve a compressed state without decoding every candidate frame. This can be cheaper when the latent transition is substantially smaller than the video generator, but the factor is implementation- and benchmark-dependent. DINO-WM itself is not a small recurrent network, and visual encoding can remain a significant cost.

The comparison is not that one always wins. Video futures may make some predicted details easier to inspect, while a chosen latent might preserve the exact features a task needs. Either family can represent uncertainty, including through ensembles, at a cost determined by its architecture. Part VIII, Chapter 1 gives decision-centric examples.

The failure modes differ too. Latent-space planners can fail because their latent cannot represent some task-critical aspect of the observation (a numeric readout, a tiny object). Video-space planners can fail because the generator's pixels are right for the wrong reasons or because the perception head that computes JJ is not robust to generated artifacts. An uncertainty penalty can help when its score is informative about task-relevant model error, even if it is not perfectly calibrated. It cannot reliably protect against confidently wrong predictions assigned a low penalty; those predictions may lead the planner to choose poor actions.

Receding-horizon feedback can limit the effect of rollout error by replanning from fresh observations. It cannot make an action-independent inner model distinguish consequences of candidate actions. For tasks that require choosing among actions, matched counterfactual branching is therefore a central diagnostic, though the test must cover action differences that matter in the task.

Worked Example: An Occluded Object on a 1D Corridor

We now build a small, CPU-sized scalar-position test bed for selected ladder diagnostics. The track is unbounded in the simulation; "corridor" names the one-dimensional coordinate and fixed occluder, not physical walls.

An object moves on a 1D corridor. It starts on the left and receives recorded actions from a right-biased behavioral policy; many, but not all, sampled trajectories pass through a fixed occluder band. A state xt∈Rx_t \in \mathbb{R} is a scalar position; actions are in {−1,0,+1}\{-1, 0, +1\}. The dynamics are a damped velocity response:

vt+1=d vt+at,xt+1=xt+g vt+1+ϵt,v_{t+1} = d \, v_t + a_t, \qquad x_{t+1} = x_t + g \, v_{t+1} + \epsilon_t,

with d=0.5d = 0.5, g=0.4g = 0.4, and ϵt∼N(0,σ2)\epsilon_t \sim \mathcal{N}(0, \sigma^2) with σ=0.02\sigma = 0.02. Rewriting, the per-step displacement satisfies the ARX model

Δt  =  xt+1−xt  =  d Δt−1+g at+ϵt−d ϵt−1,\Delta_t \;=\; x_{t+1} - x_t \;=\; d \, \Delta_{t-1} + g \, a_t + \epsilon_t - d\,\epsilon_{t-1},

which is a linear input-output relation we can fit from data, with one lagged displacement (Δt−1\Delta_{t-1}) and one control input (ata_t). The composite residual ϵt−dϵt−1\epsilon_t-d\epsilon_{t-1} is an MA(1) process, not independent noise; ordinary least squares below is a small-data illustration rather than an exact recovery guarantee. This is an input-output system-identification example related to Part II, Chapter 3.

The observation function is simple:

ot={xtif xt∉[4,6]∅otherwise (occluded).o_t = \begin{cases} x_t & \text{if } x_t \notin [4, 6] \\ \varnothing & \text{otherwise (occluded)}. \end{cases}

A predictor that only receives the visible position has no new position measurement during occlusion. It can still propagate a state estimate using past motion and actions. This scalar-position surrogate tests action-conditioned extrapolation; it is not a video generator or, by itself, a complete test of visual object permanence.

We compare three predictors on the same held-out trajectories:

  • Persistence. Predicts Δt=0\Delta_t = 0: the object stays where it was.
  • Action-blind ARX. Fits Δt=c1Δt−1\Delta_t = c_1 \Delta_{t-1} with no action input.
  • Action-conditioned ARX. Fits Δt=c1Δt−1+c2at\Delta_t = c_1 \Delta_{t-1} + c_2 a_t.

The three predictors are linear with at most two fitted parameters. This controlled comparison focuses on the action input rather than differences in network capacity. They are evaluated on the same held-out trajectories; the fitted ARX models use the same visible training transitions, while persistence has no training step.

Setting up the corridor

We start with imports and constants.

In[3]:
Code
import numpy as np

OCCL_LO, OCCL_HI = 4.0, 6.0
DAMP, GAIN = 0.5, 0.4
PROCESS_NOISE = 0.02


def is_visible(x):
    return not (OCCL_LO <= x <= OCCL_HI)

The simulate function below produces a trajectory by sampling actions from a fixed behavioral policy and applying the ARX dynamics. Recorded actions are what the world model is allowed to see.

In[4]:
Code
def simulate(steps, seed):
    """Roll out the corridor dynamics and record timed actions and visibility."""
    r = np.random.default_rng(seed)
    x, v = float(r.uniform(1.0, 3.0)), 0.0
    xs, vs, acts, vis = [], [], [], []
    for _ in range(steps):
        a = int(r.choice([-1, 0, 1], p=[0.2, 0.25, 0.55]))
        xs.append(x)
        vs.append(v)
        acts.append(a)
        vis.append(is_visible(x))
        v = DAMP * v + a
        x = x + GAIN * v + r.normal(0.0, PROCESS_NOISE)
    return (
        np.array(xs),
        np.array(vs),
        np.array(acts, float),
        np.array(vis, bool),
    )

We now generate training and test sets with disjoint random seeds. This holds out full trajectories under the same simulator and policy; it does not test transfer to new scenes, mechanisms, or action policies.

In[5]:
Code
STEPS = 30
train = [simulate(STEPS, seed) for seed in range(24)]
test = [simulate(STEPS, seed) for seed in range(100, 110)]

occluded_fraction = float(np.mean([~t[3] for t in test]))
Out[6]:
Console
Held-out trajectories: 10
Fraction of held-out timesteps with the object occluded: 0.250

The aggregate fraction tells us whether the held-out set contains hidden timesteps; it does not establish that every trajectory crosses the band.

Fitting the three predictors

We fit the ARX coefficients by ordinary least squares, using only transitions whose three required positions are visible. This is an observation-availability restriction in the scalar surrogate, not a claim that a video system sees exact positions. Persistence has no fitted parameters.

In[7]:
Code
def arx_dataset(trajs):
    rows, targets = [], []
    for xs, _, acts, vis in trajs:
        for t in range(1, len(xs) - 1):
            if vis[t - 1] and vis[t] and vis[t + 1]:
                rows.append([xs[t] - xs[t - 1], acts[t]])
                targets.append(xs[t + 1] - xs[t])
    return np.array(rows), np.array(targets)


X_all, y_all = arx_dataset(train)
coefs_act = np.linalg.lstsq(X_all, y_all, rcond=None)[0]
coefs_blind = np.linalg.lstsq(X_all[:, :1], y_all, rcond=None)[0]
c1_act, c2_act = float(coefs_act[0]), float(coefs_act[1])
c1_blind = float(coefs_blind[0])
Out[8]:
Console
Fitted coefficients (training set only)
  action-blind ARX:          c1 = 0.610                 (true damping d = 0.5)
  action-conditioned ARX:    c1 = 0.499, c2 = 0.399   (true damping d = 0.5, true gain g = 0.4)

The action-conditioned fit recovers the true damping and gain closely on this training set. The action-blind fit has no action term and estimates a single lag coefficient of about 0.610, above the true damping of 0.5. That coefficient minimizes squared prediction error under a misspecified model and this data-collection policy; it is not an estimate of either physical parameter. The action-blind predictor cannot distinguish two different actions applied after the same observed history.

The coefficient estimates can be compared directly to the true dynamics. The cell below prepares the comparison, and the figure then shows the fitted and true coefficients side by side.

In[9]:
Code
coefficient_groups = ["action-blind ARX", "action-conditioned ARX"]
coefficient_bars = {
    "damping c1": np.array([c1_blind, c1_act]),
    "gain c2": np.array([0.0, c2_act]),
}
true_coefficients = {"damping d": DAMP, "gain g": GAIN}
coefficient_x = np.arange(len(coefficient_groups))
coefficient_width = 0.35
Out[10]:
Visualization
Grouped bar chart comparing fitted and true ARX damping and gain coefficients.
Fitted ARX coefficients compared with the true damping and action-gain values. The action-blind predictor omits the action input; its single fitted lag coefficient is about 0.610 and is not the true damping of 0.5. The action-conditioned fit is near the true damping of 0.5 and gain of 0.4 on this training set.

Recursive rollout and diagnostics

For an ARX predictor, the transition function is Δ^=c1Δ^prev+c2a\hat{\Delta} = c_1 \hat{\Delta}_{\text{prev}} + c_2 a. Multi-step rollouts feed the predictor's own estimate back in:

In[11]:
Code
MODELS = {
    "persistence": (0.0, 0.0),
    "action-blind ARX": (c1_blind, 0.0),
    "action-conditioned ARX": (c1_act, c2_act),
}


def rollout(xs, acts, coeffs, start=1):
    """Recursive prediction of x[start+1..end] from x[start], Delta[start-1], and actions."""
    x_hat = xs[start]
    d_hat = xs[start] - xs[start - 1]
    preds = []
    for k in range(start, len(xs) - 1):
        d_hat = coeffs[0] * d_hat + coeffs[1] * acts[k]
        x_hat = x_hat + d_hat
        preds.append(x_hat)
    return np.array(preds)

We compute one-step errors on held-out transitions with visible input and target positions, then recursive errors from a visible initial pair. The one-step metric is a local prediction diagnostic with reference state supplied at each eligible transition. The recursive metric propagates the predictor's own state without later position updates; its errors can accumulate.

In[12]:
Code
def one_step_errors(trajs):
    out = {}
    for name, coeffs in MODELS.items():
        errs = []
        for xs, _, acts, vis in trajs:
            for t in range(1, len(xs) - 1):
                if not (vis[t - 1] and vis[t] and vis[t + 1]):
                    continue
                d_prev = xs[t] - xs[t - 1]
                d_hat = coeffs[0] * d_prev + coeffs[1] * acts[t]
                errs.append(abs(d_hat - (xs[t + 1] - xs[t])))
        out[name] = float(np.mean(errs))
    return out


one_step = one_step_errors(test)

Recursive errors are averaged per horizon index and grouped by whether the true object was visible at that step. Because group membership also varies with trajectory and horizon, this descriptive split does not isolate an occlusion effect.

In[13]:
Code
HORIZON = 20
recursive = {name: np.zeros(HORIZON) for name in MODELS}
visible_err = {name: [] for name in MODELS}
occluded_err = {name: [] for name in MODELS}

for xs, _, acts, vis in test:
    for name, coeffs in MODELS.items():
        preds = rollout(xs, acts, coeffs, start=1)
        for k in range(min(HORIZON, len(preds))):
            t = k + 2
            e = abs(preds[k] - xs[t])
            recursive[name][k] += e**2
            (visible_err if vis[t] else occluded_err)[name].append(e)

for name in MODELS:
    recursive[name] = np.sqrt(recursive[name] / len(test))

visible_mean = {
    n: (float(np.mean(v)) if v else 0.0) for n, v in visible_err.items()
}
occluded_mean = {
    n: (float(np.mean(v)) if v else 0.0) for n, v in occluded_err.items()
}
Out[14]:
Console
One-step mean absolute error
  persistence                0.3695
  action-blind ARX           0.2994
  action-conditioned ARX     0.0205

Recursive RMSE at horizon 20
  persistence                5.6288
  action-blind ARX           5.1456
  action-conditioned ARX     0.0924

Mean error during occluded steps (recursive)
  persistence                2.3070
  action-blind ARX           1.7684
  action-conditioned ARX     0.0341

On these simulated held-out trajectories, the action-conditioned fit has lower one-step and recursive position error than the action-blind baselines. Since the models are comparably small, the action input is an important controlled difference here. The size and shape of this gap are properties of this simulator, policy, and fitted sample—not a forecast for real video generators.

Counterfactual action response

We now run the matched-initial-state counterfactual test from a specific state just left of the occluder. The object starts at x0=3.0x_0 = 3.0 with velocity v0=0.3v_0 = 0.3, and we apply a constant action of +1+1 versus −1-1 for 12 steps. The reference true_branch sets the zero-mean process noise ϵt\epsilon_t to zero, so its paths are conditional means for this linear simulator. The ARX model_branch initializes its prior displacement to gv0=0.4(0.3)=0.12g v_0=0.4(0.3)=0.12, consistent with that noiseless reference. Under each model, we compare the final position gap.

In[15]:
Code
CF_STEPS = 12
X0, D0 = 3.0, 0.3


def true_branch(a_const, steps=CF_STEPS):
    x, d = X0, D0
    traj = [x]
    for _ in range(steps):
        d = DAMP * d + a_const
        x = x + GAIN * d
        traj.append(x)
    return np.array(traj)


def model_branch(coeffs, a_const, steps=CF_STEPS):
    x, d = X0, GAIN * D0
    traj = [x]
    for _ in range(steps):
        d = coeffs[0] * d + coeffs[1] * a_const
        x = x + d
        traj.append(x)
    return np.array(traj)


true_sep = float(true_branch(1.0)[-1] - true_branch(-1.0)[-1])
pred_sep = {
    name: float(model_branch(coeffs, 1.0)[-1] - model_branch(coeffs, -1.0)[-1])
    for name, coeffs in MODELS.items()
}
Out[16]:
Console
True final-position separation between right and left branches: 17.6004
  persistence                predicted separation: 0.0000
  action-blind ARX           predicted separation: 0.0000
  action-conditioned ARX     predicted separation: 17.5450

Persistence and action-blind ARX produce zero branch separation because neither uses the action input. The action-conditioned model reproduces the reference separation closely in this toy. The branches share an initial state and drift velocity and differ only in the constant applied action. Predicting no difference here fails to represent this intervention's known effect; it is not a general verdict about every action or environment.

The separation numbers can also be unpacked into the full branch trajectories. The cell below evaluates the true and model branches for the right-push and left-push counterfactuals, and the figure overlays them.

In[17]:
Code
cf_steps_axis = np.arange(CF_STEPS + 1)
cf_true_right = true_branch(1.0)
cf_true_left = true_branch(-1.0)
cf_act_right = model_branch(MODELS["action-conditioned ARX"], 1.0)
cf_act_left = model_branch(MODELS["action-conditioned ARX"], -1.0)
cf_blind_right = model_branch(MODELS["action-blind ARX"], 1.0)
cf_blind_left = model_branch(MODELS["action-blind ARX"], -1.0)
Out[18]:
Visualization
Line chart of true and predicted counterfactual branches for right and left actions.
Matched-initial-state counterfactual branch trajectories for constant right and left actions over twelve steps. The true branches diverge, the action-conditioned model follows both closely, and the action-blind model predicts one trajectory for both actions in this toy.

A rollout picture

Before plotting, we precompute the horizon curves and pick one held-out trajectory to display. All computation happens here, in plain Python; the figure cells below are presentation-only. This particular trajectory enters the occluder band only at its final two timesteps; the aggregate test below covers longer occluded intervals.

In[19]:
Code
example = test[3]
xs_ex, _, acts_ex, vis_ex = example
example_preds = {
    name: rollout(xs_ex, acts_ex, coeffs, start=1)
    for name, coeffs in MODELS.items()
}
example_t = np.arange(2, 2 + len(next(iter(example_preds.values()))))

horizons = np.arange(1, HORIZON + 1)
Out[20]:
Visualization
Line chart of true object position and three model predictions along a 1D corridor with a shaded occluder band.
Recursive predictions on one held-out corridor trajectory for persistence, action-blind ARX, and action-conditioned ARX predictors. The action-conditioned model tracks the true position, including the two final timesteps in the occluder band. Persistence holds its start-of-rollout position; the action-blind model drifts without responding to the recorded actions.

The action-conditioned curve stays close to the truth on this trajectory, including its two occluded timesteps. Persistence holds its start-of-rollout position and misses the subsequent motion. The action-blind predictor extrapolates from its prior displacement but does not respond to the recorded actions, so it too fails to keep up. The visibility-stratified comparison comes from the aggregate results below, not this one example.

Out[21]:
Visualization
Line chart of recursive RMSE vs rollout horizon for three linear predictors.
Recursive root-mean-square error as a function of rollout horizon, averaged over ten held-out trajectories, for persistence, action-blind ARX, and action-conditioned ARX predictors. The action-conditioned model stays far below the two action-blind baselines. The baseline errors rise overall across the displayed horizons, with local dips rather than strictly monotone growth.

The baseline curves rise overall as recursive predictions drift away from the true trajectory, although neither rises at every horizon. Persistence repeats an old position; action-blind ARX extrapolates motion without the recorded actions. The action-conditioned predictor stays much closer to the held-out trajectories. These curves compare predictors on this toy data; their shapes do not establish a universal error-growth law.

Occlusion memory and counterfactual response

We summarize the two most diagnostic scores in a single side-by-side figure.

Out[22]:
Visualization
Grouped bar chart of mean position error for three models on visible vs occluded timesteps.
Mean recursive position error for persistence, action-blind ARX, and action-conditioned ARX predictors, grouped by whether the true object was visible or occluded at each timestep. The action-conditioned predictor has much lower error in both groups. Since these are accumulated rollout errors rather than matched one-step errors, this split does not isolate the effect of occlusion.
Bar chart of counterfactual branch separation for three models with a dashed true-value line.
Final-position separation between the right-push and left-push counterfactuals from a matched initial state for each predictor. Persistence and action-blind ARX produce zero separation, while only the action-conditioned model matches the true value marked by the dashed line.

The left panel shows lower error for the action-conditioned predictor in both the visible and occluded groups. Its visible-step mean is about 0.06, compared with about 2.91 for action-blind ARX and 3.32 for persistence; the same ordering holds behind the occluder. Because errors accumulate over different trajectories and horizons, this grouping cannot by itself attribute the gap to occlusion. The right panel tests a different question: with the initial state held fixed, only the action-conditioned predictor separates the two commanded action branches by roughly the true amount. The toy permits this controlled comparison because its state and action mechanism are known.

This scalar surrogate does not benchmark a video system. It demonstrates horizon-resolved position error, a visibility-stratified descriptive check, and matched counterfactual action branches where reference state and dynamics are known. A proper occlusion-memory probe would additionally control horizon and motion regime, or score reappearance after a matched hidden interval. Real evaluations need task-specific baselines, held-out conditions, and reference outcomes appropriate to the claim.

Limitations and Impact

Video generators as world models have real potential and real limits, and the two are difficult to disentangle by looking at clips.

Many video generators are trained primarily on observational clips, which need not contain the timed control signals required by a planner. Action-labelled interaction data can help, but its causal value depends on coverage, the behavior policy, confounding, and whether resets or matched interventions are possible. Collecting repeated interventions can be expensive in physical settings; simulation can provide greater control but introduces its own transfer problem. Neither a conditioning token nor a large video corpus alone resolves these issues.

The second limitation is compute. Sampling many high-dimensional futures within a control deadline may be costly, but latency depends on architecture, sampler, candidate count, horizon, hardware, and batching. A 10 Hz controller has a 100 ms cycle; whether a particular denoising budget fits that cycle requires measurement on its target hardware. Latent rollouts may avoid repeatedly decoding full frames, while video rollouts expose details that a compressed state might discard. These are design choices to benchmark against the task's latency and information requirements, not a fixed speed hierarchy.

The third limitation is diagnostic coverage. FVD, PSNR, and LPIPS measure different aspects of generated clips; none is designed to establish correct action effects or task success. Their sensitivity to a small but decisive event depends on the event, representation, aggregation, and metric. Pair perceptual scores with targeted state/event checks, matched interventions where feasible, and closed-loop outcomes for planning claims.

Each diagnostic has a baseline and failure mode. A generator may look calibrated on coarse event bins while missing conditional structure; a planner may exploit simulator error or an evaluation artifact. The ladder is a scaffold, not a proof or a requirement to pass all seven tests. Choose the rungs that match the intended use, disclose their limits, and test policy and scene shifts when relevant. Parts X and XI develop applications and evaluation questions.

Rendered futures can make some errors easy to spot: an object may vanish across an occlusion or a predicted contact may look wrong. Latent models are not inherently opaque, though; probes, decoded predictions, and decision tests can reveal errors that a clip review misses. Visual inspection is a useful diagnostic alongside quantitative and interventional evaluation. Part IX, Chapter 4 turns to interactive generated environments and latent-action interfaces.

Summary

  • An action-conditioned video generator samples pθ(o~t+1:t+K∣ht,at:t+K−1,c)p_\theta(\tilde{o}_{t+1:t+K} \mid h_t, a_{t:t+K-1}, c). A timed, calibrated action channel is needed for the intervention-based planning tests in this chapter; pixels alone do not provide it.
  • Observational conditioning p(ot+1∣ht,at)p(o_{t+1} \mid h_t, a_t) need not equal the interventional quantity p(ot+1∣do⁡(at),ht)p(o_{t+1} \mid \operatorname{do}(a_t), h_t) when action assignment is confounded. Narrow or absent action support instead limits what can be estimated from logged data; it does not by itself prove that the two quantities differ.
  • Perceptual coherence, physical plausibility, and decision usefulness are different properties. Perceptual scores provide partial evidence about appearance; task-relevant constraints and counterfactual probes test specific physical or causal behavior; closed-loop evaluation tests whether the predictions help decisions. None alone certifies the other properties.
  • The qualification ladder has seven rungs: held-out one-step prediction, open-loop rollouts, occlusion-memory probes, matched counterfactual action branches, intervention tests, uncertainty calibration, and closed-loop replanning. Each has its own baseline, false positive, and failure interpretation.
  • A video-space planner can score generated futures with a task cost and estimated uncertainty, optionally filtering candidates by model-estimated risk. Visual inspectability may help, but runtime, perception errors, and model calibration constrain its use.
  • The scalar corridor toy separates action-conditioned and action-blind linear predictors on this policy and simulator. It shows horizon-dependent errors and matched action-branch response; its visible-versus-occluded error grouping is descriptive, not a controlled proof of memory. It makes no direct claim about a named video generator.
  • The remaining frontier is persistent, controllable, low-latency interaction. Generating a convincing clip is not the same as running an environment you can play inside. The next chapter picks up exactly there.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about video generators as candidate world models.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026videogenerators, author = {Michael Brenndoerfer}, title = {Video Generators as Candidate World Models}, year = {2026}, url = {https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Video Generators as Candidate World Models. Retrieved from https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models
MLAAcademic
Michael Brenndoerfer. "Video Generators as Candidate World Models." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models>.
CHICAGOAcademic
Michael Brenndoerfer. "Video Generators as Candidate World Models." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Video Generators as Candidate World Models'. Available at: https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Video Generators as Candidate World Models. https://mbrenndoerfer.com/writing/video-generators-as-candidate-world-models

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.