Part of World Models Handbook
World-model evaluation uses a five-rung ladder covering one-step prediction, open-loop rollout, intervention, closed-loop planning, and transfer.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
The World-Model Evaluation Ladder
Suppose two dynamics models are evaluated on the same held-out robot transitions. One has lower one-step squared error. Is that enough to choose the model for a controller? The worked example below gives a concrete counterexample: the one-step winner is not the best tracker, and changing the environment changes the control ranking again. The example is synthetic; it is not a report of a deployed robot.
The difference comes from how a model is used. A one-step test supplies a real state as input. An open-loop simulator instead feeds predictions back into the model. A controller queries candidate actions, executes one in the real environment, and replans after observing the successor. These uses expose different errors.
For a model trained by one-step teacher forcing, the training pairs contain observed inputs and successors. Deployment need not use that same protocol: a rollout can visit predicted states, a planner can consider alternative actions, and a decision objective can weight errors differently. Other world models train on sequences or multi-step objectives, and many deployed controllers retain feedback. The distinction here is between evaluation protocols, not a universal account of world-model training.
Think of checking individual steps of a worked solution versus checking the assembled solution. Both checks are useful, but one does not replace the other. In a world model, recursion and action selection are part of that assembly.
The world-model evaluation ladder organizes five complementary questions.
- One-step prediction. Given a real transition , how accurate is a point prediction or predictive distribution under the stated test law? In the point-prediction test here, every input is a real state.
- Open-loop rollout. From a common initial state and fixed actions, how far does a recursively predicted trajectory depart from the true trajectory, without observation corrections?
- Counterfactual intervention. How accurately does the model predict the response to changing the actions? Observational identification and unit-specific counterfactuals need assumptions beyond predictive accuracy.
- Closed-loop planning. How well does the model-planner combination act in the true environment when feedback is available?
- Transfer. Does that combination remain useful under a declared change in dynamics, observations, initial-state regime, objective, or constraints?
These are not strictly nested tests. A good result on one protocol does not certify another protocol with different inputs or objectives. A scalar score can be useful for a declared deployment objective. The ladder supplies a diagnostic vector alongside it, so that a scalar does not conceal untested uses.
The five rungs are teacher-forced one-step prediction, recursive open-loop rollout, intervention-response comparison, closed-loop decision quality, and transfer. Each result needs its distribution, horizon, action protocol, model inputs, and evaluation budget.
Ranking models requires that context. A prediction winner, simulator winner, and control winner can be different models; they need not differ in every experiment.
We keep observation (), state (), action (), model (), planner, and objective separate. Unless stated otherwise, is the fully observed state of a controlled dynamical system, as in [Part I, Ch 2: What Is a World Model?] and [Part III, Ch 2: Sufficient State and State Abstraction]. Under partial observation, a belief or learned latent state must also be evaluated: an inaccurate upstream state estimate can affect the planner even when the transition model is accurate on its intended inputs. The ladder makes the object being scored explicit.
One-step prediction
The one-step test scores a prediction at an observed input without feeding the model's previous prediction back into it. That makes it convenient for regression-style models, but the resulting score does not test recursive simulation or action selection. Teacher forcing also does not make transitions statistically independent; neighboring observations can remain correlated.
Let be a controlled Markov transition kernel. It returns a probability law over successors, not a successor state vector. Let be a common evaluation distribution over , and let be the conditional mean. For a deterministic point predictor with finite squared risk, assume the successor and predictor have finite second moments under this joint law. Then
The first term measures conditional-mean prediction error. The second averages conditional squared noise over the same evaluation law. It depends on and , but not on the deterministic point predictor. To obtain the identity, write and expand the squared norm. The conditional expectation of the cross term is zero because . Noise is already part of ; it is not added again to an expected squared error against sampled successors.
Where:
- and : the joint state-action input drawn from ;
- : the sampled successor drawn from ;
- : the conditional mean successor;
- : a deterministic point prediction;
- : the average conditional squared noise norm.
Training data may be collected by a behavior policy and a state-estimation pipeline. The evaluation law can be chosen to match that collection regime or to represent a different use. The conditional-mean decomposition applies to the selected law; it does not establish that this law covers deployment.
On held-out transitions the empirical score is
Here indexes the evaluation transitions, is their count, and each is the observed successor of input . With independent evaluation draws from the specified joint law, conditional on a fixed fitted predictor, this sample average is unbiased for its population risk. Disjoint row indices alone do not establish independence from the fitted predictor or independence among evaluation rows. For trajectory data, the sampling and splitting unit needs to be stated.
Three quantities help interpret the score. Let be a fixed hypothesis class, let minimize over that class when such a minimizer exists, and let be the fitted model.
- Approximation error is , the best conditional-mean error available in the class.
- Fitting excess is . Finite-data estimation and imperfect optimization can both contribute to it.
- Evaluation error is , a signed discrepancy due to the evaluation sample.
These definitions give an exact bookkeeping identity when the minimizer exists:
The final bracket can be negative; the first two excess terms cannot. More training data does not remove approximation error in a fixed misspecified class. Under suitable sampling, identifiability, and consistency assumptions it can reduce estimation excess over the best member, even when the class is misspecified.
For example, let with and , but allow only constant predictors . The best constant is zero and its approximation risk is one. Fitting the sample mean on independent draws gives expected excess risk . The approximation error remains while estimation excess vanishes. There is no universal rule that greater capacity always increases estimation error; that tradeoff depends on the class, algorithm, regularization, and data.
For a comparison on common evaluation transitions, uncertainty belongs to the paired loss differences. The marginal standard error of one model's MSE is not the standard error of the difference. A numeric gap such as cannot be judged from alone.
Finally, averaging signed prediction errors is not the same as averaging squared errors. Positive and negative mistakes do not cancel in MSE. A conditional mean can be the optimal point predictor yet fail to represent multimodal futures needed by a particular planner. That is a reason to evaluate a predictive distribution or the downstream task separately, not to claim that smoothing is automatically low-error.
A small two-component oscillator makes the distinction concrete without copying the answer into the predictor. Its true state is , and its transition rotates that vector by . The same misspecified model is used in both tests: it rotates by the correct angle but expands the state by a factor of . Teacher forcing resets each input to the true state; open-loop rollout feeds the model's previous prediction back into itself.
import numpy as np
t = np.arange(41)
omega = 2 * np.pi / 20
rotation = np.array(
[[np.cos(omega), np.sin(omega)], [-np.sin(omega), np.cos(omega)]]
)
true_z = np.column_stack((np.sin(omega * t), np.cos(omega * t)))
oscillator_model = 1.02 * rotation
teacher_z = (oscillator_model @ true_z[:-1].T).T
rollout_z = np.zeros_like(true_z)
rollout_z[0] = true_z[0]
for step in range(40):
rollout_z[step + 1] = oscillator_model @ rollout_z[step]
teacher_mse = float(np.mean(np.sum((teacher_z - true_z[1:]) ** 2, axis=1)))
oscillator_rollout_rmse = float(
np.sqrt(np.mean(np.sum((rollout_z[1:] - true_z[1:]) ** 2, axis=1)))
)
constant_mean_mse = float(np.mean((true_z[:, 0].mean() - true_z[:, 0]) ** 2))Teacher-forced vector MSE: 0.000400 Open-loop vector RMSE: 0.643858 Constant-mean scalar MSE over all 41 states: 0.487805
The teacher-forced vector MSE is , because every true state has norm one. In open loop the norm becomes , so the error grows even though the rotation angle is correct. A constant-mean predictor does not track this sine wave under teacher forcing: its scalar MSE across the 41 plotted states is , not zero.

Three confounds deserve special attention.
- Distribution coverage. The evaluation law need not match the state-action occupancy induced by deployment. If data cover only near-upright robot states, predictions far outside that region are extrapolations unless other data or structural constraints inform them. A state absent from a finite sample is not necessarily outside the support of the population distribution.
- Stochasticity. Point-prediction MSE and distributional scores answer different questions. A probabilistic model can be scored by its mean for MSE, or by a suitable proper scoring rule for its distribution. Comparing a random sample from one model with a point estimate from another changes the scoring convention. Uncertainty-aware planning needs separate tests. [Perceptual and Generative Evaluation] is planned to examine observation-quality scoring, not certify a planning system.
- Observation error. A state estimated from an image or point cloud can be inaccurate before the dynamics predictor receives it. Separating perception and transition errors needs component-specific measurements, identifiable model/noise assumptions, or interventions. Aggregate prediction error alone does not distinguish those components. Their units, directions, dependence, and propagation matter; two nominal error magnitudes alone do not determine the combined prediction error.
A one-step result is therefore a distribution-specific check. It can be decisive if that is the intended task. It cannot, by itself, establish rollout quality, intervention accuracy, or useful control.
Open-loop rollout
An open-loop rollout feeds predicted states back into the model while using a fixed action sequence. This tests the composed transition map, including how errors at one input affect later predictions. It does not imply that the one-step evaluation samples were independent.
First distinguish two environments. In a stochastic environment, . For the pathwise bound below, restrict attention to deterministic dynamics with a state-valued map :
Where:
- : deterministic true dynamics, returning a successor state;
- : a transition probability kernel; in this deterministic special case it places all probability at ;
- and : the true and predicted states;
- : the common action at step ;
- : the positive integer rollout horizon.
Both branches use the same initial state and action sequence , with no observation corrections during prediction. Trajectory RMSE and endpoint error summarize different parts of the path:
The RMSE averages the successor states, excluding the common zero-error initial state. Endpoint error scores only the final state. Neither dominates as a metric: an application that needs the whole path may need intermediate errors, while terminal decisions may prioritize the endpoint. For stochastic rollouts, state whether the score compares distributions, expected quantities, or coupled sampled paths; unrelated noise draws can contribute to a pathwise difference.
To derive an error envelope, let and assume:
- common fixed actions and a common initial state;
- a uniform satisfying for every compared true/predicted input pair and action along the trajectories;
- a uniform local error bound on every true input encountered.
A sufficient stronger condition is that both bounds hold throughout a reachable region containing those trajectories. Add and subtract , then use the triangle inequality:
Here , bounds amplification in the stated norm, and bounds local transition error. Iterating this inequality yields
These are upper envelopes, not observed growth laws. If the same uniform assumptions hold for every horizon and , the envelope does not exceed . At it is linear, and at it is exponential. Actual errors can remain zero even under a loose bound with .
The distinction matters in three ways.
- Stability: one-step average loss does not measure the amplification of perturbations. For differentiable maps, a suitable uniform bound on the induced Jacobian norm can establish the Lipschitz condition on an appropriate region. A spectral radius at one state is not a general bound on products of changing Jacobians.
- Horizon: report the horizon and the trajectory statistic together. The same models can change order as the horizon changes.
- Ranking: similar held-out average risks do not determine rollout ranking or supply the uniform used above. A contractive model can sometimes outperform a less biased expansive model in rollout, but neither stability nor local accuracy alone fixes the ranking.
The following plot holds fixed and draws the envelope for three chosen values of . It plots the bound, not measured trajectory errors.
import matplotlib.pyplot as plt
import numpy as np
from book_plot_style import PALETTE, polish_axes, theme_color, use_book_style
horizon = np.arange(1, 41)
e1 = 0.01
L_values = [0.9, 1.0, 1.1]
bounds = {}
for L in L_values:
if L == 1.0:
bounds[L] = e1 * horizon
else:
bounds[L] = e1 * (L**horizon - 1) / (L - 1)
The envelope assumes accuracy at the inputs reached by the rollout. A held-out average error under does not establish that condition. Predicted trajectories may enter regions with few observed samples, or regions outside the support of altogether. These are different cases: a finite sample can omit inputs that still have positive population probability.
This mismatch between observed training inputs and model-generated inputs is one mechanism behind compounding-error problems. It motivates training methods that explicitly confront model-generated states, such as Talvitie's self-correcting models. The paper analyzes that mechanism under its stated assumptions; it does not imply that every high-capacity model fails at long horizons. [Part VII, Ch 4: Policies Learned in Imagination] is planned to cover the planning use of such rollouts.
Counterfactual intervention
The intervention rung asks how predictions change when the action protocol changes. A controlled simulator comparison and an observationally identified causal effect are different ways of obtaining an answer. Predictive accuracy on collected transitions is not enough to establish the latter.
For a deterministic environment , fix an initial state and two action sequences, baseline and intervention . Run both sequences from that common state in the true and model systems. The final response is
The response error is . Superscripts identify the action branch; the two branches share . Changing only actions isolates their effect within this deterministic experiment. If the two model branch errors have an equal additive component, that component cancels in the response difference. Unequal or action-dependent errors need not cancel.
In a stochastic system, a transition kernel specifies each branch's marginal law, not how their exogenous randomness is paired. A unit-specific counterfactual comparison needs that additional causal structure. For example, two branches can each be Bernoulli while sharing the same outcome, taking complementary outcomes, or using independent outcomes. Their marginal means coincide, but their paired differences have different distributions. A structural transition can specify a shared sequence across the branches; a kernel alone does not supply it. Pearl's structural-counterfactual account makes this dependence on background variables explicit.
Alternatively, define an average intervention response as . That difference of marginal expectations does not require a cross-branch noise coupling, though estimating it from observational data still requires identification assumptions. Neither definition should be silently substituted for the other.
For nonparametric adjustment-based observational identification, specify a causally adequate pre-action state, consistency, conditional exchangeability (no unmeasured action-outcome confounding), and positivity for the actions being compared. Multi-step policies need corresponding sequential conditions. State-dependent action selection is not automatically confounding after conditioning on an adequate state. A deterministic behavior policy can nevertheless leave alternative actions unsupported. Extrapolating their effects requires additional structural assumptions. The roles of consistency, exchangeability, and positivity are developed in Hernán and Robins' Causal Inference: What If.
An observational conditional describes successors in collected data. An interventional distribution describes successors when actions are set. Equality requires causal assumptions. Coverage concerns which action responses can be estimated; it is distinct from absence of confounding. A controlled deterministic simulator test does not establish observational identification.
Two practical distinctions help interpret the score.
- Data and assumptions: unsupported action effects cannot be recovered from observational accuracy alone. Additional experiments, structural knowledge, or scientific priors can constrain them. A poor response result may reflect data coverage, incorrect assumptions, or model error; changing architecture alone does not establish identification. [Part IV, Ch 3: Causal Models and Interventions] treats these choices.
- Response features: sign, magnitude, and timing can matter differently to a controller. A final-state norm does not separately diagnose all three. Report the response itself alongside a scalar error when these distinctions matter.
A successful response test covers the specified states and interventions. It does not guarantee useful optimization gradients, safe control, or accuracy on all alternative actions. Those need further tests.
Closed-loop planning
Closed-loop planning evaluates the model as part of a decision system. At each step the agent observes the real state, queries the model to choose an action, executes that action in the true environment, and observes again. For a deterministic environment the resulting state is ; with stochastic dynamics it is sampled from .
Re-observation prevents a predicted state from directly replacing the real next state. It does not sever the effect of model error on the trajectory: predictions influence the action, and the action influences future real states. Feedback can correct some errors, but an unstable or poorly chosen controller can still fail.
Small prediction perturbations need not cause small action changes. Consider actions and , true next state , and target zero. An exact predictor gives equal squared costs, so an optimizer with a tie rule can choose . Perturb only the predicted successor from to , with . The prediction changes by only , but the unique preferred action is now , changing the real successor by two. Smooth action sensitivity or stable correction requires additional conditions on the optimizer, objective, constraints, and closed-loop dynamics.
Let be expected cumulative cost in the true environment, with a common objective, initial-state distribution, noise law, horizon, and policy class . If , define
Assume the costs and infimum are finite. If an optimal policy attains that infimum, the gap is . For rewards being maximized, use instead. The sign depends on whether the objective is a cost or a reward.
Model approximation, finite-data estimation, planner suboptimality, and distribution mismatch are diagnostic categories, not universally independent additive quantities. An exact decomposition would need defined intermediate models and policies. We use the categories to investigate the observed gap.
A planner can exploit model errors by selecting actions whose predicted cost is favorable but whose real cost is not. This is one possible failure mechanism of optimizing through an inaccurate model. Greater search capacity does not invariably worsen performance: it can also find better actions. The direction depends on the errors, objective, constraints, and search procedure.
Three practical points follow.
- Feedback and rollout test different uses: replanning from observations can compensate for recursive model drift. Good tracking in that setting does not certify a long open-loop simulation.
- Report the whole decision system: the result belongs to the model, planner, objective, constraints, observation convention, and compute budget together. Changing any of them can change the ranking.
- Separate objective validity from decision optimization: a model can predict transitions accurately while the controller optimizes the wrong objective for deployment. State the intended objective and the scored objective before attributing a decision failure to dynamics.
The worked example uses observed states and a one-step grid search. It does not test belief-state inference, long-horizon planning, or safety.
Transfer
Transfer repeats an evaluation under a specified change. Several changes should be distinguished.
- A dynamics shift replaces by a different transition law in the queried region. A predictor accurate for can become inaccurate there.
- An initial-state shift changes occupancy and can expose poorly covered regions without changing the transition law.
- A changed objective or action constraint changes which decisions are useful even when the dynamics predictor remains correct.
- An observation shift can change the state or belief supplied to the controller while the physical transition law stays fixed.
Fitting data from alone establishes neither exact correctness under nor inevitable failure under . State what changed, what remains fixed, and whether retraining, adaptation, or extra observations are allowed. The experiment below changes only one true dynamics coefficient and keeps the candidate models, controller, target, initial state, and action grid fixed.
A scientific prior can help if it encodes structure that persists under the chosen shift; it can hurt if that structure changes. Better in-distribution fit does not universally imply worse transfer, and structural bias does not universally imply better transfer. [Part XII, Ch 2: Robustness, Calibration, and Safe Control] is planned to examine these questions more systematically.
Observation quality remains a separate part of the system, as discussed in [Part III, Ch 1: Observations, Sensors, and Multimodal Data]. Controlled changes to the perception pipeline can help distinguish observation failures from transition failures. The next chapters are planned to cover state and physics diagnostics.
A worked example: the model that wins one-step loses closed-loop
The example below is deliberately misspecified and small, so that everything is on the table. It is a compact illustration, not evidence that any particular architecture is inferior.
We use a fully observable linear system with a deterministic state-coupling structure plus additive Gaussian process noise:
with , , and i.i.d. across .
Where:
-
: the scalar state at time
-
: the scalar control action at time
-
: the true state-coupling coefficient
-
: the true action gain
-
: i.i.d. Gaussian process noise with standard deviation
-
State: scalar (fully observed).
-
Action: scalar , with a hard bound for the planning tests.
-
Planning objective: squared predicted next-state distance to , with no action penalty.
-
Control score: mean absolute target error over true states 31–40. This diagnostic is not the cumulative squared-cost gap defined above.
-
Episode horizon: 40 steps for rollouts, 40 steps for closed loop.
-
Data policy: for one-step evaluation, actions are drawn i.i.d. from . For rollout evaluation, we use the fixed sequence we specify in each rung.
-
Evaluation data: 800 independent synthetic transitions. The candidates are hand specified before sampling; no training set or fitted-model claim is involved.
-
Noise regime: Gaussian process noise is present in the one-step sample only. Rollout, intervention, and control tests use the noiseless transition, so their errors measure misspecification rather than different noise draws.
-
Model state: the Delayed candidate also carries the previous action. Its recurrent state is , initialized with in rollout and control tests.
We compare three candidate models of the form
where:
- : the model's next-state prediction;
- : the input state, true for teacher forcing and replanning, predicted for an open-loop rollout;
- : the previous applied action, , with for the trajectory experiments;
- : the model's state-coupling coefficient
- : the model's instantaneous action gain
- : the model's action-lag coefficient, capturing the influence of the previous action
The following candidates are chosen to expose different failure modes.
- Aggressive (): captures the state coupling closely but overshoots it into instability.
- Conservative (): underestimates the state coupling.
- Delayed (, ): gets the state coupling right but adds a spurious action lag.
These are hand-constructed misspecifications, not fitted solutions. Our goal is not to argue that any particular model class fails; it is to show that each rung rewards different inductive biases and that a lower rung does not by itself guarantee a higher one without additional linking assumptions. Notice that each model is a plausible way to be wrong about the same system. The Aggressive model is approximately correct as a one-step predictor because its coefficient error is small; the Conservative model trades a larger bias for more conservative behavior; the Delayed model gets the coupling right but invents a spurious dependence on the previous action. These hand-selected deviations let us isolate what happens when different coefficient errors are pushed through the evaluation tests.
import numpy as np
# Ground-truth one-dimensional controlled dynamics:
# s_{t+1} = A_TRUE * s_t + B_TRUE * u_t + noise
A_TRUE = 0.95
B_TRUE = 0.50
NOISE_STD = 0.02
# Candidate models of the form s_hat_{t+1} = a * s_t + b * u_t + c * u_{t-1}.
MISS = {
"Aggressive (a=1.02)": {"a": 1.02, "b": 0.50, "c": 0.00},
"Conservative (a=0.80)": {"a": 0.80, "b": 0.50, "c": 0.00},
"Delayed (a=0.95, c=0.20)": {"a": 0.95, "b": 0.50, "c": 0.20},
}
def model_step(model, s, u, u_prev):
return model["a"] * s + model["b"] * u + model["c"] * u_prevNow the rung-one evaluation, on 800 i.i.d. transitions.
rng = np.random.default_rng(20240501)
N_EVAL = 800
s_t = rng.normal(0.0, 1.0, size=N_EVAL)
u_prev_t = rng.normal(0.0, 1.0, size=N_EVAL)
u_t = rng.normal(0.0, 1.0, size=N_EVAL)
s_next_t = A_TRUE * s_t + B_TRUE * u_t + rng.normal(0.0, NOISE_STD, size=N_EVAL)
one_step_mse = {}
one_step_losses = {}
expected_one_step_mse = {}
gaussian_mse_se = {}
for name, model in MISS.items():
pred = model["a"] * s_t + model["b"] * u_t + model["c"] * u_prev_t
losses = (pred - s_next_t) ** 2
one_step_losses[name] = losses
one_step_mse[name] = float(np.mean(losses))
expected_risk = (
(model["a"] - A_TRUE) ** 2
+ (model["b"] - B_TRUE) ** 2
+ model["c"] ** 2
+ NOISE_STD**2
)
expected_one_step_mse[name] = expected_risk
gaussian_mse_se[name] = expected_risk * np.sqrt(2 / N_EVAL)
paired_mse_se = {}
names = list(MISS)
for left in range(len(names)):
for right in range(left + 1, len(names)):
pair = (names[left], names[right])
differences = one_step_losses[pair[0]] - one_step_losses[pair[1]]
paired_mse_se[pair] = float(
np.std(differences, ddof=1) / np.sqrt(N_EVAL)
) Aggressive (a=1.02) sample MSE = 0.005843; population MSE = 0.005300; Gaussian marginal SE = 0.000265
Conservative (a=0.80) sample MSE = 0.024686; population MSE = 0.022900; Gaussian marginal SE = 0.001145
Delayed (a=0.95, c=0.20) sample MSE = 0.040841; population MSE = 0.040400; Gaussian marginal SE = 0.002020
Paired Aggressive (a=1.02) minus Conservative (a=0.80): sample gap -0.018842; estimated SE 0.000938
Paired Aggressive (a=1.02) minus Delayed (a=0.95, c=0.20): sample gap -0.034998; estimated SE 0.002047
Paired Conservative (a=0.80) minus Delayed (a=0.95, c=0.20): sample gap -0.016156; estimated SE 0.002321On these independent standard-normal inputs, each model's residual is a zero-mean Gaussian linear combination of state, action, previous action, and independent process noise. Its population MSE is therefore
The resulting risks are for Aggressive, for Conservative, and for Delayed. Aggressive's state coefficient error is smaller than Conservative's; Delayed matches that coefficient exactly. Aggressive has lower total risk than Delayed because Delayed's spurious previous-action term adds variance . The sampled MSEs differ from these population risks, but retain the same order for the stated seed. No candidate was fitted to the sample.
For independent zero-mean Gaussian residuals of variance , the variance of their squared value is . Thus the sample MSE has standard error : at , five percent of the population MSE. The code prints that marginal calculation and separately estimates the standard error of each paired loss difference. The latter includes the covariance between model losses scored on the same transitions.
Now for rung two. We fix and apply constant actions over 40 steps, then measure trajectory RMSE.
HORIZON = 40
S_INIT = 0.3
actions_const = np.full(HORIZON, 0.20)
true_traj_const = np.zeros(HORIZON + 1)
true_traj_const[0] = S_INIT
for t in range(HORIZON):
true_traj_const[t + 1] = (
A_TRUE * true_traj_const[t] + B_TRUE * actions_const[t]
)
rollout_rmse_const = {}
for name, model in MISS.items():
traj = np.zeros(HORIZON + 1)
traj[0] = S_INIT
u_prev = 0.0
for t in range(HORIZON):
traj[t + 1] = model_step(model, traj[t], actions_const[t], u_prev)
u_prev = actions_const[t]
rollout_rmse_const[name] = float(
np.sqrt(np.mean((traj[1:] - true_traj_const[1:]) ** 2))
) Delayed (a=0.95, c=0.20) RMSE = 0.4936
Conservative (a=0.80) RMSE = 0.8969
Aggressive (a=1.02) RMSE = 2.3856Under a constant positive push , the true system's noiseless steady state is . The Aggressive model has , so its rollout diverges. The Conservative model has steady state , well below the truth. The Delayed model acquires an effective action gain , giving steady state , incidentally closer to the true value than the Conservative model's. Delayed wins rung two, despite losing rung one. The gain is not the true gain . Delayed wins this finite-horizon comparison because its state decay coefficient matches the truth and its accumulated gain error is smaller than the competing models' trajectory errors for this initial state and input. This is not a claim that the lagged model has correct constant-input dynamics.
For a controlled-intervention trajectory diagnostic, change the constant input to over the same 40 steps. This is another open-loop trajectory test under a deliberately changed action sequence. It is not the intervention-response difference defined earlier, and it does not establish causal identification. The simulator supplies the true dynamics by construction.
def paired_rollouts(model, actions, initial=0.3):
truth = np.zeros(len(actions) + 1)
predicted = np.zeros_like(truth)
truth[0] = predicted[0] = initial
previous = 0.0
for step, action in enumerate(actions):
truth[step + 1] = A_TRUE * truth[step] + B_TRUE * action
predicted[step + 1] = model_step(
model, predicted[step], action, previous
)
previous = action
return truth, predicted
actions_alt = 0.4 * (-1.0) ** np.arange(HORIZON)
rollout_rmse_alt = {}
long_alt_rmse = {}
alternating_amplitude = {"Truth": 0.4 * B_TRUE / (1 + A_TRUE)}
for name, model in MISS.items():
truth, predicted = paired_rollouts(model, actions_alt)
rollout_rmse_alt[name] = float(
np.sqrt(np.mean((predicted[1:] - truth[1:]) ** 2))
)
truth_long, predicted_long = paired_rollouts(
model, 0.4 * (-1.0) ** np.arange(500)
)
long_alt_rmse[name] = float(
np.sqrt(np.mean((predicted_long[1:] - truth_long[1:]) ** 2))
)
if abs(model["a"]) < 1:
alternating_amplitude[name] = (
0.4 * abs(model["b"] - model["c"]) / (1 + model["a"])
)Aggressive (a=1.02): alternating RMSE H=40 0.504376; H=500 1807.149336 Conservative (a=0.80): alternating RMSE H=40 0.135162; H=500 0.039716 Delayed (a=0.95, c=0.20): alternating RMSE H=40 0.045496; H=500 0.041403 Truth: steady alternating amplitude 0.1025641 Conservative (a=0.80): steady alternating amplitude 0.1111111 Delayed (a=0.95, c=0.20): steady alternating amplitude 0.0615385
Delayed wins the 40-step alternating-trajectory test, but not because its steady alternating response is most accurate. After the initial step, , so its effective alternating gain is . For , substituting into the forced recurrence gives the signed coefficient . The amplitude magnitude is . All displayed stable candidates have , so their coefficients and magnitudes agree. The true amplitude is , Conservative's is , and Delayed's is . Delayed instead benefits at this short horizon from matching the true transient decay coefficient. At 500 steps Conservative has lower trajectory RMSE than Delayed, as the printed results show. Aggressive is unstable, so its formal forced periodic solution is not an attracting steady response.
We also compute the actual response-difference test. Use the constant input as baseline and the alternating input as intervention, starting both trajectories at with previous action zero. Compare their final-state differences, not merely the intervened trajectory:
intervention_error = {}
intervention_responses = {}
for name, model in MISS.items():
truth_base, predicted_base = paired_rollouts(model, actions_const)
truth_changed, predicted_changed = paired_rollouts(model, actions_alt)
delta_true = truth_changed[-1] - truth_base[-1]
delta_model = predicted_changed[-1] - predicted_base[-1]
intervention_responses[name] = (float(delta_true), float(delta_model))
intervention_error[name] = float(abs(delta_model - delta_true))Delayed (a=0.95, c=0.20): true response -1.832359; model response -2.477563; final response error 0.645204 Conservative (a=0.80): true response -1.832359; model response -0.611030; final response error 1.221329 Aggressive (a=1.02): true response -1.832359; model response -5.920590; final response error 4.088231
This response metric and the alternating-trajectory RMSE answer different questions, even if they select the same winner here. Equal additive trajectory biases cancel in the difference, but general dynamical errors need not. The metric tests these two known simulator interventions only; it does not certify unseen interventions or an observationally learned causal model.
Rung four is closed-loop planning. We use a one-step MPC that searches 201 equally spaced actions in and minimizes squared predicted next-state distance to . It then executes the chosen action in the true deterministic environment and reobserves the state. This costs 201 candidate one-step predictions per control step, with no action penalty and no lookahead beyond one step. The grid enforces the action constraint by construction; no state-safety constraint is imposed.
U_GRID = np.linspace(-0.5, 0.5, 201)
def mpc_choose(model, s, u_prev, target):
preds = model["a"] * s + model["b"] * U_GRID + model["c"] * u_prev
idx = int(np.argmin((preds - target) ** 2))
return float(U_GRID[idx])
def closed_loop(model, s_start, target, steps, a_env, b_env):
s = s_start
u_prev = 0.0
traj = [s]
for _ in range(steps):
u = mpc_choose(model, s, u_prev, target)
s = a_env * s + b_env * u
traj.append(s)
u_prev = u
return np.array(traj)
STEPS = 40
TARGET = 1.0
closed_loop_error = {}
for name, model in MISS.items():
traj = closed_loop(
model,
s_start=0.0,
target=TARGET,
steps=STEPS,
a_env=A_TRUE,
b_env=B_TRUE,
)
tail = traj[-10:]
closed_loop_error[name] = float(np.mean(np.abs(tail - TARGET)))
transfer_error = {}
A_SHIFT = 0.80
for name, model in MISS.items():
traj = closed_loop(
model,
s_start=0.0,
target=TARGET,
steps=STEPS,
a_env=A_SHIFT,
b_env=B_TRUE,
)
tail = traj[-10:]
transfer_error[name] = float(np.mean(np.abs(tail - TARGET)))Closed-loop tracking error (same environment):
Delayed (a=0.95, c=0.20) |s - s*| = 0.0200
Aggressive (a=1.02) |s - s*| = 0.0654
Conservative (a=0.80) |s - s*| = 0.1765
Transfer to a changed environment (a shifts from 0.95 to 0.80):
Conservative (a=0.80) |s - s*| = 0.0000
Aggressive (a=1.02) |s - s*| = 0.1803
Delayed (a=0.95, c=0.20) |s - s*| = 0.1875An idealized continuous-action controller helps explain the tracking bias. Assume deterministic dynamics, an inactive action bound, no action penalty, and one-step squared next-state tracking toward target . With , its unconstrained choice is
At a fixed point both state and action are constant. The true environment therefore satisfies , while the model's target equation becomes . Eliminating yields
where is the continuous controller's equilibrium, and are true environment coefficients, and are model coefficients. The denominator must be nonzero. This algebra finds an equilibrium, not a guarantee of stability or feasibility; those must be checked separately.
def continuous_equilibrium(model, a_env, b_env, target):
denominator = (model["b"] + model["c"]) * (1 - a_env) + b_env * model["a"]
return b_env * target / denominator
continuous_fixed_points = {
"same": {
name: continuous_equilibrium(model, A_TRUE, B_TRUE, TARGET)
for name, model in MISS.items()
},
"transfer": {
name: continuous_equilibrium(model, A_SHIFT, B_TRUE, TARGET)
for name, model in MISS.items()
},
}same continuous-action equilibria: Aggressive (a=1.02): 0.9345794393 Conservative (a=0.80): 1.1764705882 Delayed (a=0.95, c=0.20): 0.9803921569 transfer continuous-action equilibria: Aggressive (a=1.02): 0.8196721311 Conservative (a=0.80): 1.0000000000 Delayed (a=0.95, c=0.20): 0.8130081301
The continuous same-environment equilibria are approximately (Aggressive), (Conservative), and (Delayed). In transfer they are , , and . The lagged-action term remains in the transfer equation; dropping it would give the wrong Delayed equilibrium.
The implemented controller searches 201 discrete actions instead. It need not converge to those continuous equilibria: grid quantization can produce small periodic trajectories. We report the actual mean absolute tracking error over states 31 through 40, not an analytically inferred bias. Delayed has the lowest same-environment tail error, while Conservative has the lowest transfer tail error. Conservative is worst in the same-environment closed-loop comparison, not least accurate on every metric.
delayed_name = "Delayed (a=0.95, c=0.20)"
delayed_same_long = closed_loop(
MISS[delayed_name], 0.0, TARGET, 1000, A_TRUE, B_TRUE
)
delayed_transfer_long = closed_loop(
MISS[delayed_name], 0.0, TARGET, 1000, A_SHIFT, B_TRUE
)
delayed_same_40 = closed_loop(
MISS[delayed_name], 0.0, TARGET, STEPS, A_TRUE, B_TRUE
)
delayed_transfer_tail_mae = float(
np.mean(np.abs(delayed_transfer_long[-10:] - TARGET))
)Delayed actual state after 40 same-environment steps: 0.9794848891 Delayed last five same-environment states: [0.98050003 0.97897503 0.98002627 0.98102496 0.97947371] Delayed last two transfer states: [0.81527778 0.80972222] Delayed long-run transfer tail MAE: 0.1875000
From this initialization, Delayed approaches a five-step cycle in the same environment and a two-step cycle in transfer. The transfer cycle has states approximately and , with corresponding executed actions and ; its tail MAE is . These grid trajectories explain why a finite-run statistic differs from a continuous-action fixed point. Neither statistic is an architectural ranking.
Closing the loop also introduces a decision baseline. An oracle one-step controller using the true coefficients separates dynamics misspecification from action-grid and startup effects in this toy; a zero-action baseline shows the value of control. This is a diagnostic baseline, not a claim that oracle dynamics are available in deployments.
oracle_model = {"a": A_TRUE, "b": B_TRUE, "c": 0.0}
oracle_traj = closed_loop(oracle_model, 0.0, TARGET, STEPS, A_TRUE, B_TRUE)
oracle_tail_mae = float(np.mean(np.abs(oracle_traj[-10:] - TARGET)))
zero_action_tail_mae = float(np.mean(np.abs(np.zeros(10) - TARGET)))True-dynamics one-step controller tail MAE: 0.0002316 Zero-action baseline tail MAE: 1.0000000
The actual grid-controller trajectories show the residual tracking errors.
closed_loop_trajs = {}
for name, model in MISS.items():
closed_loop_trajs[name] = closed_loop(
model,
s_start=0.0,
target=TARGET,
steps=STEPS,
a_env=A_TRUE,
b_env=B_TRUE,
)
Now compare rankings without dividing by a near-zero winner error. The transfer winner's error is almost zero, so dividing other models' errors by it would create enormous ratios and hide differences on the other tests. We show within-test ranks, with 1 best. Ranks discard the size of the gaps; the printed raw metrics retain that information. Do not sum the ranks into a universal score.

The pattern is exactly what the ladder is designed to expose: different rungs can favour different winners.
- One-step MSE favors Aggressive on the stated independent Gaussian inputs.
- Constant-input rollout favors Delayed at the specified 40-step horizon, despite its incorrect gain.
- The response-difference metric favors Delayed for this baseline/intervention pair. The additional alternating-trajectory test is not the same metric.
- Same-environment grid control also favors Delayed; the example does not show a winner reversal between that rollout and that controller.
- Transfer favors Conservative because its coefficients match the changed environment more closely. This is no general guarantee of robust transfer.
These rankings are conditional on the test distribution, action sequence, horizon, initial state, and scoring rule. A scalar chosen for a declared deployment objective can be useful. Report that choice explicitly, and retain the other diagnostics when they describe additional intended uses.
Limitations and impact
The ladder is a diagnostic framework. It does not order models independently of a task, nor does it make the five protocols a strict staircase.
A model can perform well on a short control horizon and poorly on a long open-loop simulation. That is not a contradiction: feedback and planner inputs differ. Likewise, predicting a full state vector and controlling one task-relevant component can require different fidelity. These are reasons to record context and component-level diagnostics rather than infer an untested capability.
Statistical interpretation depends on the collection protocol. The 800-transition example uses independent Gaussian inputs and residuals, so the marginal MSE standard-error formula above applies. Real sequences may have correlated errors, non-Gaussian residuals, or variation across initial conditions and environments. A one-run comparison can be informative in a deterministic controlled experiment; claims about performance across episodes require appropriate replication and uncertainty estimates.
Use paired loss differences when models share test transitions, and choose the splitting unit to match the question. A random transition-row split can preserve dependencies across train and test or leak trajectory-specific information. It does not always do so: independently generated rows are a different protocol. If deployment concerns new initial conditions or environments, split and replicate at those levels. [Part XI, Ch 5: Datasets, Benchmarks, and Experimental Design] is planned to cover the protocol and power analysis.
The toy itself is deliberately limited. It has a scalar fully observed state, hand-specified candidates, Gaussian noise only in the one-step sample, and deterministic trajectory tests. The control experiment uses a bounded 201-point action grid, no state-safety constraints, and a one-step objective. It says nothing about the safety of an unconstrained deployment, nor about the quality of a multi-step planner or belief-state estimator.
Within those limits, the arithmetic shows why the tested protocols cannot be substituted for one another. The one-step winner differs from the constant-rollout and same-environment control winner. The alternating trajectory ranking changes between 40 and 500 steps. Changing the true state coefficient changes the control winner again. These are examples under declared conditions, not general architectural rankings.
For deployment, choose the diagnostics that match the proposed use and state the untested uses explicitly. Open-loop simulation needs appropriate horizons and path statistics. Policies learned through model rollouts need intervention and closed-loop tests under their action-selection procedure. Transfer needs a declared shift and an adaptation budget. A matched multi-rung comparison can reveal differences, but an irrelevant or unidentifiable rung should not be reported as certified.
This framework gives concrete uses to the distinctions developed in [Part II, Ch 6: Model-Based Reinforcement Learning], [Part III, Ch 2: Sufficient State and State Abstraction], and [Part VI, Ch 1: Predictive and Generative Objectives]. It separates representation quality, transition fidelity, and the way a planner uses predictions. Model exploitation is one failure mechanism to investigate rather than assume. It is a planned topic of [Part XII, Ch 1: Failure Modes and Model Exploitation].
The remaining chapters in this part are planned to specialize these questions: perceptual and generative evaluation for observation quality and scoring conventions; state, physics, and causal diagnostics for structured properties; planning and policy evaluation for decision outcomes; datasets and experimental design for reproducible comparisons; and interpretability and debugging for mechanisms behind failures.
For images, video, or point clouds, "closeness" itself requires care. Different valid images can have large pixel distances; identical point sets can have different orderings. Such metrics must state their representation and invariances. Those questions are planned topics of the next chapter, Perceptual and Generative Evaluation.
Summary
- The five protocols are one-step prediction, open-loop rollout, intervention response, closed-loop planning, and transfer. They are complementary questions, not strictly ordered guarantees.
- One-step squared risk under a common finite-moment evaluation law splits into conditional-mean error and average conditional noise. A held-out estimate needs a declared sampling protocol and uncertainty calculation.
- The deterministic rollout bound requires uniform reachable-input accuracy and Lipschitz conditions, common actions, a common initial state, and no corrections. Its geometric expression is an envelope, not a guaranteed observed growth rate.
- Deterministic branch comparisons and stochastic counterfactuals need different specifications. Nonparametric adjustment-based observational identification requires causal assumptions and action coverage. Extrapolating unsupported action effects requires additional structural assumptions; a simulator score alone does not establish identification.
- Closed-loop model errors affect real states through chosen actions. Re-observation avoids direct prediction substitution but does not guarantee smooth action response or stability.
- Transfer must specify what changed. Dynamics, observations, occupancy, objectives, and constraints can affect usefulness in different ways.
- In the stated toy, winners across the five metrics are Aggressive, Delayed, Delayed, Delayed, and Conservative. Those rankings are conditional on the horizon, input sequence, initial state, and controller.
- A scalar deployment score can be meaningful under a declared objective. Use the diagnostic vector to expose assumptions and untested behavior that the scalar would otherwise leave implicit.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the world-model evaluation ladder.
The World-Model Evaluation Ladder
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!