Part of World Models Handbook
Test world-model state, physics, causality, and memory with probes, occlusion, energy checks, and paired interventions; learn what each diagnostic can prove.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
State, Physics, Causality, and Memory Evaluation
Picture a small ball rolling across a table. For two seconds it is plainly visible. Then it passes behind a box, and for eight frames you see nothing at all. Then it emerges on the other side. Now ask a world model what happened in between.
This is a deceptively simple question. It sounds like a single question, but it is really a bundle of at least four distinct questions, and world models can answer some of them correctly while failing the others. That separation is the entire subject of this chapter.
There are at least four different things a model could do here, and they are easy to confuse. Among them, only the last is a causal question: the first three are observational or representational.
- It could generate eight frames that look like a ball continuing across the box. The pixels are plausible; the ball is in roughly the right place.
- It could maintain a hidden state that keeps the ball's position and velocity coherent, so that the ball emerges at exactly the right place at the right time.
- It could keep the ball's identity intact, so that the thing that emerges is the same object rather than a new one that happens to be nearby.
- It could do all of the above while also knowing that if you had pushed the ball differently just before it disappeared, it would have emerged somewhere else.
These are different requirements, not a strict hierarchy. Accurate positions do not guarantee persistent identity, and identity tracking does not guarantee accurate motion. An intervention asks a further question about changes to the world that observational prediction alone does not identify without assumptions. Each requirement needs its own evaluation protocol.
A model can pass the visual test while failing hidden-state, identity, or intervention tests. Even strong results across these offline diagnostics do not establish closed-loop usefulness: a planner may fail when it acts on the predictions. Chapter 64 evaluates decisions, Chapter 65 develops experimental design, and Chapter 67 examines failure modes and model exploitation.
This chapter is about how to tell these cases apart. It is in Part XI: Evaluation and Understanding. The previous chapter, "Perceptual and Generative Evaluation" (Chapter 62), handled the question of whether generated observations look right and whether their distribution matches the data. Here we develop state, constraint, and intervention diagnostics within the world-model evaluation ladder established in Chapter 61, without treating those diagnostics as nested requirements.
The chapter is organized around four evaluation families:
- State probing and representation quality. When you freeze a representation and fit a small readout, what can it recover, and what does that recovery prove?
- Object permanence and temporal consistency. Does the model maintain a coherent hidden trajectory through occlusion, and for how long?
- Physical rule and constraint tests. Does the model respect the conservation laws and boundary conditions of the regime it claims to model?
- Interventions, counterfactuals, and causal fidelity. Under a change to one mechanism, does the model predict the right change, or merely fit the observational distribution?
The four families examine separate properties. Probing measures accessibility to a readout; occlusion tests measure behavior over hidden windows; constraint tests check a named rule; interventions test a specified change under an identification design. None universally subsumes the others. Running them separately makes their different failure modes visible.
State what information the model received, what information the baseline received, and what the metric denominator counts. Compare methods on a named population, and distinguish differences in modelling from differences in information access. A privileged-state baseline can be useful, but its advantage over an observation-only method cannot be attributed entirely to better modelling.
The first three families use a point mass moving in the unit square with reflecting walls, an occluder band, and noisy position observations. The causal section then switches to a separate one-step structural model so that its confounding mechanism is explicit. The point-mass state lives in as
where:
- : position components in metres.
- : velocity components in metres per second.
We observe only the position components through , and the dynamics are the deterministic map
defined in the simulator below, where is the acceleration action.
The reflecting-wall simulator returns one next state given state and action; its random initial conditions, action sampling, and observation noise are separate draws. A structural model with a specified disturbance likewise returns one outcome once all its inputs are fixed. The later causal example introduces a shared disturbance in such a one-step model. Reusing that disturbance across counterfactual branches is an explicit coupling assumption, not a property inferred from passive observations.
The point-mass simulator fits in a few dozen lines of NumPy and lets us inspect probe, occlusion, and constraint diagnostics against known truth. Neither it nor the later causal model is a benchmark for trained foundation models. Treat them as controlled fixtures: their simplicity makes each diagnostic's target and assumptions visible, without establishing its performance on a larger system.
Physical simulator state in metres and metres per second. Observed measurement in metres. Action in metres per second squared. Representation ; the example encoder is handcrafted. The main simulator has deterministic transitions given state and action, with separate random initial conditions, sampled actions, and measurement noise. The later one-step causal model introduces an unobserved disturbance shared by its action and outcome equations. Rollout horizon . Time step s.
State Probing and Representation Quality
A probe is a small model fitted on top of a frozen representation. You take the encoder, stop its gradients, run it over a dataset, and then ask: can a simple readout predict quantity from the representation ? Saying that the representation "contains" can conflate recoverability by a chosen readout, unrestricted information content, and downstream use. We need to ask how the information is arranged and which procedure can read it out.
A frozen representation is a fixed map whose parameters do not change during evaluation. A probe is a separately fitted function trained to predict a target from . The probe's held-out score measures how accessible is to that probe family, not whether the encoder "knows" in any deeper sense.
The three questions are:
- Linear accessibility. How well does a fitted linear readout recover from on held-out data? Success demonstrates accessibility to that fitted readout on the tested distribution. Failure can reflect regularization, sample size, fitting, or distribution shift; it does not prove that no linear readout could succeed. A linear probe still learns a task-specific mapping.
- Information decodability. Can some function in a richer probe family recover ? A flexible probe can learn substantial task structure from its frozen inputs, so high performance depends jointly on the encoded information, the readout's capacity, and its fitting procedure. Capacity cannot supply target information absent from those inputs.
- Downstream functional use. Does the rest of the system use to make decisions? This is the only one of the three that connects to behaviour, and it cannot be established by a probe at all. A probe is a diagnostic instrument attached to the side of the model; it does not tell you what the model does with the information when it acts.
Comparing readouts can reveal which fitted family works better on the test distribution. Better nonlinear scores do not establish impossibility of linear recovery, and failures under both families do not prove absent information. Good decoding also does not establish causal necessity: the downstream system may ignore the recovered variable. Testing functional use requires additional experiments, such as controlled representation interventions and closed-loop comparisons.
Probe splits, controls, and preprocessing
Check these three procedural details before interpreting a probe score. Splits and preprocessing determine what information is held out; controls clarify how the selected probe behaves on an alternative task. Mistakes can bias an evaluation or limit the conclusions it supports.
The probe itself is a fitted map from the frozen representation to a target . It is a simple readout, not a full model. In the linear case, ; in the nonlinear case, the probe first maps through a fixed feature function and then applies a ridge readout. The data are the held-out trajectories, and all probe scores are computed against targets that were not used to fit either the encoder or the probe, and trajectories are held out as complete sequences rather than individual frames so that frames from one rollout never appear in two splits. The key design choice is that the encoder is frozen: the probe cannot reshape the representation to make its job easier, so any information the probe recovers must already be present in .
Split at the trajectory level, not the frame level. Consecutive frames of one trajectory are strongly correlated: neighbouring timesteps share nearly the same position. Random frame splits can put near-duplicates of test frames in the training set and inflate the probe's held-out score through memorisation. Splitting whole trajectories removes this within-trajectory leak. In our toy we hold out complete rollouts. Choose a split unit that respects the dependencies in the data; shared scenes or subjects may require a larger grouping than one trajectory.
Fit preprocessing on training data only. Standardisation statistics, whitening matrices, and PCA bases must be estimated on the training split and then applied unchanged to validation and test. Fitting them on the full dataset uses held-out information. Its effect on a score depends on the data and fitting procedure; neither inflation nor a flat train/test gap is guaranteed.
Use a control task with matched probe capacity. Hewitt and Liang (2019) developed type-level randomized control tasks for linguistic probes. Our shuffled regression targets are a simpler toy control, not a reproduction of their protocol. Report real-target and control results together, with the randomization and split specified. A control gives context about the chosen fitting procedure; it does not bound all structured-task learning or certify information content.
The toy representation
To keep this concrete, let us build a frozen representation by hand and be explicit about what it is. We will use a phase encoding of position only. In the formula and NumPy code, and denote the numerical values of observed coordinates in metres: the dimensionless phase is times the coordinate divided by a one-metre period.
where:
- : the observed horizontal position at time , in metres.
- : the observed vertical position at time , in metres.
- : the phase encoding of the horizontal coordinate as a point on the unit circle; together they determine its numerical value modulo one. Exact inversion over a full period is nonlinear; this does not preclude a useful approximate linear readout over a restricted distribution.
- : the phase encoding of the vertical coordinate as a point on the unit circle, with the same periodic-recovery property.
This is a handcrafted encoder, not a learned one. Exact recovery of a noiseless position within a specified one-period interval uses nonlinear phase inversion; positions differing by an integer share the same encoding, and noisy observations near the interval boundary complicate recovery. The encoder does not directly take velocity, but that does not make velocity statistically independent of its output. Our sampled trajectories can correlate position and velocity. The experiment tests these readouts on a named distribution, not an absolute present-versus-absent information theorem.
The toy world and the dataset
Our world is a point mass in the unit square with perfectly reflecting walls. Let us fix the conventions before writing any metric. Fixing conventions up front is not a formality; it is what makes the numbers in the tables below interpretable and comparable across sections.
- State: , metres and metres per second.
- Action: , metres per second squared, applied as an acceleration.
- Time step: s, so one trajectory of 30 steps is 3 seconds.
- Observation: with independent coordinate noise .
- Occluder: the band for , with the occluder half-width.
- Data-generating policy: m/s², independent across components and timesteps.
- Split unit: one complete trajectory. We use 60 for training, 20 for validation, 20 for test.
import numpy as np
DT = 0.1 # s
V_MAX = 0.6 # m/s
A_MAX = 0.3 # m/s^2
N_STEPS = 30 # 30 steps = 3.0 s
SIGMA_OBS = 0.004 # m
OCC_Y = (0.05, 0.95)
OCC_HALF = 0.10 # half-width of the occluder band in x
def simulate(
rng,
n_traj,
n_steps=N_STEPS,
driven=True,
damping=1.0,
restitution=1.0,
v_scale=1.0,
):
"""State s_t = (x, y, vx, vy) in metres and metres per second.
Returns positions (n, T+1, 2), velocities (n, T+1, 2),
actions (n, T, 2), and noisy observations (n, T+1, 2).
"""
pos = np.empty((n_traj, n_steps + 1, 2))
vel = np.empty((n_traj, n_steps + 1, 2))
act = np.zeros((n_traj, n_steps, 2))
pos[:, 0, 0] = rng.uniform(0.10, 0.30, n_traj)
pos[:, 0, 1] = rng.uniform(0.30, 0.70, n_traj)
vel[:, 0, 0] = v_scale * rng.uniform(0.15, 0.35, n_traj)
vel[:, 0, 1] = v_scale * rng.uniform(-0.15, 0.15, n_traj)
for t in range(n_steps):
if driven:
act[:, t] = rng.uniform(-A_MAX, A_MAX, size=(n_traj, 2))
v = damping * vel[:, t] + act[:, t] * DT
v = np.clip(v, -V_MAX, V_MAX)
p = pos[:, t] + v * DT
for axis in (0, 1):
lo = p[:, axis] < 0.0
hi = p[:, axis] > 1.0
p[lo, axis] = -p[lo, axis]
p[hi, axis] = 2.0 - p[hi, axis]
hit = lo | hi
v[hit, axis] = -restitution * v[hit, axis]
p = np.clip(p, 0.0, 1.0)
pos[:, t + 1] = p
vel[:, t + 1] = v
obs = pos + rng.normal(0.0, SIGMA_OBS, size=pos.shape)
return pos, vel, act, obs
def visibility(pos, occ_half=OCC_HALF):
"""1.0 where the ball is visible, 0.0 where it is behind the occluder."""
hidden = (
(np.abs(pos[..., 0] - 0.5) <= occ_half)
& (pos[..., 1] >= OCC_Y[0])
& (pos[..., 1] <= OCC_Y[1])
)
return (~hidden).astype(float)
def encode(pos):
"""Frozen handcrafted representation: phase encoding of position only."""
return np.stack(
[
np.cos(2 * np.pi * pos[..., 0]),
np.sin(2 * np.pi * pos[..., 0]),
np.cos(2 * np.pi * pos[..., 1]),
np.sin(2 * np.pi * pos[..., 1]),
],
axis=-1,
)Now we generate the training, validation, and test trajectories. The split is made on the trajectory axis, so no single trajectory contributes frames to two different splits. This is the discipline introduced above, applied here: the unit of the split is one whole trajectory, not one frame.
rng = np.random.default_rng(11)
N_TRAIN, N_VAL, N_TEST = 60, 20, 20
## NOTE: several diagnostics below are sensitive to the seed; keep the seed fixed.
## Probes, hidden-horizon errors and freeze-drift diagnostics share this main split.
## Width sweeps and physical-constraint tests use separately generated fixtures.
pos_all, vel_all, act_all, obs_all = simulate(
rng, N_TRAIN + N_VAL + N_TEST, driven=True
)
mask_all = visibility(pos_all)
## act_all already has shape (n, T, 2)
train_sl = slice(0, N_TRAIN)
val_sl = slice(N_TRAIN, N_TRAIN + N_VAL)
test_sl = slice(N_TRAIN + N_VAL, None)trajectories: train=60 val=20 test=20 frames per trajectory: 31 maximum absolute velocity component: 0.553 m/s (per-component clamp 0.6) fraction of hidden frames: 0.278 unit of the split: one complete trajectory
The hidden fraction tells us how much of the data sits behind the occluder. At 0.278 it is not small, and occlusion is not rare in this dataset. The reason occlusion matters out of proportion to its frequency is that a rare failure to maintain a coherent hidden state can dominate the downstream consequences, even if the model is otherwise accurate on 90% of the data.
Fitting the probes
We now build the probe dataset. Because the encoder is defined on observed positions, we keep only timesteps where the ball is visible, and we standardise the representation using training statistics only. Standardising with statistics computed on the full dataset is the leakage trap described above, so we are careful to compute the mean and standard deviation from the training split and apply the same transformation to validation and test.
def visible_arrays(sl, pos, vel, obs, mask):
Z = encode(obs[sl])
keep = mask[sl].astype(bool)
return Z[keep], pos[sl][keep], vel[sl][keep]
Z_tr, P_tr, V_tr = visible_arrays(train_sl, pos_all, vel_all, obs_all, mask_all)
Z_va, P_va, V_va = visible_arrays(val_sl, pos_all, vel_all, obs_all, mask_all)
Z_te, P_te, V_te = visible_arrays(test_sl, pos_all, vel_all, obs_all, mask_all)
mu, sd = Z_tr.mean(axis=0), Z_tr.std(axis=0) + 1e-9
Zt, Zv, Zs = (Z_tr - mu) / sd, (Z_va - mu) / sd, (Z_te - mu) / sd
## Targets: position (x, y) and velocity (vx, vy)
Y_tr = np.hstack([P_tr, V_tr])
Y_te = np.hstack([P_te, V_te])
print(f"train frames={Zt.shape[0]} val={Zv.shape[0]} test={Zs.shape[0]}")def ridge_fit(X, Y, alpha=1e-3):
return np.linalg.solve(X.T @ X + alpha * np.eye(X.shape[1]), X.T @ Y)
def r2_columns(Y, Yhat):
ss_res = np.sum((Y - Yhat) ** 2, axis=0)
ss_tot = np.sum((Y - Y.mean(axis=0)) ** 2, axis=0)
if np.any(ss_tot <= 0):
raise ValueError("R-squared requires nonconstant test targets")
return 1.0 - ss_res / ss_tot
def with_bias(X):
return np.hstack([X, np.ones((X.shape[0], 1))])
## Undefined constant-target R-squared is rejected rather than silently assigned.
try:
r2_columns(np.ones((3, 1)), np.ones((3, 1)))
except ValueError:
pass
else:
raise AssertionError("Constant-target R-squared must be rejected")We fit four probe procedures on the same frozen features. The linear and nonlinear readouts test different forms of accessibility, the shuffled-target control reveals how the selected fitting procedure behaves on that control task, and the intercept-only baseline provides a trivial reference. The control score is not a capacity floor.
- an intercept-only baseline, which predicts a fitted training-target constant and can score below zero on held-out when training and test means differ;
- a linear probe, a ridge readout directly on ;
- a nonlinear probe, a ridge readout on fixed random Fourier features . With Gaussian frequencies and uniform phases, the normalized inner product approximates a Gaussian kernel. We use the unnormalized features with the stated ridge penalty, so normalization would also change the effective regularization;
- a control probe, the same nonlinear feature map, but trained on randomly permuted targets.
## Intercept-only baseline
W_mean = ridge_fit(np.ones((Zt.shape[0], 1)), Y_tr)
## Linear probe
W_lin = ridge_fit(with_bias(Zt), Y_tr)
## Nonlinear probe: random Fourier features over the frozen representation
D_FEAT = 800
W_rff = rng.normal(0.0, 3.0, size=(D_FEAT, Zt.shape[1]))
b_rff = rng.uniform(0.0, 2 * np.pi, size=D_FEAT)
def phi(Z):
return np.cos(Z @ W_rff.T + b_rff)
W_nl = ridge_fit(with_bias(phi(Zt)), Y_tr)
## Control probe: identical pipeline, randomly permuted training targets
perm = rng.permutation(Y_tr.shape[0])
W_ctrl = ridge_fit(with_bias(phi(Zt)), Y_tr[perm])
r2_mean = r2_columns(Y_te, np.ones((Zs.shape[0], 1)) @ W_mean)
r2_lin = r2_columns(Y_te, with_bias(Zs) @ W_lin)
r2_nl = r2_columns(Y_te, with_bias(phi(Zs)) @ W_nl)
r2_ctrl = r2_columns(Y_te, with_bias(phi(Zs)) @ W_ctrl)
probe_rows = [
("x position", r2_mean[0], r2_lin[0], r2_nl[0], r2_ctrl[0]),
("y position", r2_mean[1], r2_lin[1], r2_nl[1], r2_ctrl[1]),
("vx velocity", r2_mean[2], r2_lin[2], r2_nl[2], r2_ctrl[2]),
("vy velocity", r2_mean[3], r2_lin[3], r2_nl[3], r2_ctrl[3]),
]target mean-only linear nonlinear control x position -0.034 0.826 0.963 -137.543 y position -0.284 0.499 -2.135 -32.485 vx velocity -0.001 0.012 -101.539 -109.621 vy velocity -0.053 0.285 -170.657 -210.255 All scores are held-out R^2 on whole held-out trajectories. The control column tests this probe family on shuffled training targets.
Read the table column by column. Each column has a specific interpretation, and the value of the table comes from comparing them.
- The mean-only column uses an intercept-only ridge fit, a training-target constant slightly shrunk toward zero by the penalty. It can score below zero because test compares against the test-target mean, which need not equal that fitted constant.
- The linear column measures accessibility to this linear readout. In this finite sample, and position achieve about 0.83 and 0.50, despite the phase encoding. Linear access need not vanish merely because the encoding is nonlinear.
- The nonlinear column uses a high-capacity random-feature readout. It improves position but fails badly on and velocity. Its very negative scores indicate poor held-out prediction, not a lack of information established by the test.
- The control column fits shuffled training targets. Its strongly negative scores show that this chosen probe can generalise poorly on the control task; they are not a universal capacity floor or a bound on what another probe could recover.
The four targets do not support a simple present-versus-absent classification: the results are mixed across probe families and state components. The handcrafted encoder explicitly takes only the current noisy position, yet correlations in the sampled trajectories can let a readout predict some velocity. The nonlinear readout's failure also shows why adding capacity does not guarantee a better diagnostic. Report these outcomes rather than replacing them with the intended story. The experiment illustrates probe sensitivity to fitting and data distribution; it does not prove that velocity information is absent or that position is inaccessible linearly.

Three things a probe can never tell you
Interpret the probe scores with three limits in mind: readout capacity, failed recovery, and downstream use.
A flexible probe may be doing substantial task learning from the information available in its frozen inputs, not creating target information missing from them. A nonlinear readout's gain can depend on both the encoded features and the readout's capacity and fitting procedure; the score does not apportion these contributions. Our control records this procedure's held-out performance after target shuffling. It bounds neither all learning on random targets nor learning on structured targets. Report the probe architecture, feature dimension, and regularisation strength so that the procedure can be reproduced.
Low linear performance does not show that information is absent. Failure of both chosen families still cannot rule out successful recovery by another readout or training procedure. A control task provides context, not an information-theoretic absence test. What you have shown is how these specific probe procedures performed on this split.
Good decoding does not establish causal necessity. If a fixed readout recovered velocity exactly almost surely throughout the target population, then on that population. Zero error on a finite test set does not establish that condition. Neither result shows that the transition model or policy uses the recovered velocity. Controlled representation interventions can test that use, but an intervention inside a model is not automatically an intervention in the world. Chapter 66 returns to this distinction in interpretability and world-model debugging.
There is a further trap in our toy: the encoder is handcrafted, so this experiment characterises a diagnostic procedure, not a trained world model. Its inputs are known by construction, but statistical relationships between position and velocity still depend on the trajectory distribution. A real encoder requires its own splits, controls and interpretation.
A probe measures what a separate offline readout can recover from a frozen representation. A controller measures what the live system does with that representation under feedback. These are different quantities, and the second is the subject of Chapter 64. Predicting velocity statistically is not proof that a downstream planner uses a reliable velocity estimate.
Object Permanence and Temporal Consistency
Our probe evaluates individual timesteps. Object permanence concerns an object's continued existence when it is unobserved. Tracking its hidden position through a gap is a richer operational diagnostic: a model may preserve existence or identity while remaining uncertain about location or predicting motion inaccurately. Hidden-trajectory evaluation therefore needs its own target in addition to a persistence criterion.
This is where compounding error and memory meet. The previous chapter, "Perceptual and Generative Evaluation," asked whether a generated frame looks right. Here we ask whether the hidden trajectory is right. For an occluded-object control task, position accuracy can matter to the chosen action even when the generated frames look plausible. These are complementary tests: accurate hidden positions do not guarantee good visual generation, and plausible frames do not certify accurate positions.
We should be precise about what "correct" means during an occlusion. Three distinct properties are usually bundled together:
- Hidden trajectory accuracy. The model's estimate of the ball's position while hidden is close to the true position.
- Reappearance accuracy. At the first visible timestep after the gap, the model's belief matches the observation.
- Identity persistence. The object that reappears is the same object, rather than a newly spawned one that happens to be nearby.
Multiple indistinguishable objects can exchange identities while preserving the visible set of positions. An identity-aware test needs persistent labels or another identity criterion. Our single-object toy does not test competing identity assignments; that does not make identity and trajectory accuracy equivalent. A tracked object can keep its label while its predicted position is wrong.
Observational ambiguity also matters. A noisy visible history may be consistent with several hidden continuations. A point loss can still compare predictors by expected risk over realized trajectories, but does not characterize the forecast distribution. Where the task needs uncertainty, report and evaluate a conditional distribution rather than treating one plausible continuation as the uniquely predictable outcome.
Setting up the occluded-object task
The six paths below are the first six held-out trajectories in index order, and the amber markers identify frames hidden by the band. This figure shows ground-truth paths and visibility geometry, not a predictor's errors or reappearance accuracy.
## Deterministic selection rule: the first six held-out trajectories in index order,
## with no resampling and no random draw. SHOW_IDX is fixed at construction time.
N_SHOW = 6
SHOW_IDX = np.arange(N_SHOW)
show_pos = pos_all[test_sl][
SHOW_IDX
] # (N_SHOW, N_STEPS + 1, 2) positions in metres
show_mask = mask_all[test_sl][SHOW_IDX] # 1.0 visible, 0.0 behind the occluder
occ_lo, occ_hi = 0.5 - OCC_HALF, 0.5 + OCC_HALF
We evaluate three predictors with explicitly stated information budgets. Stating the budget matters here more than in almost any other comparison, because the three predictors differ enormously in what they are allowed to see, and a naive comparison across them is meaningless.
- Freeze. Holds the last visible observation fixed: . Uses observations only, no motion model. On a free trajectory without a wall bounce during the gap, this drifts at the object's speed, so its error grows roughly linearly in and is the natural last-observation, no-motion baseline.
- Constant velocity. Estimates velocity from the two immediately preceding visible observations and extrapolates: . It ignores subsequent forces and contacts. Observation noise and those omitted dynamics can all affect its error.
- Privileged constant velocity. Extrapolates the simulator's true position and velocity at , but still ignores subsequent forces and contacts. This baseline discloses simulator-state access; it is neither a full-simulator oracle nor an irreducible-error bound.
The metric is the mean Euclidean position error in metres over the hidden window, indexed by the hidden offset . Formally, writing for the set of occluded runs that survive to offset and for a predictor's estimate on run at hidden step , the reported quantity at offset is
where:
- : the hidden offset in steps of s, running from to .
- : an index over occluded runs in the held-out split.
- : the index of the last visible timestep before run , so is a hidden timestep.
- : the predictor's estimated position at that hidden step.
- : the simulator's true position at that hidden step.
- : the denominator, the number of occluded runs that are at least steps long.
The metric is in metres and is an average over runs at each offset, with the same occluded-run denominator applied to every predictor. The denominator shrinks as grows, since fewer runs survive to large offsets, and we report it alongside the error so you can judge how thin the tail is. Every predictor is evaluated on the same , while their different information budgets remain explicit. At a fixed offset, each eligible run contributes once; this is not a pooled average over all hidden frames.
import numpy as np
## Wall bounces during occlusion are already produced by the reflecting-wall
## dynamics in simulate(); this evaluation reuses the main held-out trajectories
## used by the probes and freeze-drift diagnostic. The bounces come from the physics,
## not from a post-hoc perturbation, so no trajectory is modified after the fact.
def occluded_runs(vis_row):
"""Return maximal (start, stop) index intervals where visibility is 0."""
runs, start = [], None
for i, value in enumerate(vis_row):
if value < 0.5 and start is None:
start = i
elif value >= 0.5 and start is not None:
runs.append((start, i))
start = None
if start is not None:
runs.append((start, len(vis_row)))
return runs
def first_usable_run(vis_row, min_start=2):
"""First hidden run with two immediately preceding visible frames."""
for start, stop in occluded_runs(vis_row):
if (
start >= min_start
and stop > start
and np.all(vis_row[start - 2 : start] >= 0.5)
):
return start, stop
return None
## Regression fixtures: earlier visibility somewhere in the history is not enough.
assert first_usable_run(np.array([1, 0, 1, 0, 0, 1])) is None
assert first_usable_run(np.array([1, 1, 0, 0, 1])) == (2, 4)
assert first_usable_run(np.array([0, 0, 0, 0])) is None
assert first_usable_run(np.array([1, 1, 1, 1])) is None
assert occluded_runs(np.array([1, 1, 0, 0])) == [(2, 4)]K_MAX = 6
pos_te = pos_all[test_sl]
obs_te = obs_all[test_sl]
vel_te = vel_all[test_sl]
mask_te = mask_all[test_sl]
predictor_names = [
"freeze",
"constant velocity",
"privileged constant velocity",
]
err_sum = {name: np.zeros(K_MAX) for name in predictor_names}
n_at_k = np.zeros(K_MAX)
for i in range(pos_te.shape[0]):
for start, stop in occluded_runs(mask_te[i]):
if start < 2 or not np.all(mask_te[i, start - 2 : start] >= 0.5):
continue
p_last = obs_te[i, start - 1]
v_hat = (obs_te[i, start - 1] - obs_te[i, start - 2]) / DT
s_true_last = pos_te[i, start - 1]
v_true = vel_te[i, start - 1]
for k in range(1, min(stop - start, K_MAX) + 1):
j = start - 1 + k
if j >= pos_te.shape[1]:
break
truth = pos_te[i, j]
preds = {
"freeze": p_last,
"constant velocity": p_last + k * DT * v_hat,
"privileged constant velocity": s_true_last + k * DT * v_true,
}
for name, pred in preds.items():
err_sum[name][k - 1] += np.linalg.norm(pred - truth)
n_at_k[k - 1] += 1
mean_err = {
name: np.divide(
err_sum[name], n_at_k, out=np.full(K_MAX, np.nan), where=n_at_k > 0
)
for name in predictor_names
}
## An empty contributing set must be missing, not a fabricated zero error.
_empty_error = np.divide(
np.zeros(2), np.zeros(2), out=np.full(2, np.nan), where=np.zeros(2) > 0
)
assert np.isnan(_empty_error).all() hidden k freeze const. vel priv.CV n runs
1 0.0283 0.0100 0.0024 23
2 0.0539 0.0165 0.0050 22
3 0.0801 0.0242 0.0083 22
4 0.1066 0.0323 0.0121 22
5 0.1350 0.0412 0.0172 19
6 0.1573 0.0482 0.0223 18
Errors are mean Euclidean position error in metres, over held-out runs.
Privileged CV reads the last visible simulator state, not future forces.The table format is deliberate: the same metric, the same denominator, and the same test trajectories for all three predictors, with the information budget of each written down explicitly. This is what makes the comparison readable. Without the disclosed budgets, a reader would have no way to know whether the difference between two rows reflects a modelling improvement or a difference in what the predictors were allowed to see.

The shape of each curve is the diagnosis. The freeze predictor grows approximately linearly in here because it ignores the ball's velocity. Its error is the displacement from the last noisy observation to the realized hidden position, not the path length traveled. Under approximately straight, constant motion and small observation noise, the leading-order relation is
where:
- : the hidden offset in steps.
- : the ball's velocity vector.
- : the mean speed.
- s: the step size.
That is, the hidden offset times the mean speed times the step size, under approximately constant motion. The constant-velocity estimate improves on freezing here, but its residual can reflect observation noise, later driving actions, and contact changes. The privileged baseline extrapolates the true initial position and velocity rather than executing the full simulator. Its non-zero residual cannot be attributed uniquely to wall bounces, and it is not an irreducible-error floor.
The dashed line uses , with mean speed over all hidden held-out frames. That population differs from the eligible run set at each offset, and driving, contacts, and measurement noise also affect the measured curve. Their similarity here is a descriptive comparison, not an exact matched-population identity or a causal decomposition.
## Leading-order prediction of the freeze error: k * dt * mean speed on hidden frames.
## The mean speed is a deterministic statistic of vel_te and mask_te, with no sampling,
## using the same main held-out trajectories as the hidden-horizon predictors.
hidden_frames = mask_te < 0.5
mean_speed_hidden = float(np.linalg.norm(vel_te[hidden_frames], axis=-1).mean())
drift_pred = np.arange(1, K_MAX + 1) * DT * mean_speed_hidden
freeze_measured = mean_err["freeze"]
The comparison names each predictor's information budget and metric denominator. In this driven task, horizon-dependent error mixes initial-state estimation with omitted future forces and possible contacts. It is not a pure measure of learned memory: these predictors are handcrafted, and the constant-velocity baseline has no learned recurrent state.
Sweeping the occlusion length
Horizon is one axis; occluder geometry is another. We sweep the band half-width on fixed undriven trajectories. The metric is position error at the final hidden step of the first usable gap, not at the first visible reappearance. Wider bands change both gap lengths and which trajectories qualify, so this sweep does not isolate elapsed time from sample selection.
sweep_widths = [0.02, 0.05, 0.10, 0.20, 0.30]
rng_sw = np.random.default_rng(404)
p_sw, v_sw, a_sw, o_sw = simulate(rng_sw, 40, driven=False)
sweep_rows = []
for w in sweep_widths:
m_w = visibility(p_sw, occ_half=w)
f_tot = c_tot = t_tot = 0.0
n_used = 0
n_censored = 0
gap_lengths = []
for i in range(p_sw.shape[0]):
run = first_usable_run(m_w[i])
if run is None:
continue
start, stop = run
k = stop - start
if k < 1:
continue
truth = p_sw[i, start - 1 + k]
p_last = o_sw[i, start - 1]
v_hat = (o_sw[i, start - 1] - o_sw[i, start - 2]) / DT
f_tot += np.linalg.norm(p_last - truth)
c_tot += np.linalg.norm(p_last + k * DT * v_hat - truth)
t_tot += np.linalg.norm(
p_sw[i, start - 1] + k * DT * v_sw[i, start - 1] - truth
)
n_used += 1
n_censored += int(stop == m_w.shape[1])
gap_lengths.append(k)
denominator = n_used if n_used else np.nan
sweep_rows.append(
(
w,
float(np.mean(gap_lengths)) if n_used else np.nan,
f_tot / denominator,
c_tot / denominator,
t_tot / denominator,
n_used,
n_censored,
)
) occ. half-width mean gap freeze const. vel priv.CV n censored
0.02 1.70 0.0405 0.0177 0.0000 40 0
0.05 4.38 0.1057 0.0412 0.0000 40 0
0.10 8.68 0.2105 0.0658 0.0000 40 1
0.20 17.11 0.4097 0.0985 0.0000 36 8
0.30 21.57 0.4776 0.1263 0.0000 21 12
Error is measured at the last hidden step of the first gap in each rollout.
All three predictors are evaluated on the same trajectories and the same gaps.The freeze predictor accumulates larger errors as the band widens; the observation-based constant-velocity predictor remains better but also degrades. The privileged extrapolator is essentially exact on these evaluated gaps. This does not demonstrate a bounce-induced jump. A gap still hidden at the simulation endpoint is right-censored: its recorded length is only the observed portion. Report contributing trajectories and censored counts alongside that mean length.

Memory ablations, and what they do and do not show
Recurrent and transformer-based world models carry their hidden state across time in a learned memory. A natural diagnostic is to reset or truncate that memory at the moment the object disappears and see whether performance degrades. If it does, the memory matters for the task. This is an appealing experiment because it is easy to run, but the inference it supports is narrower than it first appears.
A mid-episode reset may be out of distribution if training never used such resets. Context truncation may likewise be unfamiliar, depending on the training protocol. Performance changes measure sensitivity to the stated manipulation, not an isolated memory circuit. Describe the training context and reset schedule before interpreting either ablation.
Four useful temporal properties need different measurements:
- Sequence coherence is a local property: successive predicted states are mutually consistent.
- Long-range dependence is about how far back in time the model's prediction depends on.
- Object identity requires a persistent identity criterion; multi-object gaps make assignment ambiguity particularly visible.
- Uncertainty should be reported, not suppressed: when several hidden continuations are consistent with the visible data, a good model produces a spread of plausible futures rather than one confident guess.
Finally, note the discipline about oracle baselines and information budgets. It is easy to build a comparison in which one method quietly reads the simulator's latent state while another reads only observations, and then report the resulting gap as a modelling improvement. Every baseline in this section is labelled with what it reads. If you introduce an oracle, introduce it loudly. The reason this discipline is worth repeating is that the temptation to compare against an oracle without disclosing it is strong: the oracle often performs better, and the difference can be presented as evidence of modelling quality rather than of information asymmetry.
Physical Rule and Constraint Tests
The previous section asked whether the hidden state is coherent. This one asks whether the dynamics obey a specified physical law under a named regime. A model can produce a smooth, coherent, plausible trajectory that travels at an unphysical speed through a dissipative medium, or that passes through a wall, or that gains energy at every bounce. Constraint tests are how you catch that. They are the natural complement to the trajectory-accuracy tests above: accuracy asks whether the model is close to the truth on average, and constraint tests ask whether the model respects the mechanisms that produce the truth.
Specify the regime before testing a rule. In this force-free, elastic toy, kinetic energy is conserved. Applied forces can do work; damping and lossy contacts can dissipate energy, so a driven dissipative system needs both contributions in its energy balance. A conservation test without a regime can penalize correct dynamics. Here a lossy contact reduces normal kinetic energy only when the incident normal velocity is nonzero.
We will evaluate three regimes, all with the same dynamics code but different parameters:
| Regime | Forcing | Damping | Restitution | Expected energy behaviour over a rollout |
|---|---|---|---|---|
| Force-free, elastic | none | 1.0 | 1.0 | exactly |
| Force-free, damped | none | 0.9 | 1.0 | exactly, since energy scales as |
| Free, lossy contact | none | 1.0 | 0.7 | constant between contacts; normal-component energy is multiplied by at an impact |
The damping factor multiplies both velocity components by each step, so total kinetic energy is multiplied by . At a wall, restitution multiplies only the normal velocity by ; its energy component is multiplied by , while tangential energy is unchanged. Total energy therefore does not generally fall by a factor of . These regimes test different mechanisms, but a residual alone cannot identify which mechanism is wrong.
The metric: end-of-rollout energy ratio residual
For a rollout of length starting from the true initial state , define the predicted energy ratio and the true ratio , where is the kinetic energy
where kg is the mass and is the velocity at step . The energy ratio residual is
with the predicted and true ratios defined as
where:
- : the rollout horizon, steps, so s.
- : the kinetic energy of the model's predicted state at step .
- : the kinetic energy of the simulator's true state at step .
- : the model's predicted energy ratio over the rollout.
- : the true energy ratio produced by the simulator's own dynamics.
These ratios require positive initial energy. The fixtures satisfy that condition; the implementation floors denominators at joules as a numerical safeguard, which changes the ratio convention near zero. In the elastic case ; with damping and no clipping ; with lossy contact the ratio depends on each impact. The violation rate uses the declared absolute ratio-error tolerance , not a universal physical threshold.
Stepwise kinetic-energy equality is the relevant rule only in the force-free, elastic regime here. A correct simulation may have floating-point discrepancies, whereas a wrong predictor may have much larger errors. Comparing with the regime's reference ratio and stating a tolerance makes the evaluated quantity explicit.
We evaluate four predictors in each regime to compare regime dependence and model-class limitations. Simulator truth is reused as an identity reference; evaluating numerical integration drift would require an independent higher-accuracy reference or a time-step convergence study, which this example does not provide.
- True integrator. Simulator truth reused as the reference prediction. Its zero residual follows by identity, not from an independent numerical-accuracy check.
- Linear model fitted on force-free data. A one-step linear map estimated by ridge regression on force-free trajectories.
- Linear model fitted on damped data. The same estimator, fitted on damped trajectories.
- No-change control. The deliberately wrong negative control: . It preserves initial velocity and kinetic energy but freezes position, illustrating that passing an energy check is not sufficient for correct dynamics.
def kinetic_energy(vel, mass=1.0):
return 0.5 * mass * np.sum(vel**2, axis=-1)
def fit_linear_dynamics(pos, vel):
"""Ridge fit of s_{t+1} ~ [s_t, 1] using all transitions in a regime."""
S = np.concatenate([pos, vel], axis=-1)
X = S[:, :-1].reshape(-1, S.shape[-1])
Y = S[:, 1:].reshape(-1, S.shape[-1])
return ridge_fit(np.hstack([X, np.ones((X.shape[0], 1))]), Y, alpha=1e-6)
def rollout(W, s0, n_steps):
"""Recursively apply the learned one-step model. s0 has shape (n, 4)."""
S = np.empty((s0.shape[0], n_steps + 1, s0.shape[1]))
S[:, 0] = s0
for t in range(n_steps):
Xb = np.hstack([S[:, t], np.ones((s0.shape[0], 1))])
S[:, t + 1] = Xb @ W
return S
def evaluate(S_pred, truth, tol=0.05):
E_pred = kinetic_energy(S_pred[..., 2:4])
E_true = kinetic_energy(truth[..., 2:4])
ratio_pred = E_pred[:, -1] / np.maximum(E_pred[:, 0], 1e-12)
ratio_true = E_true[:, -1] / np.maximum(E_true[:, 0], 1e-12)
resid = float(np.mean(np.abs(ratio_pred - ratio_true)))
viol = float(np.mean(np.abs(ratio_pred - ratio_true) > tol))
pos_err = np.linalg.norm(S_pred[..., :2] - truth[..., :2], axis=-1)
return resid, viol, np.sqrt(np.mean(pos_err**2, axis=0))H_TEST = 30
rng_e = np.random.default_rng(2024)
regimes = {
"free / elastic": simulate(
rng_e, 20, driven=False, damping=1.0, restitution=1.0
),
"free / damped": simulate(
rng_e, 20, driven=False, damping=0.9, restitution=1.0
),
"free / lossy contact": simulate(
rng_e, 20, driven=False, damping=1.0, restitution=0.7, v_scale=1.5
),
}
fit_sets = {
"free": simulate(rng_e, 20, driven=False, damping=1.0, restitution=1.0),
"damped": simulate(rng_e, 20, driven=False, damping=0.9, restitution=1.0),
}
W_free = fit_linear_dynamics(fit_sets["free"][0], fit_sets["free"][1])
W_damp = fit_linear_dynamics(fit_sets["damped"][0], fit_sets["damped"][1])
phys_results = {}
for rname, (pos_r, vel_r, act_r, obs_r) in regimes.items():
s0 = np.concatenate([pos_r[:, 0], vel_r[:, 0]], axis=-1)
truth = np.concatenate(
[pos_r[:, : H_TEST + 1], vel_r[:, : H_TEST + 1]], axis=-1
)
models = {
"true integrator": truth,
"linear fit (free)": rollout(W_free, s0, H_TEST),
"linear fit (damped)": rollout(W_damp, s0, H_TEST),
"no-change control": np.repeat(s0[:, None, :], H_TEST + 1, axis=1),
}
phys_results[rname] = {m: evaluate(S, truth) for m, S in models.items()}free / elastic model energy ratio resid violation rate RMSE h=1 RMSE h=30 true integrator 0.0000 0.000 0.00000 0.00000 linear fit (free) 0.9537 1.000 0.00100 0.16624 linear fit (damped) 0.9982 1.000 0.00254 0.49824 no-change control 0.0000 0.000 0.02536 0.71493 free / damped model energy ratio resid violation rate RMSE h=1 RMSE h=30 true integrator 0.0000 0.000 0.00000 0.00000 linear fit (free) 0.0728 0.500 0.00368 0.41512 linear fit (damped) 0.0000 0.000 0.00000 0.00000 no-change control 0.9982 1.000 0.02568 0.24587 free / lossy contact model energy ratio resid violation rate RMSE h=1 RMSE h=30 true integrator 0.0000 0.000 0.00000 0.00000 linear fit (free) 0.5171 1.000 0.00054 0.25837 linear fit (damped) 0.5730 1.000 0.00413 0.34631 no-change control 0.4252 0.850 0.04129 0.58632

The simulator trajectory is compared with itself and therefore has exactly zero residual; this is a reference identity check, not an independent numerical-accuracy test. The damped fit matches damped data closely. The free-trained global linear map cannot represent the piecewise reflection mechanism and has a large elastic-regime energy error. The no-change control preserves energy in that regime while freezing position, so its zero energy error does not certify useful dynamics. The tolerance of 0.05 is an explicit diagnostic choice, not a universal physical threshold.

Separating the sources of a residual
A non-zero constraint residual can come from at least five places, and attributing it correctly is most of the interpretive work:
- Learned dynamics error. The model misrepresents the mechanism.
- Numerical integration drift. Under-resolved integration can introduce error. Diagnosing it needs an independent numerical reference; our identity comparison does not measure it.
- Observation error. If the rollout is initialised from noisy observations rather than the true state, the initial energy estimate itself is noisy.
- Coordinate and unit errors. Inconsistent velocity conversion across initial and final states can distort ratios. A common multiplicative conversion of all velocities cancels in ; ratio agreement cannot detect every unit error.
- A mismatched regime. The model may be perfectly correct for a different regime than the one under test.
The experiment includes regime mismatch and model-class misspecification. A systematic signed residual can arise from either, from calibration bias, or from other errors; its sign does not uniquely diagnose the cause. Compare matched-regime data, held-out mechanisms and an appropriate numerical reference before assigning responsibility.
The next figure compares the free-trained and damped-trained linear fits on the same 20 elastic held-out rollouts, plotting each unit's signed energy-ratio residual instead of its average. The fits also use different realized initial-state draws in their training trajectories, so this comparison does not isolate a change in training regime alone.
## Signed energy ratio residuals on the same 20 elastic rollouts, one value per unit.
def signed_ratio_residual(S_pred, truth):
E_pred = kinetic_energy(S_pred[..., 2:4])
E_true = kinetic_energy(truth[..., 2:4])
rho_pred = E_pred[:, -1] / np.maximum(E_pred[:, 0], 1e-12)
rho_true = E_true[:, -1] / np.maximum(E_true[:, 0], 1e-12)
return rho_pred - rho_true
pos_el, vel_el, _, _ = regimes["free / elastic"]
s0_el = np.concatenate([pos_el[:, 0], vel_el[:, 0]], axis=-1)
truth_el = np.concatenate(
[pos_el[:, : H_TEST + 1], vel_el[:, : H_TEST + 1]], axis=-1
)
residual_matched = signed_ratio_residual(
rollout(W_free, s0_el, H_TEST), truth_el
)
residual_mismatched = signed_ratio_residual(
rollout(W_damp, s0_el, H_TEST), truth_el
)
## Unique trajectory IDs keep all 20 units countable even when values coincide.
trajectory_ids = np.arange(1, residual_matched.size + 1)
Constraint projection can improve a chosen residual while worsening trajectory prediction. Renormalizing a nonzero predicted velocity to its initial magnitude enforces energy conservation in the elastic regime, but is wrong for a dissipative regime's energy target. Clamping position to the box does not enforce energy conservation. A projection must target the actual constraint and satisfy its feasibility conditions; it does not universally drive every residual to zero. Report projected and unprojected errors together.
Finally, note that constraint testing is not a substitute for the state probes of the previous section. A model can respect every conservation law while tracking the wrong object, and it can track the right object while violating energy conservation. These are independent axes of evaluation. The full picture requires both.
Interventions, Counterfactuals, and Causal Fidelity
The occluded-ball question from the previous sections was observational: given what was seen, what happens next? The questions in this section are different. They are about what would happen under a change. The difference between "given what was seen, what happens next" and "what would have happened if we had done something different" is precisely the difference between prediction and causal inference, and no amount of observational data can answer the second kind of question without additional assumptions.
Three targets must be kept apart, because they are estimated differently, identify different quantities, and are routinely confused:
Observational prediction. In a state-observed setup, predict from on the distribution the data were collected under. Other observationally trained sequence models can use observation histories or latent representations as inputs and targets, rather than true simulator states.
Population intervention. Compare and over a specified population. Identification needs a suitable design and assumptions: for example randomized action assignment, or an observational adjustment design with consistency, positivity, and conditional exchangeability. Absence of hidden confounding is not necessary for every possible identification strategy.
Unit-level counterfactual. For one specific unit, estimate what would have been under a different action. In the structural causal model (SCM) used here, hold that unit's initial state and background disturbance fixed across the hypothetical branches. That is a specified cross-world coupling, not a universal ranking of how difficult the three targets are to identify.
Observational fit does not identify any causal quantity without further assumptions. A model can fit the observational distribution to machine precision and still be wrong about the effect of an intervention, and the mechanism by which this happens is confounding.
Building a confounded toy
The earlier simulator sampled actions independently of its state and transition, without a shared action/outcome disturbance. We now switch to a separate one-step structural causal model in which an unobserved disturbance influences both the action and the velocity change:
where:
- : the unobserved exogenous acceleration disturbance, with numerical values in m/s² drawn from a standard normal.
- : the acceleration action in m/s² chosen by the state-feedback policy with gain s; the position reference is in metres.
- : the position at step .
- : the next velocity, driven by the action plus the disturbance coupling .
- s: the step size.
The action is chosen by a state-feedback policy, but the policy's exploration noise is the same that enters the dynamics with coupling . An engineer observing only never sees , so the observational data are confounded by construction. This toy illustrates one possible shared-signal mechanism: the observational correlation between action and outcome mixes the direct effect of the action with the effect of the shared signal. It does not establish how often that mechanism occurs in deployed systems.
Now ask: what is the causal effect of the action on the one-step velocity change? Holding fixed and differentiating, the structural derivative is
where is the one-step velocity change, is the action, and s is the step size. The derivative is because the structural equation is linear in the action with slope . This is the quantity we want to recover, and it is a structural quantity: it describes how the next velocity would change if we intervened on the action while holding everything else fixed.
That is the ground truth we want to recover. Now see what an observational regression gives. Rearranging the policy equation gives . Substituting this into the dynamics,
where has been substituted. The coefficient on is rather than , so the observational relationship overstates the structural action effect. The empirical analogue is a regression of the observed velocity change on the observed action and position.
In the noiseless structural relationship, the observational coefficient is , twice the intervention slope . The measured regression below includes observation noise and fits nearly, not exactly, perfectly. A predictive coefficient need not be an intervention effect.
KAPPA = 2.0 # policy gain
C_U = 1.0 # strength of the unobserved disturbance coupling
DT_C = 0.1 # s
def sample_units(rng, n):
"""One-step structural units with a policy that depends on the disturbance."""
x0 = rng.uniform(0.2, 0.8, n)
u = rng.normal(0.0, 1.0, n)
a = -KAPPA * (x0 - 0.5) + u # action depends on the unobserved u
dv = (a + C_U * u) * DT_C # true one-step velocity change
return x0, u, a, dvrng_c = np.random.default_rng(7)
x_all, u_all, a_all, dv_all = sample_units(rng_c, 4000)
dv_meas = dv_all + rng_c.normal(
0.0, 0.005, dv_all.size
) # measurement standard deviation 0.005 m/s, not a relative percentage
fit_slice = slice(0, 2000)
held_slice = slice(2000, None)
## Observational regression of dv on (x - 0.5, a). No interventions were used.
X_fit = np.hstack(
[
np.ones((2000, 1)),
(x_all[fit_slice] - 0.5)[:, None],
a_all[fit_slice][:, None],
]
)
beta_obs = ridge_fit(X_fit, dv_meas[fit_slice][:, None], alpha=1e-8)
effect_obs = float(beta_obs[2, 0])
X_held = np.hstack(
[
np.ones((2000, 1)),
(x_all[held_slice] - 0.5)[:, None],
a_all[held_slice][:, None],
]
)
pred_held = X_held @ beta_obs
ss_res = float(np.sum((dv_meas[held_slice][:, None] - pred_held) ** 2))
ss_tot = float(np.sum((dv_meas[held_slice] - dv_meas[held_slice].mean()) ** 2))
r2_obs = 1.0 - ss_res / ss_tot
support_lo, support_hi = float(a_all.min()), float(a_all.max())The held-out observational fit is , while its fitted action coefficient is about instead of the structural slope . Fit alone does not reveal the confounding in this example. The nearly perfect noisy fit should not be described as machine precision.
The paired counterfactual experiment
Now we run a paired simulator intervention on the held-out units. For each unit:
- Clone the unit's initial state and its exogenous disturbance .
- Run it once with its natural action .
- Run it again with the same state and the same , but with the action shifted by a fixed .
- Subtract the two outcomes.
Because is shared, the difference isolates the direct effect of the action. Substituting the structural equation into the two branches,
For the noiseless structural outcomes, dividing by nonzero recovers exactly. The code then adds independent measurement noise to both branches and averages across units, so its estimate has sampling error. Shared structural is the simulator's explicit counterfactual coupling; passive video does not supply that coupling by itself.
DELTA = 0.5
x_h = x_all[held_slice]
u_h = u_all[held_slice]
a_h = a_all[held_slice]
dv_base = (a_h + C_U * u_h) * DT_C
dv_shift = ((a_h + DELTA) + C_U * u_h) * DT_C
noise_base = rng_c.normal(0.0, 0.02, a_h.size)
noise_shift = rng_c.normal(0.0, 0.02, a_h.size)
effect_paired = float(
np.mean((dv_shift + noise_shift - dv_base - noise_base) / DELTA)
)
effect_true = float(DT_C)
effect_rows = [
("observational regression", effect_obs),
("paired counterfactual", effect_paired),
("ground truth (simulator)", effect_true),
]estimator effect of a on dv relative error observational regression 0.1999 99.9% paired counterfactual 0.0996 0.4% ground truth (simulator) 0.1000 0.0% observational fit on held-out natural data: R^2 = 0.9993 observed action extrema: [-3.703, 3.692] m/s^2 intervention magnitude delta = 0.50; fraction above observed maximum = 0.1% no intervened data was used to fit the observational model.
The observational slope is about , whereas the paired noisy estimate is about versus structural truth . The latter has roughly relative error in this run. Pairing is sufficient for this demonstration, not necessary for every population effect: assigning independently sampled units to fixed actions by randomization can identify the difference of population mean outcomes when the intervention is well-defined and observed outcomes are consistent with their assigned actions. Here the two groups have the same disturbance distribution, so their mean structural contrast divided by the action difference is . This experiment evaluates an estimator using the simulator's equations; it does not benchmark a learned model's counterfactual predictions.

What else to report alongside a causal claim
A single effect estimate is not enough to make a causal claim legible. Report at least the following. Each item corresponds to a distinct way in which a reader might reasonably disagree with the claim, and each is a way of making the claim falsifiable rather than rhetorical.
- The exact target. Average treatment effect, conditional effect at a specific state, or a unit-level counterfactual? These are different numbers. State which one, and state the unit of the estimate.
- Intervention support and overlap. Distinguish population support from finite observed extrema and from practical data coverage. Here Gaussian disturbances give actions full real-line population support, but about of shifted held-out actions exceed the finite sample maximum. Being within the sample range does not establish dense coverage or validate a causal estimate.
- Which interventions were observed during training. Here, none: the model was fitted purely observationally, and every intervention was held out. If a model was trained on interventions, say which, and reserve a distinct held-out intervention family for evaluation.
- The feasibility of the intervention. Some interventions are not physically implementable in the source system. A simulated intervention on an unfeasible variable is a modelling exercise, not evidence about the world.
- Confounding and observability. Write down which variables are observed, which are latent, and which one is assumed to be the exogenous driver. Our pairing stipulates a shared structural across the two branches. Passive observational data alone do not generally identify that cross-world coupling, even if a dataset measures a relevant disturbance.
- Shared-noise coupling. A unit-level counterfactual needs an explicit joint model across branches. Changing that coupling can change paired differences and estimator variance. With fixed branch marginals, the population difference of expected outcomes is unchanged; do not confuse a unit counterfactual with an average effect.
- Simulator-to-world limitations. Every result in this section lives inside a simulator whose structural equations we wrote. The same protocol applied to real video would require the structural model to be justified externally, and that justification is usually where the argument lives.
Invariance across environments is diagnostic evidence, not an identification argument by itself. A stable coefficient can coexist with confounding, including an unchanged confounding mechanism across the sampled environments. State which environments changed and which identification assumptions the design supports; coefficient stability alone does not establish causal fidelity.
Action conditioning does not guarantee causal validity. Our observational regression takes action as an input yet estimates the wrong intervention slope. A causal method must represent the specified intervention somehow, but it need not be an action-conditioned prediction network; identification can use a different suitable experimental or structural design.
Prompting a learned model to generate a counterfactual specifies a desired output, not an identification argument. Validate the target effect using an appropriate intervention design and state the structural assumptions. A paired simulator experiment is one option, not a universal requirement.
Limitations and Impact
Probing, occlusion sweeps, constraint tests, and intervention designs turn broad claims into narrower questions: which targets a readout recovers, how a predictor behaves while observations are hidden, which regime-specific constraints it respects, and which action effect a design identifies. The examples illustrate these procedures; they do not establish that the same measurements transfer unchanged to a trained world model.
The limitations are equally important, and they cluster into three themes.
The first three families use the reflecting point mass under named regimes; the causal example uses a separate one-step structural model. Their mechanisms, noise scales, and sampling policies are specified by construction, so the numerical values are illustrative rather than benchmarking evidence. Different fixtures can change the probe scores and residuals. The examples make the procedures explicit; they do not establish performance on a larger system.
Each diagnostic answers a specified question. A probe report measures recovery by a named procedure on a named split, not unrestricted knowledge or downstream use. The occluded-object tables compare three handcrafted predictors, not learned recurrent memory. The energy table checks global linear fits on both matched and mismatched regimes, alongside an identity reference and frozen-state control; it does not predict how a neural dynamics model would perform. Interpret each number within that scope.
The paired example cancels the shared disturbance exactly in its structural outcomes, while its noisy sample average estimates the slope approximately. Identification relies on the stated structure and design. Simulator results do not validate that structure for a real system; real interventions can, however, provide evidence that distinguishes some candidate mechanisms. Separate assumptions that a design checks from those it leaves untested.
Constraint satisfaction alone is insufficient. A projection enforcing the correct energy target can reduce that residual while worsening positions, and the frozen-state control passes the elastic energy test despite large position error. Our reused simulator truth is an identity reference, not an integration-accuracy floor.
Choose diagnostics for the task rather than treating the families as mandatory rungs. Probe accessibility with declared capacities, controls, and splits. Evaluate hidden windows with contributing-run and censoring counts. Test constraints under named regimes, alongside trajectory error and suitable reference baselines. For a causal target, state the identification design, overlap, feasibility, and structural assumptions before the estimate.
These offline diagnostics do not by themselves establish whether a planner using the model will make better decisions, rank policies correctly, or exploit an inaccurate region of the model. Chapter 64 evaluates executed returns alongside planner compute, policy ranking, and transfer; those decision diagnostics need not demonstrate a failure in every experiment. Benchmark taxonomy and statistical power are the subject of Chapter 65, and the mechanistic question of which internal components carry the state is taken up in Chapter 66. Failure modes and model exploitation are the opening subject of Chapter 67. Exploitation can make model-predicted returns optimistic, but a model's prediction gap can also be pessimistic.
Summary
This chapter developed four evaluation families for asking what a world model represents, and each was anchored to a decision about information budgets and metric denominators.
-
State probing measures accessibility to a specified readout on a specified distribution. Different probe capacities and controls can expose fitting failures, but failed decoding does not prove absent information. Trajectory-level splits and training-only preprocessing prevent particular leakage routes. Good decoding does not establish causal necessity or downstream use.
-
Object permanence and temporal consistency concern hidden trajectories, reappearance and identity, which need separate tests. Horizon and occluder-width sweeps require explicit contributing-sample counts. Privileged baselines must disclose their information; incomplete extrapolators are not automatically error bounds. Memory ablations test sensitivity to a manipulation, not isolated memory circuits.
-
Physical tests require a named regime, units and tolerance. Energy-ratio agreement can coexist with wrong trajectories, as the frozen-state control shows. Regime mismatch and model-class misspecification can both cause residuals; the sign alone does not identify the cause. An identity comparison with simulator truth is not an integration-accuracy certification.
-
Interventions, counterfactuals, and causal fidelity separate observational prediction, population interventions, and unit counterfactuals. Here a near-perfect observational fit still gives roughly twice the structural action slope. Shared disturbance cancels in noiseless paired differences; the independently noisy paired average estimates the effect approximately. Report the target, identification design, practical overlap, feasibility, training interventions, structural coupling, and simulator-to-world limits.
The thread connecting all four is the same: state what the model received, state what the baseline received, name the quantity the metric measures, and separate offline prediction quality from closed-loop decision usefulness. That last separation is where the next chapter begins.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about state, physics, causality, and memory evaluation.
State, Physics, Causality, and Memory Evaluation
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!