State, Physics, Causality, and Memory Evaluation

Michael BrenndoerferAugust 2, 202662 min read

Part of World Models Handbook

Test world-model state, physics, causality, and memory with probes, occlusion, energy checks, and paired interventions; learn what each diagnostic can prove.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

State, Physics, Causality, and Memory Evaluation

Picture a small ball rolling across a table. For two seconds it is plainly visible. Then it passes behind a box, and for eight frames you see nothing at all. Then it emerges on the other side. Now ask a world model what happened in between.

This is a deceptively simple question. It sounds like a single question, but it is really a bundle of at least four distinct questions, and world models can answer some of them correctly while failing the others. That separation is the entire subject of this chapter.

There are at least four different things a model could do here, and they are easy to confuse. Among them, only the last is a causal question: the first three are observational or representational.

  • It could generate eight frames that look like a ball continuing across the box. The pixels are plausible; the ball is in roughly the right place.
  • It could maintain a hidden state that keeps the ball's position and velocity coherent, so that the ball emerges at exactly the right place at the right time.
  • It could keep the ball's identity intact, so that the thing that emerges is the same object rather than a new one that happens to be nearby.
  • It could do all of the above while also knowing that if you had pushed the ball differently just before it disappeared, it would have emerged somewhere else.

These are different requirements, not a strict hierarchy. Accurate positions do not guarantee persistent identity, and identity tracking does not guarantee accurate motion. An intervention asks a further question about changes to the world that observational prediction alone does not identify without assumptions. Each requirement needs its own evaluation protocol.

A model can pass the visual test while failing hidden-state, identity, or intervention tests. Even strong results across these offline diagnostics do not establish closed-loop usefulness: a planner may fail when it acts on the predictions. Chapter 64 evaluates decisions, Chapter 65 develops experimental design, and Chapter 67 examines failure modes and model exploitation.

This chapter is about how to tell these cases apart. It is in Part XI: Evaluation and Understanding. The previous chapter, "Perceptual and Generative Evaluation" (Chapter 62), handled the question of whether generated observations look right and whether their distribution matches the data. Here we develop state, constraint, and intervention diagnostics within the world-model evaluation ladder established in Chapter 61, without treating those diagnostics as nested requirements.

The chapter is organized around four evaluation families:

  • State probing and representation quality. When you freeze a representation and fit a small readout, what can it recover, and what does that recovery prove?
  • Object permanence and temporal consistency. Does the model maintain a coherent hidden trajectory through occlusion, and for how long?
  • Physical rule and constraint tests. Does the model respect the conservation laws and boundary conditions of the regime it claims to model?
  • Interventions, counterfactuals, and causal fidelity. Under a change to one mechanism, does the model predict the right change, or merely fit the observational distribution?

The four families examine separate properties. Probing measures accessibility to a readout; occlusion tests measure behavior over hidden windows; constraint tests check a named rule; interventions test a specified change under an identification design. None universally subsumes the others. Running them separately makes their different failure modes visible.

State what information the model received, what information the baseline received, and what the metric denominator counts. Compare methods on a named population, and distinguish differences in modelling from differences in information access. A privileged-state baseline can be useful, but its advantage over an observation-only method cannot be attributed entirely to better modelling.

The first three families use a point mass moving in the unit square with reflecting walls, an occluder band, and noisy position observations. The causal section then switches to a separate one-step structural model so that its confounding mechanism is explicit. The point-mass state lives in R4\mathbb{R}^4 as

st=(xt,yt,vtx,vty)s_t = (x_t, y_t, v^x_t, v^y_t)

where:

  • xt,ytx_t, y_t: position components in metres.
  • vtx,vtyv^x_t, v^y_t: velocity components in metres per second.

We observe only the position components through oto_t, and the dynamics are the deterministic map

st+1=f(st,at)s_{t+1} = f(s_t, a_t)

defined in the simulator below, where ata_t is the acceleration action.

The reflecting-wall simulator returns one next state given state and action; its random initial conditions, action sampling, and observation noise are separate draws. A structural model with a specified disturbance likewise returns one outcome once all its inputs are fixed. The later causal example introduces a shared disturbance uu in such a one-step model. Reusing that disturbance across counterfactual branches is an explicit coupling assumption, not a property inferred from passive observations.

The point-mass simulator fits in a few dozen lines of NumPy and lets us inspect probe, occlusion, and constraint diagnostics against known truth. Neither it nor the later causal model is a benchmark for trained foundation models. Treat them as controlled fixtures: their simplicity makes each diagnostic's target and assumptions visible, without establishing its performance on a larger system.

Notation used in this chapter

Physical simulator state st=(xt,yt,vtx,vty)∈R4s_t = (x_t, y_t, v^x_t, v^y_t) \in \mathbb{R}^4 in metres and metres per second. Observed measurement ot∈R2o_t \in \mathbb{R}^2 in metres. Action at∈R2a_t \in \mathbb{R}^2 in metres per second squared. Representation zt∈Rdz_t \in \mathbb{R}^d; the example encoder is handcrafted. The main simulator has deterministic transitions given state and action, with separate random initial conditions, sampled actions, and measurement noise. The later one-step causal model introduces an unobserved disturbance uu shared by its action and outcome equations. Rollout horizon HH. Time step Δt=0.1\Delta t = 0.1 s.

State Probing and Representation Quality

A probe is a small model fitted on top of a frozen representation. You take the encoder, stop its gradients, run it over a dataset, and then ask: can a simple readout predict quantity qq from the representation ztz_t? Saying that the representation "contains" qq can conflate recoverability by a chosen readout, unrestricted information content, and downstream use. We need to ask how the information is arranged and which procedure can read it out.

Frozen representation and probe

A frozen representation is a fixed map zt=enc(o≤t)z_t = \mathrm{enc}(o_{\le t}) whose parameters do not change during evaluation. A probe is a separately fitted function gθg_\theta trained to predict a target qtq_t from ztz_t. The probe's held-out score measures how accessible qtq_t is to that probe family, not whether the encoder "knows" qtq_t in any deeper sense.

The three questions are:

  1. Linear accessibility. How well does a fitted linear readout recover qtq_t from ztz_t on held-out data? Success demonstrates accessibility to that fitted readout on the tested distribution. Failure can reflect regularization, sample size, fitting, or distribution shift; it does not prove that no linear readout could succeed. A linear probe still learns a task-specific mapping.
  2. Information decodability. Can some function in a richer probe family recover qtq_t? A flexible probe can learn substantial task structure from its frozen inputs, so high performance depends jointly on the encoded information, the readout's capacity, and its fitting procedure. Capacity cannot supply target information absent from those inputs.
  3. Downstream functional use. Does the rest of the system use qtq_t to make decisions? This is the only one of the three that connects to behaviour, and it cannot be established by a probe at all. A probe is a diagnostic instrument attached to the side of the model; it does not tell you what the model does with the information when it acts.

Comparing readouts can reveal which fitted family works better on the test distribution. Better nonlinear scores do not establish impossibility of linear recovery, and failures under both families do not prove absent information. Good decoding also does not establish causal necessity: the downstream system may ignore the recovered variable. Testing functional use requires additional experiments, such as controlled representation interventions and closed-loop comparisons.

Probe splits, controls, and preprocessing

Check these three procedural details before interpreting a probe score. Splits and preprocessing determine what information is held out; controls clarify how the selected probe behaves on an alternative task. Mistakes can bias an evaluation or limit the conclusions it supports.

The probe itself is a fitted map gθg_\theta from the frozen representation ztz_t to a target qtq_t. It is a simple readout, not a full model. In the linear case, gθ(zt)=Wzt+bg_\theta(z_t) = W z_t + b; in the nonlinear case, the probe first maps ztz_t through a fixed feature function and then applies a ridge readout. The data are the held-out trajectories, and all probe scores are computed against targets that were not used to fit either the encoder or the probe, and trajectories are held out as complete sequences rather than individual frames so that frames from one rollout never appear in two splits. The key design choice is that the encoder is frozen: the probe cannot reshape the representation to make its job easier, so any information the probe recovers must already be present in ztz_t.

Split at the trajectory level, not the frame level. Consecutive frames of one trajectory are strongly correlated: neighbouring timesteps share nearly the same position. Random frame splits can put near-duplicates of test frames in the training set and inflate the probe's held-out score through memorisation. Splitting whole trajectories removes this within-trajectory leak. In our toy we hold out complete rollouts. Choose a split unit that respects the dependencies in the data; shared scenes or subjects may require a larger grouping than one trajectory.

Fit preprocessing on training data only. Standardisation statistics, whitening matrices, and PCA bases must be estimated on the training split and then applied unchanged to validation and test. Fitting them on the full dataset uses held-out information. Its effect on a score depends on the data and fitting procedure; neither inflation nor a flat train/test gap is guaranteed.

Use a control task with matched probe capacity. Hewitt and Liang (2019) developed type-level randomized control tasks for linguistic probes. Our shuffled regression targets are a simpler toy control, not a reproduction of their protocol. Report real-target and control results together, with the randomization and split specified. A control gives context about the chosen fitting procedure; it does not bound all structured-task learning or certify information content.

The toy representation

To keep this concrete, let us build a frozen representation by hand and be explicit about what it is. We will use a phase encoding of position only. In the formula and NumPy code, xtx_t and yty_t denote the numerical values of observed coordinates in metres: the dimensionless phase is 2π2\pi times the coordinate divided by a one-metre period.

zt=[ cos⁡(2πxt),  sin⁡(2πxt),  cos⁡(2πyt),  sin⁡(2πyt) ]∈R4.z_t = \big[\,\cos(2\pi x_t),\; \sin(2\pi x_t),\; \cos(2\pi y_t),\; \sin(2\pi y_t)\,\big] \in \mathbb{R}^4 .

where:

  • xtx_t: the observed horizontal position at time tt, in metres.
  • yty_t: the observed vertical position at time tt, in metres.
  • cos⁡(2πxt),sin⁡(2πxt)\cos(2\pi x_t), \sin(2\pi x_t): the phase encoding of the horizontal coordinate as a point on the unit circle; together they determine its numerical value modulo one. Exact inversion over a full period is nonlinear; this does not preclude a useful approximate linear readout over a restricted distribution.
  • cos⁡(2πyt),sin⁡(2πyt)\cos(2\pi y_t), \sin(2\pi y_t): the phase encoding of the vertical coordinate as a point on the unit circle, with the same periodic-recovery property.

This is a handcrafted encoder, not a learned one. Exact recovery of a noiseless position within a specified one-period interval uses nonlinear phase inversion; positions differing by an integer share the same encoding, and noisy observations near the interval boundary complicate recovery. The encoder does not directly take velocity, but that does not make velocity statistically independent of its output. Our sampled trajectories can correlate position and velocity. The experiment tests these readouts on a named distribution, not an absolute present-versus-absent information theorem.

The toy world and the dataset

Our world is a point mass in the unit square with perfectly reflecting walls. Let us fix the conventions before writing any metric. Fixing conventions up front is not a formality; it is what makes the numbers in the tables below interpretable and comparable across sections.

  • State: st=(xt,yt,vtx,vty)s_t = (x_t, y_t, v^x_t, v^y_t), metres and metres per second.
  • Action: ata_t, metres per second squared, applied as an acceleration.
  • Time step: Δt=0.1\Delta t = 0.1 s, so one trajectory of 30 steps is 3 seconds.
  • Observation: ot=(xt,yt)+ϵto_t = (x_t, y_t) + \epsilon_t with independent coordinate noise ϵt∼N(0,(0.004 m)2I2)\epsilon_t \sim \mathcal{N}(0, (0.004\ \text{m})^2 I_2).
  • Occluder: the band x∈[0.5−w,0.5+w]x \in [0.5 - w, 0.5 + w] for y∈[0.05,0.95]y \in [0.05, 0.95], with ww the occluder half-width.
  • Data-generating policy: at∼Uniform(−0.3,0.3)2a_t \sim \mathrm{Uniform}(-0.3, 0.3)^2 m/s², independent across components and timesteps.
  • Split unit: one complete trajectory. We use 60 for training, 20 for validation, 20 for test.
In[3]:
Code
import numpy as np

DT = 0.1  # s
V_MAX = 0.6  # m/s
A_MAX = 0.3  # m/s^2
N_STEPS = 30  # 30 steps = 3.0 s
SIGMA_OBS = 0.004  # m
OCC_Y = (0.05, 0.95)
OCC_HALF = 0.10  # half-width of the occluder band in x


def simulate(
    rng,
    n_traj,
    n_steps=N_STEPS,
    driven=True,
    damping=1.0,
    restitution=1.0,
    v_scale=1.0,
):
    """State s_t = (x, y, vx, vy) in metres and metres per second.
    Returns positions (n, T+1, 2), velocities (n, T+1, 2),
    actions (n, T, 2), and noisy observations (n, T+1, 2).
    """
    pos = np.empty((n_traj, n_steps + 1, 2))
    vel = np.empty((n_traj, n_steps + 1, 2))
    act = np.zeros((n_traj, n_steps, 2))

    pos[:, 0, 0] = rng.uniform(0.10, 0.30, n_traj)
    pos[:, 0, 1] = rng.uniform(0.30, 0.70, n_traj)
    vel[:, 0, 0] = v_scale * rng.uniform(0.15, 0.35, n_traj)
    vel[:, 0, 1] = v_scale * rng.uniform(-0.15, 0.15, n_traj)

    for t in range(n_steps):
        if driven:
            act[:, t] = rng.uniform(-A_MAX, A_MAX, size=(n_traj, 2))
        v = damping * vel[:, t] + act[:, t] * DT
        v = np.clip(v, -V_MAX, V_MAX)
        p = pos[:, t] + v * DT
        for axis in (0, 1):
            lo = p[:, axis] < 0.0
            hi = p[:, axis] > 1.0
            p[lo, axis] = -p[lo, axis]
            p[hi, axis] = 2.0 - p[hi, axis]
            hit = lo | hi
            v[hit, axis] = -restitution * v[hit, axis]
        p = np.clip(p, 0.0, 1.0)
        pos[:, t + 1] = p
        vel[:, t + 1] = v

    obs = pos + rng.normal(0.0, SIGMA_OBS, size=pos.shape)
    return pos, vel, act, obs


def visibility(pos, occ_half=OCC_HALF):
    """1.0 where the ball is visible, 0.0 where it is behind the occluder."""
    hidden = (
        (np.abs(pos[..., 0] - 0.5) <= occ_half)
        & (pos[..., 1] >= OCC_Y[0])
        & (pos[..., 1] <= OCC_Y[1])
    )
    return (~hidden).astype(float)


def encode(pos):
    """Frozen handcrafted representation: phase encoding of position only."""
    return np.stack(
        [
            np.cos(2 * np.pi * pos[..., 0]),
            np.sin(2 * np.pi * pos[..., 0]),
            np.cos(2 * np.pi * pos[..., 1]),
            np.sin(2 * np.pi * pos[..., 1]),
        ],
        axis=-1,
    )

Now we generate the training, validation, and test trajectories. The split is made on the trajectory axis, so no single trajectory contributes frames to two different splits. This is the discipline introduced above, applied here: the unit of the split is one whole trajectory, not one frame.

In[4]:
Code
rng = np.random.default_rng(11)
N_TRAIN, N_VAL, N_TEST = 60, 20, 20
## NOTE: several diagnostics below are sensitive to the seed; keep the seed fixed.
## Probes, hidden-horizon errors and freeze-drift diagnostics share this main split.
## Width sweeps and physical-constraint tests use separately generated fixtures.

pos_all, vel_all, act_all, obs_all = simulate(
    rng, N_TRAIN + N_VAL + N_TEST, driven=True
)
mask_all = visibility(pos_all)
## act_all already has shape (n, T, 2)

train_sl = slice(0, N_TRAIN)
val_sl = slice(N_TRAIN, N_TRAIN + N_VAL)
test_sl = slice(N_TRAIN + N_VAL, None)
Out[5]:
Console
trajectories: train=60  val=20  test=20
frames per trajectory: 31
maximum absolute velocity component: 0.553 m/s (per-component clamp 0.6)
fraction of hidden frames: 0.278
unit of the split: one complete trajectory

The hidden fraction tells us how much of the data sits behind the occluder. At 0.278 it is not small, and occlusion is not rare in this dataset. The reason occlusion matters out of proportion to its frequency is that a rare failure to maintain a coherent hidden state can dominate the downstream consequences, even if the model is otherwise accurate on 90% of the data.

Fitting the probes

We now build the probe dataset. Because the encoder is defined on observed positions, we keep only timesteps where the ball is visible, and we standardise the representation using training statistics only. Standardising with statistics computed on the full dataset is the leakage trap described above, so we are careful to compute the mean and standard deviation from the training split and apply the same transformation to validation and test.

In[6]:
Code
def visible_arrays(sl, pos, vel, obs, mask):
    Z = encode(obs[sl])
    keep = mask[sl].astype(bool)
    return Z[keep], pos[sl][keep], vel[sl][keep]


Z_tr, P_tr, V_tr = visible_arrays(train_sl, pos_all, vel_all, obs_all, mask_all)
Z_va, P_va, V_va = visible_arrays(val_sl, pos_all, vel_all, obs_all, mask_all)
Z_te, P_te, V_te = visible_arrays(test_sl, pos_all, vel_all, obs_all, mask_all)

mu, sd = Z_tr.mean(axis=0), Z_tr.std(axis=0) + 1e-9
Zt, Zv, Zs = (Z_tr - mu) / sd, (Z_va - mu) / sd, (Z_te - mu) / sd

## Targets: position (x, y) and velocity (vx, vy)
Y_tr = np.hstack([P_tr, V_tr])
Y_te = np.hstack([P_te, V_te])

print(f"train frames={Zt.shape[0]}  val={Zv.shape[0]}  test={Zs.shape[0]}")
In[7]:
Code
def ridge_fit(X, Y, alpha=1e-3):
    return np.linalg.solve(X.T @ X + alpha * np.eye(X.shape[1]), X.T @ Y)


def r2_columns(Y, Yhat):
    ss_res = np.sum((Y - Yhat) ** 2, axis=0)
    ss_tot = np.sum((Y - Y.mean(axis=0)) ** 2, axis=0)
    if np.any(ss_tot <= 0):
        raise ValueError("R-squared requires nonconstant test targets")
    return 1.0 - ss_res / ss_tot


def with_bias(X):
    return np.hstack([X, np.ones((X.shape[0], 1))])


## Undefined constant-target R-squared is rejected rather than silently assigned.
try:
    r2_columns(np.ones((3, 1)), np.ones((3, 1)))
except ValueError:
    pass
else:
    raise AssertionError("Constant-target R-squared must be rejected")

We fit four probe procedures on the same frozen features. The linear and nonlinear readouts test different forms of accessibility, the shuffled-target control reveals how the selected fitting procedure behaves on that control task, and the intercept-only baseline provides a trivial reference. The control score is not a capacity floor.

  • an intercept-only baseline, which predicts a fitted training-target constant and can score below zero on held-out R2R^2 when training and test means differ;
  • a linear probe, a ridge readout directly on zz;
  • a nonlinear probe, a ridge readout on fixed random Fourier features ϕ(z)=cos⁡(Wz+b)\phi(z) = \cos(Wz + b). With Gaussian frequencies and uniform phases, the normalized inner product (2/D)ϕ(z)⊤ϕ(z′)(2/D)\phi(z)^\top\phi(z') approximates a Gaussian kernel. We use the unnormalized features with the stated ridge penalty, so normalization would also change the effective regularization;
  • a control probe, the same nonlinear feature map, but trained on randomly permuted targets.
In[8]:
Code
## Intercept-only baseline
W_mean = ridge_fit(np.ones((Zt.shape[0], 1)), Y_tr)

## Linear probe
W_lin = ridge_fit(with_bias(Zt), Y_tr)

## Nonlinear probe: random Fourier features over the frozen representation
D_FEAT = 800
W_rff = rng.normal(0.0, 3.0, size=(D_FEAT, Zt.shape[1]))
b_rff = rng.uniform(0.0, 2 * np.pi, size=D_FEAT)


def phi(Z):
    return np.cos(Z @ W_rff.T + b_rff)


W_nl = ridge_fit(with_bias(phi(Zt)), Y_tr)

## Control probe: identical pipeline, randomly permuted training targets
perm = rng.permutation(Y_tr.shape[0])
W_ctrl = ridge_fit(with_bias(phi(Zt)), Y_tr[perm])

r2_mean = r2_columns(Y_te, np.ones((Zs.shape[0], 1)) @ W_mean)
r2_lin = r2_columns(Y_te, with_bias(Zs) @ W_lin)
r2_nl = r2_columns(Y_te, with_bias(phi(Zs)) @ W_nl)
r2_ctrl = r2_columns(Y_te, with_bias(phi(Zs)) @ W_ctrl)

probe_rows = [
    ("x position", r2_mean[0], r2_lin[0], r2_nl[0], r2_ctrl[0]),
    ("y position", r2_mean[1], r2_lin[1], r2_nl[1], r2_ctrl[1]),
    ("vx velocity", r2_mean[2], r2_lin[2], r2_nl[2], r2_ctrl[2]),
    ("vy velocity", r2_mean[3], r2_lin[3], r2_nl[3], r2_ctrl[3]),
]
Out[9]:
Console
target        mean-only   linear  nonlinear  control
x position       -0.034    0.826      0.963 -137.543
y position       -0.284    0.499     -2.135  -32.485
vx velocity      -0.001    0.012   -101.539 -109.621
vy velocity      -0.053    0.285   -170.657 -210.255

All scores are held-out R^2 on whole held-out trajectories.
The control column tests this probe family on shuffled training targets.

Read the table column by column. Each column has a specific interpretation, and the value of the table comes from comparing them.

  • The mean-only column uses an intercept-only ridge fit, a training-target constant slightly shrunk toward zero by the penalty. It can score below zero because test R2R^2 compares against the test-target mean, which need not equal that fitted constant.
  • The linear column measures accessibility to this linear readout. In this finite sample, xx and yy position achieve about 0.83 and 0.50, despite the phase encoding. Linear access need not vanish merely because the encoding is nonlinear.
  • The nonlinear column uses a high-capacity random-feature readout. It improves xx position but fails badly on yy and velocity. Its very negative scores indicate poor held-out prediction, not a lack of information established by the test.
  • The control column fits shuffled training targets. Its strongly negative scores show that this chosen probe can generalise poorly on the control task; they are not a universal capacity floor or a bound on what another probe could recover.

The four targets do not support a simple present-versus-absent classification: the results are mixed across probe families and state components. The handcrafted encoder explicitly takes only the current noisy position, yet correlations in the sampled trajectories can let a readout predict some velocity. The nonlinear readout's failure also shows why adding capacity does not guarantee a better diagnostic. Report these outcomes rather than replacing them with the intended story. The experiment illustrates probe sensitivity to fitting and data distribution; it does not prove that velocity information is absent or that position is inaccessible linearly.

Out[10]:
Visualization
Grouped bar chart of held-out R-squared for four probe families across four targets.
Held-out R-squared for four probe procedures and four state targets. The nonlinear readout improves x position but gives very negative scores on other targets; shuffled-target controls also generalise poorly. The symmetric-log vertical scale retains these failures rather than clipping them. These finite-probe results do not establish absent information.

Three things a probe can never tell you

Interpret the probe scores with three limits in mind: readout capacity, failed recovery, and downstream use.

A flexible probe may be doing substantial task learning from the information available in its frozen inputs, not creating target information missing from them. A nonlinear readout's gain can depend on both the encoded features and the readout's capacity and fitting procedure; the score does not apportion these contributions. Our control records this procedure's held-out performance after target shuffling. It bounds neither all learning on random targets nor learning on structured targets. Report the probe architecture, feature dimension, and regularisation strength so that the procedure can be reproduced.

Low linear performance does not show that information is absent. Failure of both chosen families still cannot rule out successful recovery by another readout or training procedure. A control task provides context, not an information-theoretic absence test. What you have shown is how these specific probe procedures performed on this split.

Good decoding does not establish causal necessity. If a fixed readout recovered velocity exactly almost surely throughout the target population, then vt=g(zt)v_t=g(z_t) on that population. Zero error on a finite test set does not establish that condition. Neither result shows that the transition model or policy uses the recovered velocity. Controlled representation interventions can test that use, but an intervention inside a model is not automatically an intervention in the world. Chapter 66 returns to this distinction in interpretability and world-model debugging.

There is a further trap in our toy: the encoder is handcrafted, so this experiment characterises a diagnostic procedure, not a trained world model. Its inputs are known by construction, but statistical relationships between position and velocity still depend on the trajectory distribution. A real encoder requires its own splits, controls and interpretation.

Probing is not control evaluation

A probe measures what a separate offline readout can recover from a frozen representation. A controller measures what the live system does with that representation under feedback. These are different quantities, and the second is the subject of Chapter 64. Predicting velocity statistically is not proof that a downstream planner uses a reliable velocity estimate.

Object Permanence and Temporal Consistency

Our probe evaluates individual timesteps. Object permanence concerns an object's continued existence when it is unobserved. Tracking its hidden position through a gap is a richer operational diagnostic: a model may preserve existence or identity while remaining uncertain about location or predicting motion inaccurately. Hidden-trajectory evaluation therefore needs its own target in addition to a persistence criterion.

This is where compounding error and memory meet. The previous chapter, "Perceptual and Generative Evaluation," asked whether a generated frame looks right. Here we ask whether the hidden trajectory is right. For an occluded-object control task, position accuracy can matter to the chosen action even when the generated frames look plausible. These are complementary tests: accurate hidden positions do not guarantee good visual generation, and plausible frames do not certify accurate positions.

We should be precise about what "correct" means during an occlusion. Three distinct properties are usually bundled together:

  • Hidden trajectory accuracy. The model's estimate of the ball's position while hidden is close to the true position.
  • Reappearance accuracy. At the first visible timestep after the gap, the model's belief matches the observation.
  • Identity persistence. The object that reappears is the same object, rather than a newly spawned one that happens to be nearby.

Multiple indistinguishable objects can exchange identities while preserving the visible set of positions. An identity-aware test needs persistent labels or another identity criterion. Our single-object toy does not test competing identity assignments; that does not make identity and trajectory accuracy equivalent. A tracked object can keep its label while its predicted position is wrong.

Observational ambiguity also matters. A noisy visible history may be consistent with several hidden continuations. A point loss can still compare predictors by expected risk over realized trajectories, but does not characterize the forecast distribution. Where the task needs uncertainty, report and evaluate a conditional distribution rather than treating one plausible continuation as the uniquely predictable outcome.

Setting up the occluded-object task

The six paths below are the first six held-out trajectories in index order, and the amber markers identify frames hidden by the band. This figure shows ground-truth paths and visibility geometry, not a predictor's errors or reappearance accuracy.

In[11]:
Code
## Deterministic selection rule: the first six held-out trajectories in index order,
## with no resampling and no random draw. SHOW_IDX is fixed at construction time.
N_SHOW = 6
SHOW_IDX = np.arange(N_SHOW)
show_pos = pos_all[test_sl][
    SHOW_IDX
]  # (N_SHOW, N_STEPS + 1, 2) positions in metres
show_mask = mask_all[test_sl][SHOW_IDX]  # 1.0 visible, 0.0 behind the occluder
occ_lo, occ_hi = 0.5 - OCC_HALF, 0.5 + OCC_HALF
Out[12]:
Visualization
Trajectories in a unit square with occluder band and hidden frames marked.
The first six held-out toy trajectories in the unit square, with the occluder band shaded and hidden frames marked in amber. These are ground-truth paths and visibility markers; the figure does not measure prediction error or reappearance accuracy.

We evaluate three predictors with explicitly stated information budgets. Stating the budget matters here more than in almost any other comparison, because the three predictors differ enormously in what they are allowed to see, and a naive comparison across them is meaningless.

  • Freeze. Holds the last visible observation fixed: p^τ+k=oτ\hat{p}_{\tau+k} = o_\tau. Uses observations only, no motion model. On a free trajectory without a wall bounce during the gap, this drifts at the object's speed, so its error grows roughly linearly in kk and is the natural last-observation, no-motion baseline.
  • Constant velocity. Estimates velocity from the two immediately preceding visible observations and extrapolates: p^τ+k=oτ+kΔt (oτ−oτ−1)/Δt\hat{p}_{\tau+k} = o_\tau + k\Delta t\,(o_\tau-o_{\tau-1})/\Delta t. It ignores subsequent forces and contacts. Observation noise and those omitted dynamics can all affect its error.
  • Privileged constant velocity. Extrapolates the simulator's true position and velocity at τ\tau, but still ignores subsequent forces and contacts. This baseline discloses simulator-state access; it is neither a full-simulator oracle nor an irreducible-error bound.

The metric is the mean Euclidean position error in metres over the hidden window, indexed by the hidden offset kk. Formally, writing Rk\mathcal{R}_k for the set of occluded runs that survive to offset kk and p^τ+k(r)\hat{p}^{(r)}_{\tau+k} for a predictor's estimate on run rr at hidden step τ+k\tau + k, the reported quantity at offset kk is

err(k)=1∣Rk∣∑r∈Rk∥p^τr+k(r)−pτr+k(r)∥2\mathrm{err}(k) = \frac{1}{|\mathcal{R}_k|} \sum_{r \in \mathcal{R}_k} \big\| \hat{p}^{(r)}_{\tau_r+k} - p^{(r)}_{\tau_r+k} \big\|_2

where:

  • kk: the hidden offset in steps of Δt=0.1\Delta t = 0.1 s, running from 11 to Kmax⁡=6K_{\max} = 6.
  • rr: an index over occluded runs in the held-out split.
  • τr\tau_r: the index of the last visible timestep before run rr, so τr+k\tau_r + k is a hidden timestep.
  • p^τr+k(r)\hat{p}^{(r)}_{\tau_r+k}: the predictor's estimated position at that hidden step.
  • pτr+k(r)p^{(r)}_{\tau_r+k}: the simulator's true position at that hidden step.
  • ∣Rk∣|\mathcal{R}_k|: the denominator, the number of occluded runs that are at least kk steps long.

The metric is in metres and is an average over runs at each offset, with the same occluded-run denominator applied to every predictor. The denominator ∣Rk∣|\mathcal{R}_k| shrinks as kk grows, since fewer runs survive to large offsets, and we report it alongside the error so you can judge how thin the tail is. Every predictor is evaluated on the same Rk\mathcal{R}_k, while their different information budgets remain explicit. At a fixed offset, each eligible run contributes once; this is not a pooled average over all hidden frames.

In[13]:
Code
import numpy as np

## Wall bounces during occlusion are already produced by the reflecting-wall
## dynamics in simulate(); this evaluation reuses the main held-out trajectories
## used by the probes and freeze-drift diagnostic. The bounces come from the physics,
## not from a post-hoc perturbation, so no trajectory is modified after the fact.


def occluded_runs(vis_row):
    """Return maximal (start, stop) index intervals where visibility is 0."""
    runs, start = [], None
    for i, value in enumerate(vis_row):
        if value < 0.5 and start is None:
            start = i
        elif value >= 0.5 and start is not None:
            runs.append((start, i))
            start = None
    if start is not None:
        runs.append((start, len(vis_row)))
    return runs


def first_usable_run(vis_row, min_start=2):
    """First hidden run with two immediately preceding visible frames."""
    for start, stop in occluded_runs(vis_row):
        if (
            start >= min_start
            and stop > start
            and np.all(vis_row[start - 2 : start] >= 0.5)
        ):
            return start, stop
    return None


## Regression fixtures: earlier visibility somewhere in the history is not enough.
assert first_usable_run(np.array([1, 0, 1, 0, 0, 1])) is None
assert first_usable_run(np.array([1, 1, 0, 0, 1])) == (2, 4)
assert first_usable_run(np.array([0, 0, 0, 0])) is None
assert first_usable_run(np.array([1, 1, 1, 1])) is None
assert occluded_runs(np.array([1, 1, 0, 0])) == [(2, 4)]
In[14]:
Code
K_MAX = 6
pos_te = pos_all[test_sl]
obs_te = obs_all[test_sl]
vel_te = vel_all[test_sl]
mask_te = mask_all[test_sl]

predictor_names = [
    "freeze",
    "constant velocity",
    "privileged constant velocity",
]
err_sum = {name: np.zeros(K_MAX) for name in predictor_names}
n_at_k = np.zeros(K_MAX)

for i in range(pos_te.shape[0]):
    for start, stop in occluded_runs(mask_te[i]):
        if start < 2 or not np.all(mask_te[i, start - 2 : start] >= 0.5):
            continue
        p_last = obs_te[i, start - 1]
        v_hat = (obs_te[i, start - 1] - obs_te[i, start - 2]) / DT
        s_true_last = pos_te[i, start - 1]
        v_true = vel_te[i, start - 1]
        for k in range(1, min(stop - start, K_MAX) + 1):
            j = start - 1 + k
            if j >= pos_te.shape[1]:
                break
            truth = pos_te[i, j]
            preds = {
                "freeze": p_last,
                "constant velocity": p_last + k * DT * v_hat,
                "privileged constant velocity": s_true_last + k * DT * v_true,
            }
            for name, pred in preds.items():
                err_sum[name][k - 1] += np.linalg.norm(pred - truth)
            n_at_k[k - 1] += 1

mean_err = {
    name: np.divide(
        err_sum[name], n_at_k, out=np.full(K_MAX, np.nan), where=n_at_k > 0
    )
    for name in predictor_names
}
## An empty contributing set must be missing, not a fabricated zero error.
_empty_error = np.divide(
    np.zeros(2), np.zeros(2), out=np.full(2, np.nan), where=np.zeros(2) > 0
)
assert np.isnan(_empty_error).all()
Out[15]:
Console
 hidden k    freeze  const. vel  priv.CV   n runs
        1    0.0283      0.0100   0.0024       23
        2    0.0539      0.0165   0.0050       22
        3    0.0801      0.0242   0.0083       22
        4    0.1066      0.0323   0.0121       22
        5    0.1350      0.0412   0.0172       19
        6    0.1573      0.0482   0.0223       18

Errors are mean Euclidean position error in metres, over held-out runs.
Privileged CV reads the last visible simulator state, not future forces.

The table format is deliberate: the same metric, the same denominator, and the same test trajectories for all three predictors, with the information budget of each written down explicitly. This is what makes the comparison readable. Without the disclosed budgets, a reader would have no way to know whether the difference between two rows reflects a modelling improvement or a difference in what the predictors were allowed to see.

Out[16]:
Visualization
Line chart of position error against hidden horizon for three predictors.
Mean Euclidean position error during hidden windows of driven held-out trajectories. Freeze prediction accumulates error faster than a finite-difference constant-velocity extrapolator. The privileged true-state extrapolator still ignores later actions and contact changes, so its residual is not an irreducible uncertainty bound. The contributing run count changes with horizon.

The shape of each curve is the diagnosis. The freeze predictor grows approximately linearly in kk here because it ignores the ball's velocity. Its error is the displacement from the last noisy observation to the realized hidden position, not the path length traveled. Under approximately straight, constant motion and small observation noise, the leading-order relation is

E[∥p^τ+k−pτ+k∥2]≈k E[∥v∥2] Δt\mathbb{E}\big[\|\hat{p}_{\tau+k} - p_{\tau+k}\|_2\big] \approx k\,\mathbb{E}[\|\mathbf{v}\|_2]\,\Delta t

where:

  • kk: the hidden offset in steps.
  • v\mathbf{v}: the ball's velocity vector.
  • E[∥v∥2]\mathbb{E}[\|\mathbf{v}\|_2]: the mean speed.
  • Δt=0.1\Delta t = 0.1 s: the step size.

That is, the hidden offset times the mean speed times the step size, under approximately constant motion. The constant-velocity estimate improves on freezing here, but its residual can reflect observation noise, later driving actions, and contact changes. The privileged baseline extrapolates the true initial position and velocity rather than executing the full simulator. Its non-zero residual cannot be attributed uniquely to wall bounces, and it is not an irreducible-error floor.

The dashed line uses k Δt vˉk\,\Delta t\,\bar v, with mean speed over all hidden held-out frames. That population differs from the eligible run set at each offset, and driving, contacts, and measurement noise also affect the measured curve. Their similarity here is a descriptive comparison, not an exact matched-population identity or a causal decomposition.

In[17]:
Code
## Leading-order prediction of the freeze error: k * dt * mean speed on hidden frames.
## The mean speed is a deterministic statistic of vel_te and mask_te, with no sampling,
## using the same main held-out trajectories as the hidden-horizon predictors.
hidden_frames = mask_te < 0.5
mean_speed_hidden = float(np.linalg.norm(vel_te[hidden_frames], axis=-1).mean())
drift_pred = np.arange(1, K_MAX + 1) * DT * mean_speed_hidden
freeze_measured = mean_err["freeze"]
Out[18]:
Visualization
Line chart of measured freeze error versus a mean-speed linear prediction.
Freeze error and a constant-speed approximation using mean speed over all hidden held-out frames. The curves are similar here but use different weighting populations: measured error averages eligible runs separately at each offset. Driving, contacts, and observation noise are additional differences, so this comparison is not an exact identity or a decomposition of the error.

The comparison names each predictor's information budget and metric denominator. In this driven task, horizon-dependent error mixes initial-state estimation with omitted future forces and possible contacts. It is not a pure measure of learned memory: these predictors are handcrafted, and the constant-velocity baseline has no learned recurrent state.

Sweeping the occlusion length

Horizon is one axis; occluder geometry is another. We sweep the band half-width on fixed undriven trajectories. The metric is position error at the final hidden step of the first usable gap, not at the first visible reappearance. Wider bands change both gap lengths and which trajectories qualify, so this sweep does not isolate elapsed time from sample selection.

In[19]:
Code
sweep_widths = [0.02, 0.05, 0.10, 0.20, 0.30]
rng_sw = np.random.default_rng(404)
p_sw, v_sw, a_sw, o_sw = simulate(rng_sw, 40, driven=False)

sweep_rows = []
for w in sweep_widths:
    m_w = visibility(p_sw, occ_half=w)
    f_tot = c_tot = t_tot = 0.0
    n_used = 0
    n_censored = 0
    gap_lengths = []
    for i in range(p_sw.shape[0]):
        run = first_usable_run(m_w[i])
        if run is None:
            continue
        start, stop = run
        k = stop - start
        if k < 1:
            continue
        truth = p_sw[i, start - 1 + k]
        p_last = o_sw[i, start - 1]
        v_hat = (o_sw[i, start - 1] - o_sw[i, start - 2]) / DT
        f_tot += np.linalg.norm(p_last - truth)
        c_tot += np.linalg.norm(p_last + k * DT * v_hat - truth)
        t_tot += np.linalg.norm(
            p_sw[i, start - 1] + k * DT * v_sw[i, start - 1] - truth
        )
        n_used += 1
        n_censored += int(stop == m_w.shape[1])
        gap_lengths.append(k)
    denominator = n_used if n_used else np.nan
    sweep_rows.append(
        (
            w,
            float(np.mean(gap_lengths)) if n_used else np.nan,
            f_tot / denominator,
            c_tot / denominator,
            t_tot / denominator,
            n_used,
            n_censored,
        )
    )
Out[20]:
Console
 occ. half-width     mean gap    freeze  const. vel  priv.CV     n  censored
            0.02         1.70    0.0405      0.0177   0.0000    40         0
            0.05         4.38    0.1057      0.0412   0.0000    40         0
            0.10         8.68    0.2105      0.0658   0.0000    40         1
            0.20        17.11    0.4097      0.0985   0.0000    36         8
            0.30        21.57    0.4776      0.1263   0.0000    21        12

Error is measured at the last hidden step of the first gap in each rollout.
All three predictors are evaluated on the same trajectories and the same gaps.

The freeze predictor accumulates larger errors as the band widens; the observation-based constant-velocity predictor remains better but also degrades. The privileged extrapolator is essentially exact on these evaluated gaps. This does not demonstrate a bounce-induced jump. A gap still hidden at the simulation endpoint is right-censored: its recorded length is only the observed portion. Report contributing trajectories and censored counts alongside that mean length.

Out[21]:
Visualization
Final-hidden-step position error against occluder half-width for three predictors.
Position error at the final hidden step as occluder half-width changes on fixed undriven trajectories. The constant-velocity predictor outperforms freezing, while the privileged initial-state extrapolator is essentially exact here. Contributing trajectory subsets and gap lengths change with width; this plot does not demonstrate bounce-induced jumps.

Memory ablations, and what they do and do not show

Recurrent and transformer-based world models carry their hidden state across time in a learned memory. A natural diagnostic is to reset or truncate that memory at the moment the object disappears and see whether performance degrades. If it does, the memory matters for the task. This is an appealing experiment because it is easy to run, but the inference it supports is narrower than it first appears.

A mid-episode reset may be out of distribution if training never used such resets. Context truncation may likewise be unfamiliar, depending on the training protocol. Performance changes measure sensitivity to the stated manipulation, not an isolated memory circuit. Describe the training context and reset schedule before interpreting either ablation.

Four useful temporal properties need different measurements:

  • Sequence coherence is a local property: successive predicted states are mutually consistent.
  • Long-range dependence is about how far back in time the model's prediction depends on.
  • Object identity requires a persistent identity criterion; multi-object gaps make assignment ambiguity particularly visible.
  • Uncertainty should be reported, not suppressed: when several hidden continuations are consistent with the visible data, a good model produces a spread of plausible futures rather than one confident guess.

Finally, note the discipline about oracle baselines and information budgets. It is easy to build a comparison in which one method quietly reads the simulator's latent state while another reads only observations, and then report the resulting gap as a modelling improvement. Every baseline in this section is labelled with what it reads. If you introduce an oracle, introduce it loudly. The reason this discipline is worth repeating is that the temptation to compare against an oracle without disclosing it is strong: the oracle often performs better, and the difference can be presented as evidence of modelling quality rather than of information asymmetry.

Physical Rule and Constraint Tests

The previous section asked whether the hidden state is coherent. This one asks whether the dynamics obey a specified physical law under a named regime. A model can produce a smooth, coherent, plausible trajectory that travels at an unphysical speed through a dissipative medium, or that passes through a wall, or that gains energy at every bounce. Constraint tests are how you catch that. They are the natural complement to the trajectory-accuracy tests above: accuracy asks whether the model is close to the truth on average, and constraint tests ask whether the model respects the mechanisms that produce the truth.

Specify the regime before testing a rule. In this force-free, elastic toy, kinetic energy is conserved. Applied forces can do work; damping and lossy contacts can dissipate energy, so a driven dissipative system needs both contributions in its energy balance. A conservation test without a regime can penalize correct dynamics. Here a lossy contact reduces normal kinetic energy only when the incident normal velocity is nonzero.

We will evaluate three regimes, all with the same dynamics code but different parameters:

Energy behaviour in the three physical regimes used for constraint tests.
RegimeForcingDampingRestitutionExpected energy behaviour over a rollout
Force-free, elasticnone1.01.0EH/E0=1E_H / E_0 = 1 exactly
Force-free, dampednone0.91.0EH/E0=0.92H=0.81HE_H / E_0 = 0.9^{2H} = 0.81^{H} exactly, since energy scales as v2v^2
Free, lossy contactnone1.00.7constant between contacts; normal-component energy is multiplied by 0.490.49 at an impact

The damping factor multiplies both velocity components by 0.90.9 each step, so total kinetic energy is multiplied by 0.810.81. At a wall, restitution multiplies only the normal velocity by −e-e; its energy component is multiplied by e2=0.49e^2=0.49, while tangential energy is unchanged. Total energy therefore does not generally fall by a factor of 0.490.49. These regimes test different mechanisms, but a residual alone cannot identify which mechanism is wrong.

The metric: end-of-rollout energy ratio residual

For a rollout of length H=30H = 30 starting from the true initial state s0s_0, define the predicted energy ratio ρ^=E^H/E^0\hat\rho = \hat{E}_H / \hat{E}_0 and the true ratio ρ=EH/E0\rho = E_H / E_0, where EE is the kinetic energy

Et=12m∥vt∥22E_t = \tfrac{1}{2} m \|\mathbf{v}_t\|_2^2

where m=1m = 1 kg is the mass and vt\mathbf{v}_t is the velocity at step tt. The energy ratio residual is

ERR=∣ρ^−ρ∣\mathrm{ERR} = \big| \hat\rho - \rho \big|

with the predicted and true ratios defined as

ρ^=E^HE^0,ρ=EHE0\hat\rho = \frac{\hat{E}_H}{\hat{E}_0}, \qquad \rho = \frac{E_H}{E_0}

where:

  • HH: the rollout horizon, 3030 steps, so HΔt=3.0H\Delta t = 3.0 s.
  • E^t\hat{E}_t: the kinetic energy of the model's predicted state at step tt.
  • EtE_t: the kinetic energy of the simulator's true state at step tt.
  • ρ^\hat\rho: the model's predicted energy ratio over the rollout.
  • ρ\rho: the true energy ratio produced by the simulator's own dynamics.

These ratios require positive initial energy. The fixtures satisfy that condition; the implementation floors denominators at 10−1210^{-12} joules as a numerical safeguard, which changes the ratio convention near zero. In the elastic case ρ=1\rho=1; with damping and no clipping ρ=0.81H\rho=0.81^H; with lossy contact the ratio depends on each impact. The violation rate uses the declared absolute ratio-error tolerance 0.050.05, not a universal physical threshold.

Why not just check Et+1=EtE_{t+1} = E_t?

Stepwise kinetic-energy equality is the relevant rule only in the force-free, elastic regime here. A correct simulation may have floating-point discrepancies, whereas a wrong predictor may have much larger errors. Comparing with the regime's reference ratio and stating a tolerance makes the evaluated quantity explicit.

We evaluate four predictors in each regime to compare regime dependence and model-class limitations. Simulator truth is reused as an identity reference; evaluating numerical integration drift would require an independent higher-accuracy reference or a time-step convergence study, which this example does not provide.

  • True integrator. Simulator truth reused as the reference prediction. Its zero residual follows by identity, not from an independent numerical-accuracy check.
  • Linear model fitted on force-free data. A one-step linear map s^t+1=Ast+b\hat{s}_{t+1} = A s_t + b estimated by ridge regression on force-free trajectories.
  • Linear model fitted on damped data. The same estimator, fitted on damped trajectories.
  • No-change control. The deliberately wrong negative control: s^t+1=st\hat{s}_{t+1} = s_t. It preserves initial velocity and kinetic energy but freezes position, illustrating that passing an energy check is not sufficient for correct dynamics.
In[22]:
Code
def kinetic_energy(vel, mass=1.0):
    return 0.5 * mass * np.sum(vel**2, axis=-1)


def fit_linear_dynamics(pos, vel):
    """Ridge fit of s_{t+1} ~ [s_t, 1] using all transitions in a regime."""
    S = np.concatenate([pos, vel], axis=-1)
    X = S[:, :-1].reshape(-1, S.shape[-1])
    Y = S[:, 1:].reshape(-1, S.shape[-1])
    return ridge_fit(np.hstack([X, np.ones((X.shape[0], 1))]), Y, alpha=1e-6)


def rollout(W, s0, n_steps):
    """Recursively apply the learned one-step model. s0 has shape (n, 4)."""
    S = np.empty((s0.shape[0], n_steps + 1, s0.shape[1]))
    S[:, 0] = s0
    for t in range(n_steps):
        Xb = np.hstack([S[:, t], np.ones((s0.shape[0], 1))])
        S[:, t + 1] = Xb @ W
    return S


def evaluate(S_pred, truth, tol=0.05):
    E_pred = kinetic_energy(S_pred[..., 2:4])
    E_true = kinetic_energy(truth[..., 2:4])
    ratio_pred = E_pred[:, -1] / np.maximum(E_pred[:, 0], 1e-12)
    ratio_true = E_true[:, -1] / np.maximum(E_true[:, 0], 1e-12)
    resid = float(np.mean(np.abs(ratio_pred - ratio_true)))
    viol = float(np.mean(np.abs(ratio_pred - ratio_true) > tol))
    pos_err = np.linalg.norm(S_pred[..., :2] - truth[..., :2], axis=-1)
    return resid, viol, np.sqrt(np.mean(pos_err**2, axis=0))
In[23]:
Code
H_TEST = 30
rng_e = np.random.default_rng(2024)

regimes = {
    "free / elastic": simulate(
        rng_e, 20, driven=False, damping=1.0, restitution=1.0
    ),
    "free / damped": simulate(
        rng_e, 20, driven=False, damping=0.9, restitution=1.0
    ),
    "free / lossy contact": simulate(
        rng_e, 20, driven=False, damping=1.0, restitution=0.7, v_scale=1.5
    ),
}
fit_sets = {
    "free": simulate(rng_e, 20, driven=False, damping=1.0, restitution=1.0),
    "damped": simulate(rng_e, 20, driven=False, damping=0.9, restitution=1.0),
}

W_free = fit_linear_dynamics(fit_sets["free"][0], fit_sets["free"][1])
W_damp = fit_linear_dynamics(fit_sets["damped"][0], fit_sets["damped"][1])

phys_results = {}
for rname, (pos_r, vel_r, act_r, obs_r) in regimes.items():
    s0 = np.concatenate([pos_r[:, 0], vel_r[:, 0]], axis=-1)
    truth = np.concatenate(
        [pos_r[:, : H_TEST + 1], vel_r[:, : H_TEST + 1]], axis=-1
    )
    models = {
        "true integrator": truth,
        "linear fit (free)": rollout(W_free, s0, H_TEST),
        "linear fit (damped)": rollout(W_damp, s0, H_TEST),
        "no-change control": np.repeat(s0[:, None, :], H_TEST + 1, axis=1),
    }
    phys_results[rname] = {m: evaluate(S, truth) for m, S in models.items()}
Out[24]:
Console
free / elastic
  model                   energy ratio resid  violation rate   RMSE h=1  RMSE h=30
  true integrator                     0.0000           0.000    0.00000    0.00000
  linear fit (free)                   0.9537           1.000    0.00100    0.16624
  linear fit (damped)                 0.9982           1.000    0.00254    0.49824
  no-change control                   0.0000           0.000    0.02536    0.71493

free / damped
  model                   energy ratio resid  violation rate   RMSE h=1  RMSE h=30
  true integrator                     0.0000           0.000    0.00000    0.00000
  linear fit (free)                   0.0728           0.500    0.00368    0.41512
  linear fit (damped)                 0.0000           0.000    0.00000    0.00000
  no-change control                   0.9982           1.000    0.02568    0.24587

free / lossy contact
  model                   energy ratio resid  violation rate   RMSE h=1  RMSE h=30
  true integrator                     0.0000           0.000    0.00000    0.00000
  linear fit (free)                   0.5171           1.000    0.00054    0.25837
  linear fit (damped)                 0.5730           1.000    0.00413    0.34631
  no-change control                   0.4252           0.850    0.04129    0.58632
Out[25]:
Visualization
Grouped bar chart of energy ratio residual across three regimes for four dynamics models.
End-of-rollout energy-ratio error relative to simulator truth across three regimes. The matched damped linear fit performs well on damped trajectories, but the free-trained global linear model fails to preserve elastic dynamics over this horizon. The no-change control passes the elastic energy test despite incorrect positions. Constraint fidelity and trajectory fidelity must be checked separately.

The simulator trajectory is compared with itself and therefore has exactly zero residual; this is a reference identity check, not an independent numerical-accuracy test. The damped fit matches damped data closely. The free-trained global linear map cannot represent the piecewise reflection mechanism and has a large elastic-regime energy error. The no-change control preserves energy in that regime while freezing position, so its zero energy error does not certify useful dynamics. The tolerance of 0.05 is an explicit diagnostic choice, not a universal physical threshold.

Out[26]:
Visualization
Line chart of recursive rollout position RMSE against horizon for four dynamics models.
Position RMSE over recursive rollouts in the elastic regime. Both global linear fits accumulate error, although the free-trained fit performs better at the final horizon. The no-change control illustrates poor trajectory fidelity despite passing the elastic energy test. Initial-state error at horizon zero is zero for every method by construction.

Separating the sources of a residual

A non-zero constraint residual can come from at least five places, and attributing it correctly is most of the interpretive work:

  • Learned dynamics error. The model misrepresents the mechanism.
  • Numerical integration drift. Under-resolved integration can introduce error. Diagnosing it needs an independent numerical reference; our identity comparison does not measure it.
  • Observation error. If the rollout is initialised from noisy observations rather than the true state, the initial energy estimate itself is noisy.
  • Coordinate and unit errors. Inconsistent velocity conversion across initial and final states can distort ratios. A common multiplicative conversion of all velocities cancels in EH/E0E_H/E_0; ratio agreement cannot detect every unit error.
  • A mismatched regime. The model may be perfectly correct for a different regime than the one under test.

The experiment includes regime mismatch and model-class misspecification. A systematic signed residual can arise from either, from calibration bias, or from other errors; its sign does not uniquely diagnose the cause. Compare matched-regime data, held-out mechanisms and an appropriate numerical reference before assigning responsibility.

The next figure compares the free-trained and damped-trained linear fits on the same 20 elastic held-out rollouts, plotting each unit's signed energy-ratio residual instead of its average. The fits also use different realized initial-state draws in their training trajectories, so this comparison does not isolate a change in training regime alone.

In[27]:
Code
## Signed energy ratio residuals on the same 20 elastic rollouts, one value per unit.
def signed_ratio_residual(S_pred, truth):
    E_pred = kinetic_energy(S_pred[..., 2:4])
    E_true = kinetic_energy(truth[..., 2:4])
    rho_pred = E_pred[:, -1] / np.maximum(E_pred[:, 0], 1e-12)
    rho_true = E_true[:, -1] / np.maximum(E_true[:, 0], 1e-12)
    return rho_pred - rho_true


pos_el, vel_el, _, _ = regimes["free / elastic"]
s0_el = np.concatenate([pos_el[:, 0], vel_el[:, 0]], axis=-1)
truth_el = np.concatenate(
    [pos_el[:, : H_TEST + 1], vel_el[:, : H_TEST + 1]], axis=-1
)

residual_matched = signed_ratio_residual(
    rollout(W_free, s0_el, H_TEST), truth_el
)
residual_mismatched = signed_ratio_residual(
    rollout(W_damp, s0_el, H_TEST), truth_el
)

## Unique trajectory IDs keep all 20 units countable even when values coincide.
trajectory_ids = np.arange(1, residual_matched.size + 1)
Out[28]:
Visualization
Signed energy ratio residuals for free-trained and damped-trained linear fits, with each of the twenty test trajectory IDs on its own row.
Signed energy-ratio error for 20 elastic test trajectories. Both global linear fits underestimate final relative energy, including the free-trained fit. The unit-level view reveals recurring error but does not identify regime mismatch as its unique cause.

Constraint projection can improve a chosen residual while worsening trajectory prediction. Renormalizing a nonzero predicted velocity to its initial magnitude enforces energy conservation in the elastic regime, but is wrong for a dissipative regime's energy target. Clamping position to the box does not enforce energy conservation. A projection must target the actual constraint and satisfy its feasibility conditions; it does not universally drive every residual to zero. Report projected and unprojected errors together.

Finally, note that constraint testing is not a substitute for the state probes of the previous section. A model can respect every conservation law while tracking the wrong object, and it can track the right object while violating energy conservation. These are independent axes of evaluation. The full picture requires both.

Interventions, Counterfactuals, and Causal Fidelity

The occluded-ball question from the previous sections was observational: given what was seen, what happens next? The questions in this section are different. They are about what would happen under a change. The difference between "given what was seen, what happens next" and "what would have happened if we had done something different" is precisely the difference between prediction and causal inference, and no amount of observational data can answer the second kind of question without additional assumptions.

Three targets must be kept apart, because they are estimated differently, identify different quantities, and are routinely confused:

Three targets, three meanings

Observational prediction. In a state-observed setup, predict st+1s_{t+1} from (st,at)(s_t, a_t) on the distribution the data were collected under. Other observationally trained sequence models can use observation histories or latent representations as inputs and targets, rather than true simulator states.

Population intervention. Compare E[sH∣do(aτ=a⋆)]\mathbb{E}[s_H \mid \mathrm{do}(a_\tau=a^\star)] and E[sH∣do(aτ=a0)]\mathbb{E}[s_H \mid \mathrm{do}(a_\tau=a_0)] over a specified population. Identification needs a suitable design and assumptions: for example randomized action assignment, or an observational adjustment design with consistency, positivity, and conditional exchangeability. Absence of hidden confounding is not necessary for every possible identification strategy.

Unit-level counterfactual. For one specific unit, estimate what sHs_H would have been under a different action. In the structural causal model (SCM) used here, hold that unit's initial state and background disturbance fixed across the hypothetical branches. That is a specified cross-world coupling, not a universal ranking of how difficult the three targets are to identify.

Observational fit does not identify any causal quantity without further assumptions. A model can fit the observational distribution to machine precision and still be wrong about the effect of an intervention, and the mechanism by which this happens is confounding.

Building a confounded toy

The earlier simulator sampled actions independently of its state and transition, without a shared action/outcome disturbance. We now switch to a separate one-step structural causal model in which an unobserved disturbance utu_t influences both the action and the velocity change:

ut∼N(0,1)at=−κ (xt−0.5)+utvt+1=vt+(at+c ut) Δt\begin{aligned} u_t &\sim \mathcal{N}(0, 1) \\ a_t &= -\kappa\,(x_t - 0.5) + u_t \\ v_{t+1} &= v_t + \big(a_t + c\,u_t\big)\,\Delta t \end{aligned}

where:

  • utu_t: the unobserved exogenous acceleration disturbance, with numerical values in m/s² drawn from a standard normal.
  • ata_t: the acceleration action in m/s² chosen by the state-feedback policy with gain κ=2\kappa=2 s−2^{-2}; the position reference 0.50.5 is in metres.
  • xtx_t: the position at step tt.
  • vt+1v_{t+1}: the next velocity, driven by the action plus the disturbance coupling cc.
  • Δt=0.1\Delta t = 0.1 s: the step size.

The action is chosen by a state-feedback policy, but the policy's exploration noise is the same utu_t that enters the dynamics with coupling cc. An engineer observing only (xt,at,vt,vt+1)(x_t, a_t, v_t, v_{t+1}) never sees utu_t, so the observational data are confounded by construction. This toy illustrates one possible shared-signal mechanism: the observational correlation between action and outcome mixes the direct effect of the action with the effect of the shared signal. It does not establish how often that mechanism occurs in deployed systems.

Now ask: what is the causal effect of the action on the one-step velocity change? Holding utu_t fixed and differentiating, the structural derivative is

∂ Δvt∂at=Δt=0.1 m/s per m/s2\frac{\partial\, \Delta v_t}{\partial a_t} = \Delta t = 0.1\ \text{m/s per m/s}^2

where Δvt=vt+1−vt\Delta v_t = v_{t+1} - v_t is the one-step velocity change, ata_t is the action, and Δt=0.1\Delta t = 0.1 s is the step size. The derivative is Δt\Delta t because the structural equation is linear in the action with slope Δt\Delta t. This is the quantity we want to recover, and it is a structural quantity: it describes how the next velocity would change if we intervened on the action while holding everything else fixed.

That is the ground truth we want to recover. Now see what an observational regression gives. Rearranging the policy equation gives ut=at+κ(xt−0.5)u_t = a_t + \kappa (x_t - 0.5). Substituting this into the dynamics,

ΔvtΔt=at+c ut=(1+c) at+cκ (xt−0.5)\frac{\Delta v_t}{\Delta t} = a_t + c\,u_t = (1 + c)\,a_t + c\kappa\,(x_t - 0.5)

where ut=at+κ(xt−0.5)u_t = a_t + \kappa (x_t - 0.5) has been substituted. The coefficient on ata_t is 1+c1 + c rather than 11, so the observational relationship overstates the structural action effect. The empirical analogue is a regression of the observed velocity change on the observed action and position.

In the noiseless structural relationship, the observational coefficient is (1+c)Δt=0.2(1+c)\Delta t=0.2, twice the intervention slope 0.10.1. The measured regression below includes observation noise and fits nearly, not exactly, perfectly. A predictive coefficient need not be an intervention effect.

In[29]:
Code
KAPPA = 2.0  # policy gain
C_U = 1.0  # strength of the unobserved disturbance coupling
DT_C = 0.1  # s


def sample_units(rng, n):
    """One-step structural units with a policy that depends on the disturbance."""
    x0 = rng.uniform(0.2, 0.8, n)
    u = rng.normal(0.0, 1.0, n)
    a = -KAPPA * (x0 - 0.5) + u  # action depends on the unobserved u
    dv = (a + C_U * u) * DT_C  # true one-step velocity change
    return x0, u, a, dv
In[30]:
Code
rng_c = np.random.default_rng(7)
x_all, u_all, a_all, dv_all = sample_units(rng_c, 4000)
dv_meas = dv_all + rng_c.normal(
    0.0, 0.005, dv_all.size
)  # measurement standard deviation 0.005 m/s, not a relative percentage

fit_slice = slice(0, 2000)
held_slice = slice(2000, None)

## Observational regression of dv on (x - 0.5, a). No interventions were used.
X_fit = np.hstack(
    [
        np.ones((2000, 1)),
        (x_all[fit_slice] - 0.5)[:, None],
        a_all[fit_slice][:, None],
    ]
)
beta_obs = ridge_fit(X_fit, dv_meas[fit_slice][:, None], alpha=1e-8)
effect_obs = float(beta_obs[2, 0])

X_held = np.hstack(
    [
        np.ones((2000, 1)),
        (x_all[held_slice] - 0.5)[:, None],
        a_all[held_slice][:, None],
    ]
)
pred_held = X_held @ beta_obs
ss_res = float(np.sum((dv_meas[held_slice][:, None] - pred_held) ** 2))
ss_tot = float(np.sum((dv_meas[held_slice] - dv_meas[held_slice].mean()) ** 2))
r2_obs = 1.0 - ss_res / ss_tot

support_lo, support_hi = float(a_all.min()), float(a_all.max())

The held-out observational fit is R2≈0.9993R^2\approx0.9993, while its fitted action coefficient is about 0.199940.19994 instead of the structural slope 0.10.1. Fit alone does not reveal the confounding in this example. The nearly perfect noisy fit should not be described as machine precision.

The paired counterfactual experiment

Now we run a paired simulator intervention on the held-out units. For each unit:

  1. Clone the unit's initial state x0x_0 and its exogenous disturbance uu.
  2. Run it once with its natural action aa.
  3. Run it again with the same state and the same uu, but with the action shifted by a fixed δ\delta.
  4. Subtract the two outcomes.

Because uu is shared, the difference isolates the direct effect of the action. Substituting the structural equation Δv(a)=(a+c u) Δt\Delta v(a) = (a + c\,u)\,\Delta t into the two branches,

Δvt(a+δ)−Δvt(a)=((a+δ)+c u) Δt−(a+c u) Δt(structural equation at each action)=(a+δ+c u−a−c u) Δt(expand)=δ Δt(the shared c u cancels).\begin{aligned} \Delta v_t(a + \delta) - \Delta v_t(a) &= \big((a + \delta) + c\,u\big)\,\Delta t - \big(a + c\,u\big)\,\Delta t && \text{(structural equation at each action)} \\ &= \big(a + \delta + c\,u - a - c\,u\big)\,\Delta t && \text{(expand)} \\ &= \delta\,\Delta t && \text{(the shared } c\,u \text{ cancels)} . \end{aligned}

For the noiseless structural outcomes, dividing by nonzero δ\delta recovers Δt=0.1\Delta t=0.1 exactly. The code then adds independent measurement noise to both branches and averages across units, so its estimate has sampling error. Shared structural uu is the simulator's explicit counterfactual coupling; passive video does not supply that coupling by itself.

In[31]:
Code
DELTA = 0.5

x_h = x_all[held_slice]
u_h = u_all[held_slice]
a_h = a_all[held_slice]

dv_base = (a_h + C_U * u_h) * DT_C
dv_shift = ((a_h + DELTA) + C_U * u_h) * DT_C

noise_base = rng_c.normal(0.0, 0.02, a_h.size)
noise_shift = rng_c.normal(0.0, 0.02, a_h.size)
effect_paired = float(
    np.mean((dv_shift + noise_shift - dv_base - noise_base) / DELTA)
)
effect_true = float(DT_C)

effect_rows = [
    ("observational regression", effect_obs),
    ("paired counterfactual", effect_paired),
    ("ground truth (simulator)", effect_true),
]
Out[32]:
Console
estimator                      effect of a on dv    relative error
observational regression                  0.1999            99.9%
paired counterfactual                     0.0996             0.4%
ground truth (simulator)                  0.1000             0.0%

observational fit on held-out natural data: R^2 = 0.9993
observed action extrema: [-3.703, 3.692] m/s^2
intervention magnitude delta = 0.50; fraction above observed maximum = 0.1%
no intervened data was used to fit the observational model.

The observational slope is about 0.199940.19994, whereas the paired noisy estimate is about 0.099570.09957 versus structural truth 0.10.1. The latter has roughly 0.4%0.4\% relative error in this run. Pairing is sufficient for this demonstration, not necessary for every population effect: assigning independently sampled units to fixed actions by randomization can identify the difference of population mean outcomes when the intervention is well-defined and observed outcomes are consistent with their assigned actions. Here the two groups have the same disturbance distribution, so their mean structural contrast divided by the action difference is Δt\Delta t. This experiment evaluates an estimator using the simulator's equations; it does not benchmark a learned model's counterfactual predictions.

Out[33]:
Visualization
Bar chart comparing observational, paired counterfactual, and ground truth causal effect estimates.
One-step action-effect estimates. Observational regression has held-out R-squared about 0.9993 but gives a slope near 0.2 instead of 0.1. Sharing the structural disturbance cancels it in noiseless branch differences; independent measurement noise makes the paired average approximate, not exact. This paired design is sufficient here, not necessary for all population-effect identification.

What else to report alongside a causal claim

A single effect estimate is not enough to make a causal claim legible. Report at least the following. Each item corresponds to a distinct way in which a reader might reasonably disagree with the claim, and each is a way of making the claim falsifiable rather than rhetorical.

  • The exact target. Average treatment effect, conditional effect at a specific state, or a unit-level counterfactual? These are different numbers. State which one, and state the unit of the estimate.
  • Intervention support and overlap. Distinguish population support from finite observed extrema and from practical data coverage. Here Gaussian disturbances give actions full real-line population support, but about 0.1%0.1\% of shifted held-out actions exceed the finite sample maximum. Being within the sample range does not establish dense coverage or validate a causal estimate.
  • Which interventions were observed during training. Here, none: the model was fitted purely observationally, and every intervention was held out. If a model was trained on interventions, say which, and reserve a distinct held-out intervention family for evaluation.
  • The feasibility of the intervention. Some interventions are not physically implementable in the source system. A simulated intervention on an unfeasible variable is a modelling exercise, not evidence about the world.
  • Confounding and observability. Write down which variables are observed, which are latent, and which one is assumed to be the exogenous driver. Our pairing stipulates a shared structural uu across the two branches. Passive observational data alone do not generally identify that cross-world coupling, even if a dataset measures a relevant disturbance.
  • Shared-noise coupling. A unit-level counterfactual needs an explicit joint model across branches. Changing that coupling can change paired differences and estimator variance. With fixed branch marginals, the population difference of expected outcomes is unchanged; do not confuse a unit counterfactual with an average effect.
  • Simulator-to-world limitations. Every result in this section lives inside a simulator whose structural equations we wrote. The same protocol applied to real video would require the structural model to be justified externally, and that justification is usually where the argument lives.

Invariance across environments is diagnostic evidence, not an identification argument by itself. A stable coefficient can coexist with confounding, including an unchanged confounding mechanism across the sampled environments. State which environments changed and which identification assumptions the design supports; coefficient stability alone does not establish causal fidelity.

Action conditioning does not guarantee causal validity. Our observational regression takes action as an input yet estimates the wrong intervention slope. A causal method must represent the specified intervention somehow, but it need not be an action-conditioned prediction network; identification can use a different suitable experimental or structural design.

Counterfactual prompting is a method, not a guarantee

Prompting a learned model to generate a counterfactual specifies a desired output, not an identification argument. Validate the target effect using an appropriate intervention design and state the structural assumptions. A paired simulator experiment is one option, not a universal requirement.

Limitations and Impact

Probing, occlusion sweeps, constraint tests, and intervention designs turn broad claims into narrower questions: which targets a readout recovers, how a predictor behaves while observations are hidden, which regime-specific constraints it respects, and which action effect a design identifies. The examples illustrate these procedures; they do not establish that the same measurements transfer unchanged to a trained world model.

The limitations are equally important, and they cluster into three themes.

The first three families use the reflecting point mass under named regimes; the causal example uses a separate one-step structural model. Their mechanisms, noise scales, and sampling policies are specified by construction, so the numerical values are illustrative rather than benchmarking evidence. Different fixtures can change the probe scores and residuals. The examples make the procedures explicit; they do not establish performance on a larger system.

Each diagnostic answers a specified question. A probe report measures recovery by a named procedure on a named split, not unrestricted knowledge or downstream use. The occluded-object tables compare three handcrafted predictors, not learned recurrent memory. The energy table checks global linear fits on both matched and mismatched regimes, alongside an identity reference and frozen-state control; it does not predict how a neural dynamics model would perform. Interpret each number within that scope.

The paired example cancels the shared disturbance exactly in its structural outcomes, while its noisy sample average estimates the slope approximately. Identification relies on the stated structure and design. Simulator results do not validate that structure for a real system; real interventions can, however, provide evidence that distinguishes some candidate mechanisms. Separate assumptions that a design checks from those it leaves untested.

Constraint satisfaction alone is insufficient. A projection enforcing the correct energy target can reduce that residual while worsening positions, and the frozen-state control passes the elastic energy test despite large position error. Our reused simulator truth is an identity reference, not an integration-accuracy floor.

Choose diagnostics for the task rather than treating the families as mandatory rungs. Probe accessibility with declared capacities, controls, and splits. Evaluate hidden windows with contributing-run and censoring counts. Test constraints under named regimes, alongside trajectory error and suitable reference baselines. For a causal target, state the identification design, overlap, feasibility, and structural assumptions before the estimate.

These offline diagnostics do not by themselves establish whether a planner using the model will make better decisions, rank policies correctly, or exploit an inaccurate region of the model. Chapter 64 evaluates executed returns alongside planner compute, policy ranking, and transfer; those decision diagnostics need not demonstrate a failure in every experiment. Benchmark taxonomy and statistical power are the subject of Chapter 65, and the mechanistic question of which internal components carry the state is taken up in Chapter 66. Failure modes and model exploitation are the opening subject of Chapter 67. Exploitation can make model-predicted returns optimistic, but a model's prediction gap can also be pessimistic.

Summary

This chapter developed four evaluation families for asking what a world model represents, and each was anchored to a decision about information budgets and metric denominators.

  • State probing measures accessibility to a specified readout on a specified distribution. Different probe capacities and controls can expose fitting failures, but failed decoding does not prove absent information. Trajectory-level splits and training-only preprocessing prevent particular leakage routes. Good decoding does not establish causal necessity or downstream use.

  • Object permanence and temporal consistency concern hidden trajectories, reappearance and identity, which need separate tests. Horizon and occluder-width sweeps require explicit contributing-sample counts. Privileged baselines must disclose their information; incomplete extrapolators are not automatically error bounds. Memory ablations test sensitivity to a manipulation, not isolated memory circuits.

  • Physical tests require a named regime, units and tolerance. Energy-ratio agreement can coexist with wrong trajectories, as the frozen-state control shows. Regime mismatch and model-class misspecification can both cause residuals; the sign alone does not identify the cause. An identity comparison with simulator truth is not an integration-accuracy certification.

  • Interventions, counterfactuals, and causal fidelity separate observational prediction, population interventions, and unit counterfactuals. Here a near-perfect observational fit still gives roughly twice the structural action slope. Shared disturbance cancels in noiseless paired differences; the independently noisy paired average estimates the effect approximately. Report the target, identification design, practical overlap, feasibility, training interventions, structural coupling, and simulator-to-world limits.

The thread connecting all four is the same: state what the model received, state what the baseline received, name the quantity the metric measures, and separate offline prediction quality from closed-loop decision usefulness. That last separation is where the next chapter begins.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about state, physics, causality, and memory evaluation.

State, Physics, Causality, and Memory Evaluation

Question 1 of 70 of 7 completed
In this chapter, which of the four evaluation questions is the only one that is genuinely causal?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026statephysics, author = {Michael Brenndoerfer}, title = {State, Physics, Causality, and Memory Evaluation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). State, Physics, Causality, and Memory Evaluation. Retrieved from https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models
MLAAcademic
Michael Brenndoerfer. "State, Physics, Causality, and Memory Evaluation." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models>.
CHICAGOAcademic
Michael Brenndoerfer. "State, Physics, Causality, and Memory Evaluation." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models.
HARVARDAcademic
Michael Brenndoerfer (2026) 'State, Physics, Causality, and Memory Evaluation'. Available at: https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). State, Physics, Causality, and Memory Evaluation. https://mbrenndoerfer.com/writing/state-physics-causality-memory-evaluation-world-models

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.