Part of World Models Handbook
Frontier questions and research roadmaps for world models: generality contracts, causal structure, persistent memory, and decision-relevant evaluation.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Frontier Questions and Research Roadmaps
Consider a hypothetical demonstration: someone hands you a model and a folder of results. The model takes a short history of camera frames and a sequence of joint commands, and predicts the next frame. It was trained on tens of thousands of hours of video and robot interaction data, it renders hand deceleration, cloth, and liquid plausibly, and its average pixel error on held-out clips is the best anyone has recorded. The folder shows the model's outputs to a viewer, and the viewer cannot reliably tell them from the real footage. Everyone in the room agrees the model is excellent.
Then someone asks a decision-relevant question for the robot: if I move my gripper 3 cm to the left, does the cup end up in it?
The reported evaluation does not establish an answer. Average error under the tested action distribution does not establish the consequence of this particular intervention, especially if it lies outside that distribution. The model might answer correctly, but the supplied evidence does not verify that capability. This distinction between a capability and evidence for it is the subject of this final chapter.
Average pixel error measures a chosen image-space discrepancy, not necessarily perceptual realism or distributional fidelity. It can detect physical errors visible in evaluated frames. But averaging can dilute errors in rare, decision-relevant situations. It supplies no distribution-free accuracy guarantee for actions outside the evaluation population. An action-effect test addresses that gap by measuring outcomes under the queried actions rather than inferring them from an aggregate visual score.
Throughout this handbook we separated four things that are easy to conflate: a model's predictions, the representation those predictions are computed in, the decisions an agent makes using them, and the evidence that any of it works on a new system. Part I introduced the world-model idea and the qualification test. Part II gave us state, belief, identification, and control. Part III, Part IV, Part V, Part VI, Part VII, Part VIII, and Part IX built up representations, architectures, learned dynamics, and foundation-scale systems. Part X and Part XI examined applications and how to evaluate them. Part XII, this part, has been about reliability: failure modes, calibration and safe control, security and governance, systems, and in Sim-to-Real Deployment and Continual Correction the machinery for moving a model from a simulator into a physical system and keeping it current.
This chapter is the closing one. It does not introduce a new architecture or a new benchmark. It converts everything before it into frontier questions you can pursue, plus three small executable experiments that make the central distinctions concrete. There is no next chapter; the roadmap here is meant to be handed to you. The experiments are deliberately small enough to run on a laptop, and that is the point: the distinctions they expose are conceptual, so they do not need scale to become visible. What they need is care in construction, and that care is what the rest of the chapter is about.
The roadmap turns broad aspirations into specified questions, measurements and comparisons. A precise protocol does not guarantee an easy answer: representation learning, optimization, identification and implementation can remain difficult. It does make clear which claim a result would address.
Three questions organize the chapter:
- What would it mean for a world model to be general-purpose rather than embodiment-specific?
- What do causal, compositional, and persistent structure require that passive predictive training does not provide?
- What must evaluation measure if prediction quality is not the same as decision usefulness?
Each question is hard on its own, and each is easy to answer vacuously. "The model generalizes across embodiments" is vacuous unless you name the source and target systems, the held-out population, the adaptation budget, and the criterion. "The model has causal structure" is vacuous unless you name the intervention and the assumption that makes the conclusion valid. "The model has memory" is vacuous until you say what information is retained, for how long, how it is refreshed, and what happens when it goes stale. Precision here is not pedantry; it is the difference between a research program and a slogan.
The three questions share a practical requirement: specify what evidence would support the claim. Generality needs meaningful query interfaces and competence on a declared population. Evidence for learned structure needs suitable experiments, representations and identification assumptions. Evaluation needs a population, a protocol and decision-relevant measurements; weighting is one part. Model capacity and training data can affect all three, but do not replace these specifications.
What Generality Would Actually Require
For this chapter, we separate two useful senses of "general"; they are an organizing convention rather than an exhaustive taxonomy.
The first sense is distributional generality: demonstrated performance over a specified range of inputs from a stated family. A broad video corpus or many manipulation trajectories can support training for that goal, but dataset size alone does not establish it. Report performance on held-out populations and any subpopulations for which the broad average hides poor results.
The second sense is functional generality: competence at answering the same kind of question across declared domains, with shared query and output semantics. A shared model may accept a new robot through an explicit embodiment map without retraining its shared dynamics. Whether that works depends on the target range, available embodiment information and evaluation results, not just a common output format.
A system can have the first without the second. A generator with no action-query interface does not directly provide the action-conditioned queries used here, even if its videos score well. A generator that does expose such an interface needs a separate assessment of its action semantics and transfer performance. Neither visual plausibility nor an interface alone establishes a universal world model; any demonstrated generality remains relative to the systems, tasks and budgets actually evaluated.
Broader data can improve both coverage and learned action semantics; interface redesign can change which data is usable. Neither intervention guarantees either form of generality. Evidence for one does not automatically establish the other, although a designed transfer experiment can support both. Bounded cross-embodiment world-model results already exist: He et al. (2026 revision) use particle states and displacement actions, map native joint commands through forward kinematics, and deploy the same model for manipulation with distinct robotic hands. This is evidence within their evaluated manipulation setting, not arbitrary-system universality.
A world model is a learned (or partially specified) transition model that, given a representation of the environment's state and a candidate action, predicts the consequences of that action in a form an agent can use to evaluate and select actions. The observation, the state, the belief, the action, the transition, the reward, and the policy are distinct objects (Part I, Chapter 2; Part II, Chapter 1). A generative model of observations is not automatically a world model in this sense, because it may offer no action-conditioned transition to query.
For the planning interface used here, an agent needs to query candidate actions from an initialized state or belief. A predictor without any action-query mechanism does not directly supply that interface. Initialization might use an observation encoder, measured state, a supplied initial state or another declared procedure; a learned encoder is not mandatory. Action conditioning makes alternative-action prediction possible, but does not by itself identify a causal intervention or an individual counterfactual. Those interpretations need assumptions and evidence. Reward and policy heads are optional components, not substitutes for those conditions.
Four Capabilities People Call "A World Model"
The opening scenario conflates four capabilities that can share a name. Separate them when interpreting a result: evidence for one need not establish another.
- Predictive video generation. In this setting, produce future observations conditioned on past observations. Fidelity, temporal coherence, and physical plausibility are possible evaluation criteria (Part IX, Chapter 3). Video generation also includes unconditional and text-conditioned tasks; none inherently requires an action interface. Ho et al., "Video Diffusion Models" study these distinct settings.
- Action-conditioned dynamics prediction. Predict the next state given the current state and an action. This supplies an interface for an alternative-action query: "what if I applied this torque instead?" Interpreting its prediction causally also requires appropriate coverage and identification assumptions; an action input alone does not supply them (Part VI, Chapters 2 and 3).
- A sufficient, actionable state representation. A representation that retains the information needed for the specified decisions. In a partially observed problem, the reference is the available observation–action history or its belief, not necessarily one observation. Any approximate sufficiency claim needs a task, policy class and tolerance. Part III, Chapter 2 developed the criteria; Part VIII, Chapter 5 showed the decision-centric version in practice.
- Closed-loop competence. The whole system, including planner or policy, achieves the objective in the real environment under noise, delay, and perturbation. This is the only capability that directly corresponds to "the robot picks up the cup."
These sit in a dependency chain, but not a strict one. Action conditioning without a good state can fail. A good state without calibrated dynamics can still support a robust policy if the planner is conservative. Closed-loop competence can sometimes be achieved without any of the others, using a reactive policy that never plans. Part IX, Chapter 6 examined vision-language-action models that take exactly that route.
The practical consequence: when someone claims a new system is "a world model," ask which of the four capabilities were measured, on what population, and with what action interface.
A compact way to keep the four capabilities separate is to maintain an evidence checklist. The following matrix is an illustrative prioritization, not a measured score or a logical implication. Its values mean primary (1), supplementary (0.5), or not selected (0) in this example. Calibration alone establishes neither state sufficiency nor closed-loop competence. Test sufficiency with task-relevant information and control comparisons, and test competence by executing the system under the stated conditions.

If only generative fidelity was measured, that evidence alone does not establish control competence. Keep the four capabilities as separate columns in your notes, and record which measurements support each. This lets you assess a paper's particular claims without assuming that evidence for one capability proves another.
The Transfer Contract
For cross-embodiment claims, a transfer contract with five clauses makes the tested shift and available information explicit:
- Source and target systems. Name them. A source of "many robot arms, mostly parallel-jaw, tabletop, 6-DoF" and a target of "a 7-DoF arm with a soft gripper and a mobile base" is a specific, falsifiable statement. "Robots" is not.
- Observation and action spaces. What are the dimensions, units, frames, and semantics before and after transfer? Does the target have the same action units (joint position deltas? end-effector velocities? torques?) or does the map require a learned or analytic adapter?
- Held-out population. Which tasks, objects, embodiments and scenes were excluded from training, including pretraining? Distinguish held-out instances from unseen categories, and document what the corpus audit can and cannot verify (Part VI, Chapter 4).
- Training information and adaptation budget. Zero-shot, few-shot with demonstrations, or fine-tuned with gradient steps and environment interactions. An adaptation report without a budget does not fully specify the transfer comparison.
- Evaluation criterion. Task success rate? Return? Sample efficiency relative to a specialist baseline? Robustness to perturbation? Each answers a different question (Part XI, Chapters 3 and 4).
Population, adaptation budget and criterion jointly define the transfer question. A broad held-out population with little adaptation tests a different claim from specialization on a narrow population with a generous budget. These settings cannot be ordered by difficulty from breadth and budget alone: the target tasks, available information and success criteria may differ. Two papers can both report "transfer" while answering different questions. State these conditions before comparing their results.
The Open X-Embodiment effort provides a multi-robot dataset and RT-X policy models. Its reported positive transfer concerns policies trained with experience from multiple platforms under the paper's evaluation protocol. That supports a defined cross-robot policy-transfer result, not a claim that an action-conditioned world-dynamics model transfers universally. A dynamics-transfer study needs its own prediction and control tests even if it uses the same dataset.
Contrast the well-specified version with the aspirational version: "a single model that can be deployed on any embodied agent." That claim has no held-out population, no adaptation budget, and no criterion, so it does not define a reproducible empirical test or a clear refutation criterion. It is a research direction, and it should be labeled as one. The difference between a research direction and a research result is not rhetorical. A direction tells you where to spend effort; a result tells you what is currently true. Confusing the two can lead teams to build on unverified capability claims and reviewers to accept claims without explicit evaluation or refutation criteria.
What a General-Purpose Model Would Actually Have to Predict
For the cross-embodiment planning system proposed here, what interfaces would make its claims testable?
Consider a system with observation , latent or environment state , belief over (Part II, Chapter 2), action , transition law , reward , and policy . One useful system-level interface contract includes:
- An encoder from heterogeneous observations to a state or belief. It must handle sensor geometry that differs between platforms: different camera intrinsics and extrinsics, different numbers of views, different proprioceptive channels.
- A transition law conditioned on actions with explicit embodiment-specific meanings and units. A canonical frame with adapters is one option; embodiment-conditioned models or separate native interfaces are others. Joint torques and base velocities require a declared interpretation rather than an assumed correspondence. Representation, training coverage and scale can all affect transfer.
- A constraint interface. Joint limits, self-collision, contact, and actuator saturation differ per embodiment. A separate checker can supply constraints instead of embedding them in the learned dynamics. Optimistic impossible predictions can mislead a planner if its objective selects those regions; exploitation is a risk, not an inevitable outcome (Part XII, Chapter 1).
- A notion of what is decision-relevant. Different tasks on identical dynamics can require different task-relevant compressions. A full Markov state, or a sufficiently rich shared representation, can still be sufficient for both (Part III, Chapter 2).
Sufficiency must be assessed relative to the declared family of tasks. A model that serves pick-and-place tasks need not serve locomotion tasks on the same robot: a compression optimized for one task may discard information needed by another. A task change can therefore require a richer or different compression, but it does not necessarily do so. The same full Markov state can support several rewards or goals; training and representation choices determine how much of that information a learned model retains.
Two intermediate designs are worth testing under a declared transfer contract:
- Broad pretraining plus adapters. Train one large dynamics model on heterogeneous data, then attach small per-embodiment modules (action encoders, constraint heads, residual correctors) that are fit with a small budget. This keeps the expensive shared part and localizes the embodiment-specific part. The research question is how much of the shared part transfers across embodiments.
- Specialist models with classical identification. Fit each platform with a structured model (Part II, Chapter 3). A justified model structure can reduce the information needed to estimate its parameters, but any data-efficiency claim needs a named comparator, matched target accuracy and measured data budgets. Include this specialist as an informative target-specific baseline rather than assuming which approach will win.
The adapter approach tests how much of a shared model remains useful after embodiment-specific adaptation. Compare adapters with the unadapted model, and report what changes under the declared budget. The specialist supplies a separate target-specific comparison, not a universal lower bound. Failing to beat it does not show that pretraining supplied no transfer: a pretrained model can improve over the same architecture trained from scratch while still losing to the specialist. Estimate that transfer benefit with matched initialization comparisons or target-data learning curves, alongside the specialist result.
Negative Transfer and Interface Semantics
Transfer is not free. Negative transfer occurs when training on the source distribution makes the target worse than training on the target alone. It shows up in practice for a few identifiable reasons:
- Action-space mismatch. Ambiguous or lossy action mappings can bias action-effect prediction on affected embodiments. Aggregate visual metrics can hide that bias; a compromise does not necessarily bias every embodiment.
- Conflicting dynamics. Without information identifying the domain, identical inputs with opposite action effects are ambiguous. A squared-loss point predictor averages the conditional outcomes, which may be inappropriate for either domain. Context conditioning or a distributional predictor can represent the ambiguity differently.
- Capacity or objective conflict. With a restricted model and training budget, fitting other domains can conflict with target performance. Shared structure can also benefit the target; capacity is not necessarily partitioned into a fixed share per domain. Experiment 1 studies a different issue, compression of nuisance versus task-relevant coordinates, not cross-domain negative transfer.
- Calibration collapse. A model that was well calibrated on the source can be badly miscalibrated on the target. Prediction accuracy and calibration can both affect safety; their relevance depends on the declared safety objective and how the controller uses predictions and uncertainty (Part XII, Chapter 2).
The countermeasure is the same as the diagnosis: make the action interface semantically explicit, evaluate action-effect prediction per embodiment rather than in aggregate, and always report the specialist baseline on the target population. Each of the four causes above suggests a specific diagnostic:
- Action-space mismatch: compare predicted and realized action effects per embodiment.
- Conflicting dynamics: inspect residuals by domain and action. Bimodality can be informative when components are sufficiently separated, but its absence does not exclude conflicting dynamics.
- Capacity dilution: test whether a performance gap shrinks under controlled enlargement. This is one diagnostic, not identification of capacity dilution; optimization or representation changes can also shrink the gap.
- Calibration collapse: inspect a calibration plot on the target distribution.
These diagnostics make specific transfer failures measurable. None uniquely identifies a cause: enlarged residuals, for example, can reflect several changes at once. Their cost depends on access to representative target data and interventions.
Causal, Compositional, and Persistent Structure
The second frontier question concerns structure: what kind of structure can learned world models acquire, and what does acquiring it require?
"Structure" is not a single property that a model either has or lacks. Temporal structure, causal structure, compositional structure, and persistent-memory structure are different properties with different requirements. A model can have one and lack another. Evidence for one property alone need not establish the others, although a well-designed test may support more than one. The next three subsections separate these questions.
What Passive Prediction Cannot Identify
An observational conditional describes the sampled data-generating process. Without additional assumptions, it need not identify an intervention distribution. More observations from that same process do not resolve an ambiguity between observationally equivalent causal models. Known structure, suitable adjustment variables, randomized actions or other identification assumptions can change what the data establish.
Three distinct notions get merged in casual discussion, and Pearl's causal ladder gives the standard separation:
- Association. "Given that the agent has been moving right, what comes next?" Temporal prediction is one example of an associational query. This is the lowest level in the causal-information hierarchy, not necessarily the easiest prediction problem computationally.
- Interventional prediction. "If I force the agent to move left, what happens?" Identifying this effect requires appropriate assumptions: randomized actions, valid adjustment for confounding, or other justified structure. An explicit causal diagram alone need not identify the effect.
- Counterfactual inference. "Given what actually happened, what would have happened had I acted differently?" This also requires a justified account of the individual circumstances or latent disturbances shared across the factual and counterfactual scenarios. Observational fit alone does not generally identify it.
Imagine a manipulation dataset collected by an expert that only pushes toward the goal. Accuracy on that distribution does not validate predictions for pushing away from it. Those predictions can be wrong or miscalibrated; confidence and error are not inevitable, since structural assumptions may support extrapolation. Coverage, architecture, uncertainty estimation and evaluation all influence the outcome. An independent intervention test makes this unsupported action query explicit (Part XII, Chapter 1).
Training loss on expert trajectories need not reveal errors at actions those trajectories never cover. A model can fit the observed distribution while extrapolating poorly elsewhere. A deterministic planner that treats its predictions as exact can then exploit those errors. Broader action coverage, calibrated uncertainty, conservative planning and appropriate structural assumptions are possible remedies. None guarantees correctness outside the conditions under which it was validated; changing the architecture alone does not supply missing evidence.
Causal Representation Learning and Its Assumptions
Causal representation learning attempts to recover variables that behave causally, so that interventions and counterfactuals become well defined. The program is laid out in Schölkopf et al., "Towards Causal Representation Learning", and interventional approaches such as Ahuja et al. show identifiability results under explicit assumptions.
State the assumptions of the particular result. Ahuja et al.'s Theorem 5.3 uses a polynomial decoder with full-column-rank coefficients, support regularity, a hard intervention fixing one latent coordinate, and nonempty interior for the remaining intervention support. The learned decoder has the prescribed polynomial degree, the encoder does not collapse the support, reconstruction is exact on observational and interventional supports, and one encoded coordinate is fixed on the intervention support. Under these conditions, that intervened latent is identified up to shift and nonzero scaling. Their block-affine result uses different intervention-support and representation constraints. These are conditional coordinate-identification results, not proof that arbitrary robot latents recover complete causal dynamics. Check the particular theorem rather than applying one generic checklist to every approach.
The blunt consequence: a learned disentangled representation is not causal evidence. If you train a -VAE and the latent dimensions align with your factors, or if an object-centric slot attention module produces one slot per object, you have found a representation with a certain factorization. You have not shown that intervening on the object corresponding to slot 3 will change the state in the way your model predicts, because you have not run the intervention. The claim and the evidence are different objects.
Separating a representation into components does not establish how interventions work. Even a probability factorization need not imply independence: the chain rule also holds for dependent variables. An independent product makes an additional statistical claim, but neither separation nor independence alone identifies the effect of forcing a component to change. A representation can align with observed factors while leaving their intervention behavior unidentified. The assumptions of a particular causal-identification result, rather than the factorization alone, establish what intervention claims it supports.
The next cell constructs a linear confounded system. The associational predictor sees only . An information-advantaged structural reference also observes the common cause . This is not identification with an unobserved confounder. We report two distinct estimands: noisy individual-outcome error on held-out observational samples, and error in the interventional mean averaged over at three fixed intervention values. Their absolute magnitudes are not directly comparable.
import numpy as np
rng = np.random.default_rng(2024)
n = 4000
Z = rng.normal(size=n)
X = Z + rng.normal(scale=0.5, size=n)
Y = 2.0 * X + Z + rng.normal(scale=0.5, size=n)
# Associational predictor: Y ~ X only.
X_design = np.column_stack([X, np.ones(n)])
w_assoc, *_ = np.linalg.lstsq(X_design, Y, rcond=None)
# Structural predictor: Y ~ X + Z, using the observed confounder.
XZ_design = np.column_stack([X, Z, np.ones(n)])
w_struct, *_ = np.linalg.lstsq(XZ_design, Y, rcond=None)
def assoc_predict(x):
return w_assoc[0] * x + w_assoc[1]
def struct_predict(x, z):
return w_struct[0] * x + w_struct[1] * z + w_struct[2]
n_test = 2000
Zt = rng.normal(size=n_test)
Xt = Zt + rng.normal(scale=0.5, size=n_test)
Yt = 2.0 * Xt + Zt + rng.normal(scale=0.5, size=n_test)
obs_mse_assoc = float(np.mean((assoc_predict(Xt) - Yt) ** 2))
obs_mse_struct = float(np.mean((struct_predict(Xt, Zt) - Yt) ** 2))
x_grid = np.array([-1.0, 0.0, 1.0])
true_inter = 2.0 * x_grid
assoc_inter = assoc_predict(x_grid)
struct_inter = np.array([struct_predict(x, 0.0) for x in x_grid])
inter_mse_assoc = float(np.mean((assoc_inter - true_inter) ** 2))
inter_mse_struct = float(np.mean((struct_inter - true_inter) ** 2))
obs_errors = np.array([obs_mse_assoc, obs_mse_struct])
inter_errors = np.array([inter_mse_assoc, inter_mse_struct])
model_names = ["Associational", "Structural"]
Planners select actions, and their model queries can include actions absent from the recorded data. Test the claimed action effects with an appropriate intervention protocol: apply specified actions and compare repeated outcomes with the corresponding predicted quantities. A discrepancy challenges the prediction or its assumptions; it does not identify whether the cause is a wrong causal coefficient, unmodeled context, sampling error or another failure. Agreement on the tested interventions supplies scoped evidence, not proof of every structural claim.
Compositional Generalization and Held-Out Combinations
Compositional generalization is the ability to handle novel combinations of factors that were each observed separately. The clean experimental design is a held-out-combination split: every factor level appears in training, but selected combinations do not.
Sequence-to-sequence studies found substantial generalization gaps under compositional evaluations (Keysers et al.; Hupkes et al.; Lake and Baroni). Their protocols differ: compound-divergence splits, excluded function-pair combinations, and primitive or length splits are not identical tests. An analogous factor-combination split can be designed for dynamics learning: cover each factor level in training, hold out selected friction–payload or shape–surface pairs, and evaluate damping, contact or action-effect predictions separately on the held-out cells. This is a proposed dynamics evaluation, not a dynamics-transfer result from those sequence studies.
Two questions keep the interpretation of this split clear:
- Are separate factor effects identifiable under the stated model and assumptions? If payload co-varies perfectly with friction, their separate effects need not be identifiable without additional assumptions. For a specified linear-in-parameters family, full column rank establishes coefficient uniqueness and conditioning diagnoses numerical or noise sensitivity; neither establishes identifiability for every nonlinear model. This question matters when interpreting recovered factor effects, but separate-effect identification is not a prerequisite for every held-out-combination performance test: a known structural law can supply those effects without estimating them from the training split.
- The evaluation must distinguish seen from unseen combinations. Report both separately. A single aggregate over seen and unseen cells can conceal their contrast, even when it shows some overall performance change.
The interesting failure mode is when an additive inductive bias is wrong. If the true dynamics multiply two context factors, an additive model can fit the training combinations well and be systematically wrong off the diagonal. The second experiment below constructs exactly this case, and then deliberately breaks the multiplicative assumption too, so that the interaction model's success is not unconditional either.
A held-out-combination split probes the interaction of data coverage, model capacity, and inductive bias. A misspecified restricted family can retain error even with more samples, whereas a richer family may represent the interaction. Rank checks establish identifiability only within the chosen feature family. Additional structural assumptions can resolve ambiguities that the observed combinations alone do not resolve, but those assumptions must be stated and tested.
A model generalizes compositionally on a factor set if, having observed each factor level in some combinations, it predicts correctly on combinations of those levels that were absent from training. Empirical combination-generalization claims here are relative to the specified factors, split, evaluation population and criterion. This operational test does not exhaust abstract notions of compositional structure.
Persistent Memory: Four Different Things
"Memory" in world models covers at least four mechanically distinct mechanisms. Conflating them makes it impossible to say what an ablation measured.
- Recurrent hidden state. The internal state of an RNN or a recurrent state-space model (Part V, Chapter 1). It carries a compressed summary of history, with no guaranteed storage policy and no explicit retrieval.
- A belief state. The sufficient statistic of the history for the current latent state under the model's assumptions (Part II, Chapter 2). A belief is a filtered, uncertainty-aware summary. Under particular models a compact parametric belief can be much smaller than the stored history; a general continuous-state posterior need not have a fixed-size representation.
- External memory. This includes distinct mechanisms, not just one episodic database. Graves et al. introduced Neural Turing Machines with differentiable external memory and a learned read/write controller. Blundell et al. used non-parametric episodic state-action return memory for control. Lewis et al. provide a retrieval-augmented generation example in NLP, not a control result or an origin claim for all retrieval augmentation. A key-value or vector store can retain experience without updating model weights, but retrieval errors, staleness and storage costs still need evaluation.
- Causal finite-window transformer context. Past tokens inside a configured window are attended over directly (Part V, Chapter 2). This is one implementation, not a definition of all transformer attention. Access does not guarantee reliable retrieval. Tokens outside the window are not directly attended to unless another mechanism, such as a summary, recurrent state, or external store, retains their information.
Each mechanism has possible failure modes. A learned recurrent state can drift and can be difficult to inspect; that difficulty depends on the representation. Beliefs can be systematically miscalibrated if the model is wrong. External stores can retrieve the wrong entry, or the right entry from the wrong episode. Dropping a token removes its direct attention path; the effect on a decision depends on redundancy and any retained summaries, so behavioral forgetting need not be a hard discontinuity.
Four questions should accompany any memory claim:
- What information is retained? Episode identity, hidden context, task goal, object permanence, a recently observed cue.
- How is it refreshed? Never? On a schedule? When prediction error rises? Is there a hazard rate for context change? The third experiment shows that a memory without a refresh mechanism is worse than no memory at all once the context changes.
- How much does it cost? Storage, retrieval latency, and the compute and energy per decision (Part XII, Chapter 4).
- What are the reset semantics? When does memory clear? At episode boundaries, at task changes, at a detected distribution shift? Retention across a reset can carry stale evidence when the reset changes that information's validity. Other retained facts, such as unchanged robot geometry, may remain useful.
The four mechanisms differ not just in what they retain but in how they fail to be updated:
- A recurrent state is updated at each invoked recurrent step, which need not be every environment step. Its learned update rule may be poorly matched to the rate at which the world changes.
- A belief is updated by a filter whose assumptions may be wrong.
- A standard rolling FIFO context window drops tokens as new ones arrive, without a relevance judgment. Selective retention policies can use relevance instead.
- An external store changes through its configured write, eviction, and retention rules; an agent need not control every update. Stored information can outlast the context in which it was valid.
Each mechanism needs an explicit retention and refresh policy. The third experiment isolates fixed and updating beliefs rather than comparing these storage architectures.
Memory Can Hurt
More retained information need not improve the decisions of a fixed learned agent. The following possible failures show why retention and use must be evaluated together.
- Stale context. The goal or physics changed and memory did not. Part VI, Chapter 5 and Part XII, Chapter 5 discussed continual correction in this setting.
- Retrieval of wrong evidence. A nearest-neighbor retrieval returns a transition from a visually similar but dynamically different situation.
- Distractor amplification. Memory preserves salient but irrelevant evidence, and a learned policy trained to attend to memory can over-weight it.
- Compression loss. Summarizing history discards exactly the rare detail that made the current situation special.
For this toy, memory is an estimate of a task-relevant belief. Other memory systems may retain raw history rather than estimate a sufficient statistic. Evaluate the retained information and update mechanism under the context changes they will actually face. For an ablation of history access, match current, nonhistorical inputs and training and compute budgets as far as the design permits. Historical information is the intended treatment, not an input confound to remove. Extra parameters or different training can change the comparison; document such differences and use additional controls before attributing the effect to history access alone.
A lossy history summary can discard information; lossless coding or an invertible representation need not do so. If the discarded information is irrelevant to the declared task, a summary can reduce storage while preserving what that task needs, but computing and maintaining it also costs resources. If an omitted detail matters under a later condition, the summary may no longer be sufficient. Report performance across the expected conditions and their frequencies, alongside retention and update costs, rather than relying on an average alone.
Evaluating Beyond Visual Prediction
The third frontier question is about measurement. Auditing whether current metrics support the intended decision claim is a practical research priority proposed here; its payoff depends on the task and available evaluation data.
Part XI introduced five complementary evaluation protocols: teacher-forced one-step prediction, recursive open-loop rollout, intervention-response comparison, closed-loop planning, and transfer. These are not strictly nested tests. Results under one protocol do not by themselves establish performance under another with different inputs or objectives.
These metrics measure different quantities, not progressively more precise versions of one quantity. An image-space loss measures its chosen discrepancies; pixel MSE alone does not establish perceptual fidelity. A calibration metric measures correspondence between stated probabilities or intervals and realized events. A closed-loop metric measures performance on a declared goal. A link between these quantities needs its own assumptions or evidence; a score label does not supply that link.
The Dependency Chain of Metrics
Consider this conceptual progression of quantities. It is not a universal ordering of evaluation cost or decision relevance:
- One-step prediction loss. Mean squared error or negative log-likelihood of given the history and action. With densely recorded targets, scoring can be inexpensive relative to new data collection. Its weighting depends on the chosen loss and units; raw observation MSE can be dominated by large actual residuals, not necessarily by the highest-variance coordinates.
- Recursive rollout error. Error after autoregressive steps with the model consuming its own predictions (Part XII, Chapter 1 analyzed compounding).
- Calibration. Whether stated probabilities or prediction intervals agree with realized events. Calibration is one input to a safety argument, not validation of the complete safety filter; the decision rule and operating assumptions also need testing (Part XII, Chapter 2).
- Interventional prediction error. Error in predicted outcomes under specified interventions, which may be inside or outside the logging policy's support. Such tests directly probe action effects. Other validated prediction and rollout tests can also inform planning; an intervention mean is not an individual counterfactual.
- Action ranking accuracy. Given a state and a set of candidate action sequences, does the model order them the way the environment does? Ranking matters for an argmin over that set. Numerical value scales also matter when a planner applies thresholds, risk penalties or trade-offs between objectives.
- Closed-loop return and success. The decision-relevant outcome, measured in the real or high-fidelity environment, preferably with the evaluator blinded to the model identity where feasible.
- Safety-relevant outcomes. Constraint violations, minimum separation, peak force, recovery latency.
- Latency and adaptation cost. How long a decision takes, and what recovery after a shift requires under a declared criterion. Samples, compute, wall time and energy are possible adaptation-cost measures.
One-step observation loss alone generally does not establish the later outcomes in this list. The next experiment shows why for a particular task. Raw mean squared error weights coordinates equally in the chosen units. High-variance coordinates can dominate when their actual squared prediction errors are large; variance is not an explicit metric weight. A decision depends on the coordinates the task needs, which need not have high variance. More data can reduce prediction errors, even to zero in an appropriate setting, but does not redefine the metric's weighting or supply a universal implication from aggregate error to decision performance. A restricted model family may admit such a link; state and test its assumptions.
The Evaluator Depends on the Model
Evaluation can depend on the model without being circular. Distinguish population selection from using the model's own predictions as ground truth:
- A policy trained in the model can be evaluated validly in an independent simulator or physical system. The evaluation becomes self-scoring if the learned model also supplies the outcome used as truth.
- Resemblance between logging and evaluation policies does not establish their action-support overlap. Report which states and actions were covered, and which target queries require extrapolation; different policies can also share support.
- If the metric is a learned critic, then a metric improvement can reflect critic overfitting rather than model improvement.
Use an independent outcome evaluator, audit the target state-action coverage, and report prediction metrics alongside realized closed-loop outcomes when assessing a controller. A different collection policy is one possible coverage intervention, not a guarantee of new support. The first experiment uses a task-like evaluation policy distinct from the exploratory training policy. Obtaining additional simulator or hardware evidence has a cost that must be measured for the application.
Model-generated trajectories need not stay in already fitted regions. Scoring them against independent outcomes can reveal errors; reusing the same model as the truth source can instead hide them. Keep tuning instances separate from reporting instances, and document how the evaluation policy selects its population. A fixed independent simulator or physical measurement supplies an outcome check that the learned model does not define.
Splits, Leakage, and Correlated Repeats
Two statistical checks matter for interpreting a reported result: overlap between sampling units and the random sources varied by repeated runs.
Leakage across related transitions. Random transition splits can place strongly related samples from one rollout on both sides of the split, making the score optimistic for generalization to new episodes. Correlation alone does not prove memorization, and adjacent samples need not be near-duplicates. Use trajectory-level or episode-level splits when new-episode performance is the question, and task-level splits for new-task performance. Grouping episodes removes within-episode overlap, but shared environments, duplicated trajectories and preprocessing fitted on the test set still require separate checks.
Seeds are not automatically independent systems. Report what each seed varies: initialization, data selection, environment trajectories, task draws or combinations of these. Repeated runs on a fixed physical configuration estimate variability under the random sources actually changed; they do not by themselves sample new arms, camera mounts or floors. If the deployment claim spans those configurations, include representative configurations as sampling units. When only a few were tested, state that limitation without treating seed count as configuration count.
Keep the sampling unit tied to the claim. An error bar over initialization-only repeats answers a different question from one over changed hardware configurations. If evaluation trajectories also vary, the repeat variability includes that randomness. Report the protocol, tested configurations and conditional uncertainty rather than labeling all seed variance as either training noise or deployment risk. There is no general numerical lower-bound relation between these different quantities.
With those principles in hand, we can run three small experiments that make the abstract points concrete.
Experiment 1: Prediction Fidelity Versus Decision Usefulness
An aggregate observation metric can rank two models in the opposite order from their usefulness for a particular control task. We construct an example, not a general ranking of world-model architectures. All learned models receive the same transitions and actions. The state-focused model additionally receives the declared task-relevant coordinate set, so this is an information-abstraction comparison rather than a parameter-matched architecture benchmark.
The observation contains a small controlled state and a large, persistent background. A model can devote its limited latent capacity to reconstructing that background, reduce observation error, and still omit the coordinates the controller needs. We measure the variance hierarchy, per-coordinate prediction errors, and real-simulator control costs rather than assuming that the intended mechanism occurred.
A Stable Controlled Process and an Autonomous Background
The decision-relevant state is . Both coordinates and the scalar action use dimensionless units. The environment evolves as
Here is the positioning coordinate, is a motion coordinate, and is a normalized command. This is a discrete synthetic law, not an Euler approximation to a specified physical robot. The decay terms prevent the unbounded random walk that a long randomly accelerated free particle would exhibit. Starting from and , the rectangle , is invariant: the first update has magnitude at most , and the monotone velocity map on has magnitude at most . The small drag term makes the learned linear predictors approximations rather than exact copies of the plant.
An autonomous nuisance state follows
The initial makes this process stationary in distribution, with unit marginal variance. Its eight observed channels are , where has its first row equal to one in the first four columns and zero elsewhere, and its second row equal to one in the final four columns. Each nuisance channel therefore has marginal standard deviation 20, but the eight channels have rank two and are not independent observations. This deliberately redundant background strongly influences unnormalized observation loss. Its transition law contains neither the controlled state, action nor goal. That autonomy does not imply statistical independence from chosen actions: a learned feedback controller can use nuisance observations and thereby couple its commands, and subsequent controlled states, to the background.
The observation is . Its coordinates have different meanings and scales; averaging their squared errors embeds a measurement choice. The task is to approach an episode-specific target . Nuisance channels do not appear in the task cost.
import numpy as np
AR_N = 0.95
INNOV_N = np.sqrt(1.0 - AR_N**2)
NUISANCE_STD = 20.0
W_NUISANCE = np.zeros((2, 8))
W_NUISANCE[0, :4] = 1.0
W_NUISANCE[1, 4:] = 1.0
def true_step(x, v, action):
return 0.85 * x + 0.2 * v, 0.8 * v + 0.2 * action - 0.01 * v * np.abs(v)
def nuisance_obs(n):
return NUISANCE_STD * (n @ W_NUISANCE)
def nuisance_step(n, rng):
return AR_N * n + INNOV_N * rng.normal(size=2)The eight observation channels cannot reveal future nuisance innovations. Even the analytic conditional-mean reference must predict their mean, not substitute realized future values. That restriction will keep the reference causal.
Independent Trajectories and Declared Splits
We collect one hundred training trajectories, each with forty random-action transitions and an independent reset. The test set uses a different seed and a fixed task-like behavior policy. Every complete trajectory belongs to only one split. We concatenate transition pairs, not adjacent rows across resets, so the last state of one episode never becomes the first state of another.
def collect_transitions(n_episodes, n_steps, seed, task_like=False):
rng = np.random.default_rng(seed)
previous, following, actions = [], [], []
for _ in range(n_episodes):
x, v = rng.uniform(-1.5, 1.5), rng.uniform(-1.0, 1.0)
n = rng.normal(size=2)
target = rng.uniform(-0.8, 0.8)
for _ in range(n_steps):
obs = np.concatenate(([x, v], nuisance_obs(n)))
action = (
1.8 * (target - x) - 1.2 * v + 0.2 * rng.normal()
if task_like
else rng.uniform(-1.0, 1.0)
)
action = float(np.clip(action, -1.0, 1.0))
x, v = true_step(x, v, action)
n = nuisance_step(n, rng)
previous.append(obs)
following.append(np.concatenate(([x, v], nuisance_obs(n))))
actions.append(action)
return np.array(previous), np.array(following), np.array(actions)
train_prev, train_next, act_train = collect_transitions(100, 40, seed=1)
test_prev, test_next, act_test = collect_transitions(
100, 40, seed=2, task_like=True
)
train_std = train_prev.std(axis=0)
variance_ratio = float(np.mean(train_std[2:]) / np.mean(train_std[:2]))Train/test transition shapes: (4000, 10), (4000, 10) Relevant channel standard deviations: [0.33752297 0.23911176] Nuisance channel standard deviations: [19.3815531 19.3815531 19.3815531 19.3815531 19.47159052 19.47159052 19.47159052 19.47159052] Mean nuisance/relevant standard-deviation ratio: 67.4 Maximum training |x| and |v|: [1.48468296 0.99669796]
These diagnostics describe the actual finite sample. The analytic standard deviation of 20 is a population quantity, not a claim that a sampled trajectory has exactly that standard deviation. The stable controlled process keeps the relevant coordinates bounded without clipping them at a hidden threshold. Random actions provide variation over the allowed command interval; they do not establish coverage of every state-action pair.
Three Learned Models and a Conditional-Mean Reference
The observation models obtain a PCA basis from training inputs and fit a linear action-conditioned transition in that basis. We compare two and four components. The four-dimensional observation subspace contains both relevant coordinates and the two nuisance dimensions. The state-focused model fits only the first two coordinates and fills unmodeled background predictions with the training mean.
These models differ in what they represent and in parameter count. The state-focused model knows which coordinates the task cost uses. PCA knows neither the cost nor that coordinate assignment. The comparison tests the mismatch between an unweighted reconstruction objective and this declared task, not whether supervised coordinate selection is automatically available for a real robot.
mean_o = train_prev.mean(axis=0)
_, singular_values, Vt = np.linalg.svd(train_prev - mean_o, full_matrices=False)
def fit_latent_model(k):
P = Vt[:k].T
current = (train_prev - mean_o) @ P
following = (train_next - mean_o) @ P
features = np.column_stack([current, act_train, np.ones(len(act_train))])
coef, *_ = np.linalg.lstsq(features, following, rcond=None)
return {"P": P, "A": coef[:k].T, "B": coef[k], "c": coef[k + 1]}
def fit_state_model():
features = np.column_stack(
[train_prev[:, :2], act_train, np.ones(len(act_train))]
)
coef, *_ = np.linalg.lstsq(features, train_next[:, :2], rcond=None)
return {"A": coef[:2].T, "B": coef[2], "c": coef[3]}
m_obs2, m_obs4, m_state = (
fit_latent_model(2),
fit_latent_model(4),
fit_state_model(),
)
relevant_loading = float(np.linalg.norm(m_obs2["P"][:2]))Two-component basis loading on relevant coordinates: 0.00021 Two/four-component action coefficient norms: 0.10790, 0.22639 State-model action coefficient: [3.68779459e-17 2.00045596e-01]
A small relevant-coordinate loading supports the intended compression mechanism in this sample. It need not be exactly zero: finite-sample correlations can leak action-related information into the nuisance-dominated basis. A nonzero action coefficient alone does not establish useful action effects; we also measure predictions and control outcomes.
For a latent dimension , the fitted matrices have shapes , , . A batch of latent rows is advanced as , giving predictions. The state model has and receives state rows directly.
One-Step Error Without Future Leakage
All predictors receive the current observation and action. The analytic reference uses the true controlled law and the nuisance conditional mean . It does not observe the next nuisance innovation. Its aggregate error is therefore positive (the nuisance channels contribute the residual variance of the process) even though its prediction of the noiseless relevant transition is exact.
def latent_prediction(model, observations, actions):
latent = (observations - mean_o) @ model["P"]
following = (
latent @ model["A"].T + np.outer(actions, model["B"]) + model["c"]
)
return mean_o + following @ model["P"].T
def state_prediction(model, observations, actions):
following = (
observations[:, :2] @ model["A"].T
+ np.outer(actions, model["B"])
+ model["c"]
)
background = np.tile(mean_o[2:], (len(observations), 1))
return np.column_stack([following, background])
def oracle_prediction(observations, actions):
x, v = true_step(observations[:, 0], observations[:, 1], actions)
return np.column_stack([x, v, AR_N * observations[:, 2:]])
predictions = {
"M_obs (2 comps)": latent_prediction(m_obs2, test_prev, act_test),
"M_obs (4 comps)": latent_prediction(m_obs4, test_prev, act_test),
"M_state (relevant)": state_prediction(m_state, test_prev, act_test),
"Oracle": oracle_prediction(test_prev, act_test),
}
per_channel = {
name: np.mean((prediction - test_next) ** 2, axis=0)
for name, prediction in predictions.items()
}
agg_mse = {name: float(error.mean()) for name, error in per_channel.items()}
relevant_rmse = {
name: float(np.sqrt(error[:2].mean()))
for name, error in per_channel.items()
}
metric_names = list(predictions)
metric_values = np.array([agg_mse[name] for name in metric_names])
relevant_values = np.array([relevant_rmse[name] for name in metric_names])model aggregate MSE relevant RMSE M_obs (2 comps) 31.793 0.2638 M_obs (4 comps) 31.844 0.0004 M_state (relevant) 342.107 0.0004 Oracle 31.707 0.0000
The state-focused model's background mean is an explicit output-completion rule, not an optimal nuisance predictor. A learned nuisance predictor could change its aggregate error and ranking without changing its controller. The direction and size of that change would need measurement. The relevant-coordinate RMSE and controller cost therefore answer questions that the aggregate score alone cannot.


The observed comparison is an existence demonstration. In particular, the state-focused model's mean-completion rule produces a far larger aggregate observation error than the two-component observation model (because the mean-completion residuals on the large-scale nuisance channels dominate), even though the state-focused model has the smaller relevant-coordinate error. This shows that an aggregate score dominated by decision-irrelevant channels and a task-relevant score can rank the models oppositely. The four-component model is a capacity ablation, not a failure that must be manufactured. Retaining the controlled subspace can improve both prediction and control.
Closed-Loop Evaluation with Matched Candidate Actions
All model-based controllers use random-shooting MPC with sequences of actions. They minimize the same predicted cost,
Here are the model states after candidate action , and is the known episode goal. Each controller executes only the first action and replans from the next observation. The real simulator measures cost after that action, using the same coordinate weights. This keeps the prediction horizon's indexing consistent with evaluation.
A model-free PD controller provides a baseline. Its gains are tuned on separate simulator episodes, and the final comparison uses held-out episodes. The analytic MPC controller has the true dynamics, but finite-horizon random shooting is not an optimal-control solver; its cost is a reference, not a universal lower bound. Candidate action arrays and environmental resets use the same seeds across model-based controllers. The prediction model is the changed component; earlier selected actions can then produce different observations and subsequent decisions.
def make_linear_step(model, latent):
def step(states, actions):
following = (
states @ model["A"].T + np.outer(actions, model["B"]) + model["c"]
)
decoded = mean_o + following @ model["P"].T if latent else following
return following, decoded[:, 0], decoded[:, 1]
return step
def oracle_batch_step(states, actions):
x, v = true_step(states[:, 0], states[:, 1], actions)
return np.column_stack([x, v]), x, v
def plan(step_fn, initial, target, rng, K=200, H=8):
actions = rng.uniform(-1.0, 1.0, size=(K, H))
states = np.repeat(np.asarray(initial)[None, :], K, axis=0)
costs = np.zeros(K)
for h in range(H):
states, x, v = step_fn(states, actions[:, h])
costs += (x - target) ** 2 + 0.05 * v**2 + 0.01 * actions[:, h] ** 2
return float(actions[np.argmin(costs), 0])def model_controller_factory(step_fn, initial_fn):
def factory(seed):
rng = np.random.default_rng(seed)
return lambda obs, target: plan(step_fn, initial_fn(obs), target, rng)
return factory
def pd_factory(kp, kd):
return lambda seed: (
lambda obs, target: float(
np.clip(kp * (target - obs[0]) - kd * obs[1], -1.0, 1.0)
)
)
def evaluate(factory, n_episodes=20, n_steps=40, seed=7):
rng = np.random.default_rng(seed)
costs = []
for episode in range(n_episodes):
controller = factory(1000 + episode)
x, v = rng.uniform(-1.5, 1.5), rng.uniform(-1.0, 1.0)
target = rng.uniform(-0.8, 0.8)
n, total = rng.normal(size=2), 0.0
for _ in range(n_steps):
observation = np.concatenate(([x, v], nuisance_obs(n)))
action = float(np.clip(controller(observation, target), -1.0, 1.0))
x, v = true_step(x, v, action)
n = nuisance_step(n, rng)
total += (x - target) ** 2 + 0.05 * v**2 + 0.01 * action**2
costs.append(total)
return np.array(costs)
pd_grid = [
(kp, kd) for kp in (0.5, 1.0, 1.5, 2.0) for kd in (0.5, 1.0, 1.5, 2.0)
]
pd_scores = {
g: evaluate(pd_factory(*g), n_episodes=10, seed=11).mean() for g in pd_grid
}
best_pd = min(pd_scores, key=pd_scores.get)
controller_factories = {
"PD (no model)": pd_factory(*best_pd),
"MPC M_obs (2 comps)": model_controller_factory(
make_linear_step(m_obs2, True), lambda obs: (obs - mean_o) @ m_obs2["P"]
),
"MPC M_obs (4 comps)": model_controller_factory(
make_linear_step(m_obs4, True), lambda obs: (obs - mean_o) @ m_obs4["P"]
),
"MPC M_state (relevant)": model_controller_factory(
make_linear_step(m_state, False), lambda obs: obs[:2]
),
"MPC oracle": model_controller_factory(
oracle_batch_step, lambda obs: obs[:2]
),
}
cl_costs = {
name: evaluate(factory) for name, factory in controller_factories.items()
}
cl_mean = {name: float(costs.mean()) for name, costs in cl_costs.items()}
cl_sem = {
name: float(costs.std(ddof=1) / np.sqrt(len(costs)))
for name, costs in cl_costs.items()
}
controller_names = list(cl_mean)
controller_values = np.array([cl_mean[name] for name in controller_names])
controller_errors = np.array([cl_sem[name] for name in controller_names])
prediction_to_controller = {
"M_obs (2 comps)": "MPC M_obs (2 comps)",
"M_obs (4 comps)": "MPC M_obs (4 comps)",
"M_state (relevant)": "MPC M_state (relevant)",
}
scatter_names = list(prediction_to_controller)
scatter_x = np.array([agg_mse[name] for name in scatter_names])
scatter_y = np.array(
[cl_mean[prediction_to_controller[name]] for name in scatter_names]
)PD gains selected on tuning episodes: (2.0, 0.5)
controller mean cost SEM
PD (no model) 3.222 0.795
MPC M_obs (2 comps) 9.769 1.841
MPC M_obs (4 comps) 2.596 0.810
MPC M_state (relevant) 2.596 0.810
MPC oracle 2.596 0.810
Pairwise aggregate error: {'M_obs (2 comps)': 31.793339997706347, 'M_state (relevant)': 342.10720636320696}
Pairwise control cost: {'M_obs (2 comps)': 9.769208180140867, 'M_state (relevant)': 2.596215849016851}The SEMs describe sampling variability over the twenty simulated episodes, conditional on these fitted models and this fixed experiment. They do not measure variation over real systems or independent training runs. Common random numbers can reduce comparison noise when the paired outcomes have positive covariance; they do not remove environmental randomness or guarantee smaller variance for every pair. Estimate uncertainty in the paired cost differences when assessing a particular comparison.


Reading the Result Without Overclaiming
The two-component observation model retains little of the controlled state and its controller incurs greater cost than the state-focused model in this test. Yet completing the state-focused model's nuisance outputs with the training mean penalizes it heavily in aggregate observation error. This pair reverses the error-versus-control ordering. The four-component model is a useful counterpoint: retaining the relevant subspace can make an observation model useful for control too. There is no claim that the aggregate metric is flat across dimension counts, or that its globally best scorer must be the worst controller.
Several design choices matter. We specified a large background whose transition law is unaffected by actions; gave the state-focused learner task-coordinate knowledge; used a particular output-completion rule; and chose one goal family and planner. Changing any of these can change the rankings. In particular, adding a nuisance predictor to the state-focused model can lower its aggregate error without changing its decisions. Normalizing channels or using task-sensitive loss weights are possible alternatives, but this notebook has not trained and tested those alternatives, so it does not claim they recover a ranking.
Report coordinate-wise errors alongside the full observation score and evaluate actions in an independent environment. Good one-step predictions do not certify multi-step reliability, calibration, or robustness after a dynamics change. Conversely, an aggregate score containing coordinates outside a model's declared prediction contract may say little about its intended consumer. The example makes that distinction inspectable with code; it supplies no hardware evidence or architecture-wide conclusion.
Experiment 2: Held-Out Combinations and When Factorization Fails
The second experiment asks whether a factorized representation can predict combinations of factors it never saw together. We construct a transition family in which the true interaction between two context factors is multiplicative, so that an additive model is structurally wrong, and then we add a stress condition in which even the multiplicative assumption fails.
The product law supplies a specific interaction that the additive feature family cannot represent exactly across the chosen factor levels. Even when an additive law is correct, its finite-data estimation and held-out predictions can still be tested. Here we use a normalized quadratic damping term, , and let the coefficient depend on two context factors; this is a declared toy law, not a universal physical drag model. The capped regime asks a separate question: how does the product-feature model behave when that law changes? A test supporting a model is not thereby rigged, but testing a violated assumption makes the boundary of the result explicit.
The Transition Family
The system is a scalar velocity with quadratic damping:
where:
- : scalar velocity at step
- : action (normalized acceleration) at step
- : integration timestep
- : damping coefficient determined by the surface factor and payload factor
- : surface factor level
- : payload factor level
- : additive noise at step
where is a surface factor, is a payload factor, and both take levels in . The coefficient is what we will try to learn. In the main regime,
In the stress regime, the damping saturates:
Here is the saturation ceiling: at high the damping stops growing, so the true law is no longer a pure product.
The stress regime specifies a capped coefficient, . This is a deliberate alternative law, not a claim that every physical damping coefficient saturates. The held-out products are 3 and 6, so the test includes uncapped and capped combinations. A pure-product feature family is therefore misspecified in this regime. The coefficients and factors here are normalized toy quantities; the experiment does not establish a physical stability margin.
Rather than fit a full transition network, we estimate the single scalar for each context by least squares, then study how meta-models map context to . This controls the feature families and makes the scalar identifiability condition explicit. It does not remove estimation error: the finite noisy identification samples produce uncertain coefficient labels, and that uncertainty propagates into the meta-model errors alongside approximation error. Known ground truth lets us measure those errors, not declare the identification step exact. A real system may also permit scalar fitting, or may require a richer model with additional measurement and model errors.
FACTOR_LEVELS = (1.0, 2.0, 3.0)
SEEN_COMBOS = [(1, 1), (1, 2), (2, 1), (2, 2), (3, 3)]
HELD_OUT_COMBOS = [(1, 3), (2, 3), (3, 1), (3, 2)]
DT_C = 0.2
KAPPA_CAP = 5.0
IDENT_NOISE = 0.4
N_SAMPLES = 60
def kappa_true(s, p, saturated):
k = float(s) * float(p)
return min(k, KAPPA_CAP) if saturated else k
def identify_kappa(s, p, saturated, rng, n_samples=N_SAMPLES):
v = rng.normal(scale=1.2, size=n_samples)
a = rng.uniform(-1.0, 1.0, size=n_samples)
k = kappa_true(s, p, saturated)
v_next = (
v
+ DT_C * (a - k * v * np.abs(v))
+ IDENT_NOISE * rng.normal(size=n_samples)
)
regressor = -DT_C * v * np.abs(v)
target = v_next - v - DT_C * a
return float(regressor @ target / (regressor @ regressor))The identification samples pairs independently rather than rolling the system forward. For this one-parameter least-squares fit, a unique coefficient estimate requires . More spread can improve information under the stated noise model, but a visual velocity range is not a general persistent-excitation test. Gaussian velocity samples have unbounded support; the code imposes no hard limit.
The Split and the Two Meta-Models
Every factor level appears in training. What is held out is specific combinations: , , , and . The training set contains the diagonal and two symmetric off-diagonal cells and .
The structure of the split is deliberate. Every level of and every level of appears in training, so the model has seen each factor individually. What it has not seen is the specific way high combines with high and vice versa. If the factors combine additively, the model can predict the held-out cells by extrapolating each factor's effect. If they combine multiplicatively, the model needs to know the product, and the training cells do not uniquely determine the product unless the model's functional form includes it. This is the sense in which the split tests the inductive bias rather than the data coverage.
Two meta-models map context to the damping coefficient:
- Factorized (additive) baseline. . No interaction term.
- Interaction-capable baseline. . Characterizes the multiplicative structure.
Both are fit by least squares on the training cells. We check the design matrix rank and condition number so we are not fitting an arbitrary solution.
def fit_additive(pairs):
X = np.array([[1.0, s, p] for s, p, _ in pairs])
y = np.array([k for _, _, k in pairs])
w, *_ = np.linalg.lstsq(X, y, rcond=None)
return (lambda s, p: w[0] + w[1] * s + w[2] * p), X
def fit_interaction(pairs):
X = np.array([[1.0, s, p, s * p] for s, p, _ in pairs])
y = np.array([k for _, _, k in pairs])
w, *_ = np.linalg.lstsq(X, y, rcond=None)
return (lambda s, p: w[0] + w[1] * s + w[2] * p + w[3] * s * p), X
_rng_check = np.random.default_rng(3)
_check_seen = [
(s, p, identify_kappa(s, p, False, _rng_check)) for s, p in SEEN_COMBOS
]
_, X_add = fit_additive(_check_seen)
_, X_int = fit_interaction(_check_seen)
print(
f"Additive design rank: {np.linalg.matrix_rank(X_add)} of {X_add.shape[1]}"
)
print(
f"Interaction design rank: {np.linalg.matrix_rank(X_int)} of {X_int.shape[1]}"
)
print(f"Interaction condition number: {np.linalg.cond(X_int):.1f}")Additive design rank: 3 of 3 Interaction design rank: 4 of 4 Interaction condition number: 58.4
Both design matrices have full column rank, so least-squares coefficients are uniquely determined within each selected feature family. The reported interaction-design condition number diagnoses sensitivity for that matrix. These checks support the earlier section's model-family-specific identification discussion; they are not prerequisites for every held-out-combination performance test and do not establish identifiability for arbitrary nonlinear models.
Before comparing meta-models, we check the scalar identification procedure against analytic truth on the held-out cells. This diagnostic measures its error in the chosen unsaturated contexts; it does not make finite noisy labels exact. Downstream error still combines identification noise with estimation and functional-form error. Keep that contribution visible rather than attributing every discrepancy to the meta-model's structure.
_rng_verify = np.random.default_rng(77)
verify_rows = []
for combo in HELD_OUT_COMBOS:
s, p = combo
est = np.mean([identify_kappa(s, p, False, _rng_verify) for _ in range(20)])
verify_rows.append((s, p, kappa_true(s, p, False), est))
print(f"{'s':>3}{'p':>3}{'true kappa':>13}{'identified':>13}")
for s, p, truth, est in verify_rows:
print(f"{s:>3}{p:>3}{truth:>13.3f}{est:>13.3f}")s p true kappa identified 1 3 3.000 2.984 2 3 6.000 5.982 3 1 3.000 3.028 3 2 6.000 6.009
These diagnostic averages track the analytic coefficients closely in the unsaturated regime. Identification noise still enters every fitted meta-model, so held-out error combines that noise, estimation variance, and functional-form mismatch. The repeated study below measures this combined variability rather than attributing every error exclusively to structure.
Running Both Regimes with Repeats
We repeat the whole procedure eight times with fresh identification noise, so the error bars capture variation from sampling, identification and meta-model fitting under this fixed split. They are standard errors of each model's mean, not standard deviations of individual runs. Assess a model difference against uncertainty in its estimated mean difference; it need not exceed individual-run variability. The models share each repeat's training data, so paired repeat differences would provide a direct uncertainty estimate for their comparison. The separate marginal bars below do not supply that estimate.
def run_composition_study(saturated, n_repeats=8, seed=0):
rng = np.random.default_rng(seed)
add_e = np.zeros((n_repeats, len(HELD_OUT_COMBOS)))
int_e = np.zeros_like(add_e)
seen_add_e = np.zeros((n_repeats, len(SEEN_COMBOS)))
seen_int_e = np.zeros_like(seen_add_e)
for r in range(n_repeats):
seen = [
(s, p, identify_kappa(s, p, saturated, rng)) for s, p in SEEN_COMBOS
]
f_add, _ = fit_additive(seen)
f_int, _ = fit_interaction(seen)
for j, (s, p) in enumerate(SEEN_COMBOS):
truth = kappa_true(s, p, saturated)
seen_add_e[r, j] = abs(f_add(s, p) - truth)
seen_int_e[r, j] = abs(f_int(s, p) - truth)
for j, (s, p) in enumerate(HELD_OUT_COMBOS):
truth = kappa_true(s, p, saturated)
add_e[r, j] = abs(f_add(s, p) - truth)
int_e[r, j] = abs(f_int(s, p) - truth)
return add_e, int_e, seen_add_e, seen_int_e
add_main, int_main, seen_add_main, seen_int_main = run_composition_study(
saturated=False, seed=101
)
add_stress, int_stress, seen_add_stress, seen_int_stress = (
run_composition_study(saturated=True, seed=202)
)
def summarize(e):
per_repeat = e.mean(axis=1)
return float(per_repeat.mean()), float(
per_repeat.std(ddof=1) / np.sqrt(len(per_repeat))
)
summary = {
("Factorized", "Seen, multiplicative"): summarize(seen_add_main),
("Interaction", "Seen, multiplicative"): summarize(seen_int_main),
("Factorized", "Seen, saturating"): summarize(seen_add_stress),
("Interaction", "Seen, saturating"): summarize(seen_int_stress),
("Factorized", "Held-out, multiplicative"): summarize(add_main),
("Interaction", "Held-out, multiplicative"): summarize(int_main),
("Factorized", "Held-out, saturating"): summarize(add_stress),
("Interaction", "Held-out, saturating"): summarize(int_stress),
}model regime mean |err| SEM Factorized Seen, multiplicative 0.450 0.005 Interaction Seen, multiplicative 0.072 0.005 Factorized Seen, saturating 0.321 0.012 Interaction Seen, saturating 0.312 0.011 Factorized Held-out, multiplicative 0.978 0.022 Interaction Held-out, multiplicative 0.131 0.016 Factorized Held-out, saturating 0.492 0.023 Interaction Held-out, saturating 0.629 0.016

Interpreting the Two Regimes
In the multiplicative regime the interaction model has much lower held-out coefficient error than the additive model. The table also reports error against analytic truth on seen combinations: training-side approximation error is part of this comparison, not something to assume away. The product feature represents the specified law, while the additive family cannot represent it over the full grid. Under the capped law that advantage does not persist. These results concern two chosen feature families and splits, not every factorized representation or a universal compositionality advantage.
Under saturation, the interaction model's held-out error increases, while the additive model's error decreases relative to its own multiplicative-regime result. The mean ranking reverses in this seeded study. Both remain misspecified: neither functional form represents the capped product exactly. This is not evidence that saturation always favors an additive model. The product feature has lower error under the product law here, and that advantage does not survive this cap. Misspecification alone does not establish operational unreliability: deciding whether either error is acceptable requires a task-specific tolerance. Test the target regime and report the measured errors against that requirement.
The practical implications are direct:
- Test compositional claims with held-out combinations and report seen and unseen separately.
- When interpreting recovered factor effects, establish identifiability under the chosen feature family and report its design rank. Keep this identification claim separate from measured held-out performance.
- Include a violation condition so that the reader knows what happens when the assumed structure breaks.
- Prefer a model whose inductive bias matches the regime you will deploy in. Justify that regime from known structure or an appropriately powered target-system study; a small trial need not detect rare or later changes.
A predeclared held-out comparison can be a valid test even when it supports the proposed method. Adding a violation condition asks a separate question: how does performance change when the assumed law no longer holds? The capped regime supplies that boundary check here; it does not make all positive-only studies invalid. For deployment, test the target regime and ask whether the measured errors meet its requirements rather than seeking an unconditional model ranking.
An illustrative controlled example with a known ground-truth transition family, not a benchmark result.
Experiment 3: Memory Ablation Under Changing Context
The third experiment compares specified cue-retention and belief-update rules, then changes the goal or corrupts the cue. It is an information-policy comparison, not a pure storage-capacity ablation: the memoryless baseline deliberately ignores the initial cue and later rewards, whereas the updating agent uses both.
To study the benefit of historical information, use a task where the current observation omits relevant information. A fully observed task can still test memory overhead, distraction or harm, but cannot establish that history is necessary. Here the current position omits the hidden goal. The initial cue and later rewards provide information about that goal, and the flip condition changes whether retained cue information remains valid.
A Partially Observed Goal Task
The agent controls a scalar position with dynamics
where:
- : scalar position at step
- : control action at step
- : position decay factor
- : control gain
- : process noise at step
- : projection onto the interval
The dynamics coefficients are shared across episodes. What differs is the goal: in mode 0 and in mode 1. The non-oracle agents do not observe the mode directly. They can receive an initial cue and subsequent noisy rewards:
- A cue at . A noisy reading of the goal, with . For the non-oracle agents, this is the only initial observation carrying goal information; they also observe position.
- A reward signal after each transition. with . This is slow and noisy, but it is the only channel that can reveal a change.
The cue supplies evidence before the first action, but is available only at the start and can be corrupted. A cue-only agent that freezes its belief cannot correct it without another signal or update. Rewards arrive after transitions and can support later correction when their goal-conditioned likelihoods differ. At those likelihoods coincide, so that reward alone does not distinguish the two goals. The updating agent combines cue and reward evidence under its assumed model; correction speed depends on the observed evidence and chosen actions.
Part II, Chapter 2's belief-state machinery supplies a useful representation: a distribution over the hidden mode, updated from observations under a specified model. Its adequacy depends on that model and the decision rule. A belief representation alone does not guarantee robustness to noisy cues or changing goals; this experiment tests one particular filter and controller.
Four agents:
- Memoryless. Ignores the cue entirely and starts from a uniform prior . It uses the prior-mean goal, which is .
- Cue memory, frozen. Uses the cue to set its belief at and never updates it.
- Cue memory with Bayesian update. Uses the cue, then updates the belief from each observed reward, with a hazard rate of representing the prior probability that the context changed on any step.
- Oracle. Knows the mode.
All four share a certainty-equivalent controller: given , the expected goal is , and is clipped to . The map from belief to action is fixed; access to information and belief updates differ. This minimizes the one-step expected squared error under the stated prior and zero-mean process noise when clipping is inactive, but it is not a Bayes-optimal multi-step controller. Actions can influence how informative future rewards are, and this controller does not optimize that information-gathering trade-off.
GAIN = 0.8
DECAY3 = 0.9
T_STEPS = 20
FLIP_AT = 10
GOALS = (1.0, -1.0)
SIG_PROC = 0.05
SIG_REW = 1.5
SIG_CUE = 0.7
HAZARD = 0.05
X_LIMIT = 5.0
def env_step(x, a, proc_noise):
return float(np.clip(DECAY3 * x + GAIN * a + proc_noise, -X_LIMIT, X_LIMIT))
def prior_from_cue(cue):
lik = [np.exp(-0.5 * ((cue - g) / SIG_CUE) ** 2) for g in GOALS]
return float(lik[1] / (lik[0] + lik[1]))
def belief_update(p, x_next, r_obs):
lik = [
np.exp(-0.5 * ((r_obs + (x_next - g) ** 2) / SIG_REW) ** 2)
for g in GOALS
]
post = p * lik[1] / (p * lik[1] + (1.0 - p) * lik[0] + 1e-300)
return float((1.0 - HAZARD) * post + HAZARD * (1.0 - post))
def belief_action(x, p):
x_star_hat = (1.0 - p) * GOALS[0] + p * GOALS[1]
return float(np.clip((x_star_hat - DECAY3 * x) / GAIN, -1.0, 1.0))The reward likelihood uses the observed next state and the hypothetical mean reward . Rewards can distinguish the goals when their predicted means differ; at those means coincide, so this reward alone is uninformative about the mode. belief_update first conditions on the current reward, then applies the symmetric flip hazard. Its returned value is the prior for the next decision, not the current-time posterior. The plotted trace stores that next-decision prior.
The assumed hazard is the probability of a symmetric mode flip between steps. For hazards between zero and one half, applying it moves the reward-conditioned belief toward equal probabilities. This can prevent extreme commitment under anticipated change, but does not ensure faster recovery or lower cost. The frozen agent retains its soft cue belief, the memoryless agent uses an equal-probability belief, and the filter updates from reward before applying the hazard. The sweep measures the trade-off for the selected episodes and decision rule.
def make_episode(rng, corrupt_cue=False, flip_goal=False):
mode = int(rng.integers(0, 2))
corrupted = bool(corrupt_cue and rng.random() < 0.5)
cue_sign = -GOALS[mode] if corrupted else GOALS[mode]
return {
"mode": mode,
"corrupted": corrupted,
"flip_goal": flip_goal,
"x0": float(rng.uniform(-1.0, 1.0)),
"cue": cue_sign + float(rng.normal(scale=SIG_CUE)),
"proc": rng.normal(scale=SIG_PROC, size=T_STEPS),
"rew": rng.normal(scale=SIG_REW, size=T_STEPS),
}
AGENTS = {
"Memoryless": {"use_cue": False, "update": False, "oracle": False},
"Cue memory, frozen": {"use_cue": True, "update": False, "oracle": False},
"Cue memory, Bayesian": {"use_cue": True, "update": True, "oracle": False},
"Oracle": {"use_cue": True, "update": True, "oracle": True},
}
def rollout(episode, agent):
x = episode["x0"]
if agent["oracle"]:
p = float(episode["mode"] == 1)
elif agent["use_cue"]:
p = prior_from_cue(episode["cue"])
else:
p = 0.5
total = 0.0
trace = []
for t in range(T_STEPS):
mode_t = episode["mode"]
if episode["flip_goal"] and t >= FLIP_AT:
mode_t = 1 - mode_t
x_star = GOALS[mode_t]
if agent["oracle"]:
p = float(mode_t == 1)
a = belief_action(x, p)
x_next = env_step(x, a, episode["proc"][t])
r_obs = -((x_next - x_star) ** 2) + episode["rew"][t]
total += (x_next - x_star) ** 2
if agent["update"] and not agent["oracle"]:
p = belief_update(p, x_next, r_obs)
x = x_next
trace.append(p)
return total, np.array(trace)Conditions and Common Random Numbers
Three conditions exercise the memory in different ways:
- Stable goal, clean cue. The base case. Memory should help.
- Goal flips at . The cue is now stale. A memory that does not update is actively harmful.
- Corrupted cue. Each episode independently has probability 0.5 of reversing the cue-generating mean. The Gaussian cue noise remains unchanged. Under the agents' assumed clean-cue likelihood, a reversed mean can favor the wrong goal, but realized cue direction and confidence vary.
The stable condition compares the specified memories when the goal remains fixed. The flip condition introduces stale goal information; the corrupted-cue condition introduces misleading initial information in a mixture of episodes. These test two failure modes, not an exhaustive memory taxonomy. Stable-condition performance alone does not establish performance under either stress condition.
Episodes are generated with a fixed seed and shared across agents, so exogenous cue and noise draws and the initial state match across agents. Actions, subsequent states, and observed rewards can differ; sharing additive reward noise does not make reward observations identical. This common-random-numbers design aligns exogenous randomness. It can reduce paired-difference variance when the agent costs have positive covariance, but does not remove environment variance or guarantee variance reduction for every pair. Unpaired draws can add comparison noise; whether they do so depends on that covariance. Use the paired episode-cost differences when estimating uncertainty in this matched comparison.
CONDITIONS = {
"Stable goal": {"corrupt_cue": False, "flip_goal": False},
"Goal flips at t=10": {"corrupt_cue": False, "flip_goal": True},
"Corrupted cue": {"corrupt_cue": True, "flip_goal": False},
}
N_EPISODES = 300
def evaluate_condition(agent, cond, n_episodes=N_EPISODES, seed=5):
rng = np.random.default_rng(seed)
episodes = [make_episode(rng, **cond) for _ in range(n_episodes)]
return np.array([rollout(ep, agent)[0] for ep in episodes])
results = {
agent_name: {
cond_name: evaluate_condition(agent, cond)
for cond_name, cond in CONDITIONS.items()
}
for agent_name, agent in AGENTS.items()
}
mem_mean = {
(a, c): float(v.mean()) for a, r in results.items() for c, v in r.items()
}
mem_sem = {
(a, c): float(v.std(ddof=1) / np.sqrt(len(v)))
for a, r in results.items()
for c, v in r.items()
}agent Stable goal Goal flips at t=10 Corrupted cue Memoryless 20.086 +/- 0.027 20.038 +/- 0.024 20.033 +/- 0.024 Cue memory, frozen 5.220 +/- 0.812 35.617 +/- 0.363 36.051 +/- 2.020 Cue memory, Bayesian 5.611 +/- 0.362 11.724 +/- 0.287 9.780 +/- 0.456 Oracle 0.323 +/- 0.022 1.616 +/- 0.024 0.329 +/- 0.022

What the Numbers Say
In the stable condition, both cue-based agents have substantially lower mean cost than the memoryless agent. The frozen agent has a slightly lower sample mean than the filter at hazard 0.05, but the reported standard errors do not establish a reliable difference between these two means. The hazard shrinks beliefs toward equal probability; its contribution is better examined with the matched hazard sweep below than attributed from this comparison alone.
The frozen agent has a lower stable-condition sample mean than the filter at hazard 0.05; that is neither a significance result nor an optimality claim. An updating filter with hazard zero can learn from noisy rewards and has lower stable cost in the sweep below. That sweep separates belief updates from the positive-hazard shrinkage toward equal probability: a small positive hazard improves the tested goal-flip mean cost but raises stable-condition mean cost. The chosen update rule and change condition both matter.
In the seeded flip condition, the Bayesian agent has lower mean cost than the frozen agent. The table measures episode cost, not a distribution of detection or recovery times. The frozen agent never updates its cue belief; its mean cost exceeds the memoryless baseline here. This is not a universal ranking of updating and non-updating memories.
The memoryless controller uses an equal-probability goal belief and therefore targets the midpoint. That hedge avoids strong commitment to the wrong endpoint, but also fails to track either endpoint accurately. The frozen controller retains a concentrated, not certain, cue belief. After a change, that concentration can make its actions more costly than the midpoint hedge. A stored estimate can become harmful when its context changes and the controller continues to trust it; the size and direction of the effect depend on the change process and decision rule.
In the corrupted-cue condition, the updating agent has lower mean cost than the frozen agent. Some initial cues favor the wrong goal, with varying confidence. Rewards can then supply evidence for correction when their two goal-conditioned likelihoods differ; at those likelihoods coincide. The selected trace below illustrates correction from a wrong initial belief, not guaranteed recovery in every episode.
trace_rng = np.random.default_rng(99)
trace_episode = None
for _ in range(200):
candidate = make_episode(trace_rng, corrupt_cue=True, flip_goal=False)
if candidate["corrupted"]:
trace_episode = candidate
break
_, trace_bayes = rollout(trace_episode, AGENTS["Cue memory, Bayesian"])
_, trace_frozen = rollout(trace_episode, AGENTS["Cue memory, frozen"])
_, trace_oracle = rollout(trace_episode, AGENTS["Oracle"])
true_mode = float(trace_oracle[0] > 0.5)
steps = np.arange(1, T_STEPS + 1)
Scope and Limitations of the Toy
This task is a sufficient-statistic toy. The belief is over a two-element set, the update is a one-dimensional filter, and the assumed hazard rate is a single scalar. It is a heuristic prior, not a correctly matched generator: the flip condition changes deterministically at step ten rather than sampling independent flips with probability 0.05. Open-world deployments may involve unknown, high-dimensional or drifting context and an unknown change process, rather than this known binary state space. Other deployments can have small known modes or stationary context. Part XII, Chapter 5 discusses monitoring and correction under changing operating conditions.
The agent knows the two modes and the Gaussian reward-noise model. Its assumed hazard does not match the deterministic flip schedule. In the corrupted-cue condition, the cue likelihood is also misspecified: half the cues come from the opposite goal's Gaussian, but prior_from_cue still assumes the clean-cue model. Recovery here therefore concerns a specified controller and misspecified filter, not a generally correct Bayesian solution. A real deployment can require learning or validating its state space, observation model, change process and update rule separately.
What transfers from the toy is the structural lesson:
- State explicitly what information the memory is supposed to carry. Here it is "which goal is active."
- Pair memory with an uncertainty-aware update. Likelihood updates revise an uncertain goal belief as new reward evidence arrives, including when the hazard is zero. A positive hazard separately models possible goal changes by moving the next-decision prior toward equal probability; it is not required for evidence-based updating and does not guarantee better performance.
- Design the reset boundary. Our episodes reset at . Persisting the previous episode's belief into a new episode can start the controller with misleading evidence when the mode changes. Its effect depends on the initialization and update rule; cross-episode memory need not fail if those rules account for the boundary.
- Evaluate under change. Report the stable condition and the changed condition. Evaluation only on stable contexts cannot establish recovery after change; it also does not guarantee good stable performance.
The memory experiment above fixed the hazard rate. Sweeping it exposes a non-monotone trade-off. Moving from zero to a small positive hazard sharply lowers cost in the deterministic-flip condition, but increasing it further eventually raises that cost again. Stable-condition cost rises over the tested grid; near 0.5, repeated hedging discards useful accumulated information.
import numpy as np
def belief_update_hazard(p, x_next, r_obs, hazard):
lik = [
np.exp(-0.5 * ((r_obs + (x_next - g) ** 2) / SIG_REW) ** 2)
for g in GOALS
]
post = p * lik[1] / (p * lik[1] + (1.0 - p) * lik[0] + 1e-300)
return float((1.0 - hazard) * post + hazard * (1.0 - post))
def rollout_hazard(episode, hazard):
x = episode["x0"]
p = prior_from_cue(episode["cue"])
total = 0.0
for t in range(T_STEPS):
mode_t = episode["mode"]
if episode["flip_goal"] and t >= FLIP_AT:
mode_t = 1 - mode_t
x_star = GOALS[mode_t]
a = belief_action(x, p)
x_next = env_step(x, a, episode["proc"][t])
r_obs = -((x_next - x_star) ** 2) + episode["rew"][t]
total += (x_next - x_star) ** 2
p = belief_update_hazard(p, x_next, r_obs, hazard)
x = x_next
return total
hazard_grid = np.array([0.0, 0.01, 0.05, 0.10, 0.20, 0.50])
stable_rng = np.random.default_rng(5)
stable_eps = [
make_episode(stable_rng, corrupt_cue=False, flip_goal=False)
for _ in range(200)
]
flip_rng = np.random.default_rng(6)
flip_eps = [
make_episode(flip_rng, corrupt_cue=False, flip_goal=True)
for _ in range(200)
]
stable_costs = np.array(
[np.mean([rollout_hazard(ep, h) for ep in stable_eps]) for h in hazard_grid]
)
flip_costs = np.array(
[np.mean([rollout_hazard(ep, h) for ep in flip_eps]) for h in hazard_grid]
)
An illustrative controlled example with a hand-specified two-mode environment, not a benchmark result or evidence about a physical system.
Research, Engineering, and Study Roadmaps
The experiments above are local reproducibility exercises with specified mechanisms and scoped numerical performance comparisons. Everything in this section is either a defined exercise to build or test understanding, or a proposed research design seeking additional evidence. This is a distinction between these activities, not an exhaustive definition of research: a careful replication can add evidence about an already studied question. State what was tested, what was reproduced and what new claim the evidence supports.
A Learner's Path
If you are crossing into this field, consider the following learning sequence:
- Reproduce the three experiments above and then break them deliberately. Add a third model to experiment 1 that weights channels by task relevance. Add a third factor to experiment 2 and see how the held-out error grows with the number of held-out cells. Give experiment 3 a memory that persists across episode boundaries and see what breaks.
- Implement a small recurrent state-space model on a simple partially observed control task (Part V, Chapter 1). Compare its belief-like hidden state with a hand-built Bayesian filter under matched information and declared budgets. Differences can reflect representation, assumptions, training or optimization; the comparison does not isolate learning versus analytic filtering by itself.
- Implement one model-based planner end to end: random shooting, then cross-entropy method (Part VII, Chapter 1). Measure how planning performance changes as you change the model, and check whether its ranking matches the prediction-error ranking. The rankings can differ; measure whether they do in your experiment.
- Reproduce a published evaluation protocol from a benchmark you can run on one GPU or CPU, and run the leakage checks from Part XI, Chapter 5 on it. Finding a leakage issue in a published protocol is a contribution.
- Write a transfer contract for a paper you admire and check whether the paper's claims exceed it.
This suggested sequence develops complementary skills; the steps are not strict prerequisites. Step 1 practices controlled comparisons. Step 2 compares learned and analytic beliefs. Step 3 connects model quality to measured planning quality. Step 4 practices evaluation audits, and step 5 practices claim audits. You can adapt the order to your background and available resources.
Engineering Readiness
The engineering roadmap specifies artifacts to build and test for a deployment contract. Producing them can involve established engineering practice or unresolved research, depending on the system:
- A tested deployment contract. Named observation and action spaces, units, frames, rate, latency budget, and the conditions under which the system refuses to act. The refusal path is part of the contract.
- A monitored prediction stream. Running prediction error and calibration checks against live data, with drift alarms and requirement-based error thresholds where appropriate (Part XII, Chapter 2). Stationary errors can still violate a deployment requirement; a drift check alone does not test that requirement.
- A fallback controller. A reactive policy that does not depend on the model, plus a tested switch that engages it. An unexercised backup remains an unvalidated fallback candidate, not established recovery evidence.
- Scoped continual correction. Bounded update rules with a defined trigger, a defined data budget, and a rollback path (Part XII, Chapter 5).
- Resource limits. Measured latency, memory and energy over declared workloads, plus a justified bound where a worst-case guarantee is required. A maximum observed in finite trials is not the worst case over all possible inputs (Part XII, Chapter 4).
- Provenance and reproducibility. Model version, data version, environment version, seeds, and hardware recorded with every evaluation. Missing provenance can make attribution or reproduction difficult or ambiguous; controlled replications and partial records may still identify a cause.
- A documented failure catalog. Known failure cases, their operating conditions and reproductions. This is a practical recommendation, not a measured ranking of team documents.
These are readiness requirements for the proposed deployment contract. Some also pose research questions, especially for high-dimensional or safety-critical systems. An unvalidated plan alone does not establish a deployment guarantee. State which conditional properties are supported by justified proofs and which need empirical validation. Exercise the fallback under its declared trigger conditions, and support resource limits with measurements and justified analytic bounds where applicable. Neither an unvalidated estimate nor the maximum of finite trials establishes a worst-case bound.
Research Projects That Seek New Evidence
Each of the four projects below is a proposed falsifiable design, not a completed study. Their illustrative budgets differ, and the order is thematic rather than a cost ranking. A single researcher should first reduce the task and model scope to a measured pilot budget before committing to the GPU-month or hardware estimates below.
Project 1: Decision-Aligned Evaluation Metrics
Question and hypothesis. Does a task-weighted prediction metric predict closed-loop control quality better than an unweighted aggregate, across a family of learned dynamics models? Hypothesis: a metric that weights prediction error by the sensitivity of the task cost to each state component will rank models much more consistently with closed-loop return than uniform mean squared error, and will do so across several tasks and model classes rather than in a single constructed example.
Prior evidence. The first experiment supplies a pairwise ranking discrepancy in a controlled setting. Lambert et al. (2020) study PETS dynamics models and planning in simulated cartpole and half-cheetah, comparing expert, on-policy and state-space data distributions. In these evaluated settings, one-step prediction likelihood is not always correlated with downstream control performance; likelihood is not interchangeable with aggregate mean squared error. MBPO examines model-generated-data bias and PETS incorporates uncertainty into model-based control. The proposed study asks which candidate metrics track return over its declared task population, without claiming that no related cross-task study exists.
Strongest feasible baseline. Uniform mean squared error on the observation, and variance-normalized mean squared error. Any proposed metric must beat both.
Training and held-out population. Train five to eight model classes (linear, MLP, recurrent state-space, small transformer, two latent-variable variants) on three task families (a positioning task, a contact-rich pushing task, and a simple navigation task) in simulation. Hold out entire task instances (different mass ranges, different floor friction, different layouts), not random transitions.
Intervention. Vary the metric, holding the model family, data, and training procedure fixed. Additionally, ablate the weighting scheme: sensitivity-derived weights, learned weights, and hand-tuned weights.
Success criterion. Before evaluation, declare a practically meaningful rank-correlation improvement margin and a paired uncertainty analysis for each baseline comparison. Specify independent sampling units, seed roles, the interval or test, power assumptions, and how the multiple comparisons will be handled. Compare the metrics on the same held-out predictions and outcomes. Count a task family as supporting the hypothesis only when the predeclared analysis supports an improvement beyond the margin against both baselines; require this in at least two of the three task families. This is a proposed acceptance rule, not a result of the toy experiment.
Failure criterion. None of the tested metrics reliably improves on the baselines under the specified uncertainty analysis. This constrains those metrics and this study population. Insufficient power, poorly estimated weights and properties omitted by the metrics remain possible explanations; it does not prove that every weighted loss is inadequate.
Budget. Illustrative planning allowance: one to two GPU-months of training, roughly environment steps per task family, and a few CPU-months of evaluation. These are not measured requirements. A pilot must specify hardware, throughput, repetitions and monetary cost. Generated data does not remove simulator, asset or code licensing obligations.
Expected artifact. A benchmark suite with held-out instance splits, a reference implementation of each metric, and a public table of metric-versus-return rank correlations.
Leakage, safety, and privacy risks. Reusing task instances to tune and report the weighting can overstate generalization. Separate tuning from reporting instances and audit preprocessing and selection for leakage. The proposed simulation uses no human data. Confirm asset licenses and bound compute costs; this protocol provides no physical-deployment safety evidence.
A null result limits claims about these metrics on the tested tasks, subject to statistical power. Follow-up studies can test uncertainty, action-effect and rollout metrics. It does not justify a field-wide conclusion that no scalar metric can be useful.
Project 2: Cross-Embodiment Transfer Under a Declared Interface
Question and hypothesis. How much of the benefit of multi-embodiment pretraining survives when the shared action interface is made explicit and the target uses a motion primitive absent from the source data? Hypothesis: shared observation encoders transfer substantially, action-effect prediction transfers only within clusters of embodiments with compatible action semantics, and negative transfer appears on the target's unseen primitive.
Prior evidence. Open X-Embodiment reports policy-transfer results from multi-robot co-training. Those results do not by themselves validate transfer of a learned dynamics interface. This project specifies an action-primitive holdout and evaluates dynamics prediction separately from policy success.
Strongest feasible baseline. A specialist model trained on target-embodiment data only, matched in parameter count and update count to the shared encoder plus adapters.
Training and held-out population. Source: a public multi-robot manipulation corpus. Target: a simulated arm with an action space that includes a primitive not present in the source (for instance, a compliant insertion primitive). Hold out the target's entire task family, not individual episodes.
Intervention. Train (a) target-only, (b) joint with a shared action interface, (c) pretrain plus frozen adapter, (d) pretrain plus fine-tuned adapter. Report action-effect prediction error separately from task success.
Success criterion. Predeclare the target-task criterion and interaction-budget comparison, a meaningful improvement margin, and a paired uncertainty analysis across the stated seeds and target-task units. Specify power assumptions and handling of comparisons across conditions. Require condition (c) or (d) to improve on (a) beyond that margin with fewer target interactions. Separately predeclare an action-effect-error tolerance and uncertainty rule for the unseen primitive; require that comparison to meet the stated noninferiority criterion. Neither a raw mean difference nor an unspecified seed-variability threshold supplies this rule.
Failure criterion. No condition beats the specialist, or the aggregate improves while the unseen primitive degrades beyond the specified uncertainty margin. The latter supports a negative-transfer finding for this protocol; publication value depends on the study's rigor and scope.
Budget. Illustrative allowance: several GPU-months for pretraining plus target fine-tuning, to be revised using a measured pilot. A simulated target avoids physical robot acquisition and operation costs, not compute-hardware costs. If a physical target is used, pilot the required robot-hours and obtain a safety review before committing to several hundred hours.
Expected artifact. A reproducible transfer benchmark with a declared interface, per-primitive error breakdowns, and adapter checkpoints.
Leakage, safety, and privacy risks. Audit the instances, categories and task families declared held out, including any source or pretraining overlap (Part VI, Chapter 4). Scene overlap violates a scene-novelty claim, but need not violate a task-family-only holdout. Document the actual split and available corpus-audit evidence. Physical deployment requires force limits and a tested fallback.
Failure to detect a transfer benefit constrains this source corpus, adapter, target population, and budget. It does not identify the cause. Follow-up ablations of action interfaces, encoder capacity, optimization, and source coverage are needed before assigning the bottleneck.
Project 3: Interventional Evidence for Learned Structure
Question and hypothesis. When a model is trained on observational trajectories from an environment with known causal structure, how well does it predict a designated object's intervention response? Observational loss alone does not generally identify that response without resolving assumptions; randomization, valid adjustment or justified structural restrictions can change this. For a declared model and intervention population, test the hypothesis that observational score differences weakly predict intervention performance. The loss-matched comparison below separately asks whether architectures with similar observational scores differ under intervention.
Prior evidence. Schölkopf et al. survey the causal-representation research program; Ahuja et al. give conditional identification results and interventional learning procedures. This project seeks a controlled comparison of learned control models with matched observational loss. The cited work motivates that comparison but does not establish a literature-wide absence of similar studies.
Strongest feasible baseline. Observational-loss-matched models with different architectures: an object-centric slot model, a plain latent state-space model, and a pixel-level predictor. All tuned to comparable held-out prediction loss.
Training and held-out population. A simulated scene with controllable objects. Train on trajectories from a fixed set of action policies. Hold out object configurations that require understanding an interaction the policies never produced (for instance, stacking after training only on pushing).
Intervention. Execute a controlled intervention (apply a known impulse to object ) in the simulator, and compare the model's predicted post-intervention trajectory against the realized one. Repeat for each object and each of several impulse magnitudes.
Success criterion. At least one architecture achieves interventional error low enough to support planning (concretely: MPC using the model reaches the post-intervention target at a rate within a predeclared noninferiority margin of the planner with simulator access, using a confidence interval with adequate power), while an observational-loss-matched baseline does not.
Failure criterion. No architecture meets the predeclared planning threshold, or any observed architecture advantage is smaller than the study can resolve. Matched interventional errors alone do not distinguish uniformly good from uniformly bad models. Tight observational/interventional correlation would challenge the proposed weak-association hypothesis in this population, not support it.
Budget. Single-GPU scale for a simple 2D or low-resolution 3D scene; one to two GPU-months.
Expected artifact. An intervention test suite, plus a table of observational loss against interventional error for each architecture.
Leakage, safety, and privacy risks. The intervention must be applied in the environment, not in the model, or the test is circular. The held-out configurations must not appear in any training data, including pretraining data if any is used.
If observational loss tracks interventional error in this study, report that association within the tested model families, interventions and task population. It does not remove theoretical non-identifiability or establish the same association outside those conditions.
Project 4: Memory Under Changing Context at Deployment Scale
Question and hypothesis. In a deployment with an unknown, slowly changing context, does a memory system with an explicit change model outperform both a memoryless system and a memory system with a fixed assumed context? Hypothesis: adaptive memory lowers post-change recovery cost under the declared change population, while a misspecified hazard can erase that advantage or make memory worse than the memoryless baseline. The performance comparisons below test this scoped hypothesis; they do not establish that every benefit is causally mediated by detection accuracy.
Prior evidence. The third experiment demonstrates stale-context behavior in a two-state toy. Kirkpatrick et al. study parametric-network forgetting in sequential permuted-MNIST classification and Atari learning, comparing elastic weight consolidation with their stated baselines. Dohare et al., Nature, 2024 study loss of plasticity in Continual ImageNet, class-incremental CIFAR-100 and simulated ant locomotion with PPO, including changing-friction and stationary conditions. These evaluated learning failures are distinct from a stale deployment belief. The proposed project tests its declared memory/change-model comparison; those citations do not establish that comparable control studies are absent.
Strongest feasible baseline. Memoryless with matched parameters, and memory with a fixed context (the "frozen" agent in the toy, scaled up).
Training and held-out population. A simulated or physical manipulation task with changing payload, friction, or camera calibration. Evaluate on change events the system has never seen: changes in a parameter outside the training range of the change model.
Intervention. Vary the hazard rate and the change-detection signal (reward prediction error, prediction error on the observation, an explicit change classifier). Include a version with no update.
Success criterion. Predeclare the post-change cost window, unseen-event population, practically meaningful improvement margins, and a powered paired uncertainty analysis over the stated runs and events. Specify seed roles, sampling units, and handling of the two baseline comparisons. Require the analysis to support lower recovery cost beyond the declared margin against both baselines. This is an acceptance rule for the proposed study, not a guarantee for every event.
Failure criterion. The adaptive memory is not better, or the ranking depends on the change magnitude in a way that reverses. A reversal would show that this tested update rule is not uniformly best across these changes; it would not prove that no single hazard can be safe.
Budget. Simulated: one to two GPU-months. Physical: several hundred robot-hours plus a safety review and a tested fallback controller.
Expected artifact. A change-detection benchmark with unseen change events, plus recovery-cost curves as a function of change magnitude.
Leakage, safety, and privacy risks. Any memory that stores episodes must have a defined retention policy and a reset boundary to avoid carrying personal or stale data across sessions (Part XII, Chapter 3). Physical deployments must bound force and speed during the recovery window, when the belief is wrong.
Failure to detect an adaptive-memory advantage constrains this update rule and test population. Detection quality, memory representation, control policy, change frequency, and statistical power remain possible explanations; ablate them separately before naming a bottleneck.
Comparing the Four Projects
| Project | Core comparison | Held-out unit | Proposed core measurements | Main threat to validity |
|---|---|---|---|---|
| 1. Decision-aligned metrics | Weighted vs. uniform prediction loss | Task instances | Rank correlation with closed-loop return | Tuning the weighting on the test instances |
| 2. Cross-embodiment transfer | Pretrained + adapter vs. specialist | Task families and action primitives | Unseen-primitive error, target success and target-data learning curves | Source-corpus contamination |
| 3. Interventional structure | Observational-loss-matched architectures | Object configurations requiring unseen interactions | Intervention error and powered MPC noninferiority comparison | Circular evaluation in the model |
| 4. Memory under change | Adaptive vs. frozen vs. none | Unseen change events | Post-change recovery cost | Change-detection signal leakage from the evaluation |
The table is a summary, not a substitute for the specifications or measured pilot costs. Project 1 scores fixed predictions on the same data using different metric functions; their ranking differences are conditional on that metric comparison. Project 2 compares whole transfer procedures, including pretraining data, update budgets and adapters. Interface-specific attribution needs a matched interface ablation. Project 3 matches a selected observational score but does not control every training or representation confound. Project 4 tests change-model transfer on unseen events under its stated protocol. Held-out units define the generalization target and limit overlap; matched interventions, budgets and procedures address other confounds. Generalization claims need a specified evaluation population and appropriate splits. Descriptive or randomized within-population studies can still be interpretable without a train/test holdout.
Limitations and Impact
The three experiments are hand-constructed and low-dimensional. Experiment 1 uses a known stable discrete transition law with a small drag term and an autonomous nuisance process. Experiment 2 uses a known transition family with a known coefficient. Experiment 3 uses two known modes and a heuristic assumed hazard, tested against stable and deterministic-change conditions. Known truth permits controlled comparisons, but the numerical performance results do not establish physical-system performance. In experiment 1 the aggregate metric changes yet orders two models oppositely to controller cost; it does not establish that decision consequence by itself. The examples demonstrate scoped failure mechanisms, not the magnitude of their effects on real systems.
These examples motivate requirements for the proposed studies, not a diagnosis of what limits the whole field:
- General-purpose worlds need specified action semantics and held-out evaluation populations for the transfer being claimed. A single canonical interface is one design option, not a necessary architecture. Open X-Embodiment provides bounded policy-transfer evidence across its specified robot datasets; dynamics-model transfer needs its own evaluation.
- Causal, compositional, persistent structure calls for appropriate identification assumptions, informative tests and change-aware update rules. Model capacity, optimization, data coverage and experimental design can all matter. Experiment 2 supplies one tractable combination split; project 3 proposes an interventional comparison whose implementation cost must be measured.
- Evaluation beyond visual prediction needs measurements appropriate to the decisions being claimed. One scalar can summarize a specified objective, but it does not establish every aspect of usefulness or safety. Experiment 1 separates aggregate observation error from closed-loop cost; report coordinate-wise and control measurements to explain that discrepancy.
An unspecified universal-model claim leaves the evaluation population and criterion unclear. A transfer contract makes comparisons interpretable. Identification, calibration, change detection and fallback design deserve explicit testing alongside capability development; an impressive demonstration alone does not establish deployment reliability.
The methods in Part VIII use learned models in different ways. PETS combines probabilistic ensembles with trajectory sampling for control. MBPO trains with short model-generated rollouts branched from real data. DreamerV3 improves behavior using imagined futures. MuZero uses tree search with a model predicting planning-relevant rewards, policies and values. TD-MPC2 performs local trajectory optimization in a learned latent model. These are not all instances of policy training in imagination. Their different model interfaces also require different validation tests.
Verification deserves explicit effort alongside capability development. These experiments do not establish the field's binding constraint or the feasibility of every proposed project for one researcher. They provide runnable examples and study specifications that can be reduced to pilot experiments. A useful pilot measures its actual computational cost, checks its identification assumptions and estimates what effect size its evidence can resolve.
The external claims in this chapter are drawn from primary papers, official technical reports, and benchmark documentation, cited inline where they appear. Reported empirical results are stated with their evaluation settings where the source states them, and are described as of the source's publication date. Statements about what is not yet established are the author's assessment of the literature at the time of writing, not a survey result. The three experiments are synthetic and their figures and numbers are generated by the code in this chapter; they are illustrative controlled examples and must not be read as benchmark results or as evidence about physical hardware. Executing these cells does not verify external claims or cross-chapter references; those require separate source checks and reconciliation against the canonical book structure.
Summary
Three questions, three experiments and four proposed projects connect capability claims to specified evidence. The examples show why prediction, intervention, control and recovery measurements need separate interpretations; they do not rank the field's unresolved constraints.
Generality has to be written as a contract. State the source and target systems, query semantics, observation and action spaces, held-out population, adaptation budget and criterion. Shared interfaces need demonstrated competence. Specialists with matched target data and budgets are informative baselines. Ambiguous action mappings, unobserved context, capacity constraints and changed calibration are possible transfer failures, not an exhaustive or ranked list of causes.
Structure has to be earned. Observational data alone does not generally identify interventions or counterfactuals without additional assumptions. Suitable adjustment, randomization, or structural restrictions can change what is identifiable. Causal claims require stated assumptions and sufficient identifying information. Intervention variation is needed by some approaches, while others can identify effects from observational data under justified assumptions; independent intervention tests can check the resulting predictions. A disentangled latent or a named object slot alone does not establish a causal effect. For empirical unseen-combination performance, declare the held-out split; when interpreting recovered factor effects, check identification under the stated model family and assumptions. An assumption-violation condition separately probes the limits of a matched inductive bias. Persistent memory includes several mechanisms. In the third experiment's goal-flip condition, frozen cue memory has higher mean cost than the memoryless baseline; that ranking is conditional on this task and controller.
Evaluation has to reach the decision. One-step loss, rollout error, calibration, interventional prediction, action ranking, closed-loop return, safety outcomes, latency and adaptation cost answer different questions. A measurement of one does not guarantee another. The evaluator and evaluation population can depend on the model, so use splits at the trajectory or task level, audit for leakage, and do not treat seeds as independent systems.
The three experiments, restated as conditional examples to investigate on your own models:
- In experiment 1, the two-component observation predictor has lower aggregate observation error but higher controller cost than the state-focused predictor. Report both coordinate-specific error and true-environment decisions.
- In experiment 2, the product-feature model has lower held-out error under the multiplicative law, but not under the capped law. Report seen and held-out errors for both feature families and regimes.
- In experiment 3, fixed cue memory improves on the memoryless baseline in the stable condition, but has higher cost in the tested change and corrupted-cue conditions. Report both stable and changed conditions.
Reproduce the experiments and test conditions that challenge their assumptions. Write transfer contracts for the papers you read and check whether their claims stay within them. Build the deployment artifacts, especially the failure catalog and tested fallback. For a proposed study, declare the evaluation population and failure criterion before collecting data. A rigorous null can constrain a stated hypothesis; an unfalsifiable positive claim cannot supply that test.
A well-powered null result constrains the hypothesis tested under the declared protocol; it does not prove that no effect exists. The four projects should make either outcome informative. Declare the comparison and failure criterion, run the study, and use the evidence to decide what to test next.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about frontier questions in world-model research.
Frontier Questions and Research Roadmaps
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!