Part of World Models Handbook
Benchmark selection, leakage checks, ablations, and compute reporting connect world-model claims to reproducible experimental evidence.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Datasets, Benchmarks, and Experimental Design
Consider a hypothetical robotics study. A group logs four hours of teleoperation from one arm, fits a small next-joint-state predictor, and randomly assigns 70% of the transitions to training and 30% to testing. Suppose it reports a mean absolute error of 0.004 radians. Its draft claims that the model transfers to new collection sessions. A separate, stronger claim would be transfer to an arm it has never seen. The hours, split, and error in this opening are illustrative, not results from a reported experiment.
A reviewer asks whether the test contained a new session. It did not; nor did it contain a new arm. Both partitions came from the same logged collection. A payload or calibration offset shared within that collection can create reusable context, but sharing a constant does not make every state and action pair correlated. Nearby training transitions may help a model infer that context; the split alone does not prove that it did. Even if the illustrative 0.004 error were measured correctly, it would describe performance on the chosen held-out rows. It would not by itself establish transfer to a new session. It would also not establish transfer to a new arm.
That gap, between the question a number answers and the question a claim asks, is the subject of this chapter.
The World-Model Evaluation Ladder organized five complementary protocols: one-step prediction, open-loop rollout, intervention-response comparison, closed-loop decision quality, and transfer. The three following chapters examined perceptual and generative fidelity; state, physics, causality, and memory; and planning, control, and policy evaluation. They supplied instruments for different questions, not interchangeable certificates. This chapter asks whether the experiment producing a number supports the sentence you want to write about it.
That is a different skill, and it is not learned by reading more metrics. It is learned by interrogating the design: which unit was held out, what was fit on which rows, how many independent replicates exist, what the interval means, and what compute was spent. A perfectly reasonable metric applied to a badly designed comparison produces a perfectly precise number attached to a claim it cannot support. We will see several concrete versions of that failure in this chapter.
In most empirical sciences, a number is only as good as the design that produced it, and world-model research compounds this because the objects being measured (learned dynamics, learned abstractions of physical structure, planners that query the model thousands of times per decision) are themselves complicated artifacts whose "correctness" depends on the deployment regime. A model that generalizes perfectly within one lab's robot may fail catastrophically on the same robot after a firmware update, and a model that fails to generalize at all may still be the best available option if the deployment target is exactly the collection regime. Neither of those statements can be evaluated by staring at a loss curve. They require naming the estimand, selecting the units and splits that isolate it, and reporting the resulting numbers in a form a skeptical reader can check.
We keep the notation from Planning, Control, and Policy Evaluation: is the true state, the observation, the action, the true transition function, a learned dynamics model, the current-step reward, a policy, the episode length, the planning horizon, the number of candidate action sequences per decision, and the finite-episode return. A few new symbols appear here:
- indexes a session, a collection grouping unit. A session can be one teleoperation run or one logging period. Different sessions may still share a robot, operator, or scene; group disjointness does not by itself establish independence.
- is a persistent, session-constant quantity (an offset, a calibration, a payload) that the learner is not given.
- is the number of transitions and the number of lagged states used as features, so a feature matrix has shape once the action column is appended.
- is the number of independent replicates, a paired difference, its standard deviation, and a target difference worth detecting.
The dynamics toy observes the physical scalar directly: . Its session constant is not observed. The complete Markov state is therefore , with unchanged during an episode. Relative to that complete state, the learner is partially observed even though there is no visual perception problem. History can supply information about the persistent context, so context inference and belief over unobserved state are not categorically separate problems. This toy isolates one static-context setting; it does not evaluate general perceptual robustness or memory architectures. Contextual Markov Decision Processes provide a related formal treatment of episode-level context.
The chapter proceeds in four movements. First, what a benchmark is, and why counting benchmarks is a poor substitute for mapping claims to required coverage. Second, leakage and shift: how to name the unit you are holding out, and what happens when you do not. Third, the statistics of comparison: replication, pairing, intervals, and prospective power. Fourth, reporting: a study protocol, and honest compute accounting. The recurring question is always the same. Does this comparison support the stated claim?
Benchmark Taxonomies and Coverage
We will distinguish five objects that can otherwise be conflated in a result table. These are working definitions for reading the experiment, not a claim that every research community uses the words identically.
An offline transition dataset is recorded evidence: a finite collection of tuples such as or , with optional rewards and episode-boundary flags. Other datasets may contain images, scenarios, or judgments instead of transitions. A fixed log contains the behaviors recorded by its collection process; it does not allow a new policy to request a new interaction. Without observed or reliably reconstructed action information, it cannot directly test the model's response to specified actions. Absence of an action column does not prove that actions are unrecoverable. Identifying an intervention effect from passive observations requires additional causal assumptions or structure; predictive fit alone does not establish that identification.
An environment is an interactive process: something you can step with an action and receive a new observation, reward, and an episode-boundary signal where relevant. It may be a physics simulator, a game engine, a browser sandbox, or the physical world. Underlying dynamics alone need not define an objective, although a packaged environment interface may already include a task's reward and termination rules.
A task specifies an objective and the conditions under which it is evaluated, including an initial-state distribution. An episodic task has termination rules; a continuing task need not terminate intrinsically, but its evaluation still needs a declared horizon or long-run objective. A time-limit truncation is not necessarily the environment's terminal state. One environment can host several tasks with different objectives or evaluation conditions.
A benchmark is a bundle: a set of tasks, plus the protocol for using them. The protocol includes the split, metrics, aggregation rule, resource budgets, and reporting conventions. Declaring those choices makes the intended comparison explicit.
A leaderboard is a selected summary of a benchmark: usually one number per submission, often the mean over tasks and runs, sometimes a trimmed or normalized variant. It is a compression of a much richer object, and the compression is chosen by whoever runs the leaderboard.
The practical consequence: an interactive control suite is not automatically a released offline dataset, and a released offline dataset is not automatically an interactive benchmark. If your claim is about learning from logged data without further interaction, you need the dataset and its collection policy. If your claim is about closed-loop control, you need the environment and the protocol that governs how many interactions each method gets.
The five definitions suggest questions to ask before opening a result table:
- What was the input to fitting, and can it be inspected?
- Is the model allowed to query an environment, and under what budget?
- What objective and termination define success?
- How was the split chosen?
- How was the metric chosen?
- How was the aggregation rule chosen?
- What was compressed away to produce the leaderboard number?
Trace the answers to the experiment's artifacts and protocol. An undisclosed choice can leave the claim's support uncertain; an identified mismatch can show that the number answers a different question.
A subtlety that trips people up: a dataset and an environment can be produced by the same team, packaged in the same release, and used together, and they can still be conceptually distinct objects. The dataset is a snapshot; the environment is a generative process. A method that performs well when training on the snapshot has not necessarily demonstrated anything about the process, and a method that performs well when interacting with the process has not necessarily demonstrated anything about the snapshot. Blurring the two can cause a claim to exceed what the measured object supports.
Dimensions of coverage
Before naming any particular suite, it helps to have a vocabulary for what a benchmark can and cannot cover. These dimensions are the axes along which a claim either is or is not supported.
- Observation modality. Proprietary state vectors, RGB, depth, tactile, proprioception, language, or mixtures. A model evaluated only on state vectors has not been tested on pixels.
- State accessibility. Fully observed, partially observed with a known observation model, or partially observed with an unknown one. Observation-model knowledge changes available inference assumptions; it does not require a different architecture in every case.
- Action coverage and collection policy. Which actions were taken, by whom, at what stage of training, with how much exploration, and under what safety constraints. These choices determine the observed support for action-conditioned evaluation.
- Time horizon. How many steps before the relevant event occurs. A model that is accurate for 10 steps may be useless at 500.
- Stochasticity. Deterministic dynamics, small process noise, or multimodal outcomes. A deterministic transition test does not test whether a model represents stochastic or multimodal transition uncertainty; it need not favor every deterministic model under every metric.
- Contact and physical constraints. Free-space motion versus contact-rich manipulation with friction, impacts, and non-smooth dynamics.
- Partial observability. Which decision-relevant state is unavailable to the agent. A belief or another history summary may be needed; some tasks can be solved without explicitly representing every hidden quantity.
- Task diversity. How many distinct objectives exist, and how different they are from one another.
- Out-of-distribution axes. Which factors vary between training and test: layout, appearance, object set, physics parameters, embodiment, sensor calibration, or task structure.
- Measurement tier. Whether the reported number is a prediction error, a state-estimation error, a planning success rate, or a deployment outcome on hardware. Earlier chapters built this ladder. A benchmark directly tests the tiers it measures; inference to another tier requires an explicit, separately justified structural or identification argument.
These dimensions can interact. The time resolution and event-recording scheme determine which contact events a test exposes, while a short rollout may not reach contact at all. Diverse tasks can differ in difficulty and reward scale; their sampling weights affect aggregation. The dimensions are therefore not independent checklist items. State how each relevant choice affects the event, observation, or population being measured.
Coverage utility is claim-relative: two well-tested dimensions can be enough for a narrow claim, while nine irrelevant ones may not help it. Ask whether a benchmark tests the factors your claim requires. A famous suite can be unsuitable for a specific question, and a less familiar one can be suitable. Reputation and claim fitness are different properties; this is not a claim of statistical independence between them.
Reading four well-known suites narrowly
It is worth looking at four widely used suites and stating, precisely, what question each was built to ask. None of them is a universal world-model benchmark, and describing them that way is a category error.
DeepMind Control Suite (Tassa et al., 2018) is a set of continuous-control tasks built on a physics simulator, with a standardized interface and reward terms that are individually interpretable. Its generally continuing, infinite-horizon tasks are distinct from the fixed 1,000-step evaluation episodes used in the original paper: that evaluation limit truncates the task rather than defining an intrinsic terminal state. Its original question was to give continuous control a consistent, well-specified task set so that algorithm progress could be compared and reward structure could be inspected rather than guessed. Its limitation is equally clear: it is an environment suite, not a released offline dataset with documented collection policies. Strong performance on its tasks is not by itself evidence of action-conditioned fidelity under a fixed logged dataset, nor evidence of generalization to unseen tasks or embodiments.
D4RL (Fu et al., 2020) is a collection of offline datasets whose defining feature is collection-policy diversity: demonstrator data, medium-quality data, mixtures, replay buffers, random data, and multitask datasets. Its original question was to make the collection regime part of the benchmark rather than an unstated detail. The limitation follows directly: a result on a D4RL dataset is a statement about that collection policy. Improving on "medium" data does not imply improving on expert data, and vice versa, because the data distributions differ in ways that change which algorithmic ideas help.
There is also a provenance caution worth internalizing. The Farama Minari D4RL documentation states that its reproduced group is not identical to all original D4RL datasets, even though it was generated using the same principles. Two artifacts with the same name and similar statistics can differ in details that matter. Version and provenance are part of the result, not metadata.
The D4RL note generalizes to any dataset with a shared name and a sharded lineage. When you cite a result on "D4RL halfcheetah-medium," you are citing a specific artifact with a specific history, and if you did not generate it yourself you are relying on the artifact being what its name suggests. This is why content hashes and explicit version identifiers matter: they turn a name into a specific, checkable object. It is a small piece of discipline that pays off every time someone tries to reproduce a number and finds that the artifact they downloaded produces a different curve.
Procgen (Cobbe et al., 2019) consists of sixteen procedurally generated game-like environments, designed to measure sample efficiency and generalization across levels. Its original question was whether a policy trained on some levels transfers to unseen levels of the same game. Those levels can vary layouts, appearances, entities, and event timing. Standard unseen levels drawn from the same generator are a held-out-level generalization test, not automatically an out-of-distribution test. A claim about changed physics, embodiment, or a new task family needs a protocol that varies those factors.
Physion (Bear et al., 2021) presents simulated physical scenarios and compares machine predictions to human predictions, including on scenarios designed to probe specific physical concepts. Its original question was how well models and people predict physical outcomes from visual scenes. The limitation is that this is a passive prediction protocol. The agent is not choosing actions to achieve a goal, so a Physion result is not automatically evidence about action selection, planning, or control.
These examples show why a suite's intended question matters. A strong result can be quoted as though it answers a different question that the protocol never measured. That is a category-drift mechanism to check for, not a measured claim about how frequently papers overclaim or evidence of fraud.
A claim-to-test coverage matrix
Benchmark names identify suites but do not fully specify which factors a particular experiment tested. A requirements matrix makes that scope explicit. It has claims as rows and coverage requirements as columns. To support one of these scoped empirical claims, the benchmark evidence must address every factor designated as necessary for that claim; filling required cells is not by itself a correctness guarantee.
| Claim | Action-conditioned dynamics | Hidden state | Contact / constraints | Horizon over 100 steps | Stochastic transitions | Unseen sessions or embodiment |
|---|---|---|---|---|---|---|
| " predicts noisy next-state transitions within known, fully observed sessions" | yes | no | no | no | yes | no |
| " predicts noisy dynamics on new, fully observed sessions of the same robot" | yes | no | no | no | yes | yes |
| " predicts noisy contact dynamics on an unseen robot with changed actuation limits" | yes | no | yes | no | yes | yes |
| " supports new-robot control with hidden state, contact, stochastic dynamics, and horizons over 100 steps" | yes | yes | yes | yes | yes | yes |
The last row deliberately names a deployment involving every listed factor. Closed-loop operation alone does not require hidden state, contact, stochasticity, or a horizon over 100 steps. Likewise, contact is required in the third row because its claim names contact, not merely because actuation limits change. Deployment latency is a separate hardware-and-software measurement, absent from this qualitative matrix. Adding benchmarks that duplicate the same axes does not add missing axes; a benchmark that varies a new factor can. Coverage is the union of the tested factors, not the count of suite names.
That union is only a first check. A claim about a combination of factors needs evidence on the relevant joint setting, or explicit assumptions justifying transfer from separate tests. Contact tests on familiar robots plus free-space tests on new robots do not alone test contact on a new robot. Filling every required column with separate marginal tests is not a certificate for their combination.
The coverage table is a small binary matrix. A heatmap makes the missing columns visible at a glance.

The requirements matrix is also useful alongside a second matrix of actual measurements. Compare the paper's stated requirements with the factors its experimental section says it tested. A required factor absent from those measurements is a coverage gap. The binary requirements heatmap above does not itself contain benchmark results and cannot certify any of the four claims.
One well-designed relevant test can be more informative for a claim than five unspecified or irrelevant suite names. Neither one nor five is a universal benchmark-count requirement.
Aggregating across tasks without fooling yourself
Once a benchmark has multiple tasks, you need an aggregation rule that defines what "performance on this benchmark" means. Different defensible rules can rank the same methods differently. Choosing a favorable rule after inspecting test results is adaptive analysis selection, even if neither fitted model changes. Label that choice exploratory or evaluate the selected analysis on fresh data.
Define the task distribution first. A claim about a fixed suite concerns that finite set of tasks. A claim about a task population also needs a specified population and sampling design. More variants of one mechanism do not imply broader mechanism coverage: eight pushing tasks and two cloth tasks assign pushing 80% of the weight in an equal-per-task average. That does not guarantee 80% of its numerical contribution when score scales differ.
Fixed-suite inference versus sampled-task inference. For a fixed suite, uncertainty can come from training runs and evaluation episodes while task identities are held fixed. A claim about a sampled task population adds task-sampling uncertainty. Its magnitude depends on task variability and the sampling design; it need not be the largest component. Treating sampled tasks as fixed omits that component and can understate uncertainty. Replicating more runs on the same tasks does not supply more independent task draws.
Average versus tail. A task mean rewards aggregate score, not necessarily competence on every task: narrow excellence can compensate for poor outcomes elsewhere. A minimum or the fraction below a declared threshold exposes different failure behavior. If deployment assigns substantial cost to rare failures, report a matching tail measure alongside the mean. The appropriate choice follows the objective rather than a preferred ranking.
Reward-scale sensitivity. Averaging raw returns across tasks with very different numerical scales gives the larger-scale tasks more influence for a comparable proportional change. A scale difference alone does not prove that a task dominates every comparison: the actual method differences and task weights also matter.
Normalized scores. Tasks with very different return magnitudes contribute unevenly to a raw mean, so a single large-scale task can dominate the aggregate. The common fix rescales each task's raw return onto a comparable scale. The normalized score is:
where:
- : the method's return on the task
- : a reference anchor (often a random or weak policy), used as the low end of the scale
- : a strong reference return, used as the high end of the scale
The score expresses how far the method's return sits above the weak anchor, as a fraction of the gap between the weak and strong anchors. Three things must be stated explicitly when you use this.
- What the anchors are. A fixed positive anchor gap preserves method ordering within one task. Different positive gaps across tasks can change an aggregate ranking by changing effective task weights. A negative gap reverses the within-task direction. Disclose the anchors and weighting convention.
- What happens when the denominator is zero or negative. A zero gap makes this formula undefined. A near-zero gap amplifies a fixed nonzero numerator and makes the score sensitive; it does not make every score large. Specify exclusions, clipping, or an alternative rank summary before examining favorable rankings. Clipping changes the reported statistic and can hide tails.
- Why it is not a probability. A normalized score can exceed 1 (a method can beat the "expert" reference) and can be negative (a method can be worse than random). It is not bounded in and must not be described as if it were.
Because the score can exceed 1 or fall below 0, plots drawn as if it were a probability are misleading. The distribution it summarizes is over return differences, not over event frequencies. When a figure shows box plots of normalized scores on a axis, the reader is being nudged toward an interpretation that the numbers do not support. A figure that explicitly allows scores to range past 0 and 1, with labeled reference lines at the anchors, is more honest.
A small synthetic distribution shows the point. The scores below are illustrative; they are not from the toy dynamics above.

Interquartile mean (IQM). Sort the run-level scores across all (task, run) pairs and trim each tail. The following finite-sample convention discards whole observations from each end, then averages the rest. It retains exactly the middle 50% when is divisible by four; otherwise it retains slightly more. State this rounding convention rather than assuming every implementation uses it. If denotes the -th smallest score among scores, this trimmed estimate is:
where:
- : the total number of (task, run) scores being pooled
- : the -th smallest score after sorting
- : the number of scores dropped from each tail
For , this convention trims at least one observation from each tail and is less sensitive than the ordinary mean to sufficiently extreme trimmed observations. For , it trims nothing and equals the ordinary mean. It is not invariant to every change affecting fewer than 25% of pairs: changing a retained central score changes the estimate. Nor does trimming repair systematic bias. With an appropriate sampling design, the statistic can estimate a population trimmed mean; it does not itself define that population or establish uncertainty. Pooling pairs weights tasks by their run counts. Equal counts give equal task weights; unequal counts need an explicit weighting or sampling rule if equal representation is intended. Trimming can downweight a tail-focused effect. Select the summary for the scientific question before examining which method wins.
Score distributions and performance profiles. First define the empirical cumulative distribution function (ECDF):
where:
- : the number of (task, run) pairs in the pooled set
- : the normalized score of the -th pair
- : the indicator that is 1 when its argument is true and 0 otherwise
At , this ECDF gives the fraction of pooled normalized scores at or below 0.5, not half the raw expert return. If the strong reference return exceeds the weak reference return, this counts raw returns at or below their midpoint. If their ordering is reversed, it counts raw returns at or above their midpoint, because normalization reverses the inequality. Equal anchors leave the normalization undefined. The run-score performance profile in Agarwal et al. uses the complementary tail, : the fraction strictly above the normalized threshold. It decreases as the threshold increases, and a higher curve means a greater fraction of scores exceeds that threshold. Label the convention explicitly. Curves can reveal lower-tail differences that a shared mean conceals; a deployment choice also needs the relevant task distribution, costs, and risk constraints.
Agarwal et al. (NeurIPS 2021) argued that finite training runs make aggregate estimates uncertain, and proposed interval estimates, robust aggregation, and performance profiles rather than relying on a single mean. That is a methodological argument, not a required seed count. IQM and profiles do not change the held-out unit: aggregating same-session errors does not by itself turn them into empirical evidence about transfer to new sessions. Same-session performance can still be a valid target in its own right.
An aggregation rule maps a collection of per-(task, run) scores to a summary number. Every aggregation rule embeds a weighting over tasks, a treatment of tails, and an assumption about what the collection represents. Reporting the rule, the task list, and the per-run raw scores is what makes the summary auditable.
Train-Test Leakage and Dataset Shift
In the hypothetical opening, the draft explicitly claimed transfer to new sessions while its split evaluated rows from the same logged collection. The mistake is a mismatch between that transfer claim and the estimand, the quantity the reported number estimates. Mean absolute error can be useful; a suitable metric does not make an unsuitable held-out unit answer the intended question.
Split roles
Three roles need to be distinguished before any split is drawn, because they have different rules.
- Training. Rows used to fit parameters and data-adaptive preprocessing for this experiment. Fixed transforms and independently pretrained representations need separate provenance disclosure; they are not necessarily refitted on these rows.
- Validation (model selection). Rows used to choose hyperparameters, architectures, checkpoints, early stopping, or analysis rules. They are no longer an untouched test of the selected procedure. Their descriptive scores remain meaningful; subsequent inference requires new data or a method that accounts for adaptivity.
- Final test. Rows evaluated after model and analysis choices are frozen. If their results guide a new hyperparameter, checkpoint, or metric choice, they have participated in selection. Favorable selection on noisy results can produce optimistic performance estimates; the untouched-test interpretation no longer holds.
For example, choosing the best of five observed seeds and reporting it as representative hides the selection. Unusual outcomes are not by themselves grounds for excluding a run. Prespecified validity criteria or independently confirmed corrupt execution can justify exclusion, but report every attempted run, reasons, and sensitivity to the exclusions. Ordinary training failures that belong to the procedure being evaluated should not disappear from that procedure's estimand. The choice of five in this example is illustrative, not a required run count.
A good habit is to make the report include every seed you ran, even the ones you did not like, with a short reason for each if any were excluded. When the reader sees all five seeds and where they landed, they can do the analysis you would do if you were reviewing yourself. When the reader sees one seed and a rationale about outliers, they can only trust or distrust, and trust is not the currency of empirical claims.
What is the independent unit?
Before you can hold anything out, you must decide what the independent unit is. Candidate units in world-model research include:
- transitions
- episodes or trajectories
- collection sessions (one teleoperation run, one logging day)
- subjects, patients, or participants
- objects, tools, or manipulated items
- robots, vehicles, or embodiments
- scenes, rooms, or maps
- procedural levels or game seeds
- tasks or objectives
- dates, sites, or institutions
Each candidate corresponds to a different claim. Holding out transitions supports claims about interpolation. Holding out trajectories supports claims about new trajectories. Holding out subjects supports claims about new people. The unit is not a technical detail you choose after the fact; it is the formalization of which thing the claim is about. Two groups can run "the same" experiment with the same data and get different numbers, not because one is right and one is wrong, but because they held out different units.
Here is the rule that governs everything else in this section:
The unit you hold out must match the generalization claim.
For direct empirical tests, hold out transitions for the declared within-session target, sessions for new-session transfer, and robots for new-robot transfer. A structural argument for transporting evidence is a different route and needs its own assumptions. No split label supplies that argument automatically.
Random transition partitions are not always wrong
A random transition partition can estimate a specified within-session interpolation target when its information access and joint dependence match the intended use. Matching the marginal distribution of individual rows is not enough: deployment on a fresh session can have the same row marginals but independent session context. The design must also distinguish retrospective interpolation from forward forecasting. Reporting a random-split result with that scope is not evidence of misconduct.
It becomes misleading when it is used to support claims about a different estimand: new trajectories, new sessions, new embodiments, or new environments. The mechanisms that make it misleading are specific and worth naming:
- Shared latent context. A persistent session-level quantity (calibration, payload, operator style, scene lighting) is present in both partitions. A sufficiently flexible model can exploit it.
- Overlapping history windows. Adjacent windows containing consecutive state observations share of those observations. The fraction of all input features shared depends on and any additional features, such as actions. A test row may therefore share observations with a training row.
- Near duplicates. A controller holding a pose or a looping policy can produce repeated or near-identical states. Check their presence in the actual log rather than assuming a prevalence.
- Future information. Whole-trajectory normalization, smoothing, or embeddings can depend on observations unavailable at the prediction moment. A fixed known constant need not; trace each transform's inputs and time boundary.
- Common data sources. Two sessions may share a base model, a base scene, or a base policy, so holding one out does not hold out the underlying source.
Shared latent context can be missed by checks that only compare row IDs. A correctly implemented row split can leave training and test rows from the same session. If session context affects the observed relation, the model has an opportunity to exploit it; the split alone does not prove it did. Whether that opportunity is legitimate depends on the claim. It can be appropriate for within-session interpolation and insufficient for new-session transfer.
Temporal dependence alone is not proof of label leakage, and time indexing alone does not force autocorrelation. Suppose scalar losses are covariance-stationary, with finite positive variance and lag correlation . Expanding the variance of their average gives
Here is the number of observed losses, and is the correlation between losses positions apart. There are pairs at lag , which produces the factor . When the bracket is positive, an effective sample size relative to independent equal-variance losses is divided by that bracket. Positive aggregate lag correlation reduces this effective size; a negative aggregate can increase it. Treating positively dependent losses as independent can therefore make intervals too narrow. The formula concerns the loss sequence, rather than the correlation of raw states.
Leakage is a different question: did prediction-unavailable information, or information excluded by the declared held-out-unit protocol, enter fitting or prediction? Trace feature times and fitting provenance for that check. Use dependence-aware inference for correlated outcomes. Grouped or chronological splits can enforce a target boundary, but do not by themselves eliminate every remaining dependence.
A working taxonomy
Kapoor and Narayanan catalogued leakage patterns across machine-learning-based science and argued that many published results are affected. Their taxonomy is a useful diagnostic vocabulary. Rephrased for world models:
- Task mismatch / mis-specified estimand. The split is valid for one question but reported for another. This is the opening story.
- Correlated observations. Rows can be dependent within a unit. Treating positively dependent losses as independent can overstate effective sample size. Shared nuisance context can also carry information across partitions.
- Preprocessing leakage. Scalers, imputers, embeddings, feature selectors, target encoders, augmentation statistics, deduplication thresholds, and PCA bases fit on all rows including test rows.
- Target leakage. Features that encode the target, or that are only available after the prediction moment.
- Duplicate or window overlap. The same transition, or heavily overlapping windows, appear in both partitions.
The first category is a target/reporting mismatch; the others describe data or information relationships that need separate checks. Better preprocessing alone does not change a mismatched target. Grouping original trajectories before constructing their windows prevents windows from the same trajectory crossing partitions, provided duplicate parent trajectories are handled too. Use the taxonomy to propose diagnostics, not as a guarantee that a fix has been found.
Preprocessing fit boundaries
Split raw groups first. Then fit every learned transformation on training rows only, and apply the frozen transform to validation and test.
This includes learned standardization, imputation, principal components, feature selection, vocabulary construction, target encoding, augmentation statistics, and data-adaptive deduplication thresholds. Fitting them on all rows lets test information influence fitting and violates the declared inductive protocol. The score effect is model-dependent; it is not guaranteed to be optimistic or small. For example, invertible affine feature scaling leaves unregularized least-squares predictions unchanged when an intercept is fitted and the augmented training design has full column rank. Without that identifiability condition, the fitted training values remain unchanged, but independently chosen minimum-norm coefficients can change predictions outside the training subspace. Selecting features using test targets is a different mechanism: it can favor variables that happen to correlate with those targets. Neither the information-boundary violation nor its magnitude should be inferred from a generic transform ranking.
There is a legitimate alternative: a transductive protocol, in which the model may use unlabeled test inputs at fit time. That is a different estimand, it is sometimes the right one, and it must be disclosed. It is not the same as accidental preprocessing leakage, and it should not be described as "standard."
Grouped splits and embargoes
Two practical constructions follow from the unit decision.
Grouped splits. Build a group key (session ID, robot ID, subject ID, level seed) and partition on the key. Every window from the same trajectory stays in the same partition. This is the design used in the executable example below.
Chronological splits with an embargo. For past-to-future forecasting, train on information available before the evaluation period. Derive the gap from the observations used by each row. Under this chapter's convention, a row centered at uses state values from through , and predicts every next state from through . Its inclusive observation support is . A row centered at shares an observation exactly when , assuming no additional dependency. Thus last training center and first test center have disjoint supports when . For , centers three steps apart still share one observation; four steps apart do not. An embargo may remove rows on either side to enforce this condition. Removing steps on both sides is not a universal minimum. Wider feature, label, or preprocessing dependencies need a wider support calculation. For endpoint-only or sparse targets, use the actual support set; the interval can be a conservative bound rather than an exact description. Disjoint supports remove this mechanical overlap; they do not guarantee statistical independence or eliminate every leakage mechanism.
Grouping changes which units are held out; it need not reduce the number of test rows. An embargo removes some eligible observations, which may be training observations rather than test observations. Report row counts and independent-unit counts separately. Precision also depends on within-unit dependence, beyond the split's name. A same-session target can be easier for some models than a new-session target, but that is not a universal ordering. Choose the target first, then budget enough independent units for its uncertainty.
Holding out a group does not automatically create shift
If new sessions are drawn from the same population law as training sessions, the protocol is grouped holdout without an imposed population shift. Finite empirical distributions can still differ. Removing shared session context can make prediction harder for a model that relied on it, but grouping does not guarantee an error increase. This tests "new session, same regime"; it is not by itself evidence about a new robot.
To test an OOD axis, specify how its population law differs. Examples include robots with changed mass or actuation limits, or levels generated with parameters outside the training range. A held-out site or date is a grouping choice; it is evidence of a particular shift only when the relevant changed factors are identified. State both the held-out unit and the varied factor for each evaluation set.
Executable demonstration: two splits, three estimands
Let us build the smallest experiment that makes the mechanism visible. The true dynamics are a stable one-dimensional controlled process with a persistent session intercept:
where:
- : the next state of the simulated process
- : the current state, observed directly (that is, )
- : the control input applied at time
- : shared damping coefficient across all sessions
- : shared action-gain coefficient across all sessions
- : the persistent intercept for session , drawn once per session as and held fixed across all time steps within that session
- : process noise, drawn independently at each step as with
The learner observes the physical scalar directly and is never given or . The complete predictive state includes the unobserved persistent context, so a history-based belief about can improve prediction. Session overlap changes the information a model can reuse; it does not guarantee that a particular model has inferred the context or that context inference is always necessary for the chosen loss.
The first code block sets up the simulator and the shared constants.
import numpy as np
from scipy import stats
A_TRUE = 0.90 # shared damping of the true dynamics
B_TRUE = 0.35 # shared action gain
NOISE_SD = 0.02 # process-noise standard deviation
BIAS_SD = 0.12 # spread of the per-session intercept c_j
N_SESSIONS = 40
STEPS = 60
WINDOW = 3 # number of lagged states used as features
def simulate_cohort(
seed, n_sessions=N_SESSIONS, steps=STEPS, bias_sd=BIAS_SD, bias_shift=0.0
):
"""Simulate sessions with a persistent, learner-invisible intercept.
The physical scalar is observed directly (o_t = s_t), but complete Markov
state also contains the unobserved session-constant context c_j.
"""
rng = np.random.default_rng(seed)
c = bias_shift + rng.normal(0.0, bias_sd, size=n_sessions)
s = np.zeros((n_sessions, steps + 1))
u = rng.normal(0.0, 1.0, size=(n_sessions, steps))
for j in range(n_sessions):
for t in range(steps):
s[j, t + 1] = (
A_TRUE * s[j, t]
+ B_TRUE * u[j, t]
+ c[j]
+ NOISE_SD * rng.normal()
)
return {"c": c, "s": s, "u": u}Next we flatten each cohort into transitions. The feature row for a transition at time is and the target is . The window is what lets a local model infer the recent level of the trajectory, and therefore, indirectly, something about .
def make_transitions(cohort, window=WINDOW):
"""Stack (feature row, target, session id) triples from a cohort."""
s, u = cohort["s"], cohort["u"]
n_sessions, steps_plus_one = s.shape
horizon = steps_plus_one - 1
features, targets, session_ids = [], [], []
for j in range(n_sessions):
for t in range(window - 1, horizon):
features.append([s[j, t - k] for k in range(window)] + [u[j, t]])
targets.append(s[j, t + 1])
session_ids.append(j)
return np.array(features), np.array(targets), np.array(session_ids)
def random_transition_split(session_ids, frac=0.7, seed=0):
"""Partition transitions at random, ignoring session membership."""
rng = np.random.default_rng(seed)
order = rng.permutation(len(session_ids))
cut = int(frac * len(session_ids))
return order[:cut], order[cut:]
def session_grouped_split(session_ids, frac=0.7, seed=0):
"""Partition whole sessions, so no session appears in both halves."""
rng = np.random.default_rng(seed)
unique = np.unique(session_ids)
order = rng.permutation(len(unique))
cut = int(frac * len(unique))
train_sessions = unique[order[:cut]]
test_sessions = unique[order[cut:]]
train_idx = np.flatnonzero(np.isin(session_ids, train_sessions))
test_idx = np.flatnonzero(np.isin(session_ids, test_sessions))
return train_idx, test_idx, train_sessions, test_sessionsNow we build one training set and three evaluation sets. The fitting rows come only from the 28 training sessions. The first two sets compare held-out transitions within those sessions with transitions from wholly held-out sessions under the same collection distribution. Keeping the fitting rows fixed separates the evaluation questions without pretending that grouping and distribution shift are the same manipulation.
cohort = simulate_cohort(seed=101)
X, y, sess = make_transitions(cohort)
train_idx, heldout_idx, train_sessions, heldout_sessions = (
session_grouped_split(sess, frac=0.7, seed=11)
)
# Inside the training sessions, hold out a random 30% of transitions. This is
# the "random transition partition" that the opening story used.
within_fit, within_eval = random_transition_split(
sess[train_idx], frac=0.7, seed=12
)
fit_idx = train_idx[within_fit]
interp_idx = train_idx[within_eval]
# A second independent cohort with a shifted Gaussian intercept mean.
# Gaussian distributions have overlapping support; no strict range separation is assumed.
shifted = simulate_cohort(seed=202, bias_shift=0.45)
Xs, ys, sess_shifted = make_transitions(shifted)
# Namespace independently collected cohort IDs; local row numbering restarts at 0.
sess_shifted = sess_shifted + N_SESSIONS
n_cols = STEPS - (WINDOW - 1)
row_session = np.repeat(np.arange(N_SESSIONS), n_cols)
row_position = np.tile(np.arange(n_cols), N_SESSIONS)Before interpreting errors, check the intended structure. The assertions below verify session-level disjointness for group holdout and shared session membership for within-session evaluation. Such checks can catch an experiment that completes successfully while evaluating the wrong unit. They do not measure the prevalence of that failure relative to other experimental mistakes.
Total transitions: 2320 (features per row: 4) Training sessions: 28 held-out sessions: 12 Fit rows: 1136 random-partition eval rows: 488 New-session eval rows: 696 shifted-session eval rows: 2320 Structure assertions passed: grouped split is session-disjoint; random split is not.
In this draw, within-session evaluation shares every fitting session, whereas grouped evaluation shares none. Fitting rows remain fixed. The evaluation sets nevertheless also differ in realized contexts, state samples, dependence, and row count; their error contrast is not an isolated causal estimate of session reuse.
Now we fit two deliberately different models to the same training rows.
import matplotlib.pyplot as plt
def knn_predict(X_train, y_train, X_test, k=8):
"""Average the targets of the k nearest training rows (Euclidean)."""
d2 = ((X_test[:, None, :] - X_train[None, :, :]) ** 2).sum(axis=-1)
neighbours = np.argsort(d2, axis=1)[:, :k]
return y_train[neighbours].mean(axis=1)
def affine_fit_predict(X_train, y_train, X_test):
"""Least squares affine model using the current state and action only."""
cols = [0, X_train.shape[1] - 1]
A_train = np.column_stack([X_train[:, cols], np.ones(len(X_train))])
A_test = np.column_stack([X_test[:, cols], np.ones(len(X_test))])
coef, *_ = np.linalg.lstsq(A_train, y_train, rcond=None)
return A_test @ coef
def mean_absolute_error(y_true, y_pred):
return float(np.mean(np.abs(y_true - y_pred)))
X_fit, y_fit = X[fit_idx], y[fit_idx]
eval_sets = {
"same sessions,\nunseen transitions": (X[interp_idx], y[interp_idx]),
"unseen sessions": (X[heldout_idx], y[heldout_idx]),
"unseen shifted\nsessions": (Xs, ys),
}
knn_mae, affine_mae = {}, {}
for name, (X_eval, y_eval) in eval_sets.items():
knn_mae[name] = mean_absolute_error(
y_eval, knn_predict(X_fit, y_fit, X_eval, k=8)
)
affine_mae[name] = mean_absolute_error(
y_eval, affine_fit_predict(X_fit, y_fit, X_eval)
)The next block reports the three evaluation errors for each model. Both methods use the same fitting rows, but different feature sets and estimator classes. The third evaluation also changes the session-intercept distribution; the first two do not.
evaluation set affine k-NN window same sessions, unseen transitions 0.0629 0.1247 unseen sessions 0.0699 0.1337 unseen shifted sessions 0.1630 1.8310 Same model, same training data, k-NN window error ratio (new sessions / same sessions): 1.1x
To explain the pattern, we check a structural diagnostic: for each evaluation row, is its nearest training neighbour from the same session?
def nearest_train_session(X_train, session_train, X_test):
d2 = ((X_test[:, None, :] - X_train[None, :, :]) ** 2).sum(axis=-1)
return session_train[np.argmin(d2, axis=1)]
near_interp = nearest_train_session(X_fit, sess[fit_idx], X[interp_idx])
same_session_fraction = float(np.mean(near_interp == sess[interp_idx]))
near_heldout = nearest_train_session(X_fit, sess[fit_idx], X[heldout_idx])
heldout_same_fraction = float(np.mean(near_heldout == sess[heldout_idx]))Nearest training neighbour in the same session: random-partition eval rows: 7.2% new-session eval rows: 0.0% Neighbour session set size (random partition): 28 distinct sessions
Read the diagnostic as a measured check, not a proof of a unique mechanism. The random-partition evaluation includes sessions that also supplied fitting rows, but the computed nearest-neighbour match fraction is modest. No same-session neighbour can exist in the grouped evaluation by construction. In this draw, the k-NN error increases only modestly on held-out sessions, while its much larger increase occurs in the shifted cohort. These numbers do not establish that session reuse explains the entire error gap. The affine model also changes error across evaluation sets, so it is not invariant to grouping or shift.
The table above reports three numbers per model. Plotting them together makes the structural difference visible.

Three qualifications matter here. The observed error orderings are not guarantees for every draw. The within-session result has its own scope; presenting it as new-session transfer would change the claim. Finally, the observed gap is not a demonstrated information-theoretic barrier. With known shared coefficients and past state-action transitions, residuals give noisy observations of . Averaging those residuals can reduce context uncertainty. The implemented feature window does not include past actions, however, and this chapter has not measured a fraction of the gap closed by that enlarged estimator. That would require its own matched evaluation.
Dataset shift
Leakage is about information crossing a boundary it should not. Dataset shift is about the test distribution differing from the training distribution. They are related but distinct, and they need different responses.
The shifted cohort above makes the distributional change concrete. Plotting the session intercepts shows what moved.

- Covariate shift. For one specified prediction task, the full input distribution changes while the output conditional given those inputs stays fixed. A state-transition predictor uses inputs ; an observation predictor may use observation history and action. New lighting can change the observation domain without changing physical dynamics. New layouts are pure covariate shift only if the chosen inputs contain enough layout information for the transition conditional to remain unchanged.
- Dynamics shift. The transition law at the chosen state-action description changes, for example with friction or actuator parameters. In the toy, the conditional law given retains the same coefficients, while the context population changes; the law averaged over unobserved context can change. A fixed robust or context-conditioned system may already handle some changes. Otherwise adaptation, identification, or uncertainty-aware control are candidate responses, not universal requirements. Ordinary input normalization is not a general dynamics-shift remedy.
- Action-distribution shift. A policy change can alter the visited state-action distribution, especially when it changes reachable actions. Two policies differing only on unreachable states need not induce different occupancy. Mismatch between logged support and policy visits is a challenge in offline model-based RL, discussed in Offline and Conservative Model-Based RL.
- Task shift. The objective, reward, or termination criterion changes; specify which reward and success conventions differ.
- Embodiment shift. A different robot, sensor suite, or morphology, which can change the observation and action interfaces as well.
Match the protocol to the declared factor. Appearance randomization is one option for an appearance claim, not a universal covariate-shift treatment. Changed dynamics parameters can be evaluated zero-shot or with a declared adaptation protocol. For action-distribution shift, report collection and evaluation policies and their support. Direct empirical evaluation of new-task performance uses task-level holdout; direct empirical evaluation of new-robot performance uses robot-level holdout. A separate transfer argument can support a conditional claim under explicitly justified structural or identification assumptions; it is not a measurement on those held-out units. State whether the observation/action interface changes rather than assuming every new embodiment changes its dimensions.
The practical recommendation is to make the OOD axis a named column in your results table, not an unstated property of the test set. "Unseen session," "unseen robot," and "unseen level" are different claims, and a single "test" column hides which one you measured. A well-designed results table reads like a small matrix of estimands, and each row or column tells the reader exactly which axis was varied. When the axes are named, replication becomes easier, because the next group can reproduce the same axis even if their robot or their task is different.
Pretraining contamination
When a pretrained world model's corpus is incompletely documented and an evaluation suite was publicly available before the training cutoff, evaluation material may have entered pretraining. Memorization can then be mistaken for transfer. Public availability establishes an opportunity for contamination, not that a particular model encountered the suite.
The honest treatment is a statement about uncertainty, not a verdict. Unknown pretraining membership is an uncertainty, not evidence that contamination occurred. What you can do:
- Report what the model card or technical report says about pretraining data and cutoff dates.
- Prefer evaluation suites that are versioned, dateable, or held privately.
- Where possible, evaluate on data collected after the reported pretraining cutoff.
- Report results with and without a contamination-sensitive subset if such a subset exists.
- Avoid the two extremes: do not assert contamination without evidence, and do not assert cleanliness because the model card is silent.
Repeated leaderboard submissions provide feedback on shared test data. Scores can be dependent or deterministic: submitting identical deterministic predictions twice need not produce a fresh draw. Adaptive choices can nevertheless overfit that feedback. When candidate scores estimate the same quantity with noise, selecting the maximum can produce optimism; the size depends on dependence and the selection procedure. Report submissions, adaptation, and the selection rule. A private held-back evaluation unused for those choices provides a separate test, subject to its own population and sampling assumptions.
Ablations, Baselines, and Statistical Power
We now have a training set and the right held-out unit. The remaining questions are: what should you compare against, what does an ablation establish, and how many independent replicates do you need before a difference means anything?
What each baseline compares
A baseline defines a comparison, not automatically an isolated causal contribution. Specify its inputs, fitting and tuning procedure, and resource scope before interpreting the contrast.
- Simple dynamics baseline. Predict the next state from a scaled last action or fit a low-order linear model. The contrast measures improvement over that specified predictor; it does not separate capacity, nonlinearity, optimization, and feature effects unless those are controlled.
- Action-agnostic predictor. Fit a matched predictor that omits actions. Its contrast measures the incremental predictive value of supplied action information under the protocol. It does not by itself establish how an already-trained conditioned model uses actions. Controlled action changes or interventions address that different question, with support and identification caveats.
- Well-tuned feedback or model-free comparator. Compare control systems under a declared interaction budget. Architecture, optimization, and representation may still differ, so this is a whole-system comparison. Matched component ablations or controlled model substitutions can support more specific contribution claims; identified observational analysis or a justified structural derivation can provide other routes.
- True-dynamics oracle. Give a planner the exact simulator as a privileged diagnostic. Its achieved return is not generally an upper bound when search is approximate, the horizon is short, or information access differs. A solved optimum under the same constraints is a different object and can provide a bound when its assumptions are satisfied.
Giving a planner true state or true dynamics changes its information access and must be disclosed. It can help diagnose model-versus-search limitations, but the diagnosis is conditional on the planner and evaluation protocol. For example, a one-step greedy planner can take an immediate reward of 1 instead of investing now for reward 10 on the next step, even with exact dynamics. Suppose the model also predicts rewards and incorrectly assigns the investment an immediate reward of 2 instead of its true reward of 0. The biased reward prediction induces investment and an actual two-step return of 10. Dynamics bias alone cannot change this one-step ranking if the current-state reward is known exactly. The true-model planner's realized two-step return of 1 is not an upper bound on that result; the true constrained optimum is a separate calculation. Do not present privileged access unavailable at deployment as an ordinary deployable competitor.
A useful heuristic when reading your own results draft: for each row of a results table, ask what a skeptical reader would say if this baseline were absent. If the claim depends on that comparison, keep the baseline in the main table. If removing it leaves the claim unchanged, the comparison can move to an appendix. This keeps the main table focused on comparisons that constrain the claim.
Comparing fairly
Two questions get conflated under the heading of "fair comparison":
- Declared resource caps. Each method gets the specified limits relevant to the question, such as environment interactions or training wall-clock time. This asks what each attained under those caps. Equal gradient-step counts do not imply equal compute when batch sizes or per-step work differ.
- Reasonably tuned performance. Each method is reported at the best validated configuration found within a declared search procedure. This asks what that search attained, not what the method could achieve under every possible training and tuning procedure.
Both questions are legitimate. Matching parameter counts alone does not establish equal training compute, but it can answer a parameter-count-constrained question. For weight-memory constraints, also declare parameter precision and storage. For total deployment-memory constraints, measure or bound runtime state, caches, activations and buffers under the stated input and batch scope. Disclose training cost separately. You need not equalize every resource simultaneously; name the constraint being compared. Unequal tuning effort is part of the attained system contrast, not evidence of an isolated algorithmic effect.
The minimum disclosure is: what was searched, over what space, how it was selected, and how many configurations were tried for each method. Searching 200 configurations consumes compute and can make the winning validation estimate optimistic. A final test independent of that selection need not inherit the validation winner's bias. Report the search and preserve the fitting, selection, and testing boundaries.
One implication that papers sometimes miss: if you tuned method A over 200 configurations and method B over 20, then even with equal compute budgets the comparison is asymmetric in search effort. The honest fixes are to give both methods search budgets of the same order, to report the best-config performance and the default-config performance side by side, or to explicitly declare that the comparison is between one method's tuned configuration and another method's default configuration. Any of those is fine; silently mixing them is not.
Ablations: removal versus retraining
The word "ablation" covers two experiments that answer different questions.
- Removal ablation. Take a trained model and delete or mask a component at inference. This answers: how much does the trained system rely on this component?
- Retrained ablation. Train without the component from scratch using a declared training and tuning procedure. This estimates its contribution to the performance achieved by that procedure, not everything either architecture could learn under all procedures.
These can disagree. A component can help optimization during training yet be dispensable at inference. Conversely, a trained system can rely on a component while a separately retrained system compensates for its absence. Label the intervention and training procedure clearly; these possibilities do not establish a frequency of disagreement.
A second issue is confounding. If removing a component also changes the parameter count, training budget, or optimization difficulty, the ablation measures that combination. A one-factor-at-a-time design that leaves combinations unobserved cannot generally identify their interaction without additional assumptions. For two binary components, evaluating all four combinations is a simple factorial design for estimating the interaction as a difference of differences. Replication is still needed to estimate uncertainty. Additional combinations and replicates consume runs; state which interaction the design can identify rather than assuming every ablation reveals a universal mechanism.
A third issue is identifiability. The effect of a component depends on the rest of the system. An ablation does not automatically identify a universal causal mechanism; it identifies the role of that component in this architecture, at this scale, on this data. Writing "the recurrence is essential for memory" after one ablation is a much stronger claim than the experiment supports. A more defensible phrasing would be: "In our architecture, at this scale, on this suite, removing recurrence at inference degraded held-out error by X, and retraining without recurrence degraded it by Y." The narrower statement travels well; the broader one does not.
Finally, evaluating all subsets of six binary components gives candidate configurations and consumes a search budget. Selecting the largest noisy selection-set estimate can make that estimate optimistic; the candidate count alone does not determine the magnitude of the bias. A final test independent of that selection need not inherit it. Keep the final test set separate, and report the full search, including candidates other than the winner.
Independent replication and pairing
Consider an illustrative experiment with:
- 12 independent training runs per method,
- 30 evaluation episodes per run,
- 200 transitions per episode,
- 20 tasks.
How many independent replicates do you have? It depends entirely on what you are claiming.
| Level | What it is | Independent replicate? | Illustrative count range, not a recommendation |
|---|---|---|---|
| Transition | one row | Not a new training replicate; observation dependence depends on collection | to |
| Evaluation episode | one rollout under a fixed model and policy | Conditional on the fixed model and policy, if episodes are independently sampled; identical distributions are an additional assumption for a common-population analysis | 10 to |
| Planner stream | one sequence of decisions with a fixed seed | Only under an independently generated stream design; cross-method matching is a separate choice | 1 to 100 |
| Task | one environment or task instance | For independent-replicate population inference, justify independent task or task-cluster sampling; dependence-aware designs need their own sampling and variance assumptions | 1 to 100 |
| Training run | one fit from initialization to final checkpoint | Can be independent conditional on fixed data and independent randomization; not automatic | 3 to 20 |
Ten thousand transitions evaluated under one learned model are not ten thousand training replicates. They provide within-run observations, with a dependence structure determined by their collection. A transition-level interval cannot substitute for an interval over training procedures. It omits run variability if that variability is present; whether the omission dominates precision is an empirical question. The illustrative ranges in the table are not literature frequency estimates or required seed counts.
There are two kinds of uncertainty, and they are not interchangeable:
- Conditional uncertainty. Given a fixed fitted model, how variable are the evaluation outcomes? Independently sampled evaluation episodes estimate that variation. More such episodes can reduce Monte Carlo error for the conditional mean under finite-variance assumptions, but their cost depends on the environment and protocol.
- Procedure-level uncertainty. How variable is the outcome across the sources randomized in the target procedure? State whether the dataset, initialization, minibatch order, and model-selection process are fixed or varied. Independent fits conditional on one dataset estimate a different uncertainty from fresh-data fits that rerun selection.
If independent training runs are evaluated on identical fixed cases, inference on the run-level differences is conditional on those cases unless the analysis also models case sampling. To target fresh cases as well as fresh runs, describe the case population and sampling hierarchy. Shared cases create a crossed dependence structure rather than separate nested cases for every run. Accounting for case sampling can add a variance component, but does not guarantee that every estimated interval is wider than its fixed-case counterpart.
The paired-difference variance identity
Pairing can reduce variance in a comparison when matched outcomes have positive covariance. The variance identity states that condition and also shows when pairing can hurt.
Let the experimental unit be the independent replicate, and suppose for replicate you observe under method A and under method B, on the same cases. Define the paired difference
where:
- : the index over the independent replicate (for example, the training run)
- : the metric for replicate under method A
- : the metric for the same replicate under method B, evaluated on the same matched cases
For an error metric, positive means method A has larger error. For a return metric, positive instead means A has larger return; the algebra does not decide which direction is desirable. Write for the average of replicate differences. Variance measures how outcomes vary across repetitions; covariance measures the shared movement of two outcomes. If the are independent and identically distributed with finite variance , then
Uncorrelated equal-variance differences also suffice for this variance identity. Independence is a stronger assumption used in the subsequent interval analysis; the variance identity alone does not establish its distributional validity.
To expand in terms of the individual outcomes, assume both and have finite second moments. Finite variance of their difference alone does not ensure that the marginal variances and covariance exist. Under the marginal-moment condition,
so, retaining the preceding assumptions for the replicate differences and these finite marginal second moments,
Three consequences follow.
- Positive covariance helps. At the same and marginal outcome variances, positive covariance makes the variance of the paired mean difference smaller than an independent-sample comparison. Matching can produce this benefit, but does not establish the covariance's sign or size. Estimated interval widths also depend on the interval procedure and degrees of freedom.
- Negative covariance hurts. If the methods trade off, for example when one succeeds precisely where the other fails, pairing can be worse than not pairing.
- Matching must be defined. A seed number alone does not guarantee corresponding draws across different implementations. Feedback policies can follow different state-action paths while still sharing initial conditions, task difficulty, or exogenous disturbances. Path divergence does not force outcome covariance to zero. Measure the covariance or the paired-difference variability rather than assuming the benefit vanishes. Valid interval inference additionally needs the stated independence and distributional assumptions across replicate pairs.
Pairing removes shared outcome variation, which can include variation caused by shared training-data draws or coupled initialization. If and , the shared cancels in even if it came from training. Unshared variation remains. More independent runs are needed to resolve that residual uncertainty when it is large; matched evaluation cases alone do not guarantee a small residual.
Executable comparison: run-level paired differences
We now run the second experiment. The design is deliberately structured the way a real comparison should be:
- 12 independent runs. Each run simulates a fresh cohort with its own RNG stream, so the intercepts, noise, and actions differ.
- Within each run, a session-grouped split gives 28 training sessions and 12 held-out sessions.
- Both models are fit on the same training rows.
- Both models are evaluated on exactly the same held-out rows, which is what makes the comparison paired.
- The run-level metric is the mean absolute error over those rows, so each run contributes one number per method.
- We average matched cases within a run first, then treat the run-level differences as the replicates.
The two methods are deliberately from different families: an affine model using only the current state and action, and a k-NN model using a three-step history window. Their disagreement is what we will measure.
N_RUNS = 12
def run_comparison(seed, k=8):
"""One independent run: fresh cohort, fresh split, matched evaluation rows."""
cohort = simulate_cohort(seed=seed)
Xr, yr, sr = make_transitions(cohort)
fit_i, held_i, _, _ = session_grouped_split(sr, frac=0.7, seed=seed + 1)
X_train, y_train = Xr[fit_i], yr[fit_i]
X_eval, y_eval = Xr[held_i], yr[held_i]
mae_affine = mean_absolute_error(
y_eval, affine_fit_predict(X_train, y_train, X_eval)
)
mae_knn = mean_absolute_error(
y_eval, knn_predict(X_train, y_train, X_eval, k=k)
)
return mae_affine, mae_knn
run_results = np.array(
[run_comparison(seed=1000 + 7 * r) for r in range(N_RUNS)]
)
run_ids = np.arange(1, N_RUNS + 1)
diffs = run_results[:, 1] - run_results[:, 0] # k-NN window minus affineNow we summarize with a run-level paired Student- interval and a percentile bootstrap. The interval assumes the are approximately normal and i.i.d.; with 12 runs those assumptions are doing real work. Under the exact iid Gaussian model with positive population variance, a 95% procedure would contain the population mean difference in 95% of repeated experiments of this design. That is coverage of the mean, not a claim that 95% of future run differences lie in this interval, nor a posterior probability for the realized interval. For this training-procedure target the bootstrap resamples whole independent runs, not flattened transitions or episodes.
mean_d = float(diffs.mean())
sd_d = float(diffs.std(ddof=1))
se_d = sd_d / np.sqrt(N_RUNS)
t_crit = float(stats.t.ppf(0.975, df=N_RUNS - 1))
ci_t = (mean_d - t_crit * se_d, mean_d + t_crit * se_d)
boot_rng = np.random.default_rng(7)
boot_means = np.array(
[
diffs[boot_rng.integers(0, N_RUNS, size=N_RUNS)].mean()
for _ in range(10_000)
]
)
ci_boot = tuple(float(v) for v in np.percentile(boot_means, [2.5, 97.5]))
standardized = mean_d / sd_d
run_sd_a = float(run_results[:, 0].std(ddof=1))
run_sd_b = float(run_results[:, 1].std(ddof=1))Per-run MAE (affine, k-NN window, difference): run 1 0.0563 0.0980 +0.0417 run 2 0.0718 0.1389 +0.0671 run 3 0.0764 0.1654 +0.0890 run 4 0.0653 0.1135 +0.0482 run 5 0.0692 0.1142 +0.0450 run 6 0.0577 0.1230 +0.0653 run 7 0.0573 0.1175 +0.0602 run 8 0.0659 0.1108 +0.0449 run 9 0.0515 0.1002 +0.0487 run 10 0.0555 0.1573 +0.1018 run 11 0.0846 0.1279 +0.0434 run 12 0.0689 0.1627 +0.0938 Run-level mean difference (k-NN minus affine): +0.0624 Run-level SD of differences: 0.0215 95% paired t interval (11 df): [+0.0488, +0.0761] 95% percentile bootstrap over runs: [+0.0513, +0.0746] Standardized effect size (mean / SD): +2.91 Run-level SD of MAE, affine: 0.0098 k-NN window: 0.0235 Interval contains zero: False
Read the printed interval rather than assuming an outcome. If the interval excludes zero, the data favor one method under this estimand, these assumptions, and this number of independent runs. If it contains zero, the comparison has not resolved the difference at this replication level, and that is a statement about the experiment, not evidence that the methods are equivalent. Note also the practical effect size: a difference can be statistically detectable and too small to matter, or practically meaningful and statistically unresolved. The printed mean and interval give the statistical result and its size in state units. Judging practical importance still requires an application-specific criterion; the standardized mean-to-SD ratio is not an equivalence margin.
Two cautions about the bootstrap here. Its empirical sampling distribution uses only 12 observed run differences. Many resampled means can be numerically distinct; more resamples reduce Monte Carlo error but supply no new independent information about unobserved population tails. Resampling flattened transitions instead targets the wrong unit for this procedure-level comparison. Ignoring run variation or positive within-run dependence can make that interval too narrow; the size of the error depends on the variance components.
The run-level differences are the replicates. Two separate figures preserve distinct quantities: absolute MAE for each model, then the within-run MAE difference and its uncertainty. A confidence interval for a difference must not be overlaid on an absolute-error axis as though it were an interval for either model's MAE.


Seeds help repeatability when the relevant random streams and other nondeterminism are controlled. A seed alone is not a characterization of variation across draws. Independent runs or a justified stochastic model are needed for that uncertainty. The 12 fits in this executable example are a demonstration choice, not a standard or an assurance of adequate power. A borderline interval with few replicates warrants an unresolved conclusion rather than a categorical presence-or-absence claim.
Choosing a replication level prospectively
The right time to decide how many runs you need is before you run them. For a two-sided paired comparison, a standard approximate calculation is
where is the Type I error rate, the desired power, the smallest difference worth detecting, and the assumed standard deviation of the paired differences. The symbol is the quantile of a standard normal distribution. At and power , the two quantiles are approximately and . This is a normal approximation for planning. It is not an exact small-sample guarantee, it is not post-hoc observed power, and it is not a universal seed recommendation.
We can drive it with the pilot estimated from the 12 runs above, while keeping in mind that a pilot estimate from a small number of runs is itself uncertain.
def paired_normal_n(sigma_d, delta, alpha=0.05, power=0.80):
"""Approximate runs needed for a two-sided paired normal test."""
z_alpha = stats.norm.ppf(1.0 - alpha / 2.0)
z_beta = stats.norm.ppf(power)
return ((z_alpha + z_beta) * sigma_d / delta) ** 2
sigma_pilot = max(sd_d, 1e-9)
delta_multipliers = [0.25, 0.50, 1.00]
delta_values = [m * sigma_pilot for m in delta_multipliers]
sigma_grid = np.linspace(0.25 * sigma_pilot, 2.0 * sigma_pilot, 200)
curves = {}
for m, delta in zip(delta_multipliers, delta_values):
curves[m] = np.array([paired_normal_n(s, delta) for s in sigma_grid])
planning_table = []
for m, delta in zip(delta_multipliers, delta_values):
for factor in [0.5, 1.0, 2.0]:
n_est = paired_normal_n(sigma_pilot * factor, delta)
planning_table.append(
(m, delta, factor * sigma_pilot, int(np.ceil(n_est)))
)Pilot sigma_D from 12 runs: 0.0215 state units
target delta (x pilot SD) assumed sigma_D runs needed (rounded up)
0.25x 0.0107 32
0.25x 0.0215 126
0.25x 0.0430 503
0.50x 0.0107 8
0.50x 0.0215 32
0.50x 0.0430 126
1.00x 0.0107 2
1.00x 0.0215 8
1.00x 0.0430 32
All values are planning approximations, not guarantees.
Halving the target difference roughly quadruples the required runs.The table makes the sensitivity concrete: halving the target difference quadruples the unrounded normal-planning requirement, and doubling the assumed variance doubles it. Integer rounding can alter the exact ratios in the table. Both unrounded relations follow from the squared ratio in the formula. This is why three seeds or five seeds cannot be defended as universal defaults: the right replication budget depends on the variance of the measured quantity and the difference worth detecting, neither of which is a universal constant.
The planning formula is easier to read as a curve. The plot varies the assumed paired standard deviation for three target differences. At fixed target difference, error rate, and power, planned runs grow quadratically with SD and linearly with variance.

Practical guidance that follows:
- Round up, always. Fractional runs do not exist.
- Explore plausible SD values. A larger assumed gives a larger planned at fixed target difference, error rate, and power. This is conservative relative to the pilot value, not a certified lower bound under the unknown population distribution.
- Distinguish exact inference from approximate planning. For iid Gaussian differences with positive population variance and , the Student- interval has exact nominal coverage. Small samples can make robustness to nonnormality poor and make the normal planning formula inaccurate. A noncentral- calculation under the Gaussian model, or simulation under a justified alternative model, can refine planning.
- Check heavy-tailed scenarios. Heavy tails can make finite-sample normal planning inaccurate in either direction. Infinite variance invalidates a formula requiring . Simulate plausible distributions and assess Type I error and power rather than assuming the approximation always underestimates the requirement.
- Separate episodes from runs. More independent evaluation episodes can reduce within-run measurement noise and improve procedure-level precision, but cannot remove a nonzero between-run component. When that component dominates, extra episodes buy little compared with extra independent runs.
Equivalence, practical significance, and multiplicity
Several statistical reflexes are wrong, and each has a specific correct counterpart.
Overlapping confidence intervals do not settle whether a paired difference is zero. Two marginal intervals can overlap substantially while the paired difference is confidently nonzero, because overlap ignores the covariance. Always compute the interval on the difference, not by eyeballing two separate intervals.
Absence of statistical significance is not equivalence. Failing to reject zero leaves a range of effects compatible with the estimate and uncertainty; it neither establishes equivalence nor measures power. Prespecify meaningful equivalence bounds. A symmetric choice is , although asymmetric bounds are possible. Two one-sided tests (TOST), each at level , reject the two boundary nulls when the corresponding two-sided interval lies strictly inside those bounds. At , that is a 90% interval, not the 95% interval used above to assess a difference from zero. A 95% interval wholly inside the same bounds is sufficient but more conservative. State the test, interval level, and practical bounds together. Lakens's equivalence-test primer explains this distinction.
Statistical and practical significance are different questions. With enough runs, a difference of 0.001 state units can be statistically significant. That does not make it worth deploying. Report the effect size in interpretable units alongside the interval.
Multiplicity matters. Ten independent true-null tests, each with actual Type I error probability 0.05, have probability of at least one rejection. Nominal level 0.05 alone does not imply that each test attains that probability; dependence also changes the calculation. Define the confirmatory family and error target. Holm controls familywise error, the probability of at least one false rejection. Benjamini-Hochberg controls false discovery rate under its stated dependence assumptions, an expected proportion among rejected hypotheses. These are different guarantees. Prespecifying one primary comparison limits that primary family but does not automatically control additional claims. Labeling other analyses exploratory makes their status clear; disclosure alone does not change their statistical error probabilities.
Prespecify the primary metric, comparison, split, and stopping rule before confirmatory evaluation. Outcome-informed choices can compromise an unadjusted test interpretation. Not every later action is such a choice: transparently correcting a verified implementation error under a declared protocol is different from selecting the most favorable observed result.
Reproducibility and Compute Reporting
Everything above concerns whether a comparison is valid. This section concerns whether it is auditable: whether a reader could, in principle, check it.
A study protocol
A world-model study should be reportable as a compact, checkable protocol. The following list is organized so that each item answers a question a skeptical reader would ask. Reading it as a checklist misses the point; reading it as the set of questions a reviewer will ask is closer. Every item on the list corresponds to a plausible source of an unsupported claim, and covering them is what turns a persuasive paper into a checkable one.
Data and environment.
- Dataset or environment name, version, and content hash. For D4RL-style datasets, which variant and which generation procedure.
- Realized split manifest, or a fully specified deterministic split generator with artifact version, ordering, and randomization settings. A fraction alone is insufficient, but a fraction plus a complete generator can be reproducible.
- Grouping keys and the held-out target, with independence or dependence assumptions stated separately.
- Collection policy: who or what generated the data, at what stage of training, with what exploration, under what safety constraints.
- Provenance, licensing, and consent where applicable, without drawing legal conclusions.
Model and training.
- Training and validation budgets in environment steps, gradient steps, and wall-clock.
- Preprocessing fit boundaries: which transformations were fit on which rows.
- Initial-state and task distributions.
- Number of independent runs, and the seeds, disclosed as reproducibility settings.
- Hyperparameter search record: the space searched, the number of configurations, and the selection criterion.
- Checkpoints, so that reported numbers can be traced to a specific artifact.
Evaluation.
- Termination and reward-timing conventions: when the episode ends, when reward is credited, and how truncation is handled.
- Metrics, aggregation rule, and interval method, including any normalization anchors.
- Raw per-run scores, not only the aggregate. Aggregates can conceal per-run variation and other information needed to audit the analysis.
- Exclusions: which runs were dropped, why, and sensitivity to including interpretable recorded outcomes. If an attempted run has no valid outcome, report that fact and the declared failure-handling rule rather than inventing a score.
- The number of leaderboard submissions, if the benchmark has one.
Limitations. An explicit statement of which axes were not varied, which claims the design cannot support, and what the intended deployment regime is.
Pineau et al. (JMLR 2021) formalized much of this as a reproducibility checklist and a challenge track, and argued that sharing code and data materially improves auditability. The important caveat is that shared code is not a correctness guarantee. Rerunning flawed code reproduces a flawed claim with high fidelity.
Reproducibility aids numerical verification and auditing; it guarantees neither correctness nor external validity. A repeatable result can preserve a mistaken split or leaked fitting boundary. Some audits, such as checking a derivation or a documented protocol contradiction, do not require rerunning the full experiment. Henderson et al. (2017 preprint, AAAI 2018) documented variability across seeds and implementations in deep RL and reporting practices that hindered comparison. State what was reproduced and what that check does not establish.
Compute accounting
The word "compute" can conceal different accounting units. Separate them.
- Environment steps. Actions sent to the real or simulated environment. This is the currency of sample efficiency.
- Environment frames. Individual rendered or physics-substepped frames. With action repeat , one environment step corresponds to frames. Reporting frames while claiming steps (or vice versa) silently inflates or deflates sample efficiency by a factor of .
- Evaluation rollouts. Full episodes run for measurement. These consume environment steps too, and whether they count against the method's budget must be stated.
- Model queries. Count single-candidate, one-step model-transition evaluations during planning or imagination. One rollout pass through candidate sequences of full length with one model member uses such evaluations, assuming no early termination or reused prefixes. With full optimization passes and separately evaluated model members, the corresponding count is ; extra reward or value heads need their own accounting if their costs matter. Batching can combine many evaluations into one function call. These counts are accounting proxies, not latency measurements: batching, state size, parallelization, and hardware determine actual time.
- Training cost. Gradient steps, tokens, and wall-clock, with hardware and numerical precision stated.
A static operation budget can be computed from a disclosed design without measuring runtime. Here the hypothetical planner replans at environment steps , evaluates 256 candidate sequences of length 15 at each decision, and executes the first ten selected actions before replanning. Assume one complete rollout pass with one model member at each decision, no reused candidate prefixes, and a full 200-step episode with no early termination. The block counts single-candidate one-step model-transition evaluations, not necessarily Python calls: one batched call may evaluate many candidates. The arithmetic remains conditional on those assumptions, and no controller rollout or timing measurement is performed here.
H_PLAN = 15 # planning horizon: model transitions per candidate
K_CAND = 256 # candidate action sequences per decision
REPLANS = 20 # replanning steps per episode
EPISODE_STEPS = 200 # environment steps per evaluation episode
ACTION_REPEAT = 4 # environment frames per agent step
ROLLOUTS = 30 # evaluation episodes per (task, run)
TASKS = 8 # tasks in the suite
queries_per_decision = K_CAND * H_PLAN
queries_per_episode = queries_per_decision * REPLANS
env_steps_per_task_run = EPISODE_STEPS * ROLLOUTS
queries_per_task_run = queries_per_episode * ROLLOUTS
frames_per_episode = EPISODE_STEPS * ACTION_REPEAT
budget = {
"Model queries per decision (K x H)": queries_per_decision,
"Model queries per episode": queries_per_episode,
"Environment steps per episode": EPISODE_STEPS,
"Environment frames per episode": frames_per_episode,
"Environment steps per (task, run)": env_steps_per_task_run,
"Model queries per (task, run)": queries_per_task_run,
}Static accounting from the disclosed algorithm (not a measurement): Model queries per decision (K x H) 3,840 Model queries per episode 76,800 Environment steps per episode 200 Environment frames per episode 800 Environment steps per (task, run) 6,000 Model queries per (task, run) 2,304,000 Model queries per environment step: 384 Suite totals over 8 tasks and 1 run: 48,000 env steps, 18,432,000 model queries Elapsed latency is a separate, hardware-dependent measurement.
The ratio above counts different work types. Reporting only environment steps omits planning work; reporting only model-transition evaluations omits environment interactions. More transitions do not establish that planning dominates latency, energy, or FLOPs: a physical interaction can be expensive while many batched model transitions are cheap, or the reverse. Report both counts and separately measure or estimate the resource costs relevant to the claim.
The same accounting is easier to compare as bars on a log scale.

What not to report
Distinguish measured GPU-hours, energy, FLOPs, memory, or latency from analytical counts and estimates. An estimate can be auditable if its method, assumptions, and uncertainty are given; do not present it as observed experimental timing. For measured wall-clock, disclose device, precision, batch size, and software stack. For latency, also define whether its unit is a decision, environment step, or episode, and whether preprocessing and planning are included.
If you compute a static budget from an algorithm, label it as such. "Under our disclosed planning configuration, each decision requires model transitions" is a checkable arithmetic statement. "Our method uses 4.2 GPU-hours per run" is a measurement claim that requires a measurement.
Timings produced on a CPU notebook, including any that appear in this chapter, are local illustrations only, and only where explicitly measured and labeled. Nothing in this chapter should be read as a hardware benchmark.
What valid design still cannot diagnose
Suppose the split matches the claim, adaptive preprocessing respects the fitting boundary, the comparison uses a declared pairing design, uncertainty is reported, and resource accounting is clear. You can now assess the evidence for the stated estimand under disclosed assumptions. Checklist completion does not guarantee that a broad claim is true or that the comparison has enough precision to resolve it.
You do not know why the model behaves as it does. A held-out error measures discrepancy on the evaluated sessions. That discrepancy alone does not distinguish model misspecification from irreducible outcome variation, or identify which component limits performance. It does not tell you whether the encoder discarded the relevant variable, whether the transition model is over-smoothing, whether the planner is exploiting a model artifact, or whether the latent state failed to retain a memory that the task requires. Those are component-level questions, and they need component-level instruments: probing, latent ablations, attention or attribution analysis, and targeted counterfactual tests.
The next chapter, Interpretability and World-Model Debugging, introduces instruments for investigating those internal causes. A sound evaluation makes the observed contrast interpretable within its scope; it does not establish its mechanism. Causal conclusions from interpretability tools require explicit identification assumptions and supporting evidence. Controlled interventions are one useful route; identified observational inference or a justified structural derivation can provide another.
Limitations & Impact
This framework constrains claims; it does not guarantee results or manufacture missing evidence. Splitting logs from one robot does not create observations from a second robot. An empirical new-robot claim therefore needs new-robot evaluation or a separately justified transport argument with stated structural assumptions. Narrow an unsupported empirical claim or collect the missing evidence rather than treating an aggregation or interval as a substitute.
The second limitation concerns assumptions rather than a blanket failure of uncertainty methods. The paired Student- interval has exact nominal coverage for iid Gaussian differences with positive population variance and . With nonnormal differences its finite-sample coverage need not be exact, especially when few runs reveal little about the distribution. The percentile bootstrap uses only the observed runs and cannot infer unseen population tails from extra resamples. The normal sample-size formula is approximate planning, not a guarantee for an unknown heavy-tailed population. State the assumptions, report uncertainty and practical effect sizes, and leave conclusions unresolved when the evidence does not distinguish the relevant alternatives.
An aggregate error increase on new sessions can be compatible with several mechanisms: loss of reusable context, a changed population of dynamics parameters, changed initial-state sampling, or limited fitting data for the model class. One aggregate contrast need not identify which mechanism produced it. Targeted controlled comparisons can investigate those alternatives; the earlier State, Physics, Causality, and Memory Evaluation chapter and the next chapter provide relevant instruments, not automatic causal certificates.
Aligned held-out units and uncertainty reporting improve cross-paper interpretation when populations, metrics, protocols, and resource constraints are also aligned, or their differences are accounted for. A well-designed unresolved comparison can bound what the evidence currently distinguishes, but it does not prove equality or intrinsic task difficulty. A flawed comparison can still yield diagnostics; it should not support the claim its design cannot identify. Rigorous design also has costs: reserving groups reduces fitting data relative to using every group, independent runs add compute, and prospective power calculations may show that a proposed budget cannot resolve the target effect.
"Your split was random" is not by itself a refutation; ask which estimand it targets and whether its dependence and information access match the intended use. A random transition split can be appropriate for a specified interpolation claim. Do not use it as sufficient evidence for new-unit transfer, and do not dismiss it solely because it is random.
Finally, this chapter deliberately stops short of two adjacent topics. It does not analyze why models fail internally, which belongs to the next chapter, and it does not analyze how models are exploited by planners and optimizers, which belongs to Failure Modes and Model Exploitation. Both are places where a valid experimental design can still lead you to a wrong conclusion about a system, and both deserve their own treatment.
Summary
A benchmark couples a task collection with a protocol. Its evidence must match the claim's target population, available information, and evaluation question. Not every fixed-target risk claim requires changing an experimental factor.
- Separate the objects. Dataset, environment, task, benchmark, and leaderboard describe different parts of an experiment. Logs without observed or recoverable actions do not directly test action response; causal intervention identification needs extra assumptions. An interactive suite is not automatically a released offline dataset.
- Coverage beats count. Map scoped claims to required factors and compare those with actual measurements. Duplicate benchmark axes do not add missing coverage; a new tested factor can. Claims about joint settings need joint evidence or justified transfer assumptions, not merely separate tests of each factor.
- Aggregate deliberately. Define the task distribution before averaging. Normalized scores need stated anchors, a rule for zero denominators, and the acknowledgment that they are not bounded in . IQM and performance profiles are useful robust summaries with precise interpretations, not universal cures.
- Name the held-out unit and justify independence separately. Sessions, robots, subjects, levels, tasks, dates, and sites correspond to different generalization targets. Disjoint grouping keys alone do not prove independent samples.
- Random transition partitions are not always wrong. They can estimate a specified interpolation target when information access and dependence match deployment. They do not by themselves establish transfer to a different held-out unit. Temporal dependence alone is not label leakage.
- Split raw groups before adaptive preprocessing in the declared inductive protocol. Fit learned transforms on training rows only, disclose fixed or independently pretrained transforms, and distinguish transductive test-input use. Derive any chronological embargo from the feature and target supports.
- Holding out a group does not determine whether there is shift. New groups can come from the same population law or a changed one. Name both the held-out unit and any changed population factor.
- Transitions are not training replicates. Many transitions under one fit provide within-run observations, not new fits. Independent training replicates require stated conditioning and randomization assumptions; fresh-data claims need data-draw variation too.
- Pair, and derive why. For i.i.d. replicate pairs with finite variance and , . Positive covariance helps relative to independent outcomes with the same marginals; negative covariance hurts. Diverging feedback paths do not imply zero covariance.
- Average within a run first for the training-procedure comparison. Use run-level differences as its replicates. Its bootstrap resamples complete independent runs, not flattened transitions; a different conditional-risk target needs its own sampling unit.
- Plan the replication level. is a planning approximation, not a guarantee. Halving roughly quadruples . No blanket three-seed or five-seed rule applies.
- Interpret intervals correctly. Overlapping marginal intervals do not settle a paired difference, and absence of significance is not equivalence. Equivalence needs prespecified bounds and a stated interval level; TOST at 0.05 uses the 90% interval. Statistical and practical significance are different questions.
- Choose baselines for their actual contrast. Simple dynamics, action-agnostic, and model-free baselines do not automatically isolate components. Controlled comparisons, identified observational analysis, or justified structural derivations can support component attribution under their stated assumptions. True-dynamics planners are privileged diagnostics, not automatic return upper bounds.
- Ablations have two meanings. Removal at inference and retraining without a component answer different questions. One-factor designs can miss interactions; an interaction-identifying design needs sufficient combinations and replication for its target.
- Report the protocol and resources. Include artifact versions and hashes, split manifests, collection policy, budgets, preprocessing boundaries, seeds, search records, raw run scores, and exclusions. Separate environment interactions, frames, rollouts, model-transition evaluations, and training cost. Label measurements, analytical counts, and estimates distinctly.
- Reproducibility is not external validity. Rerunning flawed code reproduces a flawed claim. A complete protocol makes a result auditable; it does not make it true.
The question to carry forward is: which estimand does this number estimate, and is it the target of my claim? Alignment makes the result relevant evidence. Its uncertainty and remaining assumptions still determine how far the conclusion can go.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about datasets, benchmarks, and experimental design.
Datasets, Benchmarks, and Experimental Design
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!