Interpretability and World-Model Debugging

Michael BrenndoerferAugust 5, 202662 min read

Part of World Models Handbook

Probes, TCAV, rollouts, and component substitution show what a world model represents, what its planner uses, and where causal claims remain unsupported.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Interpretability and World-Model Debugging

Imagine you have trained a world model for a small robot. Its one-step prediction loss is low. When you roll the model forward under a fixed action sequence, the predicted trajectories look smooth, physically plausible, and attractive. Then you hand the model to a planner and run the agent. It oscillates, overshoots, drifts into a wall, or settles in the wrong place. Return is poor.

The useful question is where the failure enters the deployed system. In the model-based control pipeline studied here, an encoder or state estimator turns observations into a latent state, a transition model predicts the next latent under an action, a cost model scores outcomes, and a planner searches for actions. Other world models need not have this architecture. Different faults can yield similar trajectories: a biased velocity estimate and an incorrect integration step can both produce drift, while a planner or cost-model change can alter goal-directed behavior. A symptom alone does not identify its source.

Similar symptoms can arise from different components, and one component can produce several symptoms. A low return therefore does not distinguish an estimator fault from a transition or search problem. Controlled experiments can test these explanations. Interface logs, targeted substitutions, and comparisons across rollout regimes provide complementary evidence; no single diagnostic guarantees a unique attribution.

We use two groups of instruments. Representation-level tools include probes for decodability, concept-direction tests for sensitivity of a chosen output, and interventions for response to a specified latent edit. They can be applied at encoder, recurrent, or intermediate network sites, not only inside an encoder. Pipeline-level tools compare teacher-forced, recursive, and closed-loop runs and substitute reference components under disclosed contracts. These tests ask which configuration changes affect the observed behavior, rather than automatically identifying the only faulty component.

Both families are easy to over-read. A probe that recovers velocity tells you the information is accessible to that readout on that distribution. It does not tell you the model uses velocity for control, that the latent is a sufficient Markov state, or that no other information is present. A large directional derivative establishes local output sensitivity to the tested direction; a saliency map needs its own attribution definition before you give it that interpretation. Neither certifies a physical cause in the world. Component substitution measures a behavior difference under a disclosed interface and distribution. It does not automatically identify a unique cause or provide an additive decomposition of error. We will carry these distinctions through the examples.

We build on earlier chapters in Part XI: Evaluation and Understanding. State, physics, causality, and memory evaluation separated observations, latent states, and simulator states. Planning and policy evaluation treated performance as a property of a model-planner-controller combination. Datasets, benchmarks, and experimental design introduced split discipline, matched budgets, paired runs, and reproducibility. We reuse those protocols here. Keep the observation oto_t, true state sts_t, estimated or latent state ztz_t, and action ata_t separate, including their time indices and coordinate systems.

By the end, you should be able to run a small, fully deterministic diagnostic laboratory on a controlled system, interpret the results, and know which conclusions the experiment does and does not support. That discipline is what motivates the next chapter on failure modes and model exploitation, where we look at what happens when a planner actively searches for the model's blind spots. Once you can localize a fault by hand, you are in a position to understand what it means for an optimizer to find faults for you.

What Debugging a World Model Means

Debugging seeks an explanation for a measured failure. A poor loss or return can reveal a problem without locating it. Validation loss remains useful for model selection when it matches the intended use; in model-based control, downstream search and feedback can make prediction loss an incomplete guide to decision quality. We therefore examine both component predictions and the composed system.

Consider the pipeline as a chain of interfaces. The encoder reads observation history (and, in an oracle variant, privileged simulator state) and writes a latent state ztz_t. The transition model reads (zt,at)(z_t, a_t) and writes a predicted zt+1z_{t+1}. The planner uses ztz_t, the transition model, the cost objective, and its search configuration to choose ata_t. The environment advances under that action and produces the next observation. These are specified contracts for our comparisons, not guarantees enforced by Python array types. We must also disclose auxiliary parameters, random streams, and privileged inputs. Substitution is controlled only when the surrounding contracts and information sources remain compatible.

A useful debugging frame is to ask three questions in order:

  1. Is the representation informative? Does ztz_t carry the information the downstream task needs? This is a question about the encoder and the observation channel, answered with probes and concept analysis.
  2. Is the dynamics correct? Given a correct state, does the transition model predict the next state accurately, and does error stay bounded as the horizon grows? This is a question about the transition model, answered with teacher-forced and recursive rollouts.
  3. Is the planner configured correctly? Given a correct model, does the search find good actions? This is a question about the planner, answered with controlled substitution.

These three questions hold the task objective fixed. Check that objective too: does the supplied cost or reward express the intended task? Correct transitions and accurate search can still produce unwanted behavior when the planner solves the wrong objective. The laboratory uses a specified task cost throughout; it does not inject or diagnose a learned reward-model fault.

This order is a useful starting point, not a prerequisite for every test. Privileged simulator states let us test a transition model before trusting the learned estimator. A reference model similarly lets us test a planner while the deployed model remains faulty. The important requirement is to specify the input information and reference assumptions for each comparison.

The questions interact. An estimator error can change a transition model's input distribution; in other cases, a learned transition is calibrated to the estimator's coordinates and remains accurate there. A transition error can change candidate rankings and selected actions. Closed-loop return combines these effects. A controlled experiment changes one specified factor while keeping the rest fixed as far as the protocol permits, and reports what other quantities changed as a consequence.

Localization vs. attribution

Here, localization means narrowing candidate explanations through a specified diagnostic or substitution. A measured replacement effect identifies sensitivity to that configuration change, not necessarily the unique defective interface or physical cause. Faults can interact, and different components can produce similar symptoms.

This distinction is the spine of the chapter. We will build tools that localize, and we will be precise about what localization does and does not license you to claim. In particular, "localize" is a relational verb: you localize a fault relative to an interface, a distribution, and a substitution protocol. Change any of those, and the localization can change too.

Probes and Concept Analysis

A useful question about a learned representation is whether a readout can recover a human-meaningful variable: position, velocity, or object identity, for example. A probe is an auxiliary model trained to predict a target from the representation. Successful held-out prediction supports decodability under that probe and evaluation protocol.

Probe

A probe is an auxiliary model gg trained to map a representation ztz_t to a target yty_t. The probe is fitted after the representation is frozen. Held-out probe accuracy measures how accessible the target is to the chosen probe family on the chosen distribution. It is not a property of the representation alone: it depends on the probe's capacity, the training data, and the preprocessing.

A linear probe tests affine readout performance; a polynomial or MLP probe tests a different function family with specified degree, architecture, and fitting procedure. Some nonlinear families contain the linear family, so their best attainable fit cannot be worse solely because their expressive capacity is larger. Fitted finite-data predictors can still perform worse through optimization, regularization, or overfitting. Report the family, distribution, and split when describing accessibility.

Probes Are Fitted Instruments, Not Oracles

It helps to think of a probe as a measuring instrument. A thermometer does not create temperature; it reads it under a calibration and a range. A probe does not reveal the "true contents" of a representation; it reads whatever is linearly (or nonlinearly) extractable under a training protocol. This analogy is worth extending, because it captures the practical stakes. If you pick up a thermometer calibrated in a laboratory and use it outdoors in freezing weather, it may read nonsense, but that does not mean the temperature is undefined; it means the instrument was used outside its calibrated range. Similarly, a probe that reports low held-out accuracy on a distribution it was never calibrated for may be telling you almost nothing about the representation.

Three consequences follow.

First, successful decoding by a chosen probe is not sufficient to establish downstream use. It is not necessary either: a downstream computation may use a variable through a transformation that the tested probe cannot recover. Later, a nuisance-coordinate example shows the sufficient-direction failure directly: the readout recovers a supplied coordinate, while the controller ignores it.

Second, a failed fitted probe is not proof of information absence or an upper bound on the best readout in its family. Poor fitting, limited data, unsuitable regularization, distribution shift, or a restricted family can all reduce performance. The negative result describes that fitted instrument under that protocol. Testing another family or checking optimization can help distinguish these explanations.

Third, the probe's own capacity is a confound. A high-capacity probe can memorize. If you use a large probe and a small dataset, you can get high held-out accuracy on targets that are only spuriously correlated with the representation. This is why control tasks matter. Without a control, a probe result is compatible with two different stories: (a) the representation encodes the target, and (b) the probe is exploiting incidental structure to fit the target from data it should not be able to use. A matched control is how you begin to separate the two.

Concept Labels, Nuisance, and Capacity

When you probe for a human-named concept, you are making a chain of assumptions. Each link can fail. Failures of assumptions needed for the intended interpretation can make the probe measure something different or limit that interpretation. The chain looks like this:

  • You have a labeling procedure that assigns concept values to examples. The procedure may be noisy or may encode your own assumptions.
  • The concept may be correlated with nuisance variables. If "has stripes" is correlated with "is a zebra," a probe for stripes may be learning zebra-ness.
  • The probe has finite capacity, which constrains the readout functions it can represent. A restricted probe family may fail to recover an available concept; a high-capacity probe can fit incidental correlations or memorize, depending on the data and fitting protocol.

A concrete example: suppose you probe a driving world model for "is the traffic light red." If your dataset contains many more red lights at night than green lights, a probe may learn "is it night" and call it "red light." The probe reports high accuracy. The concept is not what you think it is. The high accuracy is real; the interpretation is broken. This is a nuisance-confounding failure, and it is easy to introduce accidentally because the nuisance correlation may be an artifact of how the data was collected rather than a property of the representation.

A control task helps assess the probe's capacity to learn a deliberately unrelated mapping while preserving specified properties of the real task. It does not automatically preserve every dependency or eliminate every concept correlate. Similar real-task and control-task scores can indicate low selectivity under this comparison; they do not prove that the concept is absent or that every real-task prediction uses spurious information.

John Hewitt and Percy Liang (2019, https://aclanthology.org/D19-1275/) assign independently sampled random control outputs to word types, consistently reusing each type's assignment across its tokens. This is not a permutation of linguistic types. Their selectivity measure is real-task accuracy minus control-task accuracy. Hypothetical scores of 0.95 and 0.55 give selectivity 0.40, whereas two scores of 0.95 give zero. The latter shows no advantage over that memorization control, not proof of concept absence.

Control task (Hewitt and Liang, 2019)

In Hewitt and Liang's construction, random outputs are assigned consistently by word type. The shared input and output spaces and repeated-type consistency make it a comparison for type-memorization capacity. State which properties a proposed control preserves and what its score contrast tests; no universal if-and-only-if test of incidental versus conceptual information follows.

A frame-level shuffled-target refit is different from a type-consistent control. Permuting training labels preserves their multiset and can change sample-to-label associations. An identity permutation, or swapping equal labels, leaves those observed associations unchanged. Even a nontrivial reassignment need not eliminate all accidental correlation. Its outcome depends on the split, shuffle, and fitting procedure. Use it as a randomized-label baseline, and repeat shuffles when their variability matters. To report Hewitt-and-Liang selectivity, construct the corresponding type-level task and disclose what is matched.

Splits and Preprocessing Discipline

Probes are fitted models, and they inherit every discipline that fitted models require. It is tempting to skip this discipline because the probe is "just a diagnostic," but leakage can make its estimates misleading or compromise the intended held-out interpretation. A fitting-boundary violation does not guarantee that every prediction or score changes.

  • For unseen-episode generalization, split by trajectory rather than frame. Samples from one episode share history and initial conditions. Mixing them across splits answers a within-episode question instead. Choose the grouping unit for the intended generalization target; held-out objects and held-out views likewise answer different questions.
  • Fit data-adaptive preprocessing on training only in an inductive evaluation. Estimate standardization means and scales, PCA rotations, and other transformations learned from this experiment's data on training rows, then apply the frozen transform to validation and test. Fitting those statistics on the full dataset lets held-out information influence training. Fixed analytical transforms do not require fitting. Independently pretrained transforms require declared provenance and overlap checks, not compulsory refitting on this split. A disclosed transductive protocol permits different test-input access and answers a different evaluation question.
  • Select penalties on validation, then freeze. Choose ridge penalties, probe architectures, and stopping criteria on a validation split. Once chosen, freeze the probe and evaluate once on the test split. Re-tuning after looking at test results invalidates the test split as a held-out estimate.
  • Report the shuffle scope. Permuting training targets can change the fitting task. Permuting test targets can change the evaluation pairing without changing the already-fitted predictor, provided no refitting or test-driven selection follows. Identity permutations and swaps of equal targets leave the observed pairings unchanged. Specify which labels were randomized, whether assignments were consistent by type or episode, and how randomness was repeated.

We followed all four rules in the laboratory below. The split is by episode, preprocessing statistics come from training only, the ridge penalty is selected on validation, and the shuffled control shuffles training targets only.

Linear Concept Directions and TCAV

A probe estimates a target from a representation. A sensitivity test instead chooses an output and asks how that output changes along a direction in representation space. TCAV, introduced by Been Kim and colleagues in 2018 (https://proceedings.mlr.press/v80/kim18d.html), is one such concept-based sensitivity protocol.

The TCAV recipe separates score construction from validation:

  1. Define a concept by examples. Collect a set of inputs that exhibit the concept (for instance, images of striped textures) and a set that do not.
  2. Fit a concept activation vector (CAV). A linear separator distinguishes concept examples from comparison examples at a chosen layer. Orient its normal toward the concept examples. This is a fitted direction, not a guarantee that moving along it increases the physical concept.
  3. Measure directional sensitivity. Choose a class output and take its directional derivative along the CAV. For that class's examples, the original TCAV score is the fraction with a positive derivative; compute separate scores for separately chosen classes and sites.
  4. Check whether the fitted direction yields a stable score. Kim and colleagues refit CAVs against multiple random contrast sets and apply a two-sided t-test against a score of 0.5, with a stated multiple-testing correction. Disclose the repeats, null hypothesis, and testing procedure. One raw score is not this validation, and a significant score still does not certify physical causality.

The TCAV score answers a specific question: how often does moving along the concept direction increase the output of interest? It is a quantitative, user-defined concept test. Note how the emphasis falls on "user-defined." The concept does not come from the model; it comes from examples you collected, a classifier you trained, and an output you chose.

Scope of TCAV

TCAV was developed for classifier outputs. For a world model, specify a scalar output, a direction-fitting procedure, and an evaluation set before adapting the test. A predicted-state coordinate, predicted reward, or policy logit gives a different question; different choices may, but need not, yield different values. A directional derivative alone is not the original class-aggregated TCAV score or a certificate of physical causality.

A directional derivative measures local output sensitivity at the tested representation. A fitted direction can also track nuisance variation, so that sensitivity need not reflect the intended physical concept. A small derivative does not exclude finite-edit effects: the function f(z)=z2f(z)=z^2 has derivative zero at z=0z=0 but changes under a nonzero edit. Report whether the quantity is a derivative, an aggregate sign score, or a finite intervention effect.

Control Tasks and What They Do (and Don't) Bound

It is worth stating the boundary conditions explicitly. Each of the following four distinctions separates a precise claim from a loose one, and each has a concrete failure mode in practice.

Sensitivity is not decodability. A representation can be decodable for a concept (a probe recovers it) while the output is insensitive to it (moving along the concept direction does not change the output). This is the "decodability without use" pattern. Conversely, the output can be sensitive to a direction that no probe recovers as a clean concept, because the sensitivity flows through a nonlinear interaction. Decodability and sensitivity therefore probe different things; neither finding alone establishes the other without additional assumptions about the concept, readout, output, and tested direction.

Decoding selected state variables with nonzero error does not prove Markov sufficiency. However, if every component of a sufficient finite-dimensional Markov state is exactly recoverable from the same representation, the component readouts jointly recover that state, and it is sufficient for predicting the next-state distribution given the action. The distinction is between partial or approximate recovery and exact recovery of a sufficient state, not a failure of exact component readouts to compose.

Directional sensitivity is not causal fidelity. The concept direction lives in representation space. Whether moving along it corresponds to a physically realizable change in the world depends on the data manifold, the encoder, and the task. An off-manifold latent shift can produce a large derivative without a physical interpretation. Latent proximity is a screening heuristic, not a guarantee of joint support or concept isolation. Check those properties separately when making a physical claim.

Intervention effects are not unique attribution. If replacing a latent variable changes the output, you have shown that the output depends on that variable under that intervention. You have not shown that the variable is the unique cause, that no other variable matters, or that the effect would hold under a different intervention protocol.

These four distinctions recur throughout the chapter. Keep them in mind as we move from probes to rollouts to interventions. Whenever you see a claim like "the model uses X" or "the model reasons about Y," try to identify which of the four distinctions the claim is glossing over.

Rollout Visualization and Error Localization

Rollouts compare prediction and execution under specified inputs. Label the regime: teacher-forced prediction, recursive prediction under fixed actions, or closed-loop execution. A number reported merely as rollout error does not tell the reader whether true states were supplied at each step or whether the model's predictions were fed back.

Three Regimes: Teacher-Forced, Recursive, and Closed-Loop

Teacher-forced one-step prediction supplies the true state sts_t and recorded action ata_t to predict st+1s_{t+1}. Each prediction resets to a reference state, preventing feedback of preceding prediction errors. This isolates local transition accuracy at those inputs. Its deployment relevance depends on the available sensing and state-estimation errors; fully observed systems may also have reference-quality state inputs.

Recursive (open-loop) prediction. You start from the true state s0s_0, predict s^1\hat{s}_1, then feed s^1\hat{s}_1 back into the model to predict s^2\hat{s}_2, and so on. Error can compound when earlier prediction discrepancies change later model inputs. Recursive error measures prediction discrepancy at each horizon as model predictions are fed back under a fixed action sequence. That discrepancy may grow, shrink, cancel, or remain zero. It is the natural setting for asking whether the transition model is stable, and it sits between teacher-forced and closed-loop in the sense that it exposes the model to its own errors but still runs under a fixed action sequence.

Closed-loop execution. You run the full agent in the true environment. The planner reads the (possibly corrupted) latent state, chooses an action, the environment transitions, and you observe. Closed-loop behavior measures the decision usefulness of the whole pipeline. It conflates perception, dynamics, and planning, which is exactly why it needs to be decomposed.

These three regimes can disagree sharply. A model with excellent teacher-forced accuracy can have terrible recursive accuracy if it is unstable. A model with good recursive accuracy can have terrible closed-loop behavior if the planner exploits a systematic bias. And a model with mediocre teacher-forced accuracy can have good closed-loop behavior if the planner is robust to the error it makes. In the last case, the model would look bad under one metric and fine under another. Comparing the regimes can reveal this disagreement under the stated inputs and error definitions. It is not the only diagnostic route: direct code inspection, targeted transition tests, or justified structural analysis can also identify specific faults. A regime comparison alone does not establish a unique mechanism.

What Must Be Logged

An aggregate error curve needs inspectable provenance: the inputs, computation, and error definition must be reconstructable. Retained per-step records are one way to provide that provenance; deterministic generation and executable code can provide it too. For detailed fault diagnosis, specify the following fields at each step of each regime:

Proposed diagnostic fields, with reference times and coordinate mappings specified.
FieldMeaning
episode_idGrouping key for unseen-episode comparison; identity alone does not ensure independent sampling
tStep index within the episode
z_tEncoder output (latent state estimate) at step tt
state_est_tCurrent estimate mapped into reference-state coordinates, for example d(zt)d(z_t)
a_tAction applied at step tt
z_predModel's predicted next latent
s_ref_tReference state at current time tt
s_ref_nextReference state at t+1t+1 after the recorded action
state_pred_nextPredicted next state mapped to the reference coordinate system
transition_residuals_ref_next - state_pred_next; local in teacher forcing, accumulated under recursive inputs
state_est_errors_ref_t - state_est_t, using a current estimate in reference coordinates
cost_tStage cost for action ata_t; the laboratory charges it on the post-transition state at t+1t+1
modeOne of teacher_forced, recursive, closed_loop

These are fields to specify in an auditable logging protocol. Subtract states only after aligning time, coordinates, and units. For a learned latent zz, use a disclosed diagnostic readout d(z)d(z) into reference-state coordinates; if no valid readout exists, report a latent-space error against a compatible latent reference instead. A recursive next-state residual includes earlier prediction errors and is not an isolated one-step transition error. In filtering, "innovation" usually means an observation minus its predicted observation; do not silently equate it with a simulator-state residual. The laboratory below stores state and error arrays for its focused comparisons, not a complete implementation of this logging schema.

Compare state-estimation and transition residuals to generate hypotheses, not unique diagnoses. A large estimator residual suggests checking the observation model or estimator. A large teacher-forced transition residual motivates checking the transition implementation, reference states, and timing. Low error on recorded actions with poor closed-loop cost may reflect search, cost specification, state estimation, or errors on planner-selected trajectories absent from those records. Use a controlled substitution or another targeted check to distinguish the explanations.

Trace identity matters too. Every rollout must record which episode, which initial state, which action sequence, which random seeds, and which timestamps it used. Without trace identity, you cannot align a teacher-forced trace with a recursive trace, and you cannot make paired comparisons. A paired difference cancels a shared additive noise term when it enters both outcomes with the same coefficient. With finite outcome variances and the same marginal distributions, positive covariance reduces difference variance relative to independently sampled outcomes. Pairing alone guarantees neither variance reduction nor isolation of a single causal change. Without trace alignment, the intended pairs cannot be reconstructed.

Reading a Rollout Plot

For prediction errors with aligned reference states, horizons, and units, plot error against horizon and overlay teacher-forced and recursive traces, with separate curves for no-fault and faulty conditions. Show closed-loop outcomes separately, as in the laboratory below, unless a common aligned metric has been defined for all regimes. These comparisons make the observed differences visible; they do not identify a unique fault by themselves. A few patterns can guide further checks:

  • Flat teacher-forced error with growing recursive error shows accumulation under repeated prediction. It does not by itself prove instability: a constant one-step bias in an integrator also accumulates linearly.
  • Growing teacher-forced error shows changing local prediction errors along the recorded trajectory. Systematic bias is one explanation; changing noise, state scales, or input difficulty are others.
  • Recursive error non-monotonic. This happens. State error does not have to grow monotonically. A stable model can recover from a perturbation if the dynamics contract toward the true trajectory. A model near the stability boundary can oscillate. Do not assume monotone growth; plot the trace.
  • Closed-loop cost high, rollout error low. Possible explanations include inadequate search, model error on planner-selected trajectories absent from the recorded rollouts, state-estimation error, or a cost objective that does not match the intended task. The pattern does not by itself distinguish them.

Low rollout error and high closed-loop cost are not contradictory when the rollout evaluates a different distribution from the planner's choices. Robustness does not require avoiding every state with large model error: some errors are irrelevant to the objective, and some visits are unavoidable. Assess errors on the states and action sequences relevant to the decision, alongside cost and feasibility.

The Projection Trap

A low-dimensional view of a high-dimensional latent space omits distinctions. PCA plots, nonlinear embeddings, and selected coordinate pairs can reveal patterns, but a visible cluster is not proof of a concept or causal structure. Check labels, episode identities, nuisance variables, and the full-space geometry before interpreting a projected grouping.

A concept-association claim needs evidence beyond a visually appealing cluster, such as held-out label prediction and nuisance controls. A causal-use claim additionally needs a suitable intervention or other causal identification protocol. These are different questions; an intervention is not the only way to establish an association.

Causal Interventions on Representations

Probe fitting and fixed-action rollout comparison can be observational diagnostics. Rollouts can also be part of an intervention experiment when actions or components are deliberately changed. A representation intervention changes a specified internal quantity and measures the response. It tests an internal computational effect under that protocol, not automatically a physical-world causal mechanism.

Interchange Interventions and Causal Abstraction

Geiger, Lu, Icard, and Potts applied causal-abstraction theory to neural-network analysis in 2021 (https://arxiv.org/abs/2106.02997). Their framework relates an interpretable high-level model to network computations through an alignment and interchange interventions; the broader theory predates this application.

The idea is to posit a high-level interpretable model and ask whether the network implements it. Formally, you specify an alignment between high-level variables and network representations. Then you run interchange interventions: you take a base input and a source input, run the network on both, and at a specified location you swap the representation from the source into the base. If the network's output then matches what the high-level model predicts under the corresponding counterfactual, the alignment is supported.

Interchange intervention

An interchange intervention takes the representation at a chosen site from a source run and substitutes it into a base run, then observes the output. If the output matches the high-level model's counterfactual prediction, the site is a candidate implementation of the aligned high-level variable.

For each test, specify the high-level model, network site, alignment, and intervention rule. Candidate alignments may be proposed or found by search, as in Geiger and colleagues' study. Searching and validating on the same cases can favor a post-selected explanation; a prespecified held-out check helps test the selected alignment without turning that selection into independent evidence. An interchange result is a counterfactual computation under the stated intervention.

Geiger and colleagues demonstrated the method on MQNLI, a natural-language inference task. Their experiment is a text-classification study, not a robot-control study. The framework transfers, but the specific empirical findings do not. When you apply interchange interventions to a world model, you are making an analogy, and you should say so. The analogy is productive when the correspondences are made explicit and unproductive when they are not.

The Othello-GPT Case Study

A well-known sequential-model example is the Othello-GPT work by Kenneth Li and colleagues (ICLR 2023, https://arxiv.org/abs/2210.13382). They trained a transformer to predict legal Othello moves from move sequences, then probed the internal representations for board state and intervened on them.

Two findings are commonly cited. A probe recovers board-state information from internal representations, and targeted representation edits change move predictions consistently with the tested board interpretation. These results concern the investigated models, inputs, and intervention sites.

This is a concrete, instructive case. But it is one synthetic task, and its scope is limited. Probe recovery of a board state is evidence of decodable board information, not evidence of general physical understanding. Intervention effects on move predictions are evidence that the model's predictions depend on the intervened representation, not evidence that the model plays optimally or reasons about the game the way a human does. And the board state is a fully observable, discrete, symbolic variable, which is a much friendlier setting than the continuous, partially observable, noisy state of a robot. The gap between the friendly setting and a real world-model is precisely where the analogy starts to strain, and it is the reason this chapter keeps returning to the point-mass toy rather than treating Othello-GPT as a template.

TCAV Revisited, and Linear Accessibility

The TCAV-style directional test and the interchange intervention ask related but distinct questions. TCAV asks whether the output is sensitive to a direction. Interchange asks whether the output matches a counterfactual. Both are useful, and they can disagree. When they disagree, the disagreement itself is informative: a sensitivity test may fire on a direction that turns out to correspond to a nuisance, while a counterfactual test may fail on a direction that the network happens to respond to in ways the high-level model did not predict.

Hazineh, Zhang, and Chiu's 2023 follow-up (https://arxiv.org/abs/2310.07582) distinguishes player-relative mine/yours/empty labels from absolute black/white/empty labels. The player-relative board description is much more accessible to their linear probes, and intervention success varies with layer. Label semantics, readout family, and site therefore matter. These results do not by themselves rule out a coherent board representation or establish a unique encoding.

The lesson is general. "The model encodes X" is shorthand for a layered, probe-dependent, distribution-dependent statement. When you read or write such claims, unpack them. A useful unpacking template is: at which site, according to which probe family, on which distribution of inputs, under which split, and with what control?

What an Intervention Does Not Prove

A representation intervention is a strong tool, and it is easy to over-claim with it. Four limits are worth stating explicitly.

Off-manifold shifts. An arbitrary edit may produce an internal state absent from normal runs. Nearness of the base and source representations does not guarantee support of their combined state. Evaluate joint support and unintended coordinate changes where possible, and distinguish a valid computational intervention from a physically realizable state change.

Coordinate dependence. Invertible linear transformations preserve the existence of a linear readout: its weights can be transformed with the inverse map. A particular fitted regularized probe need not retain the same score under every transform. Nonlinear reparameterizations can change linear accessibility. Coordinate labels alone are not invariant semantics.

Nuisance confounding. A fitted concept direction may change nuisance attributes as well. Check those attributes and compare interventions designed to preserve them when the question requires isolation. An orthogonal direction or an unrelated sham is not automatically a matched nuisance control; explain which confounds each comparison addresses.

Recurrent state and dose. A partial edit can leave other memory variables inconsistent with an alternative input history. That limits a history-consistent physical interpretation, but does not invalidate a well-defined intervention on the internal computational graph. Specify all edited sites and times. A dose sweep reports the response to multiple edit magnitudes rather than treating one magnitude as a universal effect.

The laboratory uses a no-op sham, a dose sweep, and matched positions. It discusses the support limitations explicitly; position matching is not a completed joint-support check.

Distinguishing Perception, Dynamics, and Planner Failures

We now come to the pipeline-level question: given a bad closed-loop behavior, which component is responsible? The tool is controlled component substitution.

Component Substitution

For a controlled toy, define a baseline, inject one disclosed fault, and replace the affected component with the baseline implementation. Restoring baseline verifies that reversal under the tested protocol. Failure to restore baseline may reflect another fault, an incomplete repair, interface incompatibility, or noise; it is not proof of interaction. In an unknown system, a successful replacement can compensate for another component's error, so test alternative explanations before naming a unique cause.

Three interfaces must be specified before you run the experiment:

Interface types and oracle substitutions for the three components of the pipeline.
ComponentReadsWritesOracle / reference substitution
Encoder / state estimatorobservation history; privileged simulator state for the oraclezt=(p^t,v^t)z_t = (\hat{p}_t, \hat{v}_t)true simulator state sts_t
Transition modelztz_t, ata_tpredicted zt+1z_{t+1}exact discrete dynamics
Plannerztz_t, transition modelaction ata_treference planner with a longer horizon

A controlled replacement needs compatible shapes, units, time indices, meanings, and information availability in addition to matching array types. A true-state estimator output must use coordinates that the unchanged transition and planner interpret correctly. If an adapter is required, disclose and test it. Record additional information and computation introduced by the replacement.

Oracles: Information and Budget

An oracle substitution may change information or computation as well as prediction accuracy. It need not change both: exact dynamics can use the same inputs and a similar calculation as an approximate model. Disclose what changes in each comparator.

In this noisy position-only toy, the oracle estimator uses privileged simulator position and velocity. An implementable estimator does not receive that exact pair from the sensing channel. Other fully observed deployments may expose the modeled state directly. The oracle comparison measures a configuration-specific cost gap, not an additive estimator contribution or a guaranteed achievable optimum. A rescue motivates improving state information; it does not prove that one feasible estimator can close the gap or that no alternative repair could help.

With candidate count and per-transition work held fixed, a longer-horizon random-shooting planner evaluates more transitions. Longer lookahead is not intrinsically more correct and need not reduce executed cost. Report candidate count, horizon, objective, and model-evaluation count. Matching candidate count isolates neither compute nor candidate-sequence coverage when horizon changes; a fixed-compute comparison answers a different question.

Oracle vs. upper bound

An oracle replacement is a comparator, not an optimality certificate. For cost minimization, a feasible policy's cost bounds optimal cost from above under the same objective and admissible information, even if the policy is not optimal. A privileged-state comparator may not be feasible for the deployment problem. State the quantity being bounded, the feasible policy class, and the proof before calling a result a bound. This laboratory reports matched sample costs, not rigorous bounds on expected deployment performance.

Faults Interact

Fault effects need not be additive. If two isolated faults increase cost by 10 and 15 units relative to the same baseline, their combined increase may differ from 25. These are hypothetical numbers, not a measured frequency claim. A joint experiment is needed to estimate the interaction in the tested system.

Feedback permits interactions: an estimator or transition change can alter candidate rankings, selected actions, later observations, and subsequent predictions. It need not change all of them; some perturbations leave the selected action unchanged. The combined outcome can be additive, subadditive, or superadditive relative to the specified single-fault contrasts. Measure that contrast rather than assume its sign.

Matched initial cases and random streams support paired comparisons and can reduce noise; they are not required for every valid randomized comparison and do not create additivity. Different actions lead to different trajectories. Episode cost and time-indexed state differences are both meaningful counterfactual outcomes, provided the comparison is defined; they are not necessarily same-state local transition errors. The laboratory includes a joint-fault condition to measure its outcome contrast rather than assume a sum.

A Deterministic CPU Laboratory

We build a reproducible CPU laboratory with NumPy and Matplotlib. Closed-form point-mass dynamics provide a reference state at each step, letting us compare estimates and predictions without fitting another reference model. These states are exact relative to the simulated model and floating-point implementation. On a real system, a simulator or measurement reference introduces its own errors and information limits. The code retains the arrays needed for the demonstrated comparisons; a production diagnostic would also implement the logging protocol above.

The System, the Observation, and the Representation

The true state is st=(pt,vt)s_t = (p_t, v_t): position and velocity. The dynamics are the exact constant-acceleration discretization of a point mass.

In[3]:
Code
import matplotlib.pyplot as plt
import numpy as np

DT = 0.2  # seconds per step
HORIZON = 30  # steps per episode
A_MAX = 1.0  # acceleration limit, m/s^2
SIGMA_O = 0.01  # observation noise standard deviation, m
N_TRAIN, N_VAL, N_TEST = 30, 15, 15
N_EPISODES = N_TRAIN + N_VAL + N_TEST


def rollout_true(state, actions, dt=DT):
    """Exact constant-acceleration discretization of a point mass."""
    p, v = state
    p_hist, v_hist = [p], [v]
    for a in actions:
        p = p + v * dt + 0.5 * a * dt * dt
        v = v + a * dt
        p_hist.append(p)
        v_hist.append(v)
    return np.array(p_hist), np.array(v_hist)

The observation channel exposes position only, with additive Gaussian noise. Velocity is not observed directly; it must be inferred from history. Retaining observation noise lets us examine its contribution to velocity-channel error; valid downstream conclusions still depend on the stated diagnostic protocol.

In[4]:
Code
def make_episode(seed, horizon=HORIZON):
    rng = np.random.default_rng(seed)
    p0 = rng.uniform(-1.0, 1.0)
    v0 = rng.uniform(-0.5, 0.5)
    actions = rng.uniform(-A_MAX, A_MAX, size=horizon)
    p_true, v_true = rollout_true((p0, v0), actions)
    obs = p_true + rng.normal(0.0, SIGMA_O, size=p_true.shape)
    return {"p": p_true, "v": v_true, "a": actions, "obs": obs}


episodes = [make_episode(1000 + i) for i in range(N_EPISODES)]
train_eps = episodes[:N_TRAIN]
val_eps = episodes[N_TRAIN : N_TRAIN + N_VAL]
test_eps = episodes[N_TRAIN + N_VAL :]

Each episode uses a separate fixed pseudorandom seed for reproducibility. The streams are intended to emulate independent draws; unequal seeds alone do not prove statistical independence. Time is in seconds, position in meters, velocity in meters per second, and acceleration in meters per second squared. A step applies constant acceleration, integrates, then observes position. Episode cost sums 30 undiscounted post-transition costs. Truncation omits later costs rather than increasing the coefficients on earlier costs.

These laboratory outputs were reproduced with Python 3.11.14 and NumPy 2.4.2, using the PCG64 generators created by default_rng. Exact sample values describe that execution, not a cross-version or cross-platform guarantee. Record the bit generator, seed schedule, RNG call sequence and arguments, build, and execution environment when reproducing them.

The representation is handcrafted. It estimates position directly from the observation and estimates velocity by finite differencing consecutive observations.

In[5]:
Code
def encode(obs, dt=DT):
    """Handcrafted 2-D representation: (position estimate, finite-difference velocity estimate).

    Without observation noise, the finite difference equals the interval's
    midpoint velocity (v[t-1] + v[t]) / 2 under constant acceleration.
    Observation noise adds (epsilon[t] - epsilon[t-1]) / dt.
    """
    T = len(obs)
    z = np.zeros((T, 2))
    z[:, 0] = obs
    z[1:, 1] = (obs[1:] - obs[:-1]) / dt
    return z

This is a diagnostic toy, not a discovered neural representation. Its velocity channel uses observation history. Without noise, it measures average velocity over the preceding interval, which equals the midpoint velocity under this toy's constant-acceleration update. It differs from the current endpoint velocity when that interval's acceleration is nonzero; the two coincide when acceleration is zero. Keeping this handcrafted representation fixed lets us vary downstream components without also retraining an encoder.

For t≥1t\geq1, the channel equals vt−1+12at−1Δt+(ϵt−ϵt−1)/Δtv_{t-1}+\tfrac12 a_{t-1}\Delta t+(\epsilon_t-\epsilon_{t-1})/\Delta t, where ϵt\epsilon_t is position-observation noise. Independent observation noises with standard deviation σo\sigma_o give differenced-noise standard deviation 2σo/Δt\sqrt{2}\sigma_o/\Delta t. The channel therefore combines interval averaging with amplified observation noise. Inspect it before trusting a downstream readout.

In[6]:
Code
_channel_episode = test_eps[0]
_z_channel = encode(_channel_episode["obs"])
_channel_time = np.arange(len(_z_channel)) * DT
_channel_v_true = _channel_episode["v"]
_channel_v_est = _z_channel[:, 1].copy()
_channel_v_est[0] = np.nan  # undefined at step 0 by construction
Out[7]:
Visualization
Line chart comparing true and estimated velocity from finite differencing over time.
Finite-difference velocity estimate versus true endpoint velocity over one test episode. The noiseless estimate equals the preceding interval's midpoint velocity under constant acceleration; observation noise adds scatter. Dotted references lie plus or minus three times sqrt(2) sigma_o/dt around current true velocity, showing the differenced-noise scale. They are not coverage bounds for the lagged channel because their center is the endpoint velocity rather than the midpoint. A successful probe does not make the raw channel a clean measurement.

Interval averaging follows from the finite-difference estimator. Noise originates in the observation channel and is amplified by differencing. A trained probe can partly compensate for these errors, so predictive recovery and raw-channel measurement accuracy are different quantities.

Probing the Representation

We fit a ridge probe to recover true velocity from the representation. The probe uses episode-grouped splits: no episode appears in more than one split. We group by episode to target unseen-episode generalization. Mixing frames from the same episode across partitions can expose shared history and, for adjacent finite-difference rows, shared observation noise across the fitting boundary.

In[8]:
Code
def build_probe_data(eps, dt=DT):
    Zs, Ys = [], []
    for ep in eps:
        z = encode(ep["obs"], dt=dt)
        Zs.append(z[1:])  # drop step 0, where the velocity channel is undefined
        Ys.append(ep["v"][1:])  # probe target: true velocity
    return np.concatenate(Zs, axis=0), np.concatenate(Ys, axis=0)


Z_tr, Y_tr = build_probe_data(train_eps)
Z_va, Y_va = build_probe_data(val_eps)
Z_te, Y_te = build_probe_data(test_eps)

Preprocessing statistics come from training only, and the ridge penalty is selected on validation.

In[9]:
Code
def standardize_fit(X):
    mu = X.mean(axis=0)
    sd = X.std(axis=0)
    sd = np.where(sd < 1e-8, 1.0, sd)
    return mu, sd


def standardize_apply(X, mu, sd):
    return (X - mu) / sd


mu_tr, sd_tr = standardize_fit(Z_tr)
Ztr_s = standardize_apply(Z_tr, mu_tr, sd_tr)
Zva_s = standardize_apply(Z_va, mu_tr, sd_tr)
Zte_s = standardize_apply(Z_te, mu_tr, sd_tr)


def ridge_fit(X, y, alpha):
    n, d = X.shape
    Xb = np.hstack([X, np.ones((n, 1))])
    A = Xb.T @ Xb + alpha * np.eye(d + 1)
    A[-1, -1] -= alpha  # do not penalize the intercept
    return np.linalg.solve(A, Xb.T @ y)


def ridge_predict(X, w):
    Xb = np.hstack([X, np.ones((X.shape[0], 1))])
    return Xb @ w


def r2_score(y, yhat):
    ss_res = np.sum((y - yhat) ** 2)
    ss_tot = np.sum((y - y.mean()) ** 2)
    return 1.0 - ss_res / ss_tot
In[10]:
Code
alphas = [1e-4, 1e-3, 1e-2, 1e-1, 1.0, 10.0]
val_scores = [
    r2_score(Y_va, ridge_predict(Zva_s, ridge_fit(Ztr_s, Y_tr, a)))
    for a in alphas
]
best_alpha = alphas[int(np.argmax(val_scores))]
w_probe = ridge_fit(Ztr_s, Y_tr, best_alpha)

We evaluate an intercept-only baseline, a training-target shuffled refit, the selected linear probe, and a probe on one orthogonally rotated representation. The baseline is the training-target mean. The shuffle supplies one randomized-label comparison, not a definitive structure-versus-noise classifier. The rotation tests this particular transform, not all invertible transformations.

In[11]:
Code
r2_intercept = r2_score(Y_te, np.full_like(Y_te, Y_tr.mean()))
r2_probe = r2_score(Y_te, ridge_predict(Zte_s, w_probe))

# Capacity-matched control: shuffle the training targets only, then refit the same probe.
shuffle_rng = np.random.default_rng(7)
Y_tr_shuffled = shuffle_rng.permutation(Y_tr)
w_shuffled = ridge_fit(Ztr_s, Y_tr_shuffled, best_alpha)
r2_shuffled = r2_score(Y_te, ridge_predict(Zte_s, w_shuffled))

# One orthogonal rotation; this does not test every invertible transform.
theta = np.pi / 4
R = np.array([[np.cos(theta), -np.sin(theta)], [np.sin(theta), np.cos(theta)]])
Ztr_rot = Ztr_s @ R.T
Zte_rot = Zte_s @ R.T
w_rot = ridge_fit(Ztr_rot, Y_tr, best_alpha)
r2_rot = r2_score(Y_te, ridge_predict(Zte_rot, w_rot))

probe_results = {
    "Intercept-only": r2_intercept,
    "Shuffled control": r2_shuffled,
    "Linear probe": r2_probe,
    "Rotated probe": r2_rot,
}
Out[12]:
Console
      Intercept-only: held-out R^2 = -0.378
    Shuffled control: held-out R^2 = -0.456
        Linear probe: held-out R^2 =  0.954
       Rotated probe: held-out R^2 =  0.954

Best ridge penalty selected on validation: 10.0
Probe gap over shuffled control: 1.410
Rotation gap (rotated - linear):   0.000
Out[13]:
Visualization
Bar chart of held-out R-squared for intercept-only, shuffled control, linear probe, and rotated probe.
Held-out R-squared for four probe variants on the frozen handcrafted representation. The linear and rotated probes score about 0.954; the shuffled-target control scores about -0.456, below the intercept-only baseline of about -0.378. The shuffled control does not recover velocity on these held-out episodes. The tested invertible rotation preserves recovery; this is evidence of linear accessibility, not sufficiency or downstream causal use.

Read the result carefully. The linear probe recovers velocity well above the intercept-only baseline on these held-out episodes. This does not establish a sufficient Markov state, downstream use of velocity, or absence of other information. The shuffled-target control scores about −0.456-0.456, below the intercept-only score of about −0.378-0.378; the scores are not equal. The linear and rotated probes both score about 0.9540.954 to numerical precision in this experiment. Invertible linear coordinate changes preserve the class of linear readouts, although fitting procedures with regularization need not be invariant to every such transformation. What changes under rotation is the interpretation of individual coordinates: velocity can remain linearly accessible without occupying the same coordinate.

Decodability Without Use

We add a diagnostic coordinate: a deterministic sine function of observed position. It is redundant given the existing position coordinate and can correlate with control-relevant position; it is not information-free. The controller interface explicitly discards it, allowing a designed decodability-without-use comparison.

In[14]:
Code
def encode_ext(obs, dt=DT):
    z = np.zeros((len(obs), 3))
    z[:, :2] = encode(obs, dt=dt)
    # Redundant sine of observed position; discarded at the controller interface.
    z[:, 2] = np.sin(5.0 * obs)
    return z


Zext_tr = np.vstack([encode_ext(ep["obs"]) for ep in train_eps])
Zext_te = np.vstack([encode_ext(ep["obs"]) for ep in test_eps])
nuis_tr = Zext_tr[:, 2]
nuis_te = Zext_te[:, 2]

mu_n, sd_n = standardize_fit(Zext_tr)
w_nuis = ridge_fit(standardize_apply(Zext_tr, mu_n, sd_n), nuis_tr, best_alpha)
r2_nuis = r2_score(
    nuis_te, ridge_predict(standardize_apply(Zext_te, mu_n, sd_n), w_nuis)
)

Now the matched pair: two states with identical meaningful coordinates (p,v)(p, v) but different nuisance values.

In[15]:
Code
def mpc_action(z, transition_fn, horizon, n_samples, a_max, rng, cost_weights):
    q_p, q_v, q_a = cost_weights
    best_cost = np.inf
    best_a = 0.0
    for _ in range(n_samples):
        actions = rng.uniform(-a_max, a_max, size=horizon)
        state = np.array(z, dtype=float)
        cost = 0.0
        for a in actions:
            state = transition_fn(state, a)
            cost += q_p * state[0] ** 2 + q_v * state[1] ** 2 + q_a * a**2
        if cost < best_cost:
            best_cost = cost
            best_a = actions[0]
    return best_a


COST_WEIGHTS = (1.0, 0.1, 0.01)


def transition_true(state, a, dt=DT):
    p, v = state
    return np.array([p + v * dt + 0.5 * a * dt * dt, v + a * dt])


z_pair_a = np.array([0.30, 0.10, 0.0])
z_pair_b = np.array([0.30, 0.10, 1.0])
a_a = mpc_action(
    z_pair_a[:2],
    transition_true,
    10,
    64,
    A_MAX,
    np.random.default_rng(77),
    COST_WEIGHTS,
)
a_b = mpc_action(
    z_pair_b[:2],
    transition_true,
    10,
    64,
    A_MAX,
    np.random.default_rng(77),
    COST_WEIGHTS,
)
Out[16]:
Console
Nuisance probe held-out R^2:       1.000
Action with nuisance value 0.0:   -0.826301
Action with nuisance value 1.0:   -0.826301
Action difference:                 0.000000e+00

The nuisance probe recovers its target essentially perfectly (held-out R^2 ~= 0.9999), because the coordinate is present verbatim in the representation and the probe is fitting the identity map. The controller's action is identical across the matched pair, because the controller reads only the first two coordinates. The pair illustrates the indifference already established by that input slice. A probe that recovers a variable tells you the variable is accessible, not that the pipeline uses it; inspect or test the consumer separately.

The two plotted diagnostics have different definitions: held-out R2R^2 for the readout, and the empirical action standard deviation (normalized by the number of sweep values, ddof=0) divided by the absolute empirical mean plus 10−1210^{-12}. They are not interchangeable or universally bounded between zero and one. In this designed no-op sweep, every controller input and candidate seed is identical, making the action coefficient of variation zero.

In[17]:
Code
nuisance_values = np.linspace(0.0, 1.0, 8)
z_fixed = np.array([0.30, 0.10])  # meaningful coordinates held constant
nuisance_actions = np.array(
    [
        mpc_action(
            z_fixed,
            transition_true,
            10,
            64,
            A_MAX,
            np.random.default_rng(77),
            COST_WEIGHTS,
        )
        for _ in nuisance_values
    ]
)
action_spread = nuisance_actions.std()
action_scale = abs(nuisance_actions.mean()) + 1e-12
action_sensitivity = (
    action_spread / action_scale
)  # zero when the action never moves
Out[18]:
Visualization
Bar chart showing held-out nuisance probe R-squared near one and action coefficient of variation equal to zero. The two bars represent different statistics, not a shared bounded scale.
Two distinct diagnostics for the designed nuisance-coordinate sweep: held-out probe R-squared and the action standard deviation divided by its absolute mean plus a small stabilizer. The probe recovers the supplied coordinate nearly perfectly, while the controller ignores that coordinate and returns the same action throughout the sweep. These statistics have different definitions; the displayed zero-to-one range covers these observed values, not their possible ranges.

High probe recovery and zero action sensitivity coexist without contradiction. The representation carries the coordinate, and the pipeline ignores it.

Localizing Faults with Rollouts

We now inject a dynamics fault: the transition model integrates with the wrong time step. The fault is disclosed and isolated.

In[19]:
Code
def transition_faulty(state, a):
    # Same functional form, wrong integration time step.
    return transition_true(state, a, dt=0.12)


def teacher_forced_errors(eps, transition_fn):
    out = []
    for ep in eps:
        p, v, a = ep["p"], ep["v"], ep["a"]
        errs = []
        for t in range(len(a)):
            pred = transition_fn(np.array([p[t], v[t]]), a[t])
            errs.append(abs(pred[0] - p[t + 1]))
        out.append(errs)
    return np.array(out)


def recursive_errors(eps, transition_fn):
    out = []
    for ep in eps:
        p, v, a = ep["p"], ep["v"], ep["a"]
        pp, vv = p[0], v[0]
        errs = []
        for t in range(len(a)):
            pp, vv = transition_fn(np.array([pp, vv]), a[t])
            errs.append(abs(pp - p[t + 1]))
        out.append(errs)
    return np.array(out)


tf_true = teacher_forced_errors(test_eps, transition_true)
rec_true = recursive_errors(test_eps, transition_true)
tf_fault = teacher_forced_errors(test_eps, transition_faulty)
rec_fault = recursive_errors(test_eps, transition_faulty)
Out[20]:
Visualization
Line chart of position error versus horizon for four rollout conditions.
Mean absolute position error versus horizon for teacher-forced and recursive predictions under correct and faulty dynamics. Both correct-dynamics traces coincide at zero. Under faulty dynamics, the final-step teacher-forced error is about 0.041 m, while recursive error reaches about 1.067 m. This fixed-action experiment distinguishes local mismatch from its accumulated rollout effect. Recursive errors need not grow monotonically in other systems; these traces do not establish asymptotic stability.
Out[21]:
Console
Teacher-forced final-step error, correct:  0.0000 m
Recursive final-step error, correct:       0.0000 m
Teacher-forced final-step error, faulty:   0.0406 m
Recursive final-step error, faulty:        1.0673 m

Both correct-dynamics traces are exactly zero: recursion alone does not create an error in this deterministic, matched-model experiment. With faulty dynamics, final-step teacher-forced position error averages about 0.0410.041 m, while recursive error averages about 1.0671.067 m. Small local mismatch accumulates along the fixed action sequence. These traces do not by themselves classify stability or uniquely identify a fault. The double integrator is not asymptotically stable, and recursive error need not grow monotonically in a different system.

Component Substitution

Now the closed-loop experiment. We run the full pipeline in the true environment from matched held-out initial states. Each condition regenerates its observation and planner streams on demand from the same episode and step seeds; no stream arrays are precomputed. Observation draws match at each episode step. Planner streams also match, but different horizons group their draws into different candidate sequences. These matched cases support paired episode-outcome comparisons under the disclosed component configurations; they do not hold the resulting trajectories or all search properties fixed. Unlike the noisy initial observation in make_episode, the closed-loop function starts its observation history with exact initial position and adds observation noise only after transitions. Every compared closed-loop condition shares this initialization.

In[22]:
Code
def encoder_correct(obs_hist, true_state):
    return encode(obs_hist, dt=DT)


def encoder_faulty(obs_hist, true_state):
    # Perception fault: velocity channel uses the wrong time step.
    return encode(obs_hist, dt=0.12)


def encoder_oracle(obs_hist, true_state):
    # Privileged readout of simulator state (not available to a deployed system).
    z = np.zeros((len(obs_hist), 2))
    z[-1] = true_state[:2]
    return z


def mpc_reference(z, transition_fn, rng):
    return mpc_action(z, transition_fn, 10, 64, A_MAX, rng, COST_WEIGHTS)


def mpc_faulty(z, transition_fn, rng):
    # Short-horizon alternative; whether it underperforms is measured below.
    return mpc_action(z, transition_fn, 3, 64, A_MAX, rng, COST_WEIGHTS)


def run_closed_loop(
    init_state, encoder_fn, transition_fn, mpc_fn, episode_idx, n_steps=HORIZON
):
    state = np.array(init_state, dtype=float)
    obs_hist = [state[0]]
    obs_rng = np.random.default_rng(5000 + episode_idx)
    total_cost = 0.0
    for t in range(n_steps):
        z = encoder_fn(np.array(obs_hist), state)
        z_t = z[-1]
        step_rng = np.random.default_rng(9000 + 100 * episode_idx + t)
        a = float(np.clip(mpc_fn(z_t, transition_fn, step_rng), -A_MAX, A_MAX))
        state = transition_true(state, a)
        obs_hist.append(state[0] + obs_rng.normal(0.0, SIGMA_O))
        total_cost += state[0] ** 2 + 0.1 * state[1] ** 2 + 0.01 * a**2
    return total_cost
In[23]:
Code
init_states = [np.array([ep["p"][0], ep["v"][0]]) for ep in test_eps]

condition_specs = [
    ("baseline", encoder_correct, transition_true, mpc_reference, "Baseline"),
    (
        "perception_fault",
        encoder_faulty,
        transition_true,
        mpc_reference,
        "Perception fault",
    ),
    (
        "perception_fault_oracle",
        encoder_oracle,
        transition_true,
        mpc_reference,
        "Perception fault + oracle perception",
    ),
    (
        "dynamics_fault",
        encoder_correct,
        transition_faulty,
        mpc_reference,
        "Dynamics fault",
    ),
    (
        "dynamics_fault_oracle",
        encoder_correct,
        transition_true,
        mpc_reference,
        "Dynamics fault + oracle dynamics",
    ),
    (
        "planner_fault",
        encoder_correct,
        transition_true,
        mpc_faulty,
        "Short horizon (H=3)",
    ),
    (
        "planner_fault_reference",
        encoder_correct,
        transition_true,
        mpc_reference,
        "H=3 replaced by H=10",
    ),
]

closed_loop_costs = {}
for key, enc, trans, mpc, _label in condition_specs:
    costs = [
        run_closed_loop(init_states[i], enc, trans, mpc, i)
        for i in range(len(test_eps))
    ]
    closed_loop_costs[key] = np.array(costs)
Out[24]:
Console
Condition                                  mean cost   paired diff    diff std
------------------------------------------------------------------------------
Baseline                                        1.45         +0.00        0.00
Perception fault                                1.68         +0.22        0.26
Perception fault + oracle perception            1.48         +0.03        0.30
Dynamics fault                                  1.53         +0.08        0.28
Dynamics fault + oracle dynamics                1.45         +0.00        0.00
Short horizon (H=3)                             1.17         -0.28        0.53
H=3 replaced by H=10                            1.45         +0.00        0.00
Out[25]:
Visualization
Seven cost bars comparing baseline, perception and dynamics faults, a short-horizon alternative, and three substitutions, with across-episode standard-error bars.
Mean closed-loop costs across 15 matched held-out episodes for seven component configurations. Perception and dynamics faults raise the sample mean. Oracle perception reduces perception-fault cost but remains slightly above baseline. Reusing baseline dynamics or planner reproduces baseline exactly. The H=3 alternative has lower cost than the H=10 baseline here, so it is not a demonstrated planner fault. Error bars are marginal across-episode standard errors, s/sqrt(15), not paired-difference confidence intervals.

The sample means differ: baseline costs about 1.4531.453, perception fault 1.6771.677, and oracle perception 1.4781.478. Oracle perception improves the faulty-encoder condition but does not beat baseline here. Replacing faulty dynamics with the exact baseline transition reproduces baseline bit-for-bit; that substitution is identical to the baseline configuration. The H=3 alternative costs about 1.1701.170, lower than the H=10 baseline. Restoring H=10 therefore increases cost in this sample, not repairs a demonstrated planner fault. All seven conditions remain interface-compatible, but compatibility does not guarantee improvement or unique attribution.

Oracle perception has access to simulator state unavailable to the observation-based encoder, so its result is a reference comparison, not an achievable or optimal bound. Both horizon configurations sample 64 candidate sequences, but the longer horizon requires more simulated transitions per decision. This finite search is not a global planning oracle, and greater horizon need not lower executed cost. With 64 candidates and no early termination, H=10 evaluates 640 one-step model transitions per decision, whereas H=3 evaluates 192. These counts do not measure wall-clock time. The comparison changes horizon and compute budget; a matched-compute experiment would answer a different question. These sample means alone do not establish statistically significant advantages or a unique mechanism.

Mixed Faults Can Interact

Now the interaction case: perception fault and dynamics fault together. This is the case that tests whether the localization properties remain well-behaved when two faults are active at once.

In[26]:
Code
mixed_specs = [
    (
        "perception_only",
        encoder_faulty,
        transition_true,
        mpc_reference,
        "Perception only",
    ),
    (
        "dynamics_only",
        encoder_correct,
        transition_faulty,
        mpc_reference,
        "Dynamics only",
    ),
    (
        "both_faults",
        encoder_faulty,
        transition_faulty,
        mpc_reference,
        "Both faults",
    ),
    (
        "both_oracle_dynamics",
        encoder_faulty,
        transition_true,
        mpc_reference,
        "Both + oracle dynamics",
    ),
    (
        "both_oracle_perception",
        encoder_oracle,
        transition_faulty,
        mpc_reference,
        "Both + oracle perception",
    ),
]

mixed_costs = {}
for key, enc, trans, mpc, _label in mixed_specs:
    costs = [
        run_closed_loop(init_states[i], enc, trans, mpc, i)
        for i in range(len(test_eps))
    ]
    mixed_costs[key] = np.array(costs)
Out[27]:
Console
Baseline cost:                      1.45
Perception-only cost:               1.68
Dynamics-only cost:                 1.53
Both-faults cost:                   1.63
Additive prediction:                1.76
Interaction (actual - additive):   -0.13
Out[28]:
Visualization
Bar chart comparing single-fault, combined, additive-prediction, and partial-substitution costs.
Six mixed-fault comparisons on the same episodes. Combined-fault cost is about 1.627, below the additive prediction of about 1.756; the observed interaction is about -0.129. Replacing perception reduces cost to about 1.465, near baseline 1.453. Replacing dynamics leaves cost about 1.677, equal to the perception-only condition and worse than both faults together here. Substitution effects are asymmetric; this example does not prove that every feedback system has nonadditive faults.

Combined-fault mean cost is about 1.6271.627, compared with additive prediction 1.7561.756, giving interaction about −0.129-0.129 on these seeds. Replacing perception gives cost about 1.4651.465, near baseline 1.4531.453. Replacing dynamics gives about 1.6771.677, reproducing the perception-only condition and increasing cost relative to both faults together. This illustrates asymmetric, compensating interactions. Nonadditivity is possible in a feedback loop, not inevitable in every system; the observed comparison does not identify a unique mechanism.

The joint-fault comparison measures each replacement's effect within this pipeline. It does not identify an exclusive responsible component or provide additive attribution. The factorial cost contrast is a descriptive interaction on these episodes; an additive decomposition requires further composition assumptions.

Matched Representation Interventions

Finally, we match a base and source episode at the step where their positions are closest, then sweep a dose interpolating or extrapolating the velocity coordinate. Position matching alone does not verify support of the combined representation. A dose sweep exposes how the response changes across this specified perturbation, without assuming monotonicity.

In[29]:
Code
base_ep = test_eps[0]
source_ep = test_eps[2]
t_match = int(
    np.argmin(np.abs(base_ep["p"] - source_ep["p"]))
)  # on the chapter's seeds this selects t = 15, where z_base[1] = -0.17 and z_source[1] = 0.26

z_base = encode(base_ep["obs"])[t_match]
z_source = encode(source_ep["obs"])[t_match]

doses = np.linspace(0.0, 1.5, 16)
actions_velocity = []
actions_sham = []
for d in doses:
    # Intervention on the velocity coordinate.
    z_int = z_base.copy()
    z_int[1] = z_base[1] + d * (z_source[1] - z_base[1])
    actions_velocity.append(
        mpc_action(
            z_int,
            transition_true,
            10,
            64,
            A_MAX,
            np.random.default_rng(4242),
            COST_WEIGHTS,
        )
    )
    # No-op sham: adding zero leaves the complete controller input unchanged.
    z_sham = z_base.copy()
    z_sham[0] = z_base[0] + d * 0.0
    actions_sham.append(
        mpc_action(
            z_sham,
            transition_true,
            10,
            64,
            A_MAX,
            np.random.default_rng(4242),
            COST_WEIGHTS,
        )
    )
Out[30]:
Visualization
Line chart of planned action versus intervention dose for velocity intervention and sham.
Planned action versus velocity-intervention dose and a no-op sham. Every dose resets the planner generator to seed 4242, using the same candidate bank rather than a fresh stream. Selected actions change non-monotonically as candidate rankings change. The sham leaves the entire input fixed and produces a flat action. Closely matched positions do not certify joint representation support; interpolation and extrapolation may create off-support inputs.

The velocity intervention changes the selected action non-monotonically: at doses 00, 0.10.1, and 0.20.2, actions are approximately −0.470-0.470, 0.3070.307, and −0.104-0.104. Each call resets the same planner seed, so these are candidate-ranking changes under a fixed bank, not fresh random draws at each dose. The no-op sham is flat because it changes no controller input. This establishes response to this specified representation perturbation, not a real-world causal mechanism.

The edited velocity differs from the base-step value, but that does not prove it lies outside the base trajectory's velocity range or the population's joint support. Position matching alone does not establish support of the combined state. Doses from zero to 1.5 include interpolation and extrapolation beyond the source displacement; either needs support checks if normal-operation semantics are claimed. The coordinate's velocity meaning is handcrafted, not discovered from a neural representation.

What the Laboratory Established

Let us be precise about what we showed and what we did not.

Established on these fixed episodes and seeds:

  • The linear probe recovers velocity above the intercept-only baseline.
  • The shuffled refit scores below that baseline; the tested orthogonal rotation preserves recovery.
  • The redundant diagnostic coordinate is decodable, but the controller interface discards it.
  • Correct-dynamics teacher-forced and recursive errors are zero. The faulty final recursive error exceeds the faulty final one-step error.
  • Oracle perception gives a cost near baseline, and exact baseline re-substitutions reproduce baseline. Neither statement is a significance or equivalence test.
  • H=3 gives a lower sample mean cost than H=10 under their disclosed, different transition budgets.
  • The mixed-fault cost is below its additive prediction on these episodes.
  • The fixed-bank velocity dose sweep changes the selected action non-monotonically. The no-op sham leaves it unchanged.

Not established by these comparisons:

  • Markov sufficiency or absence of other representation information.
  • Hewitt-and-Liang selectivity from the frame-target shuffle.
  • Unique fault attribution or an additive decomposition.
  • An achievable optimal bound from a privileged reference.
  • Physical-world causal fidelity of the representation edit.
  • Generalization beyond this protocol and evaluation distribution.

That list is the honest summary, and it is the summary you should write for any diagnostic experiment on a real world model. The discipline of writing the list explicitly is what keeps a diagnostic useful: it tells a reader exactly which claims are supported, and helps identify unsupported conclusions and candidate next experiments.

Limitations & Impact

The tools in this chapter are powerful and easy to misuse. Let us go through the main limitations, because knowing them is what separates a diagnostic that informs a fix from a diagnostic that manufactures false confidence.

Probes measure accessibility under a readout and fitting protocol. Held-out decoding supports an association, not unique intrinsic semantics. Specify capacity, preprocessing, splits, and target distribution. Different fitted families can give different results; a failed fitted readout does not establish an upper bound on its family's best achievable predictive performance under the same scoring protocol. On a fixed evaluation set with a higher-is-better score, a fitted family member instead supplies a lower bound on the best family score on those same cases. That is not a population-performance bound without further inference assumptions.

For Hewitt-and-Liang selectivity, use their type-consistent randomized mapping or a justified corresponding control. Such tasks retain selected properties, not all statistical structure. A frame-label shuffle is a different randomized-label baseline. Explain what is matched and compare uncertainty before interpreting the score gap.

A CAV is fitted from labeled examples and can inherit labeling noise or nuisance associations. TCAV measures class-output sensitivity along that direction and aggregates derivative signs over selected class examples. A world-model adaptation needs its own output and aggregation protocol; sensitivity does not certify physical causality.

Interchange interventions test a candidate high-level/network alignment. Search may generate candidates; independent held-out validation helps separate selection from testing. Off-support edits, nuisance changes, coordinate choices, recurrence, and dose limit semantic interpretation. A partial recurrent edit can be inconsistent with a physical input history while remaining a valid internal computational counterfactual.

Teacher-forced, recursive, and closed-loop measurements use different conditioning rules. Teacher forcing resets to reference states, recursion feeds back predicted states under fixed actions, and closed-loop execution selects actions through the controller. Their realized state or action values can coincide in particular models and cases. Recursive error need not grow monotonically. A lower-dimensional projection hides distinctions, and a cluster alone proves neither concept association nor causal use. Interpret each trace with its reference inputs, timing, and error definition.

A replacement measures a configuration effect under a specified interface and distribution. It need not improve performance, identify a unique fault, or supply an optimal bound. Disclose information and compute changes when present. Report episode costs and, where useful, time-indexed counterfactual traces without confusing them with same-state local residuals.

Repeating episodes with frozen components measures conditional episode variation, not variation across retrained models. Inferential intervals require a sampling model and a stated construction; repeated deterministic cases alone do not ensure coverage. To study retraining variation, fit multiple models and evaluate each under a common protocol.

These methods answer complementary questions. Probes quantify readout performance; control tasks contextualize memorization capacity; causal-abstraction tests compare aligned counterfactual computations. Rollout comparisons and replacements test suspected pipeline faults. Their usefulness comes from the specific measurement and controls, not a general promise of unique fault localization.

Treat every diagnostic as a measurement under a protocol. State the protocol, report the controls, and be explicit about the scope. When a probe succeeds, say "velocity is linearly accessible from the frozen representation on held-out episodes." When a substitution restores baseline, say "under this interface and distribution, replacing the faulty component restores baseline cost." When an intervention changes the output, say "changing this coordinate changes the planned action under this dose sweep." Each statement keeps its conclusion within the tested scope.

The next chapter, failure modes and model exploitation, examines a planner searching for optimistic predictions. The planned safe-control chapter will address ways to limit that risk. This is a forward link to its intended topic, not a claim that its currently unwritten methods implement these substitution protocols.

Summary

In a model-based control pipeline, aggregate return combines perception, transition, cost, and search effects. The diagnostics here narrow explanations through controlled comparisons rather than automatically separating a unique cause.

  • Probes are fitted instruments. They measure accessibility to a chosen readout on a chosen distribution, not intrinsic content, sufficiency, or downstream use.
  • CAVs define fitted concept directions. TCAV measures and aggregates directional sensitivity of chosen outputs. Decodability, sensitivity, finite intervention effects, and physical causal fidelity are distinct questions.
  • A shuffled-target refit is a randomized-label baseline, not a selectivity bound. Type-level controls preserve specified properties; disclose their construction and which confounds the comparison addresses.
  • Rollout visualization must distinguish teacher-forced prediction, recursive prediction, and closed-loop outcomes. Retain reconstructable inputs and computations; a detailed trace should specify identity, resets, actions, residual definitions, and horizons. Recursive error need not grow monotonically, and projections hide dimensions.
  • Interchange causal-abstraction tests evaluate a specified high-level/network alignment. Other intervention protocols test other hypotheses. Support, nuisance changes, coordinate choices, recurrent memory, and dose constrain semantic interpretation.
  • Component substitution compares configurations under a disclosed interface and distribution. It does not guarantee improvement or unique causal attribution; oracle comparators are not proven optimal bounds, and fault interactions need not be additive.
  • The laboratory built a deterministic point-mass system with a handcrafted representation, episode-grouped splits, a validation-selected ridge probe, a shuffled control, a rotation, a nuisance coordinate, two component faults and a short-horizon alternative, component substitutions, a mixed-fault comparison, and a representation dose sweep. Its findings apply to the specified episodes and protocol, not every world model.

The habit to take away is to write down the protocol with every diagnostic: what was frozen, what was split, what was controlled, what was substituted, and what the scope of the claim is. A diagnostic that respects its own limits is the only kind worth running.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about interpretability and world-model debugging.

Interpretability and World-Model Debugging

Question 1 of 80 of 8 completed
A linear ridge probe recovers true velocity from a frozen representation with held-out R-squared about 0.954, while a shuffled-target refit scores about negative 0.456, below the intercept-only baseline of about negative 0.378. What does this contrast support?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026interpretabilityworld, author = {Michael Brenndoerfer}, title = {Interpretability and World-Model Debugging}, year = {2026}, url = {https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Interpretability and World-Model Debugging. Retrieved from https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts
MLAAcademic
Michael Brenndoerfer. "Interpretability and World-Model Debugging." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts>.
CHICAGOAcademic
Michael Brenndoerfer. "Interpretability and World-Model Debugging." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Interpretability and World-Model Debugging'. Available at: https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Interpretability and World-Model Debugging. https://mbrenndoerfer.com/writing/interpretability-world-model-debugging-probes-rollouts

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.