Failure Modes and Model Exploitation

Michael BrenndoerferAugust 6, 202654 min read

Part of World Models Handbook

World models fail through compounding error, objective misspecification, planner exploitation, state aliasing, forgetting, and hallucinated dynamics.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Failure Modes and Model Exploitation

Imagine a vehicle moving along a one-dimensional track, with position xx, velocity vv, and thrust command aa. The actuator clips commands to [−1,1][-1,1]. Suppose your training data contain only commands in that interval. A linear transition rule and the clipped reference rule agree exactly there, so those data alone do not identify their different out-of-range behavior. We will construct these two rules directly; the chapter does not claim to have collected a dataset or fitted a model.

Give the surrogate to an open-loop random-shooting planner: sample NN four-action sequences from a box such as [−5,5]4[-5,5]^4, predict each rollout, and select the sequence with the lowest predicted cost. The stipulated task charges squared position error from x=1x=1 at each of the four steps. The experiments execute the whole selected sequence. Applying only its first action and searching again after an observation would instead make this a receding-horizon controller; we discuss that possibility later but do not execute it here.

One constructed candidate is [4,−4,0,0][4,-4,0,0]: thrust, reverse, then coast. The surrogate predicts cost 1.01.0 and a terminal position of 11. Executing the same commands in the clipped reference process gives terminal position 0.250.25 and cost 2.68752.6875. These are squared-position costs in the toy's normalized units, not currency. This candidate illustrates an optimistic score; it is not asserted to be the output of a particular random search or of a fitted confidence model.

The story links three mechanisms. The surrogate is exact on the stipulated action interval but biased outside it. Selection can favor candidates whose predicted costs are erroneously low. The resulting transition can also violate the reference's action-conditioned possibilities. These mechanisms overlap: exploitation does not exonerate the model that supplied its attractive error. The position cost is the task specified here, not a misspecified proxy; a separate example will isolate objective misspecification. No external attacker or coding accident is needed for this constructed failure.

This is the first chapter of Part XII: Reliable World Models. Part XI: Evaluation and Understanding covered prediction, perceptual similarity, physical consistency, and planning performance. Here we ask why strong measurements on tested cases need not imply good decisions on other cases. We will develop distinct, sometimes overlapping failure mechanisms, reproduce selected CPU-sized examples, and use controlled comparisons to diagnose failures without promising a unique cause from every symptom.

An average validation loss weights inputs according to its evaluation distribution. It does not by itself bound error on every decision-relevant input. Score-based selection can change the candidate distribution, and allowed search can expose optimistic model errors that validation did not cover. For iid proposals, selecting the minimum of a fixed finite predicted score reweights their law when there is more than one candidate and the score varies under that law. With one candidate, or a constant score and first-index tie-breaking, the selected law is unchanged. Other failures below need no off-distribution search: the objective may be wrong, an observation may omit a cue, or a conditional mean may violate transition support. Good average prediction error can therefore coexist with poor decisions. Their possible severity also depends on task-cost and horizon bounds; the finite toy does not establish arbitrarily large loss.

Notation and the distinctions that must survive

Before we touch a single failure, it is worth fixing the vocabulary. The whole chapter depends on keeping these objects separate.

Objects and their notation
  • st∈Ss_t \in \mathcal{S}: the reference state, the physical or otherwise "true" variable that generates the data.
  • ot∈Oo_t \in \mathcal{O}: the observation, which can omit information in sts_t. Our dynamical examples expose the state directly; the later cue example deliberately does not. Part III: Representing Agents and Worlds develops this distinction.
  • ztz_t: a learned latent state. A latent coordinate is not automatically a sufficient statistic for sts_t.
  • hth_t: hidden recurrent memory carried by a model across time. It may contain information that no single ztz_t does.
  • at∈Aa_t \in \mathcal{A}: the action the agent issues.
  • rtr_t: reference reward; ctc_t: reference cost. When comparing the cost and return conventions, we use rt=−ctr_t=-c_t. A learned reward or cost head is distinct from the intended task objective.
  • ff or pp: reference dynamics; f^θ\hat f_\theta: surrogate dynamics, which may be estimated from data or explicitly constructed as in these examples.
  • π\pi: a policy specifying an action or an action distribution from the information available to the controller, such as a state, observation history, or belief. A planner selects actions or action sequences by searching with a model.

Keep these objects separate when interpreting a failure. A lossy observation or latent representation can hide a needed distinction. A planner can preferentially select a model's optimistic errors. A supplied objective can differ from the intended task even with exact dynamics. These are different questions, not mutually exclusive labels: the opening already combines dynamics bias, selection, and a transition-support violation.

Two distinctions are especially useful when deciding what to test.

First, prediction quality is not decision quality. A fixed lower average supervised loss is better on that loss and evaluation distribution, but need not improve decisions under a different selected distribution or objective. A model with larger average error can still support better control if its decision-relevant predictions are better or its use is more conservative. Separate the predictions from how the agent uses them. Part VII: Planning and Agency develops the planning side; here we examine their interaction.

Second, model bias and objective misspecification live in different objects. Bias is error in surrogate dynamics or a predicted reward; misspecification is a supplied optimization target that differs from the intended task. Exact dynamics can optimize a wrong target, and an aligned target can still be evaluated through biased dynamics. Editing a target does not make a fixed transition map accurate, nor does refitting transitions make a fixed target aligned. Behavioral remedies can nevertheless overlap: a constraint or penalty can discourage a harmful plan while leaving its dynamics prediction wrong. Distinguish repairing the defective object from mitigating its consequences.

Let's write down the reference process and the surrogate so we can point at them.

In[3]:
Code
import itertools

import numpy as np

## --- The reference process: a 1-D point mass with a saturated actuator ---
DT = 0.5  # seconds per step
U_MAX = 1.0  # hardware saturation bound on the commanded thrust
X_GOAL = 1.0  # target position at every step of the horizon
H = 4  # planning and execution horizon (in steps)
A_PLAN = 5.0  # planner's candidate action range, deliberately wider than U_MAX


def reference_step(s, a):
    """Reference dynamics: the actuator saturates, everything else is linear."""
    x, v = s
    a_eff = float(np.clip(a, -U_MAX, U_MAX))
    return np.array([x + v * DT, v + a_eff * DT])


def surrogate_step(s, a):
    """Constructed linear surrogate: exact for in-range actions, unclipped outside."""
    x, v = s
    return np.array([x + v * DT, v + a * DT])


def rollout(step_fn, s0, actions):
    """Roll out a fixed action sequence under the given dynamics."""
    s = np.asarray(s0, dtype=float)
    traj = np.empty((len(actions) + 1, 2))
    traj[0] = s
    for i, a in enumerate(actions):
        s = step_fn(s, a)
        traj[i + 1] = s
    return traj


def trajectory_cost(traj):
    """Sum of squared position error to X_GOAL over the horizon."""
    return float(np.sum((traj[1:, 0] - X_GOAL) ** 2))


def palette_colors(palette):
    """Return the shared style palette as a plain list of colors."""
    return (
        list(palette.values()) if hasattr(palette, "values") else list(palette)
    )


s0_default = np.array([0.0, 0.0])

surrogate_step and reference_step agree exactly for every state whenever ∣a∣≤1|a|\le1. They are constructed maps, not the result of a reported fitting run. A dataset restricted to that interval cannot distinguish these two transition rules from observations alone. Known actuator constraints or an appropriate model class can distinguish them without collecting out-of-range data. Their discrepancy is an unconstrained extension beyond the stipulated action support.

Compare three hand-chosen plans before running the separate randomized search experiment.

In[4]:
Code
## A plan that looks perfect to the surrogate, and one the hardware can actually run.
plan_shortcut = np.array([4.0, -4.0, 0.0, 0.0])
plan_feasible = np.array([1.0, 1.0, -1.0, -1.0])
plan_naive = np.ones(H)

surrogate_traj = rollout(surrogate_step, s0_default, plan_shortcut)
reference_traj = rollout(reference_step, s0_default, plan_shortcut)
feasible_traj = rollout(reference_step, s0_default, plan_feasible)
naive_traj = rollout(reference_step, s0_default, plan_naive)

predicted_cost = trajectory_cost(surrogate_traj)
executed_cost = trajectory_cost(reference_traj)
feasible_cost = trajectory_cost(feasible_traj)
naive_cost = trajectory_cost(naive_traj)
Out[5]:
Console
Shortcut plan [4.0, -4.0, 0.0, 0.0]
  surrogate predicted cost :   1.0000
  reference executed cost  :   2.6875
  optimism (exec - pred)   :   1.6875

Feasible plan [1.0, 1.0, -1.0, -1.0]
  reference executed cost  :   1.6250

Naive full-thrust plan [1.0, 1.0, 1.0, 1.0]
  reference executed cost  :   1.8750

The constructed shortcut predicts cost 1.01.0 but executes at 2.68752.6875, worse than the feasible comparator's 1.6251.625 and the all-positive comparator's 1.8751.875. Calling an attractive scoring error model exploitation emphasizes selection of that error, not merely its existence. Restricting this toy's commands to [−1,1][-1,1] makes surrogate and reference scores agree for every considered plan. A finite random search can still miss the feasible comparator; no-search behavior depends on the chosen baseline. Neither restriction alone guarantees one of the displayed hand-chosen plans.

The examples that follow isolate complementary mechanisms. They are not a nested hierarchy: a wrong objective can occur with exact dynamics, while bias, propagation, and exploitation can occur together. The final workflow uses their distinctions to design comparisons, not to assign every failure exactly one label.

Compounding Rollout Error

Begin with prediction propagation under matched actions. It can contribute to a planning failure, but its presence or prevalence is not implied by the existence of other mechanisms.

Run the reference and surrogate from the same initial state using the same realized action sequence {a0,…,aT−1}\{a_0,\ldots,a_{T-1}\}. Define et=∥s^t−st∥e_t=\|\hat{s}_t-s_t\| in one fixed state norm. A discrepancy at one step changes the input to the next, so small one-step errors can accumulate or amplify. They need not do so: contraction, cancellation, and exact predictions can keep the realized discrepancy small. We will bound it under explicit domain and action conditions.

Let's derive the bound. Write f^θ\hat{f}_\theta for the surrogate and ff for the reference. Define the surrogate rollout and the reference rollout in parallel:

s^t+1=f^θ(s^t,at),st+1=f(st,at).\hat{s}_{t+1} = \hat{f}_\theta(\hat{s}_t, a_t), \qquad s_{t+1} = f(s_t, a_t).

Now subtract and insert a zero:

s^t+1−st+1=[f^θ(s^t,at)−f(s^t,at)]+[f(s^t,at)−f(st,at)].\hat{s}_{t+1} - s_{t+1} = \bigl[\hat{f}_\theta(\hat{s}_t, a_t) - f(\hat{s}_t, a_t)\bigr] + \bigl[f(\hat{s}_t, a_t) - f(s_t, a_t)\bigr].

Take norms and apply the triangle inequality:

et+1≤∥f^θ(s^t,at)−f(s^t,at)∥+∥f(s^t,at)−f(st,at)∥.e_{t+1} \le \bigl\|\hat{f}_\theta(\hat{s}_t, a_t) - f(\hat{s}_t, a_t)\bigr\| + \bigl\|f(\hat{s}_t, a_t) - f(s_t, a_t)\bigr\|.

The first term measures same-input approximation error at the surrogate's current state. The second evaluates the reference map at two different states. A uniform state-Lipschitz bound is a property of that fixed reference map, domain, action set, and norm; the value of this term depends on the states reached and therefore on the surrogate too. The surrogate states must be included in the domain, not just the reference states. Assume:

  • (A1) In the fixed norm, ε≥0\varepsilon\ge0 bounds ∥f^θ(s,a)−f(s,a)∥\|\hat{f}_\theta(s,a)-f(s,a)\| for every s∈Ds\in\mathcal D and every considered action. Both trajectories stay in D\mathcal D throughout the comparison.
  • (A2) In the same norm, L≥0L\ge0 bounds reference state sensitivity uniformly over the same actions: ∥f(s,a)−f(s′,a)∥≤L∥s−s′∥\|f(s,a)-f(s',a)\|\le L\|s-s'\| for all s,s′∈Ds,s'\in\mathcal D and all considered aa.

then the recursion collapses to a familiar scalar form:

  et+1≤Let+ε.  \boxed{\;e_{t+1} \le L e_t + \varepsilon.\;}

The recurrence bounds the next discrepancy by a propagated-error budget LetLe_t plus an approximation budget ε\varepsilon. These are upper bounds, not prescribed increments in the actual error. Directions, signs, and how tightly LL and ε\varepsilon fit the visited states determine how much of that budget is realized.

Unrolling from e0e_0:

et≤Lte0+ε∑j=0t−1Lj=Lte0+ε⋅Lt−1L−1(L≠1),e_t \le L^t e_0 + \varepsilon \sum_{j=0}^{t-1} L^j = L^t e_0 + \varepsilon \cdot \frac{L^t - 1}{L - 1} \quad (L \ne 1),

and for L=1L = 1 the sum stays linear:

et≤e0+εt.e_t \le e_0 + \varepsilon t.

The upper envelope has three regimes. These describe the bound, not three mandatory shapes for the actual error:

  • 0≤L<10\le L<1: the envelope tends to ε/(1−L)\varepsilon/(1-L) as its initial-error term vanishes. The actual discrepancy need not approach that plateau. This bounds state error under (A1)–(A2); it is not a task-safety guarantee. A smaller valid ε\varepsilon lowers the envelope, though the observed error need not change proportionally.
  • L=1L=1: the envelope is e0+εte_0+\varepsilon t. A scalar update s′=s+us'=s+u has unit state sensitivity for a fixed action. A double integrator is different: in a normalized state (x,v)(x,v) with A=[10.501]A=\left[\begin{smallmatrix}1&0.5\\0&1\end{smallmatrix}\right], its Euclidean operator norm is about 1.2807761.280776, not 11. Its unit eigenvalues do not give a unit operator-norm bound: AtA^t contains 0.5t0.5t. Constant velocity bias can produce quadratic position discrepancy. Halving ε\varepsilon halves the additive part of this envelope, not necessarily the realized error or a nonzero initial-error term.
  • L>1L>1: the envelope contains a geometric factor. To judge a horizon, compare the full bound, including ε\varepsilon and e0e_0, with a relevant state-error tolerance. For L=1.05L=1.05, ε=0.001\varepsilon=0.001, and e0=0e_0=0, the bound first exceeds 11 at step 8181, not after a generic handful of steps. An exact surrogate with matched initial state and actions has zero discrepancy at every horizon, even when the reference is expansive.

Three caveats matter for interpreting that formula.

First, the bound is not an equality. Propagated contributions can cancel, and loose constants or directions with low sensitivity can also leave the realized error well below it. For a scalar recurrence dt+1=λdt+btd_{t+1}=\lambda d_t+b_t, the contribution of bkb_k at time tt is λt−1−kbk\lambda^{t-1-k}b_k. Cancellation concerns these weighted contributions, not simply the sum of raw bias signs. Alternating biases need not remain bounded when λ>1\lambda>1; the examples below illustrate particular traces rather than a general cancellation guarantee.

Second, it is not a theorem that realized error increases monotonically. You can have a trajectory where the reference and the surrogate happen to drift back together for a few steps before separating again. The bound describes the envelope, not the path. Reporting a single number for "rollout error at horizon TT" collapses a whole error trace into one value and hides this shape.

Third, both trajectories must remain inside the domain where (A1)–(A2) hold. In a scalar coefficient-mismatch example, f^(s)−f(s)=δs\hat f(s)-f(s)=\delta s with δ≠0\delta\ne0, no finite uniform approximation bound exists over the whole real line. For a vector mismatch ΔAs\Delta A s, growth depends on direction and can vanish on the mismatch's nullspace. An additive bias instead admits a uniform bound on its norm. Leaving a validated domain invalidates its local bound unless it is extended or re-established.

Let's make it concrete with a scalar example. Take

st+1=λst,s^t+1=λs^t+ε,s_{t+1} = \lambda s_t, \qquad \hat{s}_{t+1} = \lambda \hat{s}_t + \varepsilon,

Here ε≥0\varepsilon\ge0 is the constant additive bias and L=∣λ∣L=|\lambda|. There is no action in this scalar illustration; adding the same input term BatBa_t to both maps would cancel in their difference, not be absorbed into λ\lambda. With d0=0d_0=0 and λ≥0\lambda\ge0, the constant-bias trace attains the geometric-sum envelope. Negative λ\lambda or changing bias signs can make it strict. The approximation error is state-independent, so this example separates propagation from changing local approximation quality.

In[6]:
Code
def signed_error_trace(lam, eps, T, alt_sign=False, e0=0.0):
    """Signed propagation of an additive bias eps through a scalar linear map."""
    e = e0
    trace = [e]
    for t in range(T):
        bias = (eps if (t % 2 == 0) else -eps) if alt_sign else eps
        e = lam * e + bias
        trace.append(e)
    return np.array(trace)


def geometric_bound(lam, eps, T, e0=0.0):
    """Evaluate the nonnegative geometric envelope without a near-unit branch.

    This is floating-point evaluation, not directed-rounding certification.
    """
    if (
        isinstance(T, (bool, np.bool_))
        or not isinstance(T, (int, np.integer))
        or T < 0
    ):
        raise ValueError("T must be a nonnegative integer")
    if not np.all(np.isfinite([lam, eps, e0])) or eps < 0:
        raise ValueError(
            "lam and e0 must be finite; eps must be finite and nonnegative"
        )
    L = abs(float(lam))
    bounds = np.empty(T + 1, dtype=float)
    bounds[0] = abs(float(e0))
    with np.errstate(over="ignore"):
        for t in range(T):
            bounds[t + 1] = L * bounds[t] + float(eps)
    return bounds


T_STEPS = 20
EPS_BIAS = 0.05

error_profiles = {
    lam: {
        "constant": signed_error_trace(lam, EPS_BIAS, T_STEPS),
        "alternating": signed_error_trace(
            lam, EPS_BIAS, T_STEPS, alt_sign=True
        ),
        "bound": geometric_bound(lam, EPS_BIAS, T_STEPS),
    }
    for lam in (0.9, 1.0, 1.05)
}

Now let's visualise the three regimes in separate figures, and add the alternating-sign case to show how loose the bound can be.

Out[7]:
Visualization
Constant-sign error matches a saturating bound; alternating-sign error is smaller after step one.
Contractive reference with lambda = 0.9: the constant-sign additive-bias trace equals the geometric bound at each step, while the alternating-sign trace is below it after step one. At step twenty the absolute errors are approximately 0.4392 and 0.0231, respectively, a magnitude ratio of nineteen; the asymptotic bound is 0.5.
Constant-sign error grows linearly; alternating-sign error alternates between zero and 0.05.
Reference with lambda = 1.0: constant-sign error and the upper bound grow linearly, while alternating-sign errors cancel every second step. The curves illustrate that equal per-step error magnitudes need not give equal recursive errors.
Constant-sign error grows geometrically; alternating-sign error never exceeds the same bound.
Expansive reference with lambda = 1.05: constant-sign error equals the geometric bound and exceeds 1.6 at step twenty. Alternating-sign error never exceeds the bound and is strictly lower after step one in this construction; alternation does not establish a general stability guarantee.

The contractive panel illustrates both tightness and cancellation. For λ=0.9\lambda = 0.9 and ε=0.05\varepsilon = 0.05, the constant-sign trace equals the geometric bound at each step and approaches ε/(1−λ)=0.5\varepsilon / (1-\lambda)=0.5. At step twenty, the constant-sign signed trace is approximately +0.4392+0.4392 and the alternating-sign signed trace is approximately −0.0231-0.0231. Their absolute errors, the quantities plotted, are approximately 0.43920.4392 and 0.02310.0231, with a magnitude ratio of nineteen. Both traces start at zero and coincide at step one, so the reduction is not present at every step. Using the bound as a planning constraint may be conservative when errors cancel, but it is tight in the aligned-sign construction. The same uniform per-step error magnitude can therefore accompany different recursive errors. These deliberately chosen signed-error patterns illustrate a mechanism; they do not establish where every trained model falls between them.

These observations suggest several questions to ask when choosing a rollout horizon.

  • Horizon effects depend on sensitivity, bias structure, initial error, and the relevant task tolerance. The three plotted references differ in λ\lambda; they are not one fixed model with three behaviors. Shortening a horizon can lower the envelope even in a contractive system close to L=1L=1, and an expansive upper bound need not describe the realized error. Compare proposed horizons on the actual system instead of labeling one regime harmless and another inevitably unusable.
  • One-step error magnitude is not a complete rollout description; error directions and temporal structure also matter. A related strategy in Janner et al.'s MBPO paper (NeurIPS 2019) uses short model rollouts branched from real states and studies rollout length alongside model and policy-distribution error. That is not a theorem about the particular alternating-bias experiment here, nor a guarantee that any short rollout has small error.
  • Talvitie's self-correcting-model work (2017) trains on model-generated inputs paired with next states from a parallel reference trajectory. Its analysis uses restricted conditions including deterministic reference dynamics and blind rollout policies, with shared action sequences. This hallucinated-training target is not generally the same as ∥f^(z,a)−f(z,a)∥\|\hat f(z,a)-f(z,a)\| at one common input, so improving it does not automatically lower (A1)'s uniform ε\varepsilon. It does not change a fixed reference map's uniform LL on a fixed domain and norm; those statements should not be mistaken for this bound proving the method's guarantees.
  • Matched actions, rather than how those actions were generated, are the key comparison. The derivation remains pathwise valid for a shared feedback-generated action sequence if (A1)–(A2) hold along both paths. If each path chooses its own action, additionally bound the action difference: with reference action sensitivity LaL_a and state sensitivity LsL_s, et+1≤Lset+La∥a^t−at∥+εe_{t+1}\le L_s e_t+L_a\|\hat a_t-a_t\|+\varepsilon on the required domains. Feedback may damp or amplify this mismatch. Neither inequality alone proves closed-loop stability or task safety.

Compounding can occur while evaluating candidate rollouts as well as while executing a chosen plan. Next we distinguish a wrong task objective from prediction bias, then ask what selection does to attractive scoring errors.

Model Bias and Objective Misspecification

A model-based agent can fail because its predictions are wrong, because its supplied objective is wrong, or because both interact. Bad trajectories and low task performance do not by themselves distinguish these possibilities. The next controlled example keeps dynamics exact and changes only the objective.

Model bias vs. objective misspecification

Model bias is discrepancy in surrogate dynamics or a surrogate reward head. From rest in the opening toy, commanding a=4a=4 predicts next velocity 22 under the unclipped surrogate, while reference clipping gives 0.50.5. The actuator caps the command, not the vehicle's velocity at all future times. Better identification, known actuator structure, or constraints on queried actions can address this discrepancy or its consequences.

Objective misspecification is discrepancy between the supplied optimization target and the intended task. If the task requires stopping and the cost omits terminal velocity, accurate dynamics do not add that preference. Add an appropriate velocity term when it is a preference, or impose a stopping constraint when it is a hard requirement. A finite penalty and a hard constraint are not interchangeable.

The two failure axes are logically distinct but can coexist and interact. Changing a fixed dynamics map does not rewrite a fixed objective, and changing a fixed objective does not correct the map. Either intervention can still change which plan is selected and mitigate observed behavior. The following example isolates objective mismatch with exact dynamics rather than proving repairs never overlap.

The following counterexample keeps the dynamics exactly correct and changes only the objective. There is no transition-prediction error to remove: improving transition accuracy alone cannot supply an omitted stopping preference while the objective, candidate set, and planning protocol remain fixed. Other model-side interventions can change selected behavior; this example does not rule them out. Consider a double integrator with exact dynamics

xt+1=xt+vt,vt+1=vt+at,x_{t+1} = x_t + v_t, \qquad v_{t+1} = v_t + a_t,

starting from (x0,v0)=(0,0)(x_0, v_0) = (0, 0). The action is bounded by ∣at∣≤1|a_t| \le 1. Over four steps, a plan (a0,a1,a2,a3)(a_0, a_1, a_2, a_3) produces

x4=3a0+2a1+a2,v4=a0+a1+a2+a3.x_4 = 3a_0 + 2a_1 + a_2, \qquad v_4 = a_0 + a_1 + a_2 + a_3.

With continuous ∣ai∣≤1|a_i|\le1 and four unit-time steps from rest, the reachable terminal position lies in [−6,6][-6,6]. Position alone does not specify stopping: [1,1,−1,−1][1,1,-1,-1] reaches x4=4x_4=4 with v4=0v_4=0. Indeed, under the rest condition ∑iai=0\sum_i a_i=0, x4=1.5a0+0.5a1−0.5a2−1.5a3≤4x_4=1.5a_0+0.5a_1-0.5a_2-1.5a_3\le4, with equality only at that plan. The proxy below penalizes terminal position error plus action energy but omits terminal velocity. Reaching the same position with nonzero velocity is therefore not ruled out.

Jproxy(a)=(x4−4)2+10−3∑tat2,Jintended(a)=(x4−4)2+v42+10−3∑tat2.J_{\text{proxy}}(a) = (x_4 - 4)^2 + 10^{-3} \sum_t a_t^2, \qquad J_{\text{intended}}(a) = (x_4 - 4)^2 + v_4^2 + 10^{-3} \sum_t a_t^2.

Both objectives are soft costs. The velocity term makes moving arrival more expensive; it does not reject every nonzero v4v_4. We will enumerate a grid of five action values per step, giving 54=6255^4=625 plans, and compare its two minimizing plans. On this grid the intended-cost optimum stops. That need not be the continuous optimum: [1,1,−1,−0.999][1,1,-1,-0.999] reaches x4=4x_4=4 with v4=0.001v_4=0.001 and intended cost 0.0039990010.003999001, below the stopping plan's 0.0040.004. Require v4=0v_4=0 explicitly if exact stopping is part of feasibility.

In[8]:
Code
A_GRID = np.array([-1.0, -0.5, 0.0, 0.5, 1.0])
PLAN_LEN = 4
X_TARGET = 4.0
ACTION_PENALTY = 1e-3

plans = np.array(list(itertools.product(A_GRID, repeat=PLAN_LEN)))

final_x = 3.0 * plans[:, 0] + 2.0 * plans[:, 1] + plans[:, 2]
final_v = plans.sum(axis=1)
action_penalty = ACTION_PENALTY * np.sum(plans**2, axis=1)

proxy_cost = (final_x - X_TARGET) ** 2 + action_penalty
intended_cost = (final_x - X_TARGET) ** 2 + final_v**2 + action_penalty

idx_proxy = int(np.argmin(proxy_cost))
idx_intended = int(np.argmin(intended_cost))

plan_proxy = plans[idx_proxy]
plan_intended = plans[idx_intended]
Out[9]:
Console
Number of enumerated plans: 625

Proxy-optimal plan    [1.0, 0.5, 0.0, 0.0]
  proxy cost          :   0.001250
  intended cost       :   2.251250
  final x, final v    :  4.000,  1.500

Intended-optimal plan [1.0, 1.0, -1.0, -1.0]
  proxy cost          :   0.004000
  intended cost       :   0.004000
  final x, final v    :  4.000,  0.000

The proxy's grid optimum, [1,0.5,0,0][1,0.5,0,0], reaches x4=4x_4=4 with v4=1.5v_4=1.5: proxy cost 0.001250.00125, intended cost 2.251252.25125. The intended cost's grid optimum, [1,1,−1,−1][1,1,-1,-1], stops at the target with cost 0.0040.004. This is an objective mismatch under exact dynamics. Moving arrival is penalized by the stated task; no collision or crash process is modeled. The calculation establishes the finite-grid contrast, not exact stopping for every continuous optimum.

The two objectives disagree not about position but about the velocity at arrival, and the selected plans make that tension visible.

In[10]:
Code
selection_x = final_x.copy()
selection_v = final_v.copy()
selection_proxy_index = idx_proxy
selection_intended_index = idx_intended
Out[11]:
Visualization
Scatter of final position versus velocity for 625 grid plans, highlighting the proxy grid optimum at the target but moving and the intended-cost grid optimum at the target at rest.
Final position and velocity of the 625 enumerated grid plans under exact dynamics. The proxy grid optimum reaches the target with velocity 1.5, while the intended-cost grid optimum arrives at rest. These soft objectives select different plans on this grid; the continuous optimum need not stop exactly.

The opening actuator failure and this objective failure isolate different defective objects. Adding terminal velocity to a cost does not make an unclipped map obey clipping; adding clipping does not express a missing task preference. A penalty or constraint can nevertheless steer selection away from a harmful candidate while the map remains wrong. Diagnose whether the intervention repairs predictions, aligns the target, or merely changes the chosen behavior.

A further distinction is reward-channel integrity. Everitt et al.'s analysis of reward tampering concerns inappropriate influence on the reward function or its inputs. Optimizing a fixed misspecified proxy need not involve such influence; tampering and misspecification can also coexist. Reward-channel protections, incentive design, objective specification, and model-scoring checks are complementary, not disjoint professional domains. Pan et al.'s proxy-reward study is relevant to optimization against an imperfect proxy, not evidence that every proxy/true-performance divergence is tampering.

The defects can coexist: a controller could combine an unclipped actuator model with a position-only proxy for a task that also requires stopping. This is a constructed possibility, not evidence of deployment prevalence. Separate questions lead to separate fixes:

  • Does the transition or reward head predict the reference quantities accurately on the inputs used for the decision? Test the relevant map and inputs, rather than inferring dynamics bias from a return gap alone.
  • Does the supplied objective express the intended task when evaluated on reference trajectories? Check missing terms, constraints, timing, discounts, and termination conventions.
  • Does selection preferentially favor favorable scoring errors? This can amplify consequences of an existing model defect; it is not an alternative that proves the model itself is sound.

Test prediction accuracy, task alignment, and selection effects under controlled comparisons. More than one can require attention. Changing one component at a time helps interpret a contrast, but downstream behavior and interactions mean a successful repair does not uniquely identify every cause.

Planner Exploitation and Adversarial Trajectories

Random shooting can propose and evaluate candidates uniformly in an action box while selecting them nonuniformly by predicted score. With multiple iid candidates and a fixed finite score that varies under the proposal law, minimum-score selection favors lower scores and changes that law. With one candidate, or constant scores and first-index tie-breaking, the selected law is unchanged. This concerns selected plans, not uniform candidate proposal or evaluation. An optimistic selected prediction need not mean poor absolute task performance, but it warrants comparing prediction and execution under the same conventions.

For costs, define selected optimism as Jref−J^J_{\rm ref}-\hat J after minimizing J^\hat J; for returns, use R^−Rref\hat R-R_{\rm ref} after maximizing R^\hat R. The relevant question is conditional error among selected candidates. Zero correlation between a score and its error does not imply independence or zero selected mean error. For example, let QQ be uniform on {−1,0,1}\{-1,0,1\} and E=Q2−2/3E=Q^2-2/3. Both E[E]\mathbb E[E] and Cov⁡(Q,E)\operatorname{Cov}(Q,E) are zero, yet selecting the largest of four independent QQ draws gives expected selected E=4/27E=4/27. A sufficient zero-mean condition is E[Ei∣Q1,…,QN]=0\mathbb E[E_i\mid Q_1,\ldots,Q_N]=0 for every candidate, with selection based on those scores. The toy below does not establish such a condition or a general monotonic relation between correlation and optimism.

A large prediction error, or a large return gap between predicted and executed, is not by itself proof of exploitation. Several other things can produce a gap:

  • Reward timing. If the reference emits rewards one step later than the training data assumed, even a perfect dynamics model can look wrong.
  • Discounting. A different discount factor or terminal condition can change the return without model exploitation.
  • Terminal conditions. Early termination can change return relative to a surrogate that runs to horizon, depending on omitted rewards and the terminal-payoff convention.
  • Random-stream differences. If the reference and the surrogate share a random seed but not the same noise realization, the gap can be an artifact.
  • Implementation errors. A sign flip, an off-by-one in the buffer, or a shape mismatch can produce gaps that look like exploitation but are bugs.

A protocol mismatch can be plan-dependent and budget-sensitive too. With exact dynamics, consider reward sequences (1,0)(1,0) and (0,4)(0,4). A scorer using discount 0.50.5 ranks them as 11 and 22, while reference undiscounted returns are 11 and 44. Enlarging a nested pool from the first plan to both changes the selected reference-minus-predicted return gap from 00 to 22. That sign is pessimism under the return-optimism convention, but the example suffices to refute budget invariance. Check timing, discount, termination, and reward implementations before attributing a changing gap to model exploitation.

The laboratory controls these alternatives by sharing the cost definition, initial states, candidate pools, and termination rules, and by rescoring candidates with exact reference dynamics. A gap that changes with search under these controls supports selection-sensitive scoring error in this toy. Exploitation need not be off-support, need not worsen monotonically with budget, and need not produce a nonzero slope if the same optimistic candidate wins every pool.

Let's operationalize this. We will run the same planner on two different scoring models: the surrogate and the reference (which we can use as a "scoring oracle" because we control the reference process in simulation). Both use the same action pool, the same horizon, and the same objective. The only difference is which dynamics they roll out with. And then we cross-evaluate: every plan selected by either scorers is executed against the reference.

For each initial state, draw one pool from the fixed box [−5,5]4[-5,5]^4 and reuse nested prefixes at increasing budgets. Retaining earlier candidates and keeping their scores fixed makes the minimum predicted cost nonincreasing. The proposal support does not widen with budget; more draws only provide more opportunities within that same box. Executed cost and optimism need not be monotone.

In[12]:
Code
N_GRID = np.array([8, 16, 32, 64, 128, 256, 512, 1024, 2048])
POOL_SIZE = int(N_GRID[-1])
K_CASES = 60
PLAN_SEED = 20240117

rng_planner = np.random.default_rng(PLAN_SEED)

pred_surrogate = np.zeros((K_CASES, len(N_GRID)))
exec_surrogate = np.zeros((K_CASES, len(N_GRID)))
pred_reference = np.zeros((K_CASES, len(N_GRID)))
exec_reference = np.zeros((K_CASES, len(N_GRID)))
naive_baseline = np.zeros(K_CASES)

for k in range(K_CASES):
    s0_k = rng_planner.uniform(-0.05, 0.05, size=2)
    pool = rng_planner.uniform(-A_PLAN, A_PLAN, size=(POOL_SIZE, H))

    surrogate_scores = np.array(
        [trajectory_cost(rollout(surrogate_step, s0_k, plan)) for plan in pool]
    )
    reference_scores = np.array(
        [trajectory_cost(rollout(reference_step, s0_k, plan)) for plan in pool]
    )

    naive_baseline[k] = trajectory_cost(
        rollout(reference_step, s0_k, np.ones(H))
    )

    for i, n in enumerate(N_GRID):
        idx_sur = int(np.argmin(surrogate_scores[:n]))
        pred_surrogate[k, i] = surrogate_scores[idx_sur]
        exec_surrogate[k, i] = reference_scores[idx_sur]

        idx_ref = int(np.argmin(reference_scores[:n]))
        pred_reference[k, i] = reference_scores[idx_ref]
        exec_reference[k, i] = reference_scores[idx_ref]

mean_pred_surrogate = pred_surrogate.mean(axis=0)
mean_exec_surrogate = exec_surrogate.mean(axis=0)
mean_pred_reference = pred_reference.mean(axis=0)
mean_exec_reference = exec_reference.mean(axis=0)

se_pred_surrogate = pred_surrogate.std(axis=0, ddof=1) / np.sqrt(K_CASES)
se_exec_surrogate = exec_surrogate.std(axis=0, ddof=1) / np.sqrt(K_CASES)
se_reference = pred_reference.std(axis=0, ddof=1) / np.sqrt(K_CASES)

mean_optimism = (exec_surrogate - pred_surrogate).mean(axis=0)
se_optimism = (exec_surrogate - pred_surrogate).std(axis=0, ddof=1) / np.sqrt(
    K_CASES
)
mean_naive_baseline = float(naive_baseline.mean())

The unit of replication is a case: one initial state from U([−0.05,0.05])2\mathcal U([-0.05,0.05])^2 and one candidate pool of 20482048 plans from U([−5,5])4\mathcal U([-5,5])^4. The 6060 cases are paired across scorers and nested budgets; the baseline executes (1,1,1,1)(1,1,1,1) on each case. Pairing controls these inputs rather than removing all case variation. The cost plot reports separate means and case-level standard errors, not errors across candidate draws. The optimism plot first forms each case's executed-minus-predicted difference, then reports its mean and standard error. For any paired outcomes, Var⁡(Y1−Y0)=Var⁡(Y1)+Var⁡(Y0)−2Cov⁡(Y1,Y0)\operatorname{Var}(Y_1-Y_0)=\operatorname{Var}(Y_1)+\operatorname{Var}(Y_0)-2\operatorname{Cov}(Y_1,Y_0). Common additive effects cancel, but heterogeneous effects remain; pairing reduces variance relative to independent outcomes when the corresponding covariance is positive, not unconditionally.

Now the results.

Out[13]:
Visualization
Line chart of predicted and executed task cost against log base two candidate count, with the reference-scored control tracking predicted equals executed and a naive baseline line.
Comparison of predicted and executed task cost as the planner's candidate pool grows. The surrogate's mean predicted cost falls monotonically because the pool is nested. Its mean executed cost fluctuates, peaks at 256 candidates, and then decreases through 2048; at 2048 it remains worse than the eight-candidate result and the naive baseline. The reference-scored planner predicts and executes the same value by construction, so it serves as a control, and the naive full-thrust baseline is shown as a horizontal reference line.
A line plot of mean executed-minus-predicted cost for surrogate-selected plans across nested budgets. Mean optimism initially increases, peaks at 256 candidates, then declines slightly; shading shows one paired case-level standard error.
Mean selected-plan optimism across 60 paired cases, with shading of one case-level standard error of the paired difference. The action box stays fixed. In this run optimism rises at smaller budgets, peaks at 256 candidates, and then declines slightly while remaining positive; it is not a monotonic-growth law or a measured support-tail effect.

Three patterns are visible in the figures and worth stating, together with what is not guaranteed.

With nested pools and unchanged scores, adding candidates cannot increase the minimum predicted cost. That limited guarantee does not cover arbitrary planners, noisy rescoring, or replaced pools, and it says nothing monotonic about executed cost.

Every budget uses the same [−5,5]4[-5,5]^4 action box. More candidates can reveal combinations with overly favorable surrogate scores, but do not widen support or necessarily increase chosen command magnitude. Off-range commands are not all optimistic: [−4,4,0,0][-4,4,0,0] has predicted cost 1313 and reference cost 5.68755.6875, giving cost optimism −7.3125-7.3125. The issue is selection of favorable errors, not membership outside the action interval alone.

The displayed optimism is positive at every tested budget and larger at high budgets than at the smallest, but it peaks and then declines. This particular search/scorer combination exposes optimism; it does not establish that every increase in search effort makes exploitation worse.

Exact deterministic reference scoring gives equal predicted and executed cost for each plan under the same initial state and conventions, hence zero per-plan optimism in this control. It selects the best plan in its sampled pool, not necessarily the global optimum. A zero gap does not require identical selected action sequences across all implementations or equally scored baselines, nor does statistical calibration guarantee equality for every stochastic realization.

Compare the selected plan with a declared no-search baseline on the tested objective. Worse cost is a regression on that metric; equal cost gives no measured improvement on it. Neither comparison by itself diagnoses exploitation or determines total deployment value, which may include other costs and constraints.

Ordinary optimization can produce adversarial-looking trajectories without an external attacker by selecting favorable scoring errors. The actuator example has saturated commands, not a modeled collision or damage process, so it does not establish that selected plans are physically unsafe to a human observer. Selection can amplify the consequences of existing errors without increasing each local state-prediction error.

An external adversary and ordinary score optimization differ in who chooses inputs and with what intent. Their failure mechanisms and defenses can overlap: constraints and score validation may address both. Part XII Chapter 3 is the planned security context; this example alone does not establish a disjoint class of mitigations.

A receding-horizon controller uses a new observation, state estimate, or belief to replan after executing part of a plan. It need not have direct access to the true state, and closed-loop control is not synonymous with repeated search. Replanning can change visited states and commitments enough to avoid or mitigate an exploitation pattern. It does not by itself correct a fixed scoring defect, and exploitation need not recur at every replanning call. Our laboratory executes open-loop sequences, so it does not measure that feedback effect.

State Aliasing, Forgetting, and Hallucinated Dynamics

Missing decision-relevant information can cap the best achievable context-dependent performance; it does not make every plan bad. A point prediction can also suffice when an objective depends only on a conditional mean. The cue example isolates information that is needed for the optimal action and absent from the current input. Supplying history or a belief can change the problem, rather than merely searching harder over the same insufficient input.

State aliasing

State aliasing

State or representation aliasing means that contexts collapsed to the same available representation require different predictions or different optimal action choices. Different required actions provide a control-relevant example, not an exhaustive definition of predictive aliasing.

If two equally likely cue histories produce the same junction input and require opposite actions, a deterministic current-input policy chooses the same action in both. A stochastic policy using only that input and cue-independent randomness has the same conditional action distribution, not necessarily the same realized draw. Either achieves expected reward at most 0.50.5 in this binary toy. A policy with a reliably retained cue can achieve 11.

Consider a two-room environment. At t=0t = 0 the agent observes a cue (A\text{A} or B\text{B}) that tells it which room contains the reward. From t=1t = 1 onward the agent is at a junction whose observation is identical in both cases. At t=2t = 2 the agent must choose "left" or "right"; only the matching action pays reward 11.

In[14]:
Code
CUE_TOKENS = ("A", "B")
JUNCTION = "junction"
CUE_TARGET = {"A": "left", "B": "right"}
prior = {c: 0.5 for c in CUE_TOKENS}


def observation_history(cue, steps_after_cue=2):
    """The observation stream: the cue is visible only at t = 0."""
    return [cue] + [JUNCTION] * steps_after_cue


observed_histories = {c: observation_history(c) for c in CUE_TOKENS}

## An observation-only policy sees only the current observation. At the
## junction, both histories collapse to the same input.
aliased_histories = [
    h for h in observed_histories.values() if h[-1] == JUNCTION
]
best_memoryless_value = max(
    sum(prior[c] for c in CUE_TOKENS if CUE_TARGET[c] == action)
    for action in ("left", "right")
)

## A policy that integrates history can recover the cue and act optimally.
best_memory_value = sum(prior[c] for c in CUE_TOKENS)
Out[15]:
Console
cue =  A  history = ['A', 'junction', 'junction']  correct action = left
cue =  B  history = ['B', 'junction', 'junction']  correct action = right

Histories that share the current observation 'junction': 2
Best observation-only value : 0.50
Best history-integrating value: 1.00

The current-observation-only controller cannot distinguish these cue histories. Feed-forward architecture alone is not the obstruction: a feed-forward function given an explicit history window containing the cue can solve the toy. A one-bit memory can also solve it if the cue is encoded, retained, and read out correctly; nominal recurrent capacity does not guarantee those behaviors.

Paired histories with matching present observations and different action requirements are one useful aliasing test. In a fully specified finite process, representation and transition/reward equivalence can also be checked exhaustively. Passing a sampled suite supports only its tested histories, delays, and metrics; even familiar but untested contexts may fail. Part XI Chapter 3 discusses state, memory, and occlusion diagnostics, not this specific cue-pair protocol. Aliasing and forgetting can coexist, and additional data can improve a learned encoding when the required information is available.

Forgetting

Forgetting is loss of retained information or of previously achieved performance. It need not mean complete ineffectiveness. Separate memory lost during fixed-parameter inference from old-task performance lost after parameter updates; these are different comparisons and can interact.

Within-episode forgetting occurs when information that was encoded is later lost during inference at fixed parameters. A model that never encoded the cue has a different failure. Insufficient long-sequence training or unstable memory can contribute, but neither inadequate training length nor nominal memory capacity determines the outcome. A fixed bit latch retains a cue indefinitely; a four-step input window can instead produce accuracy 11 at delays zero through three and 0.50.5 thereafter. Degradation can therefore be gradual or abrupt. Test delay with fixed parameters rather than identifying the mechanism from curve shape.

Across-update forgetting is deterioration on an old task following new parameter updates under a controlled evaluation protocol. Kirkpatrick et al.'s EWC work (2017) protects important parameters with a quadratic penalty. A decline can be gradual across many optimizer steps or abrupt, and its magnitude can depend on episode length if recurrent dynamics change. Compatible tasks need not cause a decline at all. Compare retained-task performance before and after updates; a mandatory sharp cliff or delay independence is not part of the definition.

Cross two controls: sweep cue delay while parameters are frozen, and compare a fixed retained evaluation set before and after controlled sequential updates. Stratify the latter by delay when relevant. Record encoding, retention, and readout behavior where possible. The scalar experiment below evaluates completed training phases; it does not trace every optimizer step or demonstrate a sharp update-time cliff.

Within-episode retention may benefit from stable explicit memory, appropriate capacity, or longer-sequence training. Across-update retention may benefit from parameter protection, replay, or freezing updates. These interventions are not exclusive: protecting recurrent dynamics can preserve cue retention, and replay can affect both. Diagnose which comparison changed rather than assuming one remedy can tell you nothing about the other mechanism.

The next experiment uses one scalar weight for two incompatible targets on the same inputs: task A predicts xx, task B predicts −x-x. Its importance weight q=mean⁡(x2)q=\operatorname{mean}(x^2) is a second-moment/curvature proxy for an EWC-style quadratic anchor, not a computed empirical Fisher. For Gaussian linear regression with known variance σ2\sigma^2, expected Fisher is E[x2]/σ2\mathbb E[x^2]/\sigma^2. Averaging the conditional Fisher over this toy's fixed input set gives q/σ2q/\sigma^2, hence qq at unit variance. Score-squared empirical Fisher on noiseless perfectly fitted data would instead be zero. Keep this likelihood and averaging distinction when interpreting the toy penalty.

In[16]:
Code
rng_forget = np.random.default_rng(7)
X_TASK = rng_forget.uniform(-1.0, 1.0, size=400)
Y_TASK_A = X_TASK.copy()
Y_TASK_B = -X_TASK.copy()


def fit_scalar(
    w, xs, ys, lam_ewc=0.0, w_ref=0.0, fisher=0.0, steps=2000, lr=0.05
):
    """Fit scalar regression with an optional EWC-style quadratic anchor."""
    for _ in range(steps):
        pred = w * xs
        grad = 2.0 * np.mean(xs * (pred - ys))
        if lam_ewc > 0.0:
            grad += 2.0 * lam_ewc * fisher * (w - w_ref)
        w -= lr * grad
    return w


def sequence_task_a_then_b(lam_ewc):
    w = 0.0
    w = fit_scalar(w, X_TASK, Y_TASK_A)
    loss_a_after_a = float(np.mean((w * X_TASK - Y_TASK_A) ** 2))
    w_a = w
    fisher = float(
        np.mean(X_TASK**2)
    )  # second-moment importance proxy, not empirical Fisher
    w = fit_scalar(
        w, X_TASK, Y_TASK_B, lam_ewc=lam_ewc, w_ref=w_a, fisher=fisher
    )
    loss_a_after_b = float(np.mean((w * X_TASK - Y_TASK_A) ** 2))
    loss_b_after_b = float(np.mean((w * X_TASK - Y_TASK_B) ** 2))
    return w_a, w, loss_a_after_a, loss_a_after_b, loss_b_after_b


EWC_LAMBDAS = np.array([0.0, 0.5, 1.0, 3.0, 9.0, 27.0])
forgetting_records = [sequence_task_a_then_b(lam) for lam in EWC_LAMBDAS]
Out[17]:
Console
lambda_ewc  w_after_B   loss_A|A   loss_A|B   loss_B|B
------------------------------------------------------
      0.00    -1.0000     0.0000     1.3667     0.0000
      0.50    -0.3333     0.0000     0.6074     0.1519
      1.00    -0.0000     0.0000     0.3417     0.3417
      3.00     0.5000     0.0000     0.0854     0.7687
      9.00     0.8000     0.0000     0.0137     1.1070
     27.00     0.9286     0.0000     0.0017     1.2708

With λewc=0\lambda_{\rm ewc}=0, task B is fitted to numerical precision, while retained task-A loss rises to approximately the loss of predicting its negative target. The finite optimizer run does not produce mathematically exact zero B loss. Increasing the penalty lowers the displayed A loss and raises B loss. No single scalar weight gives zero loss on both incompatible targets for nonzero inputs. That impossibility belongs to this one-weight problem; EWC can reduce or avoid forgetting in other settings and is not universally incapable of elimination. The columns loss_A|A, loss_A|B, and loss_B|B compare A immediately after its own training, A after B, and B after B. For the ideal converged quadratic problem with this anchor, the minimizing weight is (λewc−1)/(λewc+1)(\lambda_{\rm ewc}-1)/(\lambda_{\rm ewc}+1); the finite run approximates it. Choosing the penalty allocates error between the incompatible tasks rather than creating capacity to satisfy both.

This measures old-task loss after sequential training on a new target, using the same fixed inputs also used for fitting. It is not a held-out generalization result or a neural-architecture benchmark. Maintain a fixed retained evaluation protocol and check persistent changes across updates, with separate delay tests for inference-time retention. Interventions can overlap; their usefulness must be measured rather than inferred from a mechanism label.

Hallucinated dynamics

Hallucinated dynamics are model-generated transitions that violate the reference process or a relevant constraint. A state may look plausible in isolation while its transition is impossible; other violations can be obvious from a single state. Plausibility and conditional validity are different checks.

Define a three-state corridor with positions L=−1-1, C=00, R=11. From C, the reference moves to L or R with probability 0.50.5 each. L and R are absorbing: once reached, they remain there. This complete kernel has no return to C. Construct a deterministic model that outputs the reference conditional mean at each of these states. It maps C to C, L to L, and R to R, so its recursive rollout from C stays at C indefinitely. The following code propagates both the full reference kernel and that constructed mean map; no regression fitting is claimed.

In[18]:
Code
## State order is shared by the complete reference kernel and the plots.
CORRIDOR_STATES = ("L", "C", "R")
CORRIDOR_POSITIONS = {"L": -1.0, "C": 0.0, "R": 1.0}
p_left = 0.5
if p_left != 0.5:
    raise ValueError("This symmetric point-map example requires p_left = 0.5")
reference_kernel = np.array(
    [
        [1.0, 0.0, 0.0],
        [p_left, 0.0, 1.0 - p_left],
        [0.0, 0.0, 1.0],
    ]
)
assert np.all(reference_kernel >= 0.0)
np.testing.assert_allclose(reference_kernel.sum(axis=1), 1.0)
corridor_positions = np.array([CORRIDOR_POSITIONS[s] for s in CORRIDOR_STATES])
conditional_means = reference_kernel @ corridor_positions
mean_by_position = dict(zip(corridor_positions, conditional_means))
mean_next_position = float(conditional_means[CORRIDOR_STATES.index("C")])

## Propagate the constructed point map from C, not a fitted regressor.
HALLUCINATION_HORIZON = 5
mean_rollout_positions = [CORRIDOR_POSITIONS["C"]]
for _ in range(HALLUCINATION_HORIZON):
    mean_rollout_positions.append(
        float(mean_by_position[mean_rollout_positions[-1]])
    )

## Propagate the full reference distribution, including the absorbing exits.
reference_distribution = np.array([0.0, 1.0, 0.0])
reference_probability_of_midpoint = [float(reference_distribution[1])]
for _ in range(HALLUCINATION_HORIZON):
    reference_distribution = reference_distribution @ reference_kernel
    reference_probability_of_midpoint.append(float(reference_distribution[1]))
Out[19]:
Console
Conditional-mean next position: +0.0000 (this is CORRIDOR_POSITIONS['C'] = +0.0)

Deterministic conditional-mean rollout:
  positions: [0.0, 0.0, 0.0, 0.0, 0.0, 0.0]

Reference probability of being at the midpoint:
  p(state = C at step t): [1.0, 0.0, 0.0, 0.0, 0.0, 0.0]

With these absorbing exits, the reference has zero probability of C at every t≥1t\ge1 from the stated initial state. The constructed point model instead remains there. The population conditional mean minimizes expected squared error among deterministic point predictions with finite variance, but finite regression training need not attain it. A planner that treats this point forecast as a feasible next state may plan for remaining at C; a support-aware or distribution-aware planner can reject that interpretation. The unconditional claim that every planner will plan accordingly would be false.

The mismatch is visible as a difference in where probability mass sits at the next step.

In[20]:
Code
state_labels = list(CORRIDOR_STATES)
reference_mass = reference_kernel[CORRIDOR_STATES.index("C")].copy()
model_mass = np.array([0.0, 1.0, 0.0])
positions = np.arange(len(state_labels))
bar_width = 0.35
Out[21]:
Visualization
Grouped bar chart comparing reference and conditional-mean next-state mass at the corridor midpoint, with the reference spread across left and right exits and the conditional mean concentrated entirely on the center state.
The fully specified corridor has absorbing exits L and R. From C, reference next-state mass is split equally between those exits, while the constructed conditional-mean point model puts all mass at C, outside that transition's support. The mean is optimal for population pointwise squared loss, not a valid sampled reference transition or a guarantee of finite-trained regression behavior.

Two things make hallucinated dynamics dangerous.

An unconditioned check that position C is a globally valid state misses this violation. A one-step check conditioned on the present state and action detects it immediately from the specified transition support. A trusted kernel, explicit physical constraint, or finite-state enumeration can supply that check; a separately learned probabilistic model is not the only option. Other hallucinations may also fail single-state validity tests.

Squared-loss optimality does not mean small loss. The minimum point-prediction MSE equals the conditional variance: it is 11 for equiprobable positions ±1\pm1, and 100100 at scale ±10\pm10. Nevertheless the optimal mean lies outside the next-state support in both examples. This violation is already visible in a one-step conditional check; composition shows its further consequences. Part XI Chapter 2 provides the distribution/calibration context, while Chapter 3 discusses physical constraints and recursive checks; it does not implement this corridor's conditional-support protocol.

The mean is a correct answer to a population squared-loss point-prediction question, not necessarily to a transition-sampling or planning question. Change the prediction target or training loss when appropriate, represent a faithful conditional distribution, enforce justified transition constraints, or have the planner use distributions rather than treating every mean as feasible. A distribution-valued head alone does not guarantee correct support. Changing the task objective is distinct from changing the model-training objective. Part V develops modeling choices; the next chapter is planned to examine control-side responses.

The opening hand-chosen shortcut also violates action-conditioned reference dynamics. From rest its first command predicts velocity 22, but the clipped reference permits only an increment of 0.50.5 at that step. Velocity 22 is not globally impossible: successive bounded commands can reach it. The problem is that state-action transition, not an invalid standalone velocity or the mere spelling of an out-of-range command. Its position trace can look plausible while the joint transition is wrong.

A Diagnostic Workflow for Untrusted Models

The examples suggest a first-pass diagnostic workflow for messier systems. Its order favors inexpensive checks and reusable records, not a theorem that every step eliminates exactly one cause. Some mechanisms can be present together or mask one another.

Controlled experiments can rule out hypotheses and, with adequate design and causal assumptions, establish positive intervention effects. They do not automatically prove a unique mechanism. The sequence below gathers evidence about the tested failure and records what remains uncertain; it is not an empirical ranking of the most frequent causes.

1. Capture and replay the trajectory. Save the initial state, actions, random-stream state, model and environment versions, configuration, and relevant external or hidden inputs. Attempt replay through the same code path. If it fails, investigate the capture boundary, scheduling, hardware, version drift, and external state rather than concluding a unique cause or abandoning diagnosis. Exact replay supports matched-trace substitutions. When it is unavailable, controlled repeated trials can still estimate failure rates and intervention effects with uncertainty; that is population evidence, not the original trajectory's individual counterfactual.

2. Verify the objective and reward conventions. Audit whether the supplied target expresses the intended task on reference trajectories. Separately compare predicted and reference reward or cost using matched timing, discount, termination, and aggregation. A discrepancy can arise from transition bias, reward-head bias, stochastic outcomes, or a protocol mismatch; it does not alone prove objective misspecification. The double-integrator example changes the target itself with exact dynamics. Reward checks are useful early checks here, not a claimed most-common cause based on an absent incidence study.

3. Compare teacher-forced and recursive prediction. At matched reference states and actions, teacher forcing isolates local discrepancies; recursion measures their propagation along predicted states. Small teacher-forced error with large recursive error motivates a propagation analysis, not a unique diagnosis. Large local error can propagate too: a map with constant bias 0.010.01 has one-step discrepancy 0.010.01 and discrepancy 11 after 100100 unit-sensitivity steps. Bias and compounding are not mutually exclusive. Plot traces and applicable bounds in a stated norm; a ratio is undefined when its denominator is zero and does not identify a regime by itself.

4. Inspect selected errors and joint support. Compute predicted and reference outcomes for selected plans with an explicit optimism sign. Inspect visited state-action or history-action combinations, not just separate action ranges: matching every marginal range need not cover their combinations. Compare nested budgets under fixed candidate/scoring conventions and use reference scoring controls. Large off-support optimism is evidence to investigate; exploitation can occur on-support, and protocol mismatches can also vary with budget. A positive slope is neither necessary nor sufficient to uniquely identify exploitation.

5. Substitute reference components under stated controls. Compare one intervention at a time, retaining comparable initial conditions, search budgets, and task conventions where possible. Observe how downstream selected actions and outcomes change; successful substitution need not identify a unique cause:

  • Replace surrogate dynamics with validated reference dynamics while holding the reward definition fixed. Improvement supports a dynamics-related contribution in that comparison, but interactions or compensating errors can prevent unique attribution.
  • Replace a predicted reward head with reference reward while holding surrogate dynamics fixed. This tests reward prediction. Separately replace a supplied proxy with the intended objective to test specification; those are different substitutions.
  • Execute a hand-chosen feasible comparator. Success shows that the reference can achieve that outcome, not that the surrogate is sound. The opening feasible plan succeeds while the same surrogate still mispredicts the shortcut; model and selection defects remain compatible with that result.
  • Supply a validated fully observable state or sufficient history in place of the current input. Improvement supports an information or representation-use limitation. Distinguish observation aliasing from encoding, retention, readout, and inference effects before assigning a label.

A substitution needs a trustworthy reference and an understood interface. Simulators do not always expose every relevant state, reward, or hidden process; a validated proxy supports only its established scope. Changing one component can change downstream inputs and decisions. Preventing a failure establishes a tested intervention effect, not its unique cause; persistent failure does not prove every remaining defect is downstream. Several interacting or compensating defects may remain.

6. Retest on diagnostic cases and fresh cases. Cases used to select or tune a repair are no longer an untouched holdout. Retest them to detect regressions, then evaluate a separate representative set with uncertainty and the same declared metric. Improvement on diagnostic cases but not fresh cases warns of overfitting, though sampling variation or shift can also explain it. Improvement in two finite sets is evidence, not proof, of population improvement. Preserve a final test set not used to choose the repair when making a confirmatory claim.

Record supported contributions, excluded hypotheses, tested repairs, and unresolved alternatives. Passing these checks does not certify general reliability. A counterexample can refute a universal guarantee; establishing one requires an argument covering the specified domain, sometimes through exhaustive finite checks or formal assumptions. That is different from a diagnostic checklist, but not universally impossible outside research.

Limitations and Impact

These risks can arise across architectures and training regimes; none is inevitable for every finite-data agent. Exact representable dynamics, sufficient information, an aligned objective, and controlled execution can eliminate particular mechanisms on a specified domain. The examples identify what to check, not an architecture-independent theorem that every agent fails.

Some defects admit direct repairs within a known scope. Align a missing task preference or enforce a justified hard safety constraint; a soft cost term alone does not guarantee safety. Expose sufficient decision information and ensure it is encoded and used; revealing one variable need not remove all aliasing. Faithful distribution modeling or correctly enforced support constraints can address the corridor error, but a generic probabilistic output head need not. Each repair still needs a test or argument for its stated scope.

Residual one-step errors can propagate, but compounding is not inevitable. If f^=f\hat f=f and the initial states and actions match, discrepancy is zero at every horizon, even for f(s)=2sf(s)=2s. For fixed maps on a fixed domain, ε\varepsilon and LL constrain an envelope; refitting, structural constraints, changed operating regions, or feedback can change the comparison. Short rollouts and suitable self-correction can mitigate some residual errors. Calibrated uncertainty can inform decisions, not automatically repair the mean dynamics. Choose horizons using sensitivity, task requirements, and measured error rather than assuming every long-horizon model must eventually fail.

Exploitation is a joint interaction of scoring error, selection, and task consequences; this does not mean the model has no defect. Aggressive search, limited joint state-action coverage, and stochastic scores can create risks, but no universal worst-case ordering follows from those labels. Support restrictions, shorter horizons, and uncertainty penalties can involve tradeoffs. A penalty requires justified error coverage and a suitable coefficient for a conservative guarantee; arbitrary disagreement or variance is not automatically an error bound. Tradeoffs are not inevitable: implementing the actuator's known clipping removes this toy's scoring defect without removing any physically achievable transition. Ask which tested defects can be repaired, which residual risks remain, and what each proposed mitigation costs.

Prediction measurements alone need not establish reliable decisions. They must be connected to the intended task, available information, selected inputs, and execution protocol. Empirical checks are valuable when those properties are uncertain; fully specified finite or formal models may establish scoped guarantees by other means. The chapter does not quantify where most real deployment effort is spent.

The next chapter, Part XII Chapter 2 on robustness, calibration, and safe control, is planned to develop uncertainty-aware and constrained control responses. Here we have suggested directions, not validated a general mitigation package. The question is how to act under specified model and information limitations. Robust control, Bayesian decision theory, and formal verification provide relevant approaches; not every deployment requires every approach or has inevitable compounding error.

Summary

The mechanisms are related but distinct and can coexist. Keep these qualifications with the names:

  • Compounding is propagation of prediction discrepancy. With a fixed norm, common realized actions, ε≥0\varepsilon\ge0, and a uniform L≥0L\ge0 reference bound on a domain containing both trajectories, et+1≤Let+εe_{t+1}\le Le_t+\varepsilon. The envelope is Lte0+ε∑k=0t−1LkL^te_0+\varepsilon\sum_{k=0}^{t-1}L^k, including e0+εte_0+\varepsilon t at L=1L=1. Its contractive, linear, and geometric regimes are upper-envelope regimes, not necessary actual trajectories. Shared feedback-generated actions are allowed; independently chosen actions need an additional mismatch bound.
  • Model bias is discrepancy in dynamics or reward prediction; objective misspecification is discrepancy between the supplied target and intended task. Repairing one fixed object does not rewrite the other, but behavioral constraints and penalties can mitigate consequences of either. Both can be present.
  • Model exploitation is selection of erroneously attractive scores. A prediction–execution gap or budget slope alone does not prove it. Controlled candidate pools, score conventions, reference rescoring, and joint-support checks help distinguish contributions; neither off-support behavior nor increasing optimism is required in every case.
  • Adversarial-looking trajectories can arise from ordinary optimization, without external attack. That describes selected behavior, not a universal safety judgment or a disjoint class of defenses.
  • Aliasing collapses contexts requiring different predictions or actions. More data cannot recover a cue absent from a fixed input, but data can improve an encoder when the relevant information is available. History or memory must retain and expose what the decision needs.
  • Within-episode forgetting loses encoded information during fixed-parameter inference; across-update forgetting loses retained performance after parameter changes. Delay sweeps and retained-set comparisons distinguish the experimental controls, not mandatory curve shapes. Their mechanisms and remedies can overlap. The scalar EWC-style tradeoff follows from incompatible one-weight targets, not a universal impossibility of preventing forgetting.
  • Hallucinated dynamics violate reference transitions or constraints. A conditional mean can fall outside support without an attacker or off-support input, as in the absorbing corridor. Optimal pointwise MSE need not be small, and a one-step conditional-support check can already expose the violation. Not every multimodal mean is invalid.

The workflow captures replay conditions, audits objectives and reward conventions, compares local and recursive predictions, probes selected errors and joint support, substitutes validated components, and retests on fresh cases. It gathers positive and negative evidence under controls; it does not automatically identify a unique cause or certify a whole model.

The next chapter is planned to ask how a controller should respond to specified model limitations, how uncertainty should be checked and used, and when a fallback layer is justified. The starting point is assessed risk, not an assumption that every failure mode must already be present.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about failure modes and model exploitation.

Failure Modes and Model Exploitation

Question 1 of 80 of 8 completed
In the compounding-error recursion e_{t+1} ≤ L e_t + ε, what does ε represent?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026failuremodes, author = {Michael Brenndoerfer}, title = {Failure Modes and Model Exploitation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Failure Modes and Model Exploitation. Retrieved from https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation
MLAAcademic
Michael Brenndoerfer. "Failure Modes and Model Exploitation." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation>.
CHICAGOAcademic
Michael Brenndoerfer. "Failure Modes and Model Exploitation." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Failure Modes and Model Exploitation'. Available at: https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Failure Modes and Model Exploitation. https://mbrenndoerfer.com/writing/world-model-failure-modes-and-model-exploitation

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.