Planning, Control, and Policy Evaluation

Michael BrenndoerferAugust 3, 202651 min read

Part of World Models Handbook

Evaluate composed model-planner-policy systems through closed-loop return, sample efficiency, planning horizon tradeoffs, exploitation, and transfer protocols.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Planning, Control, and Policy Evaluation

Suppose a learned-model planner predicts a return of −3.2-3.2 but earns −48-48 when its controller runs in the reference environment. These numbers are hypothetical, not outputs from the notebook below. The contrast matters because a predicted rollout and an executed closed loop answer different questions: the first is about the surrogate dynamics, while the second is about the model, planner, policy, and environment acting together.

A gap can arise from model error, but it can also reveal mismatched evaluation protocols or an implementation bug. It is not inevitable: a model that is exact on the visited trajectories, with matching rewards and controller random streams, can predict return exactly. Lower average one-step loss can nevertheless coexist with worse closed-loop return. Prediction loss averages over its evaluation distribution; a planner may query a different distribution and amplify errors that matter to action selection. Good planning does not require this shift, and unsuccessful planning can shift too. Measure both predicted and executed return before attributing their difference.

This chapter develops that comparison and the resource accounting around it. Chapter 62, Perceptual and Generative Evaluation, covers image-quality and distributional diagnostics. Chapter 63, State, Physics, Causality, and Memory Evaluation, covers state readouts, physical constraints, interventions, and memory diagnostics. Chapter 61 introduced the full evaluation ladder, including closed-loop decisions and transfer. Here we expand the decision-facing tests: what return does a composed controller earn on specified tasks, distributions, and budgets?

We examine four questions. Task return and sample efficiency connect achieved performance to real-interaction budgets. Planning horizon and compute tradeoffs compare search depth and breadth. Model exploitation and policy ranking test selection-sensitive prediction errors and the ordering of fixed candidate policies. Zero-shot, few-shot, and cross-domain transfer specify what stays frozen, what adapts, and which deployment conditions change.

These questions share a composed-system framework, but they are not marginals of one fixed trajectory distribution. Changing a planner, policy, or environment can change the distribution itself; functionally equivalent controllers or changes outside visited states need not do so. Ranking compares scores across policies; it need not analyze distribution tails. Transfer can change just one environment parameter, as it does below. The diagnostics can disagree: a biased return estimate can preserve policy order, and a costly planner can still achieve good return. Joint reporting makes those distinctions visible.

A concrete lesson runs through all four faces. The evaluation target is a composed system, not a component. When we say "the new model is better," we must always be able to say "better at what return, on what initial-state distribution, at what compute and interaction budget, and compared against which baselines on which held-out cases." Without that, the numbers are decoration.

Model quality has several separable dimensions. Predictive fidelity compares dynamics and reward predictions with reference outcomes under a stated distribution. Representation quality asks what task-relevant information a state representation makes accessible. Uncertainty quality includes calibration or coverage of specified predictive quantities, sharpness, and useful discrimination between uncertain and well-supported predictions. Decision usefulness asks what a controller achieves in execution. All four can use environment-derived reference data; the decision test additionally executes actions in a closed loop. Strong predictive or uncertainty scores alone do not certify control performance.

This separates two questions: is the model accurate under the chosen predictive metric, and is it useful for this decision problem? Small distribution-averaged errors can still matter greatly if they affect reward-sensitive states or action ordering. Conversely, errors in task-irrelevant quantities may have little effect on a particular controller. Name the metric and task rather than treating accuracy and usefulness as synonyms.

The Evaluation Target: Model, Planner, Policy, Environment

Let st∈Rns_t \in \mathbb{R}^{n} be the environment state, at∈Aa_t \in \mathcal{A} the action, oto_t the observation, and rtr_t the reward. The deterministic reference transition is st+1=f(st,at)s_{t+1}=f(s_t,a_t); a surrogate rollout substitutes f^\hat f for ff. Rewards here use the current state and action, rt=r(st,at)r_t=r(s_t,a_t). A policy selects actions from available observations or states. A deterministic feedback policy can be written at=π(st)a_t=\pi(s_t); random-shooting MPC also depends on its internal random stream, which we denote by ω\omega. The planner scores candidate sequences inside its model, while the executed policy applies the selected first action and replans. Keeping these roles distinct clarifies what each reported return measures.

Classical control distinguishes the plant, controller, and predictive model. In this notebook, identification minimizes an empirical next-state prediction loss, while control evaluation measures closed-loop return. Identification can also be designed around control-specific criteria; generic prediction loss is not its only possible objective. A planner may be embedded in a controller, but candidate scoring and executed feedback remain different operations.

Fully Observed Toy Convention

Throughout this chapter we use a fully observed toy where ot=sto_t = s_t exactly. This removes perception failure so we can study planning, control, and evaluation without a second layer of confusion. A controller that underperforms here is underperforming because of the model-planner-policy trio, not because of noisy sensors or representation error. Partial observability returns in the belief-space planning settings covered in Part VII: Planning and Agency.

With partial observability, a model-based belief update uses transition and observation likelihoods, and planning can operate on the resulting belief state. The belief-state formulation of POMDPs makes that distinction explicit. Setting ot=sto_t=s_t removes sensing and belief-estimation errors from this toy. It also means the examples cannot test failures specific to partial observability. Those settings are covered in Part VII.

The ladder from Chapter 61 covers one-step prediction, fixed-action open-loop rollouts, counterfactual interventions, closed-loop decisions, and transfer. These are complementary tests, not nested certificates. A model can perform well on a prediction metric yet support poor decisions; a controller can compensate for some prediction errors.

Teacher-forced prediction uses reference conditioning states rather than recursively predicted states. A fixed-action open-loop rollout removes observation corrections, allowing prediction errors to propagate without guaranteeing monotonic growth. Intervention tests can alter actions, states, or mechanisms; a selected-state reset is only one design. Closed-loop decisions evaluate the trajectories generated by the controller's own actions and feedback. Transfer evaluates specified changed conditions. Passing an early proxy does not certify a later task outcome, and failing it does not necessarily imply task failure.

Chapters 62 and 63 develop perceptual and state-level diagnostics; Chapter 61 introduces the full ladder, and this chapter expands its decision-facing tests. A controller that succeeds on one task has not thereby demonstrated faithful dynamics or general transfer. We report three kinds of information:

  • A return metric computed by rolling the closed-loop policy in the true environment on a fixed held-out set of initial states.
  • A resource accounting of the reported real-interaction and candidate-transition counts. This notebook does not measure wall-clock latency or memory footprint.
  • A stress condition (a horizon, a shifted environment, a candidate policy set) that makes the returned number meaningful rather than survivable only in a lucky regime.

Return describes the tested outcome; resource counters describe its cost; stress conditions describe the scope of the test. Together they support interpretable comparisons. A result on one shifted gain or horizon establishes performance there, not robustness to every admissible shift.

A return without resource counters cannot establish sample or compute efficiency. A resource count can be audited independently, but a return comparison without a specified initial-state distribution is under-specified. Paired cases can improve comparison precision; appropriately replicated unpaired comparisons can also be valid.

A model-based controller can underperform because its model misrepresents relevant dynamics or rewards, or because its search, horizon, objective, or policy family is inadequate even with exact dynamics. Task infeasibility is another possibility. These causes interact. We use the same random-shooting planner with true dynamics as a reference, not as a global optimum or upper bound.

For a local model-substitution diagnostic, hold initial cases, rewards, search settings, and candidate random streams fixed while replacing learned dynamics with true dynamics. If the true-dynamics controller also performs poorly relative to a feasible baseline, investigate search, horizon, objective, and policy-family limitations. If substitution improves return, it provides evidence that the model choice matters under that protocol, not a certificate identifying a unique cause. Feedback trajectories can diverge even with matched streams. More search might help or hurt either controller, so test that remedy rather than ruling it out.

Task Return and Sample Efficiency

Task return is the sum of rewards earned by a policy in an episode of finite length TT, G(π,s0)=∑t=0T−1γt r(st,at)G(\pi, s_0) = \sum_{t=0}^{T-1} \gamma^{t}\, r(s_t, a_t) with s0∼μs_0 \sim \mu, where:

  • G(π,s0)G(\pi, s_0): the total discounted reward accumulated over the episode starting from initial state s0s_0
  • s0∼μs_0 \sim \mu: the initial state is drawn from the stated initial-state distribution μ\mu
  • TT: the finite episode horizon (number of steps)
  • γ∈(0,1]\gamma \in (0, 1]: the discount factor (here we set γ=1\gamma = 1, giving an undiscounted finite-horizon objective)
  • r(st,at)r(s_t, a_t): the scalar reward received at step tt in state sts_t under action ata_t
  • at=π(st)a_t = \pi(s_t) for a deterministic policy: randomized MPC additionally depends on its planner stream ω\omega
  • st+1=f(st,at)s_{t+1} = f(s_t, a_t): the next state produced by the true environment transition dynamics

The discount factor and episode horizon belong to the task definition. Here γ=1\gamma=1 and T=30T=30, so the code sums 30 unweighted rewards with no terminal reward. Changing discount or horizon can change policy rankings, but different reward timing does not force a ranking reversal. Comparisons should use the same objective, or explicitly present the different objectives as different tasks.

For a randomized controller, write the expected evaluation return as J(π)=Es0∼μ, ω[G(π,s0;ω)]J(\pi)=\mathbb{E}_{s_0\sim\mu,\,\omega}[G(\pi,s_0;\omega)], where ω\omega contains the planner's random draws. A stochastic environment would add transition randomness to that expectation. The deterministic PD controller needs no ω\omega; random-shooting MPC does. Our finite estimates use 20 held-out initial states and one prescribed stream per case, seeded 110+i110+i for case ii. Reconstructing those streams pairs comparisons. These estimates do not average over independent training runs or repeated planner streams at each state, and geometric state coverage alone does not establish estimation precision.

An expected-return claim must state its initial-state distribution μ\mu and any other randomness being averaged. Other legitimate policy claims concern success probability, tail risk, safety, or latency and require their own defined quantities. Different initial-state distributions answer different return questions unless the comparison explicitly accounts for them.

Initial states influence the trajectories and model queries a policy generates. Starting near the origin does not guarantee that a controller stays there, nor does a farther starting state guarantee worse return. Measure visitation and return if those mechanisms are the claim. Here a fixed uniform sample on [−1,1]2[-1,1]^2 defines the held-out cases for all source controllers and the paired target experiment.

Sample efficiency describes performance as a function of real-environment interaction budget, not a raw ratio J/NJ/N. That ratio is especially unhelpful for negative rewards, where its ordering can change simply by changing the reward offset. A learning curve evaluates systems at several disclosed interaction budgets, using either cumulative online learning or separate fits. Independent seeds and uncertainty estimates are needed to characterize variability; their number should be justified by the precision of the comparison. Track these interaction buckets separately:

  • Source pretraining interactions: data used before the target adaptation experiment. Source and target can share an environment; the domains must be stated.
  • Target adaptation interactions. Real trajectories collected in the deployment environment and used to update any parameter (model, policy, value, planner hyperparameter).
  • Validation rollouts. Interactions used to select among hyperparameters or checkpoints. This is not free: those rollouts also consumed actions.
  • Test rollouts. Interactions used to report the metric, with no update or selection step consuming their outcomes.

Separating the counters prevents two accounting errors: treating synthetic model transitions as newly collected environment data, and omitting interactions used for selection. Report each bucket so readers can compare the resources their deployment constrains.

Synthetic transitions generated by f^\hat f are compute, not real samples. A planner that rolls the model K⋅HK \cdot H times per decision did not collect K⋅HK \cdot H real environment steps. It spent K⋅HK \cdot H model evaluations. Both quantities matter, but they belong to different budgets. Real interactions may constrain a deployment when collecting data is expensive or dangerous; model evaluations may constrain it when inference is slow or costly. A method can be sample-efficient and compute-inefficient, or the reverse. Report both budgets against the deployment's actual limits.

What a Learning Curve Is Not

A sweep of planning horizon HH or candidate count KK with one fixed model is a search-configuration sensitivity test, not a sample-efficiency curve. A learning curve varies the real-interaction budget and reports the learning procedure used at each point. Use multiple independent runs when estimating variability, with the run count justified by uncertainty or power rather than a universal three-seed threshold. Chapter 65 develops that experimental-design question.

A horizon sweep can mix longer lookahead, fewer candidates, and longer model rollouts. A candidate-count sweep changes search adequacy at a fixed model. Neither alone establishes a data-efficiency advantage. To measure that advantage, vary the real-interaction budget and evaluate the resulting learner or checkpoint under a consistent test protocol.

The remaining experiments use a deterministic damped double integrator with two state variables and a known nonlinear transition. This keeps return and resource accounting inspectable. The same evaluation questions apply when studying ensemble methods, imagination-based policies, or value-equivalent models, but those methods need their own implementations and protocols; this notebook does not benchmark them.

The testbed is a deterministic point mass with position and velocity state, damped dynamics, and nonlinear quadratic drag. Its reference transition is a known formula, so we can substitute it into the same planner without claiming to solve control optimally. A neural surrogate would not remove that reference formula; an unknown physical plant would make exact-dynamics substitution unavailable.

The system is a 1D damped point mass with quadratic drag. The state is s=(p,v)∈R2s=(p,v)\in\mathbb{R}^2, the action is a scalar bounded acceleration a∈[−1,1]a\in[-1,1], and its discrete source transition is vt+1=(1−d)vt+Δt g at−cdragvt∣vt∣v_{t+1}=(1-d)v_t+\Delta t\,g\,a_t-c_{\mathrm{drag}}v_t|v_t|, followed by pt+1=pt+Δt vt+1p_{t+1}=p_t+\Delta t\,v_{t+1}. These are the toy's discrete update rules, not an exact continuous-time solution.

where:

  • vtv_t: velocity at time tt
  • ptp_t: position at time tt
  • ata_t: scalar bounded acceleration applied at time tt
  • Δt\Delta t: integration timestep, 0.10.1
  • dd: linear damping coefficient, 0.050.05
  • cdragc_{\mathrm{drag}}: quadratic drag coefficient in the discrete velocity update, 0.10.1; it is distinct from the return discount γ\gamma
  • gg: actuator gain, set to a source value of 1.01.0 and (for transfer) a target value of 0.70.7
  • ∣vt∣|v_t|: absolute velocity, capturing the sign-symmetric drag effect

The first line updates velocity from linear damping, the scaled control input, and quadratic drag. The second line updates position by semi-implicit (post-update velocity) Euler integration. The reward is a quadratic tracking-and-control penalty to the origin, r(s,a)=−(p2+0.05 v2+0.01 a2)r(s, a) = -\bigl( p^2 + 0.05\, v^2 + 0.01\, a^2 \bigr), where:

  • p2p^2: position penalty, driving the state toward the origin
  • 0.05 v20.05\, v^2: velocity penalty, discouraging high speed
  • 0.01 a20.01\, a^2: control effort penalty, discouraging large actions
  • the overall negative sign: penalizes deviations rather than rewarding them, so higher return means staying closer to the origin with less effort

The quadratic term cdragv∣v∣c_{\mathrm{drag}}v|v| makes the true transition nonlinear, while the learned surrogate is affine in state and action. Least squares solves the empirical fitting problem within that class. This is a deliberate class mismatch, not an injected corruption of the fitted coefficients.

Empirical least-squares optimality does not rule out overfitting or finite-sample estimator variation. There is no transition or observation noise here, but the sampled training trajectories still affect the fitted coefficients. The affine class also cannot represent the quadratic drag everywhere. Held-out errors and control tests measure the consequences of that fit; they do not decompose all error into a single bias term.

In[3]:
Code
import numpy as np

DT = 0.1
DAMP = 0.05
DRAG = 0.1
ACTION_MIN, ACTION_MAX = -1.0, 1.0
GAIN_SOURCE = 1.0
GAIN_TARGET = 0.7
EPISODE_HORIZON = 30


def true_step_batch(S, A, gain=GAIN_SOURCE):
    S = np.atleast_2d(S).astype(float)
    A = np.atleast_1d(A).astype(float)
    A = np.clip(A, ACTION_MIN, ACTION_MAX)
    p, v = S[:, 0], S[:, 1]
    v_next = (1.0 - DAMP) * v + DT * gain * A - DRAG * v * np.abs(v)
    p_next = p + DT * v_next
    return np.stack([p_next, v_next], axis=1)


def reward_batch(S, A):
    S = np.atleast_2d(S).astype(float)
    A = np.atleast_1d(A).astype(float)
    A = np.clip(A, ACTION_MIN, ACTION_MAX)
    p, v = S[:, 0], S[:, 1]
    return -(p * p + 0.05 * v * v + 0.01 * A * A)

We collect source training transitions with a PD behavior policy plus action noise, initialized in a narrow region. This creates a restricted collection regime; whether its errors cause exploitation must be tested rather than assumed.

In[4]:
Code
def behavior_action(S, rng):
    p, v = S[:, 0], S[:, 1]
    a = -3.0 * p - 2.0 * v + 0.3 * rng.standard_normal(len(S))
    return np.clip(a, ACTION_MIN, ACTION_MAX)


def collect_transitions(
    n_episodes, horizon, s0_low, s0_high, rng, gain=GAIN_SOURCE
):
    Ss, As, Sns = [], [], []
    for _ in range(n_episodes):
        s = rng.uniform(s0_low, s0_high, size=(1, 2))
        for _ in range(horizon):
            a = behavior_action(s, rng)
            s_next = true_step_batch(s, a, gain=gain)
            Ss.append(s[0])
            As.append(a[0])
            Sns.append(s_next[0])
            s = s_next
    return np.array(Ss), np.array(As), np.array(Sns)


train_rng = np.random.default_rng(0)
S_train, A_train, S_next_train = collect_transitions(
    n_episodes=100,
    horizon=30,
    s0_low=-0.5,
    s0_high=0.5,
    rng=train_rng,
)
PRETRAIN_INTERACTIONS = int(len(S_train))

PRETRAIN_INTERACTIONS records the 3000 source training transitions. The later learned-model comparisons share that source fit, and the final accounting separates these records from adaptation, evaluation transitions, and model-search compute.

Fitting a linear model is a two-line least-squares problem. The design matrix concatenates the two state components, the scalar action, and a bias column so the model can absorb any constant offset.

In[5]:
Code
X_train = np.column_stack([S_train, A_train, np.ones(len(S_train))])
W_lin, *_rest = np.linalg.lstsq(X_train, S_next_train, rcond=None)


def learned_step_batch(S, A, W=W_lin):
    S = np.atleast_2d(S).astype(float)
    A = np.atleast_1d(A).astype(float)
    X = np.column_stack([S, A, np.ones(len(S))])
    return X @ W

Before trusting this model for planning, we look at its one-step error on held-out transitions drawn from a wider region than training. This is the teacher-forced protocol from Chapter 61: it measures error on these sampled trajectories, not the globally weakest region or every state a planner might visit.

Out[6]:
Console
Training transitions:  3000
Held-out transitions:  600
Mean one-step error at state norm < 0.5:   0.0015
Mean one-step error at state norm >= 0.5:  0.0118
Overall one-step MSE:                      0.00005

The held-out errors tend to be larger away from the central region. This is consistent with region-dependent model accuracy, but this plot does not measure planner visitation or establish that the planner seeks high-error states. The later decision tests ask whether these prediction errors matter for the policies evaluated.

The near and far averages summarize two groups of held-out transitions. A scatterplot also shows their within-group variation.

The figure below plots prediction-error magnitude against the Euclidean norm of the numerical state coordinates. Position and velocity have different physical units, so this norm is a declared toy-coordinate diagnostic, not a unit-invariant distance. The split at 0.50.5 does not define the boundary of training support, and the figure does not measure planner visitation.

Out[7]:
Visualization
One-step error versus state distance from origin, near and far regions.
Held-out one-step prediction error against state distance from the origin. The near and far groups use a stated distance threshold; larger errors are more common farther from the origin in this dataset. Distance mixes position and velocity in the toy's chosen units and is not a measure of training density. This figure does not measure planner visitation or show that a planner prefers high-error states.

The farther group has a larger mean one-step error in this dataset. Similar group means would not prove uniformly accurate predictions, and different means do not establish optimism or exploitation. Error magnitude does not reveal the sign of a return error.

A planner could favor poorly supported trajectories if their predicted rewards are spuriously high, but this plot does not show that mechanism. The decision tests below measure actual return and return-prediction discrepancies; this particular run does not demonstrate optimistic exploitation.

A fixed PD feedback controller is our no-model comparator. It supplies a concrete return benchmark without consuming model-fitting data.

PD does not consult the learned model, so it cannot select actions because that model predicts them favorably. That does not identify why a learned-model controller might lose to PD: exploitation, poor search, horizon choice, or other effects can all matter. Beating this PD controller is the chosen return benchmark for the toy, not a universal condition for a world model to be useful.

In[8]:
Code
def make_pd_policy(kp, kd):
    def policy(s):
        p, v = s
        a = -kp * p - kd * v
        return float(np.clip(a, ACTION_MIN, ACTION_MAX))

    return policy


def closed_loop_return(env_step_batch, policy, s0, horizon=EPISODE_HORIZON):
    s = np.asarray(s0, dtype=float).reshape(1, 2)
    total = 0.0
    for _ in range(horizon):
        a = float(np.clip(policy(s[0]), ACTION_MIN, ACTION_MAX))
        assert ACTION_MIN <= a <= ACTION_MAX
        r = float(reward_batch(s, np.array([a]))[0])
        total += r
        s = env_step_batch(s, np.array([a]))
        assert s.shape == (1, 2) and np.all(np.isfinite(s))
    return total


def evaluate_policy(
    env_step_batch,
    policy_factory,
    S0_array,
    horizon=EPISODE_HORIZON,
    planner_seed=110,
):
    """Fresh policy per case; identical case seeds pair candidate streams."""
    returns = np.array(
        [
            closed_loop_return(
                env_step_batch, policy_factory(planner_seed + i), s0, horizon
            )
            for i, s0 in enumerate(S0_array)
        ]
    )
    assert np.all(np.isfinite(returns))
    return returns

We fix a held-out evaluation set of initial states and compute the PD baseline in both the true environment and the learned model. The pair of numbers is the first instance of a pattern we will repeat throughout.

In[9]:
Code
eval_rng = np.random.default_rng(7)
S0_eval = eval_rng.uniform(-1.0, 1.0, size=(20, 2))

pd_policy = lambda _seed: make_pd_policy(kp=4.0, kd=3.0)
pd_returns_true = evaluate_policy(true_step_batch, pd_policy, S0_eval)
pd_returns_model = evaluate_policy(learned_step_batch, pd_policy, S0_eval)
Out[10]:
Console
PD baseline (kp=4.0, kd=3.0), 20 held-out initial states
  actual return in true env: mean=  -3.588  std= 4.697
  model-predicted return:    mean=  -3.704  std= 4.753
  real interactions used to report this number: 600
  (reporting: 0 model rollouts counted as real interactions)

Even for a simple PD controller whose behavior is not chosen by the model, the model's prediction of PD's return differs from reality. In this run the model predicts a lower return, so its error is pessimistic rather than optimistic. Replacing PD with a model-driven planner can change the gap, but neither its sign nor its cause follows from the PD comparison alone.

This is fixed-budget accounting, not a sample-efficiency learning curve. The model fit uses 3000 source transitions. The fixed PD gains require no fitting transitions, but evaluating PD costs 20 episodes of 30 steps, or 600 environment transitions. Compare both performance and disclosed costs; do not call these 600 transitions 600 rollouts or divide the negative return by a sample count.

Planning Horizon and Compute Tradeoffs

A receding-horizon planner solves a local finite-horizon optimization at each decision step. Given the current state sts_t, it evaluates candidate action sequences of length HH, picks the one that maximizes predicted cumulative reward, executes only the first action, advances one step in the environment, and repeats. The policy emitted by this loop is not a static map; its behavior depends on where inside the model the search lands at each step.

Receding-horizon control executes the first action of a plan and refreshes the current state before planning again. This can limit the effect of trajectory prediction errors, but the chosen first action still depends on the candidate reward scores. Here each score uses HH pre-transition reward terms, evaluated at the current state and H−1H-1 recursively predicted states; the final computed successor has no terminal reward and is not scored. Replanning is not a guarantee that those forecasts or actions are sound.

Random shooting draws KK candidate sequences, each containing HH actions, scores their model rollouts, and executes the first action of the highest-scoring candidate. This implementation performs K⋅HK\cdot H candidate transition evaluations per decision, excluding reward evaluation and ancillary overhead. It is a useful matched-search-count proxy when comparing these configurations, not a complete hardware-independent compute or latency metric.

Random shooting is a simple baseline. Cross-entropy methods iteratively refit a sampling distribution around elite candidates; gradient-based shooting uses differentiable dynamics. Here the random-shooting policy is determined by its state input, model, settings, and evolving random stream. Fresh policy factories reconstruct that stream per evaluation case so comparisons can share draws.

Proxies vs Latency

For a claim such as "our planner is faster," report measured latency and its hardware and implementation conditions. For a claim about this sweep under a matched candidate-transition budget, report K⋅HK\cdot H. The proxy and latency answer different questions.

Latency depends on implementation, hardware, batch size, memory layout, and model cost. Timing claims can generalize only with supporting measurements or a justified performance model. Matching K⋅HK\cdot H answers a specific depth-versus-breadth question; latency-matched comparisons and resource-performance frontiers answer other legitimate questions.

Under a fixed K⋅HK\cdot H budget, increasing HH exposes later rewards while reducing candidate count. Longer sequences explore a larger action space and use more recursively predicted states. Prediction errors can propagate, but their magnitude need not increase monotonically. Only the chosen first action is executed before replanning; all HH candidate reward stages contribute to the score.

These effects need not balance at the same horizon across tasks or search algorithms. Longer lookahead can help delayed rewards, while finite random shooting can struggle to cover longer action sequences. A horizon sweep measures the resulting behavior, not a unique decomposition of the causes.

Talvitie (2017) analyzes composed prediction errors and self-correcting models, rather than prescribing universally short MPC horizons. MBPO (Janner et al., 2019) uses short model rollouts branched from real data for policy learning; that is distinct from this MPC horizon sweep. Wang et al. (2019) explicitly studies a planning-horizon dilemma in model-based RL. These papers motivate measuring horizon sensitivity, not a theorem that shorter always wins.

The planner below owns its seeded random generator. Each evaluated episode creates a fresh policy with seed 110+i110+i. Learned and true-dynamics planners with the same (H,K)(H,K) therefore use the same candidate arrays at corresponding decision indices. Different (H,K)(H,K) settings reshape draws into different candidate sequences, so the horizon sweep is not an identical-action-sequence intervention.

In[11]:
Code
def random_shooting_action(s, step_batch_fn, H, K, rng):
    a_candidates = rng.uniform(ACTION_MIN, ACTION_MAX, size=(K, H))
    S = np.tile(np.asarray(s, dtype=float), (K, 1))
    total = np.zeros(K)
    for h in range(H):
        a_h = a_candidates[:, h]
        total += reward_batch(S, a_h)
        S = step_batch_fn(S, a_h)
    best = int(np.argmax(total))
    return float(a_candidates[best, 0])


def make_mpc_policy(step_batch_fn, H, K, seed):
    rng = np.random.default_rng(seed)

    def policy(s):
        return random_shooting_action(s, step_batch_fn, H, K, rng)

    return policy


def make_mpc_factory(step_batch_fn, H, K):
    return lambda seed: make_mpc_policy(step_batch_fn, H, K, seed)


## Fresh factories reproduce draws; learned/true planners receive matched arrays.
check_s = np.array([0.2, -0.1])
check_factory = make_mpc_factory(learned_step_batch, H=3, K=8)
check_first, check_second = check_factory(909), check_factory(909)
assert np.array_equal(
    [check_first(check_s) for _ in range(4)],
    [check_second(check_s) for _ in range(4)],
)
check_rng_a, check_rng_b = (
    np.random.default_rng(110),
    np.random.default_rng(110),
)
assert np.array_equal(
    check_rng_a.uniform(ACTION_MIN, ACTION_MAX, (8, 3)),
    check_rng_b.uniform(ACTION_MIN, ACTION_MAX, (8, 3)),
)

We sweep HH while keeping K⋅HK\cdot H exactly 640 candidate transition evaluations per decision. This holds one search-count proxy fixed while changing depth and breadth; it does not hold wall-clock time or search-space coverage fixed.

In[12]:
Code
budget_sweep = [
    {"H": 2, "K": 320},
    {"H": 5, "K": 128},
    {"H": 10, "K": 64},
    {"H": 20, "K": 32},
    {"H": 40, "K": 16},
]

sweep_results = []
for cfg in budget_sweep:
    H, K = cfg["H"], cfg["K"]
    mpc_learned = make_mpc_factory(learned_step_batch, H, K)
    mpc_oracle = make_mpc_factory(true_step_batch, H, K)
    r_learned = evaluate_policy(true_step_batch, mpc_learned, S0_eval)
    r_oracle = evaluate_policy(true_step_batch, mpc_oracle, S0_eval)
    sweep_results.append(
        {
            "H": H,
            "K": K,
            "compute": K * H,
            "learned_mean": float(r_learned.mean()),
            "learned_std": float(r_learned.std()),
            "oracle_mean": float(r_oracle.mean()),
            "oracle_std": float(r_oracle.std()),
        }
    )
Out[13]:
Console
Matched-budget sweep on 20 held-out initial states; 30-step episodes
PD baseline (true-env) mean return: -3.588

  H     K   K*H   MPC-learned   MPC-oracle
  2   320   640        -4.635       -4.634
  5   128   640        -3.832       -3.834
 10    64   640        -4.612       -4.731
 20    32   640        -5.723       -5.985
 40    16   640        -7.693       -8.699

Both planners attain their highest observed mean at H=5H=5. Learned-model means for horizons 2,5,10,20,402,5,10,20,40 are approximately −4.635,−3.832,−4.612,−5.723,−7.693-4.635,-3.832,-4.612,-5.723,-7.693; true-dynamics means are −4.634,−3.834,−4.731,−5.985,−8.699-4.634,-3.834,-4.731,-5.985,-8.699. The fixed PD mean is about −3.588-3.588, higher than every MPC mean in this sweep. These are descriptive estimates on the 20 fixed cases and prescribed streams, not replicated training-run averages.

Both curves improve from H=2H=2 to H=5H=5 and then decline. The true-dynamics controller also declines, so transition-model error is not necessary for that pattern. Paired candidate streams remove an avoidable draw mismatch within each configuration, but the sweep still changes both horizon and candidate count and allows feedback trajectories to diverge. It does not identify one remedy or establish which controller dominates beyond these cases.

Out[14]:
Visualization
Line chart of true-environment return versus planning horizon for learned and oracle MPC plus a PD baseline.
True-environment return of random-shooting MPC with $K \\cdot H=640$ candidate transition evaluations per decision. Both learned-model and true-dynamics planners improve from horizon 2 to 5, then decline. Transition-model error is not necessary for the decline, and search breadth changes with depth. Initial cases and candidate streams are paired within each configuration. The dashed line is the fixed PD baseline; this one-run comparison does not isolate a unique cause or remedy.

Do not generalize this sweep into a rule that shorter horizons are better. Delayed-reward tasks may require deeper lookahead or a terminal value estimate. Nor is the true-dynamics heuristic an upper bound: at several horizons the learned-model heuristic earns a higher observed return. Search, horizon, and model-quality sweeps under declared budgets can test possible remedies; this figure alone does not select one.

Equal candidate-transition counts need not produce equal latency. This NumPy implementation batches over KK and loops over HH, while a neural implementation has its own batching and sequence costs. Measure wall-clock latency on the relevant hardware if latency is the claim; do not infer that the two configurations must differ on every possible machine.

Model Exploitation and Policy Ranking

Model exploitation occurs when action selection favors trajectories whose predicted value is spuriously high because of model errors. A planner can prefer such trajectories without knowing that its predictions are wrong. This is a possible failure mechanism, not an inevitable property of every optimizer. Our fitted model has region-dependent prediction errors, but the measured comparisons below show pessimistic return estimates and no PD ranking reversal. They therefore do not demonstrate optimistic exploitation in this toy. They show how to measure return-prediction gaps and policy-selection fidelity without assuming that a failure must appear.

Selection can create optimism even when individual estimates are unbiased. For fixed candidate values and integrable, conditionally unbiased score errors, the expected selected score overstates the selected candidate's true value by a nonnegative amount; it can be zero. In an equal-value example with independent identically distributed scale-family noise, expected maximum error grows with candidate count and noise scale. Correlated errors, unequal true values, and candidate composition can change that behavior. These conditions distinguish the optimizer's curse from an unconditional claim that more search worsens control.

Several mechanisms can contribute to a return-prediction discrepancy:

  • Dynamics error: f^(s,a)≠f(s,a)\hat f(s, a) \ne f(s, a) in the region the planner visits.
  • Reward-model error: r^(s,a)≠r(s,a)\hat{r}(s, a) \ne r(s, a) even if the dynamics are perfect.
  • Objective misspecification: the reward the planner optimizes is not the reward the environment's evaluation uses.

These are possible mechanisms, not an exact additive decomposition. Horizon truncation, changed feedback trajectories, planner randomness, and finite evaluation samples also affect what is compared. Different remedies address different mechanisms: improve relevant dynamics data or model class, check reward predictions, or align objectives. The toy uses the same specified reward everywhere and pairs planner streams, but this does not uniquely decompose the remaining discrepancy.

At a fixed position, −(p2+v2)-(p^2+v^2) discourages speed while −(p2−v2)-(p^2-v^2) rewards it. Accurate dynamics alone do not align those opposite objective preferences. In our experiments, r^=r\hat r=r; the MPC lookahead horizon still differs from the 30-step evaluation horizon.

Possible mitigations have different assumptions. PETS (Chua et al., 2018) propagates uncertainty from probabilistic ensembles through trajectory sampling; it is not an explicit ensemble-disagreement reward penalty. MOPO (Yu et al., 2020) uses uncertainty-penalized rewards in an offline model-based learning setting. Ensemble averaging can reduce unshared estimation variation, but shared errors can remain. State-conditioned action or state-action support constraints can restrict unsupported queries; marginal action bounds alone do not guarantee supported states. Shorter rollouts, terminal value estimates, and feedback replanning are other options, each requiring evaluation in its own setting.

Bootstrap resampling is one way to diversify ensemble fits, as in PETS. A terminal value estimate approximates reward beyond the rollout horizon instead of explicitly simulating every later step. Relative to executing one fixed open-loop plan, replanning performs a new search at each observation and therefore adds planning calls.

A penalty can sacrifice useful actions if uncertainty is miscalibrated or its scale is poorly chosen. An ensemble can agree on an incorrect prediction. A support constraint can exclude a useful unseen action. A shortened horizon can miss delayed reward, and a terminal value model can inherit training blind spots. Fresh observations let replanning react to discrepancies but do not guarantee correction of model error. These are mitigations to test, not safety certificates.

For the return-prediction comparison, run the same controller implementation in learned and true dynamics from identical initial cases. Each case reconstructs its planner stream with seed 110+i110+i before both runs. Model-closed-loop return uses predicted states as controller inputs and as reward arguments; true-closed-loop return uses reference states. This is not the planner's local HH-step candidate score. Even the controller that plans with true dynamics is evaluated inside the learned outer simulator for its model-predicted point below.

In[15]:
Code
comparison_set = [
    ("PD (kp=4,kd=3)", pd_policy),
    (
        "MPC learned H=10,K=64",
        make_mpc_factory(learned_step_batch, 10, 64),
    ),
    (
        "MPC oracle  H=10,K=64",
        make_mpc_factory(true_step_batch, 10, 64),
    ),
]

pred_vs_actual = []
for name, policy in comparison_set:
    actual = evaluate_policy(true_step_batch, policy, S0_eval)
    predicted = evaluate_policy(learned_step_batch, policy, S0_eval)
    pred_vs_actual.append(
        {
            "name": name,
            "actual_mean": float(actual.mean()),
            "actual_std": float(actual.std()),
            "predicted_mean": float(predicted.mean()),
            "gap": float(predicted.mean() - actual.mean()),
        }
    )
Out[16]:
Console
policy                      model-predicted     actual      gap
PD (kp=4,kd=3)                       -3.704     -3.588   -0.116
MPC learned H=10,K=64                -4.932     -4.612   -0.319
MPC oracle  H=10,K=64                -5.080     -4.731   -0.349

All three return estimates are pessimistic in the paired run. PD predicts about −3.704-3.704 and earns −3.588-3.588, a gap of −0.116-0.116. Learned-model MPC predicts −4.932-4.932 and earns −4.612-4.612, a gap of −0.319-0.319. True-dynamics MPC predicts −5.080-5.080 in the learned outer simulator and earns −4.731-4.731 in the reference environment, the largest negative gap here, about −0.349-0.349. These are model-closed-loop discrepancies, not evidence of optimistic exploitation or a unique causal decomposition.

The scatterplot compares true-environment and model-closed-loop mean return. Above the equality diagonal would mean optimism; all three points are below it. Streams are reconstructed before each paired run, while feedback trajectories and therefore chosen actions may still diverge between the two outer simulators.

Out[17]:
Visualization
Scatter plot of model-predicted return against true return for PD and MPC controllers, compared to a y=x line.
Model-closed-loop versus true-environment mean return for three controller implementations on the same 20 cases with reconstructed paired planner streams. Every point is below equality, indicating pessimistic estimates. True-dynamics MPC has the largest negative gap, about 0.349 return units. Feedback trajectories can still diverge between outer simulators; these are descriptive discrepancies, not a unique decomposition of model error.

Every point is below the equality diagonal. True-dynamics MPC has the largest negative discrepancy in this paired evaluation. Reporting only these model-predicted scores would understate the observed returns, so the useful lesson is to measure the discrepancy rather than assume optimism.

The diagonal is exact agreement between the two mean-return estimates. Departure measures return-prediction discrepancy; three policy points are not a full distributional calibration study.

The fixed candidate set tests scoring and ordering without adaptively generating candidates from model scores. Selecting the highest predicted score can still create selection optimism; fixing the set does not remove every selection effect. Good ordering on this finite set supports good selection within it, not on unseen policies. Poor ordering can compromise selection, but it does not prove that adding candidates or changing search cannot improve actual return.

In[18]:
Code
candidate_kps = [1.0, 2.0, 3.0, 4.0, 6.0, 8.0]
candidate_kds = [0.5, 1.5, 2.0, 3.0, 4.0, 5.0]
candidates = [
    (f"PD kp={kp},kd={kd}", lambda _seed, kp=kp, kd=kd: make_pd_policy(kp, kd))
    for kp, kd in zip(candidate_kps, candidate_kds)
]

ranking_data = []
for name, policy in candidates:
    actual = evaluate_policy(true_step_batch, policy, S0_eval).mean()
    predicted = evaluate_policy(learned_step_batch, policy, S0_eval).mean()
    ranking_data.append(
        {"name": name, "actual": float(actual), "predicted": float(predicted)}
    )


def pairwise_agreement(pred, true):
    n = len(pred)
    concord = 0
    discord = 0
    for i in range(n):
        for j in range(i + 1, n):
            dp = np.sign(pred[i] - pred[j])
            dt = np.sign(true[i] - true[j])
            if dt == 0:
                continue
            concord += int(dp == dt)
            discord += int(dp != dt)
    return concord, discord

Pairwise agreement is the fraction of pairs (i,j)(i,j) with yi≠yjy_i\ne y_j for which sign(pi−pj)=sign(yi−yj)\mathrm{sign}(p_i-p_j)=\mathrm{sign}(y_i-y_j), where pip_i is the predicted mean and yiy_i the measured reference mean. A predicted tie on non-tied truth counts as disagreement. If no true-score pair is eligible, the fraction is undefined and the code reports NaN. Selection regret is max⁡iyi−yi^\max_i y_i-y_{\hat i}, with i^=arg⁡max⁡ipi\hat i=\arg\max_i p_i and first-index tie breaking. It is nonnegative for this finite set, not a gap to the globally optimal policy. Equality of finite-sample means is a numerical tie convention, not proof of equal population returns.

Out[19]:
Console
Candidate set: 6 PD controllers (six zipped kp/kd pairs)
  PD kp=1.0,kd=0.5   model=  -4.701  actual=  -4.475
  PD kp=2.0,kd=1.5   model=  -3.847  actual=  -3.741
  PD kp=3.0,kd=2.0   model=  -3.676  actual=  -3.569
  PD kp=4.0,kd=3.0   model=  -3.704  actual=  -3.588
  PD kp=6.0,kd=4.0   model=  -3.635  actual=  -3.516
  PD kp=8.0,kd=5.0   model=  -3.613  actual=  -3.491

Pairwise ranking agreement: 1.000  (concordant 15, discordant 0)
Best true policy:           PD kp=8.0,kd=5.0
Model-selected policy:      PD kp=8.0,kd=5.0
Selection regret (finite-set): 0.000

For these six zipped PD gain pairs, all 15 eligible pairs are concordant, agreement is 1.000, and both score lists select kp=8,kd=5k_p=8,k_d=5. Finite-set regret is 0.000. This run demonstrates the diagnostic without producing a ranking failure. It does not establish order preservation outside this candidate set or the fixed evaluation cases.

The fixed candidate set can now be plotted directly, with the true best and model-selected points marked.

The scatter plot below places each PD candidate at its true return on the horizontal axis and its model-predicted return on the vertical axis. A star marks the best true candidate and an X the model-selected candidate. The two marks coincide in this run because the model selects the true best candidate. Their horizontal separation would measure finite-set selection regret if the selected candidates differed; here that regret is zero.

In[20]:
Code
candidate_pred = np.array([r["predicted"] for r in ranking_data])
candidate_true = np.array([r["actual"] for r in ranking_data])
# Ties in argmax are broken by the first candidate index.
best_true = int(np.argmax(candidate_true))
best_pred = int(np.argmax(candidate_pred))
Out[21]:
Visualization
Predicted versus true returns for fixed PD candidates, best and selected marked.
True versus model-predicted mean return for six fixed PD controllers. The star and X coincide because the model selects the same candidate as the true environment: agreement is 1 across all 15 non-tied pairs and selection regret is zero. Departures from the diagonal are return-prediction errors, not evidence of ranking reversals. This run illustrates the diagnostic without demonstrating a ranking failure.

Zero-Shot, Few-Shot, and Cross-Domain Transfer

Transfer evaluation has a protocol problem before it has a metric problem. The words "zero-shot," "few-shot," and "cross-domain" are used loosely, so let's fix operational definitions that survive peer review.

Transfer needs a protocol before a metric because the same word, "zero-shot," can mean at least three different things. It can mean the model is frozen and no target data is seen at all. It can mean the model is frozen but a small amount of target data was used to select hyperparameters. It can mean the model is frozen but the planner's hyperparameters were tuned on target rollouts. These settings use different information and can have different interaction budgets and expected performance. Calling all of them zero-shot conceals those protocol differences. The operational definitions below make the information access and budgets explicit.

  • Zero-shot means no target-domain training, adaptation, or selection information is used before evaluation. Model, policy, planner, and learned hyperparameters are frozen. Target rollouts used for checkpoint or hyperparameter selection count as validation interactions and violate this strict zero-shot definition, even without gradient updates.
  • Few-shot permits a stated, bounded number of target-environment transitions (call it NfewN_{\text{few}}) and a stated number of adaptation steps (gradient updates, closed-form refits, or meta-updates). Both numbers are reported. Define the target adaptation budget as NfewN_{\text{few}} (number of target-environment transitions allowed) and KadaptK_{\text{adapt}} (number of adaptation steps). The target evaluation rollouts are separate from the NfewN_{\text{few}} transitions used for adaptation; if an evaluation rollout is used to select among adaptation settings, it is not an evaluation rollout anymore.
  • Cross-domain transfer names the specific shift: shifted dynamics (different ff), shifted observations (different sensor model), shifted reward (different task objective), or shifted task (different target). A claim like "our method transfers well" is not operational without the shift label.

To test a source-pretraining benefit at a given target-data budget, compare against a target-only baseline with the same allowed target interactions. Matching that budget removes a data-quantity confound, not every source of variation. More target data do not guarantee better performance.

The source and target environments differ only in actuator gain, from g=1.0g=1.0 to g=0.7g=0.7; observations, reward, and task stay fixed. We also pair initial cases and planner streams for the frozen-controller comparison. The physical source actuator coefficient exceeds the target coefficient by 1/0.7−1≈43%1/0.7-1\approx43\% relative to the target. This concerns the magnitude of a nonzero action's incremental actuator effect, not every total next-state prediction. The fitted affine coefficient need not equal the physical coefficient exactly.

In[22]:
Code
def target_step_batch(S, A):
    return true_step_batch(S, A, gain=GAIN_TARGET)


S0_target = (
    S0_eval.copy()
)  # pair source/target cases; only actuator gain changes

mpc_frozen = make_mpc_factory(learned_step_batch, H=10, K=64)
pd_target = lambda _seed: make_pd_policy(kp=4.0, kd=3.0)

zero_shot_returns = evaluate_policy(target_step_batch, mpc_frozen, S0_target)
pd_target_returns = evaluate_policy(target_step_batch, pd_target, S0_target)
TARGET_EVAL_TRANSITIONS = len(S0_target) * EPISODE_HORIZON
Out[23]:
Console
Zero-shot (frozen source model, no target training):
  MPC sample mean: -6.050  std: 8.540
  PD baseline:     -4.587  std: 6.162
  target episodes per controller: 20
  target transitions per controller: 600

The frozen MPC mean is about −6.050-6.050 on the target cases, compared with the target PD mean of −4.587-4.587. Its source-environment mean was −4.612-4.612 under paired cases and streams. The gain change therefore changes this controller's measured return, but the gap to PD need not be entirely caused by the model shift; search limitations were already visible in the source sweep. Adaptation tests whether a refit improves the controller, not whether it must recover that entire gap.

We now spend Nfew=400N_{\text{few}} = 400 target-environment transitions to fit a new model. To make the transfer claim operational, we compare three configurations against the PD baseline, all evaluated on the same target evaluation cases:

  • Few-shot fine-tune: refit on the union of the source pretraining set and the new target transitions, keeping the same linear model class.
  • Target-only baseline: refit from scratch on the NfewN_{\text{few}} target transitions alone. Same budget on target transitions, no source advantages.
  • PD baseline: the chosen no-model, no-fitting-data return comparator for the target task.

The two refits use identical handcrafted features and the same 400 target transitions. The source-plus-target fit adds 3000 source records with equal per-record weighting, so old-domain data dominate its loss numerically. Paired candidate streams control search-draw differences at corresponding decisions, but one collection seed and one stream per case do not establish a population pretraining benefit. PD uses no model-fitting data, though its target evaluation still consumes interactions.

Counting the source-plus-target fit's target records as free would invalidate its stated adaptation budget. All 400 target records must be reported regardless of whether the fit also uses source data.

In[24]:
Code
FEW_SHOT_INTERACTIONS = 400

adapt_rng = np.random.default_rng(55)
S_adapt, A_adapt, S_next_adapt = collect_transitions(
    n_episodes=20,
    horizon=FEW_SHOT_INTERACTIONS // 20,
    s0_low=-0.5,
    s0_high=0.5,
    rng=adapt_rng,
    gain=GAIN_TARGET,
)
assert len(S_adapt) == FEW_SHOT_INTERACTIONS

X_adapt = np.column_stack([S_adapt, A_adapt, np.ones(len(S_adapt))])
W_target_only, *_ = np.linalg.lstsq(X_adapt, S_next_adapt, rcond=None)


def target_only_step_batch(S, A):
    S = np.atleast_2d(S).astype(float)
    A = np.atleast_1d(A).astype(float)
    X = np.column_stack([S, A, np.ones(len(S))])
    return X @ W_target_only


X_fine = np.vstack([X_train, X_adapt])
Y_fine = np.vstack([S_next_train, S_next_adapt])
W_fine, *_ = np.linalg.lstsq(X_fine, Y_fine, rcond=None)


def fine_step_batch(S, A):
    S = np.atleast_2d(S).astype(float)
    A = np.atleast_1d(A).astype(float)
    X = np.column_stack([S, A, np.ones(len(S))])
    return X @ W_fine
In[25]:
Code
mpc_fine = make_mpc_factory(fine_step_batch, H=10, K=64)
mpc_target_only = make_mpc_factory(target_only_step_batch, H=10, K=64)

fine_returns = evaluate_policy(target_step_batch, mpc_fine, S0_target)
target_only_returns = evaluate_policy(
    target_step_batch, mpc_target_only, S0_target
)

transfer_results = [
    {
        "setting": "Zero-shot (frozen source model)",
        "mean": float(zero_shot_returns.mean()),
        "std": float(zero_shot_returns.std()),
    },
    {
        "setting": f"Few-shot fine-tune (source + {FEW_SHOT_INTERACTIONS} target)",
        "mean": float(fine_returns.mean()),
        "std": float(fine_returns.std()),
    },
    {
        "setting": f"Target-only baseline ({FEW_SHOT_INTERACTIONS} target)",
        "mean": float(target_only_returns.mean()),
        "std": float(target_only_returns.std()),
    },
    {
        "setting": "PD baseline (no model)",
        "mean": float(pd_target_returns.mean()),
        "std": float(pd_target_returns.std()),
    },
]
Out[26]:
Console
Target-environment return on 20 held-out target initial states
setting                                                    mean      std
Zero-shot (frozen source model)                          -6.050    8.540
Few-shot fine-tune (source + 400 target)                 -6.050    8.540
Target-only baseline (400 target)                        -6.050    8.539
PD baseline (no model)                                   -4.587    6.162

Frozen, source-plus-target, and target-only MPC all round to −6.050-6.050; their mean differences are below 0.0010.001 in this paired run. There is no appreciable adaptation gain here and no demonstrated source-pretraining benefit. This null result does not diagnose non-transferable features: the features are identical. Mixed-domain weighting, model mismatch, and search sensitivity would need separate tests.

The bars report mean target return. Error bars are the empirical standard deviation of the 20 evaluated episode returns, using denominator 20 (NumPy's default), not standard errors or confidence intervals. MPC episodes vary both initial state and prescribed candidate stream, so this spread does not isolate either source of variation. The three model-based means are nearly identical; PD has the higher mean.

Being explicit about what this experiment cannot establish is essential. The shift here is one scalar change in a low-dimensional, deterministic, fully observed system. It cannot speak to representation mismatch, partial observability, visual distribution shift, or high-dimensional action spaces. It cannot speak to whether some other pretrained architecture would have transferred better. It is an illustration of the protocol, not a claim about transfer in general.

Out[27]:
Visualization
Bar chart comparing zero-shot, few-shot, target-only, and PD baselines on the target environment.
Mean target return after changing actuator gain from 1.0 to 0.7. Frozen, source-plus-target, and target-only MPC use paired cases and candidate streams and have nearly identical means in this run. Both refits share 400 target transitions. Error bars show empirical episode-return standard deviations, not confidence intervals; the run does not demonstrate appreciable adaptation gain or source-pretraining benefit.

Target adaptation uses 400 transitions shared by the two refits. Each target controller is evaluated for 20 episodes of 30 steps, or 600 transitions; four controllers consume 2400 target evaluation transitions in total. Those outcomes are not used to fit parameters or choose settings in this notebook. If outcomes are used for selection, label their interactions validation and obtain a fresh held-out test for an unselected performance estimate. The same reused score should not be presented as untouched test evidence.

The accounting below counts reference-simulator transition calls made by the experiment, including repeated evaluations of the same cases. It is not a count of unique observations or of physical deployments. Source-side evaluations are diagnostic comparisons at predeclared settings, not held-out evidence for a final controller chosen after inspecting their results. The target settings are fixed before their target outcomes are inspected. No separate validation collection is run. Model-rollout transitions, including true-dynamics calls used inside oracle search, belong to planning compute rather than this environment-interaction tally.

In[28]:
Code
SOURCE_EVAL_PER_CONTROLLER = len(S0_eval) * EPISODE_HORIZON
interaction_accounting = {
    "Source model training": PRETRAIN_INTERACTIONS,
    "Source one-step diagnostics": len(S_test),
    "Source PD evaluation": SOURCE_EVAL_PER_CONTROLLER,
    "Source horizon sweep (5 settings, 2 controllers)": len(budget_sweep)
    * 2
    * SOURCE_EVAL_PER_CONTROLLER,
    "Source predicted/actual comparison (3 actual runs)": len(comparison_set)
    * SOURCE_EVAL_PER_CONTROLLER,
    "Source fixed-policy ranking (6 actual runs)": len(candidates)
    * SOURCE_EVAL_PER_CONTROLLER,
    "Target adaptation records (shared by two refits)": len(S_adapt),
    "Target evaluation (4 controllers)": len(transfer_results)
    * TARGET_EVAL_TRANSITIONS,
}
assert interaction_accounting["Target evaluation (4 controllers)"] == 2400
assert sum(interaction_accounting.values()) == 18400
for bucket, count in interaction_accounting.items():
    print(f"{bucket}: {count} simulator transitions")
print(
    f"Total reference-environment interactions: {sum(interaction_accounting.values())}"
)
print("Separate validation transitions: 0")
print(
    "MPC search: 640 candidate transition evaluations per decision in the sweep"
)
print("Wall-clock latency and memory footprint: not measured")
Out[28]:
Console
Source model training: 3000 simulator transitions
Source one-step diagnostics: 600 simulator transitions
Source PD evaluation: 600 simulator transitions
Source horizon sweep (5 settings, 2 controllers): 6000 simulator transitions
Source predicted/actual comparison (3 actual runs): 1800 simulator transitions
Source fixed-policy ranking (6 actual runs): 3600 simulator transitions
Target adaptation records (shared by two refits): 400 simulator transitions
Target evaluation (4 controllers): 2400 simulator transitions
Total reference-environment interactions: 18400
Separate validation transitions: 0
MPC search: 640 candidate transition evaluations per decision in the sweep
Wall-clock latency and memory footprint: not measured

Limitations and Impact

Everything in this chapter is measured inside a low-dimensional, deterministic, fully observed toy. The system has two state variables, one scalar action, no measurement noise, no partial observability, no discrete events, and no perception stack. That is enough to make the evaluation framework concrete and to demonstrate the planning-horizon dilemma, the predicted-versus-actual gap, the ranking and selection protocols, and the transfer protocol on a system you can inspect line-by-line. It is not enough to say anything about how these effects play out in a 64×64 grayscale Atari environment, a video-based robot manipulation task, or a large-scale interactive game engine. The magnitudes of exploitation, ranking inversion, and transfer gap all depend on the difficulty of the model fit, the dimension of the state and action spaces, and the structure of the reward.

The specific learned model in this chapter is deliberately the simplest possible: linear in (s,a)(s, a) with a bias. Replacing it with a neural model, an ensemble, or a diffusion head can change the numerical results, but an equivalent predictive map or unchanged selected actions can preserve them. The evaluation question remains: what does the composed model-planner-policy-environment system earn, on which held-out cases, at which interaction budget, and against which baselines? The architecture of f^\hat f enters as one component in a system, not as an end in itself. Publications that report only predictive accuracy and image fidelity are reporting on a rung of the ladder, not on the system.

Finite-set regret does not bound global suboptimality, and high unweighted pairwise agreement can hide a costly reversal between particular policies. The true-dynamics controller is a same-family heuristic reference, not a local upper bound. Reported standard deviations describe the 20 realized episode returns conditional on the disclosed seeds. They omit variation across training datasets and repeated planner streams at each initial state. Agarwal et al. (2021) explains why independent-run variation and appropriate interval estimates matter in RL evaluation.

The computational accounting in this chapter is a proxy, not a wall-clock measurement. Vectorized NumPy over KK candidates is not the same cost as the same number of candidate evaluations in a PyTorch model with a recurrent core or a transformer over a token sequence. Wherever throughput matters, wall-clock latency and memory footprint must be measured on the target hardware. The K⋅HK \cdot H proxy is useful for relative comparisons of planning algorithms at matched search size, but not for latency claims.

Transfer here changes one coefficient of deterministic nonlinear dynamics; the learned surrogates are affine. It does not test a linear-Gaussian environment family, visual shift, partial observability, or structural mechanism changes. A positive few-shot score alone would not establish either adaptation benefit or source-pretraining benefit; frozen and matched-budget target-only baselines address different comparisons. Our observed adaptation comparison is essentially a null result.

The MPC configurations share random shooting. Learned-model variants use affine surrogates, while the true-dynamics reference uses the exact nonlinear transition. PD is a separate model-free feedback comparator. We do not compare random shooting with CEM, gradient-based shooting, or learned policy priors. Substituting those components could change the results and would require its own resource accounting and evaluation.

The final limitation is the one that motivates the next chapter. Every return, regret, agreement, and transfer number in this chapter is conditional on the specific initial-state distributions, the specific evaluation horizons, the specific candidate sets, the specific source pretraining data, the specific target adaptation data, and the specific random seeds we happened to draw. Flip any of these and the ranking can change. Comparisons between systems are only meaningful when these conditions are matched or explicitly accounted for. Constructing datasets, benchmark suites, splits, and statistical summaries that make such comparisons reproducible across labs is the subject of Chapter 65, Datasets, Benchmarks, and Experimental Design. That chapter is not a restatement of what we did here; it is the discipline that makes what we did here comparable to what everyone else will do.

Summary

Decision evaluation concerns a composed model-planner-policy-environment system. Report its executed return under stated initial states, objectives, randomness, resource budgets, and baselines. Predictive quality alone does not certify that outcome, and one successful task does not certify faithful dynamics or general transfer.

Task return is the finite reward sum along an environment rollout. Expected return averages over initial states and, where relevant, planner and environment randomness. Sample efficiency describes achieved performance as a function of real-interaction budget, not return divided by sample count. This notebook is a fixed-budget experiment rather than a learning curve. Source training, adaptation, validation, and test transitions are distinct from synthetic model-search calls.

The horizon sweep matches K⋅H=640K\cdot H=640 candidate transition evaluations per decision. Both learned and true-dynamics MPC peak at H=5H=5 in this run, then decline; PD has the higher mean throughout. This is evidence about the specified search configuration, not a universal short-horizon rule or a unique model-error diagnosis. Candidate-transition counts are a proxy; latency and memory require separate measurement.

Model exploitation is a possible selection-sensitive failure in which model errors make chosen actions look spuriously valuable. Under conditionally unbiased candidate estimates, selection can create nonnegative expected optimism, with growth claims requiring additional assumptions about candidate values and errors. Our paired controller estimates are instead pessimistic, and the six-policy ranking test has agreement 1 and finite-set regret 0. These diagnostics measure discrepancies without inventing an exploitation or ranking failure. Uncertainty propagation, penalties, support constraints, shorter rollouts, terminal values, and replanning have tradeoffs, not safety guarantees.

Zero-shot and few-shot protocols must identify frozen components, target interactions, adaptation operations, and selection data. The target experiment pairs cases and candidate streams and discloses 400 adaptation transitions plus 2400 target evaluation transitions across four controllers. Frozen and both adapted MPC means are nearly identical on this run; no appreciable adaptation gain or source-pretraining benefit is demonstrated.

The experiments make the protocol inspectable: fixed cases and seeds, explicit return calculations, paired controller streams, and separate data and compute counters. They do not measure wall-clock latency, memory, confidence intervals, or population-wide transfer. Chapter 65 develops datasets, splits, benchmarks, and statistical summaries that make such comparisons reproducible beyond a single toy run.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about planning, control, and policy evaluation.

Planning, Control, and Policy Evaluation Quiz

Question 1 of 70 of 7 completed
According to the chapter, what is the primary target when evaluating a model-based planning system?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026planningcontrol, author = {Michael Brenndoerfer}, title = {Planning, Control, and Policy Evaluation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Planning, Control, and Policy Evaluation. Retrieved from https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models
MLAAcademic
Michael Brenndoerfer. "Planning, Control, and Policy Evaluation." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models>.
CHICAGOAcademic
Michael Brenndoerfer. "Planning, Control, and Policy Evaluation." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Planning, Control, and Policy Evaluation'. Available at: https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Planning, Control, and Policy Evaluation. https://mbrenndoerfer.com/writing/planning-control-policy-evaluation-world-models

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.