Part of World Models Handbook
A reproducible sim-to-real workflow covering structural mismatch, domain randomization, delay ablations, residual monitoring, and gated continual correction.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Sim-to-Real Deployment and Continual Correction
A better fitted dynamics model need not give better tracking with a fixed planner. The synthetic experiment in this chapter demonstrates that distinction while separating parameter mismatch, delayed observations and actuator limits. We develop a reproducible deployment workflow: collect informative adaptation data, monitor prediction errors, and evaluate proposed corrections before promotion. These executable results are not hardware evidence or a safety certificate.
Simulation-to-reality transfer can expose parametric mismatch, structural mismatch, observation differences, and task or interface errors. Parametric mismatch means the chosen model family can represent the relevant plant but its numerical values are wrong, such as damping after a payload change. Identifying those values requires observations that distinguish the parameters; sufficient excitation can provide that information. Good conditioning reduces sensitivity to measurement noise and numerical error, rather than being a logical requirement for exact noiseless recovery. Structural mismatch means the relevant true map is outside the fixed hypothesis class. Retuning parameters cannot make a member of that class equal an excluded map over the whole domain. This does not imply a positive approximation-error floor or a permanent observed residual: a nonclosed class can approach the map arbitrarily closely, and finite observations can miss the discrepant region.
The remedies address different parts of this problem. Domain randomization supplies training variation within a specified family; it does not automatically add an omitted mechanism. Residual models fit discrepancies relative to a nominal predictor, with generalization depending on their structure and adaptation data. Monitoring checks for departures from a reference error pattern; detection depends on distinguishable observations and a suitable decision rule, and does not by itself identify the cause. Continual correction adds a controlled update process, including rejection, versioning and rollback.
The rest of Part XII has already given you the tools to talk about what goes wrong. In Failure Modes and Model Exploitation we saw how optimization against an imperfect model can exploit and amplify accessible errors. In Robustness, Calibration, and Safe Control we developed uncertainty calibration and runtime assurance. We will not re-derive those here. Instead we will take a small, fully executable system, deliberately break it in controlled ways, and watch which remedies help, which do not, and why. Controlled comparisons help attribute effects to specified changes. Our delay and actuator-authority conditions add one factor at a time to an already parameter-shifted deployment baseline; they are not single-discrepancy comparisons against the nominal plant. Hardware can combine multiple discrepancies, motivating matched ablations rather than attribution from an aggregate score.
Training a policy, controller, or world model with simulated data or rollouts and evaluating or deploying it on a physical system. This chapter also covers monitoring and correction after transfer as operating conditions change.
The Reality Gap, Stated Precisely
This section defines the reality gap precisely. We sort the sources of mismatch by how they behave under different remedies, and we distinguish two kinds of model error that organize the rest of the chapter.
Where the gap comes from
Jakobi, Husbands, and Harvey's 1995 work used the term in its title, Noise and the Reality Gap: The Use of Simulation in Evolutionary Robotics. This is an attribution of that published use, not a claim of priority for inventing the term. For deployment, the practical question is task-specific: which distinctions must this simulator preserve for the controller and evaluation conditions? A simulator can support one task yet fail on another using the same physical system.
That reframing matters because it tells you where to look. A simulator can be extremely accurate about contact dynamics and completely silent about motor heating. It can model the arm perfectly and forget that the camera exports frames on a different clock than the controller. The gap is not a scalar error bar; it is a structured mismatch, and different parts of it have different fixes. A scalar error bar suggests a single knob to turn, whereas a structured mismatch tells you to go find the specific component of the system whose behavior is missing from the model and decide whether to model it, randomize over it, or monitor for it.
It helps to sort the sources by the part of the system they affect. Residual patterns can suggest tests, but do not uniquely identify these categories:
- Parameter mismatch. Wrong values in a chosen model family: masses, inertias, friction coefficients, damping, gains, gear ratios, battery-voltage curves or motor constants. Estimation can correct these when the experiment identifies them. They need not produce signed bias: a damping error gives a velocity-dependent residual whose mean can be zero over symmetric motion.
- Unmodeled structure. Missing or incorrectly represented mechanisms, such as deadband, backlash, stiction, hysteresis, cable stretch or contact stick-slip, relative to the chosen class. Correction can require a different functional form, state or input. Its residual need not depend on state: an omitted constant force is a structural discrepancy in a class that lacks a constant term.
- Timing and latency. Sensor delay, actuator delay, jitter, unsynchronized clocks, overruns or dropped frames. Sensor delay can leave the physical transition valid while making an observation-only prediction insufficient. Actuator delay changes which command reaches the plant and may require action history or additional state.
- Saturation and rate limits. Delivered commands differ from requested commands when limits bind. Limits may depend on speed or temperature for a particular actuator and operating range; that dependence is not universal. Gentle tests may never activate the limit. Once it does activate, the resulting state error can persist after the actuator leaves saturation.
- Observation channel differences. Noise statistics, quantization, offsets, occlusion or modality dropout. These affect what the controller knows even when the modeled plant dynamics are accurate.
- Distribution and task mismatch. Deployment reference trajectories, initial conditions or loads can differ from those used in model fitting and evaluation. Good fit on the simulated population does not establish performance under that new population. This can coexist with omitted physical mechanisms; the categories are a guide to designing checks, not mutually exclusive diagnoses.
One note on terminology before the formal distinction. Throughout this chapter the plant is the physical system being controlled, the model is the mathematical object used to predict its behavior, and the controller is the rule that maps observations into commands. These three words appear constantly, and keeping them separate is what makes the rest of the argument readable. A great deal of confusion in practice comes from sliding between them, for instance by saying "the model is wrong" when the model is right and the observation feeding the controller is stale, or by saying "the controller is unstable" when the controller is fine and the model it plans against is systematically biased.
Parametric error and structural error
Let us be precise, because this distinction organizes the whole chapter.
Suppose the true plant is a damped oscillator whose state consists of position and velocity . The continuous-time dynamics are:
where:
- : position in metres, the first component of the state
- : velocity in metres per second, the second component of the state
- : the time derivative of position, which equals by definition
- : acceleration, or equivalently the time derivative of velocity
- : mass-normalized stiffness, in ; for , the term produces restoring acceleration
- : mass-normalized viscous damping, in ; for , its contribution dissipates mechanical energy
- : dimensionless control gain, since the applied command already has acceleration units
- : the scalar command issued by the controller at time , in metres per second squared (the units of acceleration), which plays the role of the book's action . We write the raw issued command as and the post-clip applied command as . The distinction between and matters because the model may predict a command the actuator cannot deliver. The physical acceleration is always written as to avoid collision with the control symbol.
The oscillator has linear second-order dynamics in the applied input. Clipping from issued to applied command makes the overall commanded system nonlinear. A damped mechanical mode is a useful teaching example for inertia, restoring behavior and dissipation; it is not a claim that every controlled plant has such a mode. The parametric-versus-structural distinction does not depend on the simplicity of the plant; it depends on whether the discrepancy between the true plant and the model is confined to the constants or spills into the functional form.
A mismatch confined to the numerical values of quantities that appear in the model, while the model's functional form is correct. If the true parameters are and your model uses , the parametric error is the vector .
A mismatch in the model's functional form: the true map is not in the hypothesis class your model or simulator draws from. Additive parameter drift cannot represent it, and neither can adding more parameters of the wrong kind.
A plant with when the model uses has parametric mismatch if the remaining modeled structure is correct. Sensor delay is a different layer: leaves the physical transition valid, but changes the controller's information. A predictor receiving only may lack enough history to determine the next state. Residual fitting can improve that conditional prediction, but exact current-state reconstruction needs additional assumptions and information, such as a known deterministic transition and the intervening applied actions. The delay sweep below tests a controller without that reconstruction, not the best possible time-aware estimator.
When a monitor fires, first test candidate explanations. If the discrepancy is identifiable parameter mismatch, an informative episode can support an update within the current class. If the class is misspecified, consider a different functional form, state or input. Neither category determines how difficult the correction is: an unexcited parameter can be unidentifiable, while an omitted constant term can be straightforward to estimate. The monitor alone does not make that distinction.
A rollout bound makes the distinction between local defects and propagated error explicit. Compare the true recursion with under the same issued action sequence. Fix one norm and an operating region containing both trajectories. Suppose throughout that region and is -Lipschitz in state with action held fixed. The triangle inequality gives
where:
- is endpoint error, not a sum of errors over time.
- is the initial-state discrepancy.
- is a uniform one-step defect bound in the same norm.
- bounds the true map's change with state while holding the action fixed.
- counts transitions; at the sum is empty.
For the example below, both rollouts start from the same state, so . Define the resulting upper bound as . It is a conditional guarantee for each trajectory satisfying the premises, not a prediction that a trajectory attains the bound. If either rollout leaves the operating region, this argument no longer supplies the stated guarantee.
The unamplified final term comes from the newest defect. A defect introduced one transition earlier can be multiplied once by ; an older one can be multiplied twice, and so on. These factors permit contraction when as well as expansion when . The inequalities need not be jointly tight.
For , approaches . At , it is . For fixed and , grows asymptotically as . Actual error may remain much smaller, including zero for identical maps. These open-loop bounds hold actions fixed; independently replanning controllers produce different action sequences and need a separate closed-loop analysis.
Use a dimensionless state norm for the oscillator: . The plot is a generic illustration with in this norm, not an estimated Lipschitz constant or defect bound for the fitted models below.
We plot the same bound for three values of at a fixed defect bound. Only the assumed amplification factor changes; the figure does not classify the fitted plant as contractive or expansive.
import numpy as np
from book_plot_style import use_book_style
use_book_style()
# Upper bound B_H = eps * (1 + L + ... + L^(H-1)), with matched initial states.
# Evaluation settings, declared explicitly:
# fixed dimensionless one-step defect bound eps = 0.01,
# horizons 0..30 recursive steps,
# Lipschitz constants chosen to bracket the marginal case L = 1.
EPS_BOUND = 0.01
HORIZONS = np.arange(0, 31)
LIPSCHITZ = [0.8, 1.0, 1.2]
def rollout_error_bound(eps, L, H):
"""Worst-case bound under the stated smoothness and one-step-error assumptions.
H counts the number of recursive steps, so H = 0 gives an empty sum of 0.
The marginal case L = 1 is handled separately to avoid a 0/0 form; there the
series grows linearly in H rather than geometrically.
"""
h = np.arange(H + 1)
if L == 1.0:
return eps * h
return eps * (1.0 - L**h) / (1.0 - L)
# One curve per Lipschitz constant; index order is fixed so the assignment is stable.
bound_curves = {
L: rollout_error_bound(EPS_BOUND, L, int(HORIZONS[-1])) for L in LIPSCHITZ
}
Domain Randomization: What It Buys You
Domain randomization trains across a declared distribution of environments or plants rather than a single nominal instance. The aim is to cover relevant deployment variation. Successful simulated performance across that family does not by itself guarantee physical transfer: omitted mechanisms, observations, tasks and interfaces can still differ. Whether broad training or deployment-time identification is preferable depends on what can be measured and changed.
Tobin et al. (2017) trained object localization from randomized synthetic images and used it for real robotic grasping. Sadeghi and Levine (2017) trained an indoor quadrotor visual collision-avoidance policy using randomized simulated imagery without real training images. Peng et al. (2018) varied dynamics, including mass, friction and control gains, for a robot-arm object-pushing task. Kumar et al. (2021) combined randomized locomotion training with an adaptation module in RMA. That module estimates task-relevant latent environmental context from proprioceptive state-and-action history, rather than directly identifying a physical parameter vector. They demonstrated physical quadruped deployment, including unseen terrains. These are task-specific empirical demonstrations, not universal guarantees. Successful transfer also does not prove that every property of the physical deployment was inside the declared simulated support.
Three different things called "randomization"
Randomization can be applied to at least three distinct objects. These operations have different properties and should not be treated as interchangeable.
A policy trained across variation. Randomization supplies policy-training environments. A fixed policy can learn useful behavior across the sampled family without an explicit parameter-identification module. Average performance under that distribution is not a worst-case guarantee across every plant, and randomization alone does not make the policy robust by construction. A recurrent policy may also infer context implicitly from history.
A parameter-conditioned predictor. Randomization supplies predictor-training data with context: , where . Deployment must supply that context, through measurement, known configuration or identification. Wrong context can degrade predictions, but its effect depends on the predictor and the error. Context quality is one requirement alongside model fit and coverage, not their replacement.
A parameter-blind mixture predictor. Pooled regression receives state and action but no explicit plant parameters. Under squared loss an unrestricted optimum estimates the conditional mean next state given those inputs; a finite feature class approximates it. That conditional mean need not equal a transition at the mean plant parameters, because states can contain information about the plant. A pooled fit is not itself a robustness guarantee for the controller that uses it.
Below, the blind and conditioned regressions use identical transition pairs. The conditioned model additionally receives parameter context and six feature terms, and deployment context comes from an adaptation episode. This compares those specified fitting routes; it does not isolate context from representation capacity or adaptation information.
The cost of randomization
Wider randomization can require more training data and capacity, but it does not necessarily make a policy slower or less precise. Shared structure can support several plants without a fixed capacity cost per plant. Compare broad training, measured context and adaptation under matched budgets; the right choice depends on uncertainty, coverage and the cost of gathering deployment data.
There is also a coverage question: what happens outside the declared randomization family? Training-family results alone do not establish performance there. Membership is not a guarantee either: finite sampling can leave sparse regions within the support. Extrapolation may succeed or fail depending on structure; the shifted plant below probes one outside-family case, not a universal discontinuity at a support boundary.
Randomization by itself does not provide a runtime detector. A fixed trained system may nevertheless tolerate drift covered by its training and learned behavior. Monitoring checks deployment evidence as conditions evolve, including changes outside anticipated variation. The two mechanisms therefore address different requirements.
A Controlled Deployment Proxy
We use a deterministic, CPU-sized deployment proxy: simulated dynamics stand in for a physical plant. The numerical results are synthetic, not hardware measurements. Relative to a nominal-plant check, the deployment baseline changes parameters; the delay and reduced-authority conditions each add a change to that deployment baseline. These matched comparisons distinguish specified changes without isolating every possible cause of the resulting controller behavior.
The plant, the command, and the notation
We reuse the damped oscillator from the dynamical-foundations material. The deployment plant evolves under semi-implicit Euler integration. The update advances the state by first updating velocity and then using the new velocity to update position:
Velocity is updated first from the current state and applied command, then position uses the new velocity. This semi-implicit ordering differs from explicit Euler. For the undamped oscillator it admits bounded discrete oscillations when , but it is not universally more stable for stiff damping. For this damped linear update, strict stability requires , and . For example, , and violate that condition even though explicit Euler is stable for the same case. An integrator mismatch is a discretization discrepancy; its magnitude need not be small. Here the applied command is with , where:
- : position [m] and velocity [m/s] at timestep , together forming the state
- : the applied (post-clip) command entering the plant
- : the plant parameters for stiffness, damping, and control gain
- s: the discrete timestep
The control input is a scalar command in metres per second squared, and the state is in metres and metres per second. The units are declared so that the saturation limit and the reference amplitude are interpretable, but the system is abstract: you should read it as a stand-in for any second-order plant with a bounded actuator. The appeal of a second-order system as a proxy is that it is small enough that every effect can be traced by hand, nonlinear enough once saturation is introduced to exhibit structural behavior, and physically interpretable enough that the parameters carry meaning.
Let us set up the code.
The first block defines the plant step used for true trajectories. The batched nominal predictor later implements the same nominal update for candidate planning. Those two implementations must stay aligned in clipping, timestep and integration order; otherwise the experiment would introduce an unintended discrepancy.
import numpy as np
# Abstract units, declared so the numbers are readable:
# position p [m], velocity v [m/s], command u [m/s^2], timestep dt [s].
DT = 0.05
U_MAX = 1.0
NOMINAL = {"k": 1.0, "c": 0.4, "b": 1.0}
def step_plant(s, u_issued, params, u_max=U_MAX, rng=None, process_std=0.0):
"""Advance the plant by one timestep with semi-implicit Euler.
s = (p, v). The issued command is saturated before it reaches the plant,
so the actuator limit is part of the plant, not the controller.
"""
p, v = s
u_applied = float(np.clip(u_issued, -u_max, u_max))
acc = -params["k"] * p - params["c"] * v + params["b"] * u_applied
v_next = v + DT * acc
p_next = p + DT * v_next
if process_std > 0.0 and rng is not None:
p_next += process_std * rng.normal()
v_next += process_std * rng.normal()
return np.array([p_next, v_next]), u_appliedNext, a rollout helper that records the true trajectory, and a delay helper that produces a causally delayed copy of any signal. The helper illustrates shape-preserving sensor buffering. The closed-loop experiment below uses the equivalent state-history indexing directly rather than calling this helper. Causal means the delayed signal at time depends only on samples up to time , so it is a faithful model of what a real sensor buffer delivers.
def rollout(params, u_seq, s0=None, u_max=U_MAX, rng=None, process_std=0.0):
"""Roll the true plant forward under a fixed sequence of issued commands."""
n = len(u_seq)
if s0 is None:
s0 = np.zeros(2)
states = np.zeros((n + 1, 2))
applied = np.zeros(n)
states[0] = s0
for t in range(n):
states[t + 1], applied[t] = step_plant(
states[t],
u_seq[t],
params,
u_max=u_max,
rng=rng,
process_std=process_std,
)
return states, applied
def delay_signal(x, delay_steps):
"""Delay a time-major array; repeat the first sample during startup."""
x = np.asarray(x)
if x.ndim == 0:
raise ValueError("x must have a time axis")
if isinstance(delay_steps, (bool, np.bool_)) or not isinstance(
delay_steps, (int, np.integer)
):
raise TypeError("delay_steps must be an integer")
delay = max(int(delay_steps), 0) # Negative values mean no delay.
if len(x) == 0:
return x.copy()
if delay >= len(x):
# Avoid converting an arbitrarily large Python integer to an index dtype.
return np.repeat(x[:1], len(x), axis=0).copy()
indices = np.maximum(np.arange(len(x)) - delay, 0)
return x[indices].copy()
# Shape and startup checks cover scalar signals, vector signals and long delays.
assert np.array_equal(delay_signal(np.arange(5), 2), [0, 0, 0, 1, 2])
assert delay_signal(np.arange(6).reshape(3, 2), 5).shape == (3, 2)
assert np.array_equal(
delay_signal(np.arange(6).reshape(3, 2), 5),
np.repeat([[0, 1]], 3, axis=0),
)
assert delay_signal(np.empty((0, 2)), 4).shape == (0, 2)
assert np.array_equal(delay_signal(np.arange(3), -1), np.arange(3))
vector_signal = np.arange(6).reshape(3, 2)
for long_delay in (2**63, 10**100):
delayed = delay_signal(vector_signal, long_delay)
assert np.array_equal(delayed, np.repeat([[0, 1]], 3, axis=0))
assert not np.shares_memory(delayed, vector_signal)Data splits that stay separate
Using the same trajectory to fit, select and report can bias a claimed generalization estimate. We keep four data roles separate so the reported test trajectory does not enter fitting or candidate selection:
- Training set: transitions from 24 randomized plants, used to fit the domain-randomized predictors.
- Adaptation set: a short episode with bounded commands from the deployment plant, used for system identification and residual fitting. Bounded commands alone do not establish a safe state trajectory.
- Validation set: a separate episode from the deployment plant, used only to gate residual candidates.
- Final test set: an untouched episode from the deployment plant, never used for any decision.
We also keep a plant outside the randomized support, to see what happens past the edge of the promise.
def reference_traj(n, freq=0.15, amp=0.55):
"""Smooth tracking reference for closed-loop evaluation."""
t = np.arange(n) * DT
return amp * np.sin(2.0 * np.pi * freq * t)
def exploration_actions(rng, n, u_max=U_MAX, smooth=0.85, scale=1.6):
"""Smooth, bounded, zero-mean excitation for identification data."""
raw = rng.normal(size=n)
u = np.zeros(n)
prev = 0.0
for t in range(n):
prev = smooth * prev + (1.0 - smooth) * raw[t]
u[t] = np.clip(prev * scale, -u_max, u_max)
return u
def sample_params(rng):
"""Draw one plant from the randomization distribution."""
return {
"k": rng.uniform(0.6, 1.6),
"c": rng.uniform(0.2, 0.8),
"b": rng.uniform(0.7, 1.3),
}We generate randomized training data with smooth bounded commands and Gaussian initial states. Bounded commands do not impose a hard state envelope, and the Gaussian initialization is unbounded. In this fixed seeded experiment, the executed test and closed-loop state/action points fall inside the pooled training convex hull. That check does not establish dense coverage, cover every hypothetical planner rollout, or guarantee coverage for new references and initial states. The conditioned feature class can represent this noiseless transition exactly within the applied-command range; that structure also contributes to its good fit.
rng = np.random.default_rng(20240517)
EP_LEN = 160
train_plants = [sample_params(rng) for _ in range(24)]
train_data = []
for params in train_plants:
u = exploration_actions(rng, EP_LEN)
s0 = rng.normal(scale=0.4, size=2)
states, _ = rollout(params, u, s0=s0, rng=rng)
train_data.append({"params": params, "s": states, "u": u})Randomized training plants : 24 Pooled transition pairs : 3840 k range : [0.69, 1.57] c range : [0.21, 0.74] b range : [0.73, 1.28]
The generating family is , , . The printed ranges are extrema of only twenty-four draws, not those distribution endpoints. A plant with or is outside the declared family. Membership in the generating family and containment in the observed sample box are distinct checks; neither alone establishes predictive performance.
Four Candidate Models
We compare four predictors. Each receives the same randomized training data where applicable, and each is used to drive the same controller, so differences in closed-loop cost come from the predictor and not from the planner. Holding the controller fixed is what makes the comparison a test of the predictor; if the controller changed with the model, we could not tell which change produced the difference.
- Nominal sim model: the hand-written simulator, with , , , and clipping at . It is exactly the model you would ship if you trusted your simulator.
- Domain-randomized, parameter-blind: a single predictor fitted to the pooled transitions from all 24 plants, with no context input.
- Domain-randomized, parameter-conditioned: the same pooled data, but the predictor also receives the plant parameters. At deployment we feed it parameters identified from the adaptation episode.
- Nominal plus residual: the nominal simulator plus an explicitly shaped residual estimator fitted to the adaptation episode only.
The predictor class is deliberately modest: a quadratic feature map over appended with the ones vector, followed by a linear solve. Concretely, the feature map is:
The feature map transforms the raw state and action into a fixed set of basis functions, allowing the linear predictor to capture quadratic nonlinearities. The feature map is a fixed, human-chosen transformation; the learning happens in the linear weights applied to it, so the expressive power of the predictor is exactly the span of these ten basis functions. Each element of the feature vector is defined below.
where:
- : the position component of the state
- : the velocity component of the state
- : the issued scalar command
- the bilinear terms : couplings between state components (), and between state and action ()
- the quadratic terms : curvature in each variable
- : the constant bias term
The base predictor is with : ten learned weights for each state component. The conditioned extension appends , giving sixteen features and a weight matrix. These interactions represent the per-plant semi-implicit transition exactly within the tested command range. The small explicit class removes neural optimization complexity, but does not isolate training-distribution quality from context, feature capacity or adaptation information. A neural predictor could implement the same batched interface and evaluation workflow; its fit and numerical comparisons would need fresh validation.
def build_pairs(data, context=False):
"""Return (S, U, S_next), or (S, U, C, S_next) when context=True."""
S, U, S_next, C = [], [], [], []
for d in data:
S.append(d["s"][:-1])
U.append(d["u"])
S_next.append(d["s"][1:])
if context:
p = d["params"]
C.append(np.tile([p["k"], p["c"], p["b"]], (len(d["u"]), 1)))
S, U, S_next = np.vstack(S), np.concatenate(U), np.vstack(S_next)
if context:
return S, U, np.vstack(C), S_next
return S, U, S_next
def feat_base(S, U):
p, v = S[:, 0], S[:, 1]
u = np.asarray(U, dtype=float).reshape(-1)
return np.column_stack(
[p, v, u, p * v, p * u, v * u, p * p, v * v, u * u, np.ones_like(p)]
)
def feat_cond(S, U, C):
base = feat_base(S, U)
p, v = S[:, 0], S[:, 1]
u = np.asarray(U, dtype=float).reshape(-1)
k, c, b = C[:, 0], C[:, 1], C[:, 2]
return np.column_stack([base, k, c, b, k * p, c * v, b * u])
def fit_linear(X, Y, ridge=1e-8):
A = X.T @ X + ridge * np.eye(X.shape[1])
return np.linalg.solve(A, X.T @ Y)
S_tr, U_tr, S_next_tr = build_pairs(train_data)
W_blind = fit_linear(feat_base(S_tr, U_tr), S_next_tr)
S_trc, U_trc, C_tr, S_next_trc = build_pairs(train_data, context=True)
W_cond = fit_linear(feat_cond(S_trc, U_trc, C_tr), S_next_trc)The predictors themselves expose the same batched interface predict_batch(S, U), so the planner uses the same code for every candidate. In this noiseless linear plant, the identified parameter-conditioned model and the target-fitted residual model recover essentially the same dynamics. Their predictions differ only at numerical fitting precision, and their tracking costs agree to the displayed precision. These are two fitting routes, not two empirically distinct controllers in this experiment; the comparison cannot establish that one route is better.
def nominal_predict_batch(S, U):
"""The simulator: knows the nominal dynamics and its own actuator limit."""
p, v = S[:, 0], S[:, 1]
u = np.clip(np.asarray(U, dtype=float).reshape(-1), -U_MAX, U_MAX)
acc = -NOMINAL["k"] * p - NOMINAL["c"] * v + NOMINAL["b"] * u
v_next = v + DT * acc
return np.column_stack([p + DT * v_next, v_next])
def blind_predict_batch(S, U):
"""Pooled fit: one compromise predictor for the whole randomized family."""
return feat_base(S, U) @ W_blind
def make_cond_batch(ctx):
"""Context-conditioned fit; at deployment the context must be identified."""
ctx = np.asarray(ctx, dtype=float).reshape(1, -1)
def predict_batch(S, U):
C = np.tile(ctx, (len(S), 1))
return feat_cond(S, U, C) @ W_cond
return predict_batch
def make_residual_batch(W_res):
"""Nominal dynamics plus an explicitly shaped additive residual."""
def predict_batch(S, U):
return nominal_predict_batch(S, U) + feat_base(S, U) @ W_res
return predict_batchThe deployment plant and a bounded-command adaptation episode
The deployment plant sits inside the randomized support but is not the nominal plant: , , . We collect 140 steps of smooth excitation from it. That episode stands for a bounded, supervised data-collection run, not for autonomous exploration. The distinction matters: the experiment assumes we were permitted to excite the plant gently for a fixed, short interval under supervision, and it assumes the excitation was chosen to cover the directions we care about rather than to maximize information at any cost.
From that episode we do two things. First, system identification: invert the Euler step to get an acceleration target , then solve the following linear system. The system identification step recovers the plant parameters by fitting a linear model to acceleration data. The relationship is:
where:
- : the acceleration target recovered by finite-differencing the observed velocity
- : the state components at time
- : the applied (post-clip) command at time
- : the unknown plant parameters we want to estimate
The system is solved via least squares over the adaptation episode to recover the estimates . The inversion works because the semi-implicit Euler step is affine in the parameters: given , the parameter vector appears linearly on the right-hand side. A linear least-squares problem in three unknowns is about as cheap as identification gets, which is the point: when the parametric part of the gap is identifiable and you have the excitation to see it, closing it costs almost nothing. Second, residual fitting: compute the residual target using the nominal model, and fit a shaped estimator to it.
The residual target must use the same nominal transition, timestep and input/target alignment as deployment. A different integrator or a one-step pairing error can introduce a learnable discrepancy that is numerical rather than physical. A residual estimator may fit that artifact with small training error, yet introduce error when applied to correctly aligned deployment inputs. Check the convention before interpreting a fitted correction as evidence of changed plant dynamics.
# Deployment plant (inside randomized support) and a plant outside it.
DEPLOY = {"k": 1.25, "c": 0.62, "b": 0.82}
SHIFT = {"k": 1.85, "c": 1.05, "b": 0.70}
# Bounded adaptation episode from the deployment plant.
adapt_u = exploration_actions(rng, 140)
adapt_states, _ = rollout(DEPLOY, adapt_u, s0=np.array([0.15, -0.05]), rng=rng)
S_a, U_a, S_next_a = build_pairs(
[{"params": DEPLOY, "s": adapt_states, "u": adapt_u}]
)
# System identification from the same episode.
acc_target = (S_next_a[:, 1] - S_a[:, 1]) / DT
X_sys = np.column_stack([-S_a[:, 0], -S_a[:, 1], np.clip(U_a, -U_MAX, U_MAX)])
theta_hat = np.linalg.lstsq(X_sys, acc_target, rcond=None)[0]
# Residual target, aligned to the nominal model and the same (s, u) pairing.
res_target = S_next_a - nominal_predict_batch(S_a, U_a)
W_res = fit_linear(feat_base(S_a, U_a), res_target)
# Held-out validation episode from the same plant, used only for gating.
val_u = exploration_actions(rng, 120)
val_states, _ = rollout(DEPLOY, val_u, s0=np.array([-0.2, 0.1]), rng=rng)Identification from the bounded adaptation episode k_hat = 1.250 true k = 1.250 c_hat = 0.620 true c = 0.620 b_hat = 0.820 true b = 0.820 nominal simulator: k = 1.00, c = 0.40, b = 1.00 Adaptation episode length : 140 steps Validation episode length : 120 steps
Even a simple least-squares identification recovers the deployment parameters to a useful precision from 140 steps of excitation. That is the parametric part of the gap, and it is cheap to close when you have excitation covering the directions you care about. The gap between what a cheap least-squares solve recovers here and what a wide-domain-randomized problem demands sets up the comparison in the next section. Unique noiseless recovery requires the complete design matrix with columns [-p, -v, applied u] to have rank three. Position and velocity must not be collinear, and the command column must add independent information. Varying commands are a useful excitation design, but a constant nonzero command can suffice when the transient states supply the other independent directions. Zero command makes the gain unidentifiable in this regression.
We check both the generating family and the finite sample box against the evaluation plants. The figure places plant parameters beside training draws. The parameter-blind predictor does not receive these parameters, so membership in either box does not classify its state/action predictions as interpolation or extrapolation. The conditioned predictor does receive parameter context, with coverage still depending on joint feature inputs.
# Declared generating family, separate from observed sample extrema.
distribution_bounds = {"k": (0.6, 1.6), "c": (0.2, 0.8), "b": (0.7, 1.3)}
sample_bounds = {
"k": (float(ks.min()), float(ks.max())),
"c": (float(cs.min()), float(cs.max())),
"b": (float(bs.min()), float(bs.max())),
}
def inside_parameter_box(plant, bounds):
"""Coordinate-wise interval membership, not a performance or density test."""
return all(
bounds[name][0] <= plant[name] <= bounds[name][1] for name in bounds
)
deploy_inside_family = inside_parameter_box(DEPLOY, distribution_bounds)
shift_inside_family = inside_parameter_box(SHIFT, distribution_bounds)
deploy_inside_sample_box = inside_parameter_box(DEPLOY, sample_bounds)
shift_inside_sample_box = inside_parameter_box(SHIFT, sample_bounds)
assert deploy_inside_family and deploy_inside_sample_box
assert not shift_inside_family and not shift_inside_sample_box
# A valid family member can lie outside the observed twenty-four-draw box.
assert inside_parameter_box(
{"k": 0.65, "c": 0.5, "b": 1.0}, distribution_bounds
)
assert not inside_parameter_box({"k": 0.65, "c": 0.5, "b": 1.0}, sample_bounds)
Sample-based planning with a shared actuator envelope
For closed-loop evaluation we use a short-horizon random-shooting model predictive controller. It samples candidate action sequences within the actuator bounds, rolls the predictor forward, and applies the first action of the lowest-cost sequence. The same planner and candidate-action draws are used for each predictor within a regime. Different predicted states can still produce different selected actions and subsequent trajectories; this is a comparison of each predictor coupled to that fixed planner.
At each control tick, the current reference position is held constant across all eight predicted steps. This is not lookahead tracking of the future reference sequence. The stage cost adds squared position error to an effort penalty, ; with position in metres and command in metres per second squared, the coefficient has units of seconds to the fourth power so both terms have units of square metres. The chosen weight, horizon and 64 sampled sequences are experimental settings, not optimized controller parameters.
def plan_action(
predict_batch, s, p_ref, horizon=8, n_samples=64, u_max=U_MAX, rng=None
):
"""Random-shooting MPC over a fixed actuator envelope."""
cand = rng.uniform(-u_max, u_max, size=(n_samples, horizon))
S = np.tile(np.asarray(s, dtype=float), (n_samples, 1))
cost = np.zeros(n_samples)
for h in range(horizon):
U = cand[:, h]
S = predict_batch(S, U)
cost += (S[:, 0] - p_ref) ** 2 + 0.02 * U**2
return float(cand[int(np.argmin(cost)), 0])
def close_loop(
predict_batch, params, p_ref_seq, delay_steps=0, u_max=U_MAX, seed=7
):
"""Closed-loop tracking. The controller sees a delayed observation of the true state."""
plan_rng = np.random.default_rng(seed)
n = len(p_ref_seq)
states = np.zeros((n + 1, 2))
applied = np.zeros(n)
for t in range(n):
obs = states[max(0, t - delay_steps)].copy()
u = plan_action(
predict_batch, obs, p_ref_seq[t], u_max=u_max, rng=plan_rng
)
states[t + 1], applied[t] = step_plant(
states[t], u, params, u_max=u_max
)
return states, appliedNotice what delay_steps does: the plant evolves from the true states[t], while the controller receives states[max(0, t - delay_steps)]. At zero delay, the controller has exact current-state observations. This is an idealized information condition, not a mathematical lower bound on tracking cost. The finite-horizon sampled planner is not globally optimal, and a change in delay can change its chosen actions in either direction. The zero-delay nominal-model run below is a reference for this specific controller, not an achievable-performance guarantee or a hardware baseline.
Ablations: Damping, Gain, Delay, and Saturation
We compare a nominal-plant check, a parameter-shifted deployment baseline, and two changes to that deployment baseline. The last two retain the shifted parameters: their costs should be compared with the parameter-shift baseline to study the additional factor.
- Nominal plant: plant dynamics match the nominal simulator. Predictors adapted to the different deployment plant need not match this plant.
- Parameter shift: the deployment plant, no delay, full actuator authority. Pure parametric error.
- Sensor delay: the same deployment plant with a three-step observation delay. This changes the controller's information, not the physical transition.
- Tight saturation: the same deployment plant with instead of . Both planner and plant receive this bound, so the experiment measures reduced control authority, not an unmodeled clipping limit.
A note on process noise. These loops use zero process noise, fixed initial conditions and seeded planner sampling. Across-model differences therefore concern the predictors and the actions selected using them under these conditions. We do not run a noisy tracking ablation, so neither the tracking-cost change nor preservation of the model ranking under noise is established. The clean experiment isolates a comparison; it is not a performance estimate for noisy hardware.
MODELS = {
"Nominal sim model": nominal_predict_batch,
"DR (parameter-blind)": blind_predict_batch,
"DR (parameter-conditioned)": make_cond_batch(theta_hat),
"Nominal + residual": make_residual_batch(W_res),
}
REGIMES = {
"Nominal plant": {"params": NOMINAL, "delay": 0, "u_max": 1.0},
"Parameter shift": {"params": DEPLOY, "delay": 0, "u_max": 1.0},
"Sensor delay (3 steps)": {"params": DEPLOY, "delay": 3, "u_max": 1.0},
"Tight saturation": {"params": DEPLOY, "delay": 0, "u_max": 0.5},
}
T_CL = 200
ref_seq = reference_traj(T_CL)
ENV_P, ENV_V = 0.9, 3.0
closed_loop = {}
for rname, cfg in REGIMES.items():
closed_loop[rname] = {}
for mname, pred in MODELS.items():
states, applied = close_loop(
pred,
cfg["params"],
ref_seq,
delay_steps=cfg["delay"],
u_max=cfg["u_max"],
)
err = states[1:, 0] - ref_seq
closed_loop[rname][mname] = {
"mse": float(np.mean(err**2)),
"violation": float(
np.mean(
(np.abs(states[1:, 0]) > ENV_P)
| (np.abs(states[1:, 1]) > ENV_V)
)
),
"states": states,
"applied": applied,
}Mean squared position tracking error (lower is better) regime Nominal sim model DR (parameter-blind) DR (parameter-conditioned) Nominal + residual Nominal plant 0.01459 0.01607 0.01756 0.01756 Parameter shift 0.02134 0.02134 0.02558 0.02558 Sensor delay (3 steps) 0.01564 0.01568 0.01672 0.01672 Tight saturation 0.03752 0.03705 0.03870 0.03870 Fraction of steps outside the position/velocity envelope regime Nominal sim model DR (parameter-blind) DR (parameter-conditioned) Nominal + residual Nominal plant 0.000 0.000 0.000 0.000 Parameter shift 0.000 0.000 0.000 0.000 Sensor delay (3 steps) 0.000 0.000 0.000 0.000 Tight saturation 0.000 0.000 0.000 0.000
Read the tables as separate measurements. Tracking MSE evaluates reference following; the envelope fraction evaluates breaches of the declared synthetic position and velocity limits. Every envelope fraction is zero in these runs, so this table does not distinguish the models and does not demonstrate model exploitation or certify safety. A controller could improve tracking while violating a limit in another experiment, but that event is not observed here.
Model-based planning can prefer actions whose predicted consequences are wrong, including actions with understated risk. Whether an optimizer selects them depends on its objective, constraints and search. This experiment does not demonstrate that mechanism: no declared envelope is breached. Keeping tracking and envelope metrics separate would let a broader test detect a tradeoff rather than bury it in one aggregate score.
The measured rankings differ from the simple expectation that a more accurate predictor must track better. In the parameter-shift regime, the nominal and parameter-blind models have similar tracking MSE, while the conditioned and residual models have higher tracking MSE despite fitting the deployment dynamics accurately. Two columns coincide because those fitted predictors are essentially identical. This is evidence about the combined predictor and fixed sampled MPC, not evidence that residual fitting failed to generalize. Finite search, a short prediction horizon, and the planner's effort-penalized objective are possible contributors, but this experiment does not isolate their causal contributions. Do not infer a universal ranking from this single reference, initial condition, and planner seed.
Open-loop rollout error and why it ranks models differently
Closed-loop tracking error is not the same as prediction error, and the two can rank the models differently. To see this, we roll every predictor recursively from the same initial state under the same fixed action sequence from the held-out test trajectory. This is pure open-loop prediction with no feedback to hide mistakes. Recursion is the key word: each prediction is fed back in as the next state, so a small bias at one step becomes an input to the next, and the compounding described by the error bound becomes directly visible.
def recursive_rollout(predict_batch, s0, u_seq):
"""Recursively feed the model's own prediction back in as the next state."""
S = np.asarray(s0, dtype=float)[None, :].copy()
preds = [S[0].copy()]
for t in range(len(u_seq)):
S = predict_batch(S, np.array([u_seq[t]]))
preds.append(S[0].copy())
return np.array(preds)
# Untouched final test trajectory on the deployment plant.
test_u = exploration_actions(np.random.default_rng(999), 200)
test_states, _ = rollout(DEPLOY, test_u, s0=np.array([0.3, -0.2]))
# Explicit reporting scales: one metre of position error and one metre per
# second of velocity error each contribute one unit to the dimensionless norm.
# These are metric conventions, not safety tolerances or fitted parameters.
ERROR_P_SCALE = 1.0 # metres
ERROR_V_SCALE = 1.0 # metres per second
STATE_ERROR_SCALE = np.array([ERROR_P_SCALE, ERROR_V_SCALE])
def normalized_state_error(predicted, observed):
"""Dimensionless joint position/velocity error under declared fixed scales."""
return np.linalg.norm((predicted - observed) / STATE_ERROR_SCALE, axis=1)
open_loop = {}
for mname, pred in MODELS.items():
preds = recursive_rollout(pred, test_states[0], test_u)
open_loop[mname] = normalized_state_error(preds, test_states)Recursive open-loop normalized state error [dimensionless] on held-out deployment trajectory model 1 step 10 steps 50 steps 200 steps Nominal sim model 0.0032 0.0233 0.1542 0.3498 DR (parameter-blind) 0.0031 0.0277 0.1487 0.3698 DR (parameter-conditioned) 0.0000 0.0000 0.0000 0.0000 Nominal + residual 0.0000 0.0000 0.0000 0.0000
The table reports a dimensionless joint state error, , with and . The scales specify how position and velocity are weighted; they are not safety limits. Adding their unscaled squared errors would mix incompatible units. This convention preserves the numerical values of the earlier raw norm but makes its meaning explicit, and the monitoring and held-out gates below use the same convention. Other scales could change rankings and gate decisions; no scale-sensitivity study is included here.
The fitted conditioned and residual predictors produce errors that round to zero at each displayed horizon, consistent with the exactly identifiable noiseless plant. The nominal and parameter-blind predictors have small one-step error and much larger long-horizon error on this trajectory; neither curve is guaranteed to grow monotonically, and their long-horizon values are close. This is not a measured instance of the parameter-blind model necessarily diverging fastest. Prediction error and tracking MSE ask different questions and can rank the same predictors differently. Part XI: Evaluation and Understanding discusses why one-step fit alone is insufficient for a planning model.
The curve adds intermediate horizons to the table, showing how the error evolves along this fixed trajectory.


The delay trap
We sweep observation delay from zero to eight steps and rerun the same controller for the nominal, parameter-blind, and residual predictors. This changes the information supplied to the controller without refitting any predictor. The nominal curve decreases in this particular sweep; delay does not have a guaranteed monotone relationship with closed-loop tracking MSE. Neither a lower MSE at positive delay nor a worse MSE from an accurate model establishes why the controller behaves that way. A causal explanation would require further controller and task ablations.
DELAYS = [0, 1, 2, 3, 5, 8]
delay_sweep = {}
for mname in [
"Nominal sim model",
"DR (parameter-blind)",
"Nominal + residual",
]:
pred = MODELS[mname]
errs = []
for d in DELAYS:
states, _ = close_loop(pred, DEPLOY, ref_seq, delay_steps=d, u_max=1.0)
errs.append(float(np.mean((states[1:, 0] - ref_seq) ** 2)))
delay_sweep[mname] = np.array(errs)
# Reference only: nominal predictor and exact current-state observations.
zero_delay_states, _ = close_loop(
nominal_predict_batch, DEPLOY, ref_seq, delay_steps=0, u_max=1.0
)
zero_delay_reference_mse = float(
np.mean((zero_delay_states[1:, 0] - ref_seq) ** 2)
)
The residual model has higher tracking MSE than the nominal model at zero delay, even though its fitted dynamics reproduce the deployment plant to numerical precision. Its delay curve decreases rather than rising in parallel with the other curves. There is no evidence here that its dynamics fit failed to generalize. The controller is also feeding an aged state into a one-step predictor trained on synchronized current states. An accurate transition from the aged state is not automatically an estimate of the current state; that timing mismatch is distinct from parameter error. The plotted MSE alone does not quantify how much each source contributes.
A possible remedy for known delay is to propagate the delayed observation through the known intervening action history, or to use a state estimator with time-stamped observations and actions. In a noiseless known model this can reconstruct current state; with uncertainty it needs an uncertainty-aware treatment. That estimator is not implemented or evaluated here. More synchronized adaptation data can improve a dynamics fit, but it does not by itself add the missing timing and history inputs to this controller. A plateau in improvement is not a diagnosis on its own: limited excitation, a misspecified model, optimization, and observation timing all need separate checks.
Online Monitoring: Innovation and Residual Signals
Monitoring compares predictions with observations to identify departures worth investigating. In filtering, an innovation is a new measurement minus its predicted measurement, conditioned on prior information. Our fully observed proxy uses a next-state residual instead. These are related prediction-error signals, but the innovation's noise law depends on the filtering model; it is not interchangeable with every learned dynamics residual.
The residual is defined as the difference between the observed next state and the model's prediction:
where:
- : the observed next state delivered by the sensors
- : the model's prediction of the next state from the current observation and the issued action
- : the residual, a vector of the same shape as the state
- the monitor's scalar magnitude: the norm after dividing the position and velocity residuals by the fixed 1 m and 1 m/s reporting scales, respectively
Assess residuals against the model's specified observation and disturbance law. Zero mean and whiteness are properties to justify under that law, not universal requirements. A mean shift, changing variance, spikes or temporal correlation can motivate tests of dynamics, sensors, timing or operating conditions. Such patterns can have multiple causes: zero-mean parameter errors and constant structural errors are possible, as the earlier taxonomy showed.
Sequential change detection accumulates evidence over time. A likelihood-based CUSUM uses log-likelihood increments for specified null and alternative models, resetting its running statistic when accumulated evidence becomes unfavorable. Centered tabular versions can be derived for particular distributions, such as a Gaussian mean change; an arbitrary centered residual sum need not be a likelihood ratio. A calibrated threshold can support a stated false-alarm property under its assumptions. Accumulation can improve detection of small persistent shifts, but does not make every shift distinguishable or remove the false-alarm/detection-delay tradeoff. The NIST CUSUM handbook describes the centered cumulative and tabular forms. The proxy below uses a simpler consecutive-exceedance rule, not CUSUM.
Choosing thresholds on separate reference data
The practical decision is where to set the threshold. There are two honest options, and they are often conflated.
- A calibrated sequential guarantee. For example, set a threshold to meet a target expected stopping time under a specified null law, including temporal dependence. The claim is conditional on that law or on the validity of a stated robust null class; it must be checked for the deployment process.
- An empirical threshold from reference data. A sample quantile defines a cutoff. Estimate its same-condition exceedance rate on separate held-out data and account for sampling uncertainty and serial dependence. The quantile label alone supplies neither an exact population tail probability nor a sequential false-alarm guarantee.
We will use the second, and be explicit that it is not the first. The distinction matters because a quantile threshold carries no guarantee once the operating conditions drift away from the reference period: if the reference data was collected with a warm motor and the deployment runs cold, the "99th percentile" was the 99th percentile of a different distribution. Stating which kind of threshold you used is not bookkeeping; it determines what you are entitled to claim when the monitor stays quiet.
mon_rng = np.random.default_rng(4242)
def residual_series(predict_batch, states, u_seq):
pred = predict_batch(states[:-1], np.asarray(u_seq, dtype=float))
return normalized_state_error(pred, states[1:])
def collect_residuals(plant, n_eps):
"""Dimensionless residual sequence under the nominal monitor."""
out = []
for _ in range(n_eps):
u = exploration_actions(mon_rng, 120)
s0 = mon_rng.normal(scale=0.4, size=2)
states, _ = rollout(plant, u, s0=s0)
out.append(residual_series(nominal_predict_batch, states, u))
return np.concatenate(out)
# Reference period: the plant the monitor was tuned on.
ref_res = collect_residuals(NOMINAL, 20)
# Held-out period: estimate individual-sample exceedances, not alarm events.
hold_res = collect_residuals(NOMINAL, 20)
# A plant outside the randomized support: the monitor should react.
shift_res = collect_residuals(SHIFT, 20)
THRESH = float(np.quantile(ref_res, 0.99))
false_alarm_rate = float(np.mean(hold_res > THRESH))
# Beginning of the first three-exceedance run; declaration occurs two samples later.
exceed = shift_res > THRESH
detection_delay = None
run = 0
for i, e in enumerate(exceed):
run = run + 1 if e else 0
if run >= 3:
detection_delay = i - 2
breakThreshold (normalized error, reference 99th percentile) : 0.00000 Reference residual median : 0.00000 Held-out same-plant sample exceedance rate: 0.0000 Shifted-plant exceedance rate : 1.0000 Start index of first three-exceedance run : 0 First alarm declaration residual index : 2 This is a held-out single-sample exceedance fraction, not an alarm-event rate. It supplies no sequential or distribution-shift guarantee.
In this noiseless setup, the reference plant is exactly the nominal predictor. Its reference and independent same-plant residuals are zero; with the strict > rule, the held-out sample exceedance fraction is zero, not one percent. That fraction is distinct from the alarm-event rate of the three-consecutive rule. Exchangeability alone does not make a sample 99th-percentile threshold yield exactly one percent exceedances, especially with ties. Even a continuous nondegenerate distribution leaves finite-sample quantile uncertainty. A changed exceedance rate can motivate checks, but is not by itself a diagnosis of drift. Residuals are concatenated from separate 120-step episodes; a detector used on independent runs should reset at episode boundaries. The first alarm here lies wholly inside the first episode, so boundary handling does not affect the reported index.
All shifted-plant residuals exceed this degenerate zero threshold in the current runs. Three consecutive exceedances first occur at samples 0, 1, and 2. The reported index is the start of that run, zero; the detector could only declare after observing its third residual, index 2. This is deliberately favorable evidence from an exact noiseless null, not a realistic detection-delay or false-alarm guarantee. On hardware, noise, time alignment, operating changes, and the required alarm-declaration latency need a separate calibration and safety analysis.

What a residual cannot tell you
There is a hard limit on what a residual monitor can conclude, and it is a limit of identifiability, not of engineering. A residual is a scalar (or vector) summary of everything that differs between model and plant. A parametric drift, a sensor bias, an actuator degradation, and a state that has left the region where the model was fitted can all produce the same residual signature. From the residual alone you cannot tell them apart. This is not a defect of the residual; it is a property of any scalar summary that collapses many causes into one number.
Distinguishing causes requires discriminating information, which can come from existing excitation, additional sensors or a deliberately designed test. Active probing is one option, subject to an approved operating envelope, not always a necessity. Compare explicit hypotheses under the same observation and disturbance assumptions. If they induce identical full measurement laws under the chosen experiment, more of that same data cannot distinguish them. Equal predicted means alone are insufficient for this impossibility argument because their variances or other distributional properties may differ.
The shifted plant is outside the declared randomization family, and the monitor fires. That alarm alone does not establish whether parameters, observation processing or an omitted mechanism changed. The generator is known to us, but the detector does not receive that privileged diagnosis. The next section gates residual predictors on another deployment trajectory; it does not implement a fault-hypothesis comparison or diagnose the monitor's shifted plant.
Continual Correction with a Nominal-Plus-Residual Model
Residual correction is the practice of keeping the nominal physics and learning only the correction. The corrected prediction adds a learned residual term to the nominal model:
The nominal model supplies prior structure while the learned term represents its discrepancy. Small residual magnitude does not imply an easy fitting problem: a small discontinuity can still require appropriate features. Conversely, a large constant offset can be simple to learn. Assess residual structure, coverage and uncertainty, not magnitude alone, when deciding whether the decomposition helps.
where:
- : the hand-written nominal dynamics, exactly the simulator you would ship if you trusted it
- : the learned residual map, parameterized by , intended to capture only what gets wrong
- : the corrected next-state prediction used by the planner
We fit to the next-state difference from adaptation. Alignment is necessary but not sufficient: the residual class and data must represent the discrepancy too. Here both the fitted residual and the monitor derive from that vector difference, but the monitor reduces it to a scalar normalized magnitude. The correction is trained on vector targets; it is not trained on those magnitudes.
Related hybrid designs learn corrections around analytic structure. Ajay et al. (2018) augmented analytic simulation with stochastic recurrent residual models for planar pushing and bouncing. Zeng et al.'s TossingBot (2019) added learned corrections to throwing-release velocities proposed by a ballistic controller. The latter corrects control parameters, not a next-state map identical to ours. Their task-specific results motivate a pattern, not a guarantee for this residual estimator.
There are several design decisions that determine whether it works, and each one has a failure mode. Treating them as a checklist rather than an afterthought is what separates a correction that transfers from one that merely fits.
Shaping and regularization. An expressive residual can overfit noise or behave poorly away from its data, but neither outcome is inevitable. Choose features and penalties to reflect justified mechanisms, such as velocity dependence for a friction hypothesis or clipping for a known actuator limit. A weight penalty does not automatically bound function magnitude outside the data: even a linear function with a small nonzero slope grows without bound on an unbounded domain. State any operating-domain bound separately. Our polynomial basis and ridge solve constrain the fitted class without certifying extrapolation.
Alignment. Compute residual targets with the nominal model and state-action timing used at deployment. A different nominal model or timing convention defines a different correction target. Misalignment can be absorbed by a flexible estimator and hidden by an adaptation-set fit, but that outcome is not inevitable: residual patterns or independent timing checks may expose it. Evaluate aligned held-out transitions and controller behavior before accepting a correction; good adaptation error alone does not establish transfer.
Coverage and excitation. Data that never activate saturation cannot identify its limit without additional prior information. A structurally informed model may still generalize to unvisited instances; that inference must be justified and tested. For the three-parameter regression here, unique noiseless recovery requires rank three in the full matrix . A constant nonzero input can suffice during an informative transient; zero input leaves the gain column zero. Designing informative excitation and ensuring it is permitted are separate requirements.
Identifiability limits. The delay sweep did not train a delay estimator and cannot establish a theorem about recoverability. A correct transition model does not supply the controller with the missing action history or timestamps; those inputs would be needed for the compensation procedure described above. Likewise, response outside the applied action range needs either additional data or justified prior structure. Identifiability depends on the observations, excitation, model class, and assumptions, not simply on the amount of data.
Frozen candidates and held-out gates
Now the discipline that makes updates testable: never update the model on the trajectory you are using to validate it, and never promote a candidate without a held-out comparison against a rule that was written down in advance. Writing the rule down in advance prevents changing it to fit the result. This gate measures prediction error on one trajectory; it does not establish closed-loop safety or tracking improvement.
The procedure has a standard shape.
- Collect a bounded adaptation episode under supervision.
- Fit one or more candidate corrections.
- Evaluate each candidate on a held-out validation episode from the same plant.
- Apply an acceptance rule that was fixed before evaluation, including a minimum improvement margin.
- Promote only accepted candidates, and keep the previous model available for rollback.
A candidate can improve one-step error and still make things worse. Fitting noise in adaptation data can improve training one-step error while worsening held-out rollout error; it does not necessarily do so, especially if the same correction also removes a larger real bias. That is why this gate evaluates held-out rollout error rather than accepting a training-set improvement alone. The training-set number measures fit to the data used for estimation; the held-out number measures performance on that separate episode, not a guarantee for untested operating conditions.
Candidate A reuses the fit from the deployment adaptation episode. Candidate B is fitted on synthetic transitions from a different plant. Both are evaluated on the same held-out deployment trajectory with the fixed normalized state metric. Candidate B improves the rollout mean relative to the nominal model in this run, but its one-step error is worse, so the two-condition rule rejects it. This is a concrete rejection mechanism, not proof that every mismatched candidate must fail or that accepted candidates improve closed-loop tracking.
# Candidate A reuses W_res, fitted on the deployment plant's adaptation episode.
# Candidate B is fitted on data from a different plant: a plausible but mismatched source.
shift_u = exploration_actions(rng, 140)
shift_states, _ = rollout(SHIFT, shift_u, s0=np.zeros(2), rng=rng)
S_b, U_b, S_next_b = build_pairs(
[{"params": SHIFT, "s": shift_states, "u": shift_u}]
)
W_res_B = fit_linear(
feat_base(S_b, U_b), S_next_b - nominal_predict_batch(S_b, U_b)
)
def horizon_error(predict_batch, states, u_seq):
preds = recursive_rollout(predict_batch, states[0], u_seq)
return normalized_state_error(preds, states)
# Declare the metric predicate before scoring any candidate.
GATE_ROLLOUT_FACTOR = 0.98
def accept_prediction_candidate(candidate_error, baseline_error):
"""Require more than 2% rollout improvement and no first-step regression."""
return bool(
candidate_error["rollout_mean"]
< GATE_ROLLOUT_FACTOR * baseline_error["rollout_mean"]
and candidate_error["one_step"] <= baseline_error["one_step"]
)
candidates = {
"Nominal (no residual)": nominal_predict_batch,
"Candidate A (target data)": make_residual_batch(W_res),
"Candidate B (mismatched data)": make_residual_batch(W_res_B),
}
gate_error = {}
for cname, pred in candidates.items():
e = horizon_error(pred, val_states, val_u)
gate_error[cname] = {
"one_step": float(e[1]),
"rollout_mean": float(np.mean(e[1:41])),
}
# Apply the previously declared rule to the measured candidate and baseline.
decisions = {
name: accept_prediction_candidate(
errors, gate_error["Nominal (no residual)"]
)
for name, errors in gate_error.items()
}Held-out normalized state error [dimensionless] (40-step recursive rollout) candidate 1-step rollout mean decision Nominal (no residual) 0.00202 0.10825 baseline Candidate A (target data) 0.00000 0.00000 accept Candidate B (mismatched data) 0.00427 0.09853 reject Acceptance rule: rollout mean must beat the nominal by more than 2 percent, and one-step error must not regress. The rule is declared before scoring.

The gate rejects candidate B despite its lower mean rollout error because the one-step criterion also matters. That is the mechanism this experiment exercises. An accepted candidate still needs separate closed-loop, retained-condition, and safety validation. A prediction-error gate cannot substitute for those tests, particularly when better prediction did not improve the tracking MSE above.
Preserve an accepted prior artifact and a tested switchback path so rollback does not require retraining. For evaluation, the code keeps test_states out of fitting and gating. That dataflow separation does not prove historical preregistration or that a developer never inspected test results. Repeated adaptive use can bias an apparent held-out estimate, rather than making the data wholly uninformative. Use a fresh test after development decisions, and retain the criteria and evaluation record.
Hardware-in-the-Loop and Staged Deployment
Before any of this touches a physical plant, it goes through a ladder of increasingly realistic test stages. The vocabulary is worth getting right, because the stages differ in what they can and cannot establish.
- Model-in-the-loop (MIL): the controller and the plant are both software models. You are testing the algorithm.
- Software-in-the-loop (SIL): controller code runs against a simulated plant, ordinarily on a development host. This tests executed code paths, software interfaces, and control logic. Host execution profiles can be measured, but ordinarily do not establish target real-time behavior.
- Processor-in-the-loop (PIL): target-compiled controller code runs on the target processor or an equivalent instruction-set simulator, with plant inputs supplied by a simulation harness. This tests the compiled component's numerical behavior and supported resource measurements. Target-code execution time can be measured; full-loop timing still depends on the scheduler, I/O, harness, and workload, and an instruction-set simulator needs its own timing justification.
- Hardware-in-the-loop (HIL): physical control-loop components exchange signals with simulated components in a closed loop. Controller HIL can connect target hardware to a real-time plant simulator to test timing, I/O and injected faults; engine- or power-HIL can include physical loads and power interfaces. Isermann et al. (1999) discuss engine-control HIL and sensor/actuator faults. Fathy et al. (2006) review automotive HIL and an engine-in-the-loop facility. Those configurations provide different evidence and hazards.
- Physical plant trials: the real system, in a controlled environment, under supervision, with a stated operating envelope and independent safeguards.
Each stage provides evidence within its tested conditions. MathWorks' SIL/PIL overview distinguishes host code from target-compiled execution, while its execution-time profiling documentation describes measurements available in both modes. These profiles must be tied to the measured configuration. HIL can measure deadline behavior and fault responses on target hardware, but finite trials alone do not prove that every possible workload meets a deadline. Simulator validation against physical evidence remains separate. Physical hazards can also exist in HIL setups, especially power-HIL systems; using a simulated plant is not a blanket safety guarantee. Test coverage and a formal guarantee, when available, must be stated separately.
A test configuration in which the controller under test runs on its target hardware, connected through realistic electrical interfaces to a real-time simulation of the plant and its environment, so that real-time execution, I/O timing, and fault response can be exercised before a physical plant is involved.
A timing and watchdog sketch
We cannot run HIL here. What we can do is illustrate the class of logic HIL is designed to exercise: the relationship between control period, sensor period, observation age, and a staleness watchdog. This is a software fault-injection sketch on a CPU. It is not a HIL measurement, it does not characterize real-time performance, it says nothing about hardware fidelity, and it is not evidence of deployment readiness. Its only purpose is to make the relevant quantities concrete enough to reason about.
The control loop ticks every control_dt. The sensor publishes a new sample every sensor_dt, which is slower. The watchdog measures the age of the newest observation and forces the command to zero if that age exceeds a timeout. Then we inject a sensor dropout and watch the watchdog react. The design question is how the timeout should be chosen relative to the sensor period: too small and normal operation trips it, too large and a stall goes undetected for longer than the safety budget allows.
WD_CONTROL_DT = 0.01
WD_SENSOR_DT = 0.05
WD_TIMEOUT = 0.12
WD_DROPOUT_AT = 200
def watchdog_trace(control_dt, sensor_dt, timeout, n_steps, dropout_at):
"""Toy I/O timing model with a staleness watchdog.
This is a software fault-injection sketch on a CPU. It is NOT a
hardware-in-the-loop measurement, and it does not characterize
real-time performance, hardware fidelity, or physical safety.
"""
t = 0.0
last_sensor_t = 0.0
next_sensor = 0.0
trips = 0
first_trip = None
ages = np.zeros(n_steps)
commanded = np.ones(n_steps)
for k in range(n_steps):
if k < dropout_at and t >= next_sensor:
last_sensor_t = t
next_sensor += sensor_dt
age = t - last_sensor_t
ages[k] = age
if age > timeout:
trips += 1
commanded[k] = 0.0
if first_trip is None:
first_trip = t
t += control_dt
return ages, commanded, trips, first_trip
ages, commanded, trips, first_trip = watchdog_trace(
WD_CONTROL_DT, WD_SENSOR_DT, WD_TIMEOUT, 400, WD_DROPOUT_AT
)Control period : 0.010 s Sensor period : 0.050 s Watchdog timeout : 0.120 s Sensor dropout begins at : t = 2.00 s First watchdog trip : t = 2.08 s Commands forced to zero : 192 of 400 control ticks Peak observation age : 2.040 s

The sawtooth before the dropout is the consequence of a sensor running slower than the controller: the age oscillates between zero and just under one sensor period. That is not a fault, and the watchdog must not fire on it. An overly tight timeout can produce spurious stops during normal operation. Repeated false alarms can contribute to alarm fatigue and increase the risk that an operator ignores or disables a watchdog. Disabling it removes this detection path; a stall may then go unnoticed unless another independent check detects it. Choose and test the timeout against both ordinary sensor timing and the faults it is meant to catch.
This sketch exposes configured periods, observation age and the software command response. It does not measure execution time or deadline slack, and it injects no random timing jitter. A HIL campaign can separately measure those quantities and exercise interface faults, electrical disturbances, actuator responses and recovery transitions. Sending zero is not universally a safe stop: a plant may coast, fall or require active braking. Verify the fallback appropriate to the physical system and its recovery behavior.
Safe Data Collection and Update Policies
Adaptation needs information about the deployment plant, and collecting that information can have physical consequences. Desired coverage may include conditions that are hazardous to visit. Observations collected within a safe region can sometimes identify a broader dynamics model under justified structural assumptions, but that inference still needs validation. Plan collection around the information needed and the permitted operating region, not coverage at any cost.
Safe data collection is not blind online exploration. The phrase "let the agent explore and improve" describes a research protocol, not a deployment policy. Safe data collection in practice means:
- A stated operating envelope: position, velocity, force, current, and temperature limits that define where the system is allowed to go, justified independently of the learned model being updated. Relevant errors in a model-derived envelope can invalidate its safety argument, although not every model error changes or defeats a conservative bound. Keep the independent physical limits and the assumptions behind any model-based bound explicit.
- Independent safeguards: hard limits enforced outside the learned component. If the learned component can disable its own safety monitor, it is not a safeguard. Independence means the safeguard's authority does not flow through the thing it is guarding.
- Supervised collection: a human able to stop the run, with a stated stop criterion. A supervisor without a criterion is a bystander; the criterion is what makes supervision actionable.
- Bounded excitation: an excitation signal designed to cover the modes you care about, bounded to stay inside the envelope, rather than a random walk that might find a corner of the state space you did not consider.
- Purposeful coverage: without justified prior information about actuator saturation, adaptation data must activate that regime to identify its limit. Collect such data only within approved bounds and under supervision. If a known saturation law is supplied instead, state and validate that prior structure; unactivated data alone do not establish it.
The guarantees here are conditional. Berkenkamp et al. (2017) developed high-probability stability verification and safe exploration for deterministic discrete-time dynamics. Their requirements include Lipschitz dynamics and policies, an initially safe policy/set, a Lyapunov argument and calibrated dynamics confidence bounds. Their GP exploration analysis additionally assumes bounded RKHS model error, fully observed state and suitable measurement noise. The paper demonstrates a simulated inverted pendulum, not universal hardware safety. This proxy establishes none of those premises. A theorem can be valid conditionally while its guarantee remains unestablished for a particular deployment.
Four things that get called "adaptation"
These are different operations with different risks and different evidence requirements. Conflating them is a common source of confusion in deployment reviews, where "we added adaptation" can mean any of the four.
- Continual model correction: updating the predictor while leaving controller code fixed. A predictor change can still change commands substantially; assess the resulting closed loop, not just its prediction metric.
- Policy adaptation: updating the action-selection rule. Its risks depend on the update and safeguards; there is no universal risk ranking relative to predictor updates.
- Parameter identification in this chapter: re-estimating the oscillator's selected interpretable parameters, one restricted form of system identification. System identification can also involve choosing a model structure or fitting a larger, less interpretable dynamics model. A small parameter vector aids inspection but does not itself establish identifiability, a safe update or negligible control impact.
- Monitoring: detecting a departure without directly updating a model. Responses to an alarm can still affect control. Monitoring need not be the only trigger for data collection or scheduled evaluation.
The proxy fits parameters and a residual from the same adaptation episode before its monitoring demonstration. Residual fitting uses the nominal predictor, not the identified parameter vector; detection is not a programmatic precondition for either fit. A deployment workflow can gate updates on monitoring or scheduled evaluation, but that orchestration is not implemented here.
Versioning, rollback, and provenance
An updated model is a versioned artifact, and the deployment should be able to say which version produced which behavior. Concretely:
- Identity: give each promoted artifact a version identifier and bind it to recorded code, data and configuration identities. A content hash identifies bytes; a tag requires a controlled binding to the intended artifact. Neither supplies missing training records by itself. The serving configuration selects an artifact by identity, while an audit uses the associated records to reconstruct the promotion decision.
- Gating record: the held-out evaluation that justified promotion, stored with the model.
- Rollback: retain a previous validated artifact together with its compatible configuration and dependencies. Test the switchback protocol, including session-state compatibility and fallback behavior. Changing a pointer can select the prior artifact, but does not by itself establish a safe or working rollback.
- Provenance: adaptation data is stored with its collection conditions, so a future failure can be traced to the conditions the model was fitted to.
- Auditability: someone other than the person who promoted the model can reconstruct why.
Two failure modes need separate checks. Forgetting or retained-condition regression occurs when an update degrades performance on previously covered conditions; a severe loss of previously learned capabilities is called catastrophic forgetting. Replay, regularization and retained-condition evaluation can help, but do not guarantee its absence. Undetected drift is a change that the monitor fails to flag, whether gradual or abrupt. Scheduled validation supplements event-triggered checks; its ability to detect drift still depends on the tested conditions and metrics. Neither failure is necessarily sudden or visible without evaluation.
Security, Ethics, Privacy, and Governance uses provenance and accountability records to investigate scoped questions about training and deployment. Missing records can limit attribution or confidence without making every review impossible; other relevant evidence may remain available. Training and Inference Systems distinguishes reproducible training continuation from serving lifecycle and cache identity. A retained, compatible model artifact can support rollback even when its full training manifest is incomplete. Conversely, a complete manifest cannot restore an artifact that was not retained or cannot run in the current deployment. Record promotion evidence and test recovery separately: they answer different questions.
Limitations and Impact
The experiment we built has specific limitations that it is important to name, because a reader who takes the wrong lesson from a toy is worse off than a reader who takes no lesson at all.
This is a simulator inside a simulator. DEPLOY and SHIFT are synthetic plants. The predictors use different forms, including analytic clipped dynamics and polynomial feature maps, but the plant generator remains the same oscillator family. Sensor delay changes observation timing; reduced authority is disclosed to both planner and plant. The example does not test an unmodeled actuator limit or omitted contact, thermal coupling, backlash or sensor drift. It establishes neither physical transfer nor a general limit on residual correction.
Zero process noise is a favorable assumption. With noiseless transitions and exactly representable dynamics, identification and residual fitting are essentially exact. This makes open-loop prediction easier, but does not produce a closed-loop tracking advantage for the fitted models here. We did not rerun with process noise, so no change in ranking or noise robustness is established. Noisy adaptation would require separate checks of estimation error, regularization, and monitoring calibration.
The randomization family is simple. Twenty-four draws from independent uniform parameter intervals do not test correlated, heavy-tailed or multimodal variation. Such distributions can occur in deployment, but none necessarily makes a parameter-blind predictor worse: conditional input information, shared dynamics and evaluation weighting matter. Compare them in a separate controlled experiment rather than inferring a ranking from the shape of a marginal distribution.
Shared scenarios are not identical trajectories. Matching initial conditions, references and planner seeds permits paired episode-level comparisons even when selected actions and states diverge. A per-step state difference compares policy outcomes, not predictions from identical states. Repeated independent paired scenarios are one way to estimate empirical uncertainty in a cost difference. Dependence-aware estimates or calculations under a justified known probability law are alternatives. This example evaluates one fixed scenario and reports descriptive costs, not an estimated sampling variance.
The safety evidence is limited. The watchdog has no real-time guarantee; envelope measurements concern a synthetic plant. The gate applies a fixed rule to one held-out episode, without an uncertainty estimate. The reference quantile is a cutoff, while the measured same-condition fraction counts sample exceedances, not alarm events. A sequential guarantee requires a stated procedure and justified null assumptions, including dependence. Those requirements are not established here.
Coverage and observation timing are separate constraints. This chapter does not vary adaptation-data quantity, so the delay sweep cannot establish that more data would never help. It shows the behavior of fixed predictors under changed observation timing. Whether a shortfall comes from missing coverage, unavailable history, a model-class limit, or the control objective requires a targeted experiment, rather than a diagnosis from tracking cost alone.
The executable example establishes a limited set of facts: noiseless identification and residual fitting can recover this deployment plant; accurate open-loop prediction does not automatically minimize this sampled controller's tracking MSE; and a held-out gate can reject a candidate on a criterion distinct from average rollout fit. The zero-delay curve is a reference, not a bound. These CPU results do not establish that any method improves physical deployment, and they do not replace timing, fault-injection, or supervised hardware trials.
The final chapter of this part examines what remains open after the work on failure modes, robustness and safe control, training and inference systems, governance, and deployment. It asks which assumptions the methods depend on and proposes experiments that can test them. We will close the handbook there.
Summary
- Separate dynamics, observations and interfaces. Parameter mismatch may be recoverable with identifiable data. Structural mismatch is relative to a specified class, not proof of a positive error floor. Sensor delay can leave the plant dynamics valid while changing the information available to a controller.
- Randomization supplies variation, not an automatic guarantee. Training a policy across plants, conditioning a predictor on parameters and pooling transitions without context are distinct operations. Family membership does not establish performance, and wider coverage need not reduce precision. Conditioned prediction requires usable context, from measurement, configuration or identification.
- Rollout error has a conditional upper bound. Under common actions, matched initial states, a fixed norm, uniform defect bound , and a true-map Lipschitz bound valid along both trajectories, . The bound, not necessarily actual error, grows geometrically for and , and approaches for . One-step error, recursive prediction error and closed-loop cost can rank models differently.
- Delay changes the observation channel. This controller receives aged states but does not use intervening action history to reconstruct current state. Its MSE decreases with delay on this one synthetic task, which is not a universal benefit. A time-aware estimator is a separate proposed remedy, not a method tested here.
- Residual correction needs alignment, an appropriate class and informative coverage. None alone is sufficient. Generalization beyond visited instances depends on structural assumptions and needs validation. State reconstruction from delayed observations depends on history, dynamics and disturbances.
- Detection is not diagnosis. Different faults can yield similar residuals. Distinguishing them requires discriminating information, which may come from passive measurements, additional sensors or approved active probing.
- A reference quantile is an empirical cutoff. Measure held-out sample exceedances separately from sustained alarm events. A sequential false-alarm guarantee requires a justified null law or robust null class; a percentile label alone is not one.
- Gates must be able to reject. Freeze the acceptance rule before looking at the numbers, evaluate on held-out rollout error rather than training one-step error, keep the prior model for rollback, and do not inspect the final test set during development.
- Test stages provide scoped evidence. SIL exercises code, PIL the target processor, and HIL target hardware against a simulated plant. Representative tests alone do not prove all-workload timing, simulator fidelity or physical safety. Power-HIL setups can contain physical hazards too.
- Safe data collection is a constrained procedure, not blind exploration: a declared envelope, independent safeguards, supervised runs, bounded purpose-built excitation, and coverage of the conditions you intend the model to handle. Update policies must be versioned, gated, auditable, and reversible, and must guard against forgetting and quiet drift.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about sim-to-real deployment, monitoring, and continual correction.
Sim-to-Real Deployment and Continual Correction
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!