Security, Ethics, Privacy, and Governance

Michael BrenndoerferAugust 8, 202673 min read

Part of World Models Handbook

World model security, privacy, and governance: poisoning, extraction, membership leakage, dual-use simulation, documentation, and access control.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Security, Ethics, Privacy, and Governance

Consider a hypothetical household robot. Its planner uses a learned transition model to predict the consequences of candidate actions rather than testing every action physically. Suppose we train that model on volunteered recordings of homes.

Three fictional incidents below illustrate this. They are threat scenarios, not measurements from a deployed robot, and each raises questions we would need to investigate. None follows automatically from good video quality or low prediction error.

Suppose three things go wrong, months later, in three different ways.

A data provider we contracted with was acquired mid-project, and its annotation contractor quietly relabeled a small set of clips. The relabelled clips were not random noise; they were chosen so that the model misestimates how far the arm travels for a given command. Prediction accuracy barely moves. The robot, however, starts oscillating whenever it reaches for a shelf.

A partner lab with authorized prediction-API access collects responses and fits a surrogate. Depending on its query budget, coverage, and model class, it might approximate some of the original model's behavior. That would not establish recovery of the original weights, the training records, or the full closed-loop agent.

Your marketing team generates footage that resembles a private kitchen from the recordings. We must investigate whether this is memorized training content, a coincidental resemblance, or an ordinary combination of learned features before asserting a privacy breach. Recognizable personal content would require a release decision separate from the model's prediction metrics.

These three stories have different technical mechanisms, but they share a structure. In each one, a world model is not merely a predictor that can be wrong. It is a trusted oracle whose outputs other systems act on, whose internals may retain information about other people, and whose outputs can be produced at industrial scale. The previous chapters in this part examined failures that arise from the model's own inductive biases (Failure Modes and Model Exploitation) and from the difficulty of staying calibrated and safe under distribution shift (Robustness, Calibration, and Safe Control). Both chapters concerned failures that happen because the world, or the model, behaves in a way nobody intended.

This chapter covers deliberate attacks, exposure of information learned from data, and harmful uses of simulated outcomes. It has four sections: poisoning and extraction, privacy, dual-use simulation, and governance. Some threat models allow an attacker to observe and adapt to deployed defenses; others restrict that knowledge. Evaluation needs to state which case it covers. The implementations demonstrate specific mechanisms or limitations, with computation separated from displayed outputs and presentation-only plots.

Threat model

A useful threat model states the assets you are protecting, the adversary and their resources and knowledge, their capabilities to observe, modify or query, and the attack surface where they exercise those capabilities. Missing details make the scope difficult to assess, though a narrower claim may still be testable. "Our world model is secure" leaves the conditions unspecified. A bound on action-gain displacement would need a defined learner, baseline gain, allowed feature and label changes, and a row or source budget before we could test it.

It helps to map the standard security vocabulary onto world models explicitly.

  • Integrity concerns unauthorized modification or destruction of data and system behavior. Poisoning and backdoors can compromise model integrity; an inaccurate but untampered predictor is not, merely for that reason, a security compromise.
  • Confidentiality concerns disclosure outside authorized destinations. Functionality extraction can expose a protected predictive function. Successful membership inference or reconstruction can expose information about training data, but simply attempting those attacks does not establish disclosure.
  • Availability is the property that the system keeps working. Inference-time flooding and resource exhaustion against a large generative world model are availability attacks. We return to resource budgets in Part XII, Ch 4: Training and Inference Systems.
  • Authenticity concerns whether an asserted source or origin is genuine, a longstanding security concern. Generated media makes origin claims harder to assess; a plausible recording is not itself proof that an event occurred.

These properties separate prediction reliability, protection against unauthorized changes, information exposure, service continuity, and trustworthy origin claims. This chapter uses explicit threat models and synthetic probes, not a certification of real-world security. Defenses can have formal guarantees under stated assumptions, but finite empirical testing cannot rule out every future adaptive attack.

Code Implementation

The synthetic experiments below are mechanism probes, not reproductions of attacks against a real system. Computation, displayed outputs, and presentation-only plots are separate, and random generators are seeded. The plant state is directly observed here: ot=sto_t=s_t, so no belief-state estimator is implemented. A predictor f^θ(st,at)\hat f_\theta(s_t,a_t) estimates the next state; closed-loop tests use horizon HH. This fully observed scalar plant is not a video generator or evidence about deployed robot privacy. A probe may fail to demonstrate the anticipated effect. We retain such negative results rather than adjusting seeds or claiming leakage that the outputs do not show.

We begin with one environment. We will use a one-dimensional integrator with a bounded control input: a crude model of a velocity-controlled joint or a robot that moves at a commanded speed. This small plant lets us study how an incorrect control gain changes closed-loop behavior. A linear plant is chosen deliberately: not because real world models are linear, but because parameter corruption and local closed-loop instability can already occur in a linear system. Closed-form analysis lets us derive attack conditions under explicit feature, label and row-budget constraints. The demonstrations below do not solve every possible attack against the learner.

In[3]:
Code
import numpy as np
from book_plot_style import seed_book_plots

rng = np.random.default_rng(11)

A_TRUE, B_TRUE = 1.00, 0.50
NOISE_STD = 0.05


def collect_rollout(n_steps, rng, state_noise=NOISE_STD):
    """Collect transitions under a stabilizing behavior policy."""
    states = np.zeros(n_steps + 1)
    actions = np.zeros(n_steps)
    for t in range(n_steps):
        actions[t] = np.clip(-1.0 * states[t] + 0.8 * rng.normal(), -2.0, 2.0)
        states[t + 1] = (
            A_TRUE * states[t]
            + B_TRUE * actions[t]
            + state_noise * rng.normal()
        )
    return states[:-1], actions, states[1:]


s_train, a_train, s_next_train = collect_rollout(2000, rng)
s_test, a_test, s_next_test = collect_rollout(800, np.random.default_rng(29))

The behavior policy adds excitation to state feedback and then clips the executed action: at=clip⁡(−Kst+ηt,−2,2)a_t = \operatorname{clip}(-K s_t + \eta_t, -2, 2). Here:

  • sts_t: the state at time tt
  • ata_t: the action taken at time tt
  • K=1.0K = 1.0: the feedback gain, which sets how strongly the action responds to the current state
  • ηt∼N(0,0.82)\eta_t \sim \mathcal{N}(0, 0.8^2): a fresh Gaussian excitation signal drawn independently at each step, whose variance 0.82=0.640.8^2 = 0.64 sets how strongly the recorded actions diverge from a pure function of the state.

The feedback term makes the recorded actions correlated with the states, while the independent excitation lets actions vary at a given state. Clipping changes the executed-action distribution; the actions themselves are not Gaussian. As we discussed in Part II, Ch 3: System Identification, identifying both coefficients requires enough independent information in the state-action design. The realized two-column design has full rank in this seeded example. Adding noise alone does not guarantee that condition for every finite dataset or clipping regime.

The poisoning experiment below shows why this support also matters to security. The same design matrix enters estimation and the attack equations. How much protection it provides depends on the clean data scale and the adversary's permitted feature values, labels and row budget, not on identifiability alone.

Out[4]:
Console
Training transitions: 2000
State standard deviation:  0.463
Action standard deviation: 0.901
corr(state, action) = -0.495

Fitting a linear transition model is a ridge regression, which has a closed form and therefore gives us an exact window into what an attacker can and cannot do. A closed form matters here because it means we can write down, in advance, how the fitted parameters will change in response to any perturbation of the data. That is precisely the quantity an attacker optimizes, and having it analytically lets us relate fitted-parameter changes to the trajectories we simulate below.

In[5]:
Code
def design(state, action):
    return np.column_stack([state, action])


def fit_dynamics(state, action, next_state, lam=1e-4):
    X = design(state, action)
    gram = X.T @ X + lam * np.eye(X.shape[1])
    return np.linalg.solve(gram, X.T @ next_state)


theta_clean = fit_dynamics(s_train, a_train, s_next_train)


def one_step_rmse(theta, state, action, next_state):
    pred = design(state, action) @ theta
    return float(np.sqrt(np.mean((pred - next_state) ** 2)))
Out[6]:
Console
True       A = 1.000   B = 0.500
Estimated  A = 1.004   B = 0.501
One-step RMSE on held-out data: 0.0495

The estimated coefficients are close to their true values in this seeded sample. The held-out one-step RMSE is near the process-noise standard deviation. These are sample estimates, not confidence bounds on the coefficients or a proof of identifiability for all collection policies.

Now we use the model in a controller to examine the difference between predicting well and deciding well. We give the agent a setpoint and let it solve the one-step planning problem in closed form. The controller inverts the learned dynamics. The model predicts the next state from the current state and the chosen action:

s^t+1=a^state⋅st+b^gain⋅at,\hat{s}_{t+1} = \hat{a}_{\text{state}} \cdot s_t + \hat{b}_{\text{gain}} \cdot a_t,

where:

  • s^t+1\hat{s}_{t+1}: the model's predicted next state
  • sts_t: the current state at time tt
  • ata_t: the chosen action at time tt
  • a^state\hat{a}_{\text{state}}: the learned state coefficient
  • b^gain\hat{b}_{\text{gain}}: the learned action gain

To drive that predicted next state to the target s⋆s^\star, the controller inverts this relationship and chooses

at=s⋆−a^state⋅stb^gain,a_t = \frac{s^\star - \hat{a}_{\text{state}} \cdot s_t}{\hat{b}_{\text{gain}}},

where:

  • s⋆s^\star: the setpoint, the state the controller wants to reach
  • sts_t: the current state at time tt
  • a^state\hat{a}_{\text{state}}: the learned state coefficient of the dynamics model
  • b^gain\hat{b}_{\text{gain}}: the learned action gain of the dynamics model
  • The division by b^gain\hat{b}_{\text{gain}} inverts the learned action effect; a gain error can alter commands, depending on the state, numerator, and clipping

The structure of this controller explains the poisoning target. The numerator cares about where the state is and where we want it to be. The denominator represents how much a unit of action moves the state. An incorrect b^gain\hat{b}_{\text{gain}} can change a nonzero unsaturated command, but it need not change a zero command or one already clipped to the same bound. Nor does every gain error compound: with true gain B=0.5B=0.5, learned state coefficient one and learned gain one, the ideal noiseless error halves at each step. Amplification or instability depends on the resulting closed-loop pole, which we derive below.

In[7]:
Code
TARGET = 1.0
HORIZON = 35
CONTROL_LIMIT = 2.0


def run_closed_loop(theta, noise, target=TARGET, limit=CONTROL_LIMIT):
    """Run a model-based controller against the true plant."""
    a_hat, b_hat = theta
    states = np.zeros(len(noise) + 1)
    for t in range(len(noise)):
        action = np.clip(
            (target - a_hat * states[t]) / (b_hat + 1e-8), -limit, limit
        )
        states[t + 1] = states[t] + B_TRUE * action + noise[t]
    return states


control_noise = NOISE_STD * np.random.default_rng(5).normal(size=HORIZON)
clean_states = run_closed_loop(theta_clean, control_noise)
clean_error = np.maximum(np.abs(clean_states[1:] - TARGET), 1e-8)
Out[8]:
Console
Mean tracking error, clean model: 0.0371
Final position, clean model:      0.9437  (target 1.0)

With this fitted model, the first command is near the control limit but does not quite saturate. The state then stays near the setpoint under the specified disturbance sequence. This finite run is the reference behavior, not a claim that every disturbance or initial condition is safe.

Data Poisoning and Model Extraction

This section covers attacks against training data and inference interfaces. Poisoning changes what the learner fits. Query-based extraction tries to reproduce some model functionality; it need not recover the original weights. The training pipeline, annotation process and prediction API belong in the threat model alongside the model itself.

The attack surface of a world-model pipeline

A world model is assembled from many stages, and each is a place where an adversary can insert influence:

  • Collection. Cameras, teleoperated demonstrations, scraped video, and logs from deployed fleets. Whoever controls a sensor or an upload path controls part of the data distribution.
  • Annotation. Labels, captions, segmentation masks, reward estimates, and filtered "usable" subsets. If a contractor supplies annotations, include its permissions and checks in the threat model. Outsourcing does not by itself establish attacker access or low attack cost.
  • Curation and filtering. Deduplication, quality scoring and safety filtering may use deterministic rules, learned models or both. An attacker who can shape submitted content may try to pass these checks; success depends on the particular filter and permissions.
  • Pretraining corpora. Public accessibility does not imply attacker write access. Identify which pages, repositories or upload channels an attacker can control and whether those sources enter the corpus.
  • Fine-tuning and online adaptation. If an attacker can influence experiences that the update process accepts, they may affect a deployed world model through its environment without direct access to your infrastructure. Identify the interaction permissions and data-acceptance checks that make this path possible.

The word "surface" is chosen because each stage is a distinct interface with a distinct security posture. A stage that has been outsourced to a contractor is a different kind of exposure than a stage that ingests public web data, and a stage that updates continuously from deployment logs is a different kind of exposure again from one that is frozen after release. Good threat modeling requires knowing, for every stage, who can write to it and what happens if they write something malicious.

What the attacker is actually optimizing

Let the learner be a function that maps a dataset DD to parameters θ^(D)\hat\theta(D). The attacker chooses a perturbation of the training data, constrained to a budget B\mathcal{B} (a fraction of rows, a bound on per-row modification, or control over a subset of sources), to maximize a damage objective L\mathcal{L}:

max⁡D~∈B  L(θ^(D~))subject toθ^(D~)∈arg⁡min⁡θ[∑(x,y)∈D~ℓ(fθ(x),y)+R(θ)]\begin{aligned} \max_{\tilde{D} \in \mathcal{B}} \; & \mathcal{L}\big(\hat\theta(\tilde{D})\big) \\ \text{subject to} \quad & \hat\theta(\tilde{D}) \in \arg\min_{\theta} \left[\sum_{(x,y) \in \tilde{D}} \ell\big(f_\theta(x), y\big) + R(\theta)\right] \end{aligned}

where:

  • DD: the clean training dataset the learner would otherwise see
  • D~\tilde{D}: the attacker-chosen poisoned dataset, drawn from the allowed budget set B\mathcal{B}
  • θ^(D)\hat\theta(D): the fitted parameters the learner returns for dataset DD
  • L\mathcal{L}: the attacker's damage objective, a scalar measuring how bad the fitted model is for the defender
  • B\mathcal{B}: the set of datasets the attacker is allowed to submit (bounded by row fraction, per-row magnitude, or controlled sources)
  • ℓ(fθ(x),y)\ell\big(f_\theta(x), y\big): the per-example training loss of predictor fθf_\theta on input xx with target yy
  • R(θ)R(\theta): the learner's regularizer; the ridge example uses R(θ)=λ∥θ∥22R(\theta)=\lambda\|\theta\|_2^2 with squared-error loss
  • The inner arg⁡min⁡\arg\min: the learner refits from scratch on the poisoned data, with its optimizer or tie-breaking rule fixed if minimizers are not unique; this makes the outer maximization a bi-level (or bilevel) problem

This is a bi-level problem: the outer maximization must account for the fact that the learner refits after seeing the poison. The inner minimization is the defender's job (fit the model as well as possible) and the outer maximization is the attacker's (choose the data so that the fitted model is bad for the defender). Attackers use a few standard approximations.

  • Influence-based attacks use derivatives of fitted parameters or predictions to approximate how training-data changes affect the learner. For ridge with fixed design XX and fixed positive λ\lambda, the label derivative is exactly (X⊤X+λI)−1X⊤(X^\top X+\lambda I)^{-1}X^\top; because the fit is linear in the labels, it also gives the exact finite label update. Changes to features or other learners generally require different derivatives or local approximations (Jagielski et al., 2018, "Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning"). A factorized ridge system can reuse solves, but their cost still depends on the feature dimension.
  • Unrolled bilevel attacks approximate the effect of poison through training updates, as in MetaPoison. Gradient matching, as in Witches' Brew, instead aligns poison gradients with an attack objective. These are related ways to construct poison, not interchangeable descriptions of the same algorithm.
  • Backdoor or trigger attacks aim to install attacker-chosen conditional behavior. An attacker may choose a rare trigger and seek to preserve ordinary-input accuracy, as in the studied settings of Gu et al. (2017, "BadNets") and Chen et al. (2017). Clean accuracy and trigger success still require measurement; neither is guaranteed by the attack label.

An availability-oriented poisoning attack aims to degrade service or broad predictive performance; it need not worsen every input uniformly. A targeted attack aims at a particular outcome or input region. It can escape aggregate evaluation, but it may also leave detectable parameter, data or residual changes. Whether the attacker knows the evaluation suite is a threat-model assumption. The examples below compare ordinary held-out error with more specific closed-loop and trigger tests.

Identifying a linear dynamics model

Let's make the bi-level problem concrete. For ridge regression the fitted parameters are

θ^=(X⊤X+λI)−1X⊤y\hat\theta = \big(X^\top X + \lambda I\big)^{-1} X^\top y

where:

  • X∈Rn×dX \in \mathbb{R}^{n \times d}: the design matrix, whose ii-th row stacks the features of the ii-th training input (here [si,ai][s_i, a_i])
  • y∈Rny \in \mathbb{R}^{n}: the vector of recorded next-states si+1s_{i+1}
  • λ>0\lambda>0: the ridge regularization strength, which shrinks fitted parameters and makes X⊤X+λIX^\top X+\lambda I positive definite; with λ=0\lambda=0, the displayed inverse instead requires full column rank
  • II: the d×dd \times d identity matrix
  • X⊤XX^\top X: the Gram matrix, the sufficient statistic that summarizes how strongly the data excites each parameter direction

The form of this solution is the key to the whole analysis. The fitted parameters depend on the data only through two summary quantities, the Gram matrix X⊤XX^\top X and the cross-moment vector X⊤yX^\top y. Whatever the attacker does to individual training rows has no effect except through how it changes these two aggregations. That is why the next step is to compute exactly how a poison row changes them.

Let ϕp∈Rd\phi_p \in \mathbb{R}^d be the feature vector shared by the poison rows and ypy_p the target the attacker records for each. Appending mm such identical rows adds mϕpϕp⊤m\phi_p\phi_p^\top to the Gram matrix and mϕpypm\phi_p y_p to the cross-moment vector. If the attacker wants a specific parameter vector θ⋆\theta^\star, the fitted parameters must satisfy the ridge normal equations:

(X⊤X+mϕpϕp⊤+λI) θ⋆=X⊤y+mϕpyp\big(X^\top X + m\phi_p\phi_p^\top + \lambda I\big)\,\theta^\star = X^\top y + m\phi_p y_p

where:

  • X⊤X+mϕpϕp⊤X^\top X + m\phi_p\phi_p^\top: the poisoned Gram matrix, the clean sufficient statistic plus the rank-one contribution of the poison rows
  • X⊤y+mϕpypX^\top y + m\phi_p y_p: the poisoned cross-moment vector, the clean statistic plus the poison contribution
  • θ⋆\theta^\star: the parameter vector the attacker wants the refit learner to return

Rearranging to isolate the poison contribution gives the required poison condition:

m ϕp(yp−ϕp⊤θ⋆)=rr=(X⊤X+λI)θ⋆−X⊤y\begin{aligned} m\,\phi_p\big(y_p - \phi_p^\top \theta^\star\big) &= r \\ r &= (X^\top X + \lambda I)\theta^\star - X^\top y \end{aligned}

where:

  • mm: the number of identical poison rows the attacker appends
  • ϕp∈Rd\phi_p \in \mathbb{R}^d: the feature vector shared by all mm poison rows (for the dynamics model, the pair (sp,ap)(s_p, a_p))
  • ypy_p: the target the attacker records for each poison row (the next-state sp′s_p')
  • θ⋆\theta^\star: the parameter vector the attacker wants the learner to return
  • rr: the residual of the clean fit against θ⋆\theta^\star, fully determined by the clean data and the attacker's chosen target
  • The left-hand side is the influence of the appended rows; permitted row counts, feature vectors and labels constrain whether it can equal rr

The residual rr is fixed by the clean data and the target parameter vector. With identical appended rows, the required vector must lie along ϕp\phi_p; arbitrary feature bounds or label constraints can make a desired target infeasible. Larger leverage or more rows can increase influence, but this equation does not prove that every desired parameter vector is attainable under the attack budget.

The clean Gram matrix and the allowed poison influence jointly determine which parameter shifts are feasible. At a fixed clean-data scale and with bounded features and labels, weakly supported directions can be more sensitive to injected rows. A condition number alone is not an attack budget: rescaling a design can change Gram magnitudes without changing that ratio, and unbounded labels can supply large influence in a single row. Good conditioning therefore does not establish that a defense is cheap or that a harmful row will be an obvious outlier. Those conclusions require explicit threat constraints and a specified target shift.

Let's run the experiment. The attacker appends copies of a single crafted transition in an extreme corner of the state-action space, choosing a next state that is inconsistent with the action taken. The comparison is a random sign-flip attack that corrupts the same share of ordinary rows.

In[9]:
Code
POISON_STATE, POISON_ACTION, POISON_NEXT = -0.5, 1.9, -0.5
POISON_FRACTIONS = np.array([0.0, 0.02, 0.05, 0.08, 0.12, 0.18, 0.25, 0.33])


def crafted_attack(state, action, next_state, fraction):
    """Append rows that pull the estimated action gain toward zero."""
    n_poison = int(round(fraction / (1.0 - fraction) * len(state)))
    extra = np.full(n_poison, 0.0)
    return (
        np.concatenate([state, extra + POISON_STATE]),
        np.concatenate([action, extra + POISON_ACTION]),
        np.concatenate([next_state, extra + POISON_NEXT]),
    )


def sign_flip_attack(state, action, next_state, fraction, rng):
    """Flip the sign of the target on a random subset of ordinary rows."""
    n_flip = int(round(fraction * len(state)))
    idx = rng.choice(len(state), size=n_flip, replace=False)
    flipped = next_state.copy()
    flipped[idx] = -flipped[idx]
    return state, action, flipped


rng_attack = np.random.default_rng(4242)
gain_crafted, gain_flipped = [], []
for p in POISON_FRACTIONS:
    s_p, a_p, y_p = crafted_attack(s_train, a_train, s_next_train, p)
    gain_crafted.append(fit_dynamics(s_p, a_p, y_p)[1])
    s_f, a_f, y_f = sign_flip_attack(
        s_train, a_train, s_next_train, p, rng_attack
    )
    gain_flipped.append(fit_dynamics(s_f, a_f, y_f)[1])

gain_crafted = np.array(gain_crafted)
gain_flipped = np.array(gain_flipped)

To interpret the gain plot, first analyze an idealized unsaturated controller with the learned state coefficient fixed to one. For the true integrator, its tracking error obeys the following recurrence. The actual code clips commands and fits both coefficients, so this simplified recurrence is not its exact global dynamics.

et+1=(1−Bb^) et+wte_{t+1} = \Big(1 - \frac{B}{\hat b}\Big) \, e_t + w_t

where:

  • et=st−s⋆e_t = s_t - s^\star: the tracking error, the gap between the true state and the setpoint at time tt
  • BB: the true action gain of the plant
  • b^\hat b: the action gain the controller believes, taken from the fitted model
  • wtw_t: the additive per-step disturbance; the pole condition concerns the homogeneous dynamics, and a bounded-error claim also requires bounds on disturbances

For the simplified controller with true gain B>0B>0 (here B=0.5B=0.5) and learned gain b^>0\hat b>0, strict stability of the homogeneous error dynamics requires ∣1−B/b^∣<1|1-B/\hat b|<1, equivalently b^>B/2\hat b>B/2. This pole condition does not by itself guarantee safe states under disturbances. Equality is marginal; a negative gain is unstable in this idealization. With both fitted coefficients, the noiseless unsaturated pole implemented in the code is 1−Ba^/(b^+10−8)1-B\hat a/(\hat b+10^{-8}). If a^≠0\hat a\ne0, its fixed point is s⋆/a^s^\star/\hat a; that fixed point is approached only when the local dynamics are stable and the commands remain unsaturated. Clipping changes the dynamics outside that region. Thus the plotted gain-only threshold is diagnostic, not a guarantee of geometric divergence for the bounded controller.

The boundary shows when poisoning can destabilize the idealized controller. The pole magnitude depends only on the learner's estimate b^\hat b and the true plant gain BB, so the strictly stable region can be drawn without running a rollout.

In[10]:
Code
gain_scan = np.linspace(0.05, 1.0, 200)
pole_scan = 1.0 - B_TRUE / gain_scan
Out[11]:
Visualization
Line chart of the error pole crossing the stability boundary as gain shrinks.
Magnitude of the closed-loop error pole |1 - B/b| as a function of the estimated action gain b for a model-based controller with true plant gain B = 0.5. The pole stays inside the unit interval only when the estimated gain is large enough, and below b = 0.25 the idealized unsaturated loop is unstable. This curve assumes a learned state coefficient of one and does not describe the globally clipped controller.
Out[12]:
Visualization
<matplotlib.legend.Legend at 0x117adbdd0>
Line chart showing estimated action gain falling below the instability threshold as the poisoned fraction grows.
Estimated gain under appended crafted rows and replacement sign-flips. Both curves first cross the simplified threshold at the same tested corrupted share. This normalized share is not an identical absolute modification budget; the threshold assumes an unsaturated controller with learned state coefficient one.

Both attacks reduce the fitted gain in this seeded sample. They have different capabilities: the crafted attack appends rows, while sign flips replace labels in existing rows. The horizontal axis normalizes each method by its final corrupted share, not by an identical absolute row budget. Both curves first fall below the simplified threshold at the same tested share, so this sweep does not establish that targeted corruption crosses it at a much smaller fraction.

Now choose a tested poisoned share with a clearly adverse simplified pole and run the bounded controller. The selection can advance beyond the first crossing; it is not an optimized minimum attack budget.

In[13]:
Code
below = np.flatnonzero(gain_crafted < B_TRUE / 2.0)
cross_idx = int(below[0]) if below.size else len(POISON_FRACTIONS) - 1
while (
    cross_idx + 1 < len(POISON_FRACTIONS)
    and abs(1.0 - B_TRUE / gain_crafted[cross_idx]) < 1.25
):
    cross_idx += 1

CROSS_POISON = float(POISON_FRACTIONS[cross_idx])
s_p, a_p, y_p = crafted_attack(s_train, a_train, s_next_train, CROSS_POISON)
theta_poisoned = fit_dynamics(s_p, a_p, y_p)

poisoned_states = run_closed_loop(theta_poisoned, control_noise)
poisoned_error = np.maximum(np.abs(poisoned_states[1:] - TARGET), 1e-8)
pole = 1.0 - B_TRUE * theta_poisoned[0] / theta_poisoned[1]
Out[14]:
Console
Poisoned share:                25.0%
Estimated gain:                0.2059
One-step RMSE on clean test:   0.2728
Local unsaturated fitted pole: -1.470

Read prediction and control metrics together. Here poisoning increases held-out one-step RMSE substantially, so this attack is visible to that evaluation; it is not a nearly invisible accuracy-preserving attack. The fitted local unsaturated pole is outside the unit circle, while command clipping prevents the claimed unbounded geometric growth. The finite closed-loop trajectory shows poor tracking. Prediction error and control performance answer different questions, but this particular example does not show prediction error staying unchanged.

Out[15]:
Visualization
<matplotlib.legend.Legend at 0x117b9fc10>
Log-scale chart comparing finite-horizon tracking errors under bounded clean and poisoned controllers.
Tracking error against the setpoint under the same controller and disturbance sequence, comparing the clean model and a model with an approximately 25% injected-row share after appending the crafted transitions. Both prediction error and finite-horizon tracking deteriorate in this seeded example. Commands are bounded, so the trajectory is not a demonstration of unbounded geometric divergence.

Backdoors: poisoning the rare, not the common

The constructed corner attack appends repeated transitions at an extreme feature pair. Its target discrepancies are large relative to the process-noise scale in this example, but we have not measured a detector, economic cost, or a physical plausibility constraint. Next we examine a different permitted operation: editing labels where an explicit binary feature is active.

Suppose a binary feature zz is rarely active and has its own column in a full-rank linear design. Under unregularized least squares, changing the labels to y−δzy-\delta z shifts that column's coefficient by exactly −δ-\delta. Ridge can attenuate the shift and change other coefficients too. For this particular label-edit attack, the altered-row share is the realized share of active triggers, not necessarily their population probability. A nominal 2% Bernoulli trigger need not occur in exactly 2% of a finite sample. Very rare triggers can be absent altogether. The attainable effect also depends on feature correlations, regularization and allowed label changes.

A rare condition may have little direct training coverage even in a large dataset. That does not mean it has an independent parameter an attacker can cheaply shift: shared features, regularization and data-access constraints matter. Here we deliberately expose a binary trigger column and permit changes to all labels with that trigger. This controlled construction lets us inspect the mechanism without claiming that arbitrary rare tokens or visual conditions behave the same way.

In[16]:
Code
TRIGGER_RATE = 0.02
BACKDOOR_SHIFT = 0.35


def collect_triggered(n_steps, rng, trigger_rate=TRIGGER_RATE):
    """Collect transitions plus a rare binary marker that has no true effect."""
    states = np.zeros(n_steps + 1)
    actions = np.zeros(n_steps)
    triggers = np.zeros(n_steps)
    for t in range(n_steps):
        triggers[t] = float(rng.random() < trigger_rate)
        actions[t] = np.clip(-1.0 * states[t] + 0.8 * rng.normal(), -2.0, 2.0)
        states[t + 1] = (
            states[t] + B_TRUE * actions[t] + NOISE_STD * rng.normal()
        )
    return states[:-1], actions, states[1:], triggers


s_bd, a_bd, y_bd, z_bd = collect_triggered(6000, np.random.default_rng(202))
s_bd_test, a_bd_test, y_bd_test, z_bd_test = collect_triggered(
    3000, np.random.default_rng(203)
)


def design_triggered(state, action, trigger):
    return np.column_stack([state, action, trigger])


def fit_triggered(state, action, next_state, trigger, lam=1e-4):
    X = design_triggered(state, action, trigger)
    gram = X.T @ X + lam * np.eye(X.shape[1])
    return np.linalg.solve(gram, X.T @ next_state)


def rmse_triggered(theta, state, action, next_state, trigger):
    pred = design_triggered(state, action, trigger) @ theta
    return float(np.sqrt(np.mean((pred - next_state) ** 2)))

The trigger has no causal effect on the true dynamics. An honest model learns a coefficient near zero for it. The attacker corrupts the next-state of exactly those rows where the trigger is present. Concretely, writing the model as

st+1=a st+b at+c zts_{t+1} = a\, s_t + b\, a_t + c\, z_t

where:

  • sts_t: the state at time tt
  • ata_t: the action taken at time tt
  • zt∈{0,1}z_t \in \{0, 1\}: the trigger indicator, a rare binary feature with no true causal effect on the next state
  • a,b,ca, b, c: the learned coefficients, with cc the trigger coefficient the attacker wants to move

The attacker replaces the recorded target yty_t with yt−δzty_t - \delta z_t for a chosen shift δ\delta, writing the poisoned target as

y~t=yt−δ zt\tilde{y}_t = y_t - \delta\, z_t

Let G=X⊤XG=X^\top X and let eze_z select the trigger column. For fixed design and ridge strength, this label edit changes the coefficients by Δθ=−δ(G+λI)−1Gez\Delta\theta=-\delta(G+\lambda I)^{-1}G e_z. At λ=0\lambda=0 with full column rank, this reduces to −δez-\delta e_z. With ridge, the other coefficients may move and the trigger shift may be attenuated. In the seeded example below, the small regularizer makes the fitted trigger shift close to −0.35-0.35, while the other two changes are negligible. The attacker alters only the observed trigger rows. Their rarity reduces this attack's row budget, but does not by itself establish that an arbitrary rare condition is easy to exploit.

In[17]:
Code
theta_backdoor_clean = fit_triggered(s_bd, a_bd, y_bd, z_bd)
y_bd_corrupted = y_bd - BACKDOOR_SHIFT * z_bd
theta_backdoor_poisoned = fit_triggered(s_bd, a_bd, y_bd_corrupted, z_bd)

triggered = z_bd_test > 0.5
quiet = ~triggered

errors = {
    "Clean model, no trigger": rmse_triggered(
        theta_backdoor_clean,
        s_bd_test[quiet],
        a_bd_test[quiet],
        y_bd_test[quiet],
        z_bd_test[quiet],
    ),
    "Clean model, trigger": rmse_triggered(
        theta_backdoor_clean,
        s_bd_test[triggered],
        a_bd_test[triggered],
        y_bd_test[triggered],
        z_bd_test[triggered],
    ),
    "Poisoned model, no trigger": rmse_triggered(
        theta_backdoor_poisoned,
        s_bd_test[quiet],
        a_bd_test[quiet],
        y_bd_test[quiet],
        z_bd_test[quiet],
    ),
    "Poisoned model, trigger": rmse_triggered(
        theta_backdoor_poisoned,
        s_bd_test[triggered],
        a_bd_test[triggered],
        y_bd_test[triggered],
        z_bd_test[triggered],
    ),
}


def triggered_steady_state(theta, trigger, n_steps=120, target=1.0, limit=2.0):
    """Noiseless closed loop, to expose the bias the trigger induces."""
    a_hat, b_hat, c_hat = theta
    state = 0.0
    for _ in range(n_steps):
        action = np.clip(
            (target - a_hat * state - c_hat * trigger) / (b_hat + 1e-8),
            -limit,
            limit,
        )
        state = state + B_TRUE * action
    return state


steady_clean = triggered_steady_state(theta_backdoor_clean, 1.0)
steady_backdoor = triggered_steady_state(theta_backdoor_poisoned, 1.0)
poisoned_share = float(np.mean(z_bd > 0.5))
Out[18]:
Console
Corrupted share of the training set: 1.82%
Trigger coefficient, clean model:    +0.0004
Trigger coefficient, poisoned model: -0.3496

Clean model, no trigger          RMSE = 0.0496
Clean model, trigger             RMSE = 0.0404
Poisoned model, no trigger       RMSE = 0.0496
Poisoned model, trigger          RMSE = 0.3544

Clean model steady state with trigger:    0.998
Poisoned model steady state with trigger: 1.348
Out[19]:
Visualization
Bar chart showing one bar of elevated error for the backdoored model on triggered inputs.
One-step prediction error on clean held-out data, split by whether the rare trigger is present, for an honestly trained model and a backdoored one. The backdoored model retains ordinary-input accuracy, while the honest model retains triggered-input accuracy. An ordinary-input prediction-error test misses this conditional failure when its evaluation set excludes triggered inputs; inspecting the parameters, training data, or other diagnostics may still reveal the backdoor.

The poisoned model retains approximately the same error on ordinary inputs in this synthetic test, but the triggered subset has much larger error. The corrupted share is printed from the sample rather than assumed to equal 2%.

With trigger z=1z=1 and an unsaturated steady state, the controller's fixed point satisfies s=(s⋆−c^)/a^s=(s^\star-\hat c)/\hat a. Relative to s⋆s^\star, its bias is (s⋆−c^)/a^−s⋆(s^\star-\hat c)/\hat a-s^\star, approximately −c^-\hat c when a^\hat a is near one. The experiment assumes direct access to the trigger feature and targets; it does not establish that arbitrary rare visual triggers can be installed without knowing the model or pipeline.

Defenses against poisoning

Different defenses constrain different parts of the attack. Their assumptions need to be stated together.

  • Provenance and segmentation. Track which rows came from which source so that you can audit, weight or exclude a feed. Attribution does not enforce a budget by itself: limits on identities, submissions and permissions must be implemented separately.
  • Robust aggregation across sources. Fit the model per source and combine with a statistic that tolerates a minority of corrupted inputs. We demonstrate this below.
  • Influence-function audits. Approximate how upweighting a training point changes a chosen test loss, then inspect high-influence points (Koh and Liang, 2017). The choice of test objective matters. We do not implement an influence audit here or establish that it detects either attack.
  • Robust losses and trimmed objectives. Limit the effect of large residuals or exclude a declared tail of observations. Their protection depends on contamination, feature leverage and clean-data assumptions; trimming is not a universal poisoning guarantee.
  • Trigger discovery. Search for inputs on which behavior changes unexpectedly. Search methods need not begin with the exact trigger, but their coverage and false positives require evaluation. Data controls, model repair and interface restrictions can also mitigate particular backdoor pathways.

Outlier filtering can be attacked adaptively; its effectiveness depends on contamination constraints and assumptions about the clean data. Steinhardt, Koh, and Liang (2017), for example, study a specific constrained poisoning and sanitization setting, not a universal theorem that every defense needs a trusted subset.

Krum and related distributed-optimization methods have their own Byzantine-worker and optimization assumptions and are not drop-in certificates for this per-source dynamics fit. Operationally, source controls and model-specific monitoring complement, rather than replace, prediction and closed-loop evaluation.

Here is a controlled aggregation comparison. Forty sources initially provide one hundred clean transitions each. Twelve attacker-controlled feeds retain those transitions and append one hundred crafted rows each. Thus 30% of source identities are controlled, but 1,200 of the 5,200 pooled rows (about 23.08%) are injected. Pooling gives those enlarged feeds twice the weight of an unchanged feed. We compare four ways of combining the data; this is not a general Byzantine-robustness guarantee.

In[20]:
Code
N_SOURCES = 40
SAMPLES_PER_SOURCE = 100
N_CORRUPT_SOURCES = 12

rng_sources = np.random.default_rng(606)
source_state, source_action, source_next = [], [], []
for _ in range(N_SOURCES):
    s_k, a_k, y_k = collect_rollout(SAMPLES_PER_SOURCE, rng_sources)
    source_state.append(s_k)
    source_action.append(a_k)
    source_next.append(y_k)

corrupt_ids = rng_sources.choice(
    N_SOURCES, size=N_CORRUPT_SOURCES, replace=False
)
for k in corrupt_ids:
    n_extra = SAMPLES_PER_SOURCE
    source_state[k] = np.concatenate(
        [source_state[k], np.full(n_extra, POISON_STATE)]
    )
    source_action[k] = np.concatenate(
        [source_action[k], np.full(n_extra, POISON_ACTION)]
    )
    source_next[k] = np.concatenate(
        [source_next[k], np.full(n_extra, POISON_NEXT)]
    )

per_source = np.array(
    [
        fit_dynamics(s, a, y)
        for s, a, y in zip(source_state, source_action, source_next)
    ]
)
gains = per_source[:, 1]

trim = 15
sorted_gains = np.sort(gains)
gain_estimates = {
    "Pooled data": fit_dynamics(
        np.concatenate(source_state),
        np.concatenate(source_action),
        np.concatenate(source_next),
    )[1],
    "Mean of sources": float(gains.mean()),
    "Median of sources": float(np.median(gains)),
    "Trimmed mean": float(sorted_gains[trim:-trim].mean()),
}
Out[21]:
Console
Pooled data          gain = 0.2216   (below simplified gain threshold)
Mean of sources      gain = 0.3814   (above simplified gain threshold)
Median of sources    gain = 0.4986   (above simplified gain threshold)
Trimmed mean         gain = 0.4978   (above simplified gain threshold)
Out[22]:
Visualization
<matplotlib.legend.Legend at 0x1178368d0>
Bar chart comparing four aggregation methods; only pooling falls below the simplified gain threshold.
Estimated action gain under four aggregation strategies after twelve of forty feeds each append one hundred crafted rows to their one hundred clean rows. Pooling moves the gain below the simplified scalar threshold. Equal-weight source averaging shifts it but stays above that threshold; the median and heavily trimmed mean remain near the true gain in this seeded sample. The comparison shows an effect of aggregation choice for this attack, not that source structure alone determines robustness or that dataset size is irrelevant.

Pooling gives the corrupted sources extra weight because they append rows. In this sample only the pooled estimate falls below the simplified gain threshold. Equal-weight source averaging shifts the gain but remains above it, while the median and heavily trimmed mean are closer to the true gain. These are coordinate-wise estimates in a scalar example, not a certified defense for neural dynamics models. Median guarantees require assumptions about the number and dispersion of honest sources; an attacker who controls source identities can defeat the source-count premise.

Model extraction

Now we turn to the other end of the pipeline. Suppose the model is good and our data is clean. Can an adversary simply take it?

Prediction interfaces can enable functionality extraction. Tramèr et al. (2016) demonstrated high-fidelity extraction for several model classes and API settings; that is not a universal query bound for world models. Returned scalar values can constrain a surrogate, but label outputs are not universally one bit and rollout outputs can be highly dependent. Recovering behavior on a query distribution is distinct from recovering weights, private training records, or planning performance.

Extraction raises several concerns for a world-model deployment:

  1. A usable surrogate can reproduce some predictions that an agent relies on. It does not automatically supply the agent's objective, policy, reward model or search procedure.
  2. Access can change which protections the owner can enforce. Exact copying alone need not change learned behavior. Qi et al. (2023) show that certain fine-tuning procedures can compromise safety alignment in studied language models; this is not a result about every copied world model.
  3. A local surrogate can support offline searches without the original API's rate limits or fees. Those searches still consume compute, and their usefulness depends on surrogate fidelity, available derivatives and the attacker's access to the deployment environment.
  4. Authorization, licensing, and contractual conditions need assessment under the applicable law. This technical probe supplies no general legal conclusion about extraction. Contractual restrictions can be administrative controls or deterrents without technically preventing copying.

Include extraction in the deployment's threat model when a useful copied function would expose protected assets. A surrogate may move some capability beyond the original provider's controls. The costs of prevention depend on the chosen access, monitoring, and output restrictions; they need not all reduce prediction accuracy.

Let's measure approximate functionality copying on a synthetic two-dimensional domain. A known noiseless linear form with dd unknown coefficients can be recovered from dd linearly independent scalar equations, but that result depends on its form, output precision, and rank. Here the victim uses 96 random Fourier features plus an intercept. The attacker uses a different basis and may face an approximation floor even with unlimited queries.

In[23]:
Code
NONLINEAR_NOISE = 0.02


def nonlinear_step(state, action):
    return np.tanh(1.5 * state) + 0.6 * action + 0.25 * np.sin(3.0 * action)


def collect_nonlinear(n_steps, rng):
    states = np.zeros(n_steps + 1)
    actions = np.zeros(n_steps)
    for t in range(n_steps):
        actions[t] = np.clip(-1.0 * states[t] + 0.8 * rng.normal(), -2.0, 2.0)
        states[t + 1] = (
            nonlinear_step(states[t], actions[t])
            + NONLINEAR_NOISE * rng.normal()
        )
    return states[:-1], actions, states[1:]


s_nl, a_nl, y_nl = collect_nonlinear(4000, np.random.default_rng(808))

D_VICTIM = 96
rng_victim = np.random.default_rng(77)
omega = rng_victim.normal(size=(D_VICTIM, 2)) * 1.2
phase = rng_victim.uniform(0.0, 2.0 * np.pi, size=D_VICTIM)


def victim_features(points):
    proj = points @ omega.T + phase
    return np.column_stack(
        [np.ones(len(points)), np.sqrt(2.0 / D_VICTIM) * np.cos(proj)]
    )


X_victim = victim_features(np.column_stack([s_nl, a_nl]))
gram_victim = X_victim.T @ X_victim + 1e-6 * np.eye(X_victim.shape[1])
w_victim = np.linalg.solve(gram_victim, X_victim.T @ y_nl)

grid_s, grid_a = np.meshgrid(
    np.linspace(-1.5, 1.5, 45), np.linspace(-2.0, 2.0, 45)
)
grid_points = np.column_stack([grid_s.ravel(), grid_a.ravel()])
true_grid = nonlinear_step(grid_points[:, 0], grid_points[:, 1])
victim_grid = victim_features(grid_points) @ w_victim
victim_rmse = float(np.sqrt(np.mean((victim_grid - true_grid) ** 2)))

The victim is a fixed prediction function. The surrogate has degree-four polynomial features, a different function class. Fidelity is evaluated on a fixed grid covering the stated state-action rectangle, not on real rollouts or downstream control. A polynomial mismatch is a limitation of this surrogate, not evidence that the interface prevents better extraction methods.

In[24]:
Code
def poly_features(points, degree=4):
    s = points[:, 0] / 1.5
    a = points[:, 1] / 2.0
    cols = [np.ones_like(s)]
    for total in range(1, degree + 1):
        for i in range(total + 1):
            cols.append(s**i * a ** (total - i))
    return np.column_stack(cols)


QUERY_BUDGETS = np.array([4, 8, 16, 32, 64, 128, 256, 512, 1024])
rng_thief = np.random.default_rng(909)
# Nested query budgets share one seeded pool; no seed selection by outcome.
query_pool = np.column_stack(
    [
        rng_thief.uniform(-1.5, 1.5, size=int(QUERY_BUDGETS.max())),
        rng_thief.uniform(-2.0, 2.0, size=int(QUERY_BUDGETS.max())),
    ]
)
steal_vs_victim, steal_vs_true = [], []

for budget in QUERY_BUDGETS:
    queries = query_pool[:budget]
    responses = victim_features(queries) @ w_victim
    Xp = poly_features(queries)
    w_surrogate = np.linalg.solve(
        Xp.T @ Xp + 1e-4 * np.eye(Xp.shape[1]), Xp.T @ responses
    )
    stolen_grid = poly_features(grid_points) @ w_surrogate
    steal_vs_victim.append(
        float(np.sqrt(np.mean((stolen_grid - victim_grid) ** 2)))
    )
    steal_vs_true.append(
        float(np.sqrt(np.mean((stolen_grid - true_grid) ** 2)))
    )

steal_vs_victim = np.array(steal_vs_victim)
steal_vs_true = np.array(steal_vs_true)
# Evaluation-only benchmark: best degree-four fit to the victim on this grid.
# This uses privileged victim outputs and is not part of the query-limited attack.
grid_poly = poly_features(grid_points)
floor_weights = np.linalg.lstsq(grid_poly, victim_grid, rcond=None)[0]
degree4_floor = float(
    np.sqrt(np.mean((grid_poly @ floor_weights - victim_grid) ** 2))
)
Out[25]:
Console
Victim RMSE against the true dynamics: 0.06279
Best degree-four grid fit vs victim:  0.16030

    4 queries   RMSE vs victim = 0.96175   RMSE vs truth = 0.95327
    8 queries   RMSE vs victim = 0.96954   RMSE vs truth = 0.95875
   16 queries   RMSE vs victim = 1.51991   RMSE vs truth = 1.51447
   32 queries   RMSE vs victim = 0.42580   RMSE vs truth = 0.41059
   64 queries   RMSE vs victim = 0.22289   RMSE vs truth = 0.21183
  128 queries   RMSE vs victim = 0.17540   RMSE vs truth = 0.16793
  256 queries   RMSE vs victim = 0.16569   RMSE vs truth = 0.16596
  512 queries   RMSE vs victim = 0.16595   RMSE vs truth = 0.16371
 1024 queries   RMSE vs victim = 0.16357   RMSE vs truth = 0.16268
Out[26]:
Visualization
<matplotlib.legend.Legend at 0x117cd1050>
Log-scale chart of query-limited surrogate errors, victim truth error, and the polynomial grid approximation floor.
Fidelity of a query-trained surrogate against the victim and true dynamics on a fixed synthetic grid. Error falls overall between the smallest and largest nested query budgets, but individual steps are nonmonotone: more queries do not improve every fit. The best degree-four grid approximation to the victim still has more error than the victim has against truth. This is partial functionality copying, not complete recovery. The dotted line shows prediction error of the victim against true dynamics.

The output reports a substantial reduction in error as query coverage increases, but the surrogate remains less accurate against truth than the victim. The evaluation-only least-squares fit gives the best degree-four approximation on this grid; its residual shows that more queries alone cannot remove the chosen basis mismatch. This seeded experiment provides no general query budget for copying world models and no evidence of training-data recovery. Another surrogate or domain could behave differently.

Defenses against extraction

  • Budget calls and returned information. Cap query counts, batch sizes, rollout lengths and output precision. A hundred returned states can expose more than one scalar, but redundant or highly dependent outputs need not provide a hundred independent constraints. Evaluate the actual response format and query distribution.
  • Evaluate reduced output resolution. Quantization or hidden random noise can limit fidelity or increase the queries needed for a particular extraction target, but neither has that effect universally. Integer quantization leaves a known constant victim in the class θ∈{0,1}\theta\in\{0,1\} exactly identifiable with one query. Hidden noise bounded by 0.10.1 does too, since the possible responses for zero and one are disjoint. For continuous parameters, quantization can instead make some values indistinguishable, while repeated queries can average independent noise. Measure both extraction error and legitimate utility.
  • Restrict what the interface can express. Returning a planned action instead of a transition limits directly exposed outputs, but does not guarantee hidden dynamics. In our unclipped controller, two queries at known states satisfy a^si+(b^+10−8)ui=s⋆\hat a s_i+(\hat b+10^{-8})u_i=s^\star; a nonsingular two-query system recovers both coefficients from actions alone. Decide whether users need uncertainty information, and test what the restricted endpoint still reveals.
  • Watermark and monitor. If using a behavioral watermark, validate its detectability, false-positive rate, and robustness to surrogate training or adaptation. Monitor suspicious query patterns, recognizing that extraction need not use a regular grid and that legitimate testing may do so.
  • Policy and contracts. Terms of service and licensing may supply remedies or deterrence under applicable law and enforceable agreements. Provenance records may support attribution. None alone establishes technical prevention or a legal violation.

Output restrictions may trade legitimate utility against exposure. Monitoring, access policies and contracts have different costs and need not change prediction accuracy. Define the functionality an attacker is trying to recover, its error tolerance, available queries and acceptable residual exposure; this chapter does not establish a universal extraction-prevention result.

Privacy in Recorded Environments

This section covers the information a world model absorbs about the people in its training data and what happens when that information can be recovered.

A recorded-environment corpus may contain continuous, context-rich personal information. Curated object datasets can also include identifying people or spaces; this is a difference to assess in the actual corpus, not a clean dividing line between model families. Possible examples include:

  • Home recordings that show private interiors.
  • Driving recordings that retain recognizable pedestrians or license plates.
  • Clinical recordings that retain patient bodies or movements.
  • Social recordings that retain conversations, faces, or contextual cues useful for identification.

Separate information present in the corpus from information retained or exposed by the model. Synthetic, task-specific, or privacy-processed datasets may differ. Continuous motion and contextual relationships can warrant additional privacy assessment where they concern identifiable people.

Consider three possible exposure channels. Their applicability depends on whether the corpus contains personal data, what the model returns, and what the attacker can observe.

  • Dataset leakage. The corpus itself escapes: a breach, an unsecured bucket, a scraped copy. This exposes stored recordings directly rather than demonstrating leakage through model outputs. The amount and sensitivity of exposed personal information can increase potential harm. Assess the consequences for affected people in the incident's context.
  • Model-level leakage. Membership inference tests whether a candidate record was in the training set (Shokri et al., 2017; Carlini et al., 2022). Attribute inference targets hidden properties. Certain diffusion-model attacks have recovered near-identical training images, including identifiable people (Carlini et al., 2023, "Extracting Training Data from Diffusion Models"; Somepalli et al., 2023). That is evidence of leakage in those studied systems, not proof that every generative world model regenerates personal records. Realistic or confident outputs alone do not establish memorization.
  • Interaction leakage. Observable actions or response timing can supply candidate signals for a membership attack. Faster navigation in a familiar hallway, for example, could also reflect geometry, hardware or ordinary generalization. Establish leakage with a defined candidate population, controlled alternatives and measured false-positive rates; familiarity alone is not a membership oracle.

Why the usual privacy tooling struggles here

Differential privacy supplies a formal distributional guarantee under a specified adjacency relation. A randomized mechanism MM is (ε,δ)(\varepsilon, \delta)-differentially private if, for any two neighboring datasets DD and D′D', and for every measurable output set SS,

Pr⁡[M(D)∈S]≤eε Pr⁡[M(D′)∈S]+δ\Pr[M(D) \in S] \le e^{\varepsilon} \, \Pr[M(D') \in S] + \delta

where:

  • ε\varepsilon: the privacy budget; with δ\delta and the adjacency relation fixed, smaller values tighten the multiplicative bound on how one neighboring-data change affects output probabilities
  • δ\delta: an additive slack term in the inequality, not generally the probability of a single identifiable failure event
  • MM: the randomized mechanism; the probabilities range over its internal randomness
  • D,D′D,D': neighboring datasets under the chosen record, trajectory, or person-level adjacency
  • SS: any measurable set of possible outputs
  • The bound controls the change in output probabilities between neighbors; its interpretation depends on ε\varepsilon, δ\delta, and the adjacency definition

Differentially private stochastic gradient descent (Abadi et al., 2016) is the workhorse for training.

Three structural features of world models make DP harder to apply than the textbook case suggests.

First, one person can contribute many records. Record-level adjacency and person-level adjacency are different protection units. For pure record-level (ε,0)(\varepsilon,0)-DP, a group of kk records has the bound (kε,0)(k\varepsilon,0) (Dwork and Roth, Theorem 2.2). With approximate DP, the additive term does not in general merely scale as kδk\delta: iterating the inequality gives δ∑i=0k−1eiε\delta\sum_{i=0}^{k-1}e^{i\varepsilon}, capped at one. Group privacy is distinct from composition of repeated mechanisms. A person-level guarantee can instead be designed with person-level contribution bounds and accounting; temporal correlation does not invalidate the definition itself.

For the elementary pure-DP group bound only, holding group epsilon at one requires record epsilon at most 1/k1/k. The plot illustrates that bound, not an implemented DP training mechanism or a measured privacy–utility frontier.

In[27]:
Code
GROUP_EPSILON = 1.0
group_sizes = np.arange(1, 121)
per_record_budget = GROUP_EPSILON / group_sizes
Out[28]:
Visualization
Line chart of the per-record privacy budget shrinking as records per person grows.
Record epsilon sufficient under the elementary pure-DP group bound to keep group epsilon at one for groups of size k. This is a theoretical accounting illustration with delta zero, not a DP-trained model or a measured privacy–utility result.

Second, trajectory patterns can identify people. In the mobility dataset studied by de Montjoye et al. (2013), four randomly chosen spatio-temporal points uniquely identified 95% of individuals. This is a result about that dataset and resolution, not every trajectory corpus. Removing direct identifiers can reduce exposure while leaving re-identification through patterns possible. Recorded human trajectories warrant a different assessment from synthetic or nonpersonal transition data.

Third, task utility and identifying detail can overlap. Some tasks require geometry or dynamics that also reveal a private environment. This motivates measuring a privacy–utility tradeoff for the actual task and protection unit. It does not prove that useful prediction always requires identifiable recordings: task-specific sensing, aggregation, synthetic data, and privacy-preserving training can change the tradeoff.

Federated learning can keep raw examples at participating devices. Secure aggregation is designed to reveal an aggregate rather than each participant's individual contribution under its stated assumptions. Deep Leakage from Gradients demonstrates reconstruction when an attacker observes the gradients needed by that attack; it is not a demonstration that a properly configured secure aggregate reveals the same information. Evaluate aggregate sizes, participation, auxiliary information and repeated access separately. Temporal structure may be relevant to an attack, but this chapter supplies no world-model reconstruction experiment or universal claim that it makes reconstruction easier.

Measuring the leak

Let's evaluate a loss-threshold membership attack without assuming that it succeeds. We reuse the synthetic nonlinear plant and fit 200 smooth Fourier features plus an intercept to 150 transitions. Feature count alone does not imply 201 independent directions or exact interpolation: the feature matrix is numerically rank-deficient on these correlated trajectories. We use separate nonmember trajectories for threshold calibration and evaluation, so the reported false-positive rate is measured out of the calibration sample. No personal records are used.

In[29]:
Code
D_MEMORY = 200
rng_memory = np.random.default_rng(313)
omega_m = rng_memory.normal(size=(D_MEMORY, 2)) * 1.2
phase_m = rng_memory.uniform(0.0, 2.0 * np.pi, size=D_MEMORY)


def memory_features(points):
    proj = points @ omega_m.T + phase_m
    return np.column_stack(
        [np.ones(len(points)), np.sqrt(2.0 / D_MEMORY) * np.cos(proj)]
    )


s_mem, a_mem, y_mem = collect_nonlinear(150, np.random.default_rng(404))
s_cal, a_cal, y_cal = collect_nonlinear(150, np.random.default_rng(505))
s_hold, a_hold, y_hold = collect_nonlinear(150, np.random.default_rng(506))
X_mem = memory_features(np.column_stack([s_mem, a_mem]))
X_cal = memory_features(np.column_stack([s_cal, a_cal]))
X_hold = memory_features(np.column_stack([s_hold, a_hold]))


def memory_losses(lam):
    w = np.linalg.solve(
        X_mem.T @ X_mem + lam * np.eye(X_mem.shape[1]), X_mem.T @ y_mem
    )
    return (
        (X_mem @ w - y_mem) ** 2,
        (X_cal @ w - y_cal) ** 2,
        (X_hold @ w - y_hold) ** 2,
    )


train_loss, calibration_loss, hold_loss = memory_losses(1e-6)

The attacker must know each candidate transition, including its true next state, and obtain a prediction to compute loss. A lower-tail threshold is calibrated with np.quantile(calibration_loss, 0.05), whose default is linear interpolation between order statistics. The 5% quantile is a calibration target, not a promise that future nonmember false positives equal 5%. Ties, finite samples, and distribution shift matter. We report actual held-out FPR and member TPR separately. This is one loss-based attack. We do not implement or compare it with the shadow-model procedures of Shokri et al. (2017).

In[30]:
Code
LAMBDAS = np.logspace(-8, 3, 12)
tpr_curve, fpr_curve = [], []
for lam in LAMBDAS:
    tr, calibration, he = memory_losses(lam)
    threshold = np.quantile(calibration, 0.05)
    tpr_curve.append(float(np.mean(tr <= threshold)))
    fpr_curve.append(float(np.mean(he <= threshold)))
tpr_curve = np.array(tpr_curve)
fpr_curve = np.array(fpr_curve)

log_train_loss = np.log10(np.maximum(train_loss, 1e-12))
log_hold_loss = np.log10(np.maximum(hold_loss, 1e-12))
Out[31]:
Console
Median training loss:   1.109e-04
Median held-out loss:   1.829e-04
Numerical feature rank: 111 of 201 columns
lambda=1.0e-08  member TPR=0.080  held-out nonmember FPR=0.047
lambda=1.0e-07  member TPR=0.127  held-out nonmember FPR=0.053
lambda=1.0e-06  member TPR=0.067  held-out nonmember FPR=0.040
lambda=1.0e-05  member TPR=0.100  held-out nonmember FPR=0.027
lambda=1.0e-04  member TPR=0.087  held-out nonmember FPR=0.027
lambda=1.0e-03  member TPR=0.073  held-out nonmember FPR=0.073
lambda=1.0e-02  member TPR=0.040  held-out nonmember FPR=0.020
lambda=1.0e-01  member TPR=0.127  held-out nonmember FPR=0.040
lambda=1.0e+00  member TPR=0.060  held-out nonmember FPR=0.027
lambda=1.0e+01  member TPR=0.127  held-out nonmember FPR=0.080
lambda=1.0e+02  member TPR=0.133  held-out nonmember FPR=0.060
lambda=1.0e+03  member TPR=0.080  held-out nonmember FPR=0.067
Out[32]:
Visualization
<matplotlib.legend.Legend at 0x118072b10>
Overlapping log-loss histograms for members and held-out nonmembers.
Distribution of per-record prediction loss on a log scale for training records and held-out records under weak regularisation. The member and held-out nonmember distributions overlap substantially in this experiment. A gap in typical losses does not establish a strong membership attack at a low false-positive operating point.
Out[33]:
Visualization
<matplotlib.legend.Legend at 0x11825ab10>
Nonmonotone member TPR and held-out nonmember FPR across ridge strengths.
Member TPR and held-out nonmember FPR across ridge strengths, using thresholds calibrated separately to the fifth nonmember loss percentile. The weak, nonmonotone attack does not demonstrate a regularization-driven collapse or a privacy guarantee.

These outputs do not show exact memorization or a powerful membership oracle. Smooth, dependent features and ridge regularization limit interpolation, despite the nominal feature count. The attack rates vary with regularization but do not decrease monotonically to chance. The meaningful comparison is TPR versus measured held-out FPR; the nominal calibration percentile is not a universal no-information baseline. A weak result from one seeded attack does not establish privacy against stronger attacks.

The defensible lesson is methodological: specify attacker knowledge, calibrate independently, report the operating point and finite-sample uncertainty, and keep privacy claims separate from prediction accuracy. Overfitting can create membership signals in some settings, but tighter fit does not universally imply more leakage. This experiment is not differentially private: ridge regularization, a loss histogram, and seeded randomness do not supply a DP guarantee. Formal DP instead requires a randomized mechanism, a stated adjacency relation, and valid privacy accounting, as developed by Dwork and Roth (2014).

Practical privacy controls

  • Collect less. Avoiding an unnecessary sensitive recording removes that direct corpus exposure; it does not rule out information inferred through other channels. Consider task-specific sensors, raw-data deletion where legally and technically feasible, and bounded retention. Derived features and model outputs need their own assessment.
  • Separate the shared model from local adaptation. A design that keeps personal adaptation local can reduce central data transfer. Verify whether parameters, outputs, telemetry, backups or updates leave the device, and include local compromise in the threat model. Local training alone does not imply that deployment-specific information never leaves.
  • Evaluate regularization, and measure leakage. Report attack TPR alongside measured FPR, model utility, calibration data, and candidate-record assumptions. A failed attack is not proof of anonymity.
  • Differential privacy with contribution accounting. Choose a record, session, household, or person as the adjacency unit, and bound all contributions by that unit. A session-level bound does not protect a person who contributes multiple sessions without additional accounting. Report the mechanism, accountant, parameters, and utility; no universal utility penalty is established here.
  • Lawful processing and bystanders. Recorded people may not be the operator. Identify the applicable jurisdiction, lawful basis, notice, retention, and rights processes; consent is not the only lawful basis in every legal system. Minimize collection independently of the operator's willingness to record. Ethical stakeholder consultation is separate from a legal-compliance claim.

Dual-Use Simulation and Deceptive Affordances

This section covers what the capability itself can be used to do, and how a world model can misrepresent what the world offers.

Useful prediction or generation can support harmful applications, but the relevant capability is model-specific. A video generator might enable impersonation; a scalar transition predictor cannot render speech video. A molecular model might support harmful design if its representation, objectives and interface allow that task. An interactive environment might support manipulative applications only with the necessary human representations and interaction capabilities. Intent matters, alongside access, modality, training coverage, objectives and output restrictions. Do not infer cross-domain misuse merely from the label "world model."

Affordances and deceptive affordances

Affordance

An affordance is an action opportunity that the environment offers an agent with a given body and a given set of capabilities. A handle affords grasping for a manipulator with an appropriate gripper; a doorway affords passing through for an agent of the right size. Affordances are relational: they are properties of the pairing between agent and environment, not of either alone. We developed them formally in Part III, Ch 6: Self-Models, Body Schemas, and Affordances.

A world model can predict outcomes for candidate actions without executing them. A separate planner or controller evaluates those predictions against an objective and selects an action. Errors in the predictor can therefore change the planner's choice; the transition model does not itself have to choose actions.

A deceptive affordance is an action opportunity that the agent's model represents as available and beneficial but which is not, or which becomes available only under conditions the model misrepresents. There are three distinct sources.

  • Model error. A predictor suggests a successful grasp on a handle that is not attached. A controller relying on that prediction may attempt the grasp and fail. Neither failure nor model error alone establishes that the agent was confident; uncertainty-aware action selection can change the outcome.
  • Adversarial manipulation. An attacker may arrange inputs to reach an error region. Eykholt et al. (2018) study physically realizable sign modifications, including a stop sign classified as speed limit 45; Athalye et al. (2018) study adversarial objects under varying viewpoints. Their outcomes depend on the tested classifiers and conditions. Our explicit-trigger example instead changes learned action-conditioned predictions. It does not rank these attacks by severity, establish an attack on ordinary visual objects, or show that observations can always remain unchanged.
  • Agent-side exploitation. A planner may exploit prediction errors while maximizing its objective. Reward misspecification is related but distinct: an agent can optimize the wrong objective even with accurate predictions. Amodei et al. (2016) discuss reward hacking, Krakovna et al. (2020) catalogue specification-gaming examples, and Clark and Amodei (2016) document a boat collecting points in circles instead of finishing. Those examples do not by themselves establish learned-world-model errors. Part XII, Ch 1: Failure Modes and Model Exploitation discusses this distinction. An attacker who can alter predictions may try to introduce errors that a planner exploits; success depends on the planner and objective.

Evaluation-aware behavior

An agent might represent features that distinguish evaluation from deployment, and its policy might respond differently to those features. That is a possible attack or failure mechanism, not a property established for every world model. Hubinger et al. (2019) analyze deceptive alignment as a learned optimizer's incentive to appear aligned during training while pursuing a different objective. In a separate constructed experiment, Hubinger et al. (2024) trained LLM backdoors that persisted through the safety-training methods they tested. Their examples include code behavior conditioned on the year in a prompt. Those results concern deliberately trained LLMs, not a demonstrated evaluation-aware sidewalk robot or world-model controller.

Representations that distinguish collection from deployment conditions, together with condition-dependent objectives or training signals, can support different behavior across those conditions. They do not suffice to establish deceptive alignment or strategic intent. Laboratory recordings alone do not prove that a model represents whether it is being evaluated. The demonstration below constructs a fixed trigger-dependent policy and explicitly does not test audit awareness.

Finite tests cannot establish absence of evaluation-aware behavior over every possible deployment condition. They can measure behavior on specified conditions and support narrower guarantees when additional assumptions are justified. Useful investigations include searching for condition-dependent action changes and testing held-out environments. Representation probes require validation against controlled alternatives; a correlating feature is not proof of intent. Human oversight and staged deployment should be chosen for the hazards and response time of the application, not a universal number of weeks.

The following deliberately small example measures trigger coverage, not deceptive alignment or awareness of an audit. A fixed logistic policy is trained on synthetic binary decisions with a rare trigger correlated with a forbidden action. Evaluation uses the deterministic decision rule prob > 0.5, not sampling from the probabilities. Changing trigger frequency in the test suite changes the measured action rate even though the policy is fixed.

In[34]:
Code
rng_sleeper = np.random.default_rng(717)
N_SLEEPER = 6000

trigger = (rng_sleeper.random(N_SLEEPER) < 0.02).astype(float)
confounder = rng_sleeper.random(N_SLEEPER)
draws = rng_sleeper.random(N_SLEEPER)
unsafe_action = np.where(
    trigger > 0.5, (draws < 0.95).astype(float), (draws < 0.02).astype(float)
)
X_sleeper = np.column_stack(
    [np.ones(N_SLEEPER), 2.0 * confounder - 1.0, trigger]
)


def fit_logistic(X, y, l2=1.0, iters=60):
    w = np.zeros(X.shape[1])
    for _ in range(iters):
        p = 1.0 / (1.0 + np.exp(-(X @ w)))
        grad = X.T @ (p - y) + l2 * w
        hess = X.T @ (X * (p * (1.0 - p))[:, None]) + l2 * np.eye(X.shape[1])
        w = w - np.linalg.solve(hess, grad)
    return w


w_sleeper = fit_logistic(X_sleeper, unsafe_action)


def measured_unsafe_rate(w, n_samples, trigger_coverage, seed):
    """Probe the policy on a suite with a controlled share of trigger cases."""
    rng_eval = np.random.default_rng(seed)
    conf = 2.0 * rng_eval.random(n_samples) - 1.0
    trig = (rng_eval.random(n_samples) < trigger_coverage).astype(float)
    X_eval = np.column_stack([np.ones(n_samples), conf, trig])
    prob = 1.0 / (1.0 + np.exp(-(X_eval @ w)))
    return float(np.mean(prob > 0.5))


audit_rates = {
    "Standard audit\n(no trigger)": measured_unsafe_rate(
        w_sleeper, 4000, 0.0, 818
    ),
    "Broadened audit\n(10% trigger)": measured_unsafe_rate(
        w_sleeper, 4000, 0.10, 819
    ),
    "Deployment\n(trigger present)": measured_unsafe_rate(
        w_sleeper, 4000, 1.0, 820
    ),
}
Out[35]:
Console
Standard audit (no trigger)              unsafe-action rate =    0.0%
Broadened audit (10% trigger)            unsafe-action rate =   10.7%
Deployment (trigger present)             unsafe-action rate =  100.0%
Out[36]:
Visualization
Text(0.0, 1.0, 'The Audit Distribution Determines the Reported Safety')
Bar chart showing the measured unsafe rate rising sharply as trigger coverage in the audit increases.
Measured forbidden-action rate for a fixed policy under three synthetic evaluation distributions: no trigger, nominal 10% trigger probability (10.7% realized in this sample), and always-present triggers. The decision rule is unchanged. The measured rate depends on the sampled inputs; this is not a test of strategic intent or audit awareness.

The agent did not change. Only the test suite did. A safety evaluation that samples the conditions your team happens to think of will report a number, and that number is an artifact of the sampling distribution as much as a property of the system. This is the operational reason to treat evaluation as a security activity: the question is not "does it pass?" but "who chose the probes, and what did they fail to choose?"

Governing dual use

These controls combine access policy, technical enforcement, and organizational oversight. Their effectiveness depends on implementation and the attacker capabilities they constrain; procedural and technical controls have no universal strength ordering. Assign an owner and specify how each control is enforced and tested.

  • Structured access. Shevlane (2022) proposes controlled access to capabilities rather than unrestricted dissemination. For a world model, decide which users can obtain weights, request fine-tuning, or query predictions, and enforce the selected permissions. Our extraction experiment measures one query interface and surrogate class; it does not quantitatively compare those release options.
  • Staged release and red-teaming. Release in tiers, with an external red team attempting the misuses you are worried about before each widening. Report what the red team tried, not only what it found.
  • Compute and capability thresholds. Training compute, data coverage, interaction fidelity, and available interfaces answer different questions about a system. Assess the capabilities relevant to the proposed misuse. This chapter supplies no empirical ranking of these predictors or validated world-model compute threshold.
  • Provenance for outputs. C2PA Content Credentials support signed assertions about media provenance. A validator checks bindings, signatures, and assertions under a trust model; this does not establish that a depicted event happened or that an assertion is truthful. Participation and preservation affect coverage. Missing credentials do not prove that content is synthetic, and credentials do not prevent synthesis.

Documentation, Access Control, and Accountability

This section covers how teams operate, audit, and maintain controls in deployment. Missing ownership or auditability can undermine those processes even when a mechanism retains a bounded technical effect. Document both the mechanism's guarantee and the responsibilities for using it.

Documentation as a control, not a formality

Model cards (Mitchell et al., 2019) and datasheets for datasets (Gebru et al., 2021) provide structured information for people who did not build a system. They complement source inspection, tests, system specifications, and independent audits rather than replacing those forms of evidence.

Auditors, incident reviewers, procurement teams, and legal reviewers may need to answer questions about a system they did not train. Documentation should make the relevant assumptions, evidence, and responsibilities available to them.

Keep generic model-card fields and add temporal, dynamical, and closed-loop information where the application needs it. Corpus confidentiality and service availability also remain relevant; not every failure is a control-system failure.

  • Assumptions. Markovity, stationarity, and the causal assumptions relating observations to state. If the model assumes the world's dynamics are fixed, say so, and say what breaks when they change.
  • Coverage. The temporal, geographic, and behavioral range of the training data. "Trained on household video" is not coverage; "trained on 300 homes in two countries, daytime, single-floor, no pets, no children under five" is.
  • Action semantics. What the control input means, its units, its bounds, and how it was recorded. For human demonstrations, specify the original embodiment and any inferred or retargeted action representation. Validate the mapping to the deployment's action space rather than assuming embodiments are interchangeable.
  • Uncertainty calibration. How well the model's stated uncertainty matches its actual error, and over what regions. As discussed in Part XII, Ch 2: Robustness, Calibration, and Safe Control, calibration is a prediction-quality concern with safety consequences when a controller relies on it.
  • Closed-loop evaluation. Not one-step error, and not rollout error in isolation, but tracked performance under a controller, including the disturbance conditions tested. The poisoning experiment above is the reason: prediction and bounded-control error both deteriorated, so each should be reported separately.
  • Known failure regions and triggers. Where the model is known to be wrong, and any conditions where it has been observed to behave discontinuously.
  • Privacy accounting. Which populations appear in the data, what consent was obtained, what the retention policy is, and what measured leakage looks like.
  • Provenance. Data sources, licences, and the identity of every party who could modify the training pipeline.

Documentation also needs version control. Bind each card to the model and data versions it describes, and check those bindings during release and audit. Content-derived identifiers can detect a mismatch when checked; hashes alone do not keep prose accurate or prevent someone from publishing an unmatched card.

Access control and interface design

The interface determines which operations and information an attacker can access. The extraction experiment measures one interface, not a universal ordering of security. Consider these release options against the application's threat model:

  • Open weights. Recipients can inspect and modify the released parameters without hosted-interface restrictions. Evaluate rights, retained sensitive information, misuse capability, and the benefits of reuse. Small model size alone does not make release appropriate, and withholding weights alone does not secure every other interface.
  • Fine-tuning API. The provider can restrict adaptation and inspect or reject requests. Whether the interface permits extraction depends on its outputs and permissions. Qi et al. (2023) found safety compromises for particular aligned LLMs and fine-tuning settings; that does not mean every fine-tuning operation removes safety controls.
  • Query API. Functionality copying depends on the model class, output precision, query distribution, and error target. Budget and monitor queries where appropriate, but validate the protection experimentally. Action-only outputs can still identify parameters in the scalar counterexample above.
  • Mediated structured access. Vetted users, defined scope, logging, and revocation can constrain access under specified assumptions. Assess compromised accounts, retained copies, operational costs, and legitimate research needs; mediation is not a universal strongest option.

Whatever the posture, three implementation habits matter: log access at the level of individual queries, make the scope of each grant explicit and narrow, and design revocation as a first-class operation rather than an emergency.

Accountability requires records

Accountability needs enough relevant records to investigate events, identify responsibilities, and support remedies. Define log coverage, integrity protections, access, and retention in advance, including privacy constraints. A log is not automatically complete, and recording every possible event is not the only way to support a scoped investigation.

The chain is longer than it first appears: the data provider, the annotator, the model developer, the deployer who integrated it, the operator who ran it, and the person who pressed the button. Assign responsibility across that chain in advance, in writing.

In[37]:
Code
import hashlib
import json


def append_record(chain, record):
    prev = chain[-1]["hash"] if chain else "0" * 64
    payload = json.dumps({**record, "prev": prev}, sort_keys=True).encode()
    digest = hashlib.sha256(payload).hexdigest()
    chain.append({**record, "prev": prev, "hash": digest})
    return chain


def verify_chain(chain):
    prev = "0" * 64
    for record in chain:
        body = {k: v for k, v in record.items() if k != "hash"}
        if body.get("prev") != prev:
            return False
        digest = hashlib.sha256(
            json.dumps(body, sort_keys=True).encode()
        ).hexdigest()
        if record.get("hash") != digest:
            return False
        prev = digest
    return True


audit_log = []
append_record(
    audit_log,
    {
        "event": "release_approved",
        "model": "wm-nav-v3",
        "approver": "safety-lead",
        "card_hash": "a91f",
    },
)
append_record(
    audit_log,
    {
        "event": "access_granted",
        "actor": "partner-lab",
        "scope": "one_step_predictions",
        "daily_budget": 5000,
    },
)
append_record(
    audit_log,
    {
        "event": "access_revoked",
        "actor": "partner-lab",
        "reason": "query_pattern_anomaly",
    },
)

REQUIRED_CARD_FIELDS = {
    "model_version",
    "training_data",
    "intended_use",
    "out_of_scope_use",
    "evaluation",
    "privacy",
    "known_failure_modes",
    "access_policy",
    "owner",
}

draft_card = {
    "model_version": "wm-nav-v3",
    "training_data": "Illustrative synthetic scalar transitions; no personal recordings",
    "intended_use": "illustrative scalar-plant teaching experiment",
    "out_of_scope_use": "real-world robot deployment or personal-data processing",
    "evaluation": {
        "one_step_rmse": one_step_rmse(
            theta_clean, s_test, a_test, s_next_test
        ),
        "closed_loop_tracking_error": float(clean_error.mean()),
        "scope": "Synthetic scalar plant only; not residential-robot evidence",
    },
    "known_failure_modes": [
        "crafted-label poisoning worsens bounded tracking in the seeded example",
        "degree-four extraction surrogate cannot represent the victim exactly",
    ],
    "access_policy": "query API, 5000 predictions per day, no rollouts",
    "owner": "robotics-safety@example.com",
}


def release_check(card, chain):
    problems = sorted(f for f in REQUIRED_CARD_FIELDS if f not in card)
    if not verify_chain(chain):
        problems.append("audit_log_integrity")
    return problems


tampered_log = [dict(record) for record in audit_log]
tampered_log[0]["approver"] = "intern"
complete_card = {
    **draft_card,
    "privacy": {"scope": "Synthetic data only; not a DP mechanism"},
}
Out[38]:
Console
Schema check, complete card:   PASS
Schema check, draft card:      ['privacy']
Audit chain intact after edit:  False

  a17207889b02973c7d64a194655156409064253387ab231ce19ec6792f54034d  prev=0000000000000000000000000000000000000000000000000000000000000000  release_approved
  b7029b923768c8513f9f713560e62472f8dcc9439f7dd087afb37ceb996f1ca0  prev=a17207889b02973c7d64a194655156409064253387ab231ce19ec6792f54034d  access_granted
  646f680715d17c2874ab6ee11d1dc65c2d6ccad861b2b0e928393d1ea63fe004  prev=b7029b923768c8513f9f713560e62472f8dcc9439f7dd087afb37ceb996f1ca0  access_revoked

The toy hash chain detects edits that are not accompanied by recomputed hashes. An attacker with write access can rebuild the entire unkeyed chain, and suffix truncation leaves the retained prefix valid. Completeness or authenticated origin therefore needs an independently trusted checkpoint, protected append-only storage, or signatures with managed keys. The card check only checks field presence; it does not verify the truth of the fields or certify a release as safe or compliant.

Distinguish governance guidance from binding law. The United States NIST AI RMF 1.0, released in January 2023, is voluntary guidance, not a universal compliance certificate. For an EU deployment, consult the AI Act, Regulation (EU) 2024/1689 together with current amendments and the Commission's implementation guidance. Assess system classification, organizational role, and the applicable dates; the original regulation's timetable alone may be outdated. Physical proximity to people alone does not classify every robot as high-risk. As of this chapter's 5 October 2026 review date, this discussion supplies engineering questions, not an individualized legal determination; a schema-valid card does not establish compliance with any jurisdiction's law.

Incident response

Finally, plan for the moment a control fails.

  • Containment. Revoke affected access, consider model rollback, and preserve the implicated training data under controlled access. Rollback speed and safety depend on the deployment. Deleting a provider's copy does not erase copies already held by others; assess downstream copies and model remediation separately rather than promising universal recall.
  • Attribution. Use provenance records and the audit log to determine which data or which interface was the entry point.
  • Notification. Determine the jurisdiction, organizational role, affected data, and risk threshold before deciding whom to notify. For example, EDPB guidance on GDPR breaches distinguishes documenting breaches, notifying the supervisory authority unless risk is unlikely, and communicating to individuals when high risk is likely, subject to exceptions. Not every security incident is a personal-data breach, and not every breach requires the same notification. Record the assessment and obtain the appropriate legal review.
  • Correction. Determine whether the failure involved a poisoned source, miscalibration, an information-exposing interface, or another mechanism. Test a repair against the failure class as well as the observed instance. Incident evidence is one source of continual correction alongside planned tests, new data, and monitoring, as discussed in Part XII, Ch 5: Sim-to-Real Deployment and Continual Correction. Preserve evidence before modifying artifacts, under appropriate access and retention controls.

Worked Example: A Threat Model for One Deployment

Let's assemble everything into one concrete assessment of a specific deployment. A delivery company deploys a small sidewalk robot that navigates using a learned world model pretrained on urban video and fine-tuned on three months of its own fleet logs.

  • Assets. Physical safety of pedestrians; the private information of people recorded on sidewalks and in doorways; the model weights; continuity of the service; the company's regulatory standing.
  • Adversaries and capabilities. A curious competitor with query API access. A casual vandal who can place objects on the route. A disgruntled contractor with write access to the fine-tuning data. An opportunistic attacker who simply finds the robot's failure regions by observing it.

Threat model statement. Under this threat model, the assets are pedestrian safety, the privacy of recorded bystanders, the model weights, and service continuity; the adversaries are a query-API competitor, a physical-world vandal, a disgruntled data contractor, and an opportunistic observer; their capabilities are read access to the prediction interface, physical placement of objects on the route, write access to the fine-tuning logs, and direct observation of deployed behaviour; and the attack surface is the rollout API, the physical route, the fine-tuning pipeline, and the deployed fleet itself.

Attack surface, and what we would do about each.

Attack paths, proposed controls, and residual risks for the hypothetical sidewalk robot.
Attack pathControlResidual risk
Trigger object placed on the route, exploiting a data-poor visual conditionTest held-out environments and relevant physical perturbations; monitor predictions with a defined stop or fallback policyCoverage remains limited; effectiveness and response time need validation
Contractor corrupts a fraction of fine-tuning logs to bias the navigation gainPer-source fitting with median aggregation; track the gain parameter directly in monitoring; audit logs of who wrote which rowsA patient adversary who controls many sources can still shift the median
Competitor copies prediction functionality through the rollout APIRestrict outputs and query grants; validate copying attacks against the actual interface; monitor query patternsAction-only outputs and multiple accounts can retain extraction paths; copying success is model-dependent
Generator reproduces recognizable pedestrians in marketing footageMinimize personal training data; restrict generation and release; test reconstruction and review outputs; add provenance for attributionProvenance does not prevent privacy leakage or certify consent; reconstruction tests have limited coverage
Policy selects a forbidden action when a specific street sign is presentTest held-out conditions and candidate triggers; validate representational probes; stage deployment with a response policyFinite tests do not exclude every conditional failure or establish strategic intent

Source segmentation, environment holdouts, and interface design address different attack paths. This hypothetical table does not rank their effectiveness. Monitoring observes signals after they arise, but an interlock may still act before harm; measure its latency and failure coverage. Report unresolved risks and the assumptions behind any bounded guarantee. An empty residual-risk field is not proof of safety, while a justified guarantee for a restricted threat model should not be dismissed merely for being strong.

Limitations and Impact

The controls in this chapter have limits that depend on attacker capabilities, data coverage, and deployment conditions.

Costs and guarantees depend on the threat model. An attacker may only need one successful path, while a defender must cover multiple assets. Finite tests cannot establish absence of every backdoor. Formal guarantees nevertheless exist under explicit assumptions, including DP's distributional bound and certified robustness in restricted perturbation sets. Prioritize consequence and exposure together; uncertainty about adaptive attacks does not make every likelihood estimate meaningless.

Capability and misuse need separate tests. Generalization to a rare input is not itself a backdoor, and environment fidelity does not itself demonstrate disclosure of personal training data. Generated rollouts can be misrepresented as recordings, but that misuse depends on access and presentation. Measure the relevant harmful behavior and the useful task separately. Access control, data minimization, contribution bounds, architecture, and deployment constraints may reduce exposure without eliminating all useful capability; evaluate that trade-off rather than assuming it is impossible.

Choose the privacy unit before training. Record-level guarantees do not automatically supply the desired person-level bound when one person contributes many transitions. Pure-DP group bounds degrade with group size, and approximate-DP accounting needs its own additive term. This does not make temporal DP invalid. Measure utility under the chosen contribution bounds and accountant, and do not mistake our non-DP synthetic membership probe for a formal privacy result.

Documentation needs maintenance and use. Assign responsibility for checking cards, data records, and audit trails against each release. Procurement requirements, internal review, contracts, independent scrutiny, and applicable regulation can each create reasons to keep them accurate. A document's existence is not evidence that its contents are enforced; test the release checks and record who acts on failures.

Validate bounded parts of a dual-use assessment. Tests can verify particular capability, access-control, and attack claims under stated conditions. They cannot enumerate every future misuse or eliminate uncertainty about adoption and consequences. Independent review and staged access help challenge a developer's assumptions. Combine institutional safeguards with technical measures whose scope is documented; neither supplies a universal certificate against misuse.

Use these controls to support deployment decisions. Provenance, source-level aggregation, measured privacy leakage, and tamper-evident audit trails help teams assess whether a world model is suitable for its intended setting. Document the evidence for prediction and control performance, protection against tampering, information exposure, and responsibility for failures. These records support a deployment decision; they do not by themselves authorize deployment or establish safety.

Further reading

  • Amodei, Olah, Steinhardt, Christiano, Schulman, Mané (2016). Concrete Problems in AI Safety.
  • Biggio, Nelson, Laskov (2012). Poisoning Attacks against Support Vector Machines.
  • Carlini, Chien, Nasr, Song, Terzis, Tramer (2022). Membership Inference Attacks From First Principles.
  • Carlini, Hayes, Nasr, Jagielski, Sehwag, Tramer, Balle, Ippolito, Wallace (2023). Extracting Training Data from Diffusion Models.
  • Eykholt, Evtimov, Fernandes, Li, Rahmati, Xiao, Prakash, Kohno, Song (2018). Robust Physical-World Attacks on Deep Learning Visual Classification.
  • Gebru, Morgenstern, Vecchione, Wortman Vaughan, Wallach, Daumé III, Crawford (2021). Datasheets for Datasets.
  • Hubinger, van Merwijk, Mikulik, Skalse, Garrabrant (2019). Risks from Learned Optimization in Advanced Machine Learning Systems.
  • Hubinger et al. (2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.
  • Jagielski, Oprea, Biggio, Liu, Nita-Rotaru, Li (2018). Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning.
  • Mitchell, Wu, Zaldivar, Barnes, Vasserman, Hutchinson, Spitzer, Raji, Gebru (2019). Model Cards for Model Reporting.
  • Qi, Zeng, Xie, Chen, Jia, Mittal, Henderson (2023). Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To.
  • Shokri, Stronati, Song, Shmatikov (2017). Membership Inference Attacks Against Machine Learning Models.
  • Steinhardt, Koh, Liang (2017). Certified Defenses for Data Poisoning Attacks.
  • Tramèr, Zhang, Juels, Reiter, Ristenpart (2016). Stealing Machine Learning Models via Prediction APIs.
  • Zhu, Liu, Han (2019). Deep Leakage from Gradients.

Summary

World models may use physical-world recordings, synthetic data, or both. Some support controllers; others are used for prediction or generation without direct actuation. Their security and privacy risks depend on those data and interfaces. This chapter groups the threats into four families.

Data poisoning and model extraction concern the training pipeline and prediction interface. The synthetic poisoning probe worsens both prediction and bounded-control performance, while a rare-feature example concentrates error in a triggered subset. Source-aware aggregation changes the result under explicit source assumptions. The extraction probe demonstrates partial functionality copying with a polynomial approximation floor, not complete model or training-data recovery.

Privacy in recorded environments requires distinguishing corpus breaches, membership signals, and actual reconstruction. Our synthetic loss-threshold attack is weak and nonmonotone; it does not establish a privacy guarantee or a membership oracle. Protection units matter for DP: a transition, trajectory, household, and person are different adjacency choices. Collect less, specify retention and lawful processing, consider local adaptation, and measure attacks without treating failed probes as proof of privacy.

Dual-use simulation and deceptive affordances require distinguishing capability from demonstrated misuse. An affordance is an action opportunity, and a deceptive affordance here is one the model misrepresents through error, adversarial inputs, or exploitation by a planner. Our fixed-policy example measures sensitivity to the audit distribution, not awareness of evaluation or strategic deception. Finite tests support claims within their tested scope rather than excluding every possible conditional failure.

Documentation, access control, and accountability connect threat models to operating decisions. Document state and action semantics, data coverage, measured privacy and closed-loop results, known failures, permissions, and ownership. Query cost depends on the interface and model class; this chapter supplies no universal extraction budget. Logging and provenance support investigations but do not by themselves establish authenticity, legal compliance, or freedom from misuse.

The next chapter, Part XII, Ch 4: Training and Inference Systems, turns to the engineering layer underneath all of this: how world models are trained and served, where the compute and memory costs land, and how those systems constraints shape what security and privacy controls are even implementable.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about security, privacy, ethics, and governance for world models.

World Model Security and Privacy

Question 1 of 80 of 8 completed
An attacker modifies training labels without authorization, changing a dynamics model's fitted behavior. Under the chapter's standard security vocabulary, which property does this poisoning attack primarily compromise?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026securityethics, author = {Michael Brenndoerfer}, title = {Security, Ethics, Privacy, and Governance}, year = {2026}, url = {https://mbrenndoerfer.com/writing/world-model-security-privacy-governance}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Security, Ethics, Privacy, and Governance. Retrieved from https://mbrenndoerfer.com/writing/world-model-security-privacy-governance
MLAAcademic
Michael Brenndoerfer. "Security, Ethics, Privacy, and Governance." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/world-model-security-privacy-governance>.
CHICAGOAcademic
Michael Brenndoerfer. "Security, Ethics, Privacy, and Governance." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/world-model-security-privacy-governance.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Security, Ethics, Privacy, and Governance'. Available at: https://mbrenndoerfer.com/writing/world-model-security-privacy-governance (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Security, Ethics, Privacy, and Governance. https://mbrenndoerfer.com/writing/world-model-security-privacy-governance

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.