Pretrained Visual Models and DINO-WM

Michael BrenndoerferJuly 20, 202653 min read

Part of World Models Handbook

Explains how DINO-WM freezes a pretrained visual encoder, trains action-conditioned feature dynamics, and uses goal-image planning.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Pretrained Visual Models and DINO-WM

A camera-frame predictor must account for appearance as well as motion. An unweighted pixel loss scores the wall, anti-aliased edges, and sensor noise alongside the object the agent must move. Depending on the architecture, modeling those details can use training or rollout compute without improving the decision. Higher resolution increases the number of predicted values; the fraction occupied by a task-relevant object shrinks only if the field of view or background grows relative to it. A photorealistic frame is therefore not, by itself, evidence of useful control dynamics.

DINOv2 offers a useful starting point: it was trained without labels on the 142-million-image LVD-142M corpus, and its patch features retain spatial and object-part cues in the authors' evaluations. Those cues are not a solved control state. DINO-WM tests a narrower proposition: freeze a pretrained visual encoder, then learn how its features change under recorded actions. This separates reusable visual pretraining from dynamics that still have to be learned for an environment and its action interface.

So DINO-WM splits the problem in two:

  • A frozen pretrained visual encoder turns each observation into a set of feature vectors. Its weights do not change during dynamics training.
  • A trainable predictor uses a history of those features and recorded actions to predict future features. The worked toy later simplifies this to one frame and one action.

The DINO-WM dynamics loss and planner need no image decoder or reward model; the paper trains an optional separate decoder to inspect predictions. Training still requires offline observation–action trajectories, which may include expert behavior, but a new feasible image goal needs no expert demonstration or reward label for that particular goal. Planning searches for actions whose predicted patch features approach the goal image's features. That distance is an explicit, imperfect task cost, not a reward function that appears automatically from pretraining.

This chapter explains when that separation helps, when it fails, and how to build a small version of it. It sits in Part IX: Foundation and World-Action Models after tokenized game and environment models, and prepares the ground for video generators and vision-language-action systems. Those later systems make different architectural choices; they do not all freeze a visual encoder or learn a separate small dynamics model.

The core trade is reuse against task adaptation. A frozen encoder gives the dynamics learner a fixed visual coordinate system and avoids learning that representation again from the environment's action data. It also limits what the predictor can infer from a frame: a distinction absent from the features cannot be recovered from that frame alone, although history or another sensor may help. The encoder is a hypothesis about which visual distinctions matter, and the rest of the pipeline depends on testing it.

We'll build the whole pipeline end to end on a tiny synthetic world so that every piece is inspectable. Along the way we'll examine three distinct failure modes: feature aliasing (task-distinct states look the same to the encoder), dynamics error (the learned predictor misjudges the effect of an action), and proxy misalignment (a low feature-space goal cost does not imply task success). The toy exposes these mechanisms, but it does not establish how often they occur with DINOv2 or on a robot.

Frozen Visual Representations as State

The first design decision in a world model is what information the predictor uses as state. In Part III: Representing Agents and Worlds, predictive sufficiency meant that the future is conditionally independent of earlier history given that state and the actions. A frozen encoder supplies a per-frame embedding, not a guarantee of sufficiency. DINO-WM's transition model also consumes a short history of embeddings and actions; the one-frame model in our toy is an explicit simplification. The frozen visual features are not shaped by a task reward or value signal, but the learned predictor can aggregate their temporal context.

Formally, let oto_t be the observation (an image) and let fθf_\theta be the encoder with parameters θ\theta that are never updated. Define

zt=fθ(ot).z_t = f_\theta(o_t).

For a Vision Transformer with patch size pp, we can retain a grid of patch tokens rather than pool them into one vector. The transformer divides a patch-divisible image into non-overlapping patches, embeds each patch, and lets attention mix information across positions. In the equation below, HH and WW are the dimensions after any padding and are divisible by pp. Excluding any global or register tokens, the patch output has shape

zt∈RN×D,N=Hp⋅Wp.z_t \in \mathbb{R}^{N \times D}, \qquad N = \frac{H}{p}\cdot\frac{W}{p}.

For DINOv2 ViT-S/14 at 224×224224 \times 224 resolution, that is N=256N = 256 patch tokens of dimension D=384D = 384: a 256×384256 \times 384 per-frame feature matrix. Prediction may still need earlier frames and actions.

Patch tokens versus a global embedding

A global embedding (the [CLS] token, or a pooled average over patches) summarizes the image in one vector. Patch tokens retain spatial indices, though attention lets each token contain wider scene context. Comparing corresponding patch indices can localize feature differences; a global embedding can also provide a useful goal cost. Both comparisons ultimately return a scalar to the planner, and neither guarantees a useful action gradient. DINO-WM uses patch-level feature distance for its image-goal cost.

What makes a frozen encoding good enough to be a state

Not every pretrained network produces a usable state space. The properties that matter for world modeling are subtly different from the properties that matter for classification. A classifier only needs to represent whatever distinguishes its label categories, whereas a world model needs to represent everything that changes in a way that affects the future. These are related but not identical demands, and a backbone that excels at one can be mediocre at the other.

  • Informativeness. The encoding must preserve the distinctions that matter for predicting dynamics. If situations with different relevant futures produce the same complete predictor input, including the available feature history and actions, that predictor cannot separate them. A collision in one frame's features alone may be resolved by history or another sensor; freezing only blocks recovery of distinctions absent from every available encoded input. A fixed one-frame feature cost can still alias distinct goals.
  • Smoothness. Small, task-relevant changes should ideally produce predictable feature changes. An erratic local encoding can make dynamics harder to learn, but smoothness alone does not make interpolation in feature space physically realizable.
  • Spatial structure. Patch indices expose image location and can help a predictor exploit locality. Translation equivariance means that shifting an input shifts the output in the same way; a patch grid does not confer that property on a position-embedded Vision Transformer, and real motion need not be a translation.
  • Stability across nuisance factors. Lighting, texture, exposure, and clutter may be irrelevant to one task but important to another; measure whether the feature response suppresses only the variation your task can safely ignore.
  • Cost. A 256×384256 \times 384 patch set contains 98,304 scalar values, versus 150,528 RGB channel values at 224×224224 \times 224. That is only about a 1.53-fold reduction in raw element count, not proof of a storage or compute advantage. The possible benefit is in the representation's usefulness to prediction and planning, which must be measured rather than inferred from dimension alone.

Self-supervised objectives emphasize these properties differently; no universal ordering or necessary trade-off follows from the objectives alone. The choice of pretraining method is therefore not a detail you can ignore; it helps shape the state space your predictor will inhabit.

SimCLR and MoCo contrast augmented views of the same image against other examples. Their objectives encourage agreement under chosen crops and color changes; they do not guarantee exact invariance in the features later used for control. Strong color augmentation could make a red/blue distinction harder to use, for example, if color is the variable a robot needs. The augmentation family is therefore an assumption to probe against the downstream task.

MAE reconstructs masked image patches from visible-patch encodings. Its reconstruction objective allocates training effort to pixels, but that does not establish that its learned features are merely low-level or unsuitable for dynamics. A fair encoder comparison has to hold the predictor and control task fixed.

DINO trains a student to match an exponential-moving-average teacher across augmented crops, with centering and sharpening in its collapse-prevention recipe and no explicit negative pairs. DINOv2 builds on that lineage with an iBOT-style masked-patch objective and KoLeo regularization of class-token feature spread, among other changes. Its paper reports useful spatial and object-part cues, including qualitative correspondence examples; its semantic-segmentation benchmarks train readout heads, so they are not evidence of segmentation by raw nearest neighbors alone. DINO-WM uses DINOv2 patch features, but that does not make DINOv2 the default backbone for every feature-space world model.

The invariance-versus-controllability tension

Here is the tension stated plainly. A frozen encoder was optimized to make some distinctions and ignore others. Those choices were made by whoever pretrained it, for whatever data and objective they had in mind, with no knowledge of your task. The encoder is, in effect, a frozen hypothesis about which differences in the world are worth representing, and you inherit that hypothesis wholesale.

For a world model, it helps if nuisance changes have a small feature effect while controllable changes remain distinguishable and predictable. Invariance means the features are unchanged under a transformation; equivariance means they transform in a corresponding, structured way. Neither property is automatic from pretraining, and a search-based planner does not require every feature coordinate to change monotonically with an action. It needs a goal cost and dynamics model that preserve the task-relevant differences well enough to rank candidate actions.

In Part VIII, TD-MPC illustrated a representation shaped by task reward and value objectives. Frozen visual features take a different route: keep the representation fixed and fit the predictor in that space. Task-specific interaction data can favor adaptation; broad visual pretraining can make freezing attractive when such data are limited. Neither choice wins categorically. Encoder coverage, action data, dynamics, and the evaluation metric decide.

A useful first question is whether the task variable is visible in the available frames. Object location and scene layout may be; material compliance or hidden velocity may not be recoverable from one still image. Human ability to label a variable is not proof that DINOv2 retains it. Test pairs of states that differ in the variable, compare their features against nuisance variation, and use histories or other sensors when one frame is insufficient.

Is the feature vector actually Markov?

Often not. A Markov state makes the future independent of earlier history once the present state and the relevant actions are known. A single frame may leave velocity or an occluded object unresolved. A pendulum passing the same non-extreme angle with positive versus negative angular velocity looks the same in a still image but moves in opposite directions next. The one-frame embedding cannot resolve that ambiguity by itself.

There are three standard responses, all of which appear in practice:

  • Frame stacking. Concatenate features from the last kk frames. This is a relatively simple way to expose short-term history; a short stack can sometimes disambiguate velocity and direction of motion.
  • A recurrent or causal-attention predictor. Let the predictor consume a history of feature sets, so its internal state carries the missing information, as in Part V: World-Model Architectures. It can model interactions across a longer history, at the cost of more training and debugging complexity. For a fixed finite history, this is not categorically more expressive than a sufficiently capable model over stacked features.
  • Accept the approximation. In a nearly memoryless, well-observed task, one frame may be sufficient in practice. The last action alone does not generally reveal velocity or an occluded object; measure the resulting prediction and control error rather than assuming a planner will tolerate it.

DINO-WM conditions its transition model on a short history of encoded frames and actions. Such context can reveal direction of motion without an explicit recurrent hidden state. Whether the chosen history is sufficient, and what it costs relative to another architecture, depends on the environment and should be measured.

Feature-Space Dynamics Prediction

Now that we have a state, we need dynamics. The predictor is a function

z^t+1=Fϕ(zt,at)\hat{z}_{t+1} = F_\phi(z_t, a_t)

parametrized by learnable weights ϕ\phi, where ztz_t and zt+1z_{t+1} are frozen encoder outputs and ata_t is the action. The encoder is a fixed map from pixels to vectors. Its weights do not receive gradients from the dynamics loss; the toy predictor below is trained with Adam gradient steps. The predictor operates in the coordinate system the encoder defines.

The training objective and why it is a proxy

The natural objective is plain regression on the encoder's output:

L(ϕ)=E(zt,at,zt+1)∼D[∥Fϕ(zt,at)−zt+1∥2].\mathcal{L}(\phi) = \mathbb{E}_{(z_t, a_t, z_{t+1}) \sim \mathcal{D}} \left[ \left\| F_\phi(z_t, a_t) - z_{t+1} \right\|^2 \right].

Notice what changed relative to a pixel-reconstruction objective. There is no decoder in this dynamics loss. The target is the next observation's frozen features, so the quality of the learned state depends on what the encoder retains. DINO-WM can train a separate, optional diagnostic decoder; the small example below also trains a decoder while pretraining its toy encoder. Neither decoder is used for latent rollout or planning.

This design offers three potential advantages:

  • Dynamics training does not require pixel reconstruction. The transition predictor need not carry a pixel-generation head or optimize its loss. That does not make the encoder's pretraining free, nor does it automatically reallocate a fixed number of parameters.
  • A pretrained feature target may suppress nuisance detail. Depending on the task and encoder, feature error can emphasize object-level changes more usefully than pixel error. Neither metric guarantees that it weights task-critical details correctly.
  • The goal can use the same representation. Encoding a goal image gives a computable feature-distance cost without a separately trained reward model. Choosing that distance and verifying that it tracks success are still design and evaluation work.

There are two corresponding limitations:

  • Feature-space distance is a proxy for task success. Two visually different futures can map to nearby features, and two visually similar futures can map to distant features. A representation trained for invariance may collapse distinctions that a controller cares about. Encoded history or actions can sometimes recover a state distinction absent from one frame. If a task-critical distinction is absent from every input available to the predictor, improving that predictor alone cannot restore it. Separately, a fixed goal cost built from aliased goal features cannot distinguish task-distinct goals; that requires more information or a different representation or cost.
  • Latent predictions are not directly inspectable as images. A separately trained probe or diagnostic decoder can help, but its reconstruction quality must be checked; closed-loop behavior remains the decisive test.

A common practical refinement is to normalize features before regression. Squared differences in high-variance dimensions may dominate a raw Euclidean loss or planning distance; a large constant offset alone does not. One option is

z~t=zt−μσ,\tilde{z}_t = \frac{z_t - \mu}{\sigma},

where μ\mu and σ\sigma are estimated on training data, with a nonzero floor on any per-dimension scale. Per-dimension scaling changes the relative weighting of feature errors and can amplify nearly constant noise. A single global scale changes optimization units but leaves Euclidean rankings between candidate plans unchanged; it does not balance dimensions. Whether high-variance dimensions carry task signal must be tested rather than assumed.

Architectures for the predictor

Three useful design choices illustrate the tradeoffs here. MLPs and transformers are possible predictor backbones; a residual or delta target is an output parameterization that can be used with either. Convolutional, recurrent, graph-based, and state-space predictors are further options. The backbone choice depends in part on spatial structure and the cost of modeling interactions between distant scene regions.

MLPs. A per-token MLP applies shared weights to each patch, optionally with the action broadcast to each token. It is cheap but cannot directly exchange information between patches. A global MLP instead flattens all patch tokens before prediction; it can mix locations, but its input size and weights are tied to a fixed grid and it lacks an explicit spatial-neighborhood bias.

Transformers over patch tokens. Self-attention can mix information across locations and frames. Some architectures support variable token counts, but changing resolution also changes positional structure and computational cost; it is not automatic resolution transfer. DINO-WM conditions a transformer on a short frame/action history, concatenating an action embedding to each patch feature. Its temporal masking prevents access to future frames while allowing within-frame interaction; it is not a rule that each token sees only earlier tokens.

Residual or delta predictors. Instead of predicting zt+1z_{t+1} directly, predict the change:

Fϕ(zt,at)=zt+Gϕ(zt,at).F_\phi(z_t, a_t) = z_t + G_\phi(z_t, a_t).

When transitions are small in these coordinates, learning a delta can be convenient. The residual form alone does not initialize GϕG_\phi near zero: ordinary randomly initialized linear layers still produce nonzero outputs. Zero-initializing the last layer would create an identity starting point, but the toy model below does not do that. Persistence is a separate baseline to measure, not an automatic guarantee of residual training.

Action conditioning

The action has to enter the predictor somehow, and the choice matters whenever the dynamics are strongly action-dependent. A predictor that is insensitive to the action is useless for control, and a predictor that is over-sensitive to the wrong features of the action will chase spurious correlations. Common options include:

  • Concatenate the action vector to the feature vector before the first linear layer. Simple and works well when the dynamics are globally action-linear-ish.
  • Broadcast or concatenate an action embedding to every patch token, as in DINO-WM, so each token receives the command.
  • Insert an action token in the sequence, letting attention mix action and visual information. This is an architectural alternative, not the DINO-WM mechanism just described.
  • Concatenate a short history of actions, which helps when the visible effect of an action is delayed, such as when a command starts a process that only becomes visible a few frames later.

Across robots, actions differ in shape, units, timing, and physical effect. Learning a shared action representation is one research direction discussed in Part VI: Learning World Models at Scale, but an embedding alone cannot make embodiments interchangeable; calibration and embodiment-conditioned dynamics remain necessary.

Compounding error

Prediction error can accumulate over long learned-model rollouts. Understanding when that happens helps you choose horizons, budgets, and training procedures; an exact model or a contracting system is a counterexample to inevitable degradation.

Suppose one-step additive errors are independent, zero-mean vectors with root-mean-square norm ϵ\epsilon, and the transition does not amplify perturbations. Then the root-mean-square accumulated error is at most ϵH\epsilon\sqrt H under this nonexpansion assumption. Equality holds for an identity or norm-preserving transition that carries each error forward unchanged; contraction can give a smaller value. Correlation or an expanding transition Jacobian can increase error beyond this bound. Learned rollouts also feed predicted features back into the model, potentially moving off the training distribution. Neither the rate nor a sudden-failure threshold follows from one-step error alone.

Possible mitigations, with task-dependent effectiveness:

  • Train on predicted states. Add multistep losses or mix model predictions with observed inputs during training; the latter requires a specified schedule rather than merely invoking “scheduled sampling.” This is distinct from Dreamer-style imagination, which trains actor/value components using a learned world model.
  • Keep the rollout short and replan. Feedback can limit how long an inaccurate open-loop forecast is trusted. Later observations or encoded history may reveal state hidden in one frame, but replanning cannot restore a distinction absent from every available input or from an aliased fixed goal cost.
  • Train with suitable input perturbations. Noise on inputs may help local robustness if it is paired with meaningful targets and evaluated; independent zero-mean noise on regression targets alone does not teach robustness to perturbed inputs.
  • Ensembles. Predictor disagreement can expose some model uncertainty, as in PETS, but common-mode errors and representation failures can remain invisible.

An important subtlety: better one-step feature error does not always mean better closed-loop behavior. If the representation or goal cost is misspecified for the task, feature MSE can approach zero while control still fails. History or a later observation can sometimes disambiguate a one-frame state alias; fitting the same proxy more closely cannot recover information absent from all supplied inputs or distinguish goals collapsed by the fixed cost.

Action-Conditioned Latent Rollouts

Once the predictor exists, everything else is planning. A rollout is just recursive application:

z^t=ztobs,z^t+k+1=Fϕ(z^t+k,at+k),k=0,…,H−1.\hat{z}_{t} = z_t^{\text{obs}}, \qquad \hat{z}_{t+k+1} = F_\phi(\hat{z}_{t+k}, a_{t+k}), \qquad k = 0, \ldots, H-1.

At every step after the first, the input is the model's own previous output. This recursive, self-fed prediction is often called imagination. Errors can accumulate as it rolls forward, which is why the diagnostics in the code below matter.

Specifying the task with a goal image

The elegant part of the DINO-WM recipe is how it turns a rollout into a plan. There is no reward function and no value network. Instead, the task is specified by a goal observation ogo_g, and its encoded form

zg=fθ(og)z_g = f_\theta(o_g)

becomes the target. The planning objective is

at:t+H−1∗=arg⁡min⁡at:t+H−1∈AHd ⁣(z^t+H, zg),a^*_{t:t+H-1} = \arg\min_{a_{t:t+H-1}\in\mathcal A^H} d\!\left(\hat{z}_{t+H},\, z_g\right),

where dd is a distance in feature space, such as mean squared error over patch tokens, and A\mathcal A is the feasible action set. This terminal-state cost matches the DINO-WM planning objective. The toy implementation below deliberately uses a sum of intermediate-state costs to reward progress throughout its short horizon; that is a different objective and can select a different plan.

Three design consequences follow.

A test goal can be specified by one image. The planner does not need a task-specific reward model or demonstration for that goal, but the feature distance is still a chosen cost and the dynamics model needs offline observation–action trajectories. Some DINO-WM training datasets replay noisy expert trajectories; “zero-shot” here concerns a new test goal, not an absence of behavioral data.

Patch-level costs retain spatial indexing. Comparing corresponding patches can expose where images differ, unlike one pooled vector. The cost gradient need not point toward a useful action: encoder invariances, the learned transition, and occlusions can all distort it. A global embedding can still support search if it preserves goal-relevant information, but it offers less direct spatial diagnosis.

The cost is only as good as the representation and distance. If the encoder is insensitive to a task-critical difference, the cost can miss it; if it responds strongly to irrelevant variation, the planner can pursue the wrong match. DINOv2's behavior under position, texture, and lighting changes must be tested in the intended domain rather than presumed.

One practical caveat: the goal image may not show the exact configuration the agent must reach. If the agent needs to grasp an object, an image of the object at a destination may not show grip quality or contact force. Some tasks need additional process or contact information; a goal trajectory, tactile feedback, explicit constraints, or a final-approach policy are possible ways to supply it. The limitation is structural: a single frame does not specify the path or interaction process by which the configuration is reached.

Searching for actions

With the cost defined, planning optimizes an action sequence. If the predictor and cost are differentiable, one can backpropagate through a rollout; Part VII: Planning and Agency covers that option. The DINO-WM experiments use a sampling-based planner, CEM. It accommodates clipped action candidates and does not require cost gradients, but finite samples can miss good plans too. Neither family has a general advantage for every learned cost.

The standard choice is the cross-entropy method (CEM), a derivative-free optimizer:

  1. Initialize a Gaussian distribution N(μ,σ2I)\mathcal{N}(\mu, \sigma^2 I) over action sequences of length HH.
  2. Sample MM candidate sequences from the distribution.
  3. Roll out the predictor for each candidate and evaluate the cost against zgz_g.
  4. Keep the KK lowest-cost candidates (the elites).
  5. Refit μ\mu and σ\sigma to the elite set.
  6. Repeat for a fixed number of iterations. Return μ\mu.

The intuition behind CEM is that it repeatedly refits a sampling distribution to low-cost candidates, often concentrating search around promising action sequences without needing gradients. Its sampled variance need not shrink on every iteration.

Then use the result as model predictive control: execute only the first action (or the first few), observe the new frame, re-encode it, and replan. Frequent replanning limits open-loop exposure to model error, but cannot guarantee recovery from a poor representation, bad action coverage, or an unsafe intermediate action.

The horizon HH and sample count MM are important compute and accuracy controls. Longer horizons let the planner anticipate further, but later predictions may be less accurate; their quality must be measured. More samples cover more candidate sequences, with rollout cost roughly linear in MM for a fixed horizon and model. Elite count, iteration count, covariance handling, and the learned cost also matter. A longer horizon makes each candidate rollout more expensive, so horizon and sample count compete for a finite compute budget.

Receding-horizon control

Model predictive control is the loop “plan HH steps, execute one, replan.” Feedback from new observations can correct some model error. Re-encoding an observed frame also replaces the imagined latent at the next planning cycle. Neither operation guarantees task success or cancels a systematic bias within the executed action.

Does the cost being minimized correspond to the thing you want?

Sometimes, and this is the crux of evaluation for this family of models. Feature-space distance to a goal image is a proxy, and optimization can exploit a misaligned proxy. Two failure patterns recur:

  • Latent-space shortcuts. The planner may favor an action sequence whose predicted features look close to the goal but whose real outcome does not. Representation aliasing and model error are two different causes. Per-dimension whitening can amplify noisy small-variance directions, but it is not a necessary condition for failure.
  • Proxy insensitivity. The cost may barely change across task-relevant configurations. A separate final-approach controller is one possible remedy, but task-specific validation is needed.

Evaluate closed-loop task success alongside feature prediction and goal costs. The toy experiment reports both kinds of measure, but agreement on a small sample does not establish that the proxy will rank plans correctly in new states or goals.

Generalization with Pretrained Encoders

The central hypothesis behind DINO-WM is that a frozen, general-purpose visual encoder can support useful transfer across goals and scene variation. Whether it outperforms a task-specific representation depends on the task, data, and baseline. It helps to separate the forms of generalization rather than treating them as one result.

Visual generalization. Pretraining on diverse images can help an encoder recognize useful structure outside the dynamics dataset. It does not guarantee stability under a new camera, lighting condition, wallpaper, or room. Measure both feature changes and closed-loop performance under the shifts that matter for deployment.

Task generalization. A goal image can change the target without retraining the dynamics model, provided the new goal is observable in the representation and reachable through actions covered by training. In the cursor example, different target positions fit that pattern. Arbitrary goals do not.

Structural information. Patch tokens retain positions in a grid, so corresponding-token costs can reflect local differences. That is not the same as an explicit keypoint matcher: a moving object may shift across token indices, and appearance or occlusion can break naive correspondence.

What does not come for free is equally important.

New dynamics are not free. A visual encoder alone says nothing about how actions change the scene. A predictor trained on rigid-object motion may need additional data or adaptation for cloth; whether the encoder remains useful is an empirical question.

New observation regimes are not free. Camera, resolution, field-of-view, and modality shifts can change features and transition statistics. A natural-image RGB backbone should not be assumed to work on depth, thermal, or fisheye imagery without testing or adaptation. This is the distribution-shift problem discussed in Part VI: Learning World Models at Scale, now at the encoder.

Encoder bias can be a ceiling. If a frozen encoder maps two task-distinct observations to exactly the same current features and supplies no distinguishing feature history or other input, a feature-only predictor cannot recover the difference. Near-invariance is less absolute: a head might amplify retained signal, while a new sensor, history, or backbone adaptation may be needed if the distinction is truly lost. Prediction loss alone may not expose this failure.

The practical consequence is that frozen-feature world models should be evaluated on closed-loop behavior, and qualitatively on whether the encoder's features respond to the variables your task cares about. A quick diagnostic is to collect a batch of states that differ only in one task-relevant factor, encode them, and look at how much the features move. If that movement is barely detectable relative to nuisance variation or prediction noise, a low training loss offers little reassurance about closed-loop performance.

Finally, “frozen” admits intermediate designs. A trainable adapter atop a fixed backbone can reweight information that the backbone retained, but cannot reconstruct information discarded completely. Whether to adapt the backbone itself depends on data, compute, and the task-relevant sensitivity tests above; there is no universal episode-count threshold.

A Worked Example: Reaching a Key from a Picture

Before writing code, let's walk through the whole pipeline on an example small enough to check by hand. This is the same structure the code implements, just with numbers chosen to make arithmetic possible. Working through the arithmetic by hand is worth the effort because it exposes exactly which assumptions the pipeline relies on and exactly where they break.

The world. A point cursor moves in the unit square. The state is the cursor position st∈[0,1]2s_t \in [0,1]^2. The action at∈R2a_t \in \mathbb{R}^2 is a displacement, and the dynamics are st+1=clip⁡(st+at,0.05,0.95)s_{t+1} = \operatorname{clip}(s_t + a_t, 0.05, 0.95) with ∥at∥∞≤δ\|a_t\|_\infty \le \delta, where δ\delta is the per-coordinate action limit. Each frame shows a red disk at a fixed key position kk and a blue disk at the cursor. The task is to put the cursor on the key. The analytic calculation below assumes the action does not hit the clipping boundary.

The encoder. Suppose, for the sake of the example, that the encoder produces a two-number feature vector, one number for each cursor coordinate:

zt=[α st(x)β st(y)],α,β>0.z_t = \begin{bmatrix} \alpha\, s_t^{(x)} \\ \beta\, s_t^{(y)} \end{bmatrix}, \qquad \alpha, \beta > 0.

Real encoders are vastly more complex. This analytic feature map is invertible, so reaching zero feature distance is equivalent to reaching the physical goal. Invertibility alone does not preserve rankings of nonzero, weighted costs or guarantee that a finite-horizon planner finds the goal.

The dynamics. While clipping is inactive, the features evolve as

zt+1=zt+[α00β]at.z_{t+1} = z_t + \begin{bmatrix} \alpha & 0 \\ 0 & \beta \end{bmatrix} a_t.

A residual predictor Fϕ(z,a)=z+WaF_\phi(z, a) = z + W a with W=diag(α,β)W = \mathrm{diag}(\alpha, \beta) recovers the interior transition exactly. A fitted WW can approach that matrix if the training set includes enough unclipped interior transitions to identify both action directions and noise is controlled; boundary-clipped samples do not obey this linear relation. One global linear map cannot represent clipping everywhere. This idealized encoder already supplies the two cursor coordinates, so this calculation does not establish that a real visual encoder will do the same.

The goal. The goal image is the picture with the cursor drawn on the key, so

zg=[α k(x)β k(y)].z_g = \begin{bmatrix} \alpha\, k^{(x)} \\ \beta\, k^{(y)} \end{bmatrix}.

The plan. The toy example uses the summed cost ∑k=1H∥z^t+k−zg∥2\sum_{k=1}^{H} \|\hat{z}_{t+k} - z_g\|^2, unlike DINO-WM's terminal-state cost. If a feasible action sequence reaches zero feature residual, the analytic cursor is on the key. Neither invertibility nor this objective guarantees that a finite-horizon search finds that sequence, and nonzero plan rankings can differ from those under a physical-coordinate cost.

Limits of the analytic example. The rendered toy is rasterized and can map nearby physical states to identical frames before encoding; it can also occlude the red key beneath the cursor. Its learned encoder may discard more information, and its nonlinear dynamics model is only approximate. Finally, feature directions can receive unequal cost weights: if α\alpha is small, the analytic cost penalizes horizontal error weakly. These are distinct failure mechanisms, not proof that every plan fails.

The code below builds this pipeline in a synthetic world: an encoder pretrained on frames and then frozen, a learned feature-space predictor, and a CEM planner tested on goal images it has never seen. The closed-loop test will show both progress and misses.

Code Implementation

We'll build the pipeline in five stages: the world, the frozen encoder, the feature-space dynamics model, the rollout diagnostics, and the planner. Every piece is small enough to read in one sitting, and every number in the output is computed at run time. Keeping the pieces separate makes it clear which part of the system is responsible for which behavior, which is exactly the clarity a full-scale system makes hard to achieve.

Imports and configuration

In[3]:
Code
import matplotlib.pyplot as plt
import numpy as np
import torch
import torch.nn as nn
from book_plot_style import PALETTE, polish_axes, use_book_style

torch.manual_seed(0)
rng = np.random.default_rng(0)

The synthetic world

The world is deliberately simple so that we can check the learned model against ground truth. A 32×32, three-channel image shows a red key disk and a blue cursor disk. The cursor moves by the commanded displacement inside a box. The frame shows approximate disk positions, but rasterization aliases nearby cursor locations and the blue disk can occlude the red key. Information loss can therefore occur in the renderer as well as the learned encoder; the diagnostics below do not isolate one cause perfectly.

In[4]:
Code
IMG = 32


def render(cursor_xy, key_xy):
    """Paint a 32x32 RGB frame: red key disk plus blue cursor disk."""
    yy, xx = np.mgrid[0:IMG, 0:IMG].astype(np.float32)

    def disk(center, radius):
        cx = float(center[0]) * (IMG - 1)
        cy = float(center[1]) * (IMG - 1)
        return ((xx - cx) ** 2 + (yy - cy) ** 2) <= radius**2

    img = np.zeros((3, IMG, IMG), dtype=np.float32)
    key_mask = disk(key_xy, 2.2)
    img[0][key_mask] = 0.95
    img[1][key_mask] = 0.25
    img[2][key_mask] = 0.25
    cursor_mask = disk(cursor_xy, 2.2)
    img[0][cursor_mask] = 0.25
    img[1][cursor_mask] = 0.45
    img[2][cursor_mask] = 1.00
    return img


def step(state, action, low=0.05, high=0.95):
    """Displacement dynamics inside a box; the rasterized frame is lossy."""
    return np.clip(state + action, low, high)

Before learning an encoder, check what the renderer itself can hide. With the cursor drawn over the key, two different key positions can produce exactly the same 32×32 frame:

In[5]:
Code
occluding_cursor = np.array([0.50, 0.50], dtype=np.float32)
hidden_key_a = np.array([0.50, 0.50], dtype=np.float32)
hidden_key_b = np.array([0.51, 0.50], dtype=np.float32)
print(
    "distinct key positions render identically:",
    np.array_equal(
        render(occluding_cursor, hidden_key_a),
        render(occluding_cursor, hidden_key_b),
    ),
)
Out[5]:
Console
distinct key positions render identically: True

No visual encoder can recover a distinction absent from identical input pixels. This is renderer occlusion and rasterization, not a failure introduced by frozen features.

Let's render one frame and look at it, to fix the scale of the problem: a disk of radius 2.2 pixels corresponds to about 7% of the image width.

Out[6]:
Console
frame shape: (3, 32, 32)
pixel values in [0, 1]: 0.0 to 1.0
fraction of pixels that are background: 0.970703125

The reported background fraction shows why uniform pixel MSE can be a poor emphasis for this task. It does not by itself establish where another model spends its capacity.

Collecting unlabeled trajectories

We collect random-walk trajectories. Each episode gets a fresh key position and a fresh start, so the dataset covers a range of scene layouts. The first 32 episodes are the "training" split for the dynamics model; the last 8 are held out. Holding out entire episodes rather than random frames is important, because frames within an episode are highly correlated; a random-frame split would leak near-duplicate frames across the boundary and make the evaluation optimistic.

In[7]:
Code
N_EPISODES = 40
EP_LEN = 100
N_TRAIN_EPISODES = 32
ACTION_SCALE = 0.06
MAX_ACTION = 0.08

frames, states, actions, keys, ep_ids = [], [], [], [], []
for ep in range(N_EPISODES):
    key_xy = rng.uniform(0.15, 0.85, size=2)
    state = rng.uniform(0.10, 0.90, size=2)
    for _ in range(EP_LEN):
        frames.append(render(state, key_xy))
        states.append(state.copy())
        keys.append(key_xy.copy())
        ep_ids.append(ep)
        action = np.clip(
            rng.normal(scale=ACTION_SCALE, size=2), -MAX_ACTION, MAX_ACTION
        )
        actions.append(action)
        state = step(state, action)

frames = np.stack(frames).astype(np.float32)
states = np.stack(states).astype(np.float32)
actions = np.stack(actions).astype(np.float32)
keys = np.stack(keys).astype(np.float32)
ep_ids = np.array(ep_ids)
train_episode_mask = ep_ids < N_TRAIN_EPISODES
Out[8]:
Console
frames:  (4000, 3, 32, 32)
actions: (4000, 2)
training episodes: 32, held-out episodes: 8
training frames: 3200
max displacement between consecutive frames: 0.080

Consecutive states differ by at most about 0.08 in each coordinate, roughly 2.5 pixels on this raster. That small physical step motivates trying a residual predictor. It does not prove that encoded features change smoothly at a patch boundary or that residual training beats a direct predictor; those are empirical questions for the diagnostics below.

A frozen patch encoder

This is a pedagogical substitute, not a DINOv2 reproduction. We pretrain a small convolutional encoder with a reconstruction objective and a foreground-weighted loss, then freeze it before dynamics training. DINOv2 uses a different pretraining objective and data scale; the shared feature of the two systems is only that a fixed visual encoder supplies the predictor's coordinates.

The encoder produces a 4×4 grid of 32-dimensional tokens, mirroring the patch-token structure of a ViT. The grid permits spatially indexed feature comparison for the goal-image cost; whether that cost tracks task success must be tested.

In[9]:
Code
TOKEN_DIM = 32
N_TOKENS = 16
LATENT_DIM = TOKEN_DIM * N_TOKENS


class PatchEncoder(nn.Module):
    """Frozen visual backbone: image -> 4x4 grid of patch tokens."""

    def __init__(self, token_dim=TOKEN_DIM):
        super().__init__()
        self.net = nn.Sequential(
            nn.Conv2d(3, 32, 3, stride=2, padding=1),
            nn.ReLU(),
            nn.Conv2d(32, 64, 3, stride=2, padding=1),
            nn.ReLU(),
            nn.Conv2d(64, 64, 3, stride=2, padding=1),
            nn.ReLU(),
            nn.Conv2d(64, token_dim, 1),
        )

    def forward(self, x):
        features = self.net(x)  # (B, token_dim, 4, 4)
        return features.flatten(2).transpose(1, 2)  # (B, 16, token_dim)


class PatchDecoder(nn.Module):
    """Pretrains the encoder; excluded from dynamics and planning, retained for diagnostics."""

    def __init__(self, token_dim=TOKEN_DIM):
        super().__init__()
        self.net = nn.Sequential(
            nn.ConvTranspose2d(token_dim, 64, 4, stride=2, padding=1),
            nn.ReLU(),
            nn.ConvTranspose2d(64, 32, 4, stride=2, padding=1),
            nn.ReLU(),
            nn.ConvTranspose2d(32, 16, 4, stride=2, padding=1),
            nn.ReLU(),
            nn.Conv2d(16, 3, 3, padding=1),
            nn.Sigmoid(),
        )

    def forward(self, tokens):
        b, n, d = tokens.shape
        side = int(round(n**0.5))
        grid = tokens.transpose(1, 2).reshape(b, d, side, side)
        return self.net(grid)

Now we pretrain the encoder on the training episodes' frames, using no actions or goal/outcome labels. Roughly 97% of pixels are black background, so plain pixel MSE could favor a nearly black output. We threshold the rendered RGB frames to identify nonblack foreground pixels and upweight them. This is a hand-designed synthetic-domain foreground prior, not a separately supplied renderer disk mask or a DINOv2-style self-supervised objective.

In[10]:
Code
imgs = torch.tensor(frames)
imgs_train = imgs[train_episode_mask]

encoder, decoder = PatchEncoder(), PatchDecoder()
ae_optimizer = torch.optim.Adam(
    list(encoder.parameters()) + list(decoder.parameters()), lr=2e-3
)

AE_STEPS, AE_BATCH = 600, 64
ae_history = []
for _ in range(AE_STEPS):
    batch = imgs_train[torch.randint(0, imgs_train.shape[0], (AE_BATCH,))]
    reconstruction = decoder(encoder(batch))
    foreground = (batch.amax(dim=1, keepdim=True) > 0.1).float()
    pixel_weight = 1.0 + 20.0 * foreground
    loss = (pixel_weight * (reconstruction - batch).square()).mean()
    ae_optimizer.zero_grad()
    loss.backward()
    ae_optimizer.step()
    ae_history.append(loss.item())

Now freeze. From this point forward, no gradient ever flows into encoder.

Out[12]:
Console
encoder pretraining loss: 0.31562 -> 0.00231
frozen parameters in encoder: 58,400
tokens per frame: 16, token dim: 32, latent size: 512

The decoder remains in memory for the reconstruction diagnostic below, but is excluded from dynamics training and planning. We do not use it again after that diagnostic.

Out[13]:
Visualization
Sample one original red-key and blue-cursor scene above the decoded approximation.
Sample 1: original frame above its decoder reconstruction from frozen 4×4 patch tokens, on the same 0–1 pixel scale. The decoder is a diagnostic, not part of the transition model.
Sample two original red-key and blue-cursor scene above the decoded approximation.
Sample 2: original frame above its decoder reconstruction, using the same fixed 0–1 display scale.
Sample three original red-key and blue-cursor scene above the decoded approximation.
Sample 3: original frame above its decoder reconstruction, using the same fixed 0–1 display scale.
Sample four original red-key and blue-cursor scene above the decoded approximation.
Sample 4: original frame above its decoder reconstruction, using the same fixed 0–1 display scale.

The common display scale lets us check whether the 512-number latent retains approximate disk positions, color, and intensity instead of merely inspecting a contrast-enhanced image. Even a visually plausible reconstruction would not establish that these features are sufficient for control; that is what the following rollout and planning tests examine.

Extracting and normalizing features

We encode every frame once and apply a global normalization. This changes the numerical scale of the regression loss; unlike per-dimension whitening, it does not change relative feature weights or candidate-plan rankings under Euclidean distance.

We estimate the normalization statistics from training frames only. Per-dimension whitening is an alternative, but it can amplify nearly constant noisy dimensions. With one scalar scale, high-variance directions retain their relative influence; whether that influence is useful requires a task-level check.

In[14]:
Code
with torch.no_grad():
    Z = encoder(imgs).numpy()  # (N, 16, TOKEN_DIM)

Z_train = Z[train_episode_mask]
feature_mean = Z_train.reshape(-1, TOKEN_DIM).mean(axis=0)
feature_scale = (Z_train - feature_mean).std()
Z_norm = (Z - feature_mean) / feature_scale
Z_flat = Z_norm.reshape(len(Z_norm), LATENT_DIM)
Out[15]:
Console
raw feature tensor: (4000, 16, 32)
global feature scale: 1.3882
per-dimension std after normalization:
  min 0.587, median 0.885, max 1.982

The spread of per-dimension standard deviations after global scaling is a diagnostic of relative variation. Dominance in variation can affect a squared-error cost, but variance alone does not tell us whether the changing directions matter for the goal.

The cost is easier to inspect as a landscape. We render the cursor over a fixed grid with a fixed key, encode each frame, and evaluate feature distance to a goal image drawn at the key. Rasterization can make nearby positions identical, and the learned Conv–ReLU encoder is not known to be invertible or globally smooth. The observed landscape is a diagnostic for this trained toy, not a general theorem.

Out[17]:
Visualization
Heatmap of feature-space cost to a goal over cursor positions with goal star.
Feature-space cost to a goal image over a sampled grid of cursor positions. The measured minimum is near the key; rasterization and a learned encoder do not guarantee an exact or unique physical-state minimum.

Building the transition dataset and learning dynamics

Transition pairs are consecutive frames within the same episode. Cross-episode pairs must be excluded, since they are not real transitions. This is the same leakage concern as the train/test split: joining the last frame of one episode to the first frame of the next would create a pair that no action could produce, and it would corrupt the learned dynamics.

In[18]:
Code
same_episode = np.r_[ep_ids[1:] == ep_ids[:-1], False]
train_pairs = np.where(train_episode_mask & same_episode)[0]

Z_t = torch.tensor(Z_flat[train_pairs])
Z_next = torch.tensor(Z_flat[train_pairs + 1])
A_t = torch.tensor(actions[train_pairs])
Out[19]:
Console
training transitions: 3,168
input dim: 514 (latent plus action), output dim: 512

The predictor is a residual MLP: it predicts a change in feature space rather than the next feature directly. The code does not zero-initialize its final layer, so the initial residual is not guaranteed to be small. We measure improvement against persistence below.

The action enters by concatenation. The physical update is action-affine only away from clipping boundaries; whether this MLP handles boundary transitions is empirical.

In[20]:
Code
class LatentDynamics(nn.Module):
    """Residual predictor in normalized feature space."""

    def __init__(self, latent_dim=LATENT_DIM, action_dim=2, hidden=256):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(latent_dim + action_dim, hidden),
            nn.SiLU(),
            nn.Linear(hidden, hidden),
            nn.SiLU(),
            nn.Linear(hidden, latent_dim),
        )

    def forward(self, z, a):
        return z + self.net(torch.cat([z, a], dim=-1))


dynamics = LatentDynamics()
dyn_optimizer = torch.optim.Adam(dynamics.parameters(), lr=1e-3)

DYN_STEPS, DYN_BATCH = 800, 256
dyn_history = []
for _ in range(DYN_STEPS):
    idx = torch.randint(0, Z_t.shape[0], (DYN_BATCH,))
    prediction = dynamics(Z_t[idx], A_t[idx])
    loss = nn.functional.mse_loss(prediction, Z_next[idx])
    dyn_optimizer.zero_grad()
    loss.backward()
    dyn_optimizer.step()
    dyn_history.append(loss.item())

dynamics.eval()
Out[21]:
Console
dynamics loss: 0.31021 -> 0.16331
predicting no change at all would give: 0.28829
full-train fitted MSE: 0.16575
paired full-train reduction over persistence: 42.5%

The paired full-training-set comparison against persistence checks whether the predictor improves on a simple baseline for this one-step metric. The preceding last-minibatch loss is a training trace, not the numerator of that comparison. Held-out transition error and closed-loop task success are needed before calling the learned dynamics useful.

Rollout diagnostics: one-step versus recursive error

Now the key measurement. We roll the predictor forward on held-out episodes and compare two error curves:

  • One-step error: always start from a true feature vector and predict one step ahead. This measures the model in the regime it was trained on.
  • Recursive error: start from a true feature vector and then feed predictions back in. This measures the model in the regime the planner will actually use it.
  • Persistence error: compare each future observation with the fixed starting feature. This is a reference curve, not a mathematical lower bound for a learned model or a policy.
In[22]:
Code
HORIZON_EVAL = 15
STRIDE = 12

start_idx = []
for ep in range(N_TRAIN_EPISODES, N_EPISODES):
    for t in range(0, EP_LEN - HORIZON_EVAL, STRIDE):
        start_idx.append(ep * EP_LEN + t)
start_idx = np.array(start_idx)

z0_eval = torch.tensor(Z_flat[start_idx])
src_idx = start_idx[:, None] + np.arange(HORIZON_EVAL)[None, :]
targets_eval = torch.tensor(Z_flat[src_idx + 1])

with torch.no_grad():
    rolling = []
    z = z0_eval
    for h in range(HORIZON_EVAL):
        z = dynamics(z, torch.tensor(actions[src_idx[:, h]]))
        rolling.append(z)
    rollout_pred = torch.stack(rolling, dim=1)

    one_step_pred = dynamics(
        torch.tensor(Z_flat[src_idx]),
        torch.tensor(actions[src_idx]),
    )

recursive_mse = (
    ((rollout_pred - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)
one_step_mse = (
    ((one_step_pred - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)
persistence_curve = (
    ((z0_eval[:, None, :] - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)
Out[23]:
Console
held-out start states: 64
one-step MSE at horizon 1:      0.33500
recursive MSE at horizon 1:     0.33500
recursive MSE at horizon 15:     2.22436
persistence MSE at horizon 15:   1.14052
recursive error growth factor (h=15 compared with h=1): 6.6x
Out[24]:
Visualization
Line chart of three error curves rising from a low flat line to a higher compounding curve.
Feature-space error against rollout horizon on held-out episodes. One-step predictions use observed features at each step; recursive predictions feed back model outputs. In this run, recursive error rises and crosses the fixed-start persistence reference between horizons six and seven.

The gap between the one-step and recursive curves shows how feedback of model predictions affects this held-out set. Where recursive MSE crosses the fixed-start persistence reference, the model predicts these recorded open-loop trajectories less accurately under this feature metric. That is not a comparison of action policies, a safe-horizon certificate, or a statement about every initial state.

This diagnostic motivates testing shorter horizons, alternative models, and closed-loop control. A CEM planner optimizes goal cost rather than the random-trajectory prediction MSE plotted here, so the crossing cannot establish that doing nothing is better or that search cannot help.

Planning to a goal image with the cross-entropy method

The toy planner searches over action sequences and sums feature distance to the encoded goal at every predicted step. This favors earlier approach to the goal, but it differs from DINO-WM's terminal-cost objective stated above.

In[25]:
Code
def cem_plan(
    z_start, z_goal, horizon=8, n_samples=128, n_elite=12, n_iters=3, seed=0
):
    """Cross-entropy method over action sequences, scored in feature space."""
    plan_rng = np.random.default_rng(seed)
    mean = np.zeros((horizon, 2), dtype=np.float32)
    std = np.full((horizon, 2), MAX_ACTION, dtype=np.float32)
    z0 = torch.tensor(z_start[None, :])
    zg = torch.tensor(z_goal[None, :])

    for _ in range(n_iters):
        noise = plan_rng.standard_normal((n_samples, horizon, 2)).astype(
            np.float32
        )
        candidates = np.clip(
            mean[None] + std[None] * noise, -MAX_ACTION, MAX_ACTION
        )
        candidate_t = torch.tensor(candidates)

        with torch.no_grad():
            z = z0.repeat(n_samples, 1)
            cost = torch.zeros(n_samples)
            for h in range(horizon):
                z = dynamics(z, candidate_t[:, h])
                cost = cost + ((z - zg) ** 2).mean(dim=1)

        elite = candidates[torch.argsort(cost)[:n_elite].numpy()]
        mean = elite.mean(axis=0)
        std = elite.std(axis=0) + 1e-3

    return mean

The MPC loop re-encodes a rendered observation at every real step, so the next search starts from the encoder's output rather than the previous imagined feature. That observation is still lossy because of rasterization and possible occlusion.

In[26]:
Code
def encode_state(state_xy, key_xy):
    """Render, encode with the frozen backbone, and normalize."""
    frame = render(state_xy, key_xy)
    with torch.no_grad():
        z = encoder(torch.tensor(frame[None])).numpy()[0]
    return ((z - feature_mean) / feature_scale).reshape(-1).astype(np.float32)


def run_mpc(key_xy, start_xy, n_steps=20, horizon=8, seed=0):
    """Receding-horizon control. Execute one action, observe, replan."""
    goal_latent = encode_state(key_xy, key_xy)
    state = np.asarray(start_xy, dtype=np.float32).copy()
    trajectory = [state.copy()]
    for k in range(n_steps):
        current_latent = encode_state(state, key_xy)
        plan = cem_plan(
            current_latent, goal_latent, horizon=horizon, seed=seed + k
        )
        state = step(state, plan[0])
        trajectory.append(state.copy())
    return np.array(trajectory)

Evaluating closed-loop success

This is where proxy and task outcomes can be compared. We track final physical distance from cursor to key and final observed-frame feature distance to the goal. The latter is related to, but not identical to, the summed predicted feature cost CEM optimized. A random-action baseline runs on the same scenarios with the same number of steps.

In[27]:
Code
N_SCENARIOS = 10
eval_rng = np.random.default_rng(1234)
eval_keys = eval_rng.uniform(0.20, 0.80, size=(N_SCENARIOS, 2)).astype(
    np.float32
)
eval_starts = eval_rng.uniform(0.15, 0.85, size=(N_SCENARIOS, 2)).astype(
    np.float32
)

mpc_trajectories, mpc_final_distance = [], []
for i in range(N_SCENARIOS):
    traj = run_mpc(
        eval_keys[i], eval_starts[i], n_steps=20, horizon=8, seed=100 * i
    )
    mpc_trajectories.append(traj)
    mpc_final_distance.append(float(np.linalg.norm(traj[-1] - eval_keys[i])))
mpc_final_distance = np.array(mpc_final_distance)

random_final_distance = []
for i in range(N_SCENARIOS):
    state = eval_starts[i].copy()
    for _ in range(20):
        noise = np.clip(
            eval_rng.normal(scale=ACTION_SCALE, size=2), -MAX_ACTION, MAX_ACTION
        )
        state = step(state, noise)
    random_final_distance.append(float(np.linalg.norm(state - eval_keys[i])))
random_final_distance = np.array(random_final_distance)
Out[28]:
Console
initial distance to key, mean:      0.397
latent MPC final distance, mean:    0.378
random walk final distance, mean:   0.396
latent MPC success rate (<= 0.10):  10%
random success rate (<= 0.10):      0%

The toy model was trained without goal/outcome labels or demonstrations, but its encoder pretraining used a foreground mask derived by thresholding rendered RGB pixels. It does not consistently reach the key in these held-out scenarios. It sometimes approaches the target and sometimes stalls far away. The predictor learned from random-walk transitions in 32 episodes; the test goal was supplied as an image. Reporting misses alongside successes matters because feature-space fit is not task success.

For each held-out scenario we compare final physical distance with final observed-frame feature cost. This is a descriptive cross-scenario check: each point has a different key and start, and the plotted final-state cost is not the exact summed predicted cost optimized within that scenario. Rank inversions across different goals do not prove that CEM misranks candidate plans for one goal. A within-goal candidate study would be needed for that conclusion.

Out[30]:
Visualization
Scatter of final task distance versus feature-space cost across held-out scenarios.
Final observed-frame feature cost against final physical distance for ten distinct held-out scenarios. Each point has a different goal and start; cross-scenario ranks do not test candidate-plan ranking within one goal.
Out[31]:
Visualization
Scenario one square plot with cursor path, start circle, and goal star separated at the end.
Held-out scenario 1: latent MPC cursor trajectory from the circled start toward the starred goal supplied as an image. The trajectory finishes visibly short of the goal.
Scenario two square plot with cursor path, start circle, and goal star.
Held-out scenario 2: latent MPC cursor trajectory from the circled start toward the starred goal supplied as an image.
Scenario three square plot with cursor path, start circle, and goal star separated at the end.
Held-out scenario 3: latent MPC cursor trajectory from the circled start toward the starred goal supplied as an image. The trajectory finishes visibly short of the goal.
Scenario four square plot with cursor path, start circle, and goal star.
Held-out scenario 4: latent MPC cursor trajectory from the circled start toward the starred goal supplied as an image.
Out[32]:
Visualization
Bar chart of mean final goal distance for latent MPC and random actions, with ten horizontally jittered individual outcomes over each bar.
Final distance to the goal key for the learned latent planner and a random-action baseline with the same number of steps on ten scenarios. Bars show sample means; dots show individual outcomes. The small mean gap and wide spread do not establish a general advantage.

A note on what just happened

The encoder in this example is not DINOv2. It is a tiny convolutional network trained by foreground-weighted reconstruction on a few thousand 32×32 RGB images. What it shares with DINOv2 is a system role: a frozen visual map supplies features for dynamics learning and goal comparison. The toy uses a one-frame MLP, global normalization, and a summed-horizon CEM cost; a real-backbone implementation must revisit those choices rather than copy them uncritically.

Scaling up changes the data distribution, patch grid, feature dimension, action dynamics, and compute budget. A pretrained backbone may improve useful visual invariance, but that must be verified for the task. A Transformer is one way to model token interactions, not a requirement imposed by dimension alone. The toy remains useful for understanding the interfaces and diagnostics, not for predicting real-robot performance.

Limitations and Impact

DINO-WM combines a frozen pretrained visual encoder, action-conditioned feature prediction, and goal-image planning. Its dynamics objective and planner do not require a pixel decoder or a reward head. That does not mean arbitrary tasks are covered, that behavioral data contain no demonstrations, or that no diagnostic decoder can be trained. Compared with decision-centric methods such as Dreamer, MuZero, and TD-MPC, the distinguishing design choice is to use an externally pretrained visual representation as the fixed prediction and goal-comparison space.

The methodological division is useful: generic images can pretrain an encoder, while environment-specific observation–action trajectories train the transition model. The latter data are not the same as unlabeled video, and collecting them may be costly. This decomposition is one research option, not a universal template for later video generators or vision-language-action systems, some of which train representations and dynamics jointly.

The limitations are equally real, and they cluster into five groups.

Representation sensitivity matters. If the frozen encoder completely merges task-distinct observations and the predictor has no other input or history, downstream layers cannot reconstruct that distinction. A targeted paired-state probe, representation sensitivity analysis, and closed-loop evaluation can expose such failures. Small but nonzero retained differences might be amplified by a trainable head; complete collapse may require backbone adaptation, a new sensor, or temporal context. Loss curves alone are insufficient.

Feature distance is a proxy. The planning cost is a squared distance in a high-dimensional space—512 coordinates for the toy, and 98,304 raw patch-feature coordinates for the stated ViT-S/14 grid and width. Dimensionality alone does not prove an exploitable shortcut. Patch-level costs and short replanning horizons may help in some tasks, but neither guarantees alignment. Test within-goal candidate rankings and report closed-loop task outcomes.

Compounding error limits open-loop trust. In this toy run, recursive error grows over fifteen steps and crosses the fixed-start persistence reference between steps six and seven, while one-step error stays lower. That argues for testing shorter horizons here, not prescribing a universal horizon. Closed-loop trials determine whether replanning compensates for this open-loop error. A separately trained decoder or probe can aid inspection, subject to its own fidelity limits.

A goal image omits process constraints. A final configuration does not specify “move the block without knocking over the cup,” nor does it show an internal state invisible to the camera. Multiple images, a goal video, language, explicit constraints, or a learned task cost may supply more information; each adds its own modeling and validation burden.

Latent predictions need diagnostics. The transition model does not itself output video. A separate decoder or probe can visualize aspects of a predicted feature state, but neither is automatically cheap, faithful, or sufficient for safety assurance. In consequential deployments, validate the actual closed-loop behavior and safety constraints directly.

Frozen pretrained encoders can make goal-image planning practical without a pixel-reconstruction loss in the dynamics model. Their value relative to task-specific alternatives is empirical. Treat the encoder as a hypothesis about which distinctions matter, then test that hypothesis against the task before trusting the planner.

Summary

  • Frozen visual representations as state. A fixed encoder such as DINOv2 supplies per-frame patch features; DINO-WM's transition model also uses short observation and action histories. Patch-level features preserve spatial indexing, but not guaranteed object correspondence.
  • The trade. Pretraining can provide useful visual structure without task-specific encoder training, while freezing can hide distinctions needed for control. Test the representation in the intended domain.
  • Feature-space dynamics prediction. The dynamics objective predicts next-step frozen features from action-conditioned context. It has no pixel-reconstruction term; an optional diagnostic decoder is separate. Feature MSE remains a proxy for control.
  • Normalization. A single global scale changes numeric units without changing Euclidean plan rankings. Per-dimension scaling changes relative weights and can amplify noisy nearly constant dimensions.
  • Action-conditioned rollouts. Recursive predictions can drift. Compare one-step, recursive, and fixed-start persistence errors on held-out trajectories, then test horizons in closed loop.
  • Planning with goal images. DINO-WM's CEM planner uses terminal feature distance to an encoded goal; the toy uses a summed intermediate cost. A new test goal can be supplied as one image, but dynamics training still requires aligned observation–action trajectories.
  • Generalization. Goal-image reuse applies only where the representation and learned dynamics retain the necessary information and action coverage. Scene, camera, modality, and dynamics shifts require empirical checks or adaptation.
  • Evaluate closed-loop. Report task success and inspect feature sensitivity to task-critical variables; neither low feature MSE nor a cross-goal scatter plot proves plan ranking within a goal.

The next chapter turns to video generators, which can produce visible future frames rather than only latent predictions. Whether those frames are action-consistent enough to support planning is a separate empirical question.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about pretrained visual models and DINO-WM.

Pretrained Visual Models and DINO-WM

Question 1 of 80 of 8 completed
In the DINO-WM recipe described in this chapter, which component is trained and which is frozen?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026pretrainedvisual, author = {Michael Brenndoerfer}, title = {Pretrained Visual Models and DINO-WM}, year = {2026}, url = {https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Pretrained Visual Models and DINO-WM. Retrieved from https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models
MLAAcademic
Michael Brenndoerfer. "Pretrained Visual Models and DINO-WM." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models>.
CHICAGOAcademic
Michael Brenndoerfer. "Pretrained Visual Models and DINO-WM." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Pretrained Visual Models and DINO-WM'. Available at: https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Pretrained Visual Models and DINO-WM. https://mbrenndoerfer.com/writing/dino-wm-pretrained-visual-models-world-models

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.