Part of World Models Handbook
Explains how DINO-WM freezes a pretrained visual encoder, trains action-conditioned feature dynamics, and uses goal-image planning.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Pretrained Visual Models and DINO-WM
A camera-frame predictor must account for appearance as well as motion. An unweighted pixel loss scores the wall, anti-aliased edges, and sensor noise alongside the object the agent must move. Depending on the architecture, modeling those details can use training or rollout compute without improving the decision. Higher resolution increases the number of predicted values; the fraction occupied by a task-relevant object shrinks only if the field of view or background grows relative to it. A photorealistic frame is therefore not, by itself, evidence of useful control dynamics.
DINOv2 offers a useful starting point: it was trained without labels on the 142-million-image LVD-142M corpus, and its patch features retain spatial and object-part cues in the authors' evaluations. Those cues are not a solved control state. DINO-WM tests a narrower proposition: freeze a pretrained visual encoder, then learn how its features change under recorded actions. This separates reusable visual pretraining from dynamics that still have to be learned for an environment and its action interface.
So DINO-WM splits the problem in two:
- A frozen pretrained visual encoder turns each observation into a set of feature vectors. Its weights do not change during dynamics training.
- A trainable predictor uses a history of those features and recorded actions to predict future features. The worked toy later simplifies this to one frame and one action.
The DINO-WM dynamics loss and planner need no image decoder or reward model; the paper trains an optional separate decoder to inspect predictions. Training still requires offline observation–action trajectories, which may include expert behavior, but a new feasible image goal needs no expert demonstration or reward label for that particular goal. Planning searches for actions whose predicted patch features approach the goal image's features. That distance is an explicit, imperfect task cost, not a reward function that appears automatically from pretraining.
This chapter explains when that separation helps, when it fails, and how to build a small version of it. It sits in Part IX: Foundation and World-Action Models after tokenized game and environment models, and prepares the ground for video generators and vision-language-action systems. Those later systems make different architectural choices; they do not all freeze a visual encoder or learn a separate small dynamics model.
The core trade is reuse against task adaptation. A frozen encoder gives the dynamics learner a fixed visual coordinate system and avoids learning that representation again from the environment's action data. It also limits what the predictor can infer from a frame: a distinction absent from the features cannot be recovered from that frame alone, although history or another sensor may help. The encoder is a hypothesis about which visual distinctions matter, and the rest of the pipeline depends on testing it.
We'll build the whole pipeline end to end on a tiny synthetic world so that every piece is inspectable. Along the way we'll examine three distinct failure modes: feature aliasing (task-distinct states look the same to the encoder), dynamics error (the learned predictor misjudges the effect of an action), and proxy misalignment (a low feature-space goal cost does not imply task success). The toy exposes these mechanisms, but it does not establish how often they occur with DINOv2 or on a robot.
Frozen Visual Representations as State
The first design decision in a world model is what information the predictor uses as state. In Part III: Representing Agents and Worlds, predictive sufficiency meant that the future is conditionally independent of earlier history given that state and the actions. A frozen encoder supplies a per-frame embedding, not a guarantee of sufficiency. DINO-WM's transition model also consumes a short history of embeddings and actions; the one-frame model in our toy is an explicit simplification. The frozen visual features are not shaped by a task reward or value signal, but the learned predictor can aggregate their temporal context.
Formally, let be the observation (an image) and let be the encoder with parameters that are never updated. Define
For a Vision Transformer with patch size , we can retain a grid of patch tokens rather than pool them into one vector. The transformer divides a patch-divisible image into non-overlapping patches, embeds each patch, and lets attention mix information across positions. In the equation below, and are the dimensions after any padding and are divisible by . Excluding any global or register tokens, the patch output has shape
For DINOv2 ViT-S/14 at resolution, that is patch tokens of dimension : a per-frame feature matrix. Prediction may still need earlier frames and actions.
A global embedding (the [CLS] token, or a pooled average over patches) summarizes the image in one vector. Patch tokens retain spatial indices, though attention lets each token contain wider scene context. Comparing corresponding patch indices can localize feature differences; a global embedding can also provide a useful goal cost. Both comparisons ultimately return a scalar to the planner, and neither guarantees a useful action gradient. DINO-WM uses patch-level feature distance for its image-goal cost.
What makes a frozen encoding good enough to be a state
Not every pretrained network produces a usable state space. The properties that matter for world modeling are subtly different from the properties that matter for classification. A classifier only needs to represent whatever distinguishes its label categories, whereas a world model needs to represent everything that changes in a way that affects the future. These are related but not identical demands, and a backbone that excels at one can be mediocre at the other.
- Informativeness. The encoding must preserve the distinctions that matter for predicting dynamics. If situations with different relevant futures produce the same complete predictor input, including the available feature history and actions, that predictor cannot separate them. A collision in one frame's features alone may be resolved by history or another sensor; freezing only blocks recovery of distinctions absent from every available encoded input. A fixed one-frame feature cost can still alias distinct goals.
- Smoothness. Small, task-relevant changes should ideally produce predictable feature changes. An erratic local encoding can make dynamics harder to learn, but smoothness alone does not make interpolation in feature space physically realizable.
- Spatial structure. Patch indices expose image location and can help a predictor exploit locality. Translation equivariance means that shifting an input shifts the output in the same way; a patch grid does not confer that property on a position-embedded Vision Transformer, and real motion need not be a translation.
- Stability across nuisance factors. Lighting, texture, exposure, and clutter may be irrelevant to one task but important to another; measure whether the feature response suppresses only the variation your task can safely ignore.
- Cost. A patch set contains 98,304 scalar values, versus 150,528 RGB channel values at . That is only about a 1.53-fold reduction in raw element count, not proof of a storage or compute advantage. The possible benefit is in the representation's usefulness to prediction and planning, which must be measured rather than inferred from dimension alone.
Self-supervised objectives emphasize these properties differently; no universal ordering or necessary trade-off follows from the objectives alone. The choice of pretraining method is therefore not a detail you can ignore; it helps shape the state space your predictor will inhabit.
SimCLR and MoCo contrast augmented views of the same image against other examples. Their objectives encourage agreement under chosen crops and color changes; they do not guarantee exact invariance in the features later used for control. Strong color augmentation could make a red/blue distinction harder to use, for example, if color is the variable a robot needs. The augmentation family is therefore an assumption to probe against the downstream task.
MAE reconstructs masked image patches from visible-patch encodings. Its reconstruction objective allocates training effort to pixels, but that does not establish that its learned features are merely low-level or unsuitable for dynamics. A fair encoder comparison has to hold the predictor and control task fixed.
DINO trains a student to match an exponential-moving-average teacher across augmented crops, with centering and sharpening in its collapse-prevention recipe and no explicit negative pairs. DINOv2 builds on that lineage with an iBOT-style masked-patch objective and KoLeo regularization of class-token feature spread, among other changes. Its paper reports useful spatial and object-part cues, including qualitative correspondence examples; its semantic-segmentation benchmarks train readout heads, so they are not evidence of segmentation by raw nearest neighbors alone. DINO-WM uses DINOv2 patch features, but that does not make DINOv2 the default backbone for every feature-space world model.
The invariance-versus-controllability tension
Here is the tension stated plainly. A frozen encoder was optimized to make some distinctions and ignore others. Those choices were made by whoever pretrained it, for whatever data and objective they had in mind, with no knowledge of your task. The encoder is, in effect, a frozen hypothesis about which differences in the world are worth representing, and you inherit that hypothesis wholesale.
For a world model, it helps if nuisance changes have a small feature effect while controllable changes remain distinguishable and predictable. Invariance means the features are unchanged under a transformation; equivariance means they transform in a corresponding, structured way. Neither property is automatic from pretraining, and a search-based planner does not require every feature coordinate to change monotonically with an action. It needs a goal cost and dynamics model that preserve the task-relevant differences well enough to rank candidate actions.
In Part VIII, TD-MPC illustrated a representation shaped by task reward and value objectives. Frozen visual features take a different route: keep the representation fixed and fit the predictor in that space. Task-specific interaction data can favor adaptation; broad visual pretraining can make freezing attractive when such data are limited. Neither choice wins categorically. Encoder coverage, action data, dynamics, and the evaluation metric decide.
A useful first question is whether the task variable is visible in the available frames. Object location and scene layout may be; material compliance or hidden velocity may not be recoverable from one still image. Human ability to label a variable is not proof that DINOv2 retains it. Test pairs of states that differ in the variable, compare their features against nuisance variation, and use histories or other sensors when one frame is insufficient.
Is the feature vector actually Markov?
Often not. A Markov state makes the future independent of earlier history once the present state and the relevant actions are known. A single frame may leave velocity or an occluded object unresolved. A pendulum passing the same non-extreme angle with positive versus negative angular velocity looks the same in a still image but moves in opposite directions next. The one-frame embedding cannot resolve that ambiguity by itself.
There are three standard responses, all of which appear in practice:
- Frame stacking. Concatenate features from the last frames. This is a relatively simple way to expose short-term history; a short stack can sometimes disambiguate velocity and direction of motion.
- A recurrent or causal-attention predictor. Let the predictor consume a history of feature sets, so its internal state carries the missing information, as in Part V: World-Model Architectures. It can model interactions across a longer history, at the cost of more training and debugging complexity. For a fixed finite history, this is not categorically more expressive than a sufficiently capable model over stacked features.
- Accept the approximation. In a nearly memoryless, well-observed task, one frame may be sufficient in practice. The last action alone does not generally reveal velocity or an occluded object; measure the resulting prediction and control error rather than assuming a planner will tolerate it.
DINO-WM conditions its transition model on a short history of encoded frames and actions. Such context can reveal direction of motion without an explicit recurrent hidden state. Whether the chosen history is sufficient, and what it costs relative to another architecture, depends on the environment and should be measured.
Feature-Space Dynamics Prediction
Now that we have a state, we need dynamics. The predictor is a function
parametrized by learnable weights , where and are frozen encoder outputs and is the action. The encoder is a fixed map from pixels to vectors. Its weights do not receive gradients from the dynamics loss; the toy predictor below is trained with Adam gradient steps. The predictor operates in the coordinate system the encoder defines.
The training objective and why it is a proxy
The natural objective is plain regression on the encoder's output:
Notice what changed relative to a pixel-reconstruction objective. There is no decoder in this dynamics loss. The target is the next observation's frozen features, so the quality of the learned state depends on what the encoder retains. DINO-WM can train a separate, optional diagnostic decoder; the small example below also trains a decoder while pretraining its toy encoder. Neither decoder is used for latent rollout or planning.
This design offers three potential advantages:
- Dynamics training does not require pixel reconstruction. The transition predictor need not carry a pixel-generation head or optimize its loss. That does not make the encoder's pretraining free, nor does it automatically reallocate a fixed number of parameters.
- A pretrained feature target may suppress nuisance detail. Depending on the task and encoder, feature error can emphasize object-level changes more usefully than pixel error. Neither metric guarantees that it weights task-critical details correctly.
- The goal can use the same representation. Encoding a goal image gives a computable feature-distance cost without a separately trained reward model. Choosing that distance and verifying that it tracks success are still design and evaluation work.
There are two corresponding limitations:
- Feature-space distance is a proxy for task success. Two visually different futures can map to nearby features, and two visually similar futures can map to distant features. A representation trained for invariance may collapse distinctions that a controller cares about. Encoded history or actions can sometimes recover a state distinction absent from one frame. If a task-critical distinction is absent from every input available to the predictor, improving that predictor alone cannot restore it. Separately, a fixed goal cost built from aliased goal features cannot distinguish task-distinct goals; that requires more information or a different representation or cost.
- Latent predictions are not directly inspectable as images. A separately trained probe or diagnostic decoder can help, but its reconstruction quality must be checked; closed-loop behavior remains the decisive test.
A common practical refinement is to normalize features before regression. Squared differences in high-variance dimensions may dominate a raw Euclidean loss or planning distance; a large constant offset alone does not. One option is
where and are estimated on training data, with a nonzero floor on any per-dimension scale. Per-dimension scaling changes the relative weighting of feature errors and can amplify nearly constant noise. A single global scale changes optimization units but leaves Euclidean rankings between candidate plans unchanged; it does not balance dimensions. Whether high-variance dimensions carry task signal must be tested rather than assumed.
Architectures for the predictor
Three useful design choices illustrate the tradeoffs here. MLPs and transformers are possible predictor backbones; a residual or delta target is an output parameterization that can be used with either. Convolutional, recurrent, graph-based, and state-space predictors are further options. The backbone choice depends in part on spatial structure and the cost of modeling interactions between distant scene regions.
MLPs. A per-token MLP applies shared weights to each patch, optionally with the action broadcast to each token. It is cheap but cannot directly exchange information between patches. A global MLP instead flattens all patch tokens before prediction; it can mix locations, but its input size and weights are tied to a fixed grid and it lacks an explicit spatial-neighborhood bias.
Transformers over patch tokens. Self-attention can mix information across locations and frames. Some architectures support variable token counts, but changing resolution also changes positional structure and computational cost; it is not automatic resolution transfer. DINO-WM conditions a transformer on a short frame/action history, concatenating an action embedding to each patch feature. Its temporal masking prevents access to future frames while allowing within-frame interaction; it is not a rule that each token sees only earlier tokens.
Residual or delta predictors. Instead of predicting directly, predict the change:
When transitions are small in these coordinates, learning a delta can be convenient. The residual form alone does not initialize near zero: ordinary randomly initialized linear layers still produce nonzero outputs. Zero-initializing the last layer would create an identity starting point, but the toy model below does not do that. Persistence is a separate baseline to measure, not an automatic guarantee of residual training.
Action conditioning
The action has to enter the predictor somehow, and the choice matters whenever the dynamics are strongly action-dependent. A predictor that is insensitive to the action is useless for control, and a predictor that is over-sensitive to the wrong features of the action will chase spurious correlations. Common options include:
- Concatenate the action vector to the feature vector before the first linear layer. Simple and works well when the dynamics are globally action-linear-ish.
- Broadcast or concatenate an action embedding to every patch token, as in DINO-WM, so each token receives the command.
- Insert an action token in the sequence, letting attention mix action and visual information. This is an architectural alternative, not the DINO-WM mechanism just described.
- Concatenate a short history of actions, which helps when the visible effect of an action is delayed, such as when a command starts a process that only becomes visible a few frames later.
Across robots, actions differ in shape, units, timing, and physical effect. Learning a shared action representation is one research direction discussed in Part VI: Learning World Models at Scale, but an embedding alone cannot make embodiments interchangeable; calibration and embodiment-conditioned dynamics remain necessary.
Compounding error
Prediction error can accumulate over long learned-model rollouts. Understanding when that happens helps you choose horizons, budgets, and training procedures; an exact model or a contracting system is a counterexample to inevitable degradation.
Suppose one-step additive errors are independent, zero-mean vectors with root-mean-square norm , and the transition does not amplify perturbations. Then the root-mean-square accumulated error is at most under this nonexpansion assumption. Equality holds for an identity or norm-preserving transition that carries each error forward unchanged; contraction can give a smaller value. Correlation or an expanding transition Jacobian can increase error beyond this bound. Learned rollouts also feed predicted features back into the model, potentially moving off the training distribution. Neither the rate nor a sudden-failure threshold follows from one-step error alone.
Possible mitigations, with task-dependent effectiveness:
- Train on predicted states. Add multistep losses or mix model predictions with observed inputs during training; the latter requires a specified schedule rather than merely invoking “scheduled sampling.” This is distinct from Dreamer-style imagination, which trains actor/value components using a learned world model.
- Keep the rollout short and replan. Feedback can limit how long an inaccurate open-loop forecast is trusted. Later observations or encoded history may reveal state hidden in one frame, but replanning cannot restore a distinction absent from every available input or from an aliased fixed goal cost.
- Train with suitable input perturbations. Noise on inputs may help local robustness if it is paired with meaningful targets and evaluated; independent zero-mean noise on regression targets alone does not teach robustness to perturbed inputs.
- Ensembles. Predictor disagreement can expose some model uncertainty, as in PETS, but common-mode errors and representation failures can remain invisible.
An important subtlety: better one-step feature error does not always mean better closed-loop behavior. If the representation or goal cost is misspecified for the task, feature MSE can approach zero while control still fails. History or a later observation can sometimes disambiguate a one-frame state alias; fitting the same proxy more closely cannot recover information absent from all supplied inputs or distinguish goals collapsed by the fixed cost.
Action-Conditioned Latent Rollouts
Once the predictor exists, everything else is planning. A rollout is just recursive application:
At every step after the first, the input is the model's own previous output. This recursive, self-fed prediction is often called imagination. Errors can accumulate as it rolls forward, which is why the diagnostics in the code below matter.
Specifying the task with a goal image
The elegant part of the DINO-WM recipe is how it turns a rollout into a plan. There is no reward function and no value network. Instead, the task is specified by a goal observation , and its encoded form
becomes the target. The planning objective is
where is a distance in feature space, such as mean squared error over patch tokens, and is the feasible action set. This terminal-state cost matches the DINO-WM planning objective. The toy implementation below deliberately uses a sum of intermediate-state costs to reward progress throughout its short horizon; that is a different objective and can select a different plan.
Three design consequences follow.
A test goal can be specified by one image. The planner does not need a task-specific reward model or demonstration for that goal, but the feature distance is still a chosen cost and the dynamics model needs offline observation–action trajectories. Some DINO-WM training datasets replay noisy expert trajectories; “zero-shot” here concerns a new test goal, not an absence of behavioral data.
Patch-level costs retain spatial indexing. Comparing corresponding patches can expose where images differ, unlike one pooled vector. The cost gradient need not point toward a useful action: encoder invariances, the learned transition, and occlusions can all distort it. A global embedding can still support search if it preserves goal-relevant information, but it offers less direct spatial diagnosis.
The cost is only as good as the representation and distance. If the encoder is insensitive to a task-critical difference, the cost can miss it; if it responds strongly to irrelevant variation, the planner can pursue the wrong match. DINOv2's behavior under position, texture, and lighting changes must be tested in the intended domain rather than presumed.
One practical caveat: the goal image may not show the exact configuration the agent must reach. If the agent needs to grasp an object, an image of the object at a destination may not show grip quality or contact force. Some tasks need additional process or contact information; a goal trajectory, tactile feedback, explicit constraints, or a final-approach policy are possible ways to supply it. The limitation is structural: a single frame does not specify the path or interaction process by which the configuration is reached.
Searching for actions
With the cost defined, planning optimizes an action sequence. If the predictor and cost are differentiable, one can backpropagate through a rollout; Part VII: Planning and Agency covers that option. The DINO-WM experiments use a sampling-based planner, CEM. It accommodates clipped action candidates and does not require cost gradients, but finite samples can miss good plans too. Neither family has a general advantage for every learned cost.
The standard choice is the cross-entropy method (CEM), a derivative-free optimizer:
- Initialize a Gaussian distribution over action sequences of length .
- Sample candidate sequences from the distribution.
- Roll out the predictor for each candidate and evaluate the cost against .
- Keep the lowest-cost candidates (the elites).
- Refit and to the elite set.
- Repeat for a fixed number of iterations. Return .
The intuition behind CEM is that it repeatedly refits a sampling distribution to low-cost candidates, often concentrating search around promising action sequences without needing gradients. Its sampled variance need not shrink on every iteration.
Then use the result as model predictive control: execute only the first action (or the first few), observe the new frame, re-encode it, and replan. Frequent replanning limits open-loop exposure to model error, but cannot guarantee recovery from a poor representation, bad action coverage, or an unsafe intermediate action.
The horizon and sample count are important compute and accuracy controls. Longer horizons let the planner anticipate further, but later predictions may be less accurate; their quality must be measured. More samples cover more candidate sequences, with rollout cost roughly linear in for a fixed horizon and model. Elite count, iteration count, covariance handling, and the learned cost also matter. A longer horizon makes each candidate rollout more expensive, so horizon and sample count compete for a finite compute budget.
Model predictive control is the loop “plan steps, execute one, replan.” Feedback from new observations can correct some model error. Re-encoding an observed frame also replaces the imagined latent at the next planning cycle. Neither operation guarantees task success or cancels a systematic bias within the executed action.
Does the cost being minimized correspond to the thing you want?
Sometimes, and this is the crux of evaluation for this family of models. Feature-space distance to a goal image is a proxy, and optimization can exploit a misaligned proxy. Two failure patterns recur:
- Latent-space shortcuts. The planner may favor an action sequence whose predicted features look close to the goal but whose real outcome does not. Representation aliasing and model error are two different causes. Per-dimension whitening can amplify noisy small-variance directions, but it is not a necessary condition for failure.
- Proxy insensitivity. The cost may barely change across task-relevant configurations. A separate final-approach controller is one possible remedy, but task-specific validation is needed.
Evaluate closed-loop task success alongside feature prediction and goal costs. The toy experiment reports both kinds of measure, but agreement on a small sample does not establish that the proxy will rank plans correctly in new states or goals.
Generalization with Pretrained Encoders
The central hypothesis behind DINO-WM is that a frozen, general-purpose visual encoder can support useful transfer across goals and scene variation. Whether it outperforms a task-specific representation depends on the task, data, and baseline. It helps to separate the forms of generalization rather than treating them as one result.
Visual generalization. Pretraining on diverse images can help an encoder recognize useful structure outside the dynamics dataset. It does not guarantee stability under a new camera, lighting condition, wallpaper, or room. Measure both feature changes and closed-loop performance under the shifts that matter for deployment.
Task generalization. A goal image can change the target without retraining the dynamics model, provided the new goal is observable in the representation and reachable through actions covered by training. In the cursor example, different target positions fit that pattern. Arbitrary goals do not.
Structural information. Patch tokens retain positions in a grid, so corresponding-token costs can reflect local differences. That is not the same as an explicit keypoint matcher: a moving object may shift across token indices, and appearance or occlusion can break naive correspondence.
What does not come for free is equally important.
New dynamics are not free. A visual encoder alone says nothing about how actions change the scene. A predictor trained on rigid-object motion may need additional data or adaptation for cloth; whether the encoder remains useful is an empirical question.
New observation regimes are not free. Camera, resolution, field-of-view, and modality shifts can change features and transition statistics. A natural-image RGB backbone should not be assumed to work on depth, thermal, or fisheye imagery without testing or adaptation. This is the distribution-shift problem discussed in Part VI: Learning World Models at Scale, now at the encoder.
Encoder bias can be a ceiling. If a frozen encoder maps two task-distinct observations to exactly the same current features and supplies no distinguishing feature history or other input, a feature-only predictor cannot recover the difference. Near-invariance is less absolute: a head might amplify retained signal, while a new sensor, history, or backbone adaptation may be needed if the distinction is truly lost. Prediction loss alone may not expose this failure.
The practical consequence is that frozen-feature world models should be evaluated on closed-loop behavior, and qualitatively on whether the encoder's features respond to the variables your task cares about. A quick diagnostic is to collect a batch of states that differ only in one task-relevant factor, encode them, and look at how much the features move. If that movement is barely detectable relative to nuisance variation or prediction noise, a low training loss offers little reassurance about closed-loop performance.
Finally, “frozen” admits intermediate designs. A trainable adapter atop a fixed backbone can reweight information that the backbone retained, but cannot reconstruct information discarded completely. Whether to adapt the backbone itself depends on data, compute, and the task-relevant sensitivity tests above; there is no universal episode-count threshold.
A Worked Example: Reaching a Key from a Picture
Before writing code, let's walk through the whole pipeline on an example small enough to check by hand. This is the same structure the code implements, just with numbers chosen to make arithmetic possible. Working through the arithmetic by hand is worth the effort because it exposes exactly which assumptions the pipeline relies on and exactly where they break.
The world. A point cursor moves in the unit square. The state is the cursor position . The action is a displacement, and the dynamics are with , where is the per-coordinate action limit. Each frame shows a red disk at a fixed key position and a blue disk at the cursor. The task is to put the cursor on the key. The analytic calculation below assumes the action does not hit the clipping boundary.
The encoder. Suppose, for the sake of the example, that the encoder produces a two-number feature vector, one number for each cursor coordinate:
Real encoders are vastly more complex. This analytic feature map is invertible, so reaching zero feature distance is equivalent to reaching the physical goal. Invertibility alone does not preserve rankings of nonzero, weighted costs or guarantee that a finite-horizon planner finds the goal.
The dynamics. While clipping is inactive, the features evolve as
A residual predictor with recovers the interior transition exactly. A fitted can approach that matrix if the training set includes enough unclipped interior transitions to identify both action directions and noise is controlled; boundary-clipped samples do not obey this linear relation. One global linear map cannot represent clipping everywhere. This idealized encoder already supplies the two cursor coordinates, so this calculation does not establish that a real visual encoder will do the same.
The goal. The goal image is the picture with the cursor drawn on the key, so
The plan. The toy example uses the summed cost , unlike DINO-WM's terminal-state cost. If a feasible action sequence reaches zero feature residual, the analytic cursor is on the key. Neither invertibility nor this objective guarantees that a finite-horizon search finds that sequence, and nonzero plan rankings can differ from those under a physical-coordinate cost.
Limits of the analytic example. The rendered toy is rasterized and can map nearby physical states to identical frames before encoding; it can also occlude the red key beneath the cursor. Its learned encoder may discard more information, and its nonlinear dynamics model is only approximate. Finally, feature directions can receive unequal cost weights: if is small, the analytic cost penalizes horizontal error weakly. These are distinct failure mechanisms, not proof that every plan fails.
The code below builds this pipeline in a synthetic world: an encoder pretrained on frames and then frozen, a learned feature-space predictor, and a CEM planner tested on goal images it has never seen. The closed-loop test will show both progress and misses.
Code Implementation
We'll build the pipeline in five stages: the world, the frozen encoder, the feature-space dynamics model, the rollout diagnostics, and the planner. Every piece is small enough to read in one sitting, and every number in the output is computed at run time. Keeping the pieces separate makes it clear which part of the system is responsible for which behavior, which is exactly the clarity a full-scale system makes hard to achieve.
Imports and configuration
import matplotlib.pyplot as plt
import numpy as np
import torch
import torch.nn as nn
from book_plot_style import PALETTE, polish_axes, use_book_style
torch.manual_seed(0)
rng = np.random.default_rng(0)The synthetic world
The world is deliberately simple so that we can check the learned model against ground truth. A 32×32, three-channel image shows a red key disk and a blue cursor disk. The cursor moves by the commanded displacement inside a box. The frame shows approximate disk positions, but rasterization aliases nearby cursor locations and the blue disk can occlude the red key. Information loss can therefore occur in the renderer as well as the learned encoder; the diagnostics below do not isolate one cause perfectly.
IMG = 32
def render(cursor_xy, key_xy):
"""Paint a 32x32 RGB frame: red key disk plus blue cursor disk."""
yy, xx = np.mgrid[0:IMG, 0:IMG].astype(np.float32)
def disk(center, radius):
cx = float(center[0]) * (IMG - 1)
cy = float(center[1]) * (IMG - 1)
return ((xx - cx) ** 2 + (yy - cy) ** 2) <= radius**2
img = np.zeros((3, IMG, IMG), dtype=np.float32)
key_mask = disk(key_xy, 2.2)
img[0][key_mask] = 0.95
img[1][key_mask] = 0.25
img[2][key_mask] = 0.25
cursor_mask = disk(cursor_xy, 2.2)
img[0][cursor_mask] = 0.25
img[1][cursor_mask] = 0.45
img[2][cursor_mask] = 1.00
return img
def step(state, action, low=0.05, high=0.95):
"""Displacement dynamics inside a box; the rasterized frame is lossy."""
return np.clip(state + action, low, high)Before learning an encoder, check what the renderer itself can hide. With the cursor drawn over the key, two different key positions can produce exactly the same 32×32 frame:
occluding_cursor = np.array([0.50, 0.50], dtype=np.float32)
hidden_key_a = np.array([0.50, 0.50], dtype=np.float32)
hidden_key_b = np.array([0.51, 0.50], dtype=np.float32)
print(
"distinct key positions render identically:",
np.array_equal(
render(occluding_cursor, hidden_key_a),
render(occluding_cursor, hidden_key_b),
),
)distinct key positions render identically: True
No visual encoder can recover a distinction absent from identical input pixels. This is renderer occlusion and rasterization, not a failure introduced by frozen features.
Let's render one frame and look at it, to fix the scale of the problem: a disk of radius 2.2 pixels corresponds to about 7% of the image width.
frame shape: (3, 32, 32) pixel values in [0, 1]: 0.0 to 1.0 fraction of pixels that are background: 0.970703125
The reported background fraction shows why uniform pixel MSE can be a poor emphasis for this task. It does not by itself establish where another model spends its capacity.
Collecting unlabeled trajectories
We collect random-walk trajectories. Each episode gets a fresh key position and a fresh start, so the dataset covers a range of scene layouts. The first 32 episodes are the "training" split for the dynamics model; the last 8 are held out. Holding out entire episodes rather than random frames is important, because frames within an episode are highly correlated; a random-frame split would leak near-duplicate frames across the boundary and make the evaluation optimistic.
N_EPISODES = 40
EP_LEN = 100
N_TRAIN_EPISODES = 32
ACTION_SCALE = 0.06
MAX_ACTION = 0.08
frames, states, actions, keys, ep_ids = [], [], [], [], []
for ep in range(N_EPISODES):
key_xy = rng.uniform(0.15, 0.85, size=2)
state = rng.uniform(0.10, 0.90, size=2)
for _ in range(EP_LEN):
frames.append(render(state, key_xy))
states.append(state.copy())
keys.append(key_xy.copy())
ep_ids.append(ep)
action = np.clip(
rng.normal(scale=ACTION_SCALE, size=2), -MAX_ACTION, MAX_ACTION
)
actions.append(action)
state = step(state, action)
frames = np.stack(frames).astype(np.float32)
states = np.stack(states).astype(np.float32)
actions = np.stack(actions).astype(np.float32)
keys = np.stack(keys).astype(np.float32)
ep_ids = np.array(ep_ids)
train_episode_mask = ep_ids < N_TRAIN_EPISODESframes: (4000, 3, 32, 32) actions: (4000, 2) training episodes: 32, held-out episodes: 8 training frames: 3200 max displacement between consecutive frames: 0.080
Consecutive states differ by at most about 0.08 in each coordinate, roughly 2.5 pixels on this raster. That small physical step motivates trying a residual predictor. It does not prove that encoded features change smoothly at a patch boundary or that residual training beats a direct predictor; those are empirical questions for the diagnostics below.
A frozen patch encoder
This is a pedagogical substitute, not a DINOv2 reproduction. We pretrain a small convolutional encoder with a reconstruction objective and a foreground-weighted loss, then freeze it before dynamics training. DINOv2 uses a different pretraining objective and data scale; the shared feature of the two systems is only that a fixed visual encoder supplies the predictor's coordinates.
The encoder produces a 4×4 grid of 32-dimensional tokens, mirroring the patch-token structure of a ViT. The grid permits spatially indexed feature comparison for the goal-image cost; whether that cost tracks task success must be tested.
TOKEN_DIM = 32
N_TOKENS = 16
LATENT_DIM = TOKEN_DIM * N_TOKENS
class PatchEncoder(nn.Module):
"""Frozen visual backbone: image -> 4x4 grid of patch tokens."""
def __init__(self, token_dim=TOKEN_DIM):
super().__init__()
self.net = nn.Sequential(
nn.Conv2d(3, 32, 3, stride=2, padding=1),
nn.ReLU(),
nn.Conv2d(32, 64, 3, stride=2, padding=1),
nn.ReLU(),
nn.Conv2d(64, 64, 3, stride=2, padding=1),
nn.ReLU(),
nn.Conv2d(64, token_dim, 1),
)
def forward(self, x):
features = self.net(x) # (B, token_dim, 4, 4)
return features.flatten(2).transpose(1, 2) # (B, 16, token_dim)
class PatchDecoder(nn.Module):
"""Pretrains the encoder; excluded from dynamics and planning, retained for diagnostics."""
def __init__(self, token_dim=TOKEN_DIM):
super().__init__()
self.net = nn.Sequential(
nn.ConvTranspose2d(token_dim, 64, 4, stride=2, padding=1),
nn.ReLU(),
nn.ConvTranspose2d(64, 32, 4, stride=2, padding=1),
nn.ReLU(),
nn.ConvTranspose2d(32, 16, 4, stride=2, padding=1),
nn.ReLU(),
nn.Conv2d(16, 3, 3, padding=1),
nn.Sigmoid(),
)
def forward(self, tokens):
b, n, d = tokens.shape
side = int(round(n**0.5))
grid = tokens.transpose(1, 2).reshape(b, d, side, side)
return self.net(grid)Now we pretrain the encoder on the training episodes' frames, using no actions or goal/outcome labels. Roughly 97% of pixels are black background, so plain pixel MSE could favor a nearly black output. We threshold the rendered RGB frames to identify nonblack foreground pixels and upweight them. This is a hand-designed synthetic-domain foreground prior, not a separately supplied renderer disk mask or a DINOv2-style self-supervised objective.
imgs = torch.tensor(frames)
imgs_train = imgs[train_episode_mask]
encoder, decoder = PatchEncoder(), PatchDecoder()
ae_optimizer = torch.optim.Adam(
list(encoder.parameters()) + list(decoder.parameters()), lr=2e-3
)
AE_STEPS, AE_BATCH = 600, 64
ae_history = []
for _ in range(AE_STEPS):
batch = imgs_train[torch.randint(0, imgs_train.shape[0], (AE_BATCH,))]
reconstruction = decoder(encoder(batch))
foreground = (batch.amax(dim=1, keepdim=True) > 0.1).float()
pixel_weight = 1.0 + 20.0 * foreground
loss = (pixel_weight * (reconstruction - batch).square()).mean()
ae_optimizer.zero_grad()
loss.backward()
ae_optimizer.step()
ae_history.append(loss.item())Now freeze. From this point forward, no gradient ever flows into encoder.
encoder pretraining loss: 0.31562 -> 0.00231 frozen parameters in encoder: 58,400 tokens per frame: 16, token dim: 32, latent size: 512
The decoder remains in memory for the reconstruction diagnostic below, but is excluded from dynamics training and planning. We do not use it again after that diagnostic.




The common display scale lets us check whether the 512-number latent retains approximate disk positions, color, and intensity instead of merely inspecting a contrast-enhanced image. Even a visually plausible reconstruction would not establish that these features are sufficient for control; that is what the following rollout and planning tests examine.
Extracting and normalizing features
We encode every frame once and apply a global normalization. This changes the numerical scale of the regression loss; unlike per-dimension whitening, it does not change relative feature weights or candidate-plan rankings under Euclidean distance.
We estimate the normalization statistics from training frames only. Per-dimension whitening is an alternative, but it can amplify nearly constant noisy dimensions. With one scalar scale, high-variance directions retain their relative influence; whether that influence is useful requires a task-level check.
with torch.no_grad():
Z = encoder(imgs).numpy() # (N, 16, TOKEN_DIM)
Z_train = Z[train_episode_mask]
feature_mean = Z_train.reshape(-1, TOKEN_DIM).mean(axis=0)
feature_scale = (Z_train - feature_mean).std()
Z_norm = (Z - feature_mean) / feature_scale
Z_flat = Z_norm.reshape(len(Z_norm), LATENT_DIM)raw feature tensor: (4000, 16, 32) global feature scale: 1.3882 per-dimension std after normalization: min 0.587, median 0.885, max 1.982
The spread of per-dimension standard deviations after global scaling is a diagnostic of relative variation. Dominance in variation can affect a squared-error cost, but variance alone does not tell us whether the changing directions matter for the goal.
The cost is easier to inspect as a landscape. We render the cursor over a fixed grid with a fixed key, encode each frame, and evaluate feature distance to a goal image drawn at the key. Rasterization can make nearby positions identical, and the learned Conv–ReLU encoder is not known to be invertible or globally smooth. The observed landscape is a diagnostic for this trained toy, not a general theorem.

Building the transition dataset and learning dynamics
Transition pairs are consecutive frames within the same episode. Cross-episode pairs must be excluded, since they are not real transitions. This is the same leakage concern as the train/test split: joining the last frame of one episode to the first frame of the next would create a pair that no action could produce, and it would corrupt the learned dynamics.
same_episode = np.r_[ep_ids[1:] == ep_ids[:-1], False]
train_pairs = np.where(train_episode_mask & same_episode)[0]
Z_t = torch.tensor(Z_flat[train_pairs])
Z_next = torch.tensor(Z_flat[train_pairs + 1])
A_t = torch.tensor(actions[train_pairs])training transitions: 3,168 input dim: 514 (latent plus action), output dim: 512
The predictor is a residual MLP: it predicts a change in feature space rather than the next feature directly. The code does not zero-initialize its final layer, so the initial residual is not guaranteed to be small. We measure improvement against persistence below.
The action enters by concatenation. The physical update is action-affine only away from clipping boundaries; whether this MLP handles boundary transitions is empirical.
class LatentDynamics(nn.Module):
"""Residual predictor in normalized feature space."""
def __init__(self, latent_dim=LATENT_DIM, action_dim=2, hidden=256):
super().__init__()
self.net = nn.Sequential(
nn.Linear(latent_dim + action_dim, hidden),
nn.SiLU(),
nn.Linear(hidden, hidden),
nn.SiLU(),
nn.Linear(hidden, latent_dim),
)
def forward(self, z, a):
return z + self.net(torch.cat([z, a], dim=-1))
dynamics = LatentDynamics()
dyn_optimizer = torch.optim.Adam(dynamics.parameters(), lr=1e-3)
DYN_STEPS, DYN_BATCH = 800, 256
dyn_history = []
for _ in range(DYN_STEPS):
idx = torch.randint(0, Z_t.shape[0], (DYN_BATCH,))
prediction = dynamics(Z_t[idx], A_t[idx])
loss = nn.functional.mse_loss(prediction, Z_next[idx])
dyn_optimizer.zero_grad()
loss.backward()
dyn_optimizer.step()
dyn_history.append(loss.item())
dynamics.eval()dynamics loss: 0.31021 -> 0.16331 predicting no change at all would give: 0.28829 full-train fitted MSE: 0.16575 paired full-train reduction over persistence: 42.5%
The paired full-training-set comparison against persistence checks whether the predictor improves on a simple baseline for this one-step metric. The preceding last-minibatch loss is a training trace, not the numerator of that comparison. Held-out transition error and closed-loop task success are needed before calling the learned dynamics useful.
Rollout diagnostics: one-step versus recursive error
Now the key measurement. We roll the predictor forward on held-out episodes and compare two error curves:
- One-step error: always start from a true feature vector and predict one step ahead. This measures the model in the regime it was trained on.
- Recursive error: start from a true feature vector and then feed predictions back in. This measures the model in the regime the planner will actually use it.
- Persistence error: compare each future observation with the fixed starting feature. This is a reference curve, not a mathematical lower bound for a learned model or a policy.
HORIZON_EVAL = 15
STRIDE = 12
start_idx = []
for ep in range(N_TRAIN_EPISODES, N_EPISODES):
for t in range(0, EP_LEN - HORIZON_EVAL, STRIDE):
start_idx.append(ep * EP_LEN + t)
start_idx = np.array(start_idx)
z0_eval = torch.tensor(Z_flat[start_idx])
src_idx = start_idx[:, None] + np.arange(HORIZON_EVAL)[None, :]
targets_eval = torch.tensor(Z_flat[src_idx + 1])
with torch.no_grad():
rolling = []
z = z0_eval
for h in range(HORIZON_EVAL):
z = dynamics(z, torch.tensor(actions[src_idx[:, h]]))
rolling.append(z)
rollout_pred = torch.stack(rolling, dim=1)
one_step_pred = dynamics(
torch.tensor(Z_flat[src_idx]),
torch.tensor(actions[src_idx]),
)
recursive_mse = (
((rollout_pred - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)
one_step_mse = (
((one_step_pred - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)
persistence_curve = (
((z0_eval[:, None, :] - targets_eval) ** 2).mean(dim=-1).mean(dim=0).numpy()
)held-out start states: 64 one-step MSE at horizon 1: 0.33500 recursive MSE at horizon 1: 0.33500 recursive MSE at horizon 15: 2.22436 persistence MSE at horizon 15: 1.14052 recursive error growth factor (h=15 compared with h=1): 6.6x

The gap between the one-step and recursive curves shows how feedback of model predictions affects this held-out set. Where recursive MSE crosses the fixed-start persistence reference, the model predicts these recorded open-loop trajectories less accurately under this feature metric. That is not a comparison of action policies, a safe-horizon certificate, or a statement about every initial state.
This diagnostic motivates testing shorter horizons, alternative models, and closed-loop control. A CEM planner optimizes goal cost rather than the random-trajectory prediction MSE plotted here, so the crossing cannot establish that doing nothing is better or that search cannot help.
Planning to a goal image with the cross-entropy method
The toy planner searches over action sequences and sums feature distance to the encoded goal at every predicted step. This favors earlier approach to the goal, but it differs from DINO-WM's terminal-cost objective stated above.
def cem_plan(
z_start, z_goal, horizon=8, n_samples=128, n_elite=12, n_iters=3, seed=0
):
"""Cross-entropy method over action sequences, scored in feature space."""
plan_rng = np.random.default_rng(seed)
mean = np.zeros((horizon, 2), dtype=np.float32)
std = np.full((horizon, 2), MAX_ACTION, dtype=np.float32)
z0 = torch.tensor(z_start[None, :])
zg = torch.tensor(z_goal[None, :])
for _ in range(n_iters):
noise = plan_rng.standard_normal((n_samples, horizon, 2)).astype(
np.float32
)
candidates = np.clip(
mean[None] + std[None] * noise, -MAX_ACTION, MAX_ACTION
)
candidate_t = torch.tensor(candidates)
with torch.no_grad():
z = z0.repeat(n_samples, 1)
cost = torch.zeros(n_samples)
for h in range(horizon):
z = dynamics(z, candidate_t[:, h])
cost = cost + ((z - zg) ** 2).mean(dim=1)
elite = candidates[torch.argsort(cost)[:n_elite].numpy()]
mean = elite.mean(axis=0)
std = elite.std(axis=0) + 1e-3
return meanThe MPC loop re-encodes a rendered observation at every real step, so the next search starts from the encoder's output rather than the previous imagined feature. That observation is still lossy because of rasterization and possible occlusion.
def encode_state(state_xy, key_xy):
"""Render, encode with the frozen backbone, and normalize."""
frame = render(state_xy, key_xy)
with torch.no_grad():
z = encoder(torch.tensor(frame[None])).numpy()[0]
return ((z - feature_mean) / feature_scale).reshape(-1).astype(np.float32)
def run_mpc(key_xy, start_xy, n_steps=20, horizon=8, seed=0):
"""Receding-horizon control. Execute one action, observe, replan."""
goal_latent = encode_state(key_xy, key_xy)
state = np.asarray(start_xy, dtype=np.float32).copy()
trajectory = [state.copy()]
for k in range(n_steps):
current_latent = encode_state(state, key_xy)
plan = cem_plan(
current_latent, goal_latent, horizon=horizon, seed=seed + k
)
state = step(state, plan[0])
trajectory.append(state.copy())
return np.array(trajectory)Evaluating closed-loop success
This is where proxy and task outcomes can be compared. We track final physical distance from cursor to key and final observed-frame feature distance to the goal. The latter is related to, but not identical to, the summed predicted feature cost CEM optimized. A random-action baseline runs on the same scenarios with the same number of steps.
N_SCENARIOS = 10
eval_rng = np.random.default_rng(1234)
eval_keys = eval_rng.uniform(0.20, 0.80, size=(N_SCENARIOS, 2)).astype(
np.float32
)
eval_starts = eval_rng.uniform(0.15, 0.85, size=(N_SCENARIOS, 2)).astype(
np.float32
)
mpc_trajectories, mpc_final_distance = [], []
for i in range(N_SCENARIOS):
traj = run_mpc(
eval_keys[i], eval_starts[i], n_steps=20, horizon=8, seed=100 * i
)
mpc_trajectories.append(traj)
mpc_final_distance.append(float(np.linalg.norm(traj[-1] - eval_keys[i])))
mpc_final_distance = np.array(mpc_final_distance)
random_final_distance = []
for i in range(N_SCENARIOS):
state = eval_starts[i].copy()
for _ in range(20):
noise = np.clip(
eval_rng.normal(scale=ACTION_SCALE, size=2), -MAX_ACTION, MAX_ACTION
)
state = step(state, noise)
random_final_distance.append(float(np.linalg.norm(state - eval_keys[i])))
random_final_distance = np.array(random_final_distance)initial distance to key, mean: 0.397 latent MPC final distance, mean: 0.378 random walk final distance, mean: 0.396 latent MPC success rate (<= 0.10): 10% random success rate (<= 0.10): 0%
The toy model was trained without goal/outcome labels or demonstrations, but its encoder pretraining used a foreground mask derived by thresholding rendered RGB pixels. It does not consistently reach the key in these held-out scenarios. It sometimes approaches the target and sometimes stalls far away. The predictor learned from random-walk transitions in 32 episodes; the test goal was supplied as an image. Reporting misses alongside successes matters because feature-space fit is not task success.
For each held-out scenario we compare final physical distance with final observed-frame feature cost. This is a descriptive cross-scenario check: each point has a different key and start, and the plotted final-state cost is not the exact summed predicted cost optimized within that scenario. Rank inversions across different goals do not prove that CEM misranks candidate plans for one goal. A within-goal candidate study would be needed for that conclusion.






A note on what just happened
The encoder in this example is not DINOv2. It is a tiny convolutional network trained by foreground-weighted reconstruction on a few thousand 32×32 RGB images. What it shares with DINOv2 is a system role: a frozen visual map supplies features for dynamics learning and goal comparison. The toy uses a one-frame MLP, global normalization, and a summed-horizon CEM cost; a real-backbone implementation must revisit those choices rather than copy them uncritically.
Scaling up changes the data distribution, patch grid, feature dimension, action dynamics, and compute budget. A pretrained backbone may improve useful visual invariance, but that must be verified for the task. A Transformer is one way to model token interactions, not a requirement imposed by dimension alone. The toy remains useful for understanding the interfaces and diagnostics, not for predicting real-robot performance.
Limitations and Impact
DINO-WM combines a frozen pretrained visual encoder, action-conditioned feature prediction, and goal-image planning. Its dynamics objective and planner do not require a pixel decoder or a reward head. That does not mean arbitrary tasks are covered, that behavioral data contain no demonstrations, or that no diagnostic decoder can be trained. Compared with decision-centric methods such as Dreamer, MuZero, and TD-MPC, the distinguishing design choice is to use an externally pretrained visual representation as the fixed prediction and goal-comparison space.
The methodological division is useful: generic images can pretrain an encoder, while environment-specific observation–action trajectories train the transition model. The latter data are not the same as unlabeled video, and collecting them may be costly. This decomposition is one research option, not a universal template for later video generators or vision-language-action systems, some of which train representations and dynamics jointly.
The limitations are equally real, and they cluster into five groups.
Representation sensitivity matters. If the frozen encoder completely merges task-distinct observations and the predictor has no other input or history, downstream layers cannot reconstruct that distinction. A targeted paired-state probe, representation sensitivity analysis, and closed-loop evaluation can expose such failures. Small but nonzero retained differences might be amplified by a trainable head; complete collapse may require backbone adaptation, a new sensor, or temporal context. Loss curves alone are insufficient.
Feature distance is a proxy. The planning cost is a squared distance in a high-dimensional space—512 coordinates for the toy, and 98,304 raw patch-feature coordinates for the stated ViT-S/14 grid and width. Dimensionality alone does not prove an exploitable shortcut. Patch-level costs and short replanning horizons may help in some tasks, but neither guarantees alignment. Test within-goal candidate rankings and report closed-loop task outcomes.
Compounding error limits open-loop trust. In this toy run, recursive error grows over fifteen steps and crosses the fixed-start persistence reference between steps six and seven, while one-step error stays lower. That argues for testing shorter horizons here, not prescribing a universal horizon. Closed-loop trials determine whether replanning compensates for this open-loop error. A separately trained decoder or probe can aid inspection, subject to its own fidelity limits.
A goal image omits process constraints. A final configuration does not specify “move the block without knocking over the cup,” nor does it show an internal state invisible to the camera. Multiple images, a goal video, language, explicit constraints, or a learned task cost may supply more information; each adds its own modeling and validation burden.
Latent predictions need diagnostics. The transition model does not itself output video. A separate decoder or probe can visualize aspects of a predicted feature state, but neither is automatically cheap, faithful, or sufficient for safety assurance. In consequential deployments, validate the actual closed-loop behavior and safety constraints directly.
Frozen pretrained encoders can make goal-image planning practical without a pixel-reconstruction loss in the dynamics model. Their value relative to task-specific alternatives is empirical. Treat the encoder as a hypothesis about which distinctions matter, then test that hypothesis against the task before trusting the planner.
Summary
- Frozen visual representations as state. A fixed encoder such as DINOv2 supplies per-frame patch features; DINO-WM's transition model also uses short observation and action histories. Patch-level features preserve spatial indexing, but not guaranteed object correspondence.
- The trade. Pretraining can provide useful visual structure without task-specific encoder training, while freezing can hide distinctions needed for control. Test the representation in the intended domain.
- Feature-space dynamics prediction. The dynamics objective predicts next-step frozen features from action-conditioned context. It has no pixel-reconstruction term; an optional diagnostic decoder is separate. Feature MSE remains a proxy for control.
- Normalization. A single global scale changes numeric units without changing Euclidean plan rankings. Per-dimension scaling changes relative weights and can amplify noisy nearly constant dimensions.
- Action-conditioned rollouts. Recursive predictions can drift. Compare one-step, recursive, and fixed-start persistence errors on held-out trajectories, then test horizons in closed loop.
- Planning with goal images. DINO-WM's CEM planner uses terminal feature distance to an encoded goal; the toy uses a summed intermediate cost. A new test goal can be supplied as one image, but dynamics training still requires aligned observation–action trajectories.
- Generalization. Goal-image reuse applies only where the representation and learned dynamics retain the necessary information and action coverage. Scene, camera, modality, and dynamics shifts require empirical checks or adaptation.
- Evaluate closed-loop. Report task success and inspect feature sensitivity to task-critical variables; neither low feature MSE nor a cross-goal scatter plot proves plan ranking within a goal.
The next chapter turns to video generators, which can produce visible future frames rather than only latent predictions. Whether those frames are action-consistent enough to support planning is a separate empirical question.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about pretrained visual models and DINO-WM.
Pretrained Visual Models and DINO-WM
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!