Part of World Models Handbook
VLA and world-action models fuse vision, language, and control, tokenize robot actions, and transfer across embodiments under closed-loop evaluation.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Vision-Language-Action and World-Action Models
Imagine a robot arm on a cluttered desk and the instruction "put the red mug in the sink." Nothing in that sentence names a joint angle, a grasp point, or a trajectory. There is no coordinate, no pose, and no motor command in that instruction. Yet a competent human reading it would know what to look for, where the target should end up, and roughly how to move their hand. To act, the robot must find the red mug, ground "in the sink" in a partially observed scene, and emit motor commands that survive contact with a ceramic mug on a non-slip mat. Those are intertwined perception, language, and control problems. A common robotics design handled them in separate modules.
A modular pipeline might use a perception stack to produce object labels and poses, a task planner to reason over symbols, and a motion controller to track a trajectory. Its interfaces are chosen by designers and can omit task-relevant detail. A bounding box and class label may omit handle orientation needed for a grasp; a symbolic goal may need geometric refinement before control. Some systems detect such errors early, but an untested interface can allow a perception mistake to propagate through an otherwise sound plan. Debugging then requires tracing the failure across modules.
Vision-language-action (VLA) models learn a mapping from images and a language instruction, often with proprioception, to robot commands. They can reduce reliance on hand-designed perception-to-planning interfaces, though deployed systems still need controllers, safety checks, and embodiment-specific command interfaces. World-action models (WAMs) add a predictive component alongside action generation, such as a forecast of future observations or latent state. A predictor can support consequence-aware decisions, but merely training one does not mean the deployed policy plans with it.
A model that conditions robot-action generation on visual observations and a natural-language instruction, sometimes with depth or proprioception. Its components may be pretrained separately, frozen, or fine-tuned together; joint end-to-end training is an architectural choice, not a requirement of the term.
An emerging term for systems that couple action generation with explicit prediction of future states, observations, or video. The predictor may forecast the consequences of candidate actions or produce a goal-conditioned visual plan that a separate controller turns into actions. The term is not yet standardized: some papers use a joint generative video-and-action model, while others pair a policy with a predictive auxiliary head. A slower vision-language module and faster action module alone make a dual-rate VLA, not necessarily a world-action model; the predictive component is the distinguishing feature used in this chapter.
The two families are easiest to keep distinct if you hold on to the separation we have maintained since Part I: The World-Model Idea: a policy answers "what should I do?" An action-conditioned world model asks "what happens if I do this?"; a goal-conditioned video planner instead proposes what a desired future might look like. A plain VLA is a policy. A world-action system couples action generation to an explicit predictor: the components may share parameters or a training loss, or a separately trained controller may turn predicted futures into actions at deployment. Only an action-conditioned predictor directly compares candidate actions' consequences. The distinction determines what questions you can ask of the system, what failure modes you inherit, and whether prediction is used when the robot acts.
The physical body through which an agent acts and senses: its degrees of freedom, joint limits, link lengths, gripper or hand morphology, actuator dynamics, control frequency, and sensor suite. Two agents share an embodiment only if their action spaces and observation spaces are compatible in both dimensionality and physical meaning.
That vocabulary matters because this chapter sits at the junction of two promises that are often conflated. The first is generality of interface: specifying a supported robot task in language. The second is generality of knowledge: experience from one body helping control another. These are different claims with different evidence and failure modes. Language-conditioned behavior works on constrained tasks and setups, but a sentence alone does not guarantee a novel, reliable behavior. Cross-embodiment transfer has also been demonstrated under particular conditions; identifying those conditions is more useful than treating either promise as universal.
This chapter covers four interlocking topics. First, how and why vision, language, and control get fused, and what each coupling buys you. Second, how actions become tokens or continuous samples, because that decision quietly determines which architectures are even available to you. Third, how predictive models get paired with action models, including the objective-design problem of keeping both heads useful. Fourth, how models transfer across embodiments and how they should be evaluated once they are deployed in a closed loop. Each topic builds on the one before it, so the chapter is best read in order even though each section is self-contained enough to revisit.
Joint perception, language, and control
A short robot instruction such as "put the red mug in the sink" leaves the mug's location, the scene geometry, and the motion details unspecified. A controller must recover them from observations, memory, and learned priors. The engineering question is where that information is represented and how the pieces are joined. Designs range from separate modules exchanging symbols to models that fuse vision and language before producing actions. Their interpretability, data requirements, and capability depend on the implementation and task.
Why fuse the three modalities at all
One case for fusion is an information-bottleneck argument, related to the sufficient-state discussion in Part III: Representing Agents and Worlds. A symbolic modular pipeline may pass selected labels, poses, and predicates across interfaces while omitting image detail. Other modular systems pass images, maps, or uncertainty estimates, so the bottleneck depends on the interface. A bounding box may be sufficient for one grasp and omit the handle orientation needed for another. Targeted interface tests can expose such omissions before physical failure, though finding every task-relevant detail in advance remains difficult.
If the visual encoder is unfrozen and trained end-to-end with the policy, action loss can update its features. Errors on translucent objects, for example, may encourage features that distinguish useful optical cues, provided the training data expose that failure and the optimization can correct it. This is one meaning of joint training: action prediction can shape perception instead of relying only on an intermediate label objective. Frozen-encoder VLAs do not have this gradient path. Even with end-to-end training, better sample efficiency or worse interpretability relative to a modular baseline is an empirical question, not a consequence of fusion alone.
Language enters for a different reason. Natural language is a flexible, compositional interface for many task specifications. An instruction can name objects, spatial relations, and ordering constraints, and can qualify an action with words such as "carefully." Unlike a flat classifier with one label per task, it can express new combinations without adding a class for each one. Structured task languages can be compositional too, and an unfamiliar phrase still requires grounding and may require new training data. Part IV: Space, Physics, and Causality discusses how to test such combinations. Broad language pretraining may already include phrases such as "red mug," but that exposure does not establish that a robot can execute the corresponding instruction.
A formal view: language-conditioned control as a POMDP
The cleanest formalization reuses machinery we have already seen. Take a partially observable Markov decision process with state , camera observation , proprioceptive observation , action , transition kernel , and joint observation kernel . Add a fixed instruction drawn from a distribution over natural-language strings. A task-specific reward or success condition specifies what achieving that instruction means; the behavior-cloning objective below instead learns from demonstrations of that task. The task is to learn a policy
where:
- : the policy, parameterized by
- : the action at time , in the action space
- : the camera-observation history up to time , i.e. , each in
- : the proprioceptive-observation history up to time , where includes measurements such as joint angles and gripper width
- : the fixed natural-language instruction, shared across the episode. The instruction conditions the whole thing but never changes across an episode; it is a parameter of the task, not part of the state. Written as a token sequence of tokens, it is fixed for the whole episode and conditions every action.
Two limitations of the simplified policy setup are worth stating because their effects can be mistaken for model errors.
- Markov assumption. The written form admits history, but many VLA implementations condition on a limited set of frames plus the instruction. A policy without another memory input cannot recover a past event absent from that window, as discussed in Part III, Chapter 4: Temporal State and Memory. A short-context policy is not inherently defective; it is suited to tasks where its available observations carry enough information.
- Train–deployment distribution shift. A policy fitted to demonstrations may encounter different tasks, objects, and scenes at deployment. This is distinct from whether the POMDP's transition and observation kernels are stationary. Distribution shift also affects other supervised learners, but robotics can attach physical costs to the resulting errors.
Behavior cloning turns this into a supervised learning problem with a simple objective over a dataset of demonstrations:
where:
- : the dataset of expert demonstrations, each an observation-instruction-action tuple
- : a single demonstration drawn from the dataset
- : the conditional density of action for continuous actions, or probability mass for discrete action bins, given the observation, proprioception, and instruction
- : the negative log-likelihood loss, the expectation under the demonstration distribution
Everything distinctive about VLA architectures is downstream of how you represent the conditioning inputs and how you parameterize over a continuous, high-precision action space. The objective itself is unremarkable; what makes the problem interesting is the representation choice around it.
Input fusion and action output are separate choices
Two input-fusion choices and one action-output choice show how these models are assembled. They are design axes, not three mutually exclusive families: a model can fuse visual and language tokens and also emit actions as text tokens.
Late fusion with a shared embedding. Separate encoders process images and language, then their embeddings are concatenated and fed to an action head. This is a relatively simple design, though its data and compute costs depend on the encoders and which weights are trained. The visual encoder can be changed independently. Cross-modal interaction begins after each encoder has formed its features, so the visual features themselves cannot be conditioned on whether the instruction says "pick up the mug" or "avoid the mug." A sufficiently expressive downstream head may still distinguish the two tasks.
Token-level joint fusion. Visual features and language tokens enter a shared transformer context, sometimes with proprioceptive tokens. In OpenVLA specifically, DINOv2 and SigLIP visual features are fused, projected into the Llama 2 input-embedding space, and processed with text by the language-model backbone. This is joint self-attention over the resulting token sequence, not a dedicated cross-attention module at every layer, and OpenVLA does not require proprioception as an input. RT-2 also builds on pretrained vision-language backbones, but details differ across its variants. Token-level fusion permits task-dependent interaction between modalities at computational cost; an attention pattern by itself does not prove that a phrase is correctly grounded.
Action-as-text output. Independently of the input-fusion choice, a model can discretize each action dimension and produce action tokens through its language decoder rather than through a separate continuous-action head. RT-2 used existing integer tokens for action bins in its PaLI-X variant and repurposed rarely used tokens in its PaLM-E variant. In both cases, the output is decoded as an action when the model is prompted for robot control. This makes instruction understanding, visual question answering, and motor control next-token prediction over a shared backbone.
RT-2's combination of pretrained multimodal inputs and action-token output is worth dwelling on, because it supports behaviors absent from the robot demonstrations. When RT-2 was co-fine-tuned on web-scale vision-language data alongside robot trajectories, it retained vision-language capabilities that could inform control. The RT-2 paper tested instructions involving symbols, reasoning, and recognition. In its reported unseen-scenario evaluation, RT-2 reached 62% success versus RT-1's 32%; a separate Language Table simulation evaluation reported 90% for RT-2, so those figures are not endpoints of one comparison. The results are consistent with transfer from web pretraining through shared weights, although the evaluation alone does not isolate a single causal mechanism.
What the language backbone contributes
The pretrained vision-language backbone contributes three distinct things that are easy to conflate, and separating them clarifies why the same architectural choice helps some tasks and not others.
- Semantic priors. Object categories, attribute terms, common-sense relations, and some spatial relationships can be learned from broad image-text data. Such pretraining may help resolve "the fruit next to the bowl" in a new scene, but the target robot still needs to localize the actual objects and geometry.
- Broadly pretrained perceptual features. Encoders such as CLIP, SigLIP, and DINOv2 have seen much broader visual data than a small task-specific encoder trained on a few hundred demonstrations. Their features may help under variation in lighting, viewpoint, and clutter, but robustness relative to a task-specific baseline must be measured on the target setting. This is the motivation for the vision-only world models in Chapter 2: Pretrained Visual Models and DINO-WM, not a guarantee of superior control.
- A shared interface across tasks and embodiments. Language can express a new task without changing a fixed classifier vocabulary. Whether the policy executes that task still depends on its training coverage, perception, and action capabilities; new demonstrations or adaptation may be needed.
What the language backbone costs
The costs are less discussed but shape architecture decisions just as strongly, and in practice they are the constraints that architects argue about.
The first is spatial precision. Image encoders vary in patch size and training objective; an encoder optimized for broad semantic recognition may not preserve every detail needed for a particular grasp. The tolerable localization error depends on the object, gripper, and controller. Some systems combine visual pathways with complementary strengths. OpenVLA fuses DINOv2 features, which retain patch-level spatial structure, with language-aligned SigLIP features before projection into its language backbone. That design does not by itself prove that either pathway supplies sufficient grasp precision; the target task must test it.
The second is latency. At an illustrative decoding rate of 100 tokens per second, an autoregressive backbone emits one token every 10 milliseconds, already using an entire 100 Hz control period. A command with eight scalar action tokens takes about 80 milliseconds before prefill, communication, and execution overhead. We will spend a large part of the next section on action-tokenization tricks that reduce this delay, and a large part of the section after that on architectures that split the slow reasoner from the fast actor.
The third is fine-tuning instability. Fine-tuning a large pretrained backbone on a narrow robot dataset can degrade capabilities that made it valuable, a catastrophic-forgetting risk that task training loss may not expose. Parameter-efficient adaptation, such as low-rank adapters on a frozen backbone, and co-fine-tuning with a mixture of robot and pretraining data are two responses. Their trade-offs depend on the task and training regime rather than guaranteeing preservation of every pretrained capability.
Memory, history, and the limits of reactivity
A reactive VLA using one frame is memoryless with respect to its input. Its observation stream can be non-Markov even when the underlying POMDP state is Markov. Two consequences follow.
First, some tasks require information absent from the current frame. "Put the apple where the cup was" requires the earlier cup location after the cup moves. A strictly one-frame policy cannot recover that hidden event from its current image alone unless the information is available through another input or prior. The written form admits history; a short frame window helps while the relevant event remains within it, while an explicit memory module can retain information longer, as discussed in Part III, Chapter 4. Neither approach guarantees that the policy will use the information correctly.
Second, occlusion can make a single view ambiguous. A gripper may block a camera's view of the object during a grasp, while another camera might still see it. Recent frames can help recover what was visible before the occlusion, but a short window will not always disambiguate the scene. The benefit depends on camera placement, geometry, and how long the relevant information remains available.
The practical compromise used by many current systems is action chunking with a short observation window. The model sees the last one or two frames and emits a block of future actions. Chunking amortizes policy inference across several control steps, while the window supplies recent context for partially observed scenes.
Does the model use what it sees?
A question worth asking of any instruction-conditioned policy: is the reported success rate coming from visual grounding, or could the model be relying on familiar scene configurations and a weak instruction prior? Aggregate success can hide either mechanism.
Useful diagnostics are interventions rather than attention pictures alone. Shuffle images across otherwise matched trials and measure whether performance changes; replace the instruction with a conflicting one and check whether actions change in the intended direction. Attention visualizations can help locate candidate failure modes, but attention weights are not causal proof of visual grounding. A policy may also solve some benchmark tasks from proprioception and scene priors, so image and instruction ablations are needed before attributing success to either modality.
Action Tokenization and Robot Policies
The action space is where VLA models differ most from vision-language models, and where consequential engineering decisions get made. Language models emit discrete symbols from a finite vocabulary. Robot actions are continuous vectors, often 7 to 30 dimensional, sampled at 10 to 100 Hz. Bridging that gap is the subject of this section. Representation choice determines the loss function, inference procedure, achievable command precision, and action distributions the model can learn.
Why actions are the awkward modality
Digital images and tokenized text have discrete input encodings, though pixel bit depths vary and a word need not correspond to one token. Robot actions are often represented as continuous vectors with physical units. An autoregressive token decoder needs a discrete action encoding; a transformer with a continuous output head does not. Three common action-output choices are:
- Discretize each dimension into bins and predict a categorical distribution, optionally folding the bins into the language vocabulary as tokens.
- Regress directly, with a mean-squared-error or L1 loss on a continuous output head.
- Generate by sampling from a learned distribution over continuous actions, using diffusion or flow matching.
These choices are not exhaustive: joint categorical distributions, learned codebooks, mixtures, and other heads are possible. The selected output representation shapes the training objective, inference cost, and which action distributions the model can express.
Uniform binning and its arithmetic
The simplest scheme maps each action dimension from its operating range into equally spaced bins, then predicts a categorical distribution over bins. Encoding is
where:
- : the continuous value of the -th action dimension, in the operating range
- : the lower bound of the operating range for dimension
- : the upper bound of the operating range for dimension
- : the number of equally spaced reconstruction values, with and
- : the integer bin index assigned to , in
For the displayed endpoint-grid encoding, decode index as . Adjacent reconstruction values are apart, so rounding gives a maximum absolute error of for inputs inside the stated range. A distinct midpoint-bin convention partitions the range into intervals and decodes to each interval's center; its maximum error is . These conventions require different encoders and must not be mixed. For and , the endpoint grid has spacing and maximum error normalized units. If one normalized unit corresponds to one radian, that bound is about 3.92 milliradians; whether it is below a controller's effective resolution depends on the hardware.
Two properties of this scheme deserve attention.
The first is that the distribution of actions depends on the dataset and control convention. If demonstrations concentrate near zero or near saturation, uniform bins may spend resolution on rarely used regions; another dataset may use the range more evenly. Inspect the target distribution before choosing a grid. Per-dimension normalization using training-set statistics or percentiles can reduce scale mismatch but does not make the distribution uniform. Nonuniform bins can allocate more values to dense regions, at the cost of a less simple encoding and decoder.
The second is that discretization is not a lossless stand-in for regression. With a fixed endpoint grid, each decoded command must be one of its reconstruction values, with the per-command quantization bound derived above. That does not impose the same bound on final position: feedback, repeated relative commands, and dynamics can yield intermediate endpoints. Tasks needing finer individual commands can use more bins, a smaller action scale, or a continuous head.
From bins to tokens: the RT-2 recipe
Turning bins into language tokens requires one extra decision: which tokens. Repurposing tokens needed for ordinary text can interfere with the model's language behavior, so the mapping depends on the pretrained tokenizer and on how action outputs are decoded.
RT-2 used different mappings for its two backbones. PaLI-X already had tokens for the integers 0–999, so its action bins reused those numeric tokens. PaLM-E lacked that convenient numeric representation, so RT-2 overwrote its 256 least frequently used tokens for the action vocabulary. During robot-action decoding, RT-2 restricts sampling to valid action tokens; on ordinary vision-language tasks, the full language vocabulary remains available. The task context and decoding rule, not a universally private token namespace, determine how an output is interpreted.
The elegance comes with a cost, and the cost is sequence length. The token stream that carries the actions is much longer than the geometric action chunk it represents, and that length is what determines inference latency.
The token budget and why chunking exists
Consider a seven-joint arm with one additional gripper command, controlled at 50 Hz. Each action vector has eight scalars, so one second of behavior contains 400 scalars; with one token per scalar, that is 400 tokens per second of behavior. A seven-billion-parameter autoregressive model may decode on the order of 50 to 150 tokens per second in a particular serving setup, depending strongly on hardware, precision, batching, and implementation. At those assumed rates, generating one second of behavior takes about 2.7 to 8 seconds. This setup therefore cannot sustain 50 Hz through one-token-at-a-time decoding. Faster hardware can narrow the gap, but it does not remove the sequential decoding steps required by this representation.
Predicting a block of future actions instead of a single action . A non-autoregressive head can emit a block in one forward pass. Open-loop execution queries the policy once per chunk; ACT-style temporal ensembling instead queries it every control step and combines overlapping predictions. Chunking therefore does not, by itself, guarantee a lower query frequency. ACT introduced this approach for manipulation with around 50 to 100.
Action chunking is the first fix. If a non-autoregressive head emits a 50-step chunk in one call, the controller can query it once per second instead of fifty times. An autoregressive language decoder still needs a decoding step for each token: chunking alone leaves the 400-token count per second of behavior unchanged. It can reduce policy or control-interface calls, but it does not remove the underlying next-token decoding cost.
The second fix attacks the token count directly. Smooth robot trajectories often concentrate more of their energy in low-frequency coefficients along the time axis. Quantizing those coefficients and then applying byte-pair encoding to the resulting symbol stream can compress an action chunk. We use an illustrative 400-scalar-to-30-token ratio in the worked example below, rather than treating that ratio as universal. At 30 tokens per second of behavior, the representation fits within the assumed 50 to 150 token-per-second decoding throughput. FAST exploits this temporal regularity in a way analogous to how image compression exploits spatial regularity.
The third option uses a continuous action head rather than per-scalar tokens. A regression head can emit a chunk in one evaluation. Diffusion and flow heads generate a chunk through an iterative solver whose evaluation count depends on the model and solver; some implementations use about ten steps. These heads avoid per-scalar autoregressive decoding, but only stochastic generative heads model multimodality without an added mixture mechanism.
Continuous action heads: regression, diffusion, flow
A regression head adds a linear layer mapping the backbone's pooled representation to an action vector, trained with or loss. It is the cheapest option and the most limited. It has no mechanism to express uncertainty, no way to represent more than one answer per input, and no generative sampling procedure.
The limitation is mode averaging. Demonstration data can be multimodal: a teleoperator may go around an obstacle on either side, and both are correct. For an unconstrained predictor minimizing expected squared error, the optimum is the conditional mean, which may be valid in neither mode. For scalar loss, any conditional median is optimal; it may lie in a dominant mode, but for equally weighted, separated modes a median can also lie between them. Either deterministic regression head still represents only one action per input. Vector-valued loss yields coordinatewise medians, which need not form a feasible joint action; the coordinatewise medians of a joint distribution can lie in a region the joint distribution assigns essentially no probability.
Diffusion and flow-matching heads can address mode averaging by learning a conditional distribution from which multiple action chunks can be sampled. They do not guarantee that the fitted distribution captures every valid mode. Flow matching, used in the family, defines an interpolant between noise and a target action chunk ,
where:
- : a noise sample drawn from a standard normal, , with the same shape as the action chunk
- : the target action chunk from the demonstration data
- : the interpolation time, running from (pure noise) to (clean action)
- : the interpolated sample at time , the input the velocity field is asked to explain
The backbone context encodes the observation and instruction. A velocity field is trained to regress the straight-line target velocity .
where:
- : the velocity field, parameterized by , evaluated at the interpolated sample at time with context
- : the constant velocity along the straight-line interpolation from noise to data, which is what the field is regressed toward
- : the backbone's context, which conditions the field on the observation and instruction
- : the learnable parameters of the velocity field
A taxonomy of action heads and when to use each
| Head | Training objective | Multimodal | Inference cost | When it fits |
|---|---|---|---|---|
| regression | MSE on the action vector | No | One forward pass | Well-behaved, near-unimodal demonstrations; tight latency budgets |
| regression | MAE on the action vector | No (median) | One forward pass | Same, but with outliers or noisy labels |
| Factorized discrete bins | Per-dimension cross-entropy | Per dimension; joint modes may mix | One forward pass | You want a cheap categorical head and can test joint-action consistency |
| Vocabulary tokens | Next-token cross-entropy | Yes | Many autoregressive steps | You want one model for reasoning and control at low frequency |
| Diffusion | Denoising score matching | Yes | Solver-dependent iterative evaluations | Tasks with distinct action strategies when iterative cost fits the control loop |
| Flow matching | Conditional flow matching | Yes | Solver-dependent iterative evaluations | Same; compare quality and latency at the chosen step count |
The taxonomy in Table action-heads uses a factorized per-dimension categorical head. Its dimensions are sampled independently, so a bimodal task can combine the upper mode with the lower mode and produce a joint action absent from demonstrations. Correct marginals do not guarantee a valid joint sample. A joint or autoregressive categorical model need not have that independence, but it pays for dependency modeling in complexity or decoding steps. Chunk-level tokenizers such as FAST offer another way to encode cross-time and cross-dimension structure; which representation works better requires task-specific evaluation.
Normalization, relative actions, and action-space invariants
Three implementation choices can affect whether a heterogeneous VLA trains reliably. They are easy to miss in a high-level architecture diagram and worth checking explicitly.
Per-dimension normalization. Action dimensions can have different scales and units. Shoulder rotation in radians and gripper opening in meters can share an optimizer, but unnormalized MSE gives the numerically larger dimension much more weight: compare ranges and . Many VLA pipelines therefore normalize action dimensions using training-set statistics or quantiles; the chosen convention must be saved and inverted at execution time.
Absolute versus relative actions. Absolute joint positions depend on a specific kinematic calibration. Relative commands, such as an end-effector pose delta plus gripper command, can be easier to reuse across similar arms, but their physical effect still depends on the body and controller. A controller can apply a delta using the current measured pose; the policy need not integrate every previous delta internally if proprioception provides that pose.
Padding and masking. Heterogeneous datasets may have different action dimensionalities. One approach pads actions to a shared maximum width and masks padded entries in the loss; another maps commands into a canonical action space, as in RT-X. If padded entries are left unmasked, the loss also trains the model to predict filler values that are not real commands. The choice must be documented alongside action units and decoding.
Predictive Models Paired with Action Models
Everything to this point describes a policy: it maps observations and instructions to actions. A world-action system also includes a predictive component, but it need not be a second head on the policy. Prediction may be an auxiliary training target, a separately generated visual plan used by a controller, or a joint action-and-future output. This section examines those couplings and why prediction and control are not the same objective.
Why predict the future at all
There are four distinct motivations, and separating them prevents confusion about what a predictive component is for.
As an auxiliary representation-learning objective. Predicting future states encourages the encoder to retain information about dynamics that pure action imitation might not need. This is the motivation for the predictive objectives in Part VI: Learning World Models at Scale: learning rewards features that help predict transitions on the training distribution, while an imitation head can rely on features correlated with an expert's next move. Predictive loss does not guarantee that the model identifies every causal factor, but the distinction matters when deployment shifts the state distribution.
To exploit action-free video. Human videos and teleoperation recordings without reliable action labels are easier to collect than labeled robot trajectories. Observation-only video prediction can learn from those frames directly; an action-conditioned dynamics head needs recorded or inferred actions. VPT illustrates the latter route in Minecraft: an inverse-dynamics model trained on a smaller labeled set assigns pseudo-actions to a large unlabeled video set, which then supervises a behavior-cloned policy. It is an example of obtaining action supervision from video, not of training an action-conditioned dynamics head on frames alone.
To plan. If you can roll the model forward under candidate actions, you can select actions by predicting their consequences rather than merely imitating. For a one-step lookahead with discount factor , select : the immediate reward is undiscounted, while the estimated continuation value is discounted. Repeating this choice after each observation is a short-horizon version of the model-predictive control loop from Part VII: Planning and Agency.
To verify and to generate data. A sufficiently calibrated model could help screen candidate actions before execution, though a prediction is not a safety guarantee. A generative model can also produce synthetic experience: NVIDIA's DreamGen pipeline generates "neural trajectories" from a video world model and uses them to augment robot-policy training, an instance of the synthetic-data approach in Part VI, Chapter 4: Data Curation and Synthetic Data. Both uses are limited by fidelity, especially away from the training distribution.
Four couplings of a predictor to a policy
The literature contains at least four distinct ways to attach a predictor to an actor, and they fail differently.
Auxiliary head on a shared trunk. One encoder can feed action and prediction heads. GR-1 predicts future video frames and actions through a GPT-style transformer; the prediction loss can shape the shared representation. The objectives may compete for capacity or gradients. Loss weighting is one way to manage that tradeoff; selective gradient flow, separate training schedules, or separate modules are other design choices.
Text-and-image-conditioned video planner with a separate inverse dynamics model. UniPi conditions its video generator on the initial image and a text description of the current goal, then uses an inverse dynamics model to extract actions from the planned frames. The video-generation stage need not take actions as input; the inverse model supplies the embodiment-specific control interface. Its two-stage error chain remains important: the visual plan must be feasible, and the inverse model must turn consecutive planned frames into workable commands.
One transformer for action and future-image generation. WorldVLA trains one autoregressive transformer on action generation and action-conditioned future-image generation. Under an ordinary causal mask, later action tokens in a chunk can attend to earlier predicted action tokens, so an early action error can propagate directly through the chunk. WorldVLA's action-generation mask blocks that access: each chunk action depends on the text and visual inputs rather than preceding generated actions, allowing the chunk actions to be produced in parallel. Its world-model branch still uses causal attention to generate future images conditioned on actions. The failure mode resembles the compounding-error argument in Part IV, Chapter 5: Stochasticity and Uncertainty, but it is action-to-action feedback within a chunk, not an intermediate generated image corrupting the next action.
Latent-action generative models. Instead of predicting in a robot's motor-command space, infer a discrete latent action code from consecutive frames and model the future in that code space. Genie uses this mechanism; the planned Chapter 4 in this part will develop it further. Its learned codes are not tied to named actuator commands, but this does not establish that the same code transfers between robot bodies or can be decoded into their controls.
Objective design: the loss-weighting problem
When one network carries both a policy loss and a prediction loss, their relative weight is a hyperparameter with first-order consequences.
If the predictive term dominates, the trunk may prioritize visual fidelity over control-relevant distinctions. Photorealistic reconstruction does not by itself establish that the representation is sufficient for choosing actions. Some texture or material cues matter for grasping; the concern is spending capacity on details irrelevant to the target task. This motivates latent predictive and energy-based objectives, as discussed in Part V, Chapter 5: Predictive Embeddings, JEPA, and Energy-Based Models.
If the action term dominates, the shared trunk may learn features useful for imitation but weak for prediction. With behavior cloning, both losses can be expectations under a demonstration dataset, yet their numerical scales and gradient magnitudes are not automatically comparable. A cross-entropy action loss and an MSE prediction loss live on different scales, and their ratio can drift over training. On-policy training can also change the state distribution used by the policy objective, shifting the effective balance even if the explicit loss weights stay fixed.
There is no universal resolution, but there are three useful defaults. First, weight the prediction loss by its eventual role: planning requires predictive accuracy over the action candidates and horizons used, whereas an auxiliary prediction head may need less weight if it improves the policy's representation. Tune this weight against separate control and prediction metrics rather than assuming a fixed small value. Second, predict in a learned latent space rather than pixel space when task-irrelevant visual detail would dominate the objective. Third, evaluate the two heads separately. A single aggregate training loss can hide a strong head compensating for a weak one.
What makes a prediction useful for control
This is the most important conceptual point in the chapter, so it is worth stating bluntly: accurate future prediction is neither necessary nor sufficient for good control.
It is not necessary because a reactive policy without an explicit predictor can succeed at many manipulation tasks. It is not sufficient because low average prediction error on observed or demonstrated transitions does not establish accuracy for alternative actions a planner might consider. An action-conditioned model may reproduce likely next frames on the data distribution yet rank counterfactual action rollouts incorrectly, which is the failure that matters when using it to compare actions. A goal-conditioned visual planner needs a different test: whether its proposed future is feasible and its inverse controller can realize it.
One key diagnostic for planning is counterfactual action sensitivity: does the predicted future change in the correct direction, by roughly the correct amount, when the action changes? A model that predicts an accurate future for the demonstrated action but predicts the same future for a counterfactual action with a different true outcome cannot reliably compare those actions, no matter how low its one-step prediction error is. This test helps establish whether an action-conditioned video generator is useful as a world model; action ranking and calibration must also be checked for the intended decision. The distinction is about tested capability, not the name of the architecture. It is related to the criteria in Part I, Chapter 3: The World-Model Qualification Test.
One diagnostic is to take a held-out state and compare model rollouts under a recorded action and a perturbed action. Checking whether the true futures diverge in the same way additionally requires a resettable simulator, repeated controlled trials, or another way to observe both interventions from comparable states; an ordinary logged transition supplies only one outcome. If the model predicts nearly identical futures for actions whose true outcomes differ, its action conditioning may be too weak for planning, even if its average one-step prediction error is low.
A second criterion is value equivalence in the relevant region. If the model is used for planning, the ranking of candidate actions can matter more than pixel-perfect rollouts. A model that underestimates distance traveled might still rank "approach the mug" above "retreat from the mug" in a simple unobstructed task; around obstacles, contact thresholds, or deadlines, the same bias can reverse the preferred plan. This is the insight behind value-equivalent models such as MuZero, which the planned Part VIII chapter addresses. Accuracy and usefulness come apart when the downstream consumer is a decision rather than a reproduction.
Dual-system architectures: separating the slow reasoner from the fast actor
One response to the latency tension is to split perception and action generation across two update rates. Scene interpretation may not need to rerun at every motor-control step, but the appropriate rates depend on the task and how quickly the scene changes.
A slower vision-language module can supply embeddings to a higher-rate action module. In NVIDIA's GR00T N1, a vision-language model runs at about 10 Hz and provides intermediate embeddings to a diffusion-transformer action module running at about 120 Hz, trained with flow matching over action chunks. Those embeddings are learned features, not guaranteed explicit "what" and "where" tokens. For the split to work, they must carry enough task-relevant information between updates while the faster module responds to current state inputs.
The decomposition can reduce how often the large vision-language module runs. It also suggests useful debugging questions: did the slower representation miss a scene change, or did the action module respond poorly to a sound representation? Those causes can interact, especially when modules are trained together, so the split does not make failures automatically attributable to one timescale or guarantee easier training.
The cost is that System 1 operates on context that may become stale between System 2 updates. Fresh proprioception and frequent action-chunk replanning can reduce this risk at the fast timescale. Event-triggered System 2 refresh when a visual scene changes is a possible design, not a mechanism established by the GR00T N1 architecture cited here; without such a mechanism, the slower context is refreshed on its scheduled cadence.
Using the model: imagination, data amplification, and evaluation
Once you have a predictive model paired with a policy, three applications open up.
Imagination. Train policy updates on rollouts inside a learned model, as in Part VIII, Chapter 3: The Dreamer Family. Dreamer still collects new environment experience to update that model; imagined learning does not turn the whole process into an offline computation. A policy optimized against model predictions can exploit favorable model errors. Shorter rollouts, conservative objectives, or model disagreement checks may reduce this risk in some settings; the planned Part VIII chapter on offline and conservative model-based RL will examine these choices.
Data amplification. Rather than optimizing a policy through imagined rewards, generate trajectories with a model and add them to supervised training. This changes the training distribution and therefore the fitted objective; it does not automatically sidestep model exploitation or synthetic-data bias. If a generated trajectory violates a physical constraint and is treated as valid supervision, it can teach the wrong action. Filtering, weighting, and real-world validation are ways to limit that risk.
Evaluation in imagination. Use the model as a cheap approximate simulator to screen policies before spending real robot time. Real-robot evaluation is a major bottleneck in robot-policy research, but an imagined evaluation inherits the model-fidelity limitations described above. Part XI: Evaluation and Understanding treats it as a measurement-instrument problem: the question is whether model errors distort the policy comparisons you are trying to make.
Failure modes: drift, exploitation, and the horizon mismatch
Three failure modes are worth checking where the architecture uses a predictive model for rollout or optimization.
Compounding drift. One-step errors can accumulate when predictions are fed back as inputs, especially when rollouts leave the training distribution. Local sensitivity of the learned dynamics is one contributor, alongside model bias, action choice, and observation noise; no single Lipschitz statistic establishes long-horizon planning quality. Re-anchoring on a reliable observation interrupts the particular chain of feeding predicted state back into the next prediction, but it does not guarantee a global error bound.
Model exploitation. A planner or policy optimizing against an imperfect model can favor actions where its errors are optimistic. This can create a systematic failure that average prediction error hides, although exploitation is not inevitable in every run. Evaluate model quality on the states and actions the optimized policy visits.
Horizon mismatch. The horizon over which the model is accurate can be shorter than the task horizon. A model useful for five-step prediction can still support a fifty-step task through repeated observation and replanning; it cannot justify an uncorrected fifty-step imagined rollout. Recognizing the mismatch tells you whether to replan more often, improve the model, or use a hierarchical decomposition.
Embodiment Transfer and Closed-Loop Evaluation
A VLA trained on one robot arm and deployed on another tests a central claim of generality, and the claim is more conditional than headlines suggest. This section covers what varies across embodiments, which transfer strategies have evidence, and how to evaluate the result.
What varies across embodiments
The variation is wider than dimensionality. Two robots with the same number of joints can still differ substantially.
- Action space semantics. Joint positions command a specific kinematic chain. End-effector deltas command a pose. Either way the meaning of a dimension depends on the body. The same numerical vector means different physical motions on different arms.
- Degrees of freedom and workspace. A 6-DoF arm with a parallel gripper, a 7-DoF arm with a dexterous hand, a mobile manipulator, and a humanoid occupy different dimensionalities and different reachable sets.
- Observation space. Camera count, placement, intrinsics, resolution, frame rate, and field of view all differ. A wrist camera and a third-person camera see fundamentally different views of the same action.
- Dynamics. Mass, compliance, friction, backlash, and actuator bandwidth differ even between two units of the same model. Two nominally identical robots can behave differently under load.
- Control frequency and latency. A 3 Hz position-commanded arm and a 100 Hz torque-commanded arm cannot execute the same action chunk.
- Task distribution. The tasks a given embodiment can perform differ because the physics differs. A humanoid can step over obstacles; a fixed arm cannot.
Four transfer strategies
Canonical action space with dataset-specific normalization. The RT-X experiments on Open X-Embodiment select one canonical camera view per dataset, resize it, and map original actions into a common seven-dimensional end-effector command, with dataset-specific action normalization. This makes pooled training tractable, but a canonical vector does not erase differences in control semantics or camera geometry. Padding and loss masking are other possible implementation tools for heterogeneous action dimensions; they are not the defining RT-X convention.
Embodiment-specific adapters, stems, and heads. A shared trunk can connect to interfaces that differ by embodiment. GR00T N1 uses embodiment-specific MLP encoders and decoders for state and action while sharing its vision-language pathway. Such interfaces accommodate different state and action dimensions, but they do not guarantee that the common representation is free of embodiment-specific effects.
Language as a shared task interface with a reusable action tokenizer. Normalize action chunks by dataset-specific statistics, then train a tokenizer that can be reused across several robot datasets. FAST+ was trained on about one million real-robot action sequences and, in the reported evaluations, closely matched dataset-specific FAST tokenizers rather than consistently outperforming them. Reusability suggests shared temporal regularities, but it does not make the resulting motor tokens embodiment-independent controls.
Latent action spaces. Infer discrete latent actions from transitions between frames, without requiring motor labels in the video corpus. Genie shows that this can support an interactive visual world model. It does not make every video a training example for every robot: latent codes need not correspond to feasible commands for a particular body, and mapping them to actuators requires additional data and validation. That is why Part VI, Chapter 2: Actions, Inverse Dynamics, and Latent Actions treats latent action discovery as a research frontier rather than a solved transfer technique.
Evidence: what transfers and what does not
The Open X-Embodiment collaboration released more than a million trajectories from 22 robot embodiments. Its RT-X robot experiments used a smaller subset of nine embodiments. The results were mixed rather than uniformly positive: RT-1-X improved four of five smaller-data evaluations in the paper's comparison, while its pooled model underperformed the corresponding single-robot baselines on the Bridge and Google Robot evaluations shown in Table I. The study establishes that positive cross-embodiment transfer is possible, not that pooling always helps.
The conditions matter for interpreting those comparisons. Action semantics, camera viewpoints, task mix, and data imbalance can all affect whether pooled training helps. The observed losses on some single-robot evaluations are a reason to test transfer per target embodiment, rather than infer it from aggregate performance alone.
Task and scene diversity are plausible reasons to pool data, but the Open X-Embodiment comparisons do not isolate a universal exchange rate between diversity and episode count. Report the target robot, task mix, and amount of target-domain data alongside any transfer claim.
Why one-step action error is the wrong metric
This is the single most important evaluation lesson in the chapter, and it connects directly to the general argument in Part XI: Evaluation and Understanding.
Held-out action-prediction loss is a cheap diagnostic for behavior cloning, and sometimes a useful preliminary filter. It is an incomplete proxy for closed-loop task success for four reasons.
Compounding. A small per-step error can accumulate. If an action error produces a two-centimeter position bias in the same direction on every step under simple additive dynamics, the displacement error after one hundred steps is two meters. If the per-step displacement errors are independent, zero-mean, and have the same finite variance , their signed sum has expectation zero and standard deviation after steps. With unequal variances, the standard deviation is the square root of their sum instead. The biased and zero-mean policies can have the same one-step action MSE because squaring discards error sign; one-step MSE also omits the closed-loop state distribution induced by the policy.
The distribution shift is created by the policy. Offline validation measures error on states the demonstrator visited. Online, the policy visits states its own errors produced. Small errors early can put the policy in states rare or absent in the training distribution, where performance may degrade. This is the standard argument for on-policy correction methods such as DAgger, discussed in Part VI, Chapter 5: Online, Offline, and Continual Learning.
Loss functions and success can be non-monotonic. A policy that commits to one mode of a bimodal distribution may have worse MSE than one that averages the modes, while succeeding more often when the mean action is invalid. That outcome depends on the task: a mean action is not always unsafe or unsuccessful. The metric can reward behavior that fails at some contact-rich tasks.
The relevant errors may be rare but severe. A deployment may assign much more cost to a collision than to a small action error. Its utility can therefore be risk-sensitive rather than an average over trials; the exact aggregation depends on the application. Mean per-step action error does not express that choice.
The practical implication is to use action-prediction error for debugging or screening, not as a substitute for closed-loop selection on the target task. Closed-loop evaluation is more expensive, but an offline ranking alone can select the worse controller and overstate confidence.
Closed-loop evaluation protocols
A defensible protocol has four properties.
Match the holdout to the generalization claim. Splitting episodes of the same task can test new trajectories, starts, or objects within that task. It does not test transfer to an unseen task family. For that stronger claim, hold out tasks or task families, and specify which combinations of objects, scenes, and instructions are new. The held-out conditions should resemble the deployment variation you care about.
Evaluate generalization along named axes. Prepare explicit test sets for unseen object instances, unseen language phrasings, unseen spatial arrangements, unseen lighting, and perturbed initial conditions. Report a separate number for each axis. An aggregate success rate over a mixture of easy and hard conditions hides more than it reveals. A model that is strong on known objects and weak on novel phrasings has a specific weakness, and the named axis makes it visible.
Measure recovery, not just reach. A policy can succeed from a clean start yet fail after a small perturbation. Whether that rules out deployment depends on the operating envelope and the cost of failure. Apply a disturbance mid-episode, or start from a slightly off-distribution state, and report recovery separately from nominal success. In the code below, the perturbed-start evaluation samples each of four initial state coordinates from a Gaussian perturbation of the nominal start with standard deviation .
Report uncertainty. Robot evaluation is high-variance with small samples. Ten trials per condition gives a binomial standard error of roughly 16 percentage points at a 50% success rate. With ten binary trials per policy, observed success rates move in 10-point steps: one extra success is weak evidence of a true improvement. Run enough trials for the intended comparison, or report confidence intervals and acknowledge the uncertainty. The statistical point separates a measured improvement from sampling noise.
Simulation-based and distributed evaluation
Because real-robot evaluation is the bottleneck, two complementary approaches have emerged.
Simulation with matched appearance. Build simulated scenes and camera views that approximate the real setup, then test whether simulated policy performance predicts real performance. SimplerEnv evaluates this with both Pearson correlation and mean maximum rank violation (MMRV), so it examines score agreement as well as ranking failures. A simulation useful for screening policies need not match every absolute success rate, but its predictive value has to be established for the policies and tasks being compared.
Distributed pairwise real-robot evaluation. RoboArena distributes double-blind pairwise policy comparisons across sites. Its published evaluation used the same DROID robot platform at seven sites, not many robot types. Pairwise judgments can make relative comparisons more scalable across varied setups, but they still require consistent trial protocols and judgment criteria; distribution alone does not remove those requirements.
Matched simulation can screen policies before target-robot trials. Distributed pairwise evaluation is itself real-robot testing; it can broaden comparisons, but neither approach alone establishes performance on the target tasks and deployment conditions.
A practical checklist
Before claiming a VLA or world-action model works:
- Is the task split held out at the task-family level?
- Is perturbation recovery reported separately from clean-start success?
- Is the instruction necessary? Rerun with shuffled instructions and check for degradation.
- Is vision necessary? Rerun with shuffled images and check for degradation.
- Is the reported model-selection metric closed-loop, or is it action MSE in disguise?
- How many trials, and what is the confidence interval?
- What is the inference latency, and does it fit the control loop the robot runs?
- If an action-conditioned predictor is present, does its prediction respond correctly to counterfactual actions?
- If a goal-conditioned visual planner is present, are its planned states feasible, and can the inverse controller realize them?
Worked Example: The Arithmetic of Action Tokenization
Before writing any code, it is worth doing the arithmetic that drives real architecture choices. Take a 7-DoF arm with a 1-DoF gripper, controlled at 50 Hz, using 1-second action chunks.
Latency budget from the model. The illustrative throughput range of 50 to 150 tokens per second describes a sequential autoregressive decoder. A non-autoregressive head can emit an action chunk in one forward pass. FAST reduces the sequence length of action chunks, but when its tokens are generated by an autoregressive VLA, they still use that model's next-token decoder; fewer tokens, not a different decoding mechanism, produce the latency benefit.
Resolution. Normalize each dimension to and use endpoint-grid reconstruction values as in the equation above. Adjacent values are apart, so the worst-case error for an in-range action is normalized units. A separate midpoint-bin encoder gives spacing and worst-case error . If one normalized unit represents one radian, the endpoint-grid bound is about 3.92 milliradians. Whether discretization is a bottleneck depends on the controller, action semantics, and task tolerance.
Token count with language tokens. One second at 50 Hz contains 50 action vectors. With eight scalar dimensions per vector and one token per scalar, that is tokens, ignoring delimiters or auxiliary tokens. At an illustrative decoding throughput of 50 to 150 tokens per second, emitting the chunk takes about 2.7 to 8 seconds.
The comparison is easier to read than to recite, so let us put the required token rates and the decoding capacity on the same axes. The figure below plots required tokens per second of behavior for three representations, with the illustrative decoding band shaded to show which schemes sit inside the feasible region.
## Deterministic arithmetic only; this preparation cell has no imports and no random state.
## Control rate, chunk length, and action dimensionality are taken from the worked
## example above. The decoding throughput is the illustrative 50 to 150 tokens per
## second used throughout this section, not a measured hardware number.
CONTROL_HZ = 50
ACTION_DIMS = 8
DECODE_TOKENS_PER_SECOND_LOW = 50
DECODE_TOKENS_PER_SECOND_HIGH = 150
tokens_per_scalar_second = (
CONTROL_HZ * ACTION_DIMS
) # 400 tokens per second of behavior
tokens_fast_per_second = 30 # illustrative FAST-compressed token rate
decode_low = DECODE_TOKENS_PER_SECOND_LOW
decode_high = DECODE_TOKENS_PER_SECOND_HIGH
token_scheme_labels = [
"one token per scalar",
"chunked, uncompressed",
"frequency-domain (FAST)",
]
token_scheme_rates = [
tokens_per_scalar_second,
tokens_per_scalar_second,
tokens_fast_per_second,
]
Chunking alone. Replanning once per chunk rather than once per physical control step reduces policy-query frequency from 50 Hz to 1 Hz. It does not reduce the autoregressive decoding steps needed to emit 400 tokens. Under the illustrative throughput above, generation still takes about 2.7 to 8 seconds, longer than the one-second chunk. Chunking alone is insufficient in that serving setup.
Frequency-domain tokenization. FAST applies a discrete cosine transform along the time axis, quantizes coefficients, and compresses the result into tokens. Smooth trajectories often concentrate energy in low-frequency coefficients, reducing sequence length.
Token counts become intuition only once they are converted into time, so the second figure restates the same arithmetic as latency against the control period.
## Pure-Python arithmetic, matching the worked example above. Bars span the low and
## high ends of the illustrative 50 to 150 token per second decoding throughput; the
## chunk period is one second because chunks are decoded once per second of behavior.
CONTROL_HZ = 50
ACTION_DIMS = 8
DECODE_LOW = 50
DECODE_HIGH = 150
chunk_period_seconds = 1.0
tokens_uncompressed = (
CONTROL_HZ * ACTION_DIMS
) # 400 tokens per second of behavior
tokens_fast = 30 # illustrative compressed rate for this worked example
latency_labels = ["one token per scalar", "frequency-domain (FAST)"]
latency_low = [tokens_uncompressed / DECODE_HIGH, tokens_fast / DECODE_HIGH]
latency_high = [tokens_uncompressed / DECODE_LOW, tokens_fast / DECODE_LOW]
latency_mid = [
0.5 * (low + high) for low, high in zip(latency_low, latency_high)
]
latency_err = [
0.5 * (high - low) for low, high in zip(latency_low, latency_high)
]
## The uncompressed and compressed streams differ by more than an order of magnitude
## in this latency view, which makes the gap hard to read on the same axes. Switch to a
## nonlinear vertical scale so both the slow and fast generations remain visible.
LATENCY_SCALE = "log"
Read against the dashed line, the uncompressed stream cannot produce a chunk as fast as the chunk is consumed, whereas the compressed stream finishes well before its one-second budget expires. If a particular 400-scalar chunk compressed to 30 tokens, its required average output rate would be 30 tokens per second of behavior; at 50 to 150 decoded tokens per second, generation would take about 0.2 to 0.6 seconds, excluding encoding, prefill, communication, and execution overhead. Actual compression and latency must be measured.
Continuous heads. Alternatively, train a flow-matching head that produces the whole chunk in about ten network evaluations on a small transformer, with no autoregression at all. The cost is a separate generative module and a training objective that no longer shares the language model's next-token machinery.
The flow-matching policy and its FAST-tokenized counterpart illustrate the two approaches. The arithmetic showing that chunking alone leaves autoregressive per-token decoding work unchanged explains why these architectures seek either parallel action generation or fewer action tokens.
Code Implementation: A Miniature World-Action Model
We will now build a complete, tiny example on a language-conditioned reaching task. The task is deliberately simple, because the point is to observe the mechanisms, not to reproduce a benchmark: a language token selects one of four goals, and a world-action model must both pick actions and predict the consequences of actions. Along the way we will see action quantization, mode averaging, compounding prediction error, and the gap between one-step error and closed-loop success. Every mechanism we isolate here has a scaled-up counterpart in the real systems discussed above.
Simulating language-conditioned demonstrations
We start with a two-dimensional point mass with velocity damping, which is a linear system we can reason about exactly. The state is , and the action is a velocity command.
import matplotlib.pyplot as plt
import numpy as np
import torch
import torch.nn as nnDT = 0.1
DAMP = 0.9
START = np.array([-1.0, 0.0])
GOALS = np.array([[1.0, -0.6], [1.0, -0.2], [1.0, 0.2], [1.0, 0.6]])
COLORS = ["red", "blue", "green", "yellow"]
VOCAB = {"reach": 0, "red": 1, "blue": 2, "green": 3, "yellow": 4}
SUCCESS_RADIUS = 0.15
def step(state, action):
"""Double integrator with velocity damping: s' = f(s, a)."""
velocity = DAMP * state[2:] + (1.0 - DAMP) * action
position = state[:2] + DT * velocity
return np.concatenate([position, velocity])
def expert_action(state, goal, kp=3.0, kv=3.0, v_max=1.0):
"""Cascaded controller: clip a desired velocity, then track it."""
desired = np.clip(kp * (goal - state[:2]), -v_max, v_max)
return np.clip(kv * (desired - state[2:]), -1.0, 1.0)The dynamics are deterministic and linear, so this network has enough capacity to represent the transition rule. A learned fit can still have optimization and finite-data error, while free-running use adds the distinct problem of feeding model-predicted states back into the input. The simple system lets us separate those effects more clearly than a contact-rich robot would.
Next we collect demonstrations with feedback. At each realized state, the controller recomputes its command before execution noise is added. The recorded state-action-next-state triples therefore follow the noisy trajectory the controller sees, rather than replaying commands from a clean trajectory open-loop.
def rollout(action_fn, steps, start=None):
state = (
np.concatenate([START, np.zeros(2)])
if start is None
else np.asarray(start, dtype=float).copy()
)
states = [state.copy()]
for _ in range(steps):
state = step(state, np.clip(action_fn(state), -1.0, 1.0))
states.append(state.copy())
return np.array(states)
def instruction_tokens(color):
return np.array([VOCAB["reach"], VOCAB[color]], dtype=np.int64)rng = np.random.default_rng(0)
N_EPISODES = 400
STEPS_PER_EPISODE = 60
ACTION_NOISE = 0.35
state_log, token_log, action_log, next_log, episode_log = [], [], [], [], []
for episode in range(N_EPISODES):
color = COLORS[episode % len(COLORS)]
goal = GOALS[COLORS.index(color)]
state = np.concatenate([START, np.zeros(2)])
for t in range(STEPS_PER_EPISODE):
commanded = expert_action(state, goal)
executed = np.clip(
commanded + rng.normal(0.0, ACTION_NOISE, size=2), -1.0, 1.0
)
nxt = step(state, executed)
state_log.append(state.copy())
token_log.append(instruction_tokens(color))
action_log.append(executed)
next_log.append(nxt)
episode_log.append(episode)
state = nxt
states = np.array(state_log)
tokens = np.array(token_log)
actions = np.array(action_log)
next_states = np.array(next_log)
episodes = np.array(episode_log)
DEMO_EPISODES_PER_COLOR = 5
demo_examples = {}
for index, color in enumerate(COLORS):
paths = []
for episode in range(index, N_EPISODES, len(COLORS)):
if len(paths) == DEMO_EPISODES_PER_COLOR:
break
block = states[
episode * STEPS_PER_EPISODE : (episode + 1) * STEPS_PER_EPISODE
]
paths.append(block[:, :2])
demo_examples[color] = paths
final_positions = next_states.reshape(N_EPISODES, STEPS_PER_EPISODE, 4)[
:, -1, :2
]
final_goal_distance = np.linalg.norm(
final_positions - GOALS[np.arange(N_EPISODES) % len(GOALS)], axis=1
)
The demonstration set has 24,000 transitions. The fraction within the goal tolerance uses the final state reached after noisy, feedback-corrected commands. Every episode starts from the same state, so the recorded paths cover only part of the state space.
The final positions come from the recorded next state after each episode's last noisy action. The following summary reports distance to the instructed goal for those realized trajectories.
transitions: 24000 state dim: 4 action dim: 2 instruction vocabulary size: 5 expert episodes settling within 0.15: 100.0% mean expert settling distance: 0.0173
The four instruction colors map to four separated goals, and the spread within each color shows the effect of action noise and feedback correction. The trajectories still cover a narrow region around the demonstrated tasks, not the full state space. The policy has no training examples for most states outside those paths.
The action-space choice: regression versus discrete bins
Before training anything, let us quantify the cost of discretizing actions and the cost of regressing them.
BIN_COUNTS = np.array([4, 8, 16, 32, 64, 128, 256, 512, 1024])
worst_case_error = 1.0 / (BIN_COUNTS - 1)
empirical_rms = []
for bins in BIN_COUNTS:
index = np.clip(np.round((actions + 1.0) / 2.0 * (bins - 1)), 0, bins - 1)
decoded = index / (bins - 1) * 2.0 - 1.0
empirical_rms.append(float(np.sqrt(((decoded - actions) ** 2).mean())))
empirical_rms = np.array(empirical_rms)
Two bin counts show the scale of the tradeoff. At the RMS error is around 0.0171 normalized action units, and at it is around 0.0021, roughly eight times smaller. These are action-space errors, not position errors; translating them into a millimeter tolerance requires the action semantics and the closed-loop controller.
Discretization also changes what the policy can represent when demonstrations contain two equally valid action strategies.
bimodal_rng = np.random.default_rng(3)
mode_upper = np.clip(bimodal_rng.normal(0.55, 0.08, 700), -1.0, 1.0)
mode_lower = np.clip(bimodal_rng.normal(-0.55, 0.08, 400), -1.0, 1.0)
bimodal_samples = np.concatenate([mode_upper, mode_lower])
bimodal_mean = float(bimodal_samples.mean())
bimodal_median = float(np.median(bimodal_samples))
bimodal_edges = np.linspace(-1.0, 1.0, 25)
bimodal_counts, _ = np.histogram(bimodal_samples, bins=bimodal_edges)
bimodal_centers = 0.5 * (bimodal_edges[:-1] + bimodal_edges[1:])
The mean sits in the empty valley between the two strategies, an action nobody ever demonstrated. A categorical head over bins, by construction, can place probability mass on both peaks, and a diffusion or flow head can sample from either. That is the strongest argument for tokenized or generative action heads, and it has nothing to do with precision.
A miniature world-action model
Now we build the model. A shared state encoder feeds two paths: the policy combines its state features with instruction tokens to choose an action, while the dynamics path combines the same state features with an action to predict the next state. Both losses update the shared encoder, making the predictive objective capable of changing the policy representation.
class WorldActionModel(nn.Module):
"""A language-conditioned policy plus an action-conditioned dynamics model."""
def __init__(self, vocab_size, n_bins=None, d_model=64):
super().__init__()
self.n_bins = n_bins
self.instruction = nn.Embedding(vocab_size, d_model)
self.state_encoder = nn.Sequential(
nn.Linear(4, d_model),
nn.ReLU(),
nn.Linear(d_model, d_model),
nn.ReLU(),
)
self.policy_trunk = nn.Sequential(
nn.Linear(d_model, d_model), nn.ReLU()
)
self.dynamics_trunk = nn.Sequential(
nn.Linear(d_model + 2, d_model),
nn.ReLU(),
nn.Linear(d_model, d_model),
nn.ReLU(),
)
self.action_head = nn.Linear(
d_model, 2 if n_bins is None else 2 * n_bins
)
self.next_state_head = nn.Linear(d_model, 4)
def policy_output(self, states, token_ids):
embedded = self.instruction(token_ids).mean(dim=1)
hidden = self.policy_trunk(self.state_encoder(states)) + embedded
out = self.action_head(hidden)
return out if self.n_bins is None else out.view(-1, 2, self.n_bins)
def next_state(self, states, action):
return self.next_state_head(
self.dynamics_trunk(
torch.cat([self.state_encoder(states), action], dim=-1)
)
)Two design notes. The instruction embedding is added to the policy-path output rather than concatenated. policy_output returns continuous actions or per-dimension logits shaped (batch, dim, n_bins), depending on whether n_bins is set. next_state predicts the state at from the state at and the action at ; it does not receive the instruction. The shared state encoder is the only component updated by both losses, so an auxiliary prediction loss can influence the policy without making the two output heads identical.
The split below holds out 80 episodes, but both sides retain the same four goals, instruction tokens, and nominal start. Its offline metrics test new trajectories within this synthetic task, not transfer to an unseen task family. The perturbed-start rollout later tests a different, local form of generalization.
TRAIN_EPISODES = 320
BATCH = 256
N_BINS = 64
train_mask = episodes < TRAIN_EPISODES
test_mask = ~train_mask
train_states = torch.tensor(states[train_mask], dtype=torch.float32)
train_tokens = torch.tensor(tokens[train_mask], dtype=torch.long)
train_actions = torch.tensor(actions[train_mask], dtype=torch.float32)
train_next_states = torch.tensor(next_states[train_mask], dtype=torch.float32)
test_states = torch.tensor(states[test_mask], dtype=torch.float32)
test_tokens = torch.tensor(tokens[test_mask], dtype=torch.long)
test_actions = torch.tensor(actions[test_mask], dtype=torch.float32)
test_next_states = torch.tensor(next_states[test_mask], dtype=torch.float32)def quantize(action, n_bins):
return torch.clamp(
((action + 1.0) / 2.0 * (n_bins - 1)).round().long(), 0, n_bins - 1
)
def decode_bins(bins, n_bins):
return bins.float() / (n_bins - 1) * 2.0 - 1.0
def train_model(variant, epochs=30, aux_weight=1.0, seed=0):
torch.manual_seed(seed)
model = WorldActionModel(
vocab_size=len(VOCAB),
n_bins=None if variant == "continuous" else N_BINS,
)
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
history = []
for _ in range(epochs):
epoch_loss = 0.0
for batch in torch.randperm(train_states.shape[0]).split(BATCH):
s, tok = train_states[batch], train_tokens[batch]
a, ns = train_actions[batch], train_next_states[batch]
out = model.policy_output(s, tok)
if variant == "continuous":
action_loss = ((out - a) ** 2).mean()
else:
target = quantize(a, N_BINS)
action_loss = nn.functional.cross_entropy(
out.reshape(-1, N_BINS), target.reshape(-1)
)
dynamics_loss = ((model.next_state(s, a) - ns) ** 2).mean()
loss = action_loss + aux_weight * dynamics_loss
optimizer.zero_grad()
loss.backward()
optimizer.step()
epoch_loss += loss.detach().item() * batch.numel()
history.append(epoch_loss / train_states.shape[0])
return model, history
continuous_model, continuous_history = train_model("continuous")
token_model, token_history = train_model("tokenized")The loss has an action term and a next-state term. aux_weight controls how strongly the next-state error updates the shared state encoder, as well as the separate dynamics path. Each call to train_model initializes and trains a separate model with the same dataset, optimizer settings, and seed, but a different action-output parameterization and loss. An observed performance difference therefore cannot be attributed to action representation alone.
variant action MSE next-state MSE continuous 0.106577 0.00005836 tokenized 0.165972 0.00009249
The two independently trained models both achieve low next-state error, but their dynamics heads are separate and their next-state MSEs differ slightly. The continuous policy also has lower one-step action MSE in this run. Neither metric alone tells us how often either policy will reach its goal when its own actions determine the states it visits. The closed-loop experiment below measures that outcome directly; neither policy consults its learned dynamics head during those rollouts.
We can also run the counterfactual action-sensitivity check described earlier. At one held-out state, reset the toy transition to the same starting point, apply two different actions, and compare the true change in next state with the learned predictor's change. This is a local diagnostic, not evidence that the predictor ranks every useful action correctly.
probe_state = test_states[0].numpy()
probe_actions = np.array([[0.0, 0.0], [0.8, -0.8]], dtype=np.float32)
probe_states = np.repeat(probe_state[None, :], 2, axis=0)
true_next = np.stack([step(probe_state, action) for action in probe_actions])
with torch.no_grad():
predicted_next = continuous_model.next_state(
torch.tensor(probe_states), torch.tensor(probe_actions)
).numpy()
print("true action-effect delta:", np.round(true_next[1] - true_next[0], 4))
print(
"predicted action-effect delta:",
np.round(predicted_next[1] - predicted_next[0], 4),
)true action-effect delta: [ 0.008 -0.008 0.08 -0.08 ] predicted action-effect delta: [ 0.0018 -0.001 0.0632 -0.0792]
The true action-effect delta is ; the fitted model predicts approximately . It captures the velocity-change direction and approximate magnitude, but understates the position change at this state. This one comparison can reveal weak action conditioning at the probe state; it does not measure error across other states or longer counterfactual rollouts.
Closed-loop evaluation in the true environment
Now we roll the policies out in the true environment, including from perturbed initial states so that recovery behavior is tested rather than just nominal tracking. A trial counts as a success if the policy reaches the goal tolerance at least once during its final 20 states; unlike the demonstration-set summary, this is not a final-state settling criterion.
def make_policy_fn(model, token_tensor):
def act(state):
with torch.no_grad():
out = model.policy_output(
torch.tensor(state, dtype=torch.float32).unsqueeze(0),
token_tensor.unsqueeze(0),
)
if model.n_bins is None:
return out.squeeze(0).numpy()
return decode_bins(out.argmax(dim=-1), model.n_bins).squeeze(0).numpy()
return act
def evaluate_policy(
action_fn_factory, trials_per_goal=25, steps=120, start_noise=0.10, seed=11
):
eval_rng = np.random.default_rng(seed)
successes = []
for index, color in enumerate(COLORS):
act = action_fn_factory(index, color)
hits = 0
for _ in range(trials_per_goal):
start = np.concatenate([START, np.zeros(2)]) + eval_rng.normal(
0.0, start_noise, size=4
)
path = rollout(act, steps, start=start)
distance = np.linalg.norm(
path[-20:, :2] - GOALS[index], axis=1
).min()
hits += float(distance < SUCCESS_RADIUS)
successes.append(hits / trials_per_goal)
return float(np.mean(successes))success_rates = {}
success_rates["expert"] = evaluate_policy(
lambda index, color: lambda s: expert_action(s, GOALS[index]), seed=11
)
for name, model, seed in [
("continuous policy", continuous_model, 11),
("tokenized policy", token_model, 11),
]:
success_rates[name] = evaluate_policy(
lambda index, color, model=model: make_policy_fn(
model, torch.tensor(instruction_tokens(color), dtype=torch.long)
),
seed=seed,
)
def wilson_interval(rate, trials, z=1.96):
denom = 1.0 + z**2 / trials
center = (rate + z**2 / (2.0 * trials)) / denom
half_width = (
z
/ denom
* np.sqrt(rate * (1.0 - rate) / trials + z**2 / (4.0 * trials**2))
)
return max(0.0, center - half_width), min(1.0, center + half_width)
success_intervals = {
name: wilson_interval(rate, 25 * len(COLORS))
for name, rate in success_rates.items()
}policy closed-loop success expert 100% continuous policy 100% tokenized policy 84% One-step action MSE (lower is better): continuous < tokenized Closed-loop success (higher is better): continuous > tokenized
Read that output carefully. The one-step MSE ranking and the closed-loop success ranking need not agree in general, because closed-loop success measures whether errors compound or self-correct as a policy visits states induced by its own actions. In this seeded run, the continuous and tokenized policies achieve 100% and 84% success, respectively, across 100 matched perturbed starts each. That is a 16-percentage-point observed gap, not an equivalence result. The separate model fits, one seed, and small synthetic task do not establish that the output parameterization caused the gap. The 100% expert and continuous-policy rates in this finite trial set also do not imply perfect reliability outside it.
Open-loop versus re-anchored model rollouts
The final experiment isolates the compounding problem directly. We take held-out demonstration action sequences and compare two ways of using the learned dynamics model: free-running, where each prediction is fed back as the next input, and re-anchored, where every prediction starts from the true state.
HORIZONS = np.arange(1, 41)
def model_rollout_error(model, episode_list, horizons):
free_running = np.zeros(len(horizons))
reanchored = np.zeros(len(horizons))
count = 0
for episode in episode_list:
block_states = states[
episode * STEPS_PER_EPISODE : (episode + 1) * STEPS_PER_EPISODE
]
block_actions = actions[
episode * STEPS_PER_EPISODE : (episode + 1) * STEPS_PER_EPISODE
]
for t in range(0, STEPS_PER_EPISODE - horizons[-1]):
predicted = torch.tensor(
block_states[t], dtype=torch.float32
).unsqueeze(0)
for step, _ in enumerate(horizons):
action = torch.tensor(
block_actions[t + step], dtype=torch.float32
).unsqueeze(0)
with torch.no_grad():
predicted = model.next_state(predicted, action)
anchored = model.next_state(
torch.tensor(
block_states[t + step], dtype=torch.float32
).unsqueeze(0),
action,
)
target = block_states[t + step + 1]
free_running[step] += float(
np.linalg.norm(predicted.squeeze(0).numpy() - target)
)
reanchored[step] += float(
np.linalg.norm(anchored.squeeze(0).numpy() - target)
)
count += 1
return np.maximum(free_running / count, 1e-12), np.maximum(
reanchored / count, 1e-12
)
held_out_episodes = sorted(set(episodes[test_mask].tolist()))[:20]
free_running_error, reanchored_error = model_rollout_error(
continuous_model, held_out_episodes, HORIZONS
)
The re-anchored curve stays low and declines from about 0.018 to 0.009 across the plotted horizons. Each of its predictions starts from a true state, avoiding recursive prediction feedback; the downward trend also reflects the particular states and actions sampled at later horizons. Free-running error grows from about 0.018 at one step to about 0.654 at forty steps. Its consecutive-step growth ratios trend downward overall, with local increases, rather than remaining constant. This is accumulated drift in this experiment, not a measured constant-factor exponential growth law. The result is specific to this near-linear system and its training coverage; a contact-rich robot may have a different error horizon.

The offline table and closed-loop chart answer different questions. Here the continuous policy has a lower action MSE and a 16-percentage-point higher observed success rate, but this single seeded toy comparison cannot identify a general advantage for either head. Real-robot tasks may change the ranking, especially when actions are multimodal or dynamics include contact discontinuities. The methodological point holds regardless of which side wins: report closed-loop outcomes with uncertainty, and use offline error as a debugging aid.
Limitations and Impact
One durable limitation of current VLA and world-action models is the cost and coverage of embodied data. Teleoperation requires an operator, robot, workspace, and resets; the resulting trajectories are not interchangeable with language-model text tokens, so comparing their raw counts gives no meaningful scale factor. Open X-Embodiment assembled more than a million trajectories across many contributors, yet each target robot and task may still have sparse coverage. Language backbones contribute visual and semantic priors, predictive objectives can use observation-only video when designed for it, and cross-embodiment pooling may help in compatible settings. None of these removes the need to measure control performance on the target distribution.
The second limitation is evaluation. Real-robot trials are slow and variable across setups. Action prediction error is a tempting substitute because it is cheap and automatic, but the code above shows that it measures a different quantity from closed-loop success. Matched simulation and distributed pairwise testing can broaden evaluation, while still needing real-robot validation and explicit uncertainty. Benchmark contamination is a further concern: repeated tuning on a task suite weakens later held-out claims. The chapter's checklist is a partial defense, not a guarantee.
Safety is the third limitation. A trained VLA's offline loss alone gives no general guarantee of safe behavior on a physical system that can break objects, damage itself, or injure a person. Limited guarantees may come from a separately verified controller, monitor, or constrained operating envelope; they do not follow from average action-prediction accuracy. A predictive head might help screen actions, but its errors need testing in the off-distribution situations where safety matters. Optimization against an imperfect predictor can also favor its blind spots. Runtime monitoring, conservative action limits, and human fallback are measures to consider, not features that every research VLA deployment already has. Part XII: Reliable World Models treats their systematic development as an open problem.
The gains are conditional but useful. A language-conditioned interface can reduce the need to redesign task-specific symbols or code for some supported instructions; a new sentence alone does not guarantee a new skill. Cross-embodiment training can reuse experience when observation and action conventions are compatible. Open X-Embodiment released data from 22 robot embodiments, while the reported RT-X training experiments used nine and showed both gains and losses across target evaluations. Predictive models add imagined rollouts and auxiliary signals to existing options such as offline demonstrations, simulation, and model-based learning; they did not invent policy learning without physical execution. The next part of this book asks how these approaches hold up as tasks grow longer and failures become more costly.
Summary
- Joint fusion is an information-bottleneck decision. Modular pipelines pass selected information across interfaces. End-to-end fusion can let action loss shape unfrozen visual features, while spatial precision, latency, and fine-tuning behavior depend on the architecture and training regime.
- Action representation shapes architecture. Uniform endpoint-grid binning has a worst-case per-command error of within the operating range. For the worked autoregressive token decoder, token count and decoding throughput constrain the control rate; continuous heads have different costs.
- Action chunking plus a compressed tokenizer can make autoregressive control feasible. A seven-joint arm with a gripper command at 50 Hz needs 400 tokens per second of behavior if each scalar takes one token, exceeding the illustrative 50 to 150 token-per-second decoding rate. The worked-example frequency-domain representation uses 30 tokens per second; continuous flow or diffusion heads avoid per-scalar autoregression.
- Regression hides multimodality. For an unrestricted deterministic predictor, the population optimum is the conditional mean, which can lie between demonstrated strategies; a trained network is not guaranteed to reach that optimum. Categorical, diffusion, and flow heads can represent multiple action modes, while independent per-dimension categorical sampling can still produce inconsistent combinations.
- Predictive heads must earn their place. Future prediction is neither necessary nor sufficient for control. Counterfactual action sensitivity is a key diagnostic; candidate-action rankings and calibration also matter for the intended decision.
- Compounding error limits open-loop rollouts. In the toy model, free-running prediction error rises with horizon while re-anchoring on true observations keeps error near the one-step level. Prefer re-anchoring when reliable observations are available.
- Dual-rate architectures trade semantic refresh for action rate. GR00T N1 pairs a roughly 10 Hz vision-language module with a roughly 120 Hz action module. That split reduces how often the larger module runs, but grounding quality and stale context must be measured.
- Cross-embodiment transfer works under conditions. Pooling can help under compatible observation and action conventions and can hurt in some settings. The Open X-Embodiment comparisons do not establish a universal tradeoff between task diversity and episode count.
- Validate with closed-loop evaluation as well as action error. Hold out task families when claiming task transfer, evaluate named generalization axes, test perturbation recovery where relevant, and report confidence intervals. Use one-step error as a diagnostic or screening metric, not the sole measure of policy quality.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about vision-language-action models, action tokenization, predictive world-action coupling, and closed-loop evaluation.
Vision-Language-Action and World-Action Models
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!