Part of World Models Handbook
Tokenized game and environment models turn frames into discrete codes. Tests check whether learned transition laws preserve the details agents need.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Tokenized Game and Environment Models
Picture a small arcade game. There is an agent in a room, a locked door on the far side, a key glinting in a corner, and a hazard that ends the run. You record hours of play and train a model to predict the next screen given the current screen and the buttons you press. The model becomes strangely good at this. Feed it a logged sequence and it reproduces the run almost frame for frame. Feed it a novel button sequence and it still produces plausible screens: walls in the right places, sprites in the right positions, a score counter that ticks up.
Then you ask the model to help you plan. One candidate action sequence heads straight to the door without touching the key. In the real game, the locked door blocks that route. If the model instead shows the door opening, it has reproduced the look of the game while getting a rule wrong. Later we will test a different sequence that deliberately collects the key first.
This is the central tension of tokenized game and environment models. If the representation discards the key state, fitting a larger dynamics model to those same current-state tokens cannot recover it without additional history or information. A model can reproduce what the world looks like and still be wrong about what the world does. We will make that distinction precise and then test it in a small runnable example.
Setting up the question
Before we can talk about a model, we need to keep five things apart. In the arcade metaphor:
- The real environment is the actual game engine: it has hidden internal state (agent position, whether the key is held, the score, the frame counter) and rules that only it fully knows.
- The observation is the image of the screen that you receive. It depends on the hidden state, but usually does not reveal all of it.
- The discrete code is the symbol (or short vector of symbols) that a tokenizer assigns to that observation. It is a compressed, learnable stand-in for the frame and the dynamics model's input.
- The policy is the decision rule that maps something (frames, codes, or hidden state) to a control.
- The learned dynamics model is what we are training: a function that predicts the next code (and possibly reward and termination) from the current code plus an action.
These five objects are easy to blur together in conversation, and confusing their roles can make failures hard to diagnose. The environment owns the truth. The observation is a partial window onto that truth. The code is a learned and therefore opinionated summary of the observation. The policy is the thing that acts. The dynamics model is the thing that guesses what comes next. When a model misbehaves, ask which part of this chain is responsible.
As we saw in Part V: World-Model Architectures and again in Part II: Inference, Decisions, and Control, this chapter evaluates a world model by whether it helps an agent act. The question here is not whether the generated screen looks like the game. It is whether the learned transition law preserves the distinctions an agent needs to pick actions several steps ahead.
That distinction, between looking right and acting right, will organize what follows. A model that produces plausible open-loop video can still have a broken action interface, a lossy tokenizer, or errors on rare events. These are possible failure modes with different causes, not inevitable outcomes of tokenization. In Part VIII: Decision-Centric Research Lineages we saw how a policy can exploit errors in a model learned on narrow data. Tokenized game models can inherit that problem, so we will make the representation, data, and action interface explicit.
The plan for the chapter is:
- Convert frames into environment tokens, and understand what that commits us to.
- Build an action-conditioned autoregressive dynamics model over those tokens.
- Understand what interactive generation does and does not give an agent.
- Look at scaling and transfer across games, following a few primary papers.
- Work through a small executable toy world that exposes the failure modes.
- Discuss limitations, impact, and the bridge to the next chapter.
Each step builds on the one before it. Step 1 defines the symbol alphabet. Step 2 defines the probability model over that alphabet. Step 3 asks what the model is good for once it exists. Step 4 asks how far the same recipe stretches when the alphabet and the rules are shared across many games. Step 5 grounds all of it in code, so the abstractions become something you can run and poke. Step 6 draws the boundary of what this family of models can and cannot promise.
Throughout, the guiding principle is the one we set up in Part I: The World-Model Idea: a world model earns its name when it makes decisions better, not when it makes pictures prettier.
From frames to environment tokens
The starting point is a trajectory: a sequence of interactions with an environment. If we write the observation, action, reward, and termination flag at time , the natural notation for a trajectory is
where:
- : the rendered observation at time in the image-based setting studied here. It may be an RGB or grayscale frame, or a stack of frames, possibly with a HUD overlay. Other interfaces, such as RAM observations, require different notation.
- : the discrete control chosen at time . In the Arcade Learning Environment, the full discrete Atari action set has 18 controls; a game's reduced action set may contain far fewer. A multi-game dataset may use a union of the included games' actions, with each game using its applicable subset.
- : the reward received after the transition into time . Depending on the game, rewards may be sparse or frequent; pickups, scores, and deaths are possible reward events.
- : the termination flag for the transition into time . When , the episode ends; a subsequent reset begins a new episode.
- : the length of the episode.
We use for the observation and action indices and for the reward index on purpose. Reward is associated with the transition after action ; in a partially observed environment it may depend on hidden state, not solely on the two rendered frames. Part II, Ch 1: Markov and Partially Observable Decision Processes used this indexing convention. A model that misaligns rewards with actions can learn the wrong decision signal even if frame reconstruction looks good.
The tuple above is a single trajectory. The index labels the observation, labels its reward and termination, and the episode ends at , so a trajectory contains observations and action-reward bundles. Any model we build will be asked to predict the next element of this sequence given the preceding ones.
It is worth pausing on how much structure the raw trajectory leaves implicit. The observation is high-dimensional and redundant. The action is a small discrete label with no meaning until we connect it to an effect. The reward and termination carry explicit task feedback, yet a naive "predict the next frame" objective does not train separate predictions for them. They may also be rare compared with ordinary visual changes. The rest of this section is about the first compression step, and the following section is about the second.
Why tokenize at all
A preprocessed RGB frame contains channel values; this is not the native resolution of every game interface. A sixty-second clip at 30 frames per second has frames. Feeding all pixel values directly into a sequence model is costly, while adjacent frames share much of their content. The model may devote capacity to predicting unchanged walls instead of small events that alter the game's state.
A tokenizer can map each observation (or short clip) into a shorter sequence of discrete symbols. If the token count is reduced by relative to a specified raw-pixel tokenization, the dynamics model sees roughly one-hundredth as many symbols per frame. Its predictions become class labels rather than continuous intensities. A softmax over those labels represents a categorical distribution over code outcomes; whether that distribution is calibrated must be tested separately.
An environment tokenizer maps an observation, or a short temporal window of observations, to a sequence of discrete codes with each . A corresponding decoder attempts to reconstruct an observation from the codes. The quadruple is the tokenization scheme. The dynamics model is trained on the sequence of rather than on .
The token count per frame and vocabulary size are two design knobs. is the number of codes emitted per observation; is the number of entries in the shared vocabulary. At a fixed token-count context window, increasing shortens the frame history that fits. Increasing at fixed does not change that token count, but changes the information capacity per token and the model's embedding and output-classification costs. These are separate tradeoffs, not interchangeable ways of setting the context length.
The combinatorial space of possible code sequences grows with both knobs. An autoregressive factorization is one tractable way to model a distribution over that space; it is not the only structured alternative.

The tokenizer can be one of several kinds:
- Deterministic and fixed: a hand-designed encoder, such as a color histogram or a downsample plus palette quantization.
- Deterministic and learned: a VQ-VAE or similar autoencoder whose encoder is trained to assign a specific code to each input under a nearest-neighbor rule.
- Stochastic and learned: a variational autoencoder with a discrete or relaxed-discrete bottleneck, where the encoder outputs a distribution over codes.
- Temporal: a video tokenizer that maps a short clip into a sequence of codes, so the dynamics model predicts a chunk of frames at once.
Which one you pick changes what the dynamics model sees. If the tokenizer is trained separately and frozen, the dynamics model inherits its quirks. If they are trained jointly, the dynamics objective can shape the tokenizer to keep information that matters for prediction, which is sometimes good and sometimes leads to degenerate codes. We will come back to this tradeoff.
A concrete tokenizer route: vector quantization
Vector quantization is a common discrete-tokenizer design. Walking through its codebook lookup helps explain representation errors that matter later in this chapter.
Suppose an encoder maps an observation to a spatial feature map of shape : . Think of it as a grid of vectors, each of dimension . We maintain a codebook of vectors . For each spatial position , we replace the feature vector with the nearest codebook entry. This is the quantize step, and it is the only place where the continuous encoder output becomes a discrete symbol:
where:
- : the encoder output at spatial position , a vector in that we want to quantize
- : the -th entry of the learned codebook, a representative vector in the same -dimensional space
- : the squared Euclidean distance between the encoder output at position and codebook entry , which measures how close the two vectors are
- : the index of the nearest codebook entry, an integer in that becomes the token for position
- : the quantized feature at position , obtained by looking up the nearest codebook vector , so the feature map is replaced by codebook entries before decoding
The discrete codes are the tokens. The decoder reconstructs an observation from the quantized features: . Training minimizes a reconstruction loss plus a commitment term that keeps the encoder close to the codebook and a codebook term that keeps the codebook close to the encoder. The two auxiliary terms, and the straight-through estimator that lets gradients pass through the non-differentiable , are engineering details; the important point for us is that the tokens are category labels, not numbers, and that the codebook is a jointly learned dictionary. A code is not "a little bit like" another code in the way two nearby real numbers are close; the codes are unordered categories, and the model has to learn any similarity structure between them from data.
The quantization step partitions the encoder's feature space by nearest codebook vector. The figure below shows the assignments; it does not depict codebook collapse, because every illustrated entry is used.

Several knobs matter for the downstream world model:
- Codebook size . A small codebook can merge visually different inputs. A very large one increases the output vocabulary and may leave some codes poorly sampled; the actual usage distribution must be measured.
- Tokens per frame . A larger can preserve more spatial detail, but always lengthens the dynamics sequence. Whether it retains task-relevant detail depends on the encoder and its training objective. The longer sequence also uses more of the sequence model's context window.
- Codebook collapse. During training, only a small fraction of the codebook may receive gradient, so most codes go unused. The effective vocabulary shrinks, and the tokenizer starts aliasing inputs that should have been distinct.
- Reconstruction vs. task sufficiency. The training objective rewards reconstructing the whole frame accurately. It does not care whether the reconstructed frame preserves the small detail that decides the game.
That last point deserves an example. Consider a key that occupies a pixel region on an screen. It covers pixels out of , or about of the pixels. A tokenizer trained to minimize mean squared reconstruction error may devote more capacity to large, frequently visible regions such as the floor and walls than to this small key. But the game cares intensely about the key: its presence can determine whether the door opens. The tokenizer can have low average reconstruction error and still lose a variable needed for this decision. The reconstruction loss and the control task measure different things.
This is why we should not equate reconstruction error with world-model quality. As we noted in Part III, Ch 2: Sufficient State and State Abstraction, a statistic sufficient for painting need not be sufficient for decisions. A tokenizer can discard a feature needed for control while still reconstructing most pixels well. That loss can propagate downstream and may escape reconstruction-only checks; targeted event probes or control tests can reveal it earlier than a failed deployment.
A token ID is a symbol from a learned dictionary. It has no intrinsic causal interpretation. A Markov state makes earlier history unnecessary for predicting the future once the current state and relevant actions are given. An object is a persistent entity in the world with identity, properties, and relations. The three can be aligned, but they are not the same thing, and a tokenized model does not automatically produce either states or objects. This is a recurring source of confusion in the literature.
A tokenized frame is an encoding that may be learned and lossy. Whether it is useful depends on the downstream task; reconstruction loss alone cannot settle whether decision-relevant distinctions survived.
Action-conditioned autoregressive dynamics
With tokens in hand, the next step is to model how they change. An autoregressive sequence model conditioned on the action history is one common choice. A frame contains tokens plus reward and termination variables; a naive categorical table over all token sequences alone would have entries. Autoregression replaces that table with a chain of per-token conditionals. Other structured model families are possible, but this factorization gives us a precise example to analyze.
Here is one factorization, which we will use as a running example. It decomposes the joint prediction of the next frame's tokens, reward, and termination into a product over tokens within the frame, multiplied by a single conditional for the reward and termination given the completed frame:
where:
- : the joint distribution the model assigns to the next frame's tokens, reward, and termination flag, conditioned on the history and the current action
- : the product of per-token conditionals that generates the next frame token by token, each token conditioned on the history, the action, and the tokens already emitted for that frame
Continuing the list of terms:
- : a bounded history of previously encoded observations and actions. Depending on the model, might be for some window , or it might be a recurrent summary.
- : the current action.
- : the tokens of the next frame that have already been generated in the autoregressive order.
- and : the reward and termination flag for the transition.
Every term in this factorization deserves a comment.
The product means the model generates the next frame token by token, conditioning each prediction on the tokens already emitted for that frame. The order is a design choice: a spatial grid can be flattened in raster order, for example. The chain rule permits any fixed order to represent the same joint distribution if the conditionals are sufficiently expressive. With finite data and model capacity, however, the order can affect both training and generated behavior.
The final factor asks for a reward and termination flag given the generated next frame. Unlike the per-token factors, it does not need a partial next-frame prefix: all frame tokens are available. Implementations may use separate heads with their own classification or regression losses. Modeling these variables matters because a visually plausible frame with the wrong reward or termination can mislead a planner.
The specific factorization above is convenient but not universal. Some models generate tokens in parallel with a masked decoder, some predict reward and termination from a hidden state shared with the token predictor, and some interleave reward prediction into the token stream itself. Treat the factorization as an engineering template, not a law.
Training: teacher forcing and cross-entropy
Training an autoregressive token model is a supervised learning problem once the tokenizer is fixed. You take a batch of trajectories, convert each observation to a token sequence, and ask the model to predict the next token sequence given the previous tokens and the observed action.
Concretely, for a training pair , the loss is the sum of cross-entropies over the tokens:
where:
- : the token-prediction loss as a function of the model parameters
- : the model's predicted probability of the -th token of the next frame given the history, the action, and the already-emitted earlier tokens
- : the number of tokens per frame
plus analogous loss terms for and . A continuous reward may use a likelihood model or a regression loss; a small discrete reward set may use cross-entropy. Token cross-entropy is the negative log-likelihood of the observed categorical target under the model. It uses ground-truth categorical targets to train the model's conditional probabilities, but it does not eliminate the train–rollout distribution shift or make greedy, temperature-adjusted, or truncated sampling match the training distribution.
The tokens used as inputs are the actual tokens from the recorded trajectory, not the model's own samples. This is teacher forcing: we condition on ground-truth history and predict one step ahead. It is a common training regime for autoregressive models and it is what Part V, Ch 3: Autoregressive Token World Models described. Two implications matter for evaluation:
- It trains the model to be accurate on data distribution histories, not on its own histories. During a planning rollout, small errors can compound and move the model's inputs away from the training distribution.
- It supplies ground-truth during training. A model may exploit cues in that history that its own rollouts later fail to preserve; teacher forcing does not itself require such reliance, and need not determine a stochastic next frame.
The second point is subtle. Suppose the next frame depends on whether the agent holds the key, and that fact is visible in because the key sprite has disappeared. The model may predict observed next frames accurately by using that cue. Such accuracy alone does not show that its own generated histories will preserve the cue or that it will handle a new pickup sequence. During rollout, an error in the generated history can remove the information on which later predictions depend. That is one way strong one-step accuracy can coexist with weak multi-step simulation; it is not proof that every teacher-forced model failed to learn the rule.
Causal masking enforces this autoregressive structure inside a transformer: token attends only to tokens and to the context . If training exposed later target tokens to that predictor, the stated next-token likelihood would leak information. The mask prevents that leakage; other generative architectures can enforce a valid factorization in different ways.
Sampling, temperature, and rollout
Once trained, we generate a future by sampling from the factorized distribution. Which sampling mode we choose changes the distribution of the futures we observe, and it changes how the model behaves inside a planner. There are several sampling modes:
- Greedy: take the argmax at each step. Deterministic but can produce degenerate, repetitive outputs when the model is uncertain.
- Temperature sampling: sample from , where is the temperature parameter: a scalar that scales the logits before the softmax is applied. Small approaches greedy; large flattens the distribution. Temperature changes concentration and diversity; its effect on fidelity depends on the task, metric, and calibration.
- Top- or nucleus (top-) sampling: when their thresholds exclude tokens, restrict sampling to a high-probability subset. This can remove rare but decision-relevant events such as a pickup or death.
Sampling mode changes the futures a planner considers. Greedy decoding gives one model future. Repeated temperature samples expose several model futures at additional cost, but they are only useful if the model distribution is accurate and calibrated. A top- or top- threshold that excludes a low-probability pickup or death removes that event from sampled plans, biasing the evaluation of those plans.
When the environment is stochastic, the same conditioned history and action can have several valid next outcomes. The observed sample still supplies a valid negative-log-likelihood training target; across repeated samples, cross-entropy can learn the conditional distribution. A single point prediction or an overconfident distribution would instead misrepresent outcome variation, such as random enemy spawns or item drops.
Where reward and termination signals fit
Token prediction alone supplies code labels, not explicit reward or termination predictions. If a known reward or stopping rule can be evaluated from the predicted trajectory, including its actions and transitions where relevant, it can provide those signals; otherwise the model needs learned heads or another estimator. A decoded picture may look like the game, but visual fidelity alone does not tell a planner the return or remaining horizon.
A planner cannot reliably score candidate trajectories by return without a reliable reward mapping for its predicted trajectories, whether learned or known. Without a reliable termination signal, a model rollout may continue past an event that ends the real episode. Frame reconstruction alone does not test either signal.
A visually convincing token stream is not a value function, and a good next-token likelihood does not establish accurate reward or termination signals. When those signals are learned as heads, evaluate their predictions separately. When known rules supply them, test whether predicted trajectories retain the variables those rules need. Either route can fail despite plausible frames.
Time alignment, action repeat, and control latency
Game trajectories have several temporal conventions that a dynamics model must respect.
- Action repeat. Many agents operate on a frame-level control rate of 60 Hz, but the agent chooses an action every frames, and the environment executes that action for the next consecutive frames. The value of becomes part of the effective time step of the model: one action corresponds to consecutive raw frames. If you forget this in the data pipeline, the model learns to predict an action's effect over an inconsistent time horizon.
- No-op controls. Some datasets include a "do nothing" action, sometimes as a legal action and sometimes as a technique for skipping events. The model treats it as just another action, so it must see it in the data.
- Frame skip and image stacking. The observation is often a stack of frames or a max-pooled pair, so the model's input is not a single image. The dynamics model learns a different causal structure than it would from single frames.
- Control latency. In some games and many robots, the action takes effect one or more steps after being issued. A model that assumes zero latency will systematically mis-predict transitions where the true effect lags.
These conventions are common in benchmark pipelines, though their exact settings vary. The preprocessing pipeline is part of the model's action interface. If an action means "press fire for four frames" during training but "press fire for one frame" at test time, the learned conditional transition no longer matches the queried process. The discrepancy can appear as systematic prediction error even when the model's reported uncertainty is low.
One game or many?
There are two scales at which you can train a tokenized game model. A single-game model is trained on trajectories from one environment. A multi-game model is trained on a union of trajectories from many environments, sometimes with an explicit game identifier and sometimes with environment-specific action adapters. The two regimes face different problems.
For a single game, has a fixed meaning within the chosen action interface, and the tokenizer and dynamics can be fitted to that game's data. Its tokens may specialize to the game's art style. The codebook size, training cost, and accuracy still depend on the architecture, data, and task; specialization alone guarantees none of them.
For multi-game, several complications arise:
- Observation representations must be compatible with the shared dynamics model. Differing resolutions, aspect ratios, or palettes may be handled by a shared tokenizer with preprocessing, or by game-specific encoders or adapters that produce compatible model inputs. Downsampling and style normalization are options, not requirements.
- Game identity or context is required. Because the dynamics depend on the game rules, the model must condition on something that identifies the game. This can be a one-hot game embedding, a learned "environment ID" token prepended to the sequence, or natural language. Text-prompt conditioning is taken up in Chapter 53 on Cosmos and omnimodal world foundation models, in Part IX: Foundation and World-Action Models.
- Action sets differ. A shared integer ID for "button 3" need not mean the same underlying effect in every game. "Jump" in one platformer might be "dash" in another and "no-op" in a third. Action masking can constrain the legal choices; per-game adapters or game-conditioned shared dynamics can model differing effects. The interface should state which semantics are shared and which depend on context.
The scaling opportunity is real: a shared tokenizer sees more visual variety, and a shared dynamics model sees more rule variety. Whether that leads to transfer depends on how well the model can condition on game identity and reuse behavior across games. It is a scaling hypothesis, not a guarantee. A model that shares a tokenizer but not a dynamics model can reuse visual features without necessarily transferring rules. A model that shares both needs enough game context to avoid mixing incompatible transition laws.
What interactive generation does and does not give an agent
A tokenized model can be interactive, meaning you can condition on an action and read a sample from the model, but "interactive" is not the same as "useful for control". The distinctions matter enough to spell out.
The interactive loop
The basic loop is:
- Initialize. Provide an initial observation or a history of observations.
- Choose an action. A policy, a human, or a planner picks .
- Sample the model. Generate from the model.
- Decode. If a display is needed, run through the decoder. If not (for internal use), keep it as tokens.
- Advance. Append and the chosen action to the history and go to step 2.
The loop can serve as a simulated environment, but its state is whatever the model and its context retain, not the game's inaccessible ground-truth state. A recurrent model may carry its own hidden memory; a history-based model may carry past tokens and actions. Neither guarantees that a decision-relevant variable discarded from the observations can be recovered. Different samples can also produce different futures from the same history, which is expected for a stochastic model but must be evaluated for consistency with the real environment.
Three uses, three validation targets
Interactive token models are often used for three distinct purposes, and each has its own validation metric.
- Playable generated environments. A human sits at the keyboard and plays the model as if it were a game. The validation is whether the experience is engaging, coherent, and controllable. Chapter 52 on Genie and interactive generated environments develops this framing in Part IX: Foundation and World-Action Models.
- Model rollouts for planning or policy learning. A decision-making algorithm plans over and returns an action sequence. The validation is whether acting on that sequence in the real environment produces the reward the planner expected.
- Conditional video demonstrations. The model is used to synthesize demonstration-like sequences for training or for entertainment. The validation here is a bit softer: fidelity to the style and content of real demonstrations.
The three uses share a plumbing, but their validation targets differ. A model that is a great playable environment may be a bad planner. A model that produces excellent demonstrations may not be resampled in a way that lets a planner exploit it. Conflating the three is a common mistake, because all three look like "the model generates the game", but only the middle one is a claim about decision quality.
Model exploitation and out-of-support action histories
Imagine optimizing a policy inside the learned model. Optimization can favor action histories where model error makes imagined reward look better than real reward. This is model exploitation, discussed in Part VIII, Ch 6: Offline and Conservative Model-Based RL. A token model may still assign a sharp predictive distribution to a poorly supported history; sharpness alone is not evidence that the prediction is accurate there.
Suppose the training data contains individual moves but never the complete key-to-door sequence. A rollout must predict observations after the pickup under a history with little direct support. The model will still return a distribution, but that distribution may be wrong or poorly calibrated. An optimizing policy can exploit optimistic errors if no evaluation or constraint catches them.
One-step held-out likelihood is not enough to catch this. The model can have excellent next-token accuracy on logged steps while mis-modeling a novel action sequence. Fixed-action branch rollouts test multi-step trajectory fidelity; feedback-controlled evaluations go further by testing how a policy reacts to the states the model or real environment actually produces.
Comparing evaluation regimes
The three evaluation regimes we should keep distinct are:
| Regime | What it measures | Where the model can fail |
|---|---|---|
| One-step held-out likelihood | Per-step predictive accuracy on data | Distribution shift; missing multi-step structure |
| Fixed-action branch rollout | Multi-step predictions for a specified action sequence, compared with the real environment under the same actions | Compounding error; hidden-state aliasing; action-interface mismatch |
| Closed-loop policy test | Actions selected from feedback during model and real-environment runs | Model exploitation; policy sensitivity to model errors; control failure |
The closed-loop test is the most direct measure of whether a model supports the intended policy, and it requires access to the real environment. Fixed-action branches are cheaper diagnostics of the learned transition law, but they do not measure an adaptive policy's return.
A useful fixed-action branch protocol is to pick a starting context, choose an action sequence different from the logged one, and run both the model and the real environment under that same sequence. Compare event frequencies, reward, termination, and state summaries at several branch points. This probes counterfactual transitions, but it remains open-loop because later actions do not respond to the generated observations. For a closed-loop test, run a specified policy in each system and let it choose each next action from the feedback it actually receives.
If the environment is stochastic, do not compare pixel-by-pixel to a single realized future. Instead, compare the distribution of futures under the model to the distribution under the real environment. This requires sampling multiple rollouts from the same context and comparing summaries: event frequencies, return distributions, termination rates. Pixel-matching against one sample can count ordinary stochastic variation as model error; one pair of samples cannot establish whether the predicted distribution is wrong.
The two failure modes are distinguishable if you look carefully:
- Aleatoric diversity: the model emits multiple plausible futures under the same action, and the real environment would also emit multiple plausible futures. This is healthy.
- Epistemic uncertainty about rarely observed actions: limited data leave the transition law poorly known. A model may nonetheless emit a sharp and wrong distribution, which can mislead a planner.
A single predictive distribution from one fitted model does not, by itself, separate outcome randomness from uncertainty about the learned transition law. Models with explicit parameter-uncertainty estimates, such as suitable ensembles or Bayesian approximations, can attempt that separation. Calibration and out-of-distribution checks are also relevant when planning decisions must be reliable; Part XII: Reliable World Models develops those tests.
Compounding error and horizon
Because the interactive loop feeds the model's own samples back in, small per-step errors can compound. After steps, where is the rollout horizon, the model's input distribution may differ from the data distribution. The rate at which this drift grows depends on the model's error, the degree of stochasticity, the degree of control, and the presence of absorbing events. There is no single number that characterizes it. Tight control, randomness, and feedback can change how errors propagate, but none alone determines whether a particular rollout drifts slowly or quickly.
What matters depends on the task and model. Per-step prediction errors can accumulate even over short horizons. Over longer horizons, missed events such as pickups, deaths, or terminations may dominate the decision error even when the pixels remain plausible. A planner's useful horizon must therefore be tested against the task and controller; Part VII, Ch 1: Sampling-Based Planning and Model Predictive Control describes short-horizon replanning as one way to limit exposure to model drift.
An interactive model that is spectacular at five steps and garbage at fifty can still be useful for short-horizon control and useless for planning to a distant goal. This is why the horizon must be treated as a property of the system (model plus controller plus task), not of the model alone. The same model can be entirely adequate for one planning problem and entirely inadequate for another.
The horizon at which a model becomes unusable depends on the model, controller, and task together. For illustration only, assume each step has an independent error event with the same probability . The probability of at least one error by step is . Real rollout errors need not be independent or constant-rate, so the curves below are not empirical predictions for a particular world model.

Scaling and transfer across games
A small single-game model is already interesting. The bigger story, and the one that has driven much of the recent literature, is training on many games and hoping that the resulting tokens, dynamics, and interfaces transfer.
Three research case studies
Rather than survey the literature broadly, we will look closely at three primary papers that illustrate distinct design choices. We deliberately do not quote leaderboard numbers: reported scores can depend materially on environment versions, action-set conventions, evaluation frames, and data budgets, so direct comparison requires matched protocols.
IRIS: Transformers are Sample-Efficient World Models. Micheli, Alonso, and Fleuret (2022, arXiv:2209.00588) combine a discrete autoencoder with an autoregressive Transformer world model. Its codes are the world model's prediction space, while its actor-critic is trained using imagined rollouts presented as decoded images. IRIS illustrates a single-game regime with per-game representation and dynamics models. The reported decision outcome is sample efficiency in the real game, distinct from video fidelity.
JOWA: Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining. Introduced by Cheng et al. (2024, arXiv:2410.00564), JOWA pretrains a world-action model on multi-game offline data and studies transfer. Its joint optimization refers to the shared Transformer backbone trained with world-model and action-value objectives. The paper trains a VQ-VAE tokenizer in an earlier stage and freezes it during joint world-action training; it does not show that the tokenizer itself becomes decision-sufficient through that stage. The transfer question is what gets adapted for a new game, using how much target-game data and interaction.
DART: Learning to Play Atari in a World of Tokens. Introduced by Agarwal, Andrews, and Ebrahimi Kahou (2024, arXiv:2406.01361), DART uses discrete representations, an autoregressive Transformer-decoder world model, and a Transformer-encoder behavior model. It is a contrast to IRIS and JOWA in how tokenized predictions are paired with behavior learning. These papers share a broad tokenized-world-model pattern while differing in their training regimes and policy components.
Genie: Generative Interactive Environments (Bruce et al., 2024, Proceedings of the 41st International Conference on Machine Learning) is a forward pointer rather than a peer. Its video tokenizer and autoregressive dynamics are directly relevant to this chapter, but its most distinctive contribution is the inference of latent actions from unlabeled video, the topic of Chapter 52 on Genie and interactive generated environments in Part IX: Foundation and World-Action Models. We mention it here only because it belongs to the same family and because the same questions about action interfaces and closed-loop evaluation apply.
Notice what we have not claimed: we have not claimed that any of these models transfer zero-shot to new games, that any of them solve long-horizon planning, or that they are the best models in their respective benchmarks. We have simply named a few primary sources for the broad design pattern.
A discrete autoregressive game model is a tokenizer plus an autoregressive sequence model whose outputs are discrete tokens conditioned on actions. Models with continuous latent states (for example, RSSM-style architectures from Part V, Ch 1: Recurrent State-Space Models) or with diffusion-based generators (as in Part V, Ch 4: Diffusion and Flow World Models) belong to different families. They may be stronger on some tasks, but they are not instances of the pattern we are discussing.
Sources of transfer, and how each can fail
The phrase "the model transfers across games" hides several distinct mechanisms. Each one can fail independently.
- Reusable visual codes. A tokenizer trained on many games may reuse visual features such as edges or sprite shapes. Whether this helps on a new game depends on its imagery and on whether those features preserve task-relevant details; measure visual transfer separately.
- Shared dynamics priors. The model may reuse patterns such as continued motion or contact followed by collection. This may help when games share transition rules, but transfer can fail when their rules differ, as between a projectile shooter and a turn-based puzzle.
- Transferable action semantics. Some controls, such as "up", may have aligned effects across games, while a shared action ID need not mean the same thing everywhere. Overlapping controls can transfer; ambiguous or game-specific ones require context, masking, or adaptation. A latent action alphabet or learned action embedding is another possible interface.
- A shared decision head. The policy or value head trained alongside the model can transfer if the semantics of "good" and "bad" are sufficiently similar across games. This is fragile: a decision head that plays Pac-Man well is not obviously a good head for a racing game.
Each of these can succeed while the others fail, so an ablation study should isolate them. If the visual codes help but the dynamics priors do not, you want to know, because the fix is different in the two cases. A transfer result that reports a single aggregate number hides which mechanism produced the gain.
Different transfer mechanisms rely on different kinds of shared structure. The figure below uses invented scores to illustrate hypotheses one might test as game similarity changes. It is not evidence that one mechanism transfers better than another.

A clean transfer protocol
The following protocol makes the transfer claim testable.
- Hold out environments, not just levels. Reserve entire games from pretraining. Holding out only levels tests intra-game generalization, a different question from cross-game transfer.
- Specify what is adapted at test time. Does the model fine-tune its tokenizer? Its dynamics? A decision head? An action adapter? Each option defines a different question.
- Specify adaptation data and evaluation interaction separately. A zero-shot transfer claim normally excludes target-game training or parameter updates, but a controller still interacts with the game during evaluation. Few-shot adaptation uses a stated small target-game dataset. In-context adaptation uses target-game context without weight updates. Report both the adaptation budget and the interactions used to score the resulting controller.
- Compare against a from-scratch baseline. Training a model on the new game alone helps show whether pretraining produced a transfer gain under the same evaluation budget.
- Include a no-pretraining ablation. Even a pretrained tokenizer can dominate the effect size; without an ablation you cannot tell whether pretraining the dynamics mattered.
- Distinguish three transfer questions. Transfer within a game (new levels) is different from transfer to a new game with the same action set, which is different from zero-shot control on a new game with a different action set. Solving one does not imply solving the others.
The following is not a rule, but a rule of thumb: if a cross-game transfer result does not specify the action interface, the data budget, and the level of adaptation, it is not interpretable.
A worked example: the key, the door, and the token
Now that the concepts are on the table, let us build a minimal example that exercises them. We will use a tile world with an agent, a key, a locked goal, a wall, and a hazard. It is small enough to reason about completely and rich enough to exhibit the failures we care about.
The story we will demonstrate:
- A discrete tokenizer that preserves the key sprite yields a model that behaves correctly on a plan that goes to the key first.
- A lossy tokenizer that aliases the key sprite into "empty" yields a model with similar one-step accuracy but a broken multi-step plan: the model cannot tell whether the agent has the key, so it cannot open the door.
- A restricted dataset that never includes key pickups yields a model with no in-support entry for that transition. Our rollout applies an explicit stay-put fallback, rather than estimating a confident prediction.
The dynamics model is an empirical count table over observed transitions: for each pair, we store next-observation outcomes with reward and termination sums. It is a pedagogical analogue, not an autoregressive Transformer or a smoothed estimator. The toy isolates representation and data support without large-model training. A more expressive model might use action history to recover some information hidden from the current lossy observation, but changing the model class alone does not guarantee that it will learn the missing pickup rule or handle an unobserved action branch.
Defining the toy world
We start by defining the grid, the tile alphabet, and the transition function.
from collections import defaultdict
import numpy as np
## Grid and tile alphabet
H, W = 5, 5
N = H * W
EMPTY, WALL, KEY, HAZARD, GOAL, AGENT = 0, 1, 2, 3, 4, 5
TILE_NAMES = ["empty", "wall", "key", "hazard", "goal", "agent"]
## Fixed tile positions
KEY_POS = 6 # (row 1, col 1)
HAZARD_POS = 3 # (row 0, col 3)
WALL_POS = 13 # (row 2, col 3)
GOAL_POS = 24 # (row 4, col 4)
## Actions: 0=up, 1=down, 2=left, 3=right
MOVES = np.array([[-1, 0], [1, 0], [0, -1], [0, 1]])
ACTION_NAMES = ["up", "down", "left", "right"]The transition function is deterministic. The agent moves, walls block, hazard ends the episode with a penalty, and the goal is only enterable once the agent holds the key.
def step(state, a):
"""Advance the toy world by one action. Returns (next_state, reward, done)."""
pos, has_key = state
r, c = divmod(int(pos), W)
dr, dc = MOVES[int(a)]
nr, nc = r + int(dr), c + int(dc)
if not (0 <= nr < H and 0 <= nc < W):
nr, nc = r, c
npos = nr * W + nc
if npos == WALL_POS:
npos = pos
new_has_key = has_key or (npos == KEY_POS)
reward, done = 0.0, False
if npos == HAZARD_POS:
reward, done = -1.0, True
elif npos == GOAL_POS:
if new_has_key:
reward, done = 1.0, True
else:
npos = pos # locked door
return (npos, new_has_key), reward, doneRendering writes a tile to each cell and then overlays the agent sprite. The key sprite is drawn only while the agent does not hold the key.
def render(state):
pos, has_key = state
grid = np.full(N, EMPTY, dtype=np.int64)
grid[WALL_POS] = WALL
grid[HAZARD_POS] = HAZARD
grid[GOAL_POS] = GOAL
if not has_key:
grid[KEY_POS] = KEY
grid[pos] = AGENT
return gridA glance at the world makes the layout clear.

The picture shows the layout that the tokenizer must encode. What we will manipulate next is not the picture but the tokenization scheme that reads it, because the scheme decides which of these tiles survive into the model's input.
Tokenization: full and lossy
The full observation is the integer grid. The lossy observation replaces every KEY tile with EMPTY. This mimics a tokenizer that assigns no code to the key sprite, either because the codebook was trained with a small or because the reconstruction loss did not care about a tiny detail.
def lossy(obs):
aliased = obs.copy()
aliased[aliased == KEY] = EMPTY
return aliased
def agent_pos(obs):
return int(np.where(obs == AGENT)[0][0])For reachable states in this toy world, the full frame reveals whether the agent holds the key: the key sprite disappears after pickup. The lossy tokenizer removes that sprite in both cases. At the same apparent position, a key-held and a key-unheld state can therefore produce the same lossy observation but different outcomes for an action toward the goal. The lossy current observation is not Markov-sufficient; an action history that records the pickup could restore the distinction.
Generating training data
We generate trajectories under a uniform exploration policy and record tuples.
EPISODES = 3000
HORIZON = 30
rng = np.random.default_rng(0)
transitions = []
for _ in range(EPISODES):
state = (0, False)
obs = render(state)
for _ in range(HORIZON):
a = int(rng.integers(4))
next_state, r, done = step(state, a)
next_obs = render(next_state)
transitions.append((obs, a, next_obs, float(r), bool(done)))
state, obs = next_state, next_obs
if done:
break
perm = np.random.default_rng(3).permutation(len(transitions))
transitions = [transitions[i] for i in perm]
split = int(0.85 * len(transitions))
train_trans = transitions[:split]
test_trans = transitions[split:]We also generate a restricted dataset. Under the restricted policy, the agent never takes an action that would step onto the key cell, so key pickups never occur in the training data.
def restricted_behavior(state, rng):
options = []
for a in range(4):
(npos, _), _, _ = step(state, a)
if npos != KEY_POS:
options.append(a)
return int(rng.choice(options)) if options else int(rng.integers(4))
rng_restricted = np.random.default_rng(23)
restricted_trans = []
for _ in range(EPISODES):
state = (0, False)
obs = render(state)
for _ in range(HORIZON):
a = restricted_behavior(state, rng_restricted)
next_state, r, done = step(state, a)
next_obs = render(next_state)
restricted_trans.append((obs, a, next_obs, float(r), bool(done)))
state, obs = next_state, next_obs
if done:
breakThe restricted policy still visits most of the grid. What it avoids is the single transition into KEY_POS.
A count-based dynamics model
The model is a lookup table. For each observed pair we store the empirical distribution over next observations and the mean reward and termination under that outcome.
def fit_count_model(transitions, tokenizer=lambda x: x):
table = defaultdict(lambda: defaultdict(lambda: [0, 0.0, 0.0]))
for obs, a, next_obs, r, done in transitions:
key = (tuple(tokenizer(obs).tolist()), int(a))
nkey = tuple(tokenizer(next_obs).tolist())
entry = table[key][nkey]
entry[0] += 1
entry[1] += r
entry[2] += float(done)
return table
def predict_next(table, obs, a, tokenizer=lambda x: x):
key = (tuple(tokenizer(obs).tolist()), int(a))
dist = table.get(key)
if not dist:
return None, 0.0, 0.0
nkey = max(dist, key=lambda k: dist[k][0])
count, r_sum, d_sum = dist[nkey]
total = sum(v[0] for v in dist.values())
return np.array(nkey, dtype=np.int64), r_sum / count, d_sum / countThe predict_next function tabulates complete next-frame outcomes instead of factorizing tokens within a frame. Its empirical frame mode need not match greedy token-by-token decoding, especially when a lossy observation aliases distinct hidden states. We use the table to isolate representation and coverage effects, not as an implementation of the earlier autoregressive equation. The table conditions on the current tokenized observation and action; it has no longer history.
We fit three such models: full tokenization on the rich data, whose tokens distinguish the key sprite from the empty tile; lossy tokenization on the rich data, whose tokens alias the key sprite into the empty tile; and full tokenization on the restricted data, whose transitions never include a key pickup.
model_full = fit_count_model(train_trans)
model_lossy = fit_count_model(train_trans, tokenizer=lossy)
model_restricted = fit_count_model(restricted_trans)One-step accuracy looks fine
Before we test the models on multi-step plans, let us look at the metric that is easy to compute and misleading: held-out next-token accuracy.
def one_step_accuracy(model, test_trans, tokenizer=lambda x: x):
correct = 0
total = 0
for obs, a, next_obs, r, done in test_trans:
pred, _, _ = predict_next(model, obs, a, tokenizer=tokenizer)
if pred is None:
continue
total += 1
if np.array_equal(pred, tokenizer(next_obs)):
correct += 1
return correct / max(total, 1), total
acc_full, n_full = one_step_accuracy(model_full, test_trans)
acc_lossy, n_lossy = one_step_accuracy(model_lossy, test_trans, tokenizer=lossy)Full tokenization, held-out one-step accuracy: 1.000 (10393 transitions) Lossy tokenization, held-out one-step accuracy: 0.997 (10393 transitions) Restricted data, held-out one-step accuracy: 1.000 (4230 transitions)

The full and lossy models score comparably here, but the lossy model is evaluated in its own token space, where the key sprite has been removed. The restricted model's perfect score is more selective: its denominator excludes 6,163 of the 10,393 held-out queries because it has no table entry for them. Accuracy must be read alongside coverage, and neither number alone tells us how the planned key-to-door sequence will unfold.
Multi-step plan: key first, then the door
We hand the model a plan that goes down to the key, picks it up, walks to the goal, and stops. The same plan is executed in the true environment, in the full model, in the lossy model, and in the restricted model. The four resulting distance-to-goal curves reveal which models preserve the pickup rule and which ones stall short of the goal.
start_state = (0, False)
## Actions: down, right, down, down, down, right, right, right
plan = [1, 3, 1, 1, 1, 3, 3, 3]
def rollout_true(state0, actions):
state = state0
obs_hist = [render(state)]
rewards = []
dones = []
for a in actions:
state, r, done = step(state, a)
obs_hist.append(render(state))
rewards.append(r)
dones.append(done)
if done:
break
return obs_hist, rewards, dones
def rollout_model(model, obs0, actions):
obs = obs0.copy()
hist = [obs.copy()]
rewards = []
dones = []
for a in actions:
pred, r_hat, d_hat = predict_next(model, obs, a, tokenizer=lambda x: x)
if pred is None:
pred = obs.copy()
obs = pred
hist.append(obs.copy())
rewards.append(r_hat)
dones.append(d_hat)
return hist, rewards, donesWe roll out all four trajectories and compute the L1 distance from the agent's tile to the goal at each step.
obs0 = render(start_state)
true_hist, true_rewards, true_dones = rollout_true(start_state, plan)
full_hist, full_rewards, full_dones = rollout_model(model_full, obs0, plan)
lossy_hist, lossy_rewards, lossy_dones = rollout_model(
model_lossy, lossy(obs0), plan
)
restricted_hist, restricted_rewards, restricted_dones = rollout_model(
model_restricted, obs0, plan
)
def distance_curve(hist):
goal_r, goal_c = divmod(GOAL_POS, W)
dists = []
for obs in hist:
pos = agent_pos(obs)
r, c = divmod(pos, W)
dists.append(abs(r - goal_r) + abs(c - goal_c))
return np.array(dists, dtype=float)
true_dist = distance_curve(true_hist)
full_dist = distance_curve(full_hist)
lossy_dist = distance_curve(lossy_hist)
restricted_dist = distance_curve(restricted_hist)True reward trajectory: ['+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+1.0'] Full model reward traj: ['+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+1.0'] Lossy model reward traj: ['+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0'] Restricted model reward: ['+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0', '+0.0'] True termination step: 8 Full model termination: 8 Lossy model termination: never Restricted termination: never
The true environment earns the reward and terminates at the goal. The full model predicts the same state path, reward, and termination step; the lossy and restricted models miss the pickup and do not predict a terminal goal on this branch. To compare all eight prescribed actions on the same chart, the model rollouts record predicted termination without stopping early. This is a fixed-action diagnostic, not a policy run.
The plot below shows the agent's distance to the goal under this fixed action sequence. The true environment reaches zero (goal achieved), and the full model follows it exactly. The lossy model stalls one step away at distance one because the aliased key tile makes the pickup-invisible observation indistinguishable across key-held states, so the model's argmax next-state never enters the goal. The restricted model reaches distance one one step later in the plan because the pickup transition was absent from its training data, so its prediction at the pickup step is wrong and stalls until it recovers. The two impaired models both end at distance one, but for different reasons.

The chart is a fixed-action branch test. It does not score picture similarity or next-token likelihood; it compares the predicted outcome of one specified plan with the environment's outcome. Two models miss the goal despite high conditional one-step accuracy. This is not a closed-loop test of a policy choosing new actions from feedback.
Out-of-support queries and explicit fallback
The restricted model deserves a closer look. It was trained on thousands of trajectories from the same grid, so it knows the world quite well, except that it never saw a key held. We can measure the effect crisply: what fraction of the key-held observations in the held-out test set does the restricted model have an entry for?
def has_key_from_obs(obs):
# Under the full tokenization, the key sprite is absent if and only if the agent holds it.
return not np.any(obs == KEY)
key_held_trans = [t for t in test_trans if has_key_from_obs(t[0])]
def support_fraction(model, transitions):
"""Fraction of test transitions with a matching observation-action entry."""
if not transitions:
return 0.0
hits = 0
for obs, a, next_obs, r, done in transitions:
key = (tuple(obs.tolist()), int(a))
if key in model:
hits += 1
return hits / len(transitions)
support_full = support_fraction(model_full, key_held_trans)
support_lossy = support_fraction(model_lossy, key_held_trans)
support_restricted = support_fraction(model_restricted, key_held_trans)Held-out key-held transitions: 5811 Full-model in-support fraction: 1.000 Lossy-model in-support fraction: 1.000 Restricted-model in-support fraction: 0.000

Both impaired models predict the wrong outcome for the fixed plan, but this code does not measure their confidence. The lossy model chooses a frequent next observation from aliased data. The restricted model has no entry for the pickup action; predict_next returns None, and rollout_model substitutes the unchanged observation. That fallback is deliberately simple, not a general safety guarantee. A neural autoregressive model normally emits a distribution even for a poorly supported query, so its uncertainty and calibration need separate tests.
The toy failure is benign, but the diagnostic carries over: inspect whether a proposed action branch lies within the model's data support, and test its predicted consequences against the real environment before trusting it for control.
Limitations & impact
Tokenized game models use compact discrete observation sequences and supervised predictive losses, but their reachable code space can still be large and autoregressive sampling can be costly. Their limitations arise from representation, data support, action semantics, and the way a planner uses their predictions.
The first limitation is the tokenizer. If the current code and available history contain no trace of a decision-relevant variable, the dynamics model cannot condition on it. In our toy, the count model sees only the current lossy grid, so removing the key tile aliases key-held and key-unheld states. Reconstruction-oriented objectives, especially a uniform pixel loss, can underweight small task-relevant details relative to large surfaces. A task-aware loss, longer action history, or a separate symbolic channel may help, but each needs its own test of whether the required distinction survives.
The second limitation is the action interface. A shared token model needs to interpret the action it receives. Even in one game, latency, action repeat, and variable frame rates can create a mismatch; multiple games add differing controls. Cross-game transfer results therefore need to state what is shared and what is adapted at test time. A policy trained inside a token model can earn high imagined returns yet perform poorly in the real game when it exploits an action-interface or transition-model error.
The third limitation is data support. When a model is trained offline on trajectories, its reliable coverage depends on those trajectories. Expert demonstrations may concentrate on successful paths and omit alternative actions; exploratory data may visit different states without reaching a task objective. Neither collection strategy guarantees coverage of the compound action sequences that a planner wants to evaluate. Our toy world exhibits this cleanly: the restricted dataset covers most of the grid but completely misses the pickup transition, so the model has no entry for the state it needs. Real datasets are messier: coverage is partial, action distributions are uneven, and rare events are rare for a reason. You should not assume that a large dataset implies coverage of the specific branch you care about. This is the same concern we discussed in Part VIII, Ch 6: Offline and Conservative Model-Based RL, where the mismatch between the data distribution and the query distribution motivates conservative evaluation.
The fourth limitation is evaluation. Images are easy to inspect, but a visually coherent rollout can hide a wrong reward, a missed termination, or an inconsistent level layout. No single scalar certifies a world model: reconstruction, token likelihood, event prediction, reward, action sensitivity, calibration, long-horizon consistency, and real closed-loop return answer different questions. A result reported on one metric should not be read as a result on the others.
Tokenized game models offer a useful testbed. Existing game trajectories can train predictive models, and discrete codes can reduce sequence length compared with raw pixel tokenization. Resettable environments support repeatable policy-return tests; comparing counterfactual branches from the same interior context additionally needs snapshot/restore, reproducible replay, or equivalent state control. Sampling cost still depends on model size, token count, and decoding scheme, so cheap imagined trajectories are not guaranteed. These systems let us test ideas from Part VII: Planning and Agency and Part VI: Learning World Models at Scale under controlled conditions.
The practical implication is that a tokenized game model should be described with the same care as any other world model. Which data was it trained on, with which behavior policy, over which action interface, and with what testing regime? Under what conditions does it transfer, and under what conditions does it fail? And, most importantly, what evidence do we have that the learned transition law preserves the distinctions an agent needs, rather than merely reproducing the appearance of the world?
This is the setup for the next chapter. Our toy world used a hand-designed tokenizer: the grid of tile types is already discrete, and we only had to decide whether to include the key. Real environments are not so kind. Chapter 50, Pretrained Visual Models and DINO-WM, in Part IX: Foundation and World-Action Models, changes the representation choice fundamentally: it keeps a continuous feature space learned by a pretrained visual model, and it predicts in that space instead of predicting discrete codes. The intuition we built here, that a representation is only as useful as the decisions it supports, carries over directly. The mechanisms do not.
Summary
Tokenized game and environment models encode observations as sequences of discrete symbols. An observation tokenizer maps frames to ; one possible action-conditioned model autoregressively predicts the next tokens, reward, and termination. Tokenization can shorten the sequence relative to a specified raw-pixel representation, while the cost of sampling and planning still depends on , , and the dynamics architecture.
The chapter's key points are summarized below.
- Tokenization is a representation choice. The tokenizer is not a neutral preprocessing step. It decides which features the dynamics model can reason about, and it can have low reconstruction error while discarding the variable that determines whether the environment is solvable.
- The autoregressive factorization is a design template. A discrete token model with an action-conditioned prior, plus reward and termination heads, is one convenient structure. Alternatives exist and are legitimate, and different choices trade capacity, training stability, and sampling speed.
- Time alignment is part of the model. Action repeat, no-op controls, stacking, and control latency affect the transition data. If timing information needed for prediction is absent from the model inputs, increasing model size alone cannot supply it.
- Interactive generation is not the same as decision usefulness. A model can be a great playable environment and a bad planner, or vice versa. Three uses demand three validation targets: playability, planning value, and demonstration quality.
- Separate branch tests from closed-loop tests. One-step accuracy does not certify a world model. Fixed-action branches test multi-step transition fidelity; adaptive policy runs in the real environment test control usefulness.
- Data support and out-of-distribution queries matter. A neural token model may emit a prediction even on poorly supported inputs. Coverage, calibration, uncertainty checks, or conservative planning can help determine whether a proposed plan should be trusted.
- Transfer has separable components. Visual codes, dynamics priors, action semantics, and decision heads can succeed or fail independently. Cross-game results should specify which mechanism is tested, how the model is adapted, and what baselines are used.
The next chapter takes the representation question one step further. Instead of training a discrete tokenizer and an autoregressive dynamics model over its codes, we will look at models that keep a continuous feature space from a frozen pretrained visual backbone and predict in that space. The same core question applies: does the learned transition preserve the distinctions the agent needs, or does it only preserve the ones that make the video look right?
Key Parameters
The key parameters for the count-based toy dynamics model are:
- Tokenizer: the function that maps each observation to its discrete code. Whether the key sprite is preserved determines whether the model can reason about the pickup rule.
- Codebook size : the number of distinct code symbols. A small can encourage aliasing; a large may leave some codes poorly sampled. Measure actual code usage and task-relevant distinctions rather than inferring either outcome from alone.
- Tokens per frame : the number of discrete codes emitted per observation. Larger can preserve more spatial detail, depending on the tokenizer, but it always lengthens the dynamics sequence and uses more context-window capacity.
- Training data support: the coverage of the training trajectories. A missing pickup transition leaves the count model without an entry; a neural model may produce a prediction, but its confidence and accuracy on that branch require separate assessment.
- Temporal and action interface: frame stacking determines the observation history, while action repeat, no-op controls, and control latency determine how controls align with transitions. A mismatch in either can make the learned dynamics wrong.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about tokenized game and environment models.
Tokenized Game and Environment Models
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!