Genie and Interactive Generated Environments

Michael BrenndoerferJuly 22, 202654 min read

Part of World Models Handbook

Genie discovers latent actions from unlabeled video and builds interactive generated environments, with tests for control, rollout, and persistence.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Genie and Interactive Generated Environments

Picture a short video clip of a cartoon character standing on a grassy platform. The scene looks coherent: shadows line up, the sky matches the grass, and the character's idle animation loops naturally. Now imagine trying to play that clip. You press an arrow key, and the character turns. You press another, and it jumps. You press a third, and it stops against a rock that was already in the background. You walk to the left edge of the platform, turn around, and the same rock is still there, in the same place, with the same shadow. You have not watched a video. You have interacted with a small, generated world.

That experience, a machine-learned environment that accepts a stream of actions and produces a consistent stream of observations in response, is the target of this chapter. The original Genie model (Bruce et al., 2024) learned a discrete latent-action interface from video without recorded button presses. Genie 2 (2024) and Genie 3 (2025) are later Google DeepMind releases that report image-prompted 3D environments and text-prompted real-time interaction, respectively. Their control interfaces and training details should not be assumed to match the original architecture.

The key shift is not simply that these models generate video. The generated frames sit inside an interaction loop: an action changes what the model generates next. When a real environment and a fixed action sequence are available, a generated rollout can be compared with its reference trajectory, as we do in the toy experiment below. For open-ended generated worlds there may be no unique reference for every branch, so we must also test responsiveness, persistence, and control through interventions. That evaluation difference matters as much as the architecture.

The toy seeds its random generators and computes model-comparison diagnostics and figures from fitted outputs and simulator reference trajectories. The feedback-gain figure is a separate seeded scalar illustration. Its prose describes the seeded configuration shown here; after changing the grid, horizon, or codebook size, rerun the cells and revisit the prose as well. A Key Parameters summary at the end records the settings.

We have already met video generators as candidate world models in Video Generators as Candidate World Models. That chapter used an occluded moving object, timed actions, matched-initial-state branches, and horizon-resolved error to ask whether a generator responds to actions. Genie adds the missing-label problem: how can video without recorded controls supply an action interface? We will trace the original model's answer, then test in a small grid world where inferred codes, true actions, and reference trajectories can all be compared.

It helps to separate four things that are often blurred together when people talk about the quality of a learned world model. Predictive fidelity asks whether the next observation is close to the real next observation. Representation quality asks whether the internal codes or tokens preserve the structure that matters for control. Uncertainty quality asks whether the model knows when it is guessing. Decision usefulness asks whether a planner or policy can achieve goals by using the model. Genie-style systems touch all four, but they do not optimize all four in the same way. The inverse model is mainly about representation and action abstraction. The forward model is mainly about predictive fidelity. The interaction loop is where decision usefulness is tested, and it is also where uncertainty and consistency failures become most visible. Keeping these axes separate will make the rest of the chapter easier to read.

Latent Action Discovery

Standard action-conditioned world models, the kind we studied in Recurrent State-Space Models and MuZero and Value-Equivalent Models, assume a dataset of tuples (ot,at,ot+1)(o_t, a_t, o_{t+1}) where the action ata_t is known, with

  • oto_t: the observation at time tt
  • ata_t: the action taken at time tt
  • ot+1o_{t+1}: the resulting observation at time t+1t+1

In an Atari replay buffer, ata_t is a recorded button press; in a robot demonstration it might be a recorded control command. Ordinary uploaded video, including gameplay clips, often lacks that paired control stream. Original Genie trained on such unlabeled gameplay video. One option is to restrict training to action-labeled recordings; another is to infer a compact action-like variable from the observed transitions.

The usual supervised recipe gives the model a state, a recorded action, and the resulting state, then minimizes prediction error. Without an action record, an unconditional video model cannot offer that same user-controlled branch. Hand-labeling some clips is possible but adds annotation work and may cover only selected behaviors. Genie asks what action-like structure can be inferred from the video itself.

Genie's answer to this problem is to introduce a latent action. Between two consecutive frames of a video, there is a change. Whatever caused that change is, in a loose sense, an "action." Even in a nature documentary, the next frame is different from the current one because something happened: the camera panned, a bird flapped its wings, a leaf fell. The latent action is a discrete code that summarizes that change. Crucially, it is not required to correspond to any human-meaningful command. It just has to be a useful, compact summary of what moved. The word "action" here is a modeling convenience, not a claim about agency. The code does not need to know who acted or why. It only needs to help predict what happens next.

Latent Action

A latent action is a discrete or continuous variable utu_t inferred from observed change, not supplied as a recorded control label. Our toy infers it from (ot,ot+1)(o_t,o_{t+1}); the original Genie's latent-action encoder can use the preceding frame history as well. Its code is useful when it helps reconstruct or predict the next observation, but its integer name carries no prescribed button meaning.

The simplest abstraction is an inverse model that receives two frames and returns a code from a small discrete codebook. In the original Genie, the latent-action model's own reconstruction decoder tests whether its code preserves information needed to reconstruct the next frame; the dynamics model consumes that code without its loss updating the latent-action encoder through the action input. The two are co-trained after tokenizer pretraining, with separate objectives. For our pair-based toy abstraction, write the code distribution as

qϕ(ut∣ot,ot+1)q_\phi(u_t \mid o_t, o_{t+1})

where:

  • utu_t: the latent action that summarizes the change between frame tt and frame t+1t+1
  • oto_t: the observation (frame) at time tt
  • ot+1o_{t+1}: the observation (frame) at time t+1t+1
  • qϕ(⋅∣⋅)q_\phi(\cdot \mid \cdot): a learned distribution over codes ut∈{1,…,K}u_t \in \{1, \dots, K\}
  • KK: the cardinality of the latent action codebook; Genie uses K=8K = 8 in the original report, but the exact number is a hyperparameter

The model is called "inverse" because it runs the prediction question backward: given the earlier and later frames, what action-like code describes their change? Recorded actions remove the need to discover such labels for training, though predicting an action from images can still be ambiguous. Genie has no recorded actions, so its latent-action model supplies the interface. Its reconstruction objective pressures the code to preserve next-frame information, and a small codebook limits the vocabulary a user has to explore.

The forward dynamics model predicts next-frame visual tokens given past tokens and a latent action. The original Genie first trains its video tokenizer, then co-trains the latent-action model and dynamics model; it stops the dynamics loss gradient at the latent-action input. The latent-action model itself has an encoder, quantized codebook, and reconstruction decoder. At inference the user selects a code, and the latent-action encoder and its training decoder are not needed. In this chapter we use oto_t for toy observations and ztz_t for full-model visual tokens:

pθ(zt+1∣z≤t, u≤t)p_\theta(z_{t+1} \mid z_{\leq t},\, u_{\leq t})

where:

  • ztz_t: the visual tokens for frame tt (for example, patches or discrete codes produced by a tokenizer)
  • z≤tz_{\leq t}: the sequence of visual tokens from frame 11 through frame tt, the conditioning context
  • u≤tu_{\leq t}: the sequence of latent actions from transition 11 through transition tt
  • pθ(⋅∣⋅)p_\theta(\cdot \mid \cdot): a learned distribution over the next frame's visual tokens

The latent-action model encodes change and reconstructs the future from that code and the past; this reconstruction objective makes the code informative. The dynamics model then learns to predict video tokens using the resulting codes. The original method does not simply backpropagate the dynamics loss through the latent-action input, so "co-trained" should not be mistaken for a single end-to-end inverse/forward loss. This separation also matters when we later compare the published model with our hand-engineered toy.

In the original method, the latent-action reconstruction objective learns an informative quantized code, while a separate cross-entropy objective trains the masked-token dynamics model on video tokens and stop-gradient latent actions. The tokenizer is pretrained before this stage. This is a staged training procedure, not one end-to-end optimization of the simplified objective below.

Codebook collapse remains a risk in quantized models: if all transitions use one code, the dynamics model has no varied action signal. Genie's reported latent-action training uses a VQ-VAE-style objective and a small codebook, but the paper does not establish that a separate code-usage-balancing term guarantees every code will be used. Code utilization and action responsiveness therefore need empirical checks.

As a simplified next-frame abstraction, a conditional token-prediction objective is the negative log-probability of observed next-frame tokens given prior tokens and actions:

Lforward(θ)=−∑tlog⁡pθ(zt+1∣z≤t, u≤t)\mathcal{L}_{\text{forward}}(\theta) = -\sum_{t} \log p_\theta(z_{t+1} \mid z_{\leq t},\, u_{\leq t})

where:

  • Lforward(θ)\mathcal{L}_{\text{forward}}(\theta): the training loss for the forward model's parameters θ\theta, written as the negative log-likelihood summed across all frame transitions tt
  • log⁡pθ(zt+1∣z≤t, u≤t)\log p_\theta(z_{t+1} \mid z_{\leq t},\, u_{\leq t}): the log-probability that the model assigns to the actual next-frame tokens zt+1z_{t+1} given the earlier tokens and latent actions
  • the leading minus sign: converts the log-likelihood into a loss, so a lower value means the model predicts the observed future better

The original latent-action model has a different training signal: its decoder reconstructs the next frame from the past and quantized inferred action, with the VQ-VAE objective training the encoder and codebook. Maximizing qϕq_\phi on codes chosen by that same qϕq_\phi, without reconstruction or another grounding term, would not make the codes useful. The forward expression above is likewise a teaching abstraction; Genie's dynamics model predicts masked frame tokens with cross-entropy, then samples each frame in multiple MaskGIT refinement steps. Our toy is more distant still: its KK-means codes come from hand-extracted displacements rather than a learned reconstruction model.

Why learn the action code instead of using optical flow or a hand-designed state? Flow describes motion but is not necessarily a compact control interface, while a handcrafted variable for one game may not transfer to another. The latent-action bottleneck restricts the code to a small vocabulary of changes. It makes that vocabulary available to the dynamics model, but does not force the dynamics model to use it: an action-insensitive predictor could still rely on visual history. We test responsiveness separately for precisely that reason.

The first thing to internalize about latent actions is that they have a permutation ambiguity. A code is a compressed summary of change, not a semantic label. Suppose you have trained both models jointly, and everything works. Now swap the labels of code 1 and code 2 in the codebook. The forward model never saw the labels, only the embedding vectors, so as long as you also swap the corresponding embedding rows, the loss is unchanged. The trained system is just as good as before. But the meaning you attach to "code 1" and "code 2" has changed. If you had been using code 1 to mean "left arrow," you now have to use code 2. The labels are an arbitrary naming convention, not a property of the learned model.

The permutation ambiguity is not a rare corner case. It is a structural property of a discrete latent codebook without label supervision. The assignment of integer names to learned codes is arbitrary: simultaneously permuting the encoder's assignments and the dynamics model's code embeddings leaves a given model's predictions unchanged. Retraining with another seed may change the code names, but can also change predictions for unrelated optimization reasons. Evaluation should therefore ask whether codes can be aligned with useful effects, not whether code 0 has a prescribed direction. As our toy shows, even the existence of a clean one-to-one alignment can fail when different button presses produce the same observed transition.

The second thing to internalize is that latent actions are not automatically identifiable physical controls. In a grid world with four buttons, one might hope for four corresponding codes, but blocked moves create a fifth observable outcome (no displacement) and obscure the button that caused it. In a cooking video, codes might summarize camera movement, hand movement, or both. The training objective seeks useful compression of change, not a guarantee of human button semantics.

This point is easy to underestimate because the grid-world example is so clean. In the grid world, the change between frames is exactly the agent's movement, and movement has four discrete directions. In real video, the change between frames can come from camera motion, object motion, lighting changes, occlusions, and cuts between scenes. The inverse model may allocate codes to any of these. A code might mean "the camera panned left" in one clip and "the player moved left" in another, if those two changes happen to be similar in token space. The model has no built-in notion of cause, intention, or control. It has a notion of predictive usefulness. That is enough to make a controllable latent space in favorable cases, but it is not enough to guarantee a human-interpretable action space.

The third thing to internalize is that a latent action does not become a keyboard command or a portable interface until you explicitly align it with one. If you want a human to press the left arrow and see the character move left, you need one of the following:

  • A small set of labeled interaction data that pairs latent codes with real commands.
  • A user interface that lets the user pick a code and remembers the mapping.
  • An auxiliary loss that pushes codes toward known controls.

Without such alignment, a person or outside policy can still select the opaque code IDs and learn their effects. What they cannot assume is that one ID already means "left arrow" or that its meaning transfers to another trained model. That is the difference between a model that has learned a control abstraction and one that has learned your control abstraction. The alignment step can be small, but it is separate from latent action discovery: discovery produces a vocabulary of change, while alignment maps your commands into that vocabulary.

Let us make this concrete with a tiny grid world. We will generate trajectories from a known environment, throw away the action labels, and see what latent actions the model discovers.

In[3]:
Code
import matplotlib.pyplot as plt
import numpy as np
from sklearn.cluster import KMeans
from sklearn.neural_network import MLPRegressor

We build an 8 by 8 grid with a handful of walls, four actions (up, down, left, right), and deterministic transitions. If the agent attempts to move into a wall or off the grid, the move is a no-op and the agent stays in place. This kind of blocked move will matter later, because it creates pairs of actions with the same observable effect.

In[4]:
Code
GRID = 8
WALLS = {(2, 3), (2, 4), (2, 5), (5, 1), (5, 2), (5, 3), (6, 3), (3, 6), (4, 6)}
ACTIONS = {"up": (-1, 0), "down": (1, 0), "left": (0, -1), "right": (0, 1)}
ACTION_NAMES = list(ACTIONS.keys())
K = 4  # four codes for five possible displacements, including no-op


def step(pos, action_name):
    dr, dc = ACTIONS[action_name]
    nr, nc = pos[0] + dr, pos[1] + dc
    if 0 <= nr < GRID and 0 <= nc < GRID and (nr, nc) not in WALLS:
        return (nr, nc)
    return pos


def render(pos):
    frame = np.zeros((GRID, GRID), dtype=np.float32)
    for r, c in WALLS:
        frame[r, c] = -1.0
    frame[pos[0], pos[1]] = 1.0
    return frame


def one_hot(idx, k=4):
    idx = np.atleast_1d(idx)
    out = np.zeros((len(idx), k), dtype=np.float32)
    out[np.arange(len(idx)), idx] = 1.0
    return out

We now generate a dataset of random trajectories. Each trajectory starts from a random free cell, and at each step we sample one of the four actions uniformly at random. We collect the current frame, the true action (which we will hide during latent-action inference), and the next frame.

In[5]:
Code
rng = np.random.default_rng(0)
n_traj = 200
horizon = 15

obs_list, act_list, nxt_list = [], [], []
for _ in range(n_traj):
    pos = (int(rng.integers(GRID)), int(rng.integers(GRID)))
    while pos in WALLS:
        pos = (int(rng.integers(GRID)), int(rng.integers(GRID)))
    for _ in range(horizon):
        a = ACTION_NAMES[int(rng.integers(4))]
        nxt = step(pos, a)
        obs_list.append(render(pos))
        act_list.append(a)
        nxt_list.append(render(nxt))
        pos = nxt

observations = np.array(obs_list)
next_observations = np.array(nxt_list)
true_actions = np.array(act_list)
Out[6]:
Console
Collected 3000 transitions
Frame shape: (8, 8)

First split by whole trajectory: 160 trajectories train the clusterer and predictors, and 40 remain for one-step testing. Adjacent frames from one trajectory therefore cannot appear on both sides of that split. From each frame pair we extract agent displacement. K-means fits four clusters on training displacements only; it assigns codes to the held-out transitions afterward. This is a hand-engineered, hindsight-based stand-in for qϕ(ut∣ot,ot+1)q_\phi(u_t \mid o_t,o_{t+1}): assigning a test code requires the test next frame, so the held-out latent-code score below is a reconstruction diagnostic, not a deployable next-frame forecast.

The use of KK-means here deserves a word of explanation. The published latent-action model learns an encoder, codebook, and reconstruction decoder; our toy replaces that problem with clustering. Displacement is one of five possible vectors: up, down, left, right, or zero for a blocked move. We ask for four clusters, so zero displacement must share a code with some movement case. The clustering step never sees button labels, but our hand-designed displacement extractor already builds in knowledge of agent position and motion. Its centers are movement or mixed movement/no-op prototypes, not proof that a generic inverse model would discover the same actions. K-means does not refine its codes using a learned reconstruction or next-frame prediction objective.

In[7]:
Code
import numpy as np
from sklearn.cluster import KMeans


def agent_position(frame):
    return np.array(np.unravel_index(np.argmax(frame), frame.shape))


# Keep complete trajectories together in the train/test split.
split_rng = np.random.default_rng(17)
trajectory_perm = split_rng.permutation(n_traj)
split_traj = int(0.8 * n_traj)
train_trajectories = trajectory_perm[:split_traj]
test_trajectories = trajectory_perm[split_traj:]
train_idx = np.concatenate(
    [np.arange(t * horizon, (t + 1) * horizon) for t in train_trajectories]
)
test_idx = np.concatenate(
    [np.arange(t * horizon, (t + 1) * horizon) for t in test_trajectories]
)

# Translation-invariant displacement between consecutive agent cells.
displacements = np.array(
    [
        agent_position(next_observations[i]) - agent_position(observations[i])
        for i in range(len(observations))
    ],
    dtype=np.float32,
)
kmeans = KMeans(n_clusters=K, n_init=10, random_state=0).fit(
    displacements[train_idx]
)
latent_actions = kmeans.predict(displacements)

action_index = {name: i for i, name in enumerate(ACTION_NAMES)}
true_idx = np.array([action_index[a] for a in true_actions])

confusion = np.zeros((K, K), dtype=int)
for lt, tt in zip(latent_actions, true_idx):
    confusion[lt, tt] += 1

# Calibrate on training labels only; test labels are used for evaluation.
confusion_train = np.zeros((K, K), dtype=int)
for lt, tt in zip(latent_actions[train_idx], true_idx[train_idx]):
    confusion_train[lt, tt] += 1
action_perm = np.argmax(confusion_train, axis=0)

The full-data confusion matrix below is descriptive: its held-out labels do not fit K-means or calibrate the lookup. Code labels can be permuted without changing the clustering. Blocked button presses all have zero displacement, so one code can mix those moves with a direction.

Out[8]:
Console
Confusion matrix (rows = latent code, columns = true action):
[[  0   0 606   0]
 [179 736 148 182]
 [  0   0   0 605]
 [544   0   0   0]]
Training-only calibrated true-action-to-code lookup: [3 1 0 2]

Notice that even when the pattern is clean, the naming of the codes is arbitrary. Nothing in the clustering objective forces code 0 to mean "up" rather than "down." Our action_perm lookup uses only training action labels to assign one code to each button by column majority. It happens to be one-to-one in this seeded run, but that rule does not guarantee a permutation in other datasets. This supervised calibration supplies extra information for prospective tests; it is not part of label-free code discovery.

The confusion matrix is also useful because it exposes the no-op wrinkle directly. If many zero-delta transitions are assigned to a single latent code, that code may look like a direction in the confusion matrix while actually absorbing all the blocked moves. In a larger model, this could mean that one latent action becomes a "nothing happened" code, which is useful but not a direction. That is not a failure. It is the inverse model doing its job: summarizing change, including the absence of change. The evaluation question is whether the codes are useful for prediction and control, not whether they line up one-to-one with human button names.

Playable Environment Generation

Latent action discovery is only half of what Genie does. The other half is generating the next frame. In the original Genie report, the model is a stack of three components:

  • Spatiotemporal tokenizer. A VQ-VAE style model that turns a video clip into a compact grid of discrete visual tokens. This is what makes the video cheap to model autoregressively, in the same spirit as the tokenized game and environment models from Tokenized Game and Environment Models.
  • Latent action model. Its encoder reads raw video frames, including preceding history, and proposes a discrete code for each transition. The pair-conditioned qϕq_\phi above is our simpler abstraction; the original paper reports raw-pixel input and treats tokenized input as a separate ablation.
  • Dynamics model. An autoregressive transformer that predicts the next set of visual tokens from the available history of visual tokens and latent actions. At inference time, this is what produces the next frame.

These three components play distinct roles. The tokenizer compresses video into discrete tokens; the latent-action model supplies a learned conditioning code; and the dynamics model predicts future tokens. In the original Genie training schedule, the tokenizer is trained first, then the latent-action and dynamics models are co-trained with stopped gradients on the dynamics model's action input. The tokenizer is not jointly updated with the other two throughout training.

Conceptually, the model is a conditional video generator: the conditioning signal is a sequence of latent actions, and the generated content is a sequence of visual tokens. If the codes preserve useful change information and the dynamics model uses them, changing a code can change the video in structured ways. An informative code can still be ignored by the dynamics model, leaving a decorative side channel. That distinction is not visible from attractive samples alone: the video prior may generate plausible frames without responding to actions. This is why action responsiveness has to be tested with branches and interventions rather than inferred from sample quality.

For an externally supplied, open-loop sequence of actions u1:T−1u_{1:T-1}, let pθrollp_\theta^{\mathrm{roll}} denote the model's causal rollout distribution, not the observational conditional from a policy that chooses later actions after seeing frames. The first equality below applies the conditional chain rule to that rollout distribution. The second is the model's nonanticipating generation rule: only actions already applied enter the next-frame kernel.

pθroll(z1:T∣u1:T−1)=∏t=1Tpθroll(zt∣z<t,u1:T−1)(conditional chain rule)=∏t=1Tpθ(zt∣z<t,u<t)(causal rollout kernel)\begin{aligned} p_\theta^{\mathrm{roll}}(z_{1:T}\mid u_{1:T-1}) &= \prod_{t=1}^{T}p_\theta^{\mathrm{roll}}(z_t\mid z_{<t},u_{1:T-1}) && \text{(conditional chain rule)} \\ &= \prod_{t=1}^{T}p_\theta(z_t\mid z_{<t},u_{<t}) && \text{(causal rollout kernel)} \end{aligned}

Here TT is the number of frames, z<tz_{<t} is the visual-token history before frame tt, and u<tu_{<t} contains the actions already applied. At t=1t=1, the frame is initialized from a prompt in interactive use; the factorization is a useful model abstraction, not a claim that Genie samples an unconstrained first frame.

The distinction between the two equalities matters: conditioning on actions is not something the ordinary chain rule can simply add to an unconditional p(z1:T)p(z_{1:T}). We define a causal rollout model and impose temporal ordering. Thus an action supplied after frame tt affects the distribution of frame t+1t+1, not the already generated frame tt. Under a feedback policy, a future chosen action may reveal information about frame tt, so the observational conditional p(z1:T∣u1:T−1)p(z_{1:T}\mid u_{1:T-1}) need not have this second factorization; interactive generation instead applies the one-step causal kernel after each chosen action.

To make the dataflow legible, let us write the forward model as a factorization:

pθroll(z1:T∣u1:T−1)=∏t=1Tpθ(zt∣z<t, u<t)p_\theta^{\mathrm{roll}}(z_{1:T} \mid u_{1:T-1}) = \prod_{t=1}^{T} p_\theta(z_t \mid z_{<t},\, u_{<t})

where:

  • z1:Tz_{1:T}: the visual tokens for frames 11 through TT
  • u1:T−1u_{1:T-1}: the latent actions for the transitions between those frames
  • pθrollp_\theta^{\mathrm{roll}}: the joint rollout law when the action sequence is externally supplied
  • pθ(zt∣z<t, u<t)p_\theta(z_t \mid z_{<t},\, u_{<t}): the next-frame conditional distribution, itself implemented with a masked-token dynamics model in the original Genie

Here ztz_t contains the visual tokens of frame tt, and u<tu_{<t} is the sequence of actions already applied. The rollout is autoregressive across frames, while Genie's MaskGIT dynamics samples tokens within a frame through iterative masked refinement. Recursive generation can propagate errors, though whether they grow depends on the learned dynamics and the states visited.

The strength of this factorization is that it can represent long-range dependencies in principle. If a frame contains an object that reappears later, the model can condition on the earlier frame and reproduce the object. The weakness is that the dependencies are only as good as the learned attention pattern. If the model does not learn to attend to the right earlier frames, the object can disappear. The autoregressive form also means that generation is sequential: frame t+1t+1 cannot be produced until frame tt is available. That is fine for a video, but for real-time interaction it imposes a latency budget that the model has to meet at every step.

The original Genie trains its tokenizer first, then co-trains a reconstruction-based latent-action model and masked-token dynamics model with separate losses. Our toy dispenses with tokenization and joint training. It trains a small predictor of ot+1o_{t+1} from oto_t and either a displacement-cluster code or the recorded action, using the same whole-trajectory split for both predictors.

The recorded actions permit two controls that a video-only dataset lacks. First, we train a true-action predictor as an oracle-label baseline. Second, we calibrate a button-to-code lookup using training labels only. At test time, the hindsight code inferred from a frame pair measures how well the code represents an observed transition; the calibrated code available from a chosen button measures prospective prediction. Comparing those two latent-code inputs is essential: the first knows the next frame when constructing its input, whereas the second does not.

In[9]:
Code
from sklearn.metrics import mean_squared_error
from sklearn.neural_network import MLPRegressor

obs_flat = observations.reshape(len(observations), -1)
nxt_flat = next_observations.reshape(len(next_observations), -1)

X_latent = np.concatenate([obs_flat, one_hot(latent_actions)], axis=1)
X_true = np.concatenate([obs_flat, one_hot(true_idx)], axis=1)
Y = nxt_flat

We train two models with the same architecture. One sees the original action labels as an oracle ata_t, and the other sees the displacement-cluster codes utu_t. Both have the same input dimension and training split. A difference in held-out error can indicate differences in the conditioning signal, but stochastic optimization and limited capacity also matter; one split and seed cannot isolate the representation's causal effect.

The toy predictor is an MLP that takes a flattened frame and a one-hot code and predicts a flattened next frame. It is not a transformer over tokens. Holding its architecture and split fixed makes the comparison informative, but two separately fitted networks can differ because of optimization as well as their conditioning codes. Nor does a small MSE gap guarantee that either model supports control.

In[10]:
Code
model_latent = MLPRegressor(
    hidden_layer_sizes=(32, 32),
    max_iter=150,
    random_state=0,
    early_stopping=True,
)
model_latent.fit(X_latent[train_idx], Y[train_idx])  # noqa: E501

model_true = MLPRegressor(
    hidden_layer_sizes=(32, 32),
    max_iter=150,
    random_state=0,
    early_stopping=True,
)
model_true.fit(X_true[train_idx], Y[train_idx])
In[11]:
Code
pred_latent = model_latent.predict(X_latent[test_idx])
calibrated_test_codes = action_perm[true_idx[test_idx]]
X_latent_calibrated_test = np.concatenate(
    [obs_flat[test_idx], one_hot(calibrated_test_codes)], axis=1
)
pred_latent_calibrated = model_latent.predict(X_latent_calibrated_test)
pred_true = model_true.predict(X_true[test_idx])
pred_persist = obs_flat[test_idx]

err_latent = mean_squared_error(Y[test_idx], pred_latent)
err_latent_calibrated = mean_squared_error(Y[test_idx], pred_latent_calibrated)
err_true = mean_squared_error(Y[test_idx], pred_true)
err_persist = mean_squared_error(Y[test_idx], pred_persist)
Out[12]:
Console
One-step frame MSE on held-out trajectories:
  Persistence:           0.0232
  True action:           0.0130
  Latent code, hindsight: 0.0131
  Latent code, calibrated:0.0134

The persistence baseline predicts that nothing changes. It is competitive with a weak predictor because some sampled actions are blocked by walls or grid boundaries. Its error reflects the moved-versus-stayed outcomes, though MSE also depends on the frame encoding. Beating persistence shows improved one-step prediction on this split; it does not by itself show useful control.

On the held-out trajectories, persistence has MSE around 0.02320.0232, while all learned variants are near 0.0130.013. The hindsight latent code is inferred from the actual next frame and is therefore a representation diagnostic, not a fair deployable forecasting input. Using the training-calibrated button-to-code lookup instead gives a slightly higher MSE than the true-action predictor on this split. These small differences still do not establish reliable control; the later action-adherence and goal-reaching tests probe that directly.

Out[13]:
Visualization
Bar chart of frame MSE for persistence, true-action, hindsight latent-code, and calibrated latent-code inputs on held-out trajectories.
One-step frame MSE on trajectory-disjoint test data. Hindsight latent codes use the true next frame to infer the code and are diagnostic only; calibrated latent codes use a train-only button-to-code lookup and are available prospectively. All learned variants beat persistence, but small MSE gaps do not establish control quality.

The cluster codes differ from four clean directional labels. One large code mixes blocked moves and a movement direction, while three others align more cleanly. Similar one-step MSE hides that difference, especially when the code was inferred with hindsight. In video-only settings without matched action labels or a reference simulator, evaluation relies more heavily on interventions and behavior tests. Our toy has the true simulator, so we compare both one-step and multi-step predictions against it.

Interactive Video Rollouts

An interactive environment adds a feedback loop to prediction. At each step the model receives an action and produces an observation; the user or agent sees that observation before choosing the next action. An ordinary generated clip can also feed predicted frames forward, but its action sequence is not chosen in response to those frames. This two-way exchange is what the following rollout tests.

Interactive Rollout

An interactive rollout is a sequence of model-generated frames produced by feeding the model's own previous output back as the next input, with user or policy actions interleaved. It is a closed-loop system, not a single conditional prediction.

The loop has a recurring structure:

  1. Initialize from an observed frame or an image prompt. In our toy, this is a rendered grid cell.
  2. Receive an input from a user or agent. The original Genie accepts one of its learned latent-action codes. Google DeepMind reports keyboard and mouse inputs for Genie 2 and real-time navigation for Genie 3, but their interface and training mechanisms should not be inferred from Genie's latent-action architecture.
  3. Interpret the input in the model's control representation. In the original Genie, a human must learn or calibrate the meaning of a latent code; the later releases do not publicly document an equivalent code-to-key lookup.
  4. Generate the next observation within a latency budget. Genie 3's announcement reports 24 frames per second at 720p, with consistency for a few minutes. Genie 2's announcement also describes a lower-quality distilled variant playable in real time; the later release's headline figures should not be projected backward onto it.
  5. Append the generated observation to the model's context and repeat.

Each of these steps can fail in a different way. Initialization can fail if the prompt is out of distribution. Input mapping can fail if the alignment from human commands to latent codes is wrong. Generation can fail if the model produces an incoherent frame. Context append can fail if the model's context window is too short or if the generated frame is treated as ground truth when it is not. Repetition can fail through drift, where small errors accumulate over many steps. The loop is simple, but it has more failure surfaces than a one-step predictor.

Both an autoregressive video clip and an interactive rollout feed generated outputs back into the model. Interaction adds action choices that can move the process into new states. An error at step tt can then affect step t+1t+1, but need not always be amplified. As in Video Generators as Candidate World Models, let o^t\hat{o}_t be the predicted observation and oto_t the reference under the same action sequence. Their difference is εt=o^t−ot\varepsilon_t = \hat{o}_t-o_t. A local linearization illustrates one possible error propagation pattern:

εt+1≈Jt εt+νt\varepsilon_{t+1} \approx J_t \, \varepsilon_t + \nu_t

where:

  • εt\varepsilon_t: the error at step tt between the model's predicted observation and the true observation
  • JtJ_t: the Jacobian of the model's prediction map with respect to the previous observation; it scales how strongly errors are transferred forward
  • νt\nu_t: new noise injected at step tt from stochastic sampling and approximation error

The Jacobian JtJ_t captures local sensitivity to the previous observation. A bound below one contracts an existing perturbation without new error, whereas fresh error νt\nu_t can sustain a nonzero error even in a contractive system. A norm above one permits amplification but does not ensure monotonic growth along every direction or noisy trajectory. Any multi-step autoregressive clip can exhibit this feedback problem; persistent interaction raises the stakes because the model must continue responding consistently to new actions. The recursion is a local illustration, not a measured stability result for Genie.

We can compare three scalar gains under the same seeded noise sequence, so differences between the plotted paths come from the gain rather than different random shocks.

In[14]:
Code
error_rng = np.random.default_rng(3)
n_steps = 30
noise_std = 0.05
jacobians = [0.9, 1.0, 1.05]
shared_noise = error_rng.normal(0, noise_std, size=n_steps)
error_paths = {}
for J in jacobians:
    eps = np.zeros(n_steps)
    eps[0] = 0.1
    for t in range(1, n_steps):
        eps[t] = J * eps[t - 1] + shared_noise[t]
    error_paths[J] = eps
Out[15]:
Visualization
Line chart of signed error under three scalar feedback multipliers.
Signed scalar error in three recursions over 30 steps, with multipliers 0.9, 1.0, and 1.05 driven by the same seeded noise. The common shocks isolate the effect of feedback gain on these sample paths; even the subunit multiplier need not make total noisy error decay monotonically.

There is also a statistical version of the problem. During one-step training the toy predictor sees real frames; during rollout it feeds its own generated frames back. Those frames may be outside the training distribution, so the one-step training loss has not directly penalized errors on them. This train–rollout mismatch is often called exposure bias. It applies to autoregressive video generation as well as interactive use; sustained interaction makes it especially consequential.

Interactive use makes persistence and action consequences easier to probe. If you walk past a rock and then turn around, the rock should still be there; a wall should block movement, and a collected item should remain collected. Video can contain and teach these patterns, including revisits, but a plausible short clip does not establish that a model will preserve them under long, user-chosen action sequences. They need direct tests at the relevant horizon.

These are rules about transitions under actions, often discussed as affordances or invariants. A wall blocks movement; a rock remains in place when the camera turns; an inventory item persists. A model can render a plausible wall but still predict movement through it. Interaction lets us test that rule directly, whereas a good-looking frame alone cannot.

Let us build the rollout loop for our toy and see how it behaves over a fixed horizon.

In[16]:
Code
rng_test = np.random.default_rng(1)
n_test_traj = 30

test_obs_seq = []
test_act_seq = []
for _ in range(n_test_traj):
    pos = (int(rng_test.integers(GRID)), int(rng_test.integers(GRID)))
    while pos in WALLS:
        pos = (int(rng_test.integers(GRID)), int(rng_test.integers(GRID)))
    obs_seq = [render(pos)]
    act_seq = []
    for _ in range(horizon):
        a = ACTION_NAMES[int(rng_test.integers(4))]
        pos = step(pos, a)
        act_seq.append(a)
        obs_seq.append(render(pos))
    test_obs_seq.append(np.array(obs_seq))
    test_act_seq.append(act_seq)
In[17]:
Code
def rollout_model(model, start_obs, action_seq, use_perm):
    obs = start_obs.copy()
    seq = [obs.copy()]
    for a in action_seq:
        a_idx = action_index[a]
        if use_perm:
            a_idx = action_perm[a_idx]
        x = np.concatenate(
            [obs.reshape(1, -1), one_hot(np.array([a_idx]))], axis=1
        )
        obs = model.predict(x).reshape(start_obs.shape)
        seq.append(obs.copy())
    return np.array(seq)


horizon_errors_latent = np.zeros(horizon + 1)
horizon_errors_true = np.zeros(horizon + 1)
horizon_errors_persist = np.zeros(horizon + 1)

for obs_seq, act_seq in zip(test_obs_seq, test_act_seq):
    pred_true = rollout_model(model_true, obs_seq[0], act_seq, False)
    pred_latent = rollout_model(model_latent, obs_seq[0], act_seq, True)
    pred_persist = np.tile(obs_seq[0][None, :, :], (horizon + 1, 1, 1))

    for t in range(horizon + 1):
        horizon_errors_true[t] += np.mean((pred_true[t] - obs_seq[t]) ** 2)
        horizon_errors_latent[t] += np.mean((pred_latent[t] - obs_seq[t]) ** 2)
        horizon_errors_persist[t] += np.mean(
            (pred_persist[t] - obs_seq[t]) ** 2
        )

horizon_errors_true /= n_test_traj
horizon_errors_latent /= n_test_traj
horizon_errors_persist /= n_test_traj

The rollout curves expose the gap between one-step and recursive prediction. In this seeded run, both learned predictors start at zero error, rise after a few steps, and finish around 0.0160.016–0.0170.017. Persistence repeats the first frame, yet its error is not flat: it changes as the true agent moves, remaining around 0.0230.023–0.0290.029. The curves average 30 fresh trajectories; they do not imply that every trajectory's error grows monotonically.

Out[18]:
Visualization
Line chart of frame MSE over rollout horizon for three models.
Recursive rollout frame MSE against the toy simulator as a function of horizon. The two learned predictors' errors rise early and then vary slowly; persistence repeats its initial frame but its error varies as the true agent moves. These curves do not show guaranteed monotonic growth.

Now let us isolate the concept of an action branch: two rollouts from the same initial observation, using different action sequences. A world model that is truly responsive should produce visibly different trajectories for different actions. A world model that is action-insensitive will produce nearly the same output regardless of what you ask for, and the branches will stay close together.

Action branches test whether output depends on input actions. A model that ignores the action produces identical branches from the same start, but a responsive model's separation need not increase monotonically: branches can converge, hit walls, or collapse under an imperfect predictor. Separation is not a test of correctness, either; both branches can be wrong. We therefore compare it with simulator-grounded metrics rather than treating any nonzero gap as proof of playability.

In[19]:
Code
branch_start = (4, 4)
branch_a = ["left"] * 4
branch_b = ["right"] * 4

initial = render(branch_start)
rollout_a = rollout_model(model_true, initial, branch_a, False)
rollout_b = rollout_model(model_true, initial, branch_b, False)

branch_sep = np.array(
    [
        np.mean((rollout_a[t] - rollout_b[t]) ** 2)
        for t in range(len(branch_a) + 1)
    ]
)
Out[20]:
Visualization
Line chart showing divergence of two predicted branches over four steps.
Branch separation for left-versus-right inputs from the same initial frame. The predicted frames differ after one step, but their MSE does not grow monotonically over this four-step rollout; responsiveness alone does not establish correct control.

Here the branches separate on the first step, then partly converge. The model responds differently to these inputs, but the tiny, nonmonotone gap does not establish accurate movement. Near-zero separation across many states and action pairs whose true outcomes differ would be evidence that the model ignores control. At a blocked corner, by contrast, two different actions should legitimately have the same outcome. This is the distinction the matched-branch test in Video Generators as Candidate World Models also needs.

If a model ignores its action channel across states where actions truly differ, a planner cannot use it to predict those differences. Branch comparison can reveal that failure, but nonzero separation is only evidence of sensitivity, not controllability: noise or systematically wrong effects can also separate branches. Pair it with reference transitions or task outcomes whenever those are available.

Evaluation of Control and Consistency

Evaluating an open-ended generated environment is harder when no reference simulator covers its possible action branches. Where a real environment and fixed action sequence exist, as in our toy, we can measure MSE against the matching reference rollout. We also use behavioral proxies because pixel similarity alone misses control and persistence. This chapter computes six measures:

  • Action adherence. Given a requested action, does the model's prediction reflect that action in a way consistent with the environment's mechanics? In our grid world, we can check whether the agent in the predicted frame ends up in the position that the true environment would produce.
  • One-step transition error. How close is a single model prediction to the true next frame, when the input is a real frame and a real action? This is the easiest metric to compute and the least informative about interactive quality.
  • Multi-step rollout error. How close is a sequence of model-generated frames to the true sequence under the same action sequence? This captures compounding drift.
  • Branch separation. How different are rollouts from the same start state under different actions? This captures action responsiveness.
  • Revisit consistency. If the agent leaves a state and returns to it, does the model predict the same observation? This captures the model's ability to maintain persistent structure.
  • Downstream control. Does a specified planner reach a goal when it selects actions using the model's predictions? This tests the predictor and planner together, not the model in isolation.

No single one of these proxies is sufficient. Action adherence can be inflated by an action-agnostic predictor of a globally common outcome, especially when no-ops dominate. One-step error can be low for a model that ignores actions entirely for the same reason. Multi-step rollout error can be dominated by drift. Branch separation can be produced by noise. Revisit consistency can be satisfied by a model that outputs a constant. The proxies are useful only together, and even together they do not fully characterize playability. They are a practical compromise: cheap to compute, interpretable, and correlated with the properties that matter.

We already computed one-step error, multi-step rollout error, and branch separation. Let us add the other three.

In[21]:
Code
def predicted_position(frame):
    r, c = np.unravel_index(np.argmax(frame), frame.shape)
    return (r, c)


def action_adherence_rate(model, obs_seq, act_seq, use_perm):
    hits = 0
    total = 0
    for t in range(len(act_seq)):
        obs = obs_seq[t]
        a = act_seq[t]
        expected_next = obs_seq[t + 1]
        a_idx = action_index[a]
        if use_perm:
            a_idx = action_perm[a_idx]
        x = np.concatenate(
            [obs.reshape(1, -1), one_hot(np.array([a_idx]))], axis=1
        )
        pred = model.predict(x).reshape(obs.shape)
        if predicted_position(pred) == predicted_position(expected_next):
            hits += 1
        total += 1
    return hits / total


adherence_true = np.mean(
    [
        action_adherence_rate(model_true, obs_seq, act_seq, False)
        for obs_seq, act_seq in zip(test_obs_seq, test_act_seq)
    ]
)
adherence_latent = np.mean(
    [
        action_adherence_rate(model_latent, obs_seq, act_seq, True)
        for obs_seq, act_seq in zip(test_obs_seq, test_act_seq)
    ]
)

Revisit consistency asks a slightly different question. Suppose the agent moves away from a cell and comes back. Does the model put the agent back in the same cell? We test this with a short sequence that leaves the start position and returns to it. Formally, if the sequence of actions a1:ka_{1:k} maps the true state back to the starting position s0s_0, then a revisiting-consistent model should predict an observation close to o0o_0 after feeding its own outputs back k times. The MSE we compute next measures how far the final prediction drifts from the original observation.

Revisit consistency is a minimal test of persistence, not proof of scene memory. Our selected action sequences return the true agent to its start; we compare each model's final recursive prediction with the initial frame. A constant-output model could score well on this comparison while failing action adherence and branch tests, so the measures belong together. The persistence baseline is exactly zero by construction and is not the best interactive model.

In[22]:
Code
revisit_starts = [(1, 5), (3, 2), (6, 5), (0, 3), (0, 0)]
seq_forward = ["right", "right", "left", "left"]

revisit_true = []
revisit_latent = []
revisit_persist = []

for start_pos in revisit_starts:
    initial = render(start_pos)
    pred_true = rollout_model(model_true, initial, seq_forward, False)[-1]
    pred_latent = rollout_model(model_latent, initial, seq_forward, True)[-1]
    revisit_true.append(np.mean((pred_true - initial) ** 2))
    revisit_latent.append(np.mean((pred_latent - initial) ** 2))
    revisit_persist.append(np.mean((initial - initial) ** 2))

revisit_true = float(np.mean(revisit_true))
revisit_latent = float(np.mean(revisit_latent))
revisit_persist = float(np.mean(revisit_persist))

Finally, we test downstream control. Give the model a target cell, use it to plan a sequence of actions greedily (each step, pick the action whose predicted next frame is closest to the target), execute those actions in the true environment, and check whether the agent actually reaches the target. Formally, for a start state s0s_0 and target s⋆s^{\star}, the planner selects at each step the action ata_t that minimizes the Manhattan distance between the predicted position pos(o^t+1)\mathrm{pos}(\hat{o}_{t+1}) and s⋆s^{\star}. This is a crude but honest test of whether the model is useful for control, not just for prediction.

The planner is greedy: at each real environment state it rerenders the true observation, scores four one-step model predictions, executes the selected action, and repeats. Thus it does replan after each step, but has no multi-step lookahead and is given privileged access to the true current state. A poor result may reflect model error, the greedy rule, target geometry, or the small trial set; it does not prove a stronger planner would fail. The test asks only whether this particular predictor-plus-controller reaches a target within 20 steps.

In[23]:
Code
def plan_and_execute(model, start_pos, target_pos, use_perm):
    pos = start_pos
    for _ in range(20):
        if pos == target_pos:
            return True
        obs = render(pos)
        best_action, best_dist = None, np.inf
        for a in ACTION_NAMES:
            if model is None:  # oracle dynamics, same greedy rule
                pred = render(step(pos, a))
            else:
                a_idx = action_index[a]
                if use_perm:
                    a_idx = action_perm[a_idx]
                x = np.concatenate(
                    [obs.reshape(1, -1), one_hot(np.array([a_idx]))], axis=1
                )
                pred = model.predict(x).reshape(obs.shape)
            r, c = predicted_position(pred)
            d = abs(r - target_pos[0]) + abs(c - target_pos[1])
            if d < best_dist:
                best_dist = d
                best_action = a
        pos = step(pos, best_action)
    return pos == target_pos


# The control-success values below are computed once and reused by the
# downstream figure cell.
rng_control = np.random.default_rng(2)
n_trials = 15
control_true = 0
control_latent = 0
control_oracle = 0
for _ in range(n_trials):
    start = (int(rng_control.integers(GRID)), int(rng_control.integers(GRID)))
    while start in WALLS:
        start = (
            int(rng_control.integers(GRID)),
            int(rng_control.integers(GRID)),
        )
    target = (int(rng_control.integers(GRID)), int(rng_control.integers(GRID)))
    while target in WALLS or target == start:
        target = (
            int(rng_control.integers(GRID)),
            int(rng_control.integers(GRID)),
        )
    if plan_and_execute(model_true, start, target, False):
        control_true += 1
    if plan_and_execute(model_latent, start, target, True):
        control_latent += 1
    if plan_and_execute(None, start, target, False):
        control_oracle += 1

control_true /= n_trials
control_latent /= n_trials
control_oracle /= n_trials

The grouped bars compare action adherence and downstream control. In this seeded run, adherence is about 0.57 for the true-action predictor and 0.52 for the calibrated latent-code predictor. Each learned-model controller reaches one of 15 targets. The same greedy rule with exact one-step simulator dynamics reaches 10 of 15. Greedy search and walls still limit that oracle. On these trials, the learned predictors add a substantial failure source beyond the planner.

Out[24]:
Visualization
Bar chart of adherence for two learned predictors and control success for those predictors plus an oracle-dynamics greedy baseline.
Action adherence is lower for the calibrated latent-code predictor. Both learned-model greedy controllers solve one of 15 targets, while the same greedy rule using exact simulator transitions solves 10 of 15. The latent-control rate uses recorded action labels to calibrate its button-to-code lookup.

The corner exposes a local identifiability limit. At (0,0)(0,0), both "up" and "left" produce no movement. A frame-pair-only inverse model cannot determine which was pressed from that transition. The labeled predictor still receives different action inputs and may output slightly different frames, so the next diagnostic measures its learned prediction gap rather than a proof of perfect ambiguity.

In[25]:
Code
corner_obs = render((0, 0))
noop_up = rollout_model(model_true, corner_obs, ["up"], False)[-1]
noop_left = rollout_model(model_true, corner_obs, ["left"], False)[-1]
noop_gap = float(np.mean((noop_up - noop_left) ** 2))
Out[26]:
Console
Action adherence rate (predicted next cell matches true next cell):
  True-action model:   0.569
  Latent-action model: 0.516

Revisit consistency MSE (lower is better):
  Persistence:         0.0000
  True-action model:   0.0150
  Latent-action model: 0.0148

Downstream control success rate:
  True-action model:   0.07
  Latent-action model: 0.07
  Oracle dynamics:     0.67

Corner no-op gap (MSE between 'up' and 'left' predictions): 0.0008

At (0,0)(0,0), both "up" and "left" are blocked and have the same true next frame. No frame-pair-only inverse model can recover which of those buttons was pressed from that single transition. The supervised true-action predictor can still receive distinct button labels, and a finite fitted model may give slightly different outputs for them; the measured gap is therefore not guaranteed to be zero. The ambiguity is about inferring the cause from the observed pair, not about whether a labeled predictor can distinguish its inputs.

This is a limit on recovering the pressed button from that frame pair, not a claim that two buttons are globally indistinguishable. They can have different effects elsewhere, and a supervised predictor can receive their labels. More subtle data-coverage problems arise when rare state–action interactions are underrepresented. Latent-action discovery cannot infer a distinction that never appears in its observations; evaluation should separate local no-op ambiguity from global action responsiveness.

We can quantify the identifiability gap by comparing the corner with a free cell where the two actions have different effects.

In[27]:
Code
free_obs = render((1, 1))
free_up = rollout_model(model_true, free_obs, ["up"], False)[-1]
free_down = rollout_model(model_true, free_obs, ["down"], False)[-1]
free_gap = float(np.mean((free_up - free_down) ** 2))
Out[28]:
Visualization
Bar chart comparing action gap at blocked corner and free cell.
MSE between true-action-model predictions for two inputs at a blocked corner and two inputs at a free cell. The first pair has identical true outcomes; a frame-pair-only inverse model cannot recover which button was pressed there. The labeled MLP still outputs a small nonzero gap. The free-cell gap is also small, so this figure does not establish clean action discrimination.
Out[29]:
Visualization
Heatmap of latent codes versus true actions with counts.
Displacement-cluster codes (rows) versus recorded actions (columns) in the toy grid. Three codes align strongly with a movement direction; the fourth mixes actions, including blocked moves with zero displacement. The largest cluster contains 41.5% of transitions, so the matrix is not a one-to-one action-label recovery.

This heatmap makes two distinct ambiguities visible. The numeric code labels are arbitrary and could be permuted without changing the clustering. More importantly, the mixed row is not an arbitrary relabeling of one button: blocked actions from several buttons produce the same zero displacement, so their intent cannot be recovered from the observed frame pair. The clustering uses displacements computed from the two frames, not the button labels. Evaluation must distinguish arbitrary code identity from the information the observation actually contains.

The heatmap also makes the no-op absorption visible. If a row has counts spread across several true actions, that latent code is not cleanly aligned with a single direction. That can happen because the code is absorbing no-op transitions, or because the clustering has merged two similar directions. In a real latent-action model, this kind of analysis is harder because the codes are not directly interpretable as directions, but the same principle applies. You want to know which codes are well used, which are collapsed, and which correspond to mixtures of changes. The confusion matrix is a simple tool for asking those questions in a setting where you happen to have the true labels.

Limitations & Impact

The original Genie paper offers a concrete route from videos without recorded actions to frame-by-frame latent control. Its platformer and robotics examples make the idea tangible. They also expose a harder question: whether the learned interface stays useful when action meaning, memory, and control are tested beyond short demonstrations.

The first limitation is possible rollout drift. In our toy, recursive error rises early and then fluctuates; it does not increase at every step. Feeding predictions back creates a train–rollout mismatch and an opportunity for error propagation, but autoregressive generation does not mathematically require error to grow without bound. Better data, objectives, representations, and context can change the behavior. Long-horizon stability must be measured, not inferred from the architecture alone.

Part of what makes drift hard is that it is not uniform. Some aspects of a scene are easy to maintain, such as a static background, and some are hard, such as a moving object that interacts with other objects. The model may hold the easy parts together while the hard parts drift. This means that the perceived quality of an interactive environment can degrade unevenly: the scene may still look like the same room, but the physics of the objects in it may have gone wrong. For an evaluation, this means that a single rollout error curve may hide important structure. It can be useful to break error down by object, region, or property, although doing so requires a way to identify those parts in the generated frames.

The second limitation is object and landmark persistence. If you walk past a tree, turn around, and walk back, an interactive world should show the same tree. Video training can contain such revisits, but ordinary short-clip objectives may provide limited pressure to preserve off-screen details over long interaction horizons. Google DeepMind reports long-horizon memory examples for Genie 2 and improved consistency for Genie 3; those are company demonstrations, not an independently established guarantee across scenes and actions. When reading a release, ask what was measured, on which examples, and with which baselines.

Persistence is partly a memory problem. A transformer cannot directly attend to a frame that has fallen outside its available context, and retaining that frame inside the window helps only if the model learns to use it. A longer effective memory or a consistency-oriented objective might improve a particular system, but neither by itself guarantees accurate revisits. The relevant test is whether objects and state remain stable after leaving and returning under controlled actions.

The third limitation is weak latent-action semantics. A latent action inferred from video need not be a human-meaningful command. Our hand-designed grid displacement clusters produce three clear directional codes and one code that mixes a direction with blocked moves; even this favorable toy does not recover four buttons cleanly. Original Genie offers a consistent latent-code interface whose meanings a user can learn, while Google DeepMind reports direct controls for Genie 2 and real-time navigation for Genie 3. The later systems' controls do not establish that they use Genie's exact latent-action training or a simple code-to-key dictionary.

For a learned code interface, possible alignment methods include labeled command–code examples or a user-defined lookup. Our toy uses the former: its action_perm is computed from true action labels. A text interface might require a separate mapping mechanism, but the Genie 3 announcement does not establish that it maps language to Genie 1-style discrete codes. These are design possibilities, not published descriptions of all three releases.

The fourth limitation is interaction latency. Real-time interaction requires that the model generate the next frame within a small budget (typically tens of milliseconds). Autoregressive transformers over a large token grid are not naturally fast. Genie 3 reports real-time interaction, but the details of how that is achieved, and at what resolution and horizon, are important to check against the primary report. Latency of a few hundred milliseconds can make interactive controls feel unresponsive, depending on the task and interface. For an agent using the model for planning, latency translates directly into planning cost. If each imagined step takes a tenth of a second, a planning horizon of a hundred steps takes ten seconds, which may be too slow for real-time control.

Latency also interacts with the rollout loop. The loop requires generating a frame, appending it to the context, and generating the next frame. As the context grows, the cost of attention can grow, so the latency may increase over the course of a rollout unless the model uses a fixed-size context or some form of recurrent state. Keeping latency bounded over long rollouts is a systems problem as much as a modeling problem. It affects how long an interaction can last and how responsive it can feel. When a release claims real-time interaction, it is worth asking at what horizon, at what resolution, and with what hardware.

The fifth limitation is proxy evaluation ambiguity. For generated worlds without a matching reference simulator, evaluation relies more on action adherence, branch separation, revisit consistency, and downstream control. Our toy does have a reference, but even there each metric misses something: one-step position accuracy can favor common no-ops, branch separation can result from incorrect or noisy effects, revisit error can favor constant output, and control success depends on the planner and task distribution. Use several measures together and state what each actually tests.

There is a deeper evaluation problem as well. The properties we care about in an interactive environment, such as persistence, responsiveness, and controllability, are not directly observable from a single rollout. They are counterfactual properties: what would have happened if the action had been different, or if the agent had turned around instead of continuing. Evaluating them requires interventions and comparisons, not just observations. This is why branch separation and matched-initial-state tests are so important. They turn a counterfactual question into an observable comparison. The same logic applies to persistence: to test whether an object is remembered, you have to revisit it and compare. The evaluation suite is essentially a set of controlled interventions on the generated world.

The original Genie paper demonstrated a division of labor: learn a compact code for observed change, then condition a video-token dynamics model on that code. Its platformer and robotics examples show controllable generation without action labels in those studied settings, not a guarantee for arbitrary videos. Reliable playability and persistence remain open tests for the foundation world models in Cosmos and Omnimodal World Foundation Models.

The natural next step is to ask what happens when you scale the paradigm up, with more data, more modalities, longer context, and richer conditioning. That is the subject of the next chapter.

Summary

The original Genie demonstrated an interactive model with latent control codes learned from video without recorded actions. Its pretrained spatiotemporal tokenizer, reconstruction-trained latent-action model, and frame-autoregressive, masked-token dynamics model have different jobs. The latent-action encoder uses video context to infer a quantized code; at inference a user selects a code whose effect can be learned through interaction. Code names are arbitrary, and no general guarantee makes them equivalent to keyboard buttons.

An autoregressive video can also feed generated frames back into the model. Interactive use adds actions selected in response to those frames, bringing persistence, geometry, responsiveness, and latency under direct test. Action adherence, one-step error, multi-step rollout error, branch separation, revisit consistency, and downstream control measure different facets. No single metric establishes playability.

The toy grid world is a teaching model, not an implementation or benchmark of Genie. A hand-designed displacement extractor feeds K-means fitted on 160 training trajectories; 40 complete trajectories are held out. Three codes align with directions and a fourth mixes movement with blocked actions. A hindsight code inferred from a held-out next frame is only a reconstruction diagnostic. For prospective control, a train-only, label-calibrated button-to-code lookup supplies the code. The learned predictors have similar one-step MSE, but the calibrated latent predictor has lower action adherence. Both learned-model greedy controllers solve one of 15 trials, compared with 10 of 15 for the same rule using exact simulator transitions. Attractive prediction error and some branch separation are therefore not evidence of a reliably playable environment.

The chapter's main lesson is that predictive accuracy and controllability differ. A model can predict well on average while responding weakly to actions; a latent code can produce visible changes without mapping cleanly to a human command. Read Genie-style results by separating code discovery, frame generation, interface calibration, and behavioral evaluation. A demonstration of one does not establish the others.

Key Parameters

The key parameters for the toy latent-action and forward models are:

  • K (latent codebook size): Number of K-means displacement clusters; set to 4 even though movement plus no-op yields five distinct observed displacements.
  • GRID: Side length of the square grid world; set to 8.
  • horizon: Number of steps per generated trajectory; set to 15 for training and reused for rollouts.
  • hidden_layer_sizes: Width and depth of the forward-model MLP; set to (32, 32).
  • max_iter: Maximum training iterations for the MLP; set to 150.
  • random_state: Seed for KK-means and the MLP; set to 0 for reproducibility.
  • early_stopping: Whether the MLP holds out a validation split to stop early; set to True.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Genie, latent actions, and interactive generated environments.

Genie and Interactive Generated Environments

Question 1 of 80 of 8 completed
What defines a latent action in the Genie-style approach described in the chapter?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026genieinteractive, author = {Michael Brenndoerfer}, title = {Genie and Interactive Generated Environments}, year = {2026}, url = {https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Genie and Interactive Generated Environments. Retrieved from https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments
MLAAcademic
Michael Brenndoerfer. "Genie and Interactive Generated Environments." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments>.
CHICAGOAcademic
Michael Brenndoerfer. "Genie and Interactive Generated Environments." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Genie and Interactive Generated Environments'. Available at: https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Genie and Interactive Generated Environments. https://mbrenndoerfer.com/writing/genie-latent-actions-interactive-generated-environments

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.