Part of World Models Handbook
Examines NVIDIA's Cosmos world foundation models: video tokenization, physical AI data curation, conditional generation, guidance calibration, and simulation.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Cosmos and Omnimodal World Foundation Models
Imagine you are responsible for validating an automatic emergency braking policy before it ships. Controlled track tests can use pedestrian surrogates, and a simulator can supply repeatable labels and rare-event variations. Neither, on its own, establishes how a perception stack will respond to every real camera condition. Road recordings help ground the evaluation but may contain too few examples of a pedestrian stepping out from behind a parked van at dusk in the rain.
The tradeoff is concrete: the simulator gives you control and labels but may have an appearance gap. Real video may undersample the scenarios you most need to test. A conditional generator could help vary scene appearance and propose candidate rare scenarios. It cannot certify their frequency, physical validity, or label fidelity merely by producing plausible pixels.
World foundation models offer a third research route, and they are the subject of this chapter. Different Cosmos releases use broad physical-world pretraining, then offer different interfaces: Transfer models can adapt a simulator render, while supported Predict variants condition on prompts or observed clips. The output can supply candidates for synthetic data or evaluation, provided downstream checks reject outputs that violate geometry, labels, or task constraints. NVIDIA's Cosmos platform, introduced in 2025, is our running example; it is a family of releases, not one model with one architecture or license.
Those jobs may benefit from related visual and temporal pretraining, but they need not share weights, representations, or an interchangeable conditioning interface. A model trained on driving video may learn how wet asphalt reflects headlights or how a van occludes a pedestrian. Whether it can photorealize a simulator render or generate a controlled rare scenario depends on the particular release, on its training, and on its supported inputs. The foundation-model proposition is reuse and adaptation across such jobs, not proof that one generator solves both by switching a prompt.
A large model pretrained on broad physical-world data with the aim of capturing useful environmental dynamics and supporting downstream tasks such as video prediction, conditional simulation, scenario generation, policy evaluation, and synthetic data production. Whether it works for any particular task requires validation. The term emphasizes intended reuse across tasks, rather than success on a single narrow benchmark.
The word omnimodal should not be projected backward onto every Cosmos release. It fits Cosmos 3, whose documented model family jointly processes and generates language, images, video, audio, and action sequences. Older Predict and Transfer models expose narrower task-specific interfaces. Even with a unified interface, accepting a modality does not guarantee faithful use of it; contradictory text, depth, segmentation, or action inputs need explicit adherence tests.
This chapter follows video generators and interactive generated environments. The Cosmos family illustrates conditional simulation and data-generation workflows, while Cosmos 3 also reaches into action modeling. The next chapter develops that action interface directly in vision-language-action and world-action models. The boundary is therefore a teaching distinction, not an architectural separation that every current model obeys.
Three ideas recur throughout and are worth stating up front, because they are the lenses through which every Cosmos component should be read.
-
Compression shapes the interface. A video tokenizer sets a nominal latent-grid size and affects which details survive. The generator may patchify those latents further and add condition tokens, so its actual sequence length, training cost, memory use, and latency also depend on architecture and serving choices. A published context length must be checked for the specific model and tokenization path before converting it into seconds of video.
-
A generator models a conditional distribution of futures. A diffusion or autoregressive model can sample from an estimated , while a point forecast summarizes it with one value. One rollout is one possible future under that learned distribution; its event frequencies cannot be read as real-world probabilities without validation.
-
For labeled training data, control adherence matters alongside appearance. A photorealistic video that no longer matches simulator labels may harm perception training. Cosmos-Transfer1 offers spatially weighted conditioning for supported controls; that mechanism motivates adherence checks, not an assumption that every Cosmos release shares the same branch or preserves labels.
Which Cosmos release are we discussing?
“Cosmos” names a platform and several related but nonidentical model families. The original Cosmos world-foundation-model platform report describes video tokenizers, video generation, transfer, curation, and guardrails. It is the source for the platform-level pipeline below. Cosmos-Predict2.5 and Transfer2.5, reported later, include a flow-based Text2World/Image2World/Video2World generator and a control-net-style transfer system; those claims belong to the 2.5 family, not retroactively to Predict1. NVIDIA reports improvements and benchmarks for those releases, but their reported scores are not independent evidence that generated clips are safe or label-preserving for a new deployment.
Cosmos 3 is the release that makes omnimodal literal: NVIDIA documents a unified Mixture-of-Transformers architecture spanning text, image, video, audio, and action sequences, with Reasoner and Generator runtime surfaces. The Cosmos 3 model card describes an autoregressive tower for discrete text tokens and a diffusion tower for continuous multimodal outputs. That is not the same as placing every modality in one undifferentiated token stream, and a supported input–output combination still needs task-specific evaluation. The model card lists the Cosmos 3 weights under OpenMDW-1.1; the older Cosmos documentation lists the platform model weights under the NVIDIA Open Model License. Check the particular artifact's license rather than treating “Cosmos” as one license.
The worked calculations in this chapter are pedagogical, not Cosmos 3 benchmarks or reproductions. When we discuss a tokenizer, flow objective, control branch, or guardrail, we will name the relevant release or mark a proposed system design. This distinction matters most for runtime safety monitors: NVIDIA documents content guardrails for generation, whereas the probability-threshold monitor implemented below is our teaching example, not a shipped, calibrated Cosmos safety component.
Physical AI Data and Tokenization
The training corpus and video tokenizer both shape what a world foundation model can learn from observations. They are not the only determinants: architecture, objective, conditioning data, and post-training matter too. A curation pass that discards fast-motion clips may conceal a temporal codec's weakness on fast events; aggressive temporal compression can discard task-relevant information even in a high-frame-rate corpus. Evaluate data slices and reconstructions together.
What makes physical AI data different
"Physical AI" is a term for systems whose intelligence is expressed through bodies in the world: vehicles, manipulators, legged robots, drones, and warehouse robots. Their data can differ from typical web video in the prevalence of ego-motion and contact, the availability of action annotations and synchronized sensors, and the continuity of trajectories. Those differences affect data collection and the tests a world model needs.
- An agentic viewpoint. Many cameras move with a vehicle or robot, so ego-motion can account for substantial pixel change; independently moving objects and lighting matter too. Video-only prediction can learn correlations along recorded trajectories without identifying how futures change under alternative actions. That distinction matters when a model is used for planning.
- Physically plausible dynamics. Objects often persist across frames, occlusions may persist or resolve, and camera trajectories should remain coherent with the scene; abrupt physical motion can still be real. For a continuous physical rollout, cuts, edits, and cross-fades are discontinuities rather than environmental dynamics. Because edited web video often contains them, a corpus assembled without care can encourage the model to reproduce discontinuities.
- Action relevance. Near-misses, contact, occlusion boundaries, and degraded visibility can be consequential yet underrepresented in collections dominated by routine operation. If they are scarce in the target corpus, targeted collection or candidate generation may help. Measure coverage and choose sampling weights for the evaluation task rather than treating every recorded second as equally informative.
- Multimodal alignment. Some physical-AI datasets include depth, LiDAR, radar, IMU, odometry, or joint readings alongside video. Available sensors vary by platform; paired signals need appropriate spatial calibration and temporal synchronization. A one-frame offset can corrupt a supervised relationship without necessarily causing visible hallucination in every output.
- Condition diversity. Time of day, weather, sensor degradation, lens dirt, and seasonality change observations; some also change physics. Wet pavement, for example, changes tire–road friction and braking behavior. Appearance conditioning alone cannot stand in for dynamics coverage, so test both.
A pipeline that reconstructs real sensor data into a simulator representation (real2sim), varies controllable factors, then generates more realistic observations (sim2real). A conditional world model may help the second hop, but generation cost and label adherence have to be measured; simulator labels do not remain valid automatically.
The curation pipeline
Large raw video collections usually need curation before training. The original Cosmos report documents stages that affect the data distribution; the order and retention rules should be chosen for the intended task and measured rather than copied wholesale.
- Session and shot splitting. The original Cosmos pipeline extracts clips of about 2 to 60 seconds and detects shot boundaries using frame changes. A hard edit between unrelated views is not an ordinary physical transition; training across it can confuse temporal prediction. Overly short clips may omit the event evolution needed for a task.
- Quality filtering. Screen clips that are blurry, over- or under-exposed, mostly static, dominated by lens occlusion, or dominated by text overlays and picture-in-picture. Retention criteria depend on the target task: static or degraded clips can be useful for some evaluations. Burned-in captions and watermarks can be copied into generated outputs, so measure their prevalence rather than assuming a particular amplification factor.
- Captioning. Use a vision-language model to annotate clips or clip segments, ideally with temporal grounding. The original Cosmos report describes a caption per 256 frames, not necessarily one exhaustive description of every full clip. Captions support text conditioning and retrieval; their coverage and errors affect which scenarios users can request reliably.
- Deduplication. Check near-duplicates within the corpus and between training and evaluation sets. Leakage can inflate a reported generalization score; the score alone cannot reveal which examples leaked.
- Diversity adjustment, shuffling, and sharding. Balance the mixture across domains, sensor types, and environments. Then write it into shards that can be streamed by a distributed loader without random seeks into object storage. Mixture balance is a modelling decision in disguise: the relative weight of highway, urban, and rural footage helps shape which environments the model handles well.
The original Cosmos report describes a 20-million-hour raw collection curated into about 100 million clips. Its documented stages include shot splitting, filtering, annotation, deduplication, and sharding. The counts belong to that release; a new corpus needs its own retention criteria, labeling plan, and evaluation split.
Data coverage and architecture both constrain physical-AI world models. Large collections of synchronized, action-labeled interaction video are difficult to curate, especially for rare events. Real recordings, structured simulators, and neural generators can complement one another, but each brings distinct limitations in coverage, fidelity, or label validity.
Tokenization: from pixels to tokens
Here a video tokenizer acts as a learned compression codec. Its encoder maps a clip into latents
and a decoder maps them back, . This shape assumes suitable padding and integer downsampling factors; learned codecs need not correspond to independent, nonoverlapping blocks. The factors are often written . The nominal token count is
This counts latent-grid cells under the stated shape assumption, not necessarily transformer tokens: a generator may patchify the latent grid and add condition tokens. More cells generally increase compute, but the relationship depends on attention pattern, implementation, precision, hardware, and serving. Dense attention has a quadratic arithmetic term in sequence length; memory-optimized exact attention need not materialize a quadratic attention matrix. Neither a fixed horizon nor a frame rate follows from the codec factor alone.
The choice of what to compress is not symmetric either. Aggressive temporal compression can lose fast events, but a learned temporal codec may integrate information across all frames in its receptive field; it is not necessarily simple frame decimation. Whether a lane crossing survives must be tested on reconstructed clips. Aggressive spatial compression can lose brake lights, fingers, distant pedestrians, or lane boundaries. As with state abstractions in Part III, the relevant question is downstream sufficiency, not average fidelity alone.
Continuous versus discrete tokens
Continuous and discrete latents are two common design choices. They affect the generator objective and interface but do not uniquely determine the generator family.
Continuous tokenizers produce floating-point latents, sometimes with KL or perceptual regularization. They pair naturally with latent diffusion and flow models. A 16-dimensional latent stored in bf16 occupies 256 storage bits per token before any entropy coding; that is not a measured rate–distortion result. The PCA toy below uses float64 in memory unless explicitly cast. Continuous latents do not preclude likelihood modeling or reuse of all sequence-model tooling, but they do not provide a native categorical next-token objective.
Discrete tokenizers map latents to integer indices in a vocabulary of size . A fixed-width representation needs bits per index; entropy coding can yield a different average rate. A learned VQ-VAE codebook is one way to construct such indices, but the original Cosmos tokenizer uses finite scalar quantization (FSQ): six scalar coordinates with levels respectively define a 64,000-entry implicit vocabulary, not a learned VQ codebook. Discrete indices support categorical autoregression, subject to the model's training and interface. Temperature, top-/nucleus sampling, and KV caching may be used when the decoder supports them; text–video interleaving and speculative decoding require compatible designs. Quantization error and autoregressive drift still matter. Learned-codebook collapse and EMA updates are VQ-specific concerns, not requirements of Cosmos's FSQ tokenizer.
The original Cosmos platform reported continuous and discrete video tokenizers and causal variants. A causal tokenizer avoids dependence on future frames and is suited to streaming. A non-causal tokenizer can still be used online with a buffering window and corresponding delay; it is not categorically unusable. Rate–distortion, latency, and model compatibility all matter to the choice.
What to measure
A tokenizer is often evaluated on reconstruction, but reconstruction alone is not a sufficient selection target for a policy-facing world model. Average pixel fidelity says little about whether the compressed representation still supports the decisions you intend to make. The useful panel of measurements looks like this.
- Distortion: PSNR and SSIM for pixel/structural fidelity, LPIPS for learned perceptual similarity, and FVD for distribution-level video comparison rather than a per-frame reconstruction score. Each can miss a small safety-critical object, so slice by task-relevant events.
- Rate: measure actual compressed bits per second including codebook or entropy-coder overhead where relevant, then relate that to latent token count and model compute. A bf16 storage estimate is not a fair bitrate comparison to entropy-coded discrete tokens.
- Downstream utility: does a policy trained or evaluated on reconstructed observations behave the same as on originals? A tokenizer can score well on average reconstruction while losing a small decision-relevant detail. For a policy-facing use case, this check should carry substantial weight, despite the extra work of running the downstream task.
- Compression-shape sensitivity: measure distortion separately on fast-motion and fine-detail slices. Average PSNR can hide small-object failures. Slicing by motion magnitude and object size takes additional annotation or detection effort but makes the operating envelope clearer.
Worked example: tokenizing a 10-second clip
The arithmetic is simple enough to do on paper, and doing it once makes every infrastructure discussion in this chapter concrete.
Take a 10-second clip at 30 frames per second and 1280×720 resolution. That is frames. With an 8x16x16 tokenizer, the temporal axis becomes groups, the height becomes , and the width becomes . The nominal token count is , before text or other modalities. With 8x8x8, it quadruples to . Compare these counts with the actual context and attention scheme of the model being evaluated rather than an unspecified language-model context window.
These latent-grid counts illustrate a resource pressure, not the input length of a particular Cosmos transformer. Patchification, attention layout, and conditioning change the actual token sequence. Full attention has quadratic arithmetic in sequence length, although exact memory-efficient implementations can avoid a quadratic attention-matrix allocation. Diffusion or flow inference performs a model evaluation per sampler step, but step counts vary by release and distillation; Predict2.5, for example, reports a four-step distilled variant. A hypothetical 30 GPU-seconds for ten generated video seconds is a throughput example, not a measured wall-clock latency or Cosmos benchmark.
Let us build a miniature version of this pipeline. We will synthesize a tiny driving clip with a synchronized depth map and a semantic segmentation map, exactly the kind of multimodal bundle that physical AI corpora contain. We will use a fixed random seed so the example is deterministic and reproducible. Then we will run a deliberately naive tokenizer over it.
import numpy as np
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
def make_scene(n_frames=24, height=32, width=32, speed=0.55, seed=0):
"""Render a tiny schematic driving clip plus aligned depth and segmentation."""
r = np.random.default_rng(seed)
frames = np.zeros((n_frames, height, width, 3))
depth = np.zeros((n_frames, height, width))
seg = np.zeros((n_frames, height, width), dtype=np.int64)
horizon = height // 2
for t in range(n_frames):
frames[t, :horizon] = (0.62, 0.72, 0.88) # sky
frames[t, horizon:] = (0.24, 0.24, 0.27) # road surface
seg[t, horizon:] = 1 # class 1: road
depth[t, :horizon] = 1.00
depth[t, horizon:] = 0.50
offset = int(round(speed * t)) % 8 # lane marker pattern scrolls up
for row in range(horizon, height):
if (row + offset) % 8 < 3:
col = width // 2
frames[t, row, col - 1 : col + 2] = (0.90, 0.90, 0.85)
seg[t, row, col - 1 : col + 2] = 2
depth[t, row, col - 1 : col + 2] = 0.45
row_c = int(horizon + 14 - 4.0 * speed * t) # car moves toward horizon
if horizon < row_c < height - 2:
frames[t, row_c - 2 : row_c + 2, 12:20] = (0.85, 0.20, 0.15)
seg[t, row_c - 2 : row_c + 2, 12:20] = 3
depth[t, row_c - 2 : row_c + 2, 12:20] = 0.18
frames[t] = np.clip(
frames[t] + r.normal(0, 0.012, frames[t].shape), 0, 1
)
return frames, depth, segWe generate one clip to fit the codec and a second, slightly slower clip with different noise to evaluate it. This is a held-out synthetic clip, not a separate driving session or a test of deployment generalization. The car's fixed drawn width and depth values are schematic, not a physically consistent receding vehicle.
train clip : (24, 32, 32, 3) depth (24, 32, 32) seg (24, 32, 32) eval clip : (24, 32, 32, 3) (different speed, different noise) raw scalar values in one clip: 73,728 segmentation classes present : [0, 1, 2, 3] depth proxy range : 0.18 to 1.00 (unitless)
Next we implement a toy patchification step: chop the clip into non-overlapping tubes and flatten each tube into a vector. Production learned tokenizers need not operate on independent, non-overlapping raw-pixel blocks. The layout is still useful for understanding the token-count arithmetic.
def patchify(video, ct=4, ch=8, cw=8):
"""Non-overlapping tubes in time-row-column order; axes must divide evenly."""
T, H, W, C = video.shape
if T % ct or H % ch or W % cw:
raise ValueError("T, H, and W must be divisible by ct, ch, and cw")
x = video.reshape(T // ct, ct, H // ch, ch, W // cw, cw, C)
x = x.transpose(0, 2, 4, 1, 3, 5, 6)
return x.reshape(-1, ct * ch * cw * C)
def unpatchify(tokens, shape, ct=4, ch=8, cw=8):
"""Inverse of patchify for a compatible shape and complete token set."""
T, H, W, C = shape
if T % ct or H % ch or W % cw:
raise ValueError("T, H, and W must be divisible by ct, ch, and cw")
x = tokens.reshape(T // ct, H // ch, W // cw, ct, ch, cw, C)
x = x.transpose(0, 3, 1, 4, 2, 5, 6)
return x.reshape(T, H, W, C)Patchification alone compresses nothing: 96 tokens of 768 values each retain the original 73,728 scalar values. (24 frames divided by ct=4 gives 6 time groups, 32/8=4 rows, and 32/8=4 columns.) Compression comes from the bottleneck. PCA and -means provide a small, runnable stand-in for a learned encoder and vector-quantized codebook; their measured tradeoff in this synthetic example should not be assumed to match a production codec.
patches = patchify(train_frames, 4, 8, 8)
mean_patch = patches.mean(axis=0, keepdims=True)
pca = PCA(n_components=16, random_state=0).fit(patches - mean_patch)
train_latents = pca.transform(patches - mean_patch)
km = KMeans(n_clusters=64, n_init=4, random_state=0).fit(train_latents)
codebook = km.cluster_centers_We can compare nominal storage bits after specifying a precision. The NumPy arrays in this experiment are float64 in memory; the following accounting instead assumes an eight-bit RGB source, 16-bit stored continuous latents, and ideal fixed-width code indices. It omits codebook storage, metadata, and entropy coding, so these are not measured compressed file sizes.
patches per clip : (96, 768) continuous latent dim : 16 (24,576 bits, 24.0x compression) discrete codebook size : 64 (576 bits, 1024x compression) Under these storage assumptions, discrete indices use fewer bits. Next: reconstruction.
We encode the held-out clip and reconstruct it both ways to answer that question.
def psnr(reference, estimate, peak=1.0):
mse = np.mean((reference - estimate) ** 2)
return np.inf if mse == 0 else 10 * np.log10(peak**2 / mse)
eval_patches = patchify(eval_frames, 4, 8, 8)
eval_latents = pca.transform(eval_patches - mean_patch)
eval_codes = km.predict(eval_latents)
recon_continuous = unpatchify(
pca.inverse_transform(eval_latents) + mean_patch, eval_frames.shape, 4, 8, 8
)
recon_discrete = unpatchify(
pca.inverse_transform(codebook[eval_codes]) + mean_patch,
eval_frames.shape,
4,
8,
8,
)
psnr_continuous = psnr(eval_frames, recon_continuous)
psnr_discrete = psnr(eval_frames, recon_discrete)held-out PSNR, continuous tokens : 25.98 dB held-out PSNR, discrete tokens : 24.34 dB quantisation penalty : 1.64 dB A single operating point says little. What matters is the whole rate-distortion curve, because that curve is what a designer trades against the generative model's cost.
Codebook usage is another diagnostic. If most entries are never selected, the nominal vocabulary overstates effective capacity; an autoregressive model can still predict the used indices, but some representational capacity is wasted. Here we have only 96 training tokens for 64 codes, so usage counts are noisy and should not be mistaken for a production collapse test.
codebook size : 64 codes never selected in training : 0 codes never selected on held-out : 22 top code share of training tokens: 6%

The rate-distortion sweep below changes PCA latent dimension and discrete codebook size. On the continuous curve, adding dimensions spends more bits and improves held-out PSNR; reading from left to right on its compression-ratio axis reverses that direction. The discrete points occupy much higher compression ratios, so this toy does not compare the two tokenizers at equal rate. It shows how to measure a tradeoff, not which tokenizer family wins. A curve that flattens as rate increases suggests checking encoder capacity and evaluation slices before buying more bits.
sweep = []
for n_dims in (4, 8, 16, 32, 64):
sub_pca = PCA(n_components=n_dims, random_state=0).fit(patches - mean_patch)
lat_tr = sub_pca.transform(patches - mean_patch)
lat_ev = sub_pca.transform(eval_patches - mean_patch)
recon = unpatchify(
sub_pca.inverse_transform(lat_ev) + mean_patch,
eval_frames.shape,
4,
8,
8,
)
cont_ratio = (eval_frames.size * 8) / (eval_patches.shape[0] * n_dims * 16)
sweep.append(
("continuous", n_dims, 0, cont_ratio, psnr(eval_frames, recon))
)
for n_codes in (16, 32, 64):
sub_km = KMeans(n_clusters=n_codes, n_init=4, random_state=0).fit(
lat_tr
)
codes = sub_km.predict(lat_ev)
recon_q = unpatchify(
sub_pca.inverse_transform(sub_km.cluster_centers_[codes])
+ mean_patch,
eval_frames.shape,
4,
8,
8,
)
disc_ratio = (eval_frames.size * 8) / (
eval_patches.shape[0] * np.ceil(np.log2(n_codes))
)
sweep.append(
(
"discrete",
n_dims,
n_codes,
disc_ratio,
psnr(eval_frames, recon_q),
)
)
cont_curve = sorted(
(r, p) for kind, _, _, r, p in sweep if kind == "continuous"
)
disc_curve = sorted((r, p) for kind, _, _, r, p in sweep if kind == "discrete")
A frame comparison can reveal spatial structure hidden by average reconstruction metrics. A tokenizer can score reasonably on whole-frame PSNR while losing a small decision-relevant object. Below we inspect one held-out frame and its reconstruction from discrete codes.
frame_index = 4 # the car is visible in this held-out frame
original_frame = eval_frames[frame_index]
reconstructed_frame = np.clip(recon_discrete[frame_index], 0, 1)In this selected frame, the red car is visible in the original but largely lost in the discrete reconstruction. The nominal code index budget saves bits at this toy operating point, yet average PSNR alone does not tell us whether a small object remains usable for a decision. An object-specific reconstruction slice and a downstream task test would make that failure measurable.


Multimodal Generative Dynamics
With latents in hand, we can state a conditional prediction task. A video generator can be used as a world model when its generated futures are fit for the intended dynamics and decision question; architecture or visual quality alone does not establish that suitability.
The conditional video model as a dynamics model
Write the conditioning context as (text, a first frame, a layout, a control signal) and the observed history as . A video world model represents
the distribution over the next observation frames. Compare this with the latent state-space form we have used throughout the book, . Two differences stand out and both have consequences.
First, this factorization has no separately defined Markov state . A video transformer still has internal latent tokens and activations, but they need not be a compact, inspectable belief state. It conditions on observation history within a finite context, so object permanence across occlusion must be learned and evaluated rather than assumed. A generated car that fails to reappear after an occlusion is a temporal-consistency error; diagnosing its cause takes more than observing pixels alone. The temporal-state discussion in Part III provides a useful comparison.
Second, an explicit action variable is not required for observational prediction. Without one, a model trained on behavior data estimates futures under the behavior mixture represented in that dataset. Such samples may aid some data-generation or perception tasks after validation, but they do not by themselves isolate a proposed braking intervention from correlated observations. Adding recorded actions makes counterfactual control studies possible only with adequate coverage, causal assumptions, and measured action adherence; conditioning alone is not a causal guarantee. We return to action interfaces in the next chapter.
Diffusion and flow matching in latent space
Diffusion and flow matching are two related ways to train continuous-latent generators; neither is simply a universally cleaner reformulation of the other. The flow construction below clarifies the number of network evaluations per solver rollout; realized latency also depends on token count, implementation, and hardware. Prediction and inpainting can share weights when the model is trained and conditioned for those tasks; they do not follow automatically from the ODE.
Let be a data sample (a latent video, all tokens at once) drawn from for context , and let be independent noise of the same shape. Define a straight-line interpolant indexed by . This teaching convention runs from noise at to data at ; the Predict2.5 report writes the opposite time orientation, which is equivalent after reversing and the velocity sign:
Along any particular path, the velocity is constant and equal to . The conditional velocity field is therefore trivial. The marginal velocity field is the conditional expectation
and regressing a network onto with a plain squared error,
has as its population squared-loss minimizer under the stated sampling construction. The network sees individual (noise, data, context) examples rather than the conditional expectation directly. Once trained, samples are produced by numerically integrating the transport equation
from to with an Euler or higher-order solver. With the ideal field and exact integration, this flow transports noise to the conditional data distribution. A trained, finite-step solver approximates it. “Probability-flow ODE” has a more specific meaning in score-based diffusion; this is a flow-matching transport ODE. Fewer steps reduce network evaluations, but latency and quality do not scale by a universal factor.
Why this matters for world models specifically, beyond sample quality:
- Generation can be parallel across frames within one solver step. A -frame clip may take the same number of network evaluations as a one-frame sample, but each evaluation processes more tokens and can cost much more compute and memory. This favors some offline batch workloads; it does not guarantee a latency win.
- Prediction and inpainting can share a masked generator when its architecture and training support that conditioning. Replacing known coordinates during sampling needs a noise-level-consistent procedure; repeatedly inserting clean values is not a universal sampler. Depth, segmentation, and first-frame controls may require separate encoders, branches, or post-training, and adherence remains measured rather than guaranteed.
- For a fixed-shape clip sampler, you choose when laying out that sample's token grid. One common longer-horizon design chains generated chunks conditioned on earlier output; it needs seam and consistency tests. Other architectures can use different extension strategies, so chunk seams are a design-specific risk, not a universal location for failure.
The autoregressive alternative
The discrete-token path has a different mix of tradeoffs, especially sequential token decoding and the opportunity to cache prior context.
Here denotes a discrete token index, not a whole frame; for Cosmos1 it can be an FSQ index. Each token is sampled from a categorical distribution conditioned on available earlier tokens and context. Generation is incremental and may reuse a KV cache, but context length, cache memory, runtime, and error accumulation bound practical horizons. Action tokens can be interleaved only when the model was designed and trained to use them. Part IX develops the foundation-model setting.
The costs are operational. Sampling token indices is sequential in this factorization, so longer sequences normally require more decoding work, although batching, caching, and serving affect wall-clock latency. Exposure bias can arise when generated histories differ from training prefixes, allowing errors to compound. A generator may synthesize plausible high-frequency detail after quantization, but it cannot infer the exact lost detail from the code alone. These are distinct failure mechanisms; quantization does not itself cause autoregressive drift.
As a first approximation, clip-oriented flow/diffusion sampling trades several full-clip network evaluations against autoregressive token-by-token latency and caching. Both families can be adapted to longer sequences or interactive workloads, with different engineering costs. The original Cosmos platform released continuous- and discrete-token video families; later releases should be described by their own interfaces rather than assumed to share that exact division.
Sampling is drawing a future, and the mean is not one
The next worked example isolates a failure of point prediction under a multimodal future.
Suppose you train a deterministic regressor to predict a future position by minimizing squared error. The Bayes-optimal solution is the conditional mean. With a multimodal conditional distribution, that mean can lie in a low-density region between likely outcomes. A point prediction then hides the risk of each mode, even though the mean itself need not have zero probability.
Multimodal futures can arise in interaction: a pedestrian may cross or stop, a vehicle may brake or continue, and a cyclist may swerve or hold a line. Their prevalence is task-dependent. For risk-sensitive decisions, a model needs a way to represent consequential alternatives; a conditional mean alone can conceal them. That does not imply a generative model is automatically calibrated or more accurate.
We use one scalar, one conditioning variable, and a two-component mixture. It illustrates a decision risk without claiming to reproduce the full behavior of a deployed system.
Worked example: the pedestrian whose mean is a ghost
Consider a single decision-relevant scalar: the pedestrian's lateral position one second in the future, measured in metres relative to the ego vehicle's path. The context is , how much the ego yields (0 means the ego does not yield at all, 1 means it yields fully). The true generative process has two outcomes:
- the pedestrian crosses: m,
- the pedestrian stops: m,
with the crossing probability increasing smoothly in , and Gaussian noise of standard deviation 0.25 m on either branch. The bimodality is the whole point: both outcomes are physically plausible, and which one occurs is uncertain given the context. Reducing the process to a single number is exactly the simplification that makes it tempting and wrong.
Let's write that process down and measure how bad the conditional mean is.
import torch
import torch.nn as nn
MODE_STOP, MODE_GO = -1.0, 2.0 # lateral position 1 s ahead, in metres
MODE_SIGMA = 0.25
def mixture_weight(yield_amount, sharpness=3.0):
"""Probability that the pedestrian completes the crossing."""
return 1.0 / (1.0 + np.exp(-sharpness * (yield_amount - 0.5)))
def sample_futures(yield_amount, n, rng):
"""Draw futures from the true bimodal process."""
complete = rng.random(n) < mixture_weight(yield_amount)
centre = np.where(complete, MODE_GO, MODE_STOP)
return centre + rng.normal(0, MODE_SIGMA, n)At an ambiguity of , the two outcomes are equally likely, and the conditional mean sits exactly between them. For a continuous outcome, any exact point has zero probability; the meaningful quantity here is the probability density near the midpoint compared with the two modes. That comparison makes the failure concrete.
conditioning c = 0.5 -> P(pedestrian crosses) = 0.50 conditional mean outcome = +0.50 m density at the mean outcome: 2.43e-08 density at a real outcome : 7.98e-01 ratio : 3.3e+07x The conditional mean lands in a region roughly seven orders of magnitude less dense than either modal centre. The midpoint is possible but very unlikely.
Now we train a generative model of the same conditional distribution using flow matching, so the model learns the shape of the uncertainty rather than a single summary. The training signal is the same data a regressor would have used; only the target changes, from the outcome to the direction that carries noise into outcomes.
class VelocityNet(nn.Module):
"""v(y_t, t, c): a learned flow-matching velocity field for one scalar future."""
def __init__(self, hidden=64):
super().__init__()
self.fc1 = nn.Linear(3, hidden)
self.fc2 = nn.Linear(hidden, hidden)
self.fc3 = nn.Linear(hidden, 1)
self.act = nn.SiLU()
def forward(self, x):
return self.fc3(self.act(self.fc2(self.act(self.fc1(x)))))
def make_batch(n, rng):
"""Sample the flow-matching regression problem: regress (y1 - y0) at the interpolant."""
c = rng.uniform(0.0, 1.0, size=n)
y1 = sample_futures(c, n, rng)
y0 = rng.normal(0.0, 1.0, size=n)
t = rng.uniform(0.0, 1.0, size=n)
y_t = (1 - t) * y0 + t * y1
return c, t, y_t, y1 - y0The training loop is short because the flow-matching objective is short. With probability 0.1 we replace the conditioning value with a null token. That training choice makes classifier-free guidance possible later, at the cost of allocating some examples to the unconditional task. The model must learn that task before we can combine conditional and unconditional velocity estimates at inference.
torch.manual_seed(0)
rng_train = np.random.default_rng(5)
NULL_COND = -1.0
model = VelocityNet()
optimizer = torch.optim.Adam(model.parameters(), lr=2e-3)
for step in range(3000):
c, t, y_t, target = make_batch(256, rng_train)
c_eff = c.copy()
drop = rng_train.random(c.shape[0]) < 0.1
c_eff[drop] = NULL_COND
features = torch.tensor(
np.stack([y_t, t, c_eff], axis=1), dtype=torch.float32
)
labels = torch.tensor(target, dtype=torch.float32)[:, None]
loss = ((model(features) - labels) ** 2).mean()
optimizer.zero_grad()
loss.backward()
optimizer.step()Sampling means integrating the learned velocity field from noise to data. We include the guidance argument now so the same function serves the next section.
def sample_model(net, c, n, steps=100, guidance=1.0, seed=0):
"""Integrate the learned flow-matching velocity from noise to futures."""
generator = torch.Generator().manual_seed(seed)
y = torch.randn(n, 1, generator=generator)
cond = torch.full((n, 1), float(c))
null = torch.full((n, 1), NULL_COND)
dt = 1.0 / steps
with torch.no_grad():
for k in range(steps):
t = torch.full((n, 1), k * dt)
v_cond = net(torch.cat([y, t, cond], dim=1))
if abs(guidance - 1.0) < 1e-9:
v = v_cond
else:
v_null = net(torch.cat([y, t, null], dim=1))
v = v_null + guidance * (v_cond - v_null)
y = y + dt * v
return y.numpy().ravel()We draw two thousand futures at several conditioning values and compare sampled event frequencies with the analytic process. A decision use needs this kind of calibration check for its own event definition and evaluation range; a visually convincing sample is not enough. For the toy monitor below, we call m the conflict-line event. That one-dimensional boundary is chosen to probe the upper crossing mode; it does not represent a vehicle footprint, pedestrian size, or physical collision geometry.
sample_counts = {}
for c_value in (0.1, 0.5, 0.9):
sample_counts[c_value] = sample_model(
model, c_value, 2000, steps=100, seed=int(c_value * 100)
)
risk_threshold = 1.9
density_grid = np.linspace(-2.5, 3.8, 400)
density_mid = mixture_pdf(density_grid, 0.5)
hist_mid, hist_edges = np.histogram(
sample_counts[0.5], bins=density_grid, density=True
)
hist_centers = 0.5 * (hist_edges[:-1] + hist_edges[1:])c=0.1 model P(crosses)=0.28 analytic=0.23 P(lateral > 1.9 m)=0.19 mean=-0.18 m c=0.5 model P(crosses)=0.48 analytic=0.50 P(lateral > 1.9 m)=0.29 mean=+0.42 m c=0.9 model P(crosses)=0.75 analytic=0.77 P(lateral > 1.9 m)=0.43 mean=+1.20 m Compare each sampled crossing-branch frequency with the analytic mixture weight; disagreement is model or sampling error. The distribution can remain bimodal even when those frequencies are not calibrated.
The histogram shows two learned outcome modes and a low-density region near the conditional mean. Its mode weights need not match the analytic mixture merely because both peaks appear.

The decision consequence depends on the planner's defined conflict region and risk threshold. In this toy, a point-estimate rule could proceed at the m mean. At , the analytic process assigns a 50% crossing probability, so a distribution-aware toy rule could brake; the fitted model's estimate must be calibrated before use. Those opposite choices illustrate what is lost by replacing a mixture with its mean. They are not a safety result: the learned sample frequencies differ from the analytic mixture, and a real planner needs validated event definitions, calibration, and operating thresholds.
A world model can have low average prediction error yet hide alternatives a controller must distinguish. Conversely, a model with worse average error might support a decision better if it preserves the relevant alternatives and their calibrated frequencies. Open-loop prediction and closed-loop decision usefulness need separate tests; the planned Part XI material develops that evaluation question.
Conditioning, Guardrails, and Customization
Conditional control, screening, and adaptation make a generative model useful in specific workflows. They are separate capabilities: having one does not establish the others. The Cosmos releases offer examples of each, with different interfaces and evidence requirements.
Conditioning modalities
Broadly, conditioning signals fall into four groups. An omnimodal family may support several input–output combinations across them, but supported combinations and optional inputs depend on the particular model and task. The groups also differ in how strongly they constrain output and how failures should be measured.
- Semantic and stylistic: text prompts can describe scene, weather, time of day, or camera behavior. Language leaves many pixels unspecified, so adherence to a particular object or trajectory needs separate checks.
- Structural and geometric: depth maps, semantic segmentation, edge maps, LiDAR returns, 3D bounding boxes, camera intrinsics and extrinsics, and trajectories from a planner. These can constrain geometry and improve label alignment, but generated pixels still need adherence checks before the original labels are reused.
- Appearance reference: a first frame or an image of a specific object can guide identity and texture. Whether the generated clip preserves either across frames must be tested.
- Control and action: steering, throttle, joint velocities, latent actions inferred from observed changes, or discrete keyboard-style actions can condition a future on a proposed choice when a compatible model was trained for them. A causal counterfactual additionally needs adequate action coverage and assumptions about confounding; supplying an action token alone is insufficient.
Two useful architectural patterns for injecting them are sequence concatenation and parallel control branches. The choice changes what can be constrained and how the model is adapted.
Sequence concatenation. Modality encoders can produce tokens that enter a shared transformer with modality identifiers and attention masks. Training on missing-modality patterns can let one model accept several subsets without changing its network shape; concatenation alone does not guarantee every subset works. Attention also gives a soft influence rather than a hard geometric or action constraint. This is a general pattern, not a claim that every Cosmos 3 modality uses one undifferentiated stream.
Parallel control branches. A side network injects spatial controls into a frozen main network through zero-initialized projections, following the ControlNet pattern. At initialization the side branch contributes zero, so the composite model matches the base model. Once trained, that branch can change outputs, and neither the pretrained prior nor control adherence is protected automatically. Branch weights can be scaled, including over space and time when the implementation supports it.
Cosmos-Transfer1 reports adaptive multimodal control: multiple spatial conditions, including depth and segmentation, can receive different weights at different locations. This can help the generator use stronger signals where they are informative, but 3D consistency and adherence are measured outcomes, not consequences guaranteed by the architecture. For a driving dataset, inspect alignment separately for near-field objects, the horizon, and small labeled regions.
For labeled training data, measure whether generated pixels still match simulator segmentation. A visually better clip with mismatched labels can damage a downstream learner. The options are to reject it, relabel it, or validate that the remaining mismatch is acceptable for the specific task; appearance quality is not a substitute for this check. Part VI discusses data quality in world-model learning.
Guidance and its calibration hazard
Classifier-free guidance is one common way to trade sample diversity for condition adherence. In its usual form, a model trained with conditioning randomly dropped combines conditional and unconditional predictions:
When this formula returns the model's plain conditional velocity. Larger extrapolates away from its unconditional velocity. In many generators that improves condition adherence at the cost of diversity and can introduce artifacts or suppress a mode. The effect depends on the trained fields and condition; no fixed guidance weight has a universal interpretation as a calibrated probability distribution.
Guidance can change event frequencies in the sampled distribution, beyond visible appearance differences. Even if unguided samples were calibrated for a particular event and evaluation set, increasing would require a new calibration check. In the toy experiment below, we call a sample a crossing-branch outcome when its lateral position exceeds m. This branch-frequency diagnostic differs from the later monitor's stricter conflict-line event at m. At , guidance drives the measured crossing-branch frequency below the analytic mixture weight. That direction and size of bias are properties of this toy model, not a general collision-risk correction factor.
guidance_values = [1.0, 2.0, 4.0, 8.0]
c_probe = 0.35
analytic_p_go = float(mixture_weight(c_probe))
guided = {}
for w in guidance_values:
s = sample_model(model, c_probe, 2000, steps=100, guidance=w, seed=7)
guided[w] = {"samples": s, "p_go": float((s > 0.5).mean())}
hist_w1, hist_edges_g = np.histogram(
guided[1.0]["samples"], bins=density_grid, density=True
)
hist_w8, _ = np.histogram(
guided[8.0]["samples"], bins=density_grid, density=True
)
hist_centers_g = 0.5 * (hist_edges_g[:-1] + hist_edges_g[1:])analytic crossing-branch mixture weight | c=0.35: 0.389 guidance w=1 sampled crossing-branch frequency=0.393 absolute difference=0.004 guidance w=2 sampled crossing-branch frequency=0.278 absolute difference=0.111 guidance w=4 sampled crossing-branch frequency=0.132 absolute difference=0.258 guidance w=8 sampled crossing-branch frequency=0.052 absolute difference=0.337 In this toy, guidance suppresses the minority crossing mode. This is useful to demonstrate a tradeoff in generation, but those samples should not be used as calibrated event probabilities without independent validation.


Guidance settings should be recorded with generated samples. If sampled futures inform risk estimates, validate event-frequency calibration on independent data for the exact model, guidance setting, context range, and sampling procedure. Using avoids extrapolative guidance but does not itself make a learned model calibrated. Appearance-tuned and decision-analysis sample streams can be kept separate; neither should be presented as a safety probability without validation.
Guardrails in three layers
"Guardrail" can refer to different checks. NVIDIA's documented Cosmos guardrails cover prompt and generated-content safety. Physical-plausibility screening and runtime risk monitoring are separate system components proposed here; the latter is not a shipped Cosmos safety feature.
Layer one: input and data filtering. Corpus preparation may screen personal or licensed material. Separately, the documented Cosmos pre-guard checks text prompts with a blocklist and content-safety classifier. Those mechanisms enforce a content policy; they do not verify physical dynamics or label fidelity, and their coverage depends on the classifier and release.
Layer two: output screening. The documented Cosmos post-guard screens generated frames for disallowed visual content and blurs faces. A physical-AI data pipeline also needs separate checks for interpenetrating objects, inconsistent scale or shadows, implausible motion, and label mismatches. These are dataset-validity or physical-plausibility problems, not proof that the shipped content filter detects them.
Layer three: runtime risk monitoring, if a world model is placed in a decision loop. An application might estimate event frequency from rollouts and compare it with warning or fallback thresholds. Such a monitor must be calibrated, stress-tested for shift, and backed by independent safeguards before deployment. It is not supplied by the content guardrails above, and a thresholded sample estimate alone is not a certified controller.
The following toy monitor only shows how distribution shift from guidance can change a thresholded decision. Its thresholds are illustrative, not operating limits for a vehicle.
from math import erf, sqrt
def normal_survival(x, mu, sigma):
"""P(N(mu, sigma^2) > x)."""
return 0.5 * (1.0 - erf((x - mu) / (sigma * sqrt(2.0))))
def analytic_conflict(c, threshold=risk_threshold):
"""Ground-truth probability that the pedestrian ends beyond the conflict line."""
p = float(mixture_weight(c))
return p * normal_survival(threshold, MODE_GO, MODE_SIGMA) + (
1 - p
) * normal_survival(threshold, MODE_STOP, MODE_SIGMA)
def safety_monitor(samples, threshold=risk_threshold, warn=0.20, block=0.45):
"""Apply illustrative thresholds to sampled toy futures, not a safety policy."""
p_conflict = float((samples > threshold).mean())
if p_conflict >= block:
action = "block"
elif p_conflict >= warn:
action = "warn"
else:
action = "proceed"
return p_conflict, action
monitor_rows = []
for c_value in (0.1, 0.35, 0.5, 0.9):
s = sample_model(model, c_value, 2000, steps=64, seed=21)
p_conflict, action = safety_monitor(s)
monitor_rows.append(
{
"c": c_value,
"p_conflict": p_conflict,
"analytic": analytic_conflict(c_value),
"action": action,
}
)
guided_probe = safety_monitor(
sample_model(model, c_probe, 2000, steps=64, guidance=8.0, seed=21)
)
unguided_probe = safety_monitor(
sample_model(model, c_probe, 2000, steps=64, guidance=1.0, seed=21)
) condition c P(conflict) analytic monitor
-----------------------------------------------
0.10 0.18 0.15 proceed
0.35 0.23 0.26 warn
0.50 0.28 0.33 warn
0.90 0.42 0.50 warn
same situation, guidance w=1: P(conflict)=0.23 -> warn
same situation, guidance w=8: P(conflict)=0.01 -> proceed
The toy threshold rule depends on the distribution it measures. Sampling with
strong guidance changed its output without changing the situation.
Two further mechanisms matter for a deployed design, although this notebook does not implement them. Independently trained ensembles can reveal some epistemic disagreement; changing a random seed in one fixed model chiefly measures sampling variability, not epistemic uncertainty. Support and shift detection asks whether the current context resembles those used to train and calibrate the model. Outside that domain, calibration evidence may not transfer. Both checks can motivate a fallback, but neither guarantees that a rollout is safe or physically correct.
Customization
Pretrained generators can be specialized in several ways. The following options differ in their training data, changed parameters, and evaluation burden; their cost ordering depends on implementation and scale.
Instruction and control tuning. Fine-tune on paired (condition, video) data to add or improve a camera, weather, time-of-day, or layout control. Whether this is cheaper than an adapter depends on the parameters updated and data required; its adherence must be tested on held-out conditions.
Parameter-efficient adaptation. Low-rank adaptation adds to a frozen weight , where and . For a large weight, is typically much smaller than ; the tiny first layer in the toy below is an exception. With initialized to zero, the combined network equals the base model at step zero. Training changes the adapter, so original-domain behavior must be rechecked. Small adapters can be stored separately for site or camera variants when the serving implementation supports swapping them.
Subject and scene personalization. Reference images and a subject identifier can adapt a generator toward a particular car, robot, or site, in the style of DreamBooth. Identity persistence across frames is a result to evaluate, not a property guaranteed by the method.
Preference and reward-based post-training. Optimize against human preferences or a learned reward for domain-specific outputs. This can encourage reward hacking or physical-distribution drift, so retain an independent fidelity check rather than treating reward improvement as proof of physical validity.
Let's implement LoRA on our toy flow model and adapt it to a shifted one-dimensional geometry: the stopping-position mean moves laterally from to m and the crossing-position mean from to m. These are lateral coordinates, not physical forward/backward distances. This is an adapter mechanics exercise, not an adaptation experiment on a world foundation model or a real site.
SHIFT_STOP, SHIFT_GO, SHIFT_SIGMA = -0.6, 2.6, 0.30
def sample_shifted(yield_amount, n, rng):
"""The new site: same decision structure, different geometry and noise."""
complete = rng.random(n) < mixture_weight(yield_amount)
centre = np.where(complete, SHIFT_GO, SHIFT_STOP)
return centre + rng.normal(0, SHIFT_SIGMA, n)class LoRAVelocity(nn.Module):
"""A frozen base flow model plus rank-r adapters on both hidden layers."""
def __init__(self, base, rank=4, alpha=8.0):
super().__init__()
self.base = base
for p in self.base.parameters():
p.requires_grad_(False)
self.a1 = nn.Linear(3, rank, bias=False)
self.b1 = nn.Linear(rank, self.base.fc1.out_features, bias=False)
self.a2 = nn.Linear(self.base.fc2.in_features, rank, bias=False)
self.b2 = nn.Linear(rank, self.base.fc2.out_features, bias=False)
nn.init.zeros_(self.b1.weight)
nn.init.zeros_(self.b2.weight)
self.scale = alpha / rank
def forward(self, x):
h1 = self.base.act(self.base.fc1(x) + self.scale * self.b1(self.a1(x)))
h2 = self.base.act(
self.base.fc2(h1) + self.scale * self.b2(self.a2(h1))
)
return self.base.fc3(h2)rng_lora = np.random.default_rng(17)
lora_model = LoRAVelocity(model, rank=4, alpha=8.0)
trainable = [p for p in lora_model.parameters() if p.requires_grad]
opt_lora = torch.optim.Adam(trainable, lr=5e-3)
for step in range(1500):
n = 256
c = rng_lora.uniform(0.0, 1.0, size=n)
y1 = sample_shifted(c, n, rng_lora)
y0 = rng_lora.normal(0.0, 1.0, size=n)
t = rng_lora.uniform(0.0, 1.0, size=n)
y_t = (1 - t) * y0 + t * y1
features = torch.tensor(np.stack([y_t, t, c], axis=1), dtype=torch.float32)
labels = torch.tensor(y1 - y0, dtype=torch.float32)[:, None]
loss = ((lora_model(features) - labels) ** 2).mean()
opt_lora.zero_grad()
loss.backward()
opt_lora.step()
n_base_params = sum(p.numel() for p in lora_model.base.parameters())
n_lora_params = sum(p.numel() for p in trainable)frozen base : mean=+0.36 m mass near +2.6 m = 0.19 base + LoRA : mean=+0.95 m mass near +2.6 m = 0.46 trainable parameters: 780 of 5,261 total (14.8%) About 15% of this toy model's parameters were trainable adapters.
This toy compares base and adapted outputs on the shifted domain, but it does not test retention on the original domain. Frozen base weights alone cannot establish retention because active adapters alter old-domain inputs too. The rank-four adapter constrains the update to the second-layer weight; for the first layer it is not low-rank relative to the base matrix and uses 268 adapter weights versus 192 base weights. Choose ranks by layer and held-out performance, and measure both domains before claiming successful specialization. The 14.8% trainable share is specific to this small network, not a full-scale adaptation-cost estimate.
Foundation-Model Infrastructure for Simulation
The model also needs a data and serving pipeline. Tokenizer rate, sampling schedule, hardware, and filtering throughput affect cost and usefulness, alongside architecture and model quality. Published generator weights are tied to their trained tokenizer and input interface; two teams cannot freely substitute arbitrary tokenizers and expect the same model to work. The calculations below isolate a few resource terms rather than ranking their real-world importance.
The cost model
We can write down two dominant matrix-multiplication terms for an illustrative full-attention transformer. This is an order-of-magnitude comparison of tokenization choices, not a latency or hardware-sizing estimate for a released Cosmos model. For tokens, model width , layers, and sampling steps, the simplified forward cost per step is
The attention score/value term is quadratic in token count; the feed-forward term is linear in tokens and quadratic in width. The estimate omits Q/K/V/output projections, embeddings, normalization, tokenizer and decoder work, and communication. Tiled exact attention can avoid materializing the quadratic attention matrix, while its arithmetic remains quadratic. Hardware utilization and memory traffic also shape wall-clock time, so the numbers below are illustrative FLOP estimates, not measured runtime or strict latency lower bounds.
def token_count(frames, height, width, ct, ch, cw):
"""Nominal latent-grid cells; not necessarily transformer tokens."""
return int(
np.ceil(frames / ct) * np.ceil(height / ch) * np.ceil(width / cw)
)
def diffusion_cost(tokens, n_layers, d_model, steps):
"""Two matmul FLOP terms for a hypothetical dense transformer sampler."""
attention = n_layers * 4 * tokens**2 * d_model
feedforward = n_layers * 16 * tokens * d_model**2
return {
"attention": attention,
"feedforward": feedforward,
"total": steps * (attention + feedforward),
}
scenarios = [
("720p, 10 s, 30 fps, 8x8x8", 300, 720, 1280, 8, 8, 8),
("720p, 10 s, 30 fps, 8x16x16", 300, 720, 1280, 8, 16, 16),
("1080p, 10 s, 30 fps, 8x16x16", 300, 1080, 1920, 8, 16, 16),
("480p, 4 s, 10 fps, 4x8x8", 40, 480, 640, 4, 8, 8),
]
scenario_tokens = [(name, token_count(*args)) for name, *args in scenarios] 720p, 10 s, 30 fps, 8x8x8 : 547,200 nominal latent-grid cells
720p, 10 s, 30 fps, 8x16x16 : 136,800 nominal latent-grid cells
1080p, 10 s, 30 fps, 8x16x16 : 310,080 nominal latent-grid cells
480p, 4 s, 10 fps, 4x8x8 : 48,000 nominal latent-grid cells
At 720p and ten seconds, 8x8x8 yields 547,200 latent-grid cells;
8x16x16 yields 136,800, before generator patchification or conditions.Now we sweep the generation horizon at a fixed resolution and compression, and watch the attention term take over.
N_LAYERS, D_MODEL, N_STEPS = 40, 2048, 35
ASSUMED_FLOPS_PER_SECOND = (
4e14 # hypothetical sustained arithmetic throughput; not a device benchmark
)
cost_rows = []
for seconds in np.arange(2, 21, 2):
frames = int(seconds * 30)
tokens_8 = token_count(frames, 720, 1280, 8, 8, 8)
tokens_16 = token_count(frames, 720, 1280, 8, 16, 16)
cost_rows.append(
{
"seconds": float(seconds),
"tokens_8x8x8": tokens_8,
"tokens_8x16x16": tokens_16,
"cost_8x8x8": diffusion_cost(tokens_8, N_LAYERS, D_MODEL, N_STEPS)[
"total"
],
"cost_8x16x16": diffusion_cost(
tokens_16, N_LAYERS, D_MODEL, N_STEPS
)["total"],
}
)8x8x8 grid cells: 115,200 at 2 s -> 1,080,000 at 20 s 8x16x16 grid cells: 28,800 at 2 s -> 270,000 at 20 s token growth over the horizon : 9.4x compute growth over the horizon : 82.7x attention share of total at 20 s : 99% Quartering the assumed sequence length (8x16x16) multiplies the two-term FLOP estimate by 0.06x at the same horizon.
The aggregate number hides which term is responsible. Under this deliberately dense full-attention formula, attention already dominates the shortest horizon plotted and its share rises further as the horizon grows.
flops_split = []
for row in cost_rows:
parts = diffusion_cost(row["tokens_8x8x8"], N_LAYERS, D_MODEL, N_STEPS)
flops_split.append(
{
"seconds": row["seconds"],
"attention": parts["attention"],
"feedforward": parts["feedforward"],
}
)
In this full-attention estimate, compute grows faster than token count because its largest term is quadratic. Quartering tokens cuts that term by sixteen, though it does not cut total serving cost by sixteen when other work remains. Memory-efficient exact attention avoids storing the full matrix; ring attention distributes long sequences across devices, while windowed attention changes the connectivity and FLOP model. A FLOP curve alone cannot establish that memory fails before compute on a particular accelerator.
For scale only, compare with the scalar toy model trained above. These calculations count matrix multiplications, not complete program operations, and the scalar future and video clip are unlike workloads.
import time
start = time.perf_counter()
_ = sample_model(model, 0.5, 256, steps=100, seed=1)
toy_seconds = time.perf_counter() - start
toy_samples_per_second = 256 / max(toy_seconds, 1e-9)
toy_matmul_flops_per_future = 100 * 2 * (3 * 64 + 64 * 64 + 64 * 1)
illustrative_one_second_clip_flops = diffusion_cost(
token_count(30, 720, 1280, 8, 16, 16), N_LAYERS, D_MODEL, N_STEPS
)["total"]
scale_gap = illustrative_one_second_clip_flops / toy_matmul_flops_per_futuretoy flow model : 29,401 scalar futures per second on CPU (870,400 matmul-only FLOPs per future) illustrative full-attention model: 3.73e+15 FLOPs for one standalone 1-second video clip unlike-workload arithmetic ratio: roughly 4e+09x one 20 s clip at 8x16x16: 861.4 PFLOPs = 2153.6 GPU-seconds at an assumed 4e+14 FLOP/s = 5982 GPU-hours for a 10,000-clip dataset ideal compute-only spread over 64 accelerators: 93.5 hours of wall clock
These assumptions are deliberately coarse. The calculated GPU-hours describe this hypothetical full-attention configuration, a chosen effective FLOP rate, and an ideal compute-only schedule; they do not predict the cost or utility of a released model. Dataset value still requires downstream validation. The calculation shows why teams test compression and organize batching, scheduling, and caching rather than selecting a tokenizer on reconstruction quality alone.


Serving, and the two throughput regimes
A deployment may need two quite different operating regimes; it need not use the same model or serve both.
Offline data generation emphasizes validated clips per unit cost. Batching, parallel generation, sampler distillation, reduced-precision inference, and caching can help if supported and measured for the chosen model. Fewer tokens can lower cost, but the resulting loss of small objects or labels must be checked against downstream utility.
Interactive simulation has a per-step latency budget imposed by its controller and sensors. Causal or autoregressive models, caching, local attention, and shorter diffusion schedules are possible responses, each requiring task-specific quality and timing checks. The full-attention curves above illustrate why one dense architecture may be expensive; they do not set a universal lookahead limit for every serving design.
Suppose a synchronous controller runs at 10 Hz and requires camera, depth, and semantic outputs at every step. The complete model-and-sensor pipeline then has at most 100 ms before the next decision. One 80-ms denoising step leaves too little time for a multi-step sampler and other work. This does not imply that Cosmos's long-video extension meets interactive latency: Predict2's documented autoregressive inference chains diffusion-generated clips and reruns denoising for each chunk. It differs from Predict1's discrete autoregressive model. Measure end-to-end rate before calling either an interactive simulator.
Simulation integration and the sim-to-real bridge
The place where all of this becomes concrete is the simulator. Four integration patterns recur.
Real2sim reconstruction. Take a real scene and estimate simulator assets, layout, trajectories, and camera poses to obtain a scene that can be varied. This is hard and largely outside the generator. Reconstruction errors can propagate into later variations, although later fitting and validation may detect or correct some of them.
Conditional photorealization. Render a reconstructed scene with simulator labels such as depth, segmentation, boxes, and optical flow; supply the render and supported controls to a generator. Adaptive multimodal control can improve alignment, but generated appearance does not automatically preserve the labels or reproduce the real-world distribution. Measure adherence, reject or relabel mismatches, and validate whether the retained clips help a downstream perception model.
Neural evaluation environments. Use the world model directly as a policy evaluation environment. This ambitious use needs especially strong checks: model errors can bias the evaluation metric, and an optimizer may exploit weaknesses in a learned simulator. Part VIII discusses this model-exploitation risk. Compare important results with independent evidence from real tests or a separately validated simulator; optimization against the learned environment alone cannot establish deployment performance.
Sensor simulation. Synthesize LiDAR or radar returns as well as camera images. Consistent cross-sensor outputs benefit from geometric structure or aligned geometric cues, but an explicitly geometric shared latent is not the only design: LiDAR diffusion can use a sensor-specific latent, and camera-conditioned radar generation uses a radar representation with geometric guidance. The evaluation question is whether generated modalities agree on scene geometry and sensor-specific effects.
What the infrastructure does not solve
It is worth being explicit about the boundary, because infrastructure discourse tends to overclaim.
More accelerators can raise generation throughput, but that change alone does not establish physical plausibility, long-horizon consistency, control adherence, or calibration. A faster generator can also increase the number of clips that require screening. A production pipeline therefore needs measured acceptance rates and downstream utility, not an assumed fraction of usable outputs. Curation continues after generation because synthetic clips can contain physical or label-consistency failures absent from the simulator source.
Limitations and Impact
The generative objectives discussed here do not by themselves enforce momentum conservation, contact constraints, or object permanence. Larger models and better data may improve physical behavior, but visual plausibility is not proof of a physical law. Rare vehicles, road geometries, or sensor faults may be poorly represented in training, so outputs in those regimes need stronger independent checks. This is a particular challenge for synthetic rare-event coverage: the intended use is often where evidence of model fidelity is thinnest. The planned evaluation material in Part XI will address ways to measure such gaps.
The second limitation is that a conditioning input alone does not impose a hard output constraint. Structural controls such as segmentation and depth can improve adherence in a trained system, but their effect depends on the model and task. For example, a tiny pedestrian in a segmentation map might be misrendered; that is a possible failure to test, not a measured rate in this chapter. Simulator labels paired with mismatched generated pixels can corrupt perception data unless adherence is measured and failures are rejected or relabeled. These checks add cost, though this chapter has not established which pipeline component dominates. In interactive use, a model can accept a brake command and still generate a clip without the expected deceleration; action adherence must also be measured.
Third, calibration needs separate validation. In our toy, strong guidance changes a sampled event frequency; another event estimate might change differently or not at all. Independently trained ensembles may expose some epistemic disagreement, while resampling one fixed model mainly shows sampling variability. Neither detects every systematic bias, especially outside evaluated data support. Large sample counts can make Monte Carlo error small while leaving model bias in a rare-event estimate. A learned world model remains a simulator, and neither its outputs nor those of a hand-built simulator are ground truth. Its errors need deliberate measurement against independent data.
Fourth, data governance and compute matter. Large video corpora can contain faces, license plates, identifiable locations, and material with licensing or consent restrictions. Pretraining a large generator requires substantial resources, making released tokenizers, weights, and adaptation methods important for teams that do not train from scratch. Safety-critical evaluation also needs a separate validation basis: a policy evaluated only against a biased neural simulator may fail where reality differs. Independent road, track, or testbed evidence can break that loop; simulator-based evaluation is not inherently circular.
These limitations do not negate the approach. They define where a released model family provides useful infrastructure and where an application still needs evidence:
- Reusable components across related releases. Text-, image-, and video-conditioned generation and spatially controlled transfer provide building blocks, but supported inputs, weights, and preprocessing differ by release. A team should validate each intended path rather than assume conditioning signals can be freely swapped within one model.
- A candidate sim-to-real bridge for perception. Conditional photorealization can reuse simulator labels only after pixel–label adherence and downstream utility have been verified on the retained clips.
- Discrete and continuous tokenizer options. The original platform released distinct tokenizer variants that expose compression–quality tradeoffs. A downstream policy or evaluation tool may use a compatible representation, but this chapter has not shown that the same exact codes were reused across those systems.
- Parameter-efficient specialization. Adapters and subject personalization can reduce the number of parameters updated, but deployment cost and old-domain retention require separate measurement.
- An infrastructure case study. The Cosmos reports put data curation, tokenization, conditioning, and serving alongside architecture choices. Those system interfaces deserve evaluation rather than being treated as implementation details.
The bridge to the next chapter is direct. Most examples here generate observations, while Cosmos 3 also exposes action outputs. Part IX, Chapter 6 will examine vision-language-action and world-action interfaces explicitly. The planned Part X applications and Part XII reliability chapters will extend the deployment and trust questions; their current outlines should not be mistaken for completed demonstrations.
Summary
The Cosmos releases provide a worked example of physical-world model infrastructure: curation and tokenization in the original platform, later controlled generation, and an omnimodal language–vision–audio–action interface in Cosmos 3. These capabilities belong to different versions and require task-specific evaluation.
Key takeaways:
- Data curation shapes what the model can learn. Physical AI data can differ from typical web video in the prevalence of agentic viewpoints and action-relevant events, and in the availability of aligned sensors and continuous trajectories. The original Cosmos report describes shot splitting, filtering, VLM captioning, and distributed data preparation; corpus counts should be tied to that release.
- Tokenization mediates pixels and computation. The nominal compression triple sets a latent-grid cell count under the stated padding assumptions; additional patchification and conditioning can change generator sequence length. Continuous latents pair naturally with diffusion or flow models; discrete codes support categorical autoregression. Our toy does not establish a winner at equal rate. Causal tokenization avoids future-frame dependence; non-causal encoding may require buffering online.
- A generative model can represent multiple futures. In the toy pedestrian mixture, density at the conditional mean is about times lower than density at a mode. Opposite toy decisions illustrate why event definitions and validated tail probabilities matter; they are not a deployed safety result.
- Flow and diffusion samplers can process a clip jointly at each network evaluation, with compute growing with clip size. Masked prediction requires a trained conditioning path and noise-consistent sampling. Autoregressive tokens are incremental, but context, runtime, and error accumulation limit useful horizons; action control requires training for action inputs.
- Conditioning architectures offer different controls, not guarantees. Shared-token interfaces can accept trained modality combinations; adaptive control branches can weight spatial signals. Zero initialization preserves base behavior at step zero only. Label adherence must be measured before synthetic data reuse.
- Guidance changes sampled futures. In the toy model, increasing suppresses a crossing mode and biases estimated event frequency. Using removes extrapolative guidance but is not proof of calibration; any risk estimate needs independent validation for its precise sampling setup.
- Guardrails must be distinguished. Documented Cosmos components screen prompts and generated visual content; physical-plausibility filters and runtime risk monitors are additional application designs, not shipped Cosmos safeguards. The toy threshold monitor is not a safety case.
- Customization needs retention tests. Zero-initialized LoRA starts from the base function, then learns an active perturbation. Our toy adapter uses 14.8% trainable parameters; its old-domain retention and real deployment cost were not measured.
- Serving budgets differ. Offline generation emphasizes throughput; interactive use has a controller-imposed latency limit. Under this chapter's illustrative full-attention formula, a 10-second 720p clip at 8x16x16 costs about 227 PFLOPs, or about 1,580 ideal GPU-hours for 10,000 such clips at an assumed FLOP/s. The 20-second scenario costs about 861 PFLOPs per clip and about 5,980 ideal GPU-hours for 10,000 clips, before omitted overheads. These are not measured Cosmos costs.
- Throughput alone does not validate correctness. Physical plausibility, horizon consistency, control adherence, and calibration need separate checks. Curation continues after generation, and neural-simulator errors can bias policy evaluation.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Cosmos and omnimodal world foundation models.
Cosmos and Omnimodal World Foundation Models
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!