Autonomous Driving

Michael BrenndoerferJuly 28, 202656 min read

Part of World Models Handbook

Examine driving world models through bird's-eye-view and occupancy state, other-agent prediction, closed-loop simulation, and rare-event safety planning.

Autonomous Driving

A van is double-parked on the near corner of a four-way intersection. From the ego vehicle's perspective, the van blocks the view down the cross street. The ego is thirty meters out, moving at forty kilometers per hour. Somewhere beyond the van, an unseen vehicle may be rolling toward the same intersection, waiting to yield, or absent. In less than three seconds at constant speed, the ego would reach the crossing; its safe response depends on what is hidden and on its available stopping distance.

Occluded intersections are a recognized challenge in autonomous driving. The scene illustrates a central problem: acting under partial observation. The van may itself obstruct travel. Uncertainty about cross traffic adds a planning hazard. A driver might slow until the view improves; an autonomous system has to handle the uncertainty computationally. A world model offers one way to organize the estimate and its consequences. Understanding when that model deserves to be trusted is what this chapter is about.

A driving world model used for planning should support action-relevant estimates of spatial state, hidden hazards, other-agent motion, and uncertainty, whether those quantities are explicit outputs or implicit in its representation. Depending on the system, some of this work may sit in separate perception and prediction modules rather than one model. A convincing rendered video does not establish that the predicted state supports a safe braking decision. The important test is what actions the system takes when its state estimate is wrong or incomplete.

That distinction organizes this chapter. In earlier chapters we developed world models as a general idea, then layered in representations, dynamics, and planning. Here we ask what changes when the world is a road: what state matters, why predicting other agents is especially difficult, why closed-loop simulation differs from open-loop prediction, and how rare events, safety, and planning interact. Each question asks what the model must get right to support a decision. A model trained for generative fidelity and one trained for decisions can share an architecture yet differ substantially in what each represents well.

Autonomous driving is an application where observability, controllability, and safety meet latent dynamics and multimodal prediction under partial observation. Observability concerns what state can be inferred from observation and action history under assumed dynamics; controllability concerns which states the ego can reach through its actions; safety concerns which outcomes are acceptable. A learned model can represent the current scene and possible evolutions, but it does not remove the limits imposed by hidden state and uncertain behavior. Part IV develops the spatial and physical foundations.

We will work through four components: BEV and occupancy state, other-agent behavior prediction, closed-loop driving simulation, and rare-event safety planning. They interact: a prediction can be compromised by a poor spatial estimate, and a planner that relies only on predicted candidate modes cannot evaluate a hazardous mode the predictor omitted. We close with a CPU-sized two-agent intersection experiment comparing a constant-velocity baseline, a single learned trajectory, a label-supervised two-mode output, and an oracle. That toy measures one decision and its simulated outcome; it is not a full closed-loop driving benchmark.

Bird's-Eye-View and Occupancy State

A spatial driving model often needs to encode where nearby hazards are, explicitly or implicitly. Raw sensors provide partial spatial measurements in their own frames: cameras provide perspective images, LiDAR samples 3D points, and automotive radar can estimate range, angle, and Doppler-derived radial velocity. A camera pixel corresponds to a ray through space rather than one 3D point. A LiDAR return has measurement error. A radar detection can be misinterpreted without its geometry and velocity convention. For planning, these observations are commonly aligned in a shared spatial frame and interpreted as evidence about occupancy, motion, and semantics.

We call the example here an ego-centric bird's-eye-view (BEV) state: a top-down, metrically scaled view of the world in a frame attached to the ego vehicle. BEV can also be expressed in a global map frame. A BEV state can contain several channels: drivable space, lane boundaries, dynamic agents with their bounding boxes and velocities, static obstacles, and free-space boundaries. One possible BEV representation is a probabilistic occupancy grid, which discretizes space into cells and assigns an occupancy probability to each cell.

Bird's-eye-view (BEV) representation

A BEV representation is a top-down, metrically scaled encoding of the surrounding scene. This chapter's local example uses an ego-anchored frame, though BEV can also use a global map frame. It can bring observations from multiple sensors into a shared spatial state in which geometry, semantics, and motion can be reasoned about together.

Transforming sensor observations into BEV requires assumptions about calibration and spatial projection. A single monocular image does not directly specify metric depth for every pixel. To place image evidence on the ground plane, the model needs depth estimation, a ground-plane homography where applicable, multi-view geometry, or a learned lifting operation. Each approach has different failure modes:

  • Flat-ground inverse perspective mapping (IPM) assumes a planar road surface and projects image pixels onto it. It is cheap and interpretable, but that projection is valid only for points on the assumed plane.
  • Depth-based lifting estimates per-pixel depth, then unprojects. Depth errors can substantially distort the resulting BEV, alongside calibration and perception errors.
  • Learned lifting can use cross-attention or a transform network to place image features into BEV, trained end-to-end against downstream tasks. Feature assignment is learned, while camera calibration and parts of the geometry may remain explicit.
  • LiDAR-first lifting bins 3D points into voxels or vertical pillars, then aggregates their features in BEV, sometimes with a learned encoder.

The choice of lifting mechanism affects what the model can localize reliably. An image-space detector may identify a pedestrian, but applying flat-ground IPM to body pixels distorts their projected ground location. Transparent and reflective surfaces can challenge depth estimation, while learned lifting can be difficult to diagnose outside its training distribution. A planner consuming BEV features needs uncertainty information as well as an estimated location; a simple transform does not make its output trustworthy by itself.

Once the scene is in BEV, the state can be discretized into an occupancy grid. A cell can carry an estimated probability that it contains an obstacle or object under the grid's chosen occupancy definition. Semantic labels and velocity estimates are optional. Classical probabilistic occupancy grids in mobile robotics are commonly updated by combining prior cell beliefs with sensor evidence. The Bayesian framing clarifies the event an occupancy probability should describe, given the observations and the chosen definition. A neural probability output should be checked for calibration before treating it as a reliable probability of that occupancy event.

In a driving world model, occupancy can serve a different role than in a static mapping system. It can describe an estimated current scene and also be a prediction target for what may occupy cells later. A planning-oriented model may forecast occupancy several seconds ahead, conditioned on proposed ego actions and estimates of other-agent behavior. Predicted occupancy is one useful input for an occupancy-based planner. In a mapping system, occupancy is an uncertain estimate of the present or previously observed scene. A future-occupancy forecast can be checked against later observations; whether it helped the planner is a separate action-outcome question.

This brings us to a distinction that we must preserve throughout the chapter. For analysis, it helps to distinguish three layers of state, even when a system estimates them jointly:

  1. Observed occupancy. The cells that sensors directly support. This is one layer a perception stack may report; it may also infer occupancy beyond current observations.
  2. Inferred state. The cells that the model believes are occupied based on prior structure, occlusion reasoning, and dynamics, but that sensors cannot currently see. The van at the corner hides cells behind it; a good model infers that something may be there.
  3. Predicted future occupancy. The cells that the model expects to be occupied at future times, conditioned on assumed actions of ego and other agents. This is one useful input to an occupancy-based planner.

Confusing these layers can make the planner overconfident. Treating an inferred hidden object as directly observed hides the uncertainty in its existence or position. Treating one predicted future as certain can trigger unnecessary braking or conceal an alternative hazardous future. Training only to reproduce observed occupancy does not by itself establish useful future prediction. Evidence about the present, belief about the hidden present, and predictions about future states have different meanings, and the planner should preserve those distinctions.

Occupancy grids are not the only spatial representation in use. Depending on the task, a driving system may combine an occupancy grid with:

  • Voxel grids, which add height and enable 3D collision reasoning.
  • Vectorized maps, which represent lane and intersection geometry as polylines and graphs, with road rules encoded in associated attributes or semantic annotations. These carry road structure that a pure occupancy grid loses.
  • Instance-level agents, which represent other vehicles as trackable entities with their own state. These are useful for agent-centric interaction-aware prediction, which we cover next; grid-based models can represent interactions differently.

The map and the grid are complementary. The grid estimates where space is occupied; a semantic map describes road structure such as lanes and crosswalks. Perception can also classify occupied cells when semantics matter for prediction or planning, such as distinguishing a parked car from a pedestrian. A grid alone may not distinguish a lane from a shoulder, while a map alone cannot establish whether that lane is blocked now. Systems can combine map and occupancy features early in an encoder or later during planning. Early fusion makes map context available while forming features; late fusion can keep the two streams separately inspectable.

A second distinction worth keeping explicit is between ego frame and world frame. The BEV state can be expressed relative to the ego vehicle, reducing dependence on global translation for nearby planning, or in a global frame anchored to a map, which supports route-level reasoning and consistency across locations. Systems may maintain local state for near-term planning alongside a global map for localization and route-following. Predictions then need a specified transform between frames whenever the planner compares them with global map features.

The resolution of the grid is a design choice with consequences. A one-meter cell can blur a curb, narrow obstacle, or lane marking; a ten-centimeter cell preserves finer geometry but greatly increases the number of cells over the same area. Resolution trades compute against representational detail, and the useful value depends on the task and sensing range. Multi-resolution schemes can allocate finer cells near the ego or around important objects. Grid-based token and convolutional models need a defined spatial discretization, even when they support more than one scale; not every spatial model uses a grid.

Finally, occupancy is not the same as semantics. A cell can be occupied without the model knowing whether the object is a car, bicycle, or cone. Class information can help predict motion, but it does not replace collision geometry; geometry alone does not reveal intent. Some driving models jointly predict occupancy, semantic class, and motion, while others use separate heads or modules. When a planner consumes separate estimates, their meanings and uncertainties should be characterized. When internal features are latent, downstream tests should still probe decisions under plausible state-estimation and prediction errors.

As we discussed in Geometry, Mapping, and 3D/4D State, the choice of spatial representation determines what downstream algorithms can use. An occupancy grid is one compromise: it encodes geometry in a planner-friendly structure, but discretization and occlusion remain difficult. For semantic occupancy, scarce examples of rare classes add a further challenge. When a driving system misjudges a scene or its future, it helps to ask whether the error came from sensing, hidden-state inference, or future prediction.

Other-Agent Behavior Prediction

If the spatial state answers "what is where," behavior prediction asks "what might other road users do?" Other agents are decision-makers with goals and information the ego cannot observe directly. Vehicle dynamics constrain their motion, but do not determine their choices. A planning model therefore needs to represent multiple plausible behaviors, often with probabilities rather than only a kinematic extrapolation. A reachable or robust uncertainty set is another possible representation.

A standard formulation is trajectory prediction: given agent histories and map context, forecast positions over a specified horizon. Recurrent networks, graph neural networks, transformers, and diffusion-based methods have all been studied. Benchmarks such as Argoverse and nuScenes report displacement errors, but their values depend on the dataset, agent class, horizon, and protocol and do not directly measure decision safety.

A mean trajectory can be a poor planning summary when the future has distinct modes. In the toy intersection, the other vehicle either yields before the box or continues through it. Their mean is a slower path that need not resemble either behavior. A planner using only that mean may miss overlap risk; a planner that always protects against every plausible mode may instead brake unnecessarily. Both errors matter. The toy's risk threshold favors avoiding predicted overlap, and we separately count extra braking rather than optimizing an explicit cost function.

For this problem, a useful output is a set of trajectory modes with associated probabilities. The following expression is a discrete approximation: it assigns probability mass to each complete candidate trajectory rather than spreading density continuously over all trajectories.

p(τ∣o1:t,m)=∑k=1Kwk(o1:t,m) δ ⁣(τ−τk(o1:t,m))p(\tau \mid o_{1:t}, m) = \sum_{k=1}^K w_k(o_{1:t},m) \, \delta\!\left(\tau - \tau^k(o_{1:t},m)\right)

where:

  • o1:to_{1:t}: the available observation history up to the current time tt; hidden agents need not have observed tracks
  • mm: the map (lane geometry, rules) that conditions the prediction
  • τk\tau^k: the kk-th mode trajectory, a full future path over the prediction horizon
  • wkw_k: the probability (mixture weight) assigned to mode kk, with wk≥0w_k\geq 0 and ∑k=1Kwk=1\sum_{k=1}^K w_k = 1
  • δ(τ−τk)\delta(\tau - \tau^k): a point mass at the kk-th mode trajectory, so the output is a discrete mixture over trajectories

In words, the predictor assigns weights to KK complete candidate trajectories. Both the weights and candidate paths depend on the available history and map. Sampling this predictive distribution selects one modeled candidate. The actual future need not equal any candidate. The planner can inspect all modeled candidates. Its expectation ∑kwkτk\sum_k w_k\tau^k generally lies outside the finite set of candidates. When the second moment exists, that mean minimizes expected squared prediction error among unrestricted point trajectories. It need not be a dynamically plausible complete path. It is also not necessarily a sufficient summary of overlap risk.

Modes may correspond to interpretable intents such as turning, yielding, or continuing, although learned components do not always align cleanly with those labels. A network can produce candidate paths and logits from a shared representation, with wk=softmax(logits)kw_k=\mathrm{softmax}(\text{logits})_k. The number KK trades candidate coverage against computation: too few candidates may miss relevant behavior, while more candidates raise inference and planning cost and call for evaluation. In fixed-output models, changing KK can also change the learned output head; sampling-based models can draw more candidates after training. If candidate weights are used as probabilities, their calibration should be checked.

Several approaches can produce multimodal predictions.

  • Mixture density networks output mixture parameters directly. They can be trained with a mixture negative log-likelihood, but optimization may require scale constraints and the modes can collapse or overlap.
  • Anchor-based methods start with reference endpoints or trajectories and learn to select or refine them. Anchors encourage coverage of specified behaviors but do not guarantee useful diversity after refinement.
  • Latent-variable methods introduce a discrete or continuous latent zz and decode a trajectory conditional on it. Diffusion-based samplers are another route to diverse predictions, not necessarily the same parameterization as an anchor or fixed latent mixture.

The choice matters for coverage, probability calibration, and planning cost. An anchor can often be refined beyond its initial path, but its placement still affects what the model learns. Mixture and latent-variable approaches can also duplicate or collapse modes; diversity must be checked empirically rather than assumed from the architecture.

A second dimension is interaction awareness. Road users may respond to one another: a vehicle approaching an occupied intersection may slow, and a pedestrian may wait for a gap. Independent forecasts can miss this coupling. One approach encodes agents jointly, using a scene graph and attention so predictions depend on the shared context. Other interaction models are possible, and their computational cost depends on how many agent relationships they represent. In dense scenes, an independent predictor should be checked for mutually inconsistent futures rather than assumed adequate.

A third dimension is action conditioning. When a planner needs to evaluate how other agents would react to alternative ego actions, its response prediction needs to condition on each proposed action as well as the observed past. If the ego accelerates, the other agent may yield; if the ego brakes, the other may continue. A simpler planner can use action-independent forecasts when those responses are outside its model, as the toy below does, but it cannot then claim to predict interactive counterfactual responses. Learning such responses is hard: logged data records only one realized ego action and outcome per situation, while a planner asks what would happen under another action. That is an interventional question, not automatically answered by an observational forecast.

Action-conditioned prediction is closely related to a response model: a function that, given the current scene and a proposed ego action, returns a distribution of other-agent responses. Training on logged scenes with logged ego actions can estimate an associational conditional distribution, P(Y∣X,A)P(Y\mid X,A), where XX is the represented scene, AA the ego action, and YY the other-agent outcome. Using that estimate as an interventional response, P(Y∣X,do⁡(A))P(Y\mid X,\operatorname{do}(A)), requires a well-defined ego action, consistency, adequate action support, and no unmeasured confounding after conditioning on a sufficient scene state. For action sequences, corresponding assumptions apply at each decision step. Hidden factors affecting both the logged action and the other agent's response can break the equivalence even for actions observed in the logs; unsupported actions add extrapolation risk. Restricting plans to well-supported actions can reduce extrapolation risk. It does not by itself identify the causal response. Interactive simulation, interventions where feasible, and sensitivity tests are ways to probe the remaining gap.

Finally, evaluation of behavior prediction is subtle. Low average displacement error does not guarantee that a predictor covers a rare hazardous future. Here, per-agent min-of-K error measures the time-averaged error of the best complete candidate. It does not assess candidate probabilities, joint interaction consistency, or the actions a planner takes. The worked example compares candidate coverage with action outcomes to show why both kinds of measurement are needed.

We'll return to evaluation in Planning, Control, and Policy Evaluation. For now, a trajectory predictor used for planning must represent futures that materially change the decision. The loss and architecture affect whether the predictor represents those futures. Whether candidate weights are calibrated affects how reliably the planner can interpret them as probabilities. The planner's risk rule affects whether represented futures influence its action.

Closed-Loop Driving Simulation

An open-loop prediction benchmark takes a recorded history and scores a forecast against the recorded future. Real driving is closed-loop: ego actions change later states and observations. Open-loop prediction quality is useful evidence. It cannot by itself establish the safety or usefulness of a policy interacting with the environment.

The reason is compounding error and distribution shift. Open-loop evaluation conditions on recorded histories; a closed-loop policy generates its own future observations. Model errors can alter the ego's actions and later states, while other agents may react to those actions. This extends the compounding-error issue discussed in Sampling-Based Planning and Model Predictive Control to a coupled multi-agent system.

Closed-loop simulation lets the ego act and measures outcomes after the simulated world responds. It complements, rather than replaces, open-loop tests and on-road evidence. Three useful evaluation setups differ in what they replay or simulate:

  • Log replay. The ego is re-driven through a logged scenario, with the other agents replayed from the log. Simple and reproducible, but the other agents ignore the ego. An alternative ego path can be checked against fixed replayed tracks, but replay cannot test how logged agents would react to that path or fully interactive counterfactual maneuvers.
  • Reactive-agent simulation. Other agents are controlled by learned or rule-based policies that respond to the ego. CARLA and Waymo's Sim Agents work illustrate different tools and protocols in this broad family. Reaction improves some counterfactual tests, but fidelity and reproducibility depend on the agent models and seeds.
  • Sensor-level simulation. The environment also renders observations for the ego's perception system. It tests a longer stack but introduces its own sensor and visual-domain gaps; more simulated components do not automatically make it more faithful to real roads.

The simulator's modeled components constrain which failure modes an evaluation can expose; scenario selection and measurements matter too. Log replay cannot evaluate how a logged agent would have responded to an alternative ego action. Reactive simulation can test modeled counterfactual responses, but they may differ from that agent's unobserved real response. Simulation can explore many rare or hazardous scenarios without exposing people to those tests, while real-world evaluation remains necessary for deployment claims.

For transfer testing, it helps to expose the policy to an interface close to the one it will use on the vehicle: compatible sensor data, maps, and actuator commands. Matching an interface does not guarantee matching real-world statistics or dynamics. Differences between simulated and real observations or responses are part of the simulation-to-reality gap and should be measured explicitly.

Simulation is also a diagnostic tool. Varying agent density, weather, or sensor noise can expose sensitivity, but a collapse under one perturbation does not identify a single faulty module without further ablation. False-positive occupancy may cause unnecessary braking; missed occupancy may create a collision risk. A useful evaluation records these distinct error types instead of reporting only one pass/fail outcome.

The simulator itself is a world model. This point is easy to miss: in a closed-loop driving simulation, the "environment" that the ego interacts with is another learned or hand-authored model. The quality of that environment model limits what the ego evaluation can establish. Overly predictable simulated agents may reward brittle planner behavior, while idealized physics may make simulated outcomes misleading about real-road dynamics. Neither mismatch guarantees a particular failure in a separately learned model, so simulator fidelity should be treated as an experimental variable rather than a fixed backdrop.

The worked example below is narrower than a reactive closed-loop simulator: it makes one decision from a prediction and rolls out the result. It still shows why an average trajectory error alone can obscure a collision-relevant mode, and why collision counts must be read alongside braking cost.

We'll return to closed-loop evaluation frameworks in The World-Model Evaluation Ladder and Datasets, Benchmarks, and Experimental Design, where we connect driving to broader evaluation methodology.

Rare Events, Safety, and Planning

Lane-keeping, vehicle-following, and intersection negotiation are recurring driving tasks, but rare hazardous situations deserve separate attention. A child entering the road unexpectedly or a cyclist swerving around a parked truck may be infrequent in a dataset while still mattering greatly to safety. Even high average prediction accuracy can conceal failures concentrated in such cases. The evaluation must therefore report more than an aggregate error rate.

This has several implications for how driving world models are built and evaluated.

Rare-event coverage. Limited or unrepresentative training examples can make a model underpredict a hazardous behavior, such as wrong-way driving; absence from the training set does not prove the model cannot generalize to it. Responses include collecting and mining more data, synthesizing counterfactual cases while checking their fidelity, and adding uncertainty-aware monitors or fallbacks. These approaches can be layered, but each needs its own evaluation: more examples do not guarantee coverage, synthetic events may differ from real ones, and an uncertainty signal does not by itself specify a safe maneuver.

Calibrated uncertainty. A planner needs to know what a model's probabilities mean for the event it uses to decide. For a stated class of cases, a calibrated 90% prediction interval should contain the outcome about 90% of the time; a calibrated 10% collision-risk estimate should correspond to collisions in about 10% of comparable cases under the specified action and horizon. Those are different calibration questions. A risk threshold is meaningful only if the collision event, operating conditions, and model probability are defined consistently. Inflated collision-risk estimates for safe cases can cause avoidable interventions under a conservative policy; underestimated risk can conceal danger.

Calibration is hard to establish for multimodal predictions. Temperature scaling can recalibrate suitable logits. Conformal methods can provide marginal coverage under their stated assumptions. Ensembles can estimate disagreement. Shift detectors can flag departures from training conditions. These methods target different quantities, and none alone supplies calibrated collision risk in every deployment condition. As we will see in Robustness, Calibration, and Safe Control, calibration measured on one distribution does not automatically hold after distribution shift.

Calibration measured only on logged ego actions also does not validate collision-risk probabilities for alternative actions. Those counterfactual estimates need the action-support and causal-response assumptions discussed above, or suitable interventional validation.

A planner that consumes a multimodal prediction has to convert the distribution into an action. Several strategies exist:

  • Worst-case. Choose an action that minimizes the maximum modeled cost over represented modes; the worst mode can differ by action. This is conservative relative to those futures, but not a guarantee against omitted hazards.
  • Expected-cost. Weight modes by their probabilities and minimize expected cost. This criterion can still select an action with a low-probability severe outcome when its modeled expected cost is lower; it does not separately bound tail risk. Computational cost depends on the model and optimizer.
  • Chance-constrained. Require the modeled probability of collision to remain below a chosen threshold. Its reliability depends on the probability model and how the collision event is defined.
  • Distributionally robust. Optimize against the worst-case distribution within an ambiguity set. This can hedge distributional misspecification when the set covers plausible models, sometimes producing more conservative actions.

These strategies differ in how they treat rare costly outcomes. Worst-case planning considers the most adverse modeled case; expected-cost planning weights cases by modeled probability and consequence; chance constraints set an upper bound on modeled violation probability; distributionally robust planning considers a set of plausible probability models. None is automatically safe when relevant hazards are missing from the model. The tradeoff between progress and conservatism depends on the operating design domain and safety requirements.

Model exploitation. A planner optimized against an imperfect learned model can favor actions where that model is overly optimistic, as discussed in Failure Modes and Model Exploitation. For example, underestimating red-light violations can make an intersection crossing look safer than it is. The planner may be following its objective correctly while acting on an incorrect risk estimate.

Ensemble disagreement, targeted stress tests, runtime monitors, and independently specified constraints can expose or limit some model errors. A fallback or minimum-risk maneuver may be part of a system design, but a confidence threshold alone cannot certify that the learned model recognizes every unsafe case. The fallback must itself be tested against sensing, dynamics, and road-rule failures.

The Role of Planning. The planner selects ego actions against predicted futures while respecting road rules, physical limits, and uncertainty. Sampling-based MPC, discussed in Sampling-Based Planning and Model Predictive Control, is one approach: propose candidate action sequences, roll them out through a model, score them, and execute a short prefix before replanning. Better predictions can improve this process, but safety also depends on the objective, constraints, fallback behavior, and errors outside the learned model.

The coupling between world model and planner is the central theme of this chapter. In an application like autonomous driving, the world model is not a neutral representation. It is a decision-making instrument. Prediction metrics and action-outcome tests answer different questions; neither alone establishes real-road safety. A useful evaluation connects errors in the model's output to decisions made with that output.

A worked example: a two-agent intersection

We now build a small, CPU-sized experiment that isolates the behavior-uncertainty part of the opening intersection. Unlike the fully occluded opening, the toy gives the ego the other agent's current position and speed but not its intent. The ego travels east along the xx-axis toward a shared box centered at the origin. The other agent travels north along the yy-axis. At t=0t=0, a predictor proposes futures and a threshold rule decides whether the ego brakes. This is a one-decision action-outcome test, not a reactive multi-agent simulator or repeated closed-loop replanning.

We make the scenario minimal so that every step is auditable. The other agent has two possible policies. In continue mode it keeps a constant velocity. In yield mode it brakes to a stop before the intersection. Both policies share the same current-state input. This toy does not simulate an observation history. A classical constant-velocity baseline extrapolates that input without learning. The single-trajectory predictor approaches the mean of the two futures. The label-supervised predictor produces one trajectory per mode. We compare prediction metrics and, after the rollouts, report sampled toy-footprint overlaps and totals of the initial braking decisions.

For the overlap test, the ego is a 3 m long by 2 m wide axis-aligned rectangle traveling east. The other vehicle is a 3 m long by 2 m wide rectangle traveling north. With centers at (xego,0)(x_{\text{ego}},0) and (0,yother)(0,y_{\text{other}}), these rectangles have positive-area overlap exactly when ∣xego∣<2.5|x_{\text{ego}}|<2.5 m and ∣yother∣<2.5|y_{\text{other}}|<2.5 m. Boundary-only touching is excluded. The shaded 5 m by 5 m box in the figures is this center-overlap region, not a claim about real intersection or vehicle dimensions. The rectangles are a toy geometry. The experiment does not model turning. It does not model shape uncertainty. It does not model contact dynamics. During rollout, overlap is checked at each 0.1-second sampled state, not continuously between steps.

Setting up the simulation

We begin with imports and physical constants. The chapter uses fixed seeds and runs on CPU; runtime depends on the machine.

In[3]:
Code
import matplotlib.pyplot as plt
import numpy as np
import torch
import torch.nn as nn

## Simulation constants
dt = 0.1  # timestep in seconds
T_pred = 30  # prediction horizon (3.0 s)
EGO_LENGTH = 3.0  # along x, meters
EGO_WIDTH = 2.0  # along y, meters
OTHER_LENGTH = 3.0  # along y, meters
OTHER_WIDTH = 2.0  # along x, meters
INTERSECTION_HALF = 2.5  # half-size of the center-overlap box in meters
assert INTERSECTION_HALF == (EGO_LENGTH + OTHER_WIDTH) / 2
assert INTERSECTION_HALF == (EGO_WIDTH + OTHER_LENGTH) / 2
EGO_START_X = -25.0  # ego initial position along x-axis
EGO_SPEED = 12.0  # ego initial speed in m/s
BRAKE_DECEL = 5.0  # other agent's yield deceleration in m/s^2

The ego is 25 meters west of the crossing center, moving east at 12 m/s. Its center will reach the near edge of the overlap box at t=(25−2.5)/12≈1.88t = (25 - 2.5)/12 \approx 1.88 s and the far edge at t≈2.29t \approx 2.29 s if it does not brake. The ego's unbraked crossing times are fixed; the other agent's arrival timing varies with its sampled initial position and speed.

Next we define the other agent's two behavior modes and a function to simulate a single rollout of either mode.

In[4]:
Code
def simulate_other(y0, v0, mode, t_total=T_pred, dt=dt):
    """Simulate the other agent along the y-axis.
    mode = 1: continue at constant velocity.
    mode = 0: yield by braking until stopped.
    Returns (positions, velocities) of length t_total + 1.
    """
    y = np.zeros(t_total + 1)
    v = np.zeros(t_total + 1)
    y[0], v[0] = y0, v0
    for t in range(t_total):
        if mode == 0:
            v[t + 1] = max(0.0, v[t] - BRAKE_DECEL * dt)
        else:
            v[t + 1] = v[t]
        y[t + 1] = y[t] + v[t] * dt
    return y, v

We sample initial states uniformly: y0∈[−26,−22]y_0 \in [-26, -22] m south of the intersection and v0∈[9,11]v_0 \in [9, 11] m/s northward. The mode is drawn uniformly from {0,1}\{0, 1\}. With the code's 0.1-second explicit position update, a yield-mode agent travels at most 12.65 m before stopping, so even one starting at y0=−22y_0=-22 m stops at y=−9.35y=-9.35 m, south of the box edge at −2.5-2.5 m. The continuous-time stopping-distance formula gives 12.1 m at 11 m/s; it is not the upper bound for this discrete rollout. Some continue-mode trajectories reach the box during the ego's crossing window, while others do not. The predictor receives only the current state (y0,v0)(y_0, v_0). That input does not reveal the sampled mode. The target contains 31 positions from the present at t=0t=0 through t=3.0t=3.0 s.

In[5]:
Code
def make_dataset(n, rng):
    X = np.zeros((n, 2))
    Y = np.zeros((n, T_pred + 1))
    modes = np.zeros(n, dtype=int)
    for i in range(n):
        mode = int(rng.integers(0, 2))
        y0 = rng.uniform(-26, -22)
        v0 = rng.uniform(9, 11)
        y, _ = simulate_other(y0, v0, mode)
        X[i] = [y0, v0]
        Y[i] = y
        modes[i] = mode
    return X, Y, modes


rng = np.random.default_rng(0)
X_train, Y_train, modes_train = make_dataset(4000, rng)
X_test, Y_test, modes_test = make_dataset(1000, rng)

## Prepare fixed trajectories outside the theme-rerunnable plotting cell.
bev_rng = np.random.default_rng(7)
bev_examples = {0: [], 1: []}
for mode in (0, 1):
    for sample_idx in range(6):
        y0 = bev_rng.uniform(-26, -22)
        v0 = bev_rng.uniform(9, 11)
        y, velocity = simulate_other(y0, v0, mode)
        bev_examples[mode].append(y)
Out[6]:
Console
Training set: (4000, 2), test set: (1000, 2)
Fraction of 'continue' mode in training: 0.50

The training set is approximately balanced between the two modes. Each sample contains a 2D input. Its target is a 31-dimensional trajectory.

Visualizing the scenario

Before training anything, we see what the situation looks like from above. We plot the center-overlap box, the ego's nominal unbraked centerline, and a few other-agent trajectories in both modes.

Out[7]:
Visualization
<matplotlib.legend.Legend at 0x112897390>
Top-down view of an intersection with an ego path from the west and separate continue and yield trajectories moving north.
Bird's-eye view of the two-agent crossing. The shaded box marks center positions at which the toy rectangular footprints can overlap. The ego approaches from the west; continue-mode paths move north through the box while yield-mode paths stop south of it. Other-agent paths have small horizontal display offsets so both clusters remain visible; all are simulated on the same centerline.

The plot shows the two clusters of trajectories clearly. Continue-mode agents move north toward or through the box; yield-mode agents stop south of it. The ego's center crosses the box at x∈[−2.5,2.5]x \in [-2.5, 2.5]. Under the stated rectangular toy footprints, simultaneous occupancy of the box's interior by the two centers is exactly the positive-area overlap condition.

Empirical center-position grids

The other agent's center position over time traces a curve in the (t,y)(t, y) plane. We discretize the lateral coordinate and estimate the probability that its center lies in each yy cell at each time step. Each time row sums to one: this is a per-time center-position distribution, not a distribution over all space-time cells. Physical footprint occupancy is different: one vehicle can cover several cells at once, so those cell-occupancy probabilities need not sum to one. This simplified display is small enough to inspect by eye.

We compute smoothed empirical center-position grids from ground-truth held-out trajectories grouped by their true modes, then combine the grids with equal weights. These dataset-averaged displays illustrate a possible output format; neither trained predictor produced them for the current scene. The Gaussian kernel only makes the sample positions visible as cell probabilities and is not a calibrated uncertainty model.

Formally, the mode-marginal position probability at time tt and lateral cell yy is

Pmix(y∣t)=∑k∈{0,1}πk Pk(y∣t)P_{\text{mix}}(y \mid t) = \sum_{k \in \{0,1\}} \pi_k \, P_k(y \mid t)

where:

  • Pk(y∣t)P_k(y \mid t): the smoothed empirical probability mass for the agent's center in lateral cell yy at time tt among held-out examples with true mode kk; its cells sum to one separately at each tt
  • πk\pi_k: the prior probability of mode kk, here π0=π1=0.5\pi_0 = \pi_1 = 0.5 since modes are drawn uniformly
  • ∑kπk Pk(y∣t)\sum_{k} \pi_k \, P_k(y \mid t): the per-time position marginal under the two-mode mixture, averaged here across the test scenes for display

This uses the same equal weights as the data-generating mode prior, but the time-indexed marginals discard temporal dependence. The grid shows where empirical center-position mass lies at each time; it cannot by itself give the probability of an overlap at any time across an interval. The later planner checks complete scene-specific predicted trajectories and their weights instead of consuming this dataset-averaged display grid.

In[8]:
Code
## Discretize the (t, y) plane
H, W = T_pred + 1, 84
t_range = (0.0, T_pred * dt)
y_range = (-30.0, 12.0)

y_centers = y_range[0] + (np.arange(W) + 0.5) * (y_range[1] - y_range[0]) / W


def build_occupancy(Y_traj, modes, mode_val, sigma=0.8):
    """Estimate per-time-step cell probabilities for one behavior mode."""
    trajectories = Y_traj[modes == mode_val]
    grid = np.zeros((H, W))
    for traj in trajectories:
        kernel = np.exp(
            -0.5 * ((y_centers[None, :] - traj[:, None]) / sigma) ** 2
        )
        kernel /= kernel.sum(axis=1, keepdims=True)
        grid += kernel
    return grid / len(trajectories)


occ_continue = build_occupancy(Y_test, modes_test, 1)
occ_yield = build_occupancy(Y_test, modes_test, 0)
occ_mix = 0.5 * occ_continue + 0.5 * occ_yield
np.testing.assert_allclose(occ_mix.sum(axis=1), np.ones(H), atol=1e-12)
assert np.max(Y_test) < y_range[1]
assert np.max(Y_test[modes_test == 0]) < -INTERSECTION_HALF
Out[9]:
Visualization
Text(0.0, 1.0, 'Empirical Center Position: Yield')
Smoothed empirical per-time center-position probabilities for held-out continue-mode trajectories. Dashed vertical lines mark the overlap-box bounds at y=±2.5 m; this is not a scene-conditioned model prediction.
Heatmap of empirical center-position probability by time for held-out continue-mode trajectories moving north.
Smoothed empirical per-time center-position probabilities for held-out yield-mode trajectories. Dashed vertical lines mark the overlap-box bounds at y=±2.5 m; this is not a scene-conditioned model prediction.
Heatmap of empirical center-position probability by time for held-out yield-mode trajectories stopping south.

The two empirical grids look different. Continue-mode center mass moves north toward and sometimes through the intersection; yield-mode center mass settles south of it. Averaging trajectories into one path would lose this distinction. Mixing mode-conditioned center-position probabilities retains mass in both regions, although it no longer identifies which behavior mode produced a given cell probability. This illustration is not itself the prediction used by the planner.

The non-oracle planner does not observe which mode is realized. The oracle baseline below receives the true mode only for comparison. For a fixed time, summing the center-position probabilities over the box gives the probability that its center is there under this empirical toy distribution. Summing those probabilities across times would double-count trajectories; an any-time footprint-overlap event needs temporal dependence and the ego's motion, which this grid discards. The later action rule therefore uses mode-specific trajectory predictions.

The time-indexed position marginal is shown below, paired with the mode-specific grids above.

Out[10]:
Visualization
Heatmap of empirical center-position probability over time and lateral position for the two-mode sample mixture.
Smoothed empirical center-position marginal at each time, formed by equally weighting held-out continue and yield groups. Dashed vertical lines mark y=±2.5 m. The grid omits trajectory identity and temporal correlation and is not a scene-conditioned predictor output.

A single-trajectory predictor

Now we train a small MLP that outputs a single trajectory. The training runs in the cell immediately below. It minimizes mean squared error, whose unrestricted population optimum is the conditional mean of the future given the input. A finite network trained for a finite number of steps only approximates that target. The architecture is deliberately simple, since the point of the experiment is the loss function and the output structure, not the network's capacity.

Out[12]:
Console
Single-trajectory training MSE: 21.9506

The training MSE is about 22 m222\,\mathrm{m}^2 per trajectory position for this seed. Under squared error, the population-optimal point prediction is the conditional mean E[τ∣o1:t,m]=∑kwkτk\mathbb{E}[\tau \mid o_{1:t},m] = \sum_k w_k\tau^k. A finite-capacity network trained for a finite number of steps need not equal that mean exactly, but here it approaches a path between the two labeled behaviors. That path is a poor summary for a threshold collision check even when its average squared error is reasonable.

A multimodal predictor

The multimodal predictor has a shared trunk and two heads. One head outputs mode logits, which softmax converts to probabilities; the other outputs one trajectory per mode. Because this synthetic dataset gives us the true behavior label for every training trajectory, we use that label to train the corresponding trajectory head and a cross-entropy objective for the mode probabilities. This makes the two output modes identifiable in this controlled example; it does not establish that an unlabeled mixture model would learn them as reliably.

Out[14]:
Console
Assigned-mode trajectory MSE: 0.9175
Mode cross-entropy:            0.6941

The training objective is

L(θ)=1N∑n=1N[1Tpred+1∑t=0Tpred(τt(n)−τtzn,(n))2(1 m)2−log⁡wzn(n)]\mathcal{L}(\theta) = \frac{1}{N} \sum_{n=1}^{N}\left[\frac{1}{T_{\rm pred}+1}\sum_{t=0}^{T_{\rm pred}}\frac{\bigl(\tau_t^{(n)}-\tau_t^{z_n,(n)}\bigr)^2}{(1\,\mathrm{m})^2} - \log w_{z_n}^{(n)}\right]

where:

  • NN: number of training samples
  • TpredT_{\rm pred}: number of prediction steps after the stored position at t=0t=0; each trajectory has Tpred+1T_{\rm pred}+1 positions
  • znz_n: the observed continue-or-yield label for sample nn
  • τ(n)\tau^{(n)}: the observed future for sample nn
  • τzn,(n)\tau^{z_n,(n)}: the predicted trajectory assigned to the labeled mode
  • wzn(n)w_{z_n}^{(n)}: the predicted probability of the labeled mode

The first term fits the trajectory head associated with the known label; the second trains the predicted mode probability. Dividing by a one-meter squared scale makes the position term dimensionless and leaves the numerical loss computed from meter-valued coordinates unchanged. It also fixes the relative weighting of position error and mode cross-entropy when interpreting the equation in physical units. Since the current input (y0,v0)(y_0,v_0) does not reveal intent, a calibrated classifier should assign approximately equal probabilities to the two modes. This labeled construction lets us examine the decisions of a two-mode predictor in this controlled example; learning modes without labels is a harder problem.

Comparing predictions

Now we evaluate the two learned predictors and a constant-velocity baseline on the test set, in a computation cell that runs once and prepares the metric values that figure cells consume. We compute seven metrics:

  • Mean squared position error (MSE) of constant-velocity extrapolation overall and separately for the true continue and yield subsets.
  • Mean squared position error (MSE) of the single predictor.
  • Mean squared position error of the mixture mean, i.e., the probability-weighted average of the two modes.
  • Min-of-K squared error for the two-mode predictor: the error of the closest candidate trajectory.
  • Mode coverage: the fraction of test trajectories for which at least one mode is within a tolerance.

These metrics compare point-prediction error with candidate-set accuracy. The constant-velocity errors expose how much the extrapolation depends on the true behavior. The single and mixture-mean errors measure average squared error; the last two metrics describe best-candidate time-averaged error and coverage at a stated threshold. None directly measures action safety or unnecessary braking, and the candidate-set metrics ignore mode probabilities.

Out[16]:
Console
Constant velocity MSE (all):            46.085
Constant velocity MSE (continue):       0.000
Constant velocity MSE (yield):          87.948
Single predictor MSE (mean trajectory): 21.951
Mixture mean MSE (expected trajectory): 22.081
Mixture min-of-K MSE (best mode):       0.887
Mixture mode coverage (within 1 m^2):   0.691

The min-of-K error is defined as

min-of-K=1N∑n=1Nmin⁡k∈{1,…,K}1Tpred+1∑t=0Tpred(τt(n)−τtk,(n))2\mathrm{min\text{-}of\text{-}K} = \frac{1}{N} \sum_{n=1}^{N} \min_{k \in \{1, \dots, K\}} \frac{1}{T_{\rm pred}+1} \sum_{t=0}^{T_{\rm pred}} \left( \tau_t^{(n)} - \tau_t^{k,(n)} \right)^2

where:

  • NN: the number of test samples
  • TpredT_{\rm pred}: the number of prediction steps after t=0t=0, giving Tpred+1T_{\rm pred}+1 stored positions
  • τt(n)\tau_t^{(n)}: the true position at time step tt for test sample nn
  • τtk,(n)\tau_t^{k,(n)}: the predicted position at time step tt from mode kk for test sample nn
  • min⁡k\min_{k}: the operation that selects the single mode closest to the truth for each sample

In words, the metric first computes the time-averaged squared error of each complete candidate trajectory, selects the best candidate per sample, then averages over samples. It measures best-candidate closeness, not pointwise coverage: the separate coverage percentage applies a 1 m21\,\mathrm{m}^2 threshold to that time-averaged error and does not require every predicted position to be within one meter.

The numbers tell a more qualified story. Constant velocity reproduces the continue path in this construction but extrapolates through the intersection when the other agent yields; its mode-specific errors show the failure hidden by one pooled score. The single predictor and mixture mean have similar MSE, as expected when both approximate the conditional mean. The mixture's min-of-K MSE is much lower, but its coverage at the stated 1 m21\,\mathrm{m}^2 tolerance is below one: it has not learned every trajectory to that precision. A planner can check each predicted mode for risk, provided the modes include the dangerous futures and their probabilities are meaningful.

Let's visualize the predictions for a specific test scenario where the true mode is "continue" and the input is ambiguous.

Out[17]:
Visualization
<matplotlib.legend.Legend at 0x1179eca50>
Line plot comparing true future trajectory with single-predictor mean and two mixture modes.
Held-out continue-mode prediction comparing a single-trajectory predictor against a two-mode, label-supervised predictor. The single predictor regresses toward the conditional mean; the two-mode predictor shows separate continue and yield paths with approximately equal weights. The later action-outcome test checks whether this representation changes the ego's decision.

The plot shows the tension. The true future moves north toward the intersection. The single predictor, forced to choose one trajectory, produces a slower mean path. The two-mode predictor shows a continue path near the truth and a separate yield path.

One-decision action-outcome evaluation

Prediction metrics are useful diagnostics but are not sufficient to evaluate action safety. What matters for safety is whether the predictor enables a useful ego decision. In this one-decision scenario, the ego observes (y0,v0)(y_0, v_0), asks the predictor for a future, checks which predicted modes overlap its planned crossing window, sums their weights, and brakes if that risk score exceeds 0.1. The choice is Boolean: if it brakes, the ego follows a fixed 5 m/s25\,\mathrm{m/s^2} deceleration profile. We then roll out both agents for up to six seconds, stopping at the first detected sampled footprint overlap. There is no repeated observation or replanning, so this is not a reactive closed-loop driving benchmark.

We now define four predictors to compare: an oracle that knows the true mode, constant-velocity extrapolation, the single-trajectory MLP, and the mixture MLP.

We now run many scenarios. For each, we draw (y0,v0)(y_0, v_0) and a true mode, and we run each predictor.

Out[21]:
Console
Scenarios: 300  (continue: 160, yield: 140)

            Oracle:  toy-footprint overlaps =   0  (continue-mode:   0, yield-mode:   0)  brakes = 105  brakes when oracle proceeds =   0
 Constant velocity:  toy-footprint overlaps =   0  (continue-mode:   0, yield-mode:   0)  brakes = 185  brakes when oracle proceeds =  80
            Single:  toy-footprint overlaps = 105  (continue-mode: 105, yield-mode:   0)  brakes =   0  brakes when oracle proceeds =   0
           Mixture:  toy-footprint overlaps =   0  (continue-mode:   0, yield-mode:   0)  brakes = 274  brakes when oracle proceeds = 169

In this seeded experiment, the oracle has zero toy-footprint overlaps and brakes in the 105 scenarios where the true continue path threatens the box. Constant-velocity extrapolation also has zero overlaps, but brakes in 185 scenarios, including 80 where the oracle proceeds: it assumes even a yielding agent continues at its initial speed. The two-mode predictor has zero overlaps but brakes in 274 of 300 scenarios, including 169 where the oracle proceeds. Thus the classical baseline is less intervention-heavy than this learned mixture under the same toy risk rule. The single predictor never brakes and has 105 overlaps in continue-mode scenarios. The risk check adds weight ∑kwk⋅1 ⁣[∃j: tj=jΔt∈[tin,tout], ∣τtjk∣<INTERSECTION_HALF]\sum_k w_k \cdot \mathbb{1}\!\left[\exists j:\, t_j=j\Delta t\in[t_{\text{in}},t_{\text{out}}],\ |\tau_{t_j}^{k}|<\texttt{INTERSECTION\_HALF}\right] over the sampled crossing window and brakes when this exceeds 0.1. A separate rule that always brakes would brake in all 300 scenarios and have no sampled overlaps here; that conservative toy rule is not a road-driving policy. These results illustrate a safety-efficiency tradeoff under the toy geometry, not a deployed performance claim.

Before inspecting a single scenario, it helps to read the aggregate. Because footprint overlap occurs only in the continue mode, the rates should be broken out by the true mode rather than pooled. The evaluation set contains the 300300 scenarios drawn above with seed 4242; each predictor receives one Boolean overlap outcome per scenario.

In[22]:
Code
predictor_names = ["Oracle", "Constant velocity", "Mixture", "Single"]
continue_collision_rate = [
    results[name]["collisions_continue"] / max(results[name]["n_continue"], 1)
    for name in predictor_names
]
yield_collision_rate = [
    results[name]["collisions_yield"] / max(results[name]["n_yield"], 1)
    for name in predictor_names
]
Out[23]:
Visualization
Grouped bar chart of toy-footprint overlap rates for oracle, constant-velocity, mixture, and single predictors.
Toy-footprint overlap rate by true behavior mode in the seeded one-decision test. Oracle, constant-velocity, and two-mode predictors have no overlaps; the mean-trajectory predictor overlaps in continue-mode cases. Overlap rate alone hides conservatism: constant velocity brakes in 185 of 300 scenarios and the two-mode predictor in 274.
Out[24]:
Visualization
<matplotlib.legend.Legend at 0x11afcfd10>
One-decision outcome for a held-out continue-mode scenario. Under the single predictor, the ego proceeds and the two shaded toy vehicle footprints overlap. The rectangles mark the simulated vehicles' positions at the first detected sampled overlap, not a separately invented collision point.
Top-down path plot with the ego and other vehicle's overlapping rectangular footprints at the first detected sampled overlap.
The same initial conditions and braking rule under the two-mode predictor cause the ego to stop before the overlap region while the other agent continues north.
Top-down path plot with the ego stopped west of the crossing and the other vehicle moving north without footprint overlap.

In the single-predictor panel, the ego crosses while the other agent is in the shared box. In the two-mode panel, the ego brakes and the other agent passes. The initial conditions and decision rule are the same. Only the predictions differ. This selected case does not show the cost of the additional braking in other cases.

What this experiment shows, and what it doesn't

The experiment is deliberately small, but it isolates a useful failure mode: an MSE-trained single trajectory can conceal a hazardous mode. Representing separate modes lets the simple planner avoid toy-footprint overlaps in this dataset, at the cost of many extra brakes compared with the oracle. Yet the constant-velocity baseline also avoids overlaps and brakes less often than the learned mixture in this particular toy. Richer prediction alone does not guarantee a better decision. The experiment therefore motivates measuring both modeled safety and progress or intervention cost, rather than prediction error alone.

Three caveats are important. The example uses two labeled behaviors even though real actions vary continuously. The other agent follows a fixed policy and never reacts to the ego; its predictions are action-independent. The ego makes one threshold decision at t=0t=0 rather than repeatedly observing and replanning. The example is therefore an action-outcome check, not an action-conditioned response model or a realistic closed-loop driving simulation.

Limitations and impact

The chapter's worked example is small enough to fit on a laptop, and that is precisely why it is useful. But the gap between this toy and a deployed driving world model is enormous, and it is worth stating clearly what that gap consists of.

The first limitation is scene complexity. Real scenes can contain many interacting agents, and factorized predictions or restricted interaction neighborhoods may omit important dependencies. A two-agent centerline example avoids occlusion, lane choices, traffic rules, realistic and varying vehicle dimensions, orientation, contact geometry, and perception error; it uses only fixed rectangular toy footprints. Success here says little about performance in a dense scene.

The second limitation is that real driving has continuous action and state spaces, and behavior need not split into two clean labels. A vehicle may accelerate slightly, drift within its lane, or yield while still moving. We supplied the two training labels in this toy problem; the model did not discover them unsupervised. The finite-mode equation is a simplified representation of trajectory uncertainty, and continuous-density or sample-based predictors are alternatives.

The third limitation is the simulation-to-reality gap. Simulated agent responses and sensor outputs may miss real behavior and environmental variation. Simulation outcomes are useful evidence but cannot by themselves establish real-road safety; deployment claims require broader validation under the applicable safety process.

The fourth limitation is the coupling between perception, prediction, and planning. The toy supplies the other agent's position and speed directly. On a vehicle, errors in object detection, depth, tracking, and maps can affect both predicted futures and the decision. Component-level ablations and end-to-end tests can complement each other when tracing how errors affect final decisions.

The value of these research directions is a richer way to represent and test driving futures. BEV encoders, multi-agent predictors, and reactive simulators let researchers study decisions under spatial uncertainty and interaction. Their practical benefits depend on the task, benchmark, and deployment setting; better open-loop prediction or more realistic-looking simulation does not automatically yield fewer real-world incidents.

The conceptual lesson extends beyond roads: when a learned model informs actions, evaluation must ask whether it preserves decision-relevant alternatives, provides uncertainty estimates that can be tested, and remains useful after actions change the state. The next chapter will apply that lens to digital environments.

Summary

Autonomous driving puts world models under physical-road constraints that shape the choices a system must make.

In many systems, sensor features are fused into a bird's-eye-view or occupancy representation of the surrounding scene. Observed, inferred, and predicted occupancy should be distinguished when interpreting model outputs, even if a system estimates them jointly; confusing their meanings can make a planner overconfident.

Other-agent behavior may be multimodal. A single average trajectory cannot represent a yield-or-continue choice. Candidate-set coverage, probability calibration, and downstream decisions all matter; no one displacement metric is sufficient.

Closed-loop simulation evaluates policy–environment interaction and complements prediction benchmarks, log replay, and real-world evidence. Reactive-agent and sensor-level setups expose different failure modes, and every simulator has a simulation-to-reality gap.

Rare events, safety, and planning tie everything together. Long-tail coverage, uncertainty calibration, planning under uncertainty, and defenses against model exploitation are important design and evaluation considerations. No single list of methods certifies deployment safety. The planner's job is to select actions that respect the map and rules while accounting for predicted alternatives and possible model error.

Our worked example showed how a single-trajectory MSE predictor regresses toward a mean path that can hide collision risk. A two-mode predictor trained with synthetic mode labels avoids overlaps between the toy vehicle footprints under the same threshold decision rule, but brakes more often than both the oracle and a constant-velocity baseline. The result is an action-outcome illustration, not evidence of real-world driving safety or of unsupervised mode discovery.

Selected Notation

Three quantities used in the toy experiment and its display are:

  • K: Number of candidate trajectory modes; this synthetic comparison uses K=2K=2.
  • sigma: Width of the display-only Gaussian smoothing kernel used to turn sampled center positions into the empirical per-time position grid; it is not a calibrated sensor-noise parameter.
  • T_pred: Prediction horizon in timesteps. Longer horizons can make regression and candidate coverage harder when behavior branches or uncertainty grows; horizon length alone does not imply more modes.

Next Chapter

The next chapter, Digital and Software Environments, will ask what happens when the environment is code rather than asphalt. Several questions will recur: what is the state, what is the action space, how do we evaluate actions, and how do we defend against model exploitation? Software transitions may be easier to replay and verify than road interactions, though hidden state and irreversible actions remain.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026autonomousdriving, author = {Michael Brenndoerfer}, title = {Autonomous Driving}, year = {2026}, url = {https://mbrenndoerfer.com/writing/autonomous-driving-world-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Autonomous Driving. Retrieved from https://mbrenndoerfer.com/writing/autonomous-driving-world-models
MLAAcademic
Michael Brenndoerfer. "Autonomous Driving." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/autonomous-driving-world-models>.
CHICAGOAcademic
Michael Brenndoerfer. "Autonomous Driving." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/autonomous-driving-world-models.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Autonomous Driving'. Available at: https://mbrenndoerfer.com/writing/autonomous-driving-world-models (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Autonomous Driving. https://mbrenndoerfer.com/writing/autonomous-driving-world-models

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.