Part of World Models Handbook
Build a miniature planar pushing world model in Python: contact features, Ridge regressors, rollout error, random-shooting plans, and affordance maps.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Robotic Manipulation
Imagine moving a mug across a table with a robot arm. A push might slide it, tip it, or miss it; a grasp introduces different contact and force questions. From a camera image alone, the controller may not know the mug's exact pose, mass distribution, or friction. This is a useful thought experiment, not a measured twelve-second robot trial. The executable example later is narrower: a two-dimensional pusher and puck with a deliberately simple contact rule.
Perception, contact, and control interact: pose error changes the selected action, and that action changes the next observation. A world model can help a controller compare possible consequences, but agreement among imperfect components is not enough for correctness. Predicted motion must be checked against observed outcomes under the intended task and sensing conditions.
Analytic mechanics remains useful for known geometry and calibrated contacts. Deformable objects, uncertain surface properties, and partial observation make those models harder to apply directly; learned predictors can complement them by estimating effects from interaction data. Neither a rigid-body model nor a learned model by itself resolves arbitrary cloth, cables, or granular materials.
Learning from data can help identify task-relevant effects, but it does not prove that color, material, texture, velocity, or history are irrelevant. Which variables may be discarded depends on the object and task. We will distinguish explicit geometric state, learned latent state, and hybrids that retain known structure while estimating uncertain effects.
The chapter asks what a manipulation model needs to predict and how a controller uses that prediction. We compare state choices and the distinct demands of grasping, pushing, and dexterity; then examine visual foresight and affordances as related approaches. The worked experiment tests only planar pushing: two small regressors, one-step and recursive errors, an open-loop random-shooting plan, and a single-action effect field. The broader concepts connect to Part III, Part IV, Part V, and Part VIII.
Contact is one recurring challenge: active constraints can change when bodies meet, slip, or separate. State choice, sensing, uncertainty, and planning horizon matter too. The toy below isolates one contact switch without claiming to model full manipulation.
Object State and Contact-Rich Dynamics
Before predicting the next state, we must decide which quantities to represent. They can be assembled into a vector, graph, image, or latent history, but that choice determines what the transition model can update and what the planner can evaluate. Missing task-relevant variables cannot be recovered from a perfect transition fit to an insufficient state alone.
One explicit geometric representation assigns poses to relevant rigid bodies: the end-effector, gripper fingers, manipulated objects, and obstacles. A rigid-body pose lives in , the special Euclidean group of rigid transformations, and can be written as:
where:
- : the homogeneous transformation matrix combining rotation and translation
- : the rotation matrix that captures orientation
- : the translation vector that captures position
The rotation part has three degrees of freedom and the translation part has three, so a pose has six degrees of freedom in total. The homogeneous form is convenient because it makes composition of transforms a matrix multiplication: applying one transform after another is just multiplying their matrices, and finding the relative transform between two bodies is a single matrix inverse followed by a product. That algebraic convenience is why pose-based state representations are so common in robotics: the geometry of a scene becomes linear algebra over matrices.
Pose alone is not a complete dynamical state. Velocities, robot joint states, shape and material properties, and contact geometry or history can matter even for rigid objects. A deformable object may require a mesh, particles, or another compact representation. Contact modes summarize useful information, but they are not the only missing variables. Two arrangements with different contact geometry generally also differ in their poses relative to object shapes; a contact graph is one representation of that distinction, not the sole source of it.
The table lists common design choices without ranking their accuracy. Any of them can succeed or fail depending on observation quality, training coverage, dynamics, and the downstream decision.
| Representation | What it may contain | Main advantage | Main limitation | Example use |
|---|---|---|---|---|
| Explicit physical state | Poses, velocities, contact modes, and relevant material or shape parameters | Interpretable variables and known constraints | Requires estimation or calibration of relevant quantities | Known objects with useful sensing |
| Learned latent state | Encoded history from images, depth, and proprioception | Can represent task-relevant information without a prescribed object list | Harder to interpret; may omit rare contacts | Partially observed scenes |
| Object-centric state | Per-object variables and interactions | Makes object identity and relations explicit | Segmentation and tracking can fail | Multi-object scenes |
| Hybrid state | Known robot kinematics plus estimated object state or learned residual dynamics | Reuses known structure while modeling uncertainty | Interface errors and missing variables remain possible | Systems with calibrated robot geometry |
Manipulation and locomotion both include contact-mode changes. In a simplified rigid-contact model, dynamics may be smooth within a fixed mode but switch at touch, slip, or separation. A physical impact, a compliant-contact model, and a discrete-time simulator need not have the same mathematical regularity. The hard switch below is a property of our toy simulator, not a theorem that every manipulation transition is discontinuous.
A one-dimensional version of the code's pushing rule makes its hard switch concrete. First move the pusher by the commanded displacement , clipping it to the workspace if necessary; then check its distance to the puck at its new position . The puck update is
where:
- and : puck and pusher positions before the step; is the pusher position after its commanded move and workspace clipping.
- : the commanded displacement, not necessarily the displacement actually achieved after clipping.
- : workspace coordinate bound; : the distance threshold in this toy contact test.
- : an ad hoc coupling gain, not a Coulomb-friction coefficient.
- : one when the post-move pusher–puck distance is below , zero otherwise.
When the post-move test is on, the puck receives in this toy rule; otherwise it stays put. The code applies the full command to the puck even if the pusher is clipped at the workspace edge. That is a simplifying artifact, not a physical contact law. For fixed nonzero , the next puck position jumps as the contact indicator changes. Away from the switch its derivative with respect to is one; at the switch the hard map is not differentiable. It is misleading to describe this as a large finite curvature.
A local linearization within the no-contact region misses the effect of an action that crosses the threshold in the same step. Smoothing the indicator can make a fitted predictor easier to optimize, while an explicit contact mode keeps the switch visible. Neither choice alone guarantees good plans near contact.
Dynamics in which contacts can change the active constraints or force law. Touch, slip, separation, and impact can make prediction difficult; how smooth the resulting transition is depends on the physical model and timescale.
This is why contact-rich dynamics receive separate treatment in Part IV. Depending on the system, they can introduce several difficulties:
- Mode changes. A hard simulator may switch between no contact, sticking, and sliding; compliant models can smooth some transitions.
- Outcome uncertainty. Uncertain state or parameters can yield several plausible outcomes near a mode boundary.
- Long-horizon dependencies. An early grasp can affect a later pour, so an error may propagate into the task result.
- Partially observed parameters. Material and contact properties may need estimation from interaction or additional sensors.
Possible responses include contact features or explicit modes, uncertainty-aware predictions, history-conditioned state estimation, and online parameter updates. Their value is empirical: a one-step error metric alone will not reveal every planning failure, so evaluate both prediction and closed-loop behavior as discussed in Part I.
Choosing the level of state representation
The four options in the table trade off observability, structure, estimation error, and computational cost. No three-question rule automatically selects a representation.
Explicit object state plus contact mode. Poses, velocities, material parameters, and contact labels can support a structured physical model when they are accurately known or estimated. Pose and mode alone are not generally a sufficient state, and explicit models are not inherently more predictive than learned alternatives.
Latent state learned from observations. An encoder can summarize RGB, depth, and proprioceptive history in , then predict a successor conditioned on action. RSSMs and Dreamer illustrate this pattern (Part V, Part VIII). It avoids a prescribed object list but may omit rare contacts; recurrence can help retain useful history without ensuring that every hidden property is recovered.
Object-centric state. Slots or entity vectors can make object identity and interaction structure explicit (Part III, Part IV). They may help maintain identity across steps, but segmentation, tracking, and the learned interactions can fail; no general long-horizon advantage over monolithic latents follows from the representation alone.
Hybrid state. A system may combine calibrated robot kinematics with learned object dynamics or residual errors. Joint encoders do not make end-effector pose exact, and known kinematics do not by themselves confer safety. The model interface still needs validation against sensing, compliance, and contact effects.
Our code below makes a smaller comparison: two linear predictors over explicit positions, one with an additional contact-proximity feature. It does not implement a latent, object-centric, or hybrid robot model.
What the transition needs to predict
The transition target depends on the task. A point predictor can be useful, as in our toy, but a distribution is valuable when the sensed state or dynamics leave multiple plausible contact outcomes. Near a threshold, uncertainty in relative position can put probability mass on both no-contact and contact branches even if each fully specified branch is deterministic. Candidate outputs include contact-mode probabilities or mixtures over pose deltas. This distinction between uncertainty about state and uncertainty in dynamics also appears in Part IV.
Uncertain friction or contact geometry may reflect missing observations or unknown parameters, not irreducible physical randomness. Better sensing or system identification may reduce some of that uncertainty. Other effects can remain stochastic at the model's chosen resolution. A planner should represent whichever uncertainties materially change its action choice rather than assume that every contact is intrinsically multimodal.
Grasping, Pushing, and Dexterous Control
Grasping, pushing, and dexterous manipulation emphasize different contacts and outcomes, but they overlap. A push can maintain contact over an interval, and a grasp can involve slipping or repositioning. The useful state, prediction horizon, and sensors depend on the actual task rather than the label alone.
Grasping
Grasp planning may estimate form or force closure from geometry, rank candidates with a learned grasp-quality predictor, or simulate what follows contact. These are different tools. A predictor of grasp success from depth or RGB-D observations need not be a component of a world model.
Dex-Net 2.0 is one influential grasp-quality example: it trained on 6.7 million synthetic point-cloud, grasp, and analytic-quality examples, rather than tens of thousands of real labeled attempts. Its grasp-quality score answers a narrower question than a rollout model. A task that requires pre-grasp rearrangement or evaluation of an object after pickup may need additional prediction or planning, but not necessarily one monolithic world model.
Depending on the task, a grasping model may need to predict some of the following:
- The contact configuration between gripper and object.
- The force distribution across the contact patches.
- The object's response: does it stay gripped, slip, rotate, or tip?
- The downstream consequences: once the gripper has the object, what task-relevant properties (pose, orientation, stability) does it now have?
An immediate success score does not by itself evaluate later use of the object. A task-aware model could compare candidate grasps by predicted downstream outcomes, provided its dynamics and reward model cover those outcomes. A single model serving different tasks by changing only the objective is a design possibility, not a general guarantee.
Pushing
Pushing moves an object through contact without closing a gripper. Contact can be brief or sustained; pose, contact point, action direction, and friction may all matter. Quasi-static planar pushing is one useful approximation, not the definition of pushing.
In slow pushing, acceleration and inertial terms can sometimes be neglected. Contact and external wrenches then approximately balance at each instant even while the object moves. Motion still depends on contact kinematics and friction; the approximation need not carry inertial velocity as a state.
Under appropriate slow-motion and frictional conditions, a quasi-static approximation can avoid explicit inertial state. It still needs relevant contact geometry and frictional assumptions. Fast motion and post-contact sliding can violate it. Our discrete toy rule below is not a quasi-static force model.
A pushing model can ask: given the pusher's displacement and the current puck pose , what is the new puck pose ? Under suitable quasi-static conditions, the answer can depend on contact geometry and friction as well as the commanded motion. Write the pusher pose as , the pusher–object contact point as , and object geometry as . Holding other friction and contact parameters implicit in , a schematic transition is:
where:
- : the new puck pose after the push, the quantity the model predicts
- : the current puck pose
- : the current pusher pose
- : the commanded pusher displacement
- : the contact point induced by the puck and pusher geometry
- : the object geometry, which determines where contact forces act
- : a schematic push transition; its form and regularity depend on the contact and friction model
Local contact features can focus a predictor on the part of the scene relevant to a push; this is one way to add physical structure (Part IV). Whether that helps data efficiency depends on the state representation, task distribution, and model. Our example compares two feature sets under one synthetic distribution.
Contact may be sensed directly with tactile or force sensors, inferred from vision, or left latent. Accurate pose prediction does not prove that a model internally represents contact geometry in a particular way. In our toy, the contact-proximity feature is supplied to one predictor, not discovered through supervision on puck motion alone.
Dexterous control
Dexterous in-hand manipulation can involve several independently moving fingers and changing contacts. An object with six pose degrees of freedom is not necessarily directly actuated in all of them: its motion is mediated by contact constraints and external forces such as gravity. In-hand rotation of a cube and rolling a ball between fingers are examples. Unlike the simple push and grasp-success examples above, contact can change repeatedly while the object remains supported. Multi-object pushing and complex grasping can also be high-dimensional.
Depending on the task and sensors, useful dexterous-model predictions may include:
- The object's pose in the hand across a motion that may involve occasional break of contact.
- The contact state on each fingertip: sticking, slipping, or rolling.
- The forces required to maintain the target pose without crushing the object.
- Task outcomes, like the final orientation of the cube or the position of a screw.
Three fingers with two contact coordinates each plus a six-degree-of-freedom object pose give twelve continuous coordinates in one explicit representation; discrete contact modes add combinations. This is an illustration, not a claim that every pushing or grasping task has fewer state variables. A latent model can compress this representation, but compression does not guarantee preservation of distinctions needed for control. The state choice depends on sensing, task, and evaluation.
Some tasks combine pushing, grasping, carrying, and in-hand reorientation. A system may use separate predictors at different resolutions, but model handoffs then introduce their own state-estimation and timing errors. Other systems use a shared policy or predictor; the appropriate decomposition remains an empirical design choice.
Visual Foresight and Action-Conditioned Video
Visual foresight predicts future observations conditioned on candidate actions rather than requiring an explicit object-pose output. Images reveal some task-relevant changes but omit hidden forces, occluded geometry, and material properties. A video model therefore does not automatically predict the full physical state or generalize to every visible scene.
Given observation history and candidate actions , a video model predicts future observations . This action-conditioned view connects to Part IX. Design choices include:
- How actions are represented: end-effector deltas, joint commands, or other controls.
- Whether image prediction preserves object motion and contacts relevant to the task.
- How to evaluate both prediction and downstream action quality.
Some manipulation systems use compact controls, though the action-space dimension and prediction difficulty vary. A visually imperfect rollout can still support useful actions if it preserves the task-relevant effects. Conversely, plausible frames need not yield effective plans, so downstream evaluation matters alongside visual prediction metrics.
Training data can come from autonomous interaction, teleoperation, scripted actions, or simulation. Finn and Levine's action-conditioned video prediction and Ebert and colleagues' visual foresight work illustrate planning from robot interaction data without explicit object-pose labels in their demonstrated settings. Avoiding pose labels does not eliminate tracking, occlusion, or out-of-distribution problems; cloth, ropes, and fluids can be particularly hard to model reliably.
Rollouts can accumulate errors, especially near uncertain contacts, and video prediction can be computationally costly. Recurrence, latent frames, or generative sampling (Part V) offer different trade-offs; none guarantees a particular useful horizon. That horizon must be measured for the data, model, and control task at hand.
Using a visual foresight model to act
Given a video model, planning can follow the template introduced in Part VII:
- Sample a set of candidate action sequences from a simple proposal distribution.
- Roll out the video model for each candidate on the conditioned current observation.
- Score each predicted clip by a task cost: distance of the object to the target in the final frame, presence of contact, stability of the predicted motion.
- Execute a short prefix of a selected sequence, observe again, and re-plan; this is a closed-loop option, not what the later toy planner implements.
An image-based goal scorer may compare pixels, track designated points, segment an object, or estimate pose; each choice can be noisy. Receding-horizon replanning or a lower-level feedback controller may correct observed errors, subject to sensing and actuation limits.
Early action-conditioned video-planning papers demonstrated useful pushing and rearrangement behavior without explicit object-pose supervision in those experiments (Finn and Levine; Ebert et al.). Latent and pixel-space approaches make different computational and diagnostic trade-offs, but neither is uniformly easier to train or better at longer horizons. Some methods decode images for scoring or inspection; others plan entirely in latent space.
Task Planning with Learned Affordances
The preceding sections mainly examined predictions of action effects, although grasp selection can also depend on later task outcomes. A manipulation task may also need to choose which object to interact with, where to reach, and in what order. Affordance estimates can help propose actions; when a world model is available, it can predict their consequences.
An affordance describes an action possibility relative to an agent, object, and task. A model may encode it as a success probability, action value, feasible contact region, or predicted effect; these quantities are not interchangeable.
Learned affordance models can condition on a scene, action, and sometimes a goal. In a rightward-push task, a contact point on one side of an object might be useful, but the outcome depends on geometry, friction, and the action direction. The toy map below instead shows a predictor's one-step displacement for a fixed rightward command; it is not a goal-conditioned success map.
Structure of a learned affordance model
One possible affordance architecture has three components:
- Feature extractor. Encodes an image or structured scene state.
- Optional goal conditioning. Specifies a target pose, instruction, or image when the output is task-specific.
- Prediction head. Estimates an explicitly defined quantity, such as success probability, action value, or displacement.
Interaction logs or simulated outcomes can supervise such a head. For a probabilistic goal-conditioned formulation, might estimate whether action at location helps achieve goal from observation . Other heads predict an effect vector instead. Language-conditioned variants connect to Part IX, but their calibration and transfer require evaluation.
How affordances plug into planning
One possible architecture is a two-level loop:
- High-level selector. Uses estimated action effects or feasibility to propose a sub-task or contact region.
- Low-level controller. Executes a candidate action and uses feedback; it may use model-predictive control, but need not.
This decomposition can help diagnose errors, but a failed push may involve perception, effect prediction, planning, or control together. Component tests and interventions, rather than the architecture alone, help assign causes (Part XI).
Task-and-motion planning can also search over symbolic predicates and call geometric or learned feasibility checks (Part VII). A PDDL-style planner is one option; the affordance predictor above does not require one.
Relating affordances to a world model
An affordance predictor can be learned directly from outcomes without an explicit world model. Where both are present, comparing predicted effects can expose disagreement: a proposed rightward push is suspect if the dynamics model instead predicts a tip. Joint training may encourage agreement but does not guarantee physical or task-level consistency. Testing the proposed action in the real system remains decisive.
A Minimal Manipulation World Model in Code
We will build a miniature 2D pushing predictor. The state contains puck and pusher positions , and the action is a commanded pusher displacement . A post-move distance test gates an ad hoc puck coupling gain. This is a deliberately synthetic rule, not a physical friction model or a reproduction of a planar-pushing benchmark. It lets us inspect feature choice, prediction error, open-loop rollout, and a one-step effect map; it does not demonstrate grasping, dexterity, or visual foresight.
We will do six things:
- Generate random-play transition data from the contact-switch simulator and inspect how rarely the pre-move proximity feature is appreciably nonzero.
- Fit a naive linear world model that ignores contact.
- Fit a contact-aware world model that has access to a contact feature.
- Compare one-step and multi-step rollout error.
- Plan a push to a target using random shooting against the learned model.
- Query the model for a one-step effect field, an affordance proxy.
The core comparison gives one regression additional contact-proximity features. We report both one-step and recursive rollout error; neither alone measures task success.
Before building the models, we fix a random seed so every experiment is reproducible, and we keep each model's input features explicit so the comparison is fair.
Everything in this chapter runs in a single notebook, and the notebook follows a strict computation and rendering boundary: data generation, model fitting, rollout evaluation, planning, and affordance-map computation all happen in ordinary (non-theme-rerun) cells that run once. The figures are produced by separate presentation-only cells, each of which declares #| theme-rerun: true, imports the shared plotting style locally, calls use_book_style(), and renders only values that were prepared in earlier cells. No cell that draws a figure trains a model, samples from a simulator, or mutates shared data.
Setup
This cell seeds the examples and locates the repository root by walking upward from the working directory. Launch it from the repository or a descendant directory; it will not find the style module from an unrelated working directory.
Simulating contact-rich dynamics
The dynamics fit on one screen: the pusher moves by the commanded action subject to workspace clipping; the puck moves by a fixed fraction of the command when the pusher's post-move position passes a distance test. The gain is not a friction coefficient.
The simulator and feature function are separate. The latter measures pre-move proximity and therefore does not encode the simulator's post-move contact decision exactly.
import matplotlib.pyplot as plt
import numpy as np
from book_plot_style import PALETTE, polish_axes, use_book_style
R_CONTACT = 0.35 # pusher-puck center-distance switch threshold
ALPHA = 0.5 # puck follows pusher at half the commanded displacement
BOUND = 2.0 # workspace half-width
def step(state, action):
px, py, qx, qy = state
ux, uy = action
# Pusher moves by the command, clipped to the toy workspace
qx_new = np.clip(qx + ux, -BOUND, BOUND)
qy_new = np.clip(qy + uy, -BOUND, BOUND)
# Post-move pusher position gates puck motion. The full command is used
# even if the pusher was clipped: a nonphysical toy artifact.
dist = np.hypot(qx_new - px, qy_new - py)
if dist < R_CONTACT:
px_new = px + ALPHA * ux
py_new = py + ALPHA * uy
else:
px_new, py_new = px, py
return np.array([px_new, py_new, qx_new, qy_new])
def contact_feature(state):
px, py, qx, qy = state
d = np.hypot(qx - px, qy - py)
return float(np.clip(1.0 - d / R_CONTACT, 0.0, 1.0))The step function advances the pusher and then checks whether its new position is within R_CONTACT of the puck; only then does the puck inherit ALPHA times the full commanded action. If clipping shortened the pusher's move, the puck update is still based on the command—a nonphysical artifact of this toy. The feature decreases linearly from one at zero pre-move separation to zero at the distance threshold. It is an input to a fitted regression, not the simulator's contact rule. A deliberately simplified soft-contact illustration would have displacement
where:
- : the puck's displacement during step , computed as the change in puck position
- : the ad hoc coupling gain introduced above
- : the contact feature, equal to 1 at zero pusher-puck separation and decreasing linearly to 0 at the distance threshold
- : the pusher's commanded displacement
This soft formula is not the fitted model's exact equation: Ridge learns coefficients for state, action, , and , including an intercept. As a function of two-dimensional relative position, the supplied proximity feature is continuous but non-differentiable at coincident positions and at the distance-threshold boundary. As a function of scalar separation , its interior kink is at . It may help a regression represent contact dependence, but timing mismatch can limit that help.
The next plot compares two hypothetical weights at pre-move separation. It illustrates the feature's shape; the simulator itself decides contact from post-move separation.
The sweep that feeds that plot is prepared here, in a computation-only cell, so the figure cell below only renders the precomputed arrays.
## Sweep pre-move separation from zero to beyond the distance threshold.
## The hard curve is a pre-move illustration, not the simulator's post-move rule.
## The ramp is continuous but has a kink at the threshold.
separation = np.linspace(0.0, 0.6, 240)
hard_switch = (separation < R_CONTACT).astype(float)
smooth_feature = np.clip(1.0 - separation / R_CONTACT, 0.0, 1.0)
We will give this feature to one of the two world models and withhold it from the other, so we can see the difference. The contact-aware model uses it as an extra input feature, while the naive model sees only the raw state and action.
Collecting a training set
We generate synthetic random-action episodes from the toy simulator. This is not a faithful model of robot data collection; in this sampling scheme, pre-move proximity is rare, which matters for an unweighted regression loss.
Each transition is stored as an independent (state, action, next_state) tuple, which keeps the regression dataset simple: random play is the data source, not a trained controller.
rng = np.random.default_rng(0)
N_EPISODES = 400
T_EPISODE = 30
states = []
actions = []
next_states = []
for _ in range(N_EPISODES):
s = np.array(
[
rng.uniform(-0.5, 0.5),
rng.uniform(-0.5, 0.5), # puck
rng.uniform(-1.5, -1.0)
if rng.random() < 0.5
else rng.uniform(1.0, 1.5),
rng.uniform(-1.5, 1.5), # pusher
]
)
for _ in range(T_EPISODE):
a = rng.normal(scale=0.15, size=2)
s_next = step(s, a)
states.append(s)
actions.append(a)
next_states.append(s_next)
s = s_next
states = np.array(states)
actions = np.array(actions)
next_states = np.array(next_states)Transitions: 12000 Transitions with contact feature > 0.1: 64 (0.5%)
The printed summary counts states whose pre-move contact feature exceeds 0.1. It is a proximity statistic, not the fraction of transitions in which the simulator's post-move contact test succeeds. Most sampled states have a feature near zero. An unweighted loss over this dataset can therefore be dominated by common non-proximal states; whether those are also the easiest transitions depends on the sampled actions and outcomes. Targeted contact evaluation or deliberate contact-rich sampling would test the rare regime more directly.
That concentration of the data is easy to see directly. Plotting the contact feature over every transition shows how few samples carry contact information.
## Contact feature value for every transition in the random-play dataset. This
## reuses the states already generated and performs no new sampling.
dataset_contact = np.array([contact_feature(s) for s in states])
Two world models
We compare a naive linear model that regresses the next state from the current state and action using only a fixed affine map, with no contact term of any kind:
where:
- : the state at time , the concatenation of the puck's position and the pusher's position
- : the pusher's commanded displacement during step
- : the learned state-transition matrix, capturing how the current state propagates in the absence of action
- : the learned control matrix, mapping the commanded displacement into the next state
- : the learned intercept vector (Ridge fits an intercept by default)
- : the model's predicted next state, whose error is measured against the simulator's true
- The affine prediction approximates the true next state; its errors are measured below.
against a contact-aware model that augments the input with the contact feature and its product with the action. The fitted coefficients remain fixed. The extra nonlinear features let this linear regression represent an action-dependent contact response without an explicit latent contact mode:
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_squared_error
def augmentation_naive(state, action):
return np.concatenate([state, action])
def augmentation_contact(state, action):
c = contact_feature(state)
return np.concatenate([state, action, [c], c * action])
## The two feature builders share the current state and action. The contact-aware
## version adds the scalar contact feature and its product with the action, so the
## regression can represent a contact-sensitive interaction without a hidden mode variable.
X_naive = np.array([augmentation_naive(s, a) for s, a in zip(states, actions)])
X_contact = np.array(
[augmentation_contact(s, a) for s, a in zip(states, actions)]
)
Y = next_states
split = int(0.8 * len(X_naive))
train = slice(0, split)
test = slice(split, len(X_naive))
model_naive = Ridge(alpha=1e-4).fit(X_naive[train], Y[train])
model_contact = Ridge(alpha=1e-4).fit(X_contact[train], Y[train])
## Both models use the same ridge penalty so the comparison isolates the effect of
## the contact feature rather than differences in regularization strength.Naive one-step RMSE: 0.0186 Contact one-step RMSE: 0.0183
Both models are small regressions on hand-crafted features; only one receives the smoothed contact feature and its action products. In this seeded run, the contact-aware model's one-step RMSE is only marginally lower. Contact is rare in the sampled random-play data, and the feature uses pre-move separation while the simulator checks post-move separation. The added feature therefore does not isolate the simulator's hard switch. This comparison tests what the feature contributes under this data distribution; it does not establish that contact-aware modeling is generally better or worse.
Multi-step rollout
One-step and recursive errors answer different questions. A planner may feed predictions back into the model; recursive rollout exposes the resulting distribution shift. The following 15-step test does not measure closed-loop control or establish an asymptotic growth law.
The first 9,600 transitions come from 320 complete episodes and the remaining 2,400 from 80 complete episodes, so the one-step train/test slices do not split episodes. The recursive evaluation below draws fresh start states and actions (50 trials), rather than reusing those test episodes. Both models see identical actions within each trial.
def rollout(model, aug_fn, s0, actions):
preds = [s0]
s = s0
for a in actions:
s = model.predict(aug_fn(s, a).reshape(1, -1))[0]
preds.append(s)
return np.array(preds)
## The earlier contiguous split aligns with 30-step episode boundaries.
## The following recursive trials are freshly sampled.
HORIZON = 15
n_eval = 50
rollout_errors = {"naive": [], "contact": []}
for _ in range(n_eval):
s0 = np.array(
[
rng.uniform(-0.5, 0.5),
rng.uniform(-0.5, 0.5),
rng.uniform(-1.2, 1.2),
rng.uniform(-1.2, 1.2),
]
)
acts = rng.normal(scale=0.15, size=(HORIZON, 2))
truth = [s0]
s = s0
for a in acts:
s = step(s, a)
truth.append(s)
truth = np.array(truth)
for name, mdl, aug in [
("naive", model_naive, augmentation_naive),
("contact", model_contact, augmentation_contact),
]:
preds = rollout(mdl, aug, s0, acts)
err = np.sqrt(mean_squared_error(truth, preds))
rollout_errors[name].append(err)
rollout_errors = {k: float(np.mean(v)) for k, v in rollout_errors.items()}naive mean 15-step RMSE: 0.0452 contact mean 15-step RMSE: 0.0480
The printed whole-trajectory RMSE averages over sixteen states: the exact initial state and fifteen recursively predicted next states. The initial zero-error row is therefore included. The separate per-step curves below show how prediction error changes across the fifteen-step horizon. In this seeded run, the contact-aware model has slightly higher mean rollout RMSE than the naive baseline despite its slightly lower one-step error. Reusing predicted states changes the input distribution at every step; an extra feature does not guarantee a better rollout when contact examples are sparse and its timing differs from the simulator's contact test. For a planner, this mismatch matters because candidate actions are scored on predicted consequences. Neither rollout RMSE nor one-step RMSE, however, substitutes for closed-loop task success.
The per-step plot makes the horizon dependence visible. Both error curves rise; their relative ordering varies across steps rather than separating consistently in favor of the contact-aware model.
## Fresh recursive rollouts to measure error across this 15-step horizon.
## Evaluation design: 30 start states and action sequences drawn from a dedicated
## seeded generator (seed 7), independent of the training and planning samples.
## Both models see identical start states and actions. Per-step error is the
## root-mean-square over the four state dimensions averaged across the 30
## rollouts. No tie-breaking is needed because each step yields one scalar error.
rng_roll = np.random.default_rng(7)
N_ROLL = 30
per_step_naive = np.zeros(HORIZON)
per_step_contact = np.zeros(HORIZON)
for _ in range(N_ROLL):
s0 = np.array(
[
rng_roll.uniform(-0.5, 0.5),
rng_roll.uniform(-0.5, 0.5),
rng_roll.uniform(-1.2, 1.2),
rng_roll.uniform(-1.2, 1.2),
]
)
acts = rng_roll.normal(scale=0.15, size=(HORIZON, 2))
truth = [s0]
s = s0
for a in acts:
s = step(s, a)
truth.append(s)
truth = np.array(truth)
preds_naive = rollout(model_naive, augmentation_naive, s0, acts)
preds_contact = rollout(model_contact, augmentation_contact, s0, acts)
for h in range(1, HORIZON + 1):
per_step_naive[h - 1] += np.sqrt(
np.mean((truth[h] - preds_naive[h]) ** 2)
)
per_step_contact[h - 1] += np.sqrt(
np.mean((truth[h] - preds_contact[h]) ** 2)
)
per_step_naive /= N_ROLL
per_step_contact /= N_ROLL
rollout_horizon = np.arange(1, HORIZON + 1)
Planning a push with random shooting
We use the contact-feature model for a small random-shooting demonstration (Part VII). It samples 400 candidate 12-step action sequences, scores each by the model-predicted final puck distance to a target, and selects one. The chosen sequence is then replayed open-loop in the simulator. This is not receding-horizon control, and the score does not establish task success.
The planner searches only over actions; the contact-aware model supplies the prediction that scores each candidate sequence, which keeps the planning loop separate from the learned dynamics.
def plan_push(
model,
aug_fn,
s0,
target,
n_samples=400,
horizon=12,
action_scale=0.15,
plan_rng=None,
):
if plan_rng is None:
plan_rng = rng
best_score = np.inf
best_actions = None
best_preds = None
for _ in range(n_samples):
acts = plan_rng.normal(scale=action_scale, size=(horizon, 2))
preds = rollout(model, aug_fn, s0, acts)
final_puck = preds[-1, :2]
score = np.linalg.norm(final_puck - target)
if score < best_score:
best_score = score
best_actions = acts
best_preds = preds
return best_actions, best_preds, best_score
start_state = np.array([0.0, 0.0, -1.0, 0.0])
target = np.array([0.6, 0.4])
best_actions, best_preds, best_score = plan_push(
model_contact, augmentation_contact, start_state, target
)Final puck distance to target: 0.4156 Puck start: [0. 0.] Puck end: [0.265 0.154] Target: [0.6 0.4]
The printed score is the model's predicted final distance to the target, not the achieved task error. In this seeded run it remains much larger than the action sampling scale, so the selected sequence is only a coarse approach, not a successful placement. The gap can reflect candidate coverage, the learned transition, and their interaction; the run does not isolate one cause. Replaying the sequence in the simulator below tests whether the predicted trajectory survives execution. A practical controller would execute a short prefix, observe the result, and replan. That closed-loop outcome, rather than the planning score alone, would determine task success.
Replaying the chosen actions through the simulator lets us compare predicted and true puck paths. The pusher path drawn below is also predicted by the learned model; the simulator's pusher path is available separately in true_preds.
## Roll the chosen action sequence through the true simulator so the learned
## prediction can be compared against ground truth. No sampling happens here; the
## actions are exactly the ones returned by the planner above.
true_preds = [start_state]
s = start_state
for a in best_actions:
s = step(s, a)
true_preds.append(s)
true_preds = np.array(true_preds)
pred_puck = best_preds[:, :2]
true_puck = true_preds[:, :2]
pusher_path = best_preds[:, 2:]
Affordance map from the same data
The fitted model can be queried for a one-step effect field: from each selected pusher start, apply the same rightward command and plot the predicted puck displacement. This is one deterministic model query per cell, not an average, probability, confidence estimate, or goal-conditioned success score. Arrow magnitudes may vary for reasons that include sparse contact data and model error; their strongest values need not occur to the left of the puck.
The map displays only start cells whose pre-move contact feature exceeds 0.05. This is a visualization filter, not a guarantee that the simulator's post-move contact test would accept or reject the same cell.
## Grid of pusher start positions around a puck at the origin
puck = np.array([0.0, 0.0])
grid_vals = np.linspace(-0.6, 0.6, 15)
push_action = np.array([0.15, 0.0]) # always push right
aff_x, aff_y, aff_dx, aff_dy = [], [], [], []
for gx in grid_vals:
for gy in grid_vals:
s0 = np.array([puck[0], puck[1], gx, gy])
# Filter by pre-move proximity; the simulator tests post-move contact.
if contact_feature(s0) <= 0.05:
continue
s1 = model_contact.predict(
augmentation_contact(s0, push_action).reshape(1, -1)
)[0]
aff_x.append(gx)
aff_y.append(gy)
aff_dx.append(s1[0] - puck[0])
aff_dy.append(s1[1] - puck[1])
aff_x = np.array(aff_x)
aff_y = np.array(aff_y)
aff_dx = np.array(aff_dx)
aff_dy = np.array(aff_dy)
max_mag = max(np.hypot(aff_dx, aff_dy).max(), 1e-6)
The arrow field queries the learned model at selected start positions without running the simulator again. Its pre-move filter keeps the display near the puck but can omit starts that would make contact after the commanded move. Each arrow encodes the model's predicted puck displacement for the same rightward action. A task planner could rank candidate starts with these predictions, then validate selected actions in closed loop. This map alone does not establish that the strongest-looking arrow corresponds to a successful real push.
The arrow magnitude is the model's predicted effect, not confidence. A large vector can be confidently right, confidently wrong, or uncertain; this plot cannot distinguish those cases. Comparing queries with held-out simulator transitions or real actions would test whether these effects are useful for choosing a push.
Key Parameters
The key parameters used in the miniature pushing experiment are:
- R_CONTACT: Pusher-puck center-distance threshold for the toy simulator's contact switch.
- N_EPISODES: Number of random-play episodes collected for training and evaluation.
- ALPHA: Fraction of the commanded displacement inherited by the puck when the simulator's contact test succeeds.
- BOUND: Workspace half-width used to clip pusher positions.
- grid_vals: Sampled pusher start positions for the illustrative affordance field. Increase this to visualize the affordance field at finer spatial resolution.
- push_action: Fixed displacement used to query that field.
- max_mag: Maximum predicted displacement magnitude, used to normalize displayed arrow lengths.
- T_EPISODE: 30 transitions per episode; with 400 episodes, the dataset contains 12,000 transitions.
- Train/test split and Ridge alpha: An episode-aligned 80/20 split (320/80 episodes); each linear model uses
Ridge(alpha=1e-4). - HORIZON and evaluation counts: Recursive rollouts use 15 actions, with 50 fresh trials for the aggregate RMSE and 30 fresh trials for the per-step plot.
- Planner search settings: Random shooting samples 400 action sequences of length 12 with Gaussian action scale 0.15.
Limitations and Impact
The miniature model isolates a feature-choice and rollout-evaluation problem. Translating those observations to a robot requires additional perception, actuation, contact, and safety tests.
The following are design questions, not measured performance claims about deployed manipulation systems.
The first limitation is observation realism. Our model sees exact positions. A physical system may have camera, depth, tactile, and proprioceptive measurements with occlusion, calibration error, and latency. Prediction should be tested under the actual sensing regime and material/lighting variation, not extrapolated from this state-observed toy. Reliability and evaluation methods are discussed in Part XII and Part XI. We have not measured a simulation-to-real error ratio here.
The second limitation is horizon. We evaluate only 15 steps from synthetic starts. Error rises over this window but the experiment does not establish exponential growth or any useful time limit on other models. Re-observing and replanning can reduce open-loop drift, while recurrent latent models (Part VIII) offer another design option. Their relative effectiveness must be measured on the intended task.
The third limitation is contact uncertainty. Our toy has one distance-gated interaction and no slip, rolling, or force sensing. A richer task may benefit from explicit modes, ensembles, or probabilistic predictions. None is universally preferable, and this chapter supplies no industrial deployment comparison.
The fourth limitation is task composition. A multi-stage system may hand object state between pushing, grasping, and placement components. Each interface needs evaluation because the receiving component may see inputs outside its training distribution. An affordance estimate alone does not solve this mismatch.
The central evaluation lesson is to keep prediction, uncertainty calibration, and decision quality separate. Our contact-feature regression slightly improves one-step RMSE but slightly worsens mean 15-step RMSE in this seeded run; the open-loop plan misses its target. Those results illustrate why a favorable local metric is insufficient evidence of a useful controller. Physical-robot impact would require separate task-success and safety evidence.
Summary
- Contact can change the relevant dynamics or make outcomes uncertain, but not every contact model has a hard discontinuity or requires distributional output.
- Explicit, latent, object-centric, and hybrid states trade off observability, structure, and estimation error; none dominates by definition.
- Grasping, pushing, and dexterous manipulation overlap and may call for different predictors or a shared model, depending on the task.
- Visual foresight predicts action-conditioned observations, not hidden physical state; its planning value must be tested downstream.
- Affordance estimates can be learned without a world model. If paired with one, compare their predicted effects rather than assuming consistency.
- The code builds a miniature planar pushing model and a naive baseline. Its contact feature gives only a marginal one-step gain and slightly worse mean rollout error in this seeded run; the imperfect model is also used for a coarse plan and an illustrative affordance map.
- Separate one-step prediction, recursive prediction, and closed-loop task success, as in Part XI.
The planned next chapter, Locomotion, Humanoids, and Navigation, examines whole-body contact and spatial control within Part X.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about robotic manipulation world models, contact-rich dynamics, and planning.
Robotic Manipulation
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!