Part of World Models Handbook
Explains how fixed datasets limit offline RL, why model exploitation can mislead planners, and how pessimism and policy constraints reduce unsupported choices.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Offline and Conservative Model-Based RL
Picture a warehouse robot that has been running for three years. It has driven several thousand kilometres, followed human-designed routes, bumped into the same four pallets, and logged every state, action, and outcome. You now have a large dataset. You have no way to collect more. Can those logs support a better controller than the one that produced them?
The answer is not obvious and not always yes. The dataset was produced by some behaviour policy, perhaps a careful human operator or an older controller. Every transition reflects choices that policy made. If it never tried to squeeze between two pallets, the dataset contains no direct evidence about what happens when you try. If it never braked hard on a wet floor, the dataset is silent on that too. Improving the controller may require different actions, and those actions may lack coverage in the logs. That is the tension we spend the chapter unpicking.
Every model-based method we have covered so far in Part VIII: Decision-Centric Research Lineages assumed a closed loop. PETS, MBPO, Dreamer, MuZero, and TD-MPC all interleave acting in the world with learning from the consequences. The world model is wrong in a hundred small ways, the agent tries something, reality corrects it, and the model improves. That correction signal is the engine.
When the agent acts, the observed successor provides a targeted new sample at the state-action pair it tried. In a stochastic environment, one outcome is not a measurement of how wrong the model's predictive distribution is; repeated observations help distinguish environmental randomness from systematic model error. Over many interactions a well-designed learning loop can concentrate model corrections where the agent visits. The state distribution induced by the policy and the region where the model improves can then move together. Offline learning loses that feedback loop.
Offline reinforcement learning removes that engine during training. The agent gets one fixed dataset and no new interaction while learning its policy. Deployment may provide feedback later, or an offline-to-online phase may collect more data, but neither can correct an unsupported decision before this offline training stage ends. The policy must therefore be selected using evidence already in the log.
That single change breaks almost everything. Two failure modes follow directly from it, and they are worth stating plainly before any mathematics:
- The log alone does not constrain the model off its support. If your dataset never contains the state-action pair , competing predictions there can fit the logged transitions equally well. A model that says "you teleport to the goal" and a model that says "you fall into a pit" may both fit the training set. Architecture, prior knowledge, or justified structural assumptions can distinguish them, but the log provides no direct target at that pair to pull the model toward the truth.
- The optimiser can exploit unsupported optimistic estimates. A policy improvement step that maximises estimated value favours regions with high estimates. If the model happens to be optimistic off-support, that step can select a policy that visits a region where the log provides little evidence. This is model exploitation, and it is not a defect of one algorithm: it is a risk whenever policy selection relies on estimates that are insufficiently grounded.
The mitigations address those two risks. We can quantify how uncertain the model is (ensembles, counts, densities). We can be pessimistic, subtracting an uncertainty penalty to make poorly supported choices less attractive. Or we can constrain the policy itself to stay near the behaviour that produced the data. The rest of the chapter develops all three ideas, illustrates uncertainty and pessimism on a gridworld, and then turns to the meta-problem that offline RL never escapes: how do you evaluate a policy you cannot run?
Learning from Fixed Datasets
This section fixes notation, states a coverage assumption that can support offline learning, and derives the simulation lemma, which is the single most useful equation in the chapter. It explains why a model that looks accurate can produce a badly wrong value function, and it draws the line between model-based and model-free approaches to the offline problem. Later we will distinguish coverage from justified structural assumptions that can also constrain unobserved outcomes.
The setup
We have a Markov decision process as introduced in Part II: Inference, Decisions, and Control. To keep the comparison with learned models clean, we will assume throughout that the state is fully observed and the MDP is stationary. Neither assumption is free. Under partial observability the effective state is a belief, and the coverage problem becomes harder because two observation histories can be indistinguishable while requiring different actions. We will flag where that matters.
The stationarity assumption matters too. If the environment drifts over time, last year's transitions may not describe the world the deployed policy will face. In the warehouse example, a reorganised floor plan changes the dynamics represented by the logs. Here we assume no such drift: training and deployment share the same environment.
Instead of interacting, we are handed a dataset
where:
- : the fixed dataset of transitions
- : the -th observed state
- : the -th observed action taken in that state
- : the observed reward, whose conditional mean is
- : the next state, drawn from the transition distribution
- : the behaviour distribution, the state-action distribution induced by whatever process generated the logs
The behaviour distribution might come from a human demonstrator, a hand-tuned controller, a previously deployed policy, or five such policies mixed together. We do not need to know which. We only need the samples. Notice that the dataset records actions and their outcomes but never explains them. Nothing tells us why the behaviour policy chose a given action, which matters later when we want to imitate it in states where the model is uncertain.
The objective is unchanged from online RL. We want a policy maximising
where:
- : the expected discounted return of policy
- : the discount factor in
- : the reward received at step
- : the initial-state distribution, with
- : the policy, giving the action distribution at state
- : the environment transition distribution
The difference from online RL is entirely in what we are allowed to do while searching for . Online, we can test , observe the return, and adjust. Offline, that loop is severed: the objective is still defined against the true environment, but we can only ever evaluate candidate policies against our beliefs about it.
These terms are used loosely in the literature and the distinctions matter.
- Off-policy learning means the update rule can use data from a policy other than the current one. It says nothing about interaction. Deep Q-learning is off-policy and can still collect fresh data.
- Offline (or batch) learning means the dataset is fixed during the offline training phase; the learner collects no new transitions in that phase. Later deployment or offline-to-online fine-tuning is a separate stage. This is a restriction on data acquisition during training, not on the update rule.
- Off-policy evaluation is the sub-problem of estimating for a candidate using data from . It arises in online settings too, but offline RL cannot avoid it.
The last point is the practical sting. In online RL you evaluate candidates by running them. Offline you must estimate their value from the same fixed data you used to train, which means your evaluation inherits your coverage problems. We return to this in the final section.
Occupancy measures and the coverage assumption
It is convenient to work with discounted state-action occupancy measures rather than raw trajectories. For a policy with initial distribution , define the mass of any measurable state-action set by
where:
- : the discounted occupancy mass that policy assigns to set
- : the probability of visiting at step under policy starting from
- : the normalising factor that makes the sum a probability distribution
The factor makes a probability distribution over . In a discrete space, abbreviates the mass of the singleton ; in a continuous space, it denotes a density only when one exists with respect to a stated reference measure. Why bother with this object rather than working with the trajectory distribution directly? Because an occupancy measure summarises "where the policy spends its time," weighting near-term visits more heavily than far-future ones. It converts the trajectory of the policy into a static picture of the space, which is exactly what we need to compare two policies and to reason about how much their visited regions overlap. With this normalisation, the objective is a simple expectation:
Now compare with . If a policy spends time where the behavior distribution has substantial mass, the data can inform those choices; where it has none, learning requires additional structure. For the displayed ratio, read both quantities as masses in a discrete space or densities with respect to a common reference measure in a continuous space. This intuition is captured by the concentrability coefficient:
where:
- : the concentrability coefficient of policy
- : the discounted occupancy of under policy
- : the behaviour distribution's density at
- : the supremum over state-action pairs in the discrete case, or the essential supremum of the density ratio in the continuous case
Define the ratio as infinite when the target occupancy assigns positive mass to a region where the behavior distribution assigns zero; zero-over-zero pairs do not affect it. In continuous spaces, this requires absolute continuity and an essentially bounded density ratio. Thus finite requires the occupancy of that policy to be supported by the behavior distribution. Some offline-RL guarantees assume coverage of an optimal policy, , together with further statistical and algorithmic conditions. Covering one optimal policy is weaker than covering every policy, but it does not by itself make any particular learning algorithm successful. A small coefficient says the behavior distribution gives reasonable weight to what the policy visits; a large one signals weak coverage.
A dataset concentrates on a policy class with coefficient if, for every ,
Roughly: no state-action pair that some policy in visits often is more than a factor rarer in the data. Small means the data covers what the policies want to do.
Coverage is one way to make offline learning identifiable. Without coverage or justified structure linking unobserved outcomes to observed ones, no method can guarantee good decisions uniformly over the unrestricted MDP class.
The coefficient compares state-action occupancy densities. The next picture simplifies them to their state marginals: if the target visits states the behaviour never reaches, then it necessarily visits unsupported state-action pairs and is infinite. A gap confined to actions at otherwise covered states would also matter, but this state-only picture does not show it.
import numpy as np
# Schematic occupancies on a one-dimensional state space s in [0, 1].
# The behaviour density is truncated to [0.0, 0.5]. Most target mass lies
# to the right of 0.5, but Gaussian tails also overlap behaviour support
# below 0.45; the exact overlap is not confined to [0.45, 0.5].
_state = np.linspace(0.0, 1.0, 501)
_behaviour_support = (_state >= 0.0) & (_state <= 0.5)
def _bump_mixture(x, centres, width):
"""Normalised sum of Gaussian bumps, used only for this schematic figure."""
bumps = np.exp(-0.5 * ((x[:, None] - np.asarray(centres)) / width) ** 2)
density = bumps.sum(axis=1)
return density / (density.sum() * (x[1] - x[0]))
rho_beta = _bump_mixture(_state, [0.15, 0.30, 0.45], 0.07) * _behaviour_support
rho_beta = rho_beta / (rho_beta.sum() * (_state[1] - _state[0]))
d_pi = _bump_mixture(_state, [0.45, 0.65, 0.85], 0.07)
unsupported_mask = (rho_beta <= 0.0) & (d_pi > 1e-4)
Consider two MDPs that agree everywhere puts mass and disagree on an unseen action. Both produce the same dataset. In one, the unseen action is better than the logged choice; in the other, it is worse. Any data-only algorithm makes the same choice in both, so it cannot guarantee near-optimality in both over this unrestricted class. Coverage or justified structural information is needed to distinguish them.
The fixed log alone cannot identify the outcome of that unseen action. An offline method needs either relevant coverage or assumptions that constrain such outcomes.
The simulation lemma
Now suppose we learn a model from and use it in place of the environment. How wrong is the resulting value function? The answer is the simulation lemma, and it deserves a careful derivation because it is the theoretical heart of the whole chapter.
Fix a deterministic policy for this derivation; a stochastic policy gives the same identity after averaging each action-dependent term over . Write for its value under the true MDP and for its value under the model. Define the gap . Expanding both Bellman equations,
Subtracting the second from the first and adding and subtracting a term that evaluates under the model,
where:
- : the value gap at state
- : the reward-model error, abbreviated
- : the transition error, evaluated using the true value function
- : the expected continuation of the gap under the model's dynamics
The adding-and-subtracting trick is the algebraic seed of the whole result. We insert not because it was there already, but to isolate the reward mismatch and the transition mismatch as a single "local inconsistency" while leaving a clean recursion for the remaining gap. Everything that follows is a matter of chasing where that local inconsistency gets weighted.
The first two terms do not depend on ; call their sum the local inconsistency . The recursion unrolls along the trajectory distribution generated by the model. Write for the model's normalised discounted state-action occupancy and for its state marginal, both induced by the learned model rather than the true MDP. Averaging the initial-state gap over gives
Here , , and . A pointwise value-gap identity would instead use a model occupancy conditioned on that particular starting state; the displayed occupancy is averaged over .
The boxed equation averages local errors over states visited by a fixed policy in the learned model, not necessarily in reality. If the model's dynamics place substantial occupancy in a region where its predictions are wrong, that region contributes heavily to the value gap. The identity does not say that error itself attracts a fixed policy toward that region.
Error compounding and model exploitation are distinct steps. The fixed-policy identity describes how one-step errors propagate through model rollouts. When a policy is then selected by maximizing learned-model value, it may favor optimistic, poorly supported regions; that selection can change the model occupancy in the identity. The two mechanisms can reinforce each other, but neither the identity nor a fixed-policy bound alone proves that the errors concentrate exactly where they are largest.
A worked example: how vacuous the worst-case bound becomes
Bound the two error terms to get a usable inequality. Suppose the reward model is off by at most and the transition model has total-variation error at most :
where bounds the reward error, bounds the total-variation transition error, and is the maximum absolute reward. Since and ,
where:
- : the maximum reward-model error,
- : the maximum total-variation transition-model error
- : the maximum absolute per-step reward
- : the discount factor
The and prefactors come from two separate sources, and it is instructive to see both. The reward error is a per-step bias paid every step, so summing it under discounting costs the horizon factor . The transition error is more expensive: total-variation error is compared against value functions bounded by , and each step of the horizon multiplies the accumulated value bound by another factor, so the sum produces . The quadratic blow-up in the transition term is exactly the compounding effect, and it is the term that decides whether the bound is usable.
Now plug in an illustrative set of numbers. Take and . If rewards lie in , then returns lie in ; under the weaker assumption, returns lie in . Let the transition model have a uniform total-variation error upper bound of per state-action pair. This is a distance between distributions, not a claim that 95% of predicted transitions are correct. The transition-error term gives
With , the reward-error term adds 1, so the full bound is 991. Even the transition term alone exceeds a range-based upper bound of 201 on the value difference when true rewards satisfy and the learned reward may differ by up to . The bound is therefore vacuous in this example: a uniform 0.05 per-step TV bound does not yield a useful guarantee at this discount factor.
Plotting the two error terms against the discount factor shows which one drives the blow-up.
import numpy as np
# Bound terms from the worked example, as functions of the discount factor.
# Reward error is bounded by eps_r / (1 - gamma); transition error by
# 2 * gamma * R_max * eps_P / (1 - gamma) ** 2. The values fix eps_r = 0.01,
# eps_P = 0.05, and R_max = 1.0, matching the numbers discussed in the text.
_gamma_grid = np.linspace(0.5, 0.999, 400)
EPS_R = 0.01
EPS_P = 0.05
R_MAX = 1.0
reward_term = EPS_R / (1.0 - _gamma_grid)
transition_term = 2.0 * _gamma_grid * R_MAX * EPS_P / (1.0 - _gamma_grid) ** 2
The moral is not that the bound is bad. The bound is correct, and it is telling you something true: worst-case, there is no useful guarantee. A useful guarantee requires additional information, such as coverage of the policies under consideration or justified structure that bounds model error where they go. Pessimistic penalties and policy constraints can use such assumptions to limit unsupported choices; they do not themselves create coverage or certify a heuristic uncertainty score. This is why offline model-based RL is not simply "model-based RL without exploration."
It is worth emphasising that the numbers here are chosen to make the point sharp, not because is a realistic worst case. A uniform TV upper bound of 0.05 permits that much transition error at any state-action pair, including those the policy visits; it does not say actual error is 0.05 everywhere. The vacuity comes from pairing that global error allowance with a long horizon. Shorten the horizon () and the bound tightens quickly; make the model extremely accurate only where the policy goes and the uniform bound overstates the damage. Both of these escapes correspond to real assumptions we can make, and both are what the algorithms in the next sections exploit.
Both families exist, and they differ in what they learn.
- Model-free offline RL learns a -function or policy directly from , and handles distribution shift by constraining the policy or penalising the critic. Representative strictly offline methods include BCQ, BEAR, BRAC, TD3+BC, IQL, and CQL. AWAC can use an offline warm start but was designed for subsequent online improvement.
- Model-based offline RL learns predictive structure from and uses it for rollouts, planning, or critic training without new interaction. MOPO, COMBO, RAMBO, and MOBILE use learned dynamics in policy learning; MOReL plans in a pessimistic learned MDP, MBOP uses model-based planning, and Trajectory Transformer models joint state-action-reward sequences for reward-guided planning rather than fitting separate and networks.
A learned transition model can generate arbitrarily many synthetic transitions, including in regions the dataset never visited. Those samples do not create new evidence about the real environment. The design problem is to use rollouts where they help policy learning without letting unsupported predictions mislead the optimiser.
Why bother with a model at all, given the risk? Because of the amplification factor. With real transitions you can train a value function on samples, or you can train a model on samples and generate many imagined transitions from it. If the model is accurate on the region the policy cares about, those rollouts can help. MBPO made this argument in the online setting, as discussed in Part VIII: Decision-Centric Research Lineages. Online interaction can expose a wrong prediction and provide data for correction, though it does not guarantee safe exploration or eventual recovery. Offline, the fixed log cannot supply that targeted correction; repeated synthetic rollouts retain the model's unsupported assumptions.
The natural worry is a mismatch between where the model is accurate and where an improved policy might go. The dataset reflects , while an unconstrained improvement step can favor actions with little coverage. Improvement need not leave support: a mixed or suboptimal behaviour dataset may already contain better actions. The methods below try to use that opportunity without letting the optimiser exploit unsupported estimates.
Out-of-Distribution Actions and Model Uncertainty
This section defines what "out of distribution" means for a learned model and why the action dimension is usually where the trouble starts. It builds the machinery for quantifying epistemic uncertainty, then constructs a gridworld where the failure is visible and unambiguous.
What is actually out of distribution
A dataset of continuous-control trajectories often covers a thin tube in joint state-action space. Even millions of transitions may leave reachable pairs poorly sampled. A robot moving through a hallway traces a narrow ribbon inside a higher-dimensional joint space. Three useful views of extrapolation arise; they overlap rather than partition the space.
- State OOD. The query state lies outside, or in a poorly covered part of, the behaviour state's distribution. This can happen when the policy drives into a new configuration; exact equality with a logged continuous state is not the test.
- Action OOD. The state is familiar but the proposed action is poorly supported conditional on that state. This is subtler: driving logs may cover a speed well but contain little evidence about full braking at that speed.
- Joint OOD. The pair is missing or rare in the logged joint distribution, even when the state and action have each appeared separately. State OOD or action OOD also makes the pair OOD; the distinctive case is a novel pairing of familiar marginals.
Action OOD is the one people underestimate. Consider a behaviour policy that is nearly deterministic, as some demonstration datasets are. In a well-covered state, the marginal may be large, but conditional on the action distribution is a narrow spike. The model has evidence for only a narrow range of actions at that state, even if it has many transitions. If an improvement step proposes an action 2 standard deviations away from the spike, the transition model may have little evidence about the consequence, but the value function still returns a number, and the optimiser may still trust it.
These views interact. In a low-dimensional example, a state far from logged states may be visibly unusual; in high dimensions, that judgment is harder. A state-only density model cannot detect a poorly supported action at a familiar state. A familiar-looking state and a familiar-looking action can also form a rare pair, so checking only cannot establish joint state-action coverage. The relevant question for a transition model is whether the pair has enough nearby evidence.
In reconstructive latent world models from Part III: Representing Agents and Worlds, an encoder maps observations into a latent space and a decoder reconstructs observations. Other architectures, including MuZero, need not have a decoder. If planning drives a latent state far beyond training support, its predicted transitions and values can become unreliable even though the vectors remain numerically valid. A model may not raise an exception or produce a NaN when this happens, so checking only numerical validity will miss the coverage problem.
Two kinds of uncertainty are routinely conflated, and conflating them causes real damage.
- Aleatoric uncertainty is irreducible randomness in the environment. A dice roll, a slipping wheel, a stochastic reward channel. It does not shrink with more data, because there is nothing more to know.
- Epistemic uncertainty is uncertainty about the model itself, arising from limited relevant evidence. It can shrink with informative data near the query point and often rises away from observed regions, though approximate or misspecified models need not show either trend reliably.
For a penalty meant to guard against model extrapolation, epistemic uncertainty is the quantity of interest: irreducible randomness alone does not show that the model lacks evidence. A risk-sensitive or safety objective may deliberately penalise stochastic outcomes for a different reason. Practical scores can mix these components. MOPO's published scale-based heuristic is one example, and its penalty coefficient must be tuned rather than read as a calibrated error bound.
Consider a robot driving on ice. If friction varies unpredictably, more examples may clarify the distribution of slides without making each slide deterministic. Penalising that randomness could make the robot avoid a well-understood route. By contrast, if the robot has never driven on a particular gravel surface, the uncertainty reflects missing evidence about that surface. A naive predictive-variance score can mix these two cases.
Quantifying epistemic uncertainty
We will compare five approaches to estimating or bounding uncertainty. They differ in cost, calibration, and how badly they can fail.
A schematic decomposition makes the distinction concrete before the individual estimator families are described.
import numpy as np
# A schematic variance decomposition over a one-dimensional input. Data cover
# the interval [2, 8]; aleatoric noise is a constant 0.04 everywhere. The
# epistemic term grows with distance from the covered interval, so the total
# predictive variance is non-zero even where the environment is well sampled.
_x = np.linspace(0.0, 10.0, 401)
_data_low, _data_high = 2.0, 8.0
distance = np.where(
_x < _data_low,
_data_low - _x,
np.where(_x > _data_high, _x - _data_high, 0.0),
)
aleatoric = np.full_like(_x, 0.04)
epistemic = 0.5 * (1.0 - np.exp(-0.5 * distance**2))
predictive = epistemic + aleatoric
No single family dominates: the choice depends on whether you need speed, a defensible statistical statement, or robustness to shared model bias.
Ensembles and disagreement. Train dynamics models with different random initialisations and different bootstrapped data subsets. Each predicts a next-state mean . Then
where:
- : the disagreement-based uncertainty score at
- : the next-state mean predicted by ensemble member with parameters
- : the mean prediction averaged over the members
- : the Euclidean norm
This is one possible disagreement score, not MOPO's practical penalty. MOPO uses the maximum Frobenius norm of a member's predicted Gaussian scale matrix across its probabilistic model ensemble, which its authors describe as a maximum standard-deviation score. Ensemble disagreement can be useful, but it is not a calibrated posterior uncertainty; members with shared architecture and data may agree even when they are all wrong. Nor does predictive variance automatically isolate epistemic uncertainty from environmental randomness. The disagreement captures how much the members differ from each other, not how far the ensemble as a whole is from the truth. If all members are biased in the same way, the disagreement stays small while the error stays large.
Count-based uncertainty. In discrete or discretisable spaces, uncertainty is a function of the visit count . A common form is
where:
- : the uncertainty score at
- : the number of times was visited in the dataset
- : a smoothing prior (set to 1 in our implementation below)
The second has the scaling of many concentration bounds. It is not, without a specified prior and observation-noise variance, an exact posterior standard deviation. Count-based scores are cheap and interpretable when counts are meaningful; we use one below to isolate the penalty mechanism from neural-network training noise. Their weakness is that "count" requires a definition of "the same event," and in continuous state spaces that definition is a discretisation, which introduces its own choices.
Density and likelihood models. Fit with a normalising flow or a VAE, then use a quantity such as as a proxy for low data density. It is not automatically a calibrated bound on transition-model error or a reliable test of support. BEAR instead uses behavior-action samples and an MMD constraint in policy improvement. FisherBRC learns a behavior density for its critic decomposition and regularizes action gradients of the offset; neither method defines this density score as a world-model reward penalty. Likelihood is not the same as support: a generative model can give high likelihood to out-of-distribution inputs, as likelihood-based OOD studies demonstrate.
Bayesian neural networks and variational inference. Place a prior over model parameters and use a posterior predictive distribution. These methods introduce inference and calibration choices that differ from those of a bootstrap ensemble; no single family wins across tasks and model classes. PILCO's Gaussian-process dynamics, discussed in Part VIII: Decision-Centric Research Lineages, provide posterior uncertainty under the GP model assumptions. Naive exact GP fitting has cubic time cost in the number of training points, so that particular implementation does not scale cheaply to large datasets.
Spectral and Lipschitz-based bounds. A smooth learned model alone does not guarantee a small error outside the data. If the error function is known to be Lipschitz, or both true and learned dynamics have compatible known Lipschitz bounds, an error measured near a query point can be extended with a distance term. Such a result is deterministic under its assumptions, though its constants may be too loose to guide planning. The needed quantity is a bound on how quickly error can change, not merely on how quickly the learned prediction changes.
Ensembles are easy to get subtly wrong. Three details matter.
- Diversity source. Varying only the initialisation often produces models that agree too much. Bootstrapping the data (sample transitions with replacement per member) or using different data orderings helps.
- Noise handling. If each member samples an independent stochastic successor during a rollout, raw sample spread mixes environmental randomness with model disagreement. Compare predictions under a controlled noise coupling, or separate disagreement among predicted means from each model's predictive variance. The two components need different interpretations; MOPO's published practical penalty uses a norm of predicted Gaussian scale (standard deviation), not the mean-disagreement formula above.
- Ensemble size. A small ensemble is a common computational compromise, but there is no universal member count at which disagreement becomes calibrated or further models cease to help. Check size against held-out model error and compute cost in the task at hand.
A gridworld where the failure is visible
Now we build something concrete. The goal is not to reproduce a benchmark. It is to construct the smallest system in which we can see the mechanism with our own eyes, compute everything exactly, and check the answer against ground truth. Minimality matters here. Offline RL failures on real benchmarks are opaque: you see a poor return, but you cannot easily separate model error from coverage. In a small grid we can isolate those two factors directly.
The environment is a grid. Row 3 is a corridor that runs the full width; the goal sits at the far right end of that row. The grid edge acts as a wall. The goal is absorbing with zero reward, and reaching it pays 1.0. The discount factor is .
The behaviour policy is a random walk along row 3, moving right with probability 0.7 and left with probability 0.3. It never leaves the corridor. This is the crucial design choice: the dataset covers one row out of seven, and covers it densely. Every other cell in the grid has never been seen.
The leftward moves matter. With a purely rightward policy, each of the 59 pre-goal corridor cells would be visited once per episode, while the absorbing goal would still receive 20 logged tail visits. The 30% leftward action instead makes the walk revisit earlier cells, producing a non-uniform pre-goal visit distribution.
The left edge clips a left action back to the same cell, while the rightward bias eventually carries most episodes to the goal. Because episodes reset at the start and stop after a short tail of goal self-transitions, the logged counts are not a stationary distribution of an endless reflecting walk. They need not be uniform along the corridor; the heatmap below shows the actual finite-episode coverage.
The optimal policy is trivially "walk right 59 times," so the return from the start state is . (The reward is paid on arrival at the goal, after 59 steps, which is step index 58 from the start.) So we can compute every quantity exactly and ask whether each method finds it.
import numpy as np
# A 7 x 60 grid. Row 3 is the corridor that the behaviour policy walks along;
# every other cell is never visited. The goal sits at the far right of row 3.
H, W = 7, 60
N_STATES = H * W
N_ACTIONS = 4
CORRIDOR_ROW = 3
START = CORRIDOR_ROW * W
GOAL = CORRIDOR_ROW * W + (W - 1)
GAMMA = 0.99
DELTAS = ((-1, 0), (1, 0), (0, -1), (0, 1))
ACTION_NAMES = ("up", "down", "left", "right")
def true_step(state, action):
"""Ground-truth transition. The goal is absorbing with zero reward."""
if state == GOAL:
return GOAL, 0.0
row, col = divmod(state, W)
d_row, d_col = DELTAS[action]
row = min(max(row + d_row, 0), H - 1) # the grid edge acts as a wall
col = min(max(col + d_col, 0), W - 1)
nxt = row * W + col
return nxt, 1.0 if nxt == GOAL else 0.0Next we roll out the behaviour policy and record the transitions. Each episode runs until the goal is reached, followed by a short tail of goal self-transitions. The logged tail includes only left and right actions at the goal, so the empirical model learns those two self-loops; up and down remain unobserved and receive the same uniform fallback as other unsupported pairs. The true goal is absorbing under every action. This distinction matters: samples of some terminal actions do not tell a purely empirical model that all terminal actions must self-loop unless that structure is supplied explicitly.
def behaviour_action(rng):
"""A random walk confined to row 3: 70% right, 30% left."""
return 3 if rng.random() < 0.7 else 2
rng = np.random.default_rng(0)
transitions = []
for _ in range(300):
state = START
for _ in range(400):
action = behaviour_action(rng)
nxt, reward = true_step(state, action)
transitions.append((state, action, reward, nxt))
state = nxt
if state == GOAL:
break
assert state == GOAL, (
"This fixed-seed episode must reach the goal before its tail"
)
for _ in range(20): # a short tail of goal self-transitions
action = behaviour_action(rng)
nxt, reward = true_step(state, action)
transitions.append((state, action, reward, nxt))
state = nxt
dataset = np.asarray(transitions, dtype=np.int64)We now tabulate the data. counts records how often each state-action pair was taken, reward_sums accumulates the observed reward, and next_counts records how often each particular successor followed. Everything downstream is built from these three arrays: the empirical model is a normalisation of them, and the uncertainty scores are functions of counts.
counts = np.zeros((N_STATES, N_ACTIONS))
reward_sums = np.zeros((N_STATES, N_ACTIONS))
next_counts = np.zeros((N_STATES, N_ACTIONS, N_STATES))
for state, action, reward, nxt in dataset:
counts[state, action] += 1
reward_sums[state, action] += reward
next_counts[state, action, nxt] += 1transitions collected : 49,937 state-action pairs observed : 120 of 1,680 grid cells visited : 60 of 420
The printed counts show 120 of 1,680 state-action pairs observed, about 7.1%. The 60 visited cells are exactly the ones in row 3. Every other cell is unobserved, and every method below has to decide what to do there. The 7.1% figure is worth pausing on: it is not an extreme number, and yet it is enough to make a naive planner fail completely, as we are about to see.
Fitting the empirical model
The model is deliberately simple: a maximum-likelihood empirical transition model with a uniform fallback for unobserved pairs. This is the tabular analogue of a neural dynamics model with no regularisation, and it has the property we want to study: off-support, it returns something arbitrary, determined by our fallback choice rather than by evidence.
where:
- : the estimated transition probability from to
- : the estimated expected reward at
- : the number of times appears in the dataset
- : the number of times followed
- : the number of states, used as the uniform fallback denominator
The uniform fallback is a stand-in for extrapolation, not a claim that other approximators behave identically. A neural network may continue a learned trend; a random forest prediction reflects its selected leaves; a Gaussian-process posterior mean approaches its prior mean when covariance with the training inputs vanishes, as can happen far away under a decaying kernel. None is automatically correct off-support. The uniform choice makes the failure mode vivid: an unvisited transition can land anywhere in the grid, including directly on the goal.
uniform = np.full(N_STATES, 1.0 / N_STATES)
P_hat = np.tile(uniform, (N_STATES, N_ACTIONS, 1))
R_hat = np.zeros((N_STATES, N_ACTIONS))
seen = counts > 0
P_hat[seen] = next_counts[seen] / counts[seen][:, None]
R_hat[seen] = reward_sums[seen] / counts[seen]Counting the uncertainty
We use a count-based uncertainty proxy because the state-action visit counts are exact and interpretable in this tabular toy. In a larger continuous-state system, an ensemble is one common alternative; the count score isolates the penalty mechanism from neural-network training noise.
where:
- : the count-based uncertainty score
- : the visit count from the dataset
- : the prior pseudo-count, fixed at 1 here
If a pair has never been observed, . If it has been observed 999 times, . This is monotone, bounded in , and it shrinks at roughly in the well-sampled regime. The version has a different decay rate, closer to common concentration-bound scaling; here we use only as a simple bounded heuristic. Choosing a bounded score also keeps the penalty comparable across state-action pairs: no single pair can swamp the sum with an arbitrarily large uncertainty value.
# Count-based uncertainty proxy in (0, 1]: exactly 1 for an unobserved pair.
UNCERTAINTY_PRIOR = 1.0
u_hat = UNCERTAINTY_PRIOR / (UNCERTAINTY_PRIOR + counts)
visits = counts.sum(axis=1).reshape(H, W)
uncertainty_grid = u_hat[:, 3].reshape(
H, W
) # right action, sampled in the corridor

Two things to notice. Right-action uncertainty is maximal off the corridor, because those pairs were never observed, and lower where rightward actions appear in the logs. The up and down actions remain maximally uncertain even on the corridor; that distinction is in the underlying state-action counts, though this heatmap shows only the right-action slice. A familiar state does not make every action at that state well supported.
A third thing to notice is that cell counts are not uniform even within the corridor. Fifty of the 59 nonterminal corridor cells have between 700 and 800 visits in this seeded run, but the two cells immediately before the goal have only 600 and 428. The absorbing goal has 6000 visits because every one of the 300 episodes adds a tail of 20 goal self-transitions. That bright goal cell should not be mistaken for stronger coverage of the approach states, and cell counts alone still say nothing about unsupported actions at those states.
Solving the model without pessimism
Now we plan. The generic recipe is value iteration on the learned model:
where:
- : the estimated action-value at under the learned model
- : the estimated state value, the maximum of over actions
- : the estimated reward
- : the estimated transition probability
- : the discount factor
iterated to convergence. This is the same fixed-point computation that underlies every model-based planner in Part VII: Planning and Agency, with a learned model in place of the true one. The reason it is dangerous offline is precisely that value iteration trusts the model everywhere: the operator steps through all four actions at every state, including states the model has never seen, and it takes whatever action looks best according to estimates that may be arbitrary.
def value_iteration(P, R, gamma=GAMMA, tol=1e-10, max_sweeps=6000):
"""Value iteration for a finite MDP given as dense (S, A, S) and (S, A) arrays."""
V = np.zeros(P.shape[0])
for _ in range(max_sweeps):
Q = R + gamma * (P @ V) # (S, A, S) @ (S,) -> (S, A)
V_next = Q.max(axis=1)
if np.max(np.abs(V_next - V)) < tol:
V = V_next
break
V = V_next
return V, R + gamma * (P @ V)To judge the result we also need the ground truth: the true MDP tensor, a direct policy evaluator, and the true optimal value. The evaluator solves the linear system directly. The selected policies' true returns therefore come from a linear solve; planner scores and the optimal value come from value iteration to the stated tolerance. This is one of the luxuries of the gridworld setting: we can compare a method's belief about a policy against the policy's actual value in the true environment with negligible numerical error on the comparison.
P_true = np.zeros((N_STATES, N_ACTIONS, N_STATES))
R_true = np.zeros((N_STATES, N_ACTIONS))
for state in range(N_STATES):
for action in range(N_ACTIONS):
nxt, reward = true_step(state, action)
P_true[state, action, nxt] = 1.0
R_true[state, action] = reward
def evaluate_policy(P, R, policy, gamma=GAMMA):
"""Exact evaluation of a fixed deterministic policy in the supplied MDP."""
grid = np.arange(N_STATES)
P_pi = P[grid, policy]
R_pi = R[grid, policy]
return np.linalg.solve(np.eye(N_STATES) - gamma * P_pi, R_pi)
def evaluate_in_true_mdp(policy):
return evaluate_policy(P_true, R_true, policy)
V_star, _ = value_iteration(P_true, R_true)PENALTY = 0.5
V_naive, Q_naive = value_iteration(P_hat, R_hat)
V_pess, Q_pess = value_iteration(P_hat, R_hat - PENALTY * u_hat)
pi_naive = Q_naive.argmax(axis=1)
pi_pess = Q_pess.argmax(axis=1)
V_true_naive = evaluate_in_true_mdp(pi_naive)
V_true_pess = evaluate_in_true_mdp(pi_pess)
V_model_naive_policy = evaluate_policy(P_hat, R_hat, pi_naive)
V_model_pess_policy = evaluate_policy(P_hat, R_hat, pi_pess)naive model planner : planner score +3.082 | real -0.000 pessimistic model planner : planner score +0.509 | real +0.558 true optimal policy : | real +0.558
There it is. The unpenalised planner estimates a return of about 3.08 from the start state. Under the true dynamics its greedy policy earns zero. The model's unsupported transition fallback has created a profitable imagined route that the policy exploits.
The reason is worth tracing. The supported route rightward along the corridor reaches the goal in 59 steps, worth in the true MDP. Moving vertically enters unsupported pairs whose fitted transition distribution is uniform over all 420 cells. Some of those random successors lie near the goal and have observed, rewarding actions. At the true absorbing goal, the logged left and right actions self-loop, but the unlogged up and down actions also get the model's uniform fallback, so imagined exits from the goal add another optimistic path. Repeated imagined opportunities make the unsupported route look better than the corridor, even though true vertical moves never teleport. Landing directly on the true goal earns no immediate reward in this implementation; the learned model can nevertheless imagine future reward after an unsupported goal action.
Once the policy steps into an unobserved cell, all four actions have identical estimated value, so the argmax picks the first one, "up." In the true grid it moves from through and to . Further up actions are clipped by the top boundary, leaving it at forever. It never reaches the goal. The value map below shows the same failure structurally: the unpenalised value function is flat across the unvisited region, a plateau of imagined reward.
The flatness is the tell. In the true grid, a policy that moves toward the goal would have values that vary with distance. Instead, the unpenalised model assigns roughly the same value to every unsampled cell because it uses the same uniform successor distribution for each unsupported pair. This plateau exposes the uniform extrapolation built into the toy model; other model errors need not produce the same visual pattern.
def greedy_trajectory(policy, steps=150):
state = START
path = [state]
for _ in range(steps):
state, _ = true_step(state, policy[state])
path.append(state)
return np.array(path)
path_naive = greedy_trajectory(pi_naive)
path_pess = greedy_trajectory(pi_pess)
rows_naive, cols_naive = np.divmod(path_naive, W)
rows_pess, cols_pess = np.divmod(path_pess, W)



The two trajectories show the mechanism. Without a penalty, the model's ignorance is read as opportunity and the policy moves off the data. With a penalty, the same model, the same dataset, and the same value iteration produce the optimal policy in this toy. A penalty of 0.5 per unobserved transition is enough to change the selected policy from zero return to the full optimum.
Pessimism and Conservative Objectives
The previous section diagnosed the disease: an unconstrained estimate is maximised by the optimiser, and off-support the estimate is arbitrary. This section describes the cure. The idea is almost embarrassingly simple, but its theoretical consequences are substantial, and the design choices around it are where most of the engineering lives.
The pessimistic principle
Let be a set of MDPs consistent with the data, in the sense that no MDP in the set is ruled out by the observed transitions. Offline, the data cannot distinguish them, so the honest thing to do is to consider the worst case within the set:
This is pessimism. It says: if the true environment lies in the candidate set, I will plan for the candidate that treats me worst. Under that condition, the pessimistic value is a lower bound on the true value of a fixed candidate policy; how informative that bound is depends on the set. The corresponding minimax formulation has the agent choose a policy, "nature" choose an MDP from the specified set, and the agent receive the worst-case value. This is a planning objective, not a guarantee of risk-free outcomes.
In practice is never enumerated. Instead, the worst case is approximated by subtracting a penalty from the reward:
where is the uncertainty function from the previous section and is the pessimism coefficient. The penalty is called "conservative" because it reduces the appeal of uncertain regions relative to their unpenalised estimates. A high predicted reward can still outweigh a high penalty; this objective does not impose a confidence threshold. The intuition is that the penalty proxies the gap between and the worst-case reward over , and where uncertainty is large that proxy can be large.
The contrast with the classical exploration literature is sharp. In a bandit or an online MDP, a well-known family of algorithms is optimistic: add a bonus to the estimated reward, , and the agent is driven toward the unfamiliar because the bonus makes it look attractive. The point is to gather information.
During offline training, an optimistic bonus can steer planning toward poorly supported actions without collecting any new evidence. This motivates pessimistic penalties in methods such as the one illustrated here. The policy must be chosen using the fixed dataset, not the prospect of gathering more information after a risky action.
The penalty must dominate the model error
The penalty is not automatically a guarantee. To obtain a lower bound, it must cover the model's possible overestimation at every relevant state-action pair. Here is a sufficient condition for a fixed deterministic policy.
Suppose the true and learned Bellman operators differ by at most when applied to the true value function:
The contraction property then gives . More locally, let the penalised model Bellman operator be . Its value, averaged over the initial distribution, is
For a pointwise lower bound, a sufficient condition is
Under this condition, . Monotonicity and contraction of the discounted Bellman operator imply
The pessimistic estimate is a certified lower bound only when the stated condition holds. Since and the true Bellman error are unknown offline, a count or ensemble score is not automatically a valid bound on that error. Policy-improvement guarantees require additional coverage, estimation, and optimisation assumptions; the displayed inequality alone does not give a suboptimality rate.
Two things follow. First, a penalty that is too small to cover overestimation lacks this lower-bound guarantee, whereas one that is too large can suppress useful policies. Second, a uniform penalty calibrated to a global worst case taxes even well-understood steps. Here is the deep problem: the condition depends on , which you do not have. A generic bound is ; tighter local bounds require additional assumptions. Using only the generic bound can produce the vacuous guarantee from the worked example earlier. More selective methods restrict where pessimism enters, and the rest of this section compares those choices.
The algorithm families
MOPO. Model-based Offline Policy Optimization learns an ensemble of probabilistic dynamics models. Its practical uncertainty score takes the largest Frobenius norm of a member's predicted Gaussian scale matrix (described in the paper as maximum standard deviation), ; it subtracts from reward and trains a policy with MBPO-style short model rollouts. That score is heuristic rather than the admissible model-error upper bound required by the paper's theorem. The theoretical comparison also uses the true reward function in its penalized model; the practical implementation learns a reward model. Under the theorem's function-class, exact-reward, and admissible-dynamics-error assumptions, the result compares the learned policy with any candidate policy by
where is the discounted sum of under policy in the learned model. This is a conditional guarantee, not a promise that the practical Gaussian scale (standard-deviation) heuristic upper-bounds real model error.
The gap between the theorem and the algorithm is instructive. The theorem assumes an admissible bound on relevant dynamics error and an exact reward; the practical algorithm estimates a Gaussian scale score and a reward model from data. That heuristic may work empirically, but it does not inherit the theorem's guarantee without additional calibration and reward-error assumptions. Read the proof as a statement about what the penalty would need to control, not a certification of the fitted ensemble.
MOReL. The Model-based Offline Reinforcement Learning algorithm of Kidambi and colleagues takes a harder line. It partitions the state-action space into known and unknown using ensemble disagreement, and then constructs a pessimistic MDP in which unknown pairs lead to an absorbing terminal state carrying a penalty:
The halt state is absorbing and carries the negative reward at every subsequent step; known pairs use the fitted transition. Thus its value under the displayed convention is , not a one-time . This hard unknown-pair rule prevents a model rollout from inventing a favorable continuation after the halt, whereas a soft penalty allows it to continue. MOReL's algorithm optionally uses a behavior-cloned policy to initialize its planner; that is not a required behavior-cloning regularizer in its objective.
COMBO. Conservative Offline Model-Based Policy Optimization learns a Q-function from a mixture of real transitions and model rollouts. Its conservative term pushes down Q-values on state-action pairs sampled under model rollouts and up Q-values on logged state-action pairs. Thus it does not apply pessimism only to real data. Unlike MOPO's practical reward penalty, this regularizer acts on the value function and does not require an explicit model-uncertainty estimate.
The COMBO idea is subtly different from the others. It does not need to know where the model is uncertain; it can simply compare where the Q-function is being trained on real data versus rollout data, and it penalises the difference. This makes COMBO less dependent on the quality of an uncertainty estimator, but more dependent on having a good critic architecture and optimiser.
RAMBO. Instead of subtracting a hand-designed uncertainty score from reward, RAMBO (Rigter and colleagues) trains the model adversarially against the value function while also fitting the data. Its model objective still has a coefficient controlling the adversarial term, so pessimism has not become hyperparameter-free. It is a value-aware model-learning approach rather than an uncertainty-penalty approach. The trade-off is that the model is trained both for data fit and for appropriately pessimistic values, which adds training complexity.
The model-free cousins, briefly
It is worth knowing that the same ideas appear on the model-free side, often with different vocabulary.
- BCQ (Batch-Constrained Q-learning) trains a generative model over the dataset's actions and restricts the policy to perturbing samples from it. The constraint is on the support of the action distribution.
- BEAR applies an MMD constraint between the learned policy's action distribution and the behaviour policy's during actor policy improvement; it does not add MMD to the critic's objective.
- TD3+BC combines Q-maximisation with a plain action-regression term in the actor objective: , with . Thus scales the Q term through , not the regression term. It is startlingly simple and competitive.
- AWAC and AWR perform advantage-weighted behaviour cloning: the policy update is a weighted maximum-likelihood fit where the weight is . Actions with low advantage get near-zero weight, so the actor update moves toward better logged actions. That statement does not guarantee that every critic backup, including AWAC's, avoids querying actions outside the dataset.
- IQL removes the OOD action query entirely by estimating the value function with expectile regression back-ups, which never need . It is arguably the cleanest statement of the underlying principle: do not ask the model a question the data cannot answer.
All of these address the same distribution-shift problem from different sides: pessimism discounts uncertain estimates, while a policy constraint limits which estimates are queried. They are complementary design choices, not generally a formal dual pair. The model-free family does not require a dynamics model, yet it still faces coverage problems when a critic is asked to value unsupported actions.
Implementing the penalty
The implementation is one line. Everything visible in the previous section came from adding - PENALTY * u_hat to the reward.
The choice of is the interesting part. Below some threshold, the penalty is too weak to overcome the model's optimism and the planner still walks off the data. Above the sampled threshold in this toy grid, the greedy policy stays on the corridor and reaches the goal; this is not a general safety guarantee. Push it much higher and you do not get a better policy, only a more pessimistic value estimate, which matters when you use that estimate to choose between candidate policies rather than to act. This is why tuning is subtler than it first appears: the optimal for shaping a policy may be different from the optimal for evaluating one.
To see the threshold, sweep the positive values of across about 3.3 orders of magnitude (with zero as an additional baseline), solve the penalised model at each value, extract the greedy policy, and evaluate that policy exactly in the true grid.
PENALTY_GRID = np.array(
[0.0, 0.005, 0.01, 0.02, 0.03, 0.05, 0.1, 0.5, 2.0, 10.0]
)
sweep_planner_score = np.zeros(len(PENALTY_GRID))
sweep_true = np.zeros(len(PENALTY_GRID))
for index, penalty in enumerate(PENALTY_GRID):
V_p, Q_p = value_iteration(P_hat, R_hat - penalty * u_hat)
sweep_planner_score[index] = V_p[START]
sweep_true[index] = evaluate_in_true_mdp(Q_p.argmax(axis=1))[START]
At the realised return is zero and the unpenalised model score is comfortably positive: the model is confidently wrong. Once the penalty makes the observed corridor more attractive than the imagined route, the realised return reaches the optimum and stays there. The penalised planner score keeps falling because model trajectories pay an additional tax. At the right of the plot that planning objective is negative while the policy remains optimal; it is not an unpenalised off-policy estimate of that policy's return.
That divergence between the penalised score and the policy is one warning sign of over-conservatism. It may not affect action quality in this grid, but it matters if penalised scores are compared across candidate policies or penalty settings. The lesson is not "pick the largest that keeps the policy optimal." The lesson is to distinguish a planning objective from a policy-value estimate and choose the penalty for the decision it must support.

Over-conservatism and how it shows up
Over-conservatism deserves more than a sentence, because it is the failure mode practitioners hit most often and it does not look like failure.
In the clean gridworld above, over-conservatism is invisible in the realised return. The corridor is the only route that pays anything, so increasing the penalty leaves nonterminal behavior and realised return unchanged in this example. In a realistic environment the pattern can differ. A dataset might cover a mediocre but safe region densely and a rewarding but thinly covered region sparsely. As grows, the latter pays a larger penalty even if the model's uncertainty score is bounded. The agent can retreat to the safe region and produce a stable-looking policy that misses attainable reward.
The mechanism is a tax that rises as logged coverage falls, regardless of the region's actual reward. Well-sampled regions pay almost nothing; thinly sampled regions pay much more. This biases the policy toward the well-sampled region even when the well-sampled region is not where the reward lives. The result is not a broken policy. It is a policy that looks reasonable, that runs cleanly, and that quietly underperforms.
There are three symptoms worth watching for:
- Return collapse relative to behaviour cloning. A policy that fails to beat a behaviour-cloning baseline may be over-constrained or over-penalised, though weak coverage, model error, or optimisation failure can produce the same symptom. The baseline is a useful diagnostic, not proof of the cause.
- Sensitivity of the chosen policy to across the middle of the range. If the argmax policy changes shape when moves from 0.3 to 0.5, the penalty is not just taxing the estimate, it is deciding the policy. That is fragile.
- A large gap between penalised and unpenalised model scores. In this fixed model, the optimal penalised score cannot increase as rises. A negative penalised score shows that the objective is heavily taxed; it does not imply the task's true return is negative or that the model believes the task is worthless.
The remedy depends on the cause. Reducing may recover a good policy but can also reintroduce model exploitation. Better data about the promising region, where collection is allowed, can reduce uncertainty; longer training on the same unsupported data cannot manufacture that missing evidence. A better-calibrated local penalty or constraint can distinguish well-supported and thinly supported regions. MOPO's penalty varies with ; MOReL halts on pairs classified as unknown. Neither automatically calibrates uncertainty to true model error.
An optimiser can seek out overestimated actions, so upper-tail value errors are especially dangerous. Underestimation also matters: it can discard a good policy or prevent improvement beyond the behaviour policy. Conservative methods trade some of that opportunity for protection against unsupported optimistic choices. Whether the trade is appropriate depends on coverage and the decision at hand.
Policy Constraints and Offline Evaluation
Pessimism and constraints are two responses to distribution shift, often used together rather than as formal duals. This section covers the constraint family, then turns to a separate problem: how can you judge a learned policy without running it?
Support constraints as an alternative to pessimism
Pessimism reduces estimated values where a penalty signals model risk, seeking to discourage unsupported choices. A support constraint instead restricts the policies considered. Neither automatically ensures that the resulting policy stays near the data: a heuristic penalty can miss model error, and an approximate constraint can admit poorly covered actions. They act on different parts of the optimization problem, and their effectiveness depends on their assumptions and implementation.
A support constraint restricts to the set of policies whose state-action distribution is covered by the dataset:
This is an idealized constraint on the full state-action occupancy, including states reached by the target policy. BCQ's conditional generative model and bounded action perturbation approximate the behavior-action support at queried states; they do not enforce this global occupancy condition, especially at unseen states. Under a continuous-density behaviour distribution, each individual sampled point has probability zero even when the distribution has full-dimensional support. That conditional support is hard to infer from a finite log, so practical methods use divergences or generative action models as approximations.
A behaviour constraint is softer: rather than requiring full support, it penalises divergence from the behaviour policy:
where denotes a discrepancy or penalty, such as KL, MMD, or squared action distance. BEAR constrains its actor with MMD, while TD3+BC uses an action-regression term. AWAC instead performs advantage-weighted behaviour cloning; its update is not literally this displayed divergence objective. The choice matters: KL is asymmetric, MMD compares distributions through a kernel, and squared distance ignores much of the action distribution's structure.
These mechanisms can be complementary, but no general equivalence follows. A support constraint may still allow optimistic errors within its permitted set; a pessimistic critic may still be poorly calibrated where data are scarce. CQL applies conservatism to critic values, MOPO penalises rewards in model rollouts, and MOReL changes unknown transitions to lead to HALT. BEAR and TD3+BC illustrate direct actor restrictions. Specific implementations may combine mechanisms, but the papers do not establish a universal constraint–penalty duality.
An exact support constraint is binary, but practical approximations usually use a divergence, a coefficient , or a generative action model. A very strong restriction can leave little room for policy improvement; a weak one can admit unsupported actions and critic overestimation. This resembles the conservatism trade-off in a penalty sweep, although constraint and penalty coefficients have different meanings and need not produce the same policies.
Evaluation is the unsolved part
Here is the part of offline RL that practitioners consistently underestimate. You have produced a policy. You cannot run it. You must decide whether to deploy it, and if you have several candidates, which one to pick.
The candidates are usually not just policies. They are (policy, model architecture, ensemble size, rollout length, penalty coefficient, learning rate) tuples, and the number of combinations can be large. Selecting among them is offline model selection, a separate problem that can be difficult with only logged data. A model-based estimate of a policy's value can inherit errors from the model used for learning. If that model overestimates a region, it may also overestimate policies that visit the region and favor them in selection. This is a related form of model exploitation on the evaluation side.
Three families of estimators exist.
Importance sampling and its variants. Given independent trajectories from a known behaviour policy , reweight each discounted trajectory return by the ratio of action probabilities:
Under matching initial-state and environment distributions, known behaviour propensities, target-policy support within behaviour support, and finite expectation, ordinary trajectory IS is unbiased for the return represented by each complete sampled trajectory. For the infinite-horizon defined above, the recorded trajectory must include all nonzero rewards (for example, by ending in a zero-reward absorbing state) or use a valid tail correction; otherwise this formula estimates a truncated return. Its variance can grow exponentially with horizon when the policies differ enough, but this is not inevitable. Per-decision IS (PDIS) applies prefix ratios to individual rewards and is unbiased for the same return target under the corresponding support, sampling, and trajectory-completeness conditions; weighted/self-normalised IS trades finite-sample bias for possible variance reduction. Doubly robust estimators combine an approximate value function with importance weights. Long horizons and continuous actions can make ratio-based estimators impractical when overlap is weak, but their suitability depends on the policies and data rather than horizon alone.
Fitted Q-evaluation. Learn by iterating fitted Bellman back-ups on the offline dataset, using actions from at the successor state rather than an argmax. The resulting value estimate depends on approximation, optimisation, and evaluation-policy coverage; a policy produced by an offline algorithm is not by construction covered by the data. FQE avoids multiplying whole-trajectory probability ratios, but fitting a critic can be computationally more expensive than computing IS weights. It is a useful OPE baseline when its assumptions and implementation are stated, not an automatic certificate for every D4RL or RL Unplugged task.
Model-based evaluation. Evaluate a fixed candidate policy in the learned model. Holding the policy fixed removes an additional maximisation within that evaluation calculation, but it does not erase how the policy was selected. A policy optimised on the same model may already exploit its weak spots, and choosing the highest of many model-based estimates can do so again. Model-based evaluation is useful when model accuracy is supported along that policy's occupancy, with uncertainty and coverage reported alongside the estimate.
It also has an important failure mode, and it is the same one. The evaluator shares the model's blind spots. If the model has a large optimistic region and the candidate policy goes there, the evaluator reports a high return for exactly the policy you should not deploy. Techniques that help include averaging over the ensemble, reporting the worst-case estimate across members, and using the uncertainty map to flag evaluations whose trajectories enter high- regions.
# Unpenalised model-based evaluation of each fixed policy, distinct from the
# penalised objective used to select the pessimistic policy.
mb_estimate_naive = V_model_naive_policy[START]
mb_estimate_pess = V_model_pess_policy[START]
pessimistic_planner_score = V_pess[START]
# The fraction of each candidate's model rollout that visits unobserved states,
# computed from the occupancy implied by rolling the policy out in the model.
def model_occupancy(policy, sweeps=4000, tolerance=1e-12):
occupancy = np.zeros(N_STATES)
state_distribution = np.zeros(N_STATES)
state_distribution[START] = 1.0
for _ in range(sweeps):
occupancy += state_distribution
state_distribution = GAMMA * (
P_hat[np.arange(N_STATES), policy].T @ state_distribution
)
if state_distribution.sum() < tolerance:
break
return occupancy / occupancy.sum()
occupancy_naive = model_occupancy(pi_naive)
occupancy_pess = model_occupancy(pi_pess)
unsupported_naive = float(occupancy_naive[(counts.sum(axis=1) == 0)].sum())
unsupported_pess = float(occupancy_pess[(counts.sum(axis=1) == 0)].sum())model-based evaluation naive candidate : estimated +3.082 pessimistic candidate : estimated +0.558 pessimistic planner score (with penalty): +0.509 discounted occupancy of unvisited states under each candidate naive candidate : 45.2% pessimistic candidate : 0.0%
The discounted unvisited-state occupancy diagnostic shows that 45.2% of the naive candidate's model-rollout mass lies in cells absent from the dataset, versus 0.0% for the pessimistic candidate. The unpenalised learned-model evaluation of the pessimistic policy is about 0.558, while its penalised planner score is about 0.509. The naive policy's model estimate relies partly on unobserved dynamics and should not be read as a guarantee about its real return.
This is the practical lesson for offline evaluation: pair estimates with coverage diagnostics and sensitivity checks. This state-only fraction flags the naive policy's exposure to unvisited cells, but misses unsupported actions within visited states. Inspect discounted state-action occupancy or logged pair counts as well; neither check by itself certifies model accuracy or the value estimate.
Benchmarks and protocols
The standard offline RL benchmarks are:
- D4RL. The benchmark includes MuJoCo locomotion, AntMaze, Franka kitchen, and Adroit. Locomotion datasets use labels such as
random,medium,medium-replay,medium-expert, andexpert; those labels do not apply uniformly to the other task families. Medium-quality locomotion data provide a useful improvement-from-imperfect-behaviour test, but no one quality label is universally the hardest across tasks and methods. - RL Unplugged. A suite spanning control and games, including continuous-control and Atari-style settings, with an emphasis on realistic logged data collection.
- NeoRL. Datasets collected by policies of controlled quality, designed for studying near-realistic offline settings where the behaviour policy is neither random nor expert.
Two protocol points matter more than the choice of benchmark. First, report the benchmark's prescribed score normalization where it is defined; D4RL commonly normalizes scores against reference random and expert returns so raw returns across tasks are more comparable. Second, report the variation across seeds and the aggregation statistic. Mean and median can differ substantially, so a ranking that changes with aggregation deserves scrutiny.
For rollout-based methods such as MOPO and COMBO, one common protocol is to fix the dataset, fit a model per seed, generate short model rollouts, train a policy or critic using real and imagined data, and choose the final policy with an offline criterion. MOReL instead plans in its pessimistic learned MDP, and Trajectory Transformer plans over predicted sequences; the rollout-mixture recipe does not describe every model-based method. In rollout methods, horizon is a consequential hyperparameter: a short rollout limits direct propagation of model error but leaves more work to the critic, while a long rollout can accumulate unsupported predictions. Policy selection needs its own reported criterion because no online return is available to settle competing settings.
Limitations and Impact
Pessimism can reduce an optimiser's incentive to exploit model error where the penalty captures the relevant risk. The gridworld made the mechanism visible: the same fitted model, dataset, and value iteration produced a zero-return policy without a penalty and an optimal policy with a reward penalty of half a unit per under-supported transition. That is an instructive example, not a universal guarantee that pessimism resolves every ill-posed offline decision problem.
Pessimism also does not manufacture information about unvisited actions. If the log contains no evidence about an action and no justified structural assumption links it to observed actions, two environments can agree on every logged transition yet require different decisions there. That rules out a uniform worst-case guarantee from the log alone. Known physics, smoothness, or transferable pretraining may provide additional information, but their assumptions and target-domain validity must be made explicit. A warehouse fleet that has only driven down the middle of the aisles has no logged outcomes from the edges. A conservative policy may avoid an edge route that is actually useful; its caution is a response to uncertainty, not proof the route is bad. Where feasible, collect relevant data; otherwise state what prior structure supports any extrapolation and how it was checked.
The practical costs are real. Pessimism introduces at least one task-dependent hyperparameter. In our gridworld policy quality changes abruptly across a threshold in ; real problems can change less cleanly, and penalties interact with rollout length, model capacity, ensemble design, and reward scale. Training multiple models increases modelling cost. Because the dataset is fixed, additional fitting or ensemble members cannot by themselves supply transition evidence in an unsupported region; new data would be required to resolve that uncertainty empirically.
There is also a subtler limitation: ensemble disagreement need not reflect shared model bias. Members trained on the same data and similar architectures may agree on a wrong extrapolation, leaving a small disagreement score where the actual transition error is large. Adding similar members does not reliably resolve that missing-data problem. Broader pretraining may help in some settings, a direction discussed in Part IX: Foundation and World-Action Models, but its uncertainty must still be checked against the target domain.
Under-dispersion limits a pessimistic method that relies on that uncertainty score, not offline RL as a whole. If the score is systematically small where model error matters, the penalty is small there too, so it cannot be treated as a certified error bound. More varied prior data or model assumptions might help, but neither certifies calibrated uncertainty without target-domain checks.
The offline setting forces coverage, model validity, and estimation bias into the open because there is no new interaction to check a candidate policy during learning. Online feedback can sometimes expose errors, but it does not automatically rescue a poor policy or make exploration safe. The same discipline matters beyond offline RL: in the evaluation protocols of Part XI: Evaluation and Understanding, in the safety arguments of Part XII: Reliable World Models, and wherever a learned model informs a decision before its consequences can be checked.
Summary
During its offline training phase, offline RL learns from a fixed dataset without further interaction. The removal of the training-time feedback loop creates two coupled problems: the log supplies no direct target where it is silent, and the optimiser can favour those places when unsupported model estimates are optimistic. Justified structural assumptions may narrow the uncertainty, but the log alone cannot test those predictions there.
- Coverage or justified structure is fundamental. Without sufficient logged coverage or informative assumptions that constrain unobserved outcomes, uniform worst-case improvement is impossible. A finite concentrability coefficient formalizes one form of coverage; different algorithms use different structural and statistical assumptions.
- The simulation lemma explains fixed-policy error. The value gap is an average of local errors weighted by the policy's learned-model rollout occupancy, which can differ from its true occupancy. Policy optimization can then favor optimistic model estimates and change where that learned-model occupancy lies. Error propagation and policy selection can reinforce each other, but the fixed-policy identity alone does not prove exploitation.
- Worst-case bounds can be vacuous. In the illustrative example with and a uniform 0.05 transition-TV bound, the transition-error term is 990 and the full bound, including , is 991. Both exceed a range-based upper bound on the possible value difference under the stated reward range. Useful guarantees need stronger information about where the policy goes and where the model is accurate.
- Uncertainty scores need interpretation. Counts and ensemble disagreement can flag missing evidence, but predictive variance may also include irreducible randomness. Shared model biases can make an ensemble agree while all its members are wrong.
- Pessimism can lower-bound value under conditions. A penalty must cover possible model overestimation at relevant pairs; a heuristic score alone is not a certificate. Too little penalty permits exploitation, while too much can discard useful policies.
- Constraints are complementary. Support constraints and behaviour regularisers restrict where a policy may go; pessimism reduces the estimated worth of uncertain choices. Some methods combine these mechanisms, but they are not generally formal duals.
- Evaluation remains difficult. Trajectory and per-decision importance sampling can be unbiased under support, sampling, and complete-return or valid tail-correction conditions, but may have severe variance under policy mismatch. Weighted IS can introduce finite-sample bias; fitted Q-evaluation and model rollouts introduce approximation or model bias. Coverage diagnostics help identify risk but do not certify an estimate.
- Model selection is a separate problem. Choosing among candidate policies and hyperparameters using offline data only needs its own evaluation and uncertainty checks; a favorable learned-model estimate is not enough.
The chapter's gridworld compressed the argument into a few hundred lines of NumPy: one row of a seven-row grid, one behaviour policy that never left it, and an additive reward penalty that changed the selected policy from zero return to the true optimum. The general mechanism is model exploitation: an optimiser can favor actions where estimates are unsupported. Beyond the toy grid, pessimism helps only to the extent that its uncertainty score tracks the relevant errors.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about offline and conservative model-based reinforcement learning.
Offline and Conservative Model-Based RL
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!