Offline and Conservative Model-Based RL

Michael BrenndoerferJuly 18, 202675 min read

Part of World Models Handbook

Explains how fixed datasets limit offline RL, why model exploitation can mislead planners, and how pessimism and policy constraints reduce unsupported choices.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Offline and Conservative Model-Based RL

Picture a warehouse robot that has been running for three years. It has driven several thousand kilometres, followed human-designed routes, bumped into the same four pallets, and logged every state, action, and outcome. You now have a large dataset. You have no way to collect more. Can those logs support a better controller than the one that produced them?

The answer is not obvious and not always yes. The dataset was produced by some behaviour policy, perhaps a careful human operator or an older controller. Every transition reflects choices that policy made. If it never tried to squeeze between two pallets, the dataset contains no direct evidence about what happens when you try. If it never braked hard on a wet floor, the dataset is silent on that too. Improving the controller may require different actions, and those actions may lack coverage in the logs. That is the tension we spend the chapter unpicking.

Every model-based method we have covered so far in Part VIII: Decision-Centric Research Lineages assumed a closed loop. PETS, MBPO, Dreamer, MuZero, and TD-MPC all interleave acting in the world with learning from the consequences. The world model is wrong in a hundred small ways, the agent tries something, reality corrects it, and the model improves. That correction signal is the engine.

When the agent acts, the observed successor provides a targeted new sample at the state-action pair it tried. In a stochastic environment, one outcome is not a measurement of how wrong the model's predictive distribution is; repeated observations help distinguish environmental randomness from systematic model error. Over many interactions a well-designed learning loop can concentrate model corrections where the agent visits. The state distribution induced by the policy and the region where the model improves can then move together. Offline learning loses that feedback loop.

Offline reinforcement learning removes that engine during training. The agent gets one fixed dataset and no new interaction while learning its policy. Deployment may provide feedback later, or an offline-to-online phase may collect more data, but neither can correct an unsupported decision before this offline training stage ends. The policy must therefore be selected using evidence already in the log.

That single change breaks almost everything. Two failure modes follow directly from it, and they are worth stating plainly before any mathematics:

  • The log alone does not constrain the model off its support. If your dataset never contains the state-action pair (s,a)(s, a), competing predictions there can fit the logged transitions equally well. A model that says "you teleport to the goal" and a model that says "you fall into a pit" may both fit the training set. Architecture, prior knowledge, or justified structural assumptions can distinguish them, but the log provides no direct target at that pair to pull the model toward the truth.
  • The optimiser can exploit unsupported optimistic estimates. A policy improvement step that maximises estimated value favours regions with high estimates. If the model happens to be optimistic off-support, that step can select a policy that visits a region where the log provides little evidence. This is model exploitation, and it is not a defect of one algorithm: it is a risk whenever policy selection relies on estimates that are insufficiently grounded.

The mitigations address those two risks. We can quantify how uncertain the model is (ensembles, counts, densities). We can be pessimistic, subtracting an uncertainty penalty to make poorly supported choices less attractive. Or we can constrain the policy itself to stay near the behaviour that produced the data. The rest of the chapter develops all three ideas, illustrates uncertainty and pessimism on a gridworld, and then turns to the meta-problem that offline RL never escapes: how do you evaluate a policy you cannot run?

Learning from Fixed Datasets

This section fixes notation, states a coverage assumption that can support offline learning, and derives the simulation lemma, which is the single most useful equation in the chapter. It explains why a model that looks accurate can produce a badly wrong value function, and it draws the line between model-based and model-free approaches to the offline problem. Later we will distinguish coverage from justified structural assumptions that can also constrain unobserved outcomes.

The setup

We have a Markov decision process M=(S,A,P,r,γ,d0)M = (\mathcal{S}, \mathcal{A}, P, r, \gamma, d_0) as introduced in Part II: Inference, Decisions, and Control. To keep the comparison with learned models clean, we will assume throughout that the state is fully observed and the MDP is stationary. Neither assumption is free. Under partial observability the effective state is a belief, and the coverage problem becomes harder because two observation histories can be indistinguishable while requiring different actions. We will flag where that matters.

The stationarity assumption matters too. If the environment drifts over time, last year's transitions may not describe the world the deployed policy will face. In the warehouse example, a reorganised floor plan changes the dynamics represented by the logs. Here we assume no such drift: training and deployment share the same environment.

Instead of interacting, we are handed a dataset

D={(si,ai,ri,si′)}i=1N,(si,ai)∼ρβ,E[ri∣si,ai]=r(si,ai),si′∼P(⋅∣si,ai),\mathcal{D} = \{(s_i, a_i, r_i, s_i')\}_{i=1}^{N}, \qquad (s_i, a_i) \sim \rho_\beta, \quad \mathbb{E}[r_i\mid s_i,a_i]=r(s_i,a_i), \quad s_i' \sim P(\cdot \mid s_i, a_i),

where:

  • D\mathcal{D}: the fixed dataset of NN transitions
  • sis_i: the ii-th observed state
  • aia_i: the ii-th observed action taken in that state
  • rir_i: the observed reward, whose conditional mean is r(si,ai)r(s_i,a_i)
  • si′s_i': the next state, drawn from the transition distribution P(⋅∣si,ai)P(\cdot\mid s_i, a_i)
  • ρβ\rho_\beta: the behaviour distribution, the state-action distribution induced by whatever process generated the logs

The behaviour distribution might come from a human demonstrator, a hand-tuned controller, a previously deployed policy, or five such policies mixed together. We do not need to know which. We only need the samples. Notice that the dataset records actions and their outcomes but never explains them. Nothing tells us why the behaviour policy chose a given action, which matters later when we want to imitate it in states where the model is uncertain.

The objective is unchanged from online RL. We want a policy π\pi maximising

J(π)=E[∑t=0∞γtr(st,at)  |  s0∼d0,  at∼π(⋅∣st),  st+1∼P(⋅∣st,at)].J(\pi) = \mathbb{E}\left[\sum_{t=0}^{\infty} \gamma^t r(s_t, a_t) \;\middle|\; s_0 \sim d_0,\; a_t \sim \pi(\cdot\mid s_t),\; s_{t+1}\sim P(\cdot\mid s_t, a_t)\right].

where:

  • J(π)J(\pi): the expected discounted return of policy π\pi
  • γ\gamma: the discount factor in [0,1)[0,1)
  • r(st,at)r(s_t, a_t): the reward received at step tt
  • d0d_0: the initial-state distribution, with s0∼d0s_0\sim d_0
  • π(⋅∣st)\pi(\cdot\mid s_t): the policy, giving the action distribution at state sts_t
  • P(⋅∣st,at)P(\cdot\mid s_t, a_t): the environment transition distribution

The difference from online RL is entirely in what we are allowed to do while searching for π\pi. Online, we can test π\pi, observe the return, and adjust. Offline, that loop is severed: the objective is still defined against the true environment, but we can only ever evaluate candidate policies against our beliefs about it.

Offline, Batch, and Off-Policy RL

These terms are used loosely in the literature and the distinctions matter.

  • Off-policy learning means the update rule can use data from a policy other than the current one. It says nothing about interaction. Deep Q-learning is off-policy and can still collect fresh data.
  • Offline (or batch) learning means the dataset is fixed during the offline training phase; the learner collects no new transitions in that phase. Later deployment or offline-to-online fine-tuning is a separate stage. This is a restriction on data acquisition during training, not on the update rule.
  • Off-policy evaluation is the sub-problem of estimating J(π)J(\pi) for a candidate π\pi using data from ρβ\rho_\beta. It arises in online settings too, but offline RL cannot avoid it.

The last point is the practical sting. In online RL you evaluate candidates by running them. Offline you must estimate their value from the same fixed data you used to train, which means your evaluation inherits your coverage problems. We return to this in the final section.

Occupancy measures and the coverage assumption

It is convenient to work with discounted state-action occupancy measures rather than raw trajectories. For a policy π\pi with initial distribution d0d_0, define the mass of any measurable state-action set BB by

dπ(B)=(1−γ)∑t=0∞γt Pr⁡((st,at)∈B∣π,d0).d^\pi(B) = (1-\gamma)\sum_{t=0}^{\infty} \gamma^{t}\, \Pr((s_t,a_t)\in B \mid \pi, d_0).

where:

  • dπ(B)d^\pi(B): the discounted occupancy mass that policy π\pi assigns to set BB
  • Pr⁡((st,at)∈B∣π,d0)\Pr((s_t,a_t)\in B\mid\pi, d_0): the probability of visiting BB at step tt under policy π\pi starting from d0d_0
  • 1−γ1-\gamma: the normalising factor that makes the sum a probability distribution

The (1−γ)(1-\gamma) factor makes dπd^\pi a probability distribution over S×A\mathcal{S}\times\mathcal{A}. In a discrete space, dπ(s,a)d^\pi(s,a) abbreviates the mass of the singleton {(s,a)}\{(s,a)\}; in a continuous space, it denotes a density only when one exists with respect to a stated reference measure. Why bother with this object rather than working with the trajectory distribution directly? Because an occupancy measure summarises "where the policy spends its time," weighting near-term visits more heavily than far-future ones. It converts the trajectory of the policy into a static picture of the space, which is exactly what we need to compare two policies and to reason about how much their visited regions overlap. With this normalisation, the objective is a simple expectation:

J(π)=11−γ E(s,a)∼dπ[r(s,a)].J(\pi) = \frac{1}{1-\gamma}\,\mathbb{E}_{(s,a)\sim d^\pi}\big[r(s,a)\big].

Now compare dπd^\pi with ρβ\rho_\beta. If a policy spends time where the behavior distribution has substantial mass, the data can inform those choices; where it has none, learning requires additional structure. For the displayed ratio, read both quantities as masses in a discrete space or densities with respect to a common reference measure in a continuous space. This intuition is captured by the concentrability coefficient:

C(π)  =  sup⁡s,a  dπ(s,a)ρβ(s,a).C(\pi) \;=\; \sup_{s,a} \; \frac{d^\pi(s,a)}{\rho_\beta(s,a)} .

where:

  • C(π)C(\pi): the concentrability coefficient of policy π\pi
  • dπ(s,a)d^\pi(s,a): the discounted occupancy of (s,a)(s,a) under policy π\pi
  • ρβ(s,a)\rho_\beta(s,a): the behaviour distribution's density at (s,a)(s,a)
  • sup⁡s,a\sup_{s,a}: the supremum over state-action pairs in the discrete case, or the essential supremum of the density ratio in the continuous case

Define the ratio as infinite when the target occupancy assigns positive mass to a region where the behavior distribution assigns zero; zero-over-zero pairs do not affect it. In continuous spaces, this requires absolute continuity and an essentially bounded density ratio. Thus finite C(π)C(\pi) requires the occupancy of that policy to be supported by the behavior distribution. Some offline-RL guarantees assume coverage of an optimal policy, C⋆=C(π⋆)<∞C^\star=C(\pi^\star)<\infty, together with further statistical and algorithmic conditions. Covering one optimal policy is weaker than covering every policy, but it does not by itself make any particular learning algorithm successful. A small coefficient says the behavior distribution gives reasonable weight to what the policy visits; a large one signals weak coverage.

Concentrability

A dataset concentrates on a policy class Π\Pi with coefficient CC if, for every π∈Π\pi \in \Pi,

sup⁡s,a dπ(s,a)ρβ(s,a)≤C.\sup_{s,a}\ \frac{d^\pi(s,a)}{\rho_\beta(s,a)} \le C .

Roughly: no state-action pair that some policy in Π\Pi visits often is more than a factor CC rarer in the data. Small CC means the data covers what the policies want to do.

Coverage is one way to make offline learning identifiable. Without coverage or justified structure linking unobserved outcomes to observed ones, no method can guarantee good decisions uniformly over the unrestricted MDP class.

The coefficient compares state-action occupancy densities. The next picture simplifies them to their state marginals: if the target visits states the behaviour never reaches, then it necessarily visits unsupported state-action pairs and C(π)C(\pi) is infinite. A gap confined to actions at otherwise covered states would also matter, but this state-only picture does not show it.

In[3]:
Code
import numpy as np

# Schematic occupancies on a one-dimensional state space s in [0, 1].
# The behaviour density is truncated to [0.0, 0.5]. Most target mass lies
# to the right of 0.5, but Gaussian tails also overlap behaviour support
# below 0.45; the exact overlap is not confined to [0.45, 0.5].
_state = np.linspace(0.0, 1.0, 501)
_behaviour_support = (_state >= 0.0) & (_state <= 0.5)


def _bump_mixture(x, centres, width):
    """Normalised sum of Gaussian bumps, used only for this schematic figure."""
    bumps = np.exp(-0.5 * ((x[:, None] - np.asarray(centres)) / width) ** 2)
    density = bumps.sum(axis=1)
    return density / (density.sum() * (x[1] - x[0]))


rho_beta = _bump_mixture(_state, [0.15, 0.30, 0.45], 0.07) * _behaviour_support
rho_beta = rho_beta / (rho_beta.sum() * (_state[1] - _state[0]))
d_pi = _bump_mixture(_state, [0.45, 0.65, 0.85], 0.07)
unsupported_mask = (rho_beta <= 0.0) & (d_pi > 1e-4)
Out[4]:
Visualization
Behaviour and target state-marginal density curves with the unsupported state region shaded; action coverage is not depicted.
Schematic state-marginal occupancies, not the full state-action densities in the definition of concentrability. The behaviour state density is confined to the left half, while the target places mass beyond that support boundary. This state-support gap is sufficient to make the state-action concentrability ratio infinite; action-only gaps are not shown.
Why uncovered actions need more assumptions

Consider two MDPs that agree everywhere ρβ\rho_\beta puts mass and disagree on an unseen action. Both produce the same dataset. In one, the unseen action is better than the logged choice; in the other, it is worse. Any data-only algorithm makes the same choice in both, so it cannot guarantee near-optimality in both over this unrestricted class. Coverage or justified structural information is needed to distinguish them.

The fixed log alone cannot identify the outcome of that unseen action. An offline method needs either relevant coverage or assumptions that constrain such outcomes.

The simulation lemma

Now suppose we learn a model M^=(P^,r^)\hat M = (\hat P, \hat r) from D\mathcal{D} and use it in place of the environment. How wrong is the resulting value function? The answer is the simulation lemma, and it deserves a careful derivation because it is the theoretical heart of the whole chapter.

Fix a deterministic policy π\pi for this derivation; a stochastic policy gives the same identity after averaging each action-dependent term over a∼π(⋅∣s)a\sim\pi(\cdot\mid s). Write VMπV^\pi_M for its value under the true MDP and VM^πV^\pi_{\hat M} for its value under the model. Define the gap Δ(s)=VMπ(s)−VM^π(s)\Delta(s) = V^\pi_M(s) - V^\pi_{\hat M}(s). Expanding both Bellman equations,

VMπ(s)=r(s,π(s))+γ Es′∼P(⋅∣s,π(s))[VMπ(s′)],VM^π(s)=r^(s,π(s))+γ Es′∼P^(⋅∣s,π(s))[VM^π(s′)].\begin{aligned} V^\pi_M(s) &= r(s,\pi(s)) + \gamma\, \mathbb{E}_{s'\sim P(\cdot\mid s,\pi(s))}\big[V^\pi_M(s')\big],\\ V^\pi_{\hat M}(s) &= \hat r(s,\pi(s)) + \gamma\, \mathbb{E}_{s'\sim \hat P(\cdot\mid s,\pi(s))}\big[V^\pi_{\hat M}(s')\big]. \end{aligned}

Subtracting the second from the first and adding and subtracting a term that evaluates VMπV^\pi_M under the model,

Δ(s)=(r−r^)⏟reward error+γ(EPVMπ−EP^ VMπ)⏟transition error, evaluated on the true value+γ EP^[Δ(s′)].\Delta(s) = \underbrace{\big(r - \hat r\big)}_{\text{reward error}} + \gamma \underbrace{\Big(\mathbb{E}_{P} V^\pi_M - \mathbb{E}_{\hat P}\, V^\pi_M\Big)}_{\text{transition error, evaluated on the true value}} + \gamma\, \mathbb{E}_{\hat P}\big[\Delta(s')\big].

where:

  • Δ(s)=VMπ(s)−VM^π(s)\Delta(s) = V^\pi_M(s) - V^\pi_{\hat M}(s): the value gap at state ss
  • r−r^r - \hat r: the reward-model error, abbreviated r(s,π(s))−r^(s,π(s))r(s,\pi(s)) - \hat r(s,\pi(s))
  • EPVMπ−EP^ VMπ\mathbb{E}_P V^\pi_M - \mathbb{E}_{\hat P}\, V^\pi_M: the transition error, evaluated using the true value function VMπV^\pi_M
  • EP^[Δ(s′)]\mathbb{E}_{\hat P}[\Delta(s')]: the expected continuation of the gap under the model's dynamics

The adding-and-subtracting trick is the algebraic seed of the whole result. We insert EPVMπ−EP^VMπ\mathbb{E}_P V^\pi_M - \mathbb{E}_{\hat P} V^\pi_M not because it was there already, but to isolate the reward mismatch and the transition mismatch as a single "local inconsistency" while leaving a clean recursion for the remaining gap. Everything that follows is a matter of chasing where that local inconsistency gets weighted.

The first two terms do not depend on Δ\Delta; call their sum the local inconsistency ℓ(s)\ell(s). The recursion Δ=ℓ+γEP^Δ\Delta = \ell + \gamma \mathbb{E}_{\hat P}\Delta unrolls along the trajectory distribution generated by the model. Write dM^π(s,a)d^\pi_{\hat M}(s,a) for the model's normalised discounted state-action occupancy and dM^π(s)d^\pi_{\hat M}(s) for its state marginal, both induced by the learned model M^\hat M rather than the true MDP. Averaging the initial-state gap over s0∼d0s_0\sim d_0 gives

  JM(π)−JM^(π)  =  11−γ Es∼dM^π[r(s,π(s))−r^(s,π(s))+γδ(s)]  \boxed{\;J_M(\pi)-J_{\hat M}(\pi) \;=\; \frac{1}{1-\gamma}\,\mathbb{E}_{s\sim d^\pi_{\hat M}}\big[r(s,\pi(s)) - \hat r(s,\pi(s)) + \gamma\delta(s)\big]\;}

Here JM(π)=Es0∼d0VMπ(s0)J_M(\pi)=\mathbb E_{s_0\sim d_0}V_M^\pi(s_0), JM^(π)=Es0∼d0VM^π(s0)J_{\hat M}(\pi)=\mathbb E_{s_0\sim d_0}V_{\hat M}^\pi(s_0), and δ(s)=EP(⋅∣s,π(s))[VMπ(s′)]−EP^(⋅∣s,π(s))[VMπ(s′)]\delta(s)=\mathbb{E}_{P(\cdot\mid s,\pi(s))}[V^\pi_M(s')] - \mathbb{E}_{\hat P(\cdot\mid s,\pi(s))}[V^\pi_M(s')]. A pointwise value-gap identity would instead use a model occupancy conditioned on that particular starting state; the displayed occupancy is averaged over d0d_0.

The boxed equation averages local errors over states visited by a fixed policy in the learned model, not necessarily in reality. If the model's dynamics place substantial occupancy in a region where its predictions are wrong, that region contributes heavily to the value gap. The identity does not say that error itself attracts a fixed policy toward that region.

Error compounding and model exploitation are distinct steps. The fixed-policy identity describes how one-step errors propagate through model rollouts. When a policy is then selected by maximizing learned-model value, it may favor optimistic, poorly supported regions; that selection can change the model occupancy in the identity. The two mechanisms can reinforce each other, but neither the identity nor a fixed-policy bound alone proves that the errors concentrate exactly where they are largest.

A worked example: how vacuous the worst-case bound becomes

Bound the two error terms to get a usable inequality. Suppose the reward model is off by at most ϵr\epsilon_r and the transition model has total-variation error at most ϵP\epsilon_P:

∣r(s,a)−r^(s,a)∣≤ϵr,∥P(⋅∣s,a)−P^(⋅∣s,a)∥TV≤ϵP.\big|r(s,a) - \hat r(s,a)\big| \le \epsilon_r, \qquad \big\|P(\cdot\mid s,a) - \hat P(\cdot\mid s,a)\big\|_{TV} \le \epsilon_P .

where ϵr\epsilon_r bounds the reward error, ϵP\epsilon_P bounds the total-variation transition error, and Rmax⁡R_{\max} is the maximum absolute reward. Since ∣VMπ(s′)∣≤Rmax⁡/(1−γ)|V^\pi_M(s')| \le R_{\max}/(1-\gamma) and ∣Epf−Eqf∣≤2∥p−q∥TVmax⁡∣f∣|\mathbb{E}_p f - \mathbb{E}_q f| \le 2\|p-q\|_{TV}\max|f|,

∣VMπ(s)−VM^π(s)∣  ≤  ϵr1−γ  +  2γRmax⁡ ϵP(1−γ)2.\big|V^\pi_M(s) - V^\pi_{\hat M}(s)\big| \;\le\; \frac{\epsilon_r}{1-\gamma} \;+\; \frac{2\gamma R_{\max}\,\epsilon_P}{(1-\gamma)^2}.

where:

  • ϵr\epsilon_r: the maximum reward-model error, ϵr≥∣r(s,a)−r^(s,a)∣\epsilon_r \ge |r(s,a) - \hat r(s,a)|
  • ϵP\epsilon_P: the maximum total-variation transition-model error
  • Rmax⁡R_{\max}: the maximum absolute per-step reward
  • γ\gamma: the discount factor

The 1/(1−γ)1/(1-\gamma) and 1/(1−γ)21/(1-\gamma)^2 prefactors come from two separate sources, and it is instructive to see both. The reward error is a per-step bias paid every step, so summing it under discounting costs the horizon factor 1/(1−γ)1/(1-\gamma). The transition error is more expensive: total-variation error is compared against value functions bounded by Rmax⁡/(1−γ)R_{\max}/(1-\gamma), and each step of the horizon multiplies the accumulated value bound by another factor, so the sum produces 1/(1−γ)21/(1-\gamma)^2. The quadratic blow-up in the transition term is exactly the compounding effect, and it is the term that decides whether the bound is usable.

Now plug in an illustrative set of numbers. Take Rmax⁡=1R_{\max}=1 and γ=0.99\gamma = 0.99. If rewards lie in [0,1][0,1], then returns lie in [0,100][0,100]; under the weaker ∣r∣≤1|r|\leq 1 assumption, returns lie in [−100,100][-100,100]. Let the transition model have a uniform total-variation error upper bound of ϵP=0.05\epsilon_P = 0.05 per state-action pair. This is a distance between distributions, not a claim that 95% of predicted transitions are correct. The transition-error term gives

2×0.99×1×0.05(1−0.99)2  =  0.09910−4  =  990,\frac{2 \times 0.99 \times 1 \times 0.05}{(1-0.99)^2} \;=\; \frac{0.099}{10^{-4}} \;=\; 990,

With ϵr=0.01\epsilon_r=0.01, the reward-error term adds 1, so the full bound is 991. Even the transition term alone exceeds a range-based upper bound of 201 on the value difference when true rewards satisfy ∣r∣≤1|r|\leq 1 and the learned reward may differ by up to ϵr=0.01\epsilon_r=0.01. The bound is therefore vacuous in this example: a uniform 0.05 per-step TV bound does not yield a useful guarantee at this discount factor.

Plotting the two error terms against the discount factor shows which one drives the blow-up.

In[5]:
Code
import numpy as np

# Bound terms from the worked example, as functions of the discount factor.
# Reward error is bounded by eps_r / (1 - gamma); transition error by
# 2 * gamma * R_max * eps_P / (1 - gamma) ** 2. The values fix eps_r = 0.01,
# eps_P = 0.05, and R_max = 1.0, matching the numbers discussed in the text.
_gamma_grid = np.linspace(0.5, 0.999, 400)
EPS_R = 0.01
EPS_P = 0.05
R_MAX = 1.0
reward_term = EPS_R / (1.0 - _gamma_grid)
transition_term = 2.0 * _gamma_grid * R_MAX * EPS_P / (1.0 - _gamma_grid) ** 2
Out[6]:
Visualization
Two curves rising sharply as discount factor approaches one.
Worst-case reward-error and transition-error terms as the discount factor approaches one, with eps_r = 0.01 and eps_P = 0.05. The plotted transition term grows through the squared horizon factor and dwarfs the plotted reward term; the possible-value range used for the vacuity comparison is discussed in the surrounding prose, not plotted here.

The moral is not that the bound is bad. The bound is correct, and it is telling you something true: worst-case, there is no useful guarantee. A useful guarantee requires additional information, such as coverage of the policies under consideration or justified structure that bounds model error where they go. Pessimistic penalties and policy constraints can use such assumptions to limit unsupported choices; they do not themselves create coverage or certify a heuristic uncertainty score. This is why offline model-based RL is not simply "model-based RL without exploration."

It is worth emphasising that the numbers here are chosen to make the point sharp, not because ϵP=0.05\epsilon_P = 0.05 is a realistic worst case. A uniform TV upper bound of 0.05 permits that much transition error at any state-action pair, including those the policy visits; it does not say actual error is 0.05 everywhere. The vacuity comes from pairing that global error allowance with a long horizon. Shorten the horizon (γ→0\gamma \to 0) and the bound tightens quickly; make the model extremely accurate only where the policy goes and the uniform bound overstates the damage. Both of these escapes correspond to real assumptions we can make, and both are what the algorithms in the next sections exploit.

Model-free and model-based offline RL

Both families exist, and they differ in what they learn.

  • Model-free offline RL learns a QQ-function or policy directly from D\mathcal{D}, and handles distribution shift by constraining the policy or penalising the critic. Representative strictly offline methods include BCQ, BEAR, BRAC, TD3+BC, IQL, and CQL. AWAC can use an offline warm start but was designed for subsequent online improvement.
  • Model-based offline RL learns predictive structure from D\mathcal D and uses it for rollouts, planning, or critic training without new interaction. MOPO, COMBO, RAMBO, and MOBILE use learned dynamics in policy learning; MOReL plans in a pessimistic learned MDP, MBOP uses model-based planning, and Trajectory Transformer models joint state-action-reward sequences for reward-guided planning rather than fitting separate P^\hat P and r^\hat r networks.

A learned transition model can generate arbitrarily many synthetic transitions, including in regions the dataset never visited. Those samples do not create new evidence about the real environment. The design problem is to use rollouts where they help policy learning without letting unsupported predictions mislead the optimiser.

Why bother with a model at all, given the risk? Because of the amplification factor. With NN real transitions you can train a value function on NN samples, or you can train a model on NN samples and generate many imagined transitions from it. If the model is accurate on the region the policy cares about, those rollouts can help. MBPO made this argument in the online setting, as discussed in Part VIII: Decision-Centric Research Lineages. Online interaction can expose a wrong prediction and provide data for correction, though it does not guarantee safe exploration or eventual recovery. Offline, the fixed log cannot supply that targeted correction; repeated synthetic rollouts retain the model's unsupported assumptions.

The natural worry is a mismatch between where the model is accurate and where an improved policy might go. The dataset reflects ρβ\rho_\beta, while an unconstrained improvement step can favor actions with little coverage. Improvement need not leave support: a mixed or suboptimal behaviour dataset may already contain better actions. The methods below try to use that opportunity without letting the optimiser exploit unsupported estimates.

Out-of-Distribution Actions and Model Uncertainty

This section defines what "out of distribution" means for a learned model and why the action dimension is usually where the trouble starts. It builds the machinery for quantifying epistemic uncertainty, then constructs a gridworld where the failure is visible and unambiguous.

What is actually out of distribution

A dataset of continuous-control trajectories often covers a thin tube in joint state-action space. Even millions of transitions may leave reachable (s,a)(s,a) pairs poorly sampled. A robot moving through a hallway traces a narrow ribbon inside a higher-dimensional joint space. Three useful views of extrapolation arise; they overlap rather than partition the space.

  • State OOD. The query state lies outside, or in a poorly covered part of, the behaviour state's distribution. This can happen when the policy drives into a new configuration; exact equality with a logged continuous state is not the test.
  • Action OOD. The state is familiar but the proposed action is poorly supported conditional on that state. This is subtler: driving logs may cover a speed well but contain little evidence about full braking at that speed.
  • Joint OOD. The pair (s,a)(s,a) is missing or rare in the logged joint distribution, even when the state and action have each appeared separately. State OOD or action OOD also makes the pair OOD; the distinctive case is a novel pairing of familiar marginals.

Action OOD is the one people underestimate. Consider a behaviour policy that is nearly deterministic, as some demonstration datasets are. In a well-covered state, the marginal ρβ(s)\rho_\beta(s) may be large, but conditional on ss the action distribution is a narrow spike. The model has evidence for only a narrow range of actions at that state, even if it has many transitions. If an improvement step proposes an action 2 standard deviations away from the spike, the transition model may have little evidence about the consequence, but the value function still returns a number, and the optimiser may still trust it.

These views interact. In a low-dimensional example, a state far from logged states may be visibly unusual; in high dimensions, that judgment is harder. A state-only density model cannot detect a poorly supported action at a familiar state. A familiar-looking state and a familiar-looking action can also form a rare pair, so checking only p(s)p(s) cannot establish joint state-action coverage. The relevant question for a transition model is whether the pair has enough nearby evidence.

In reconstructive latent world models from Part III: Representing Agents and Worlds, an encoder maps observations into a latent space and a decoder reconstructs observations. Other architectures, including MuZero, need not have a decoder. If planning drives a latent state far beyond training support, its predicted transitions and values can become unreliable even though the vectors remain numerically valid. A model may not raise an exception or produce a NaN when this happens, so checking only numerical validity will miss the coverage problem.

Aleatoric versus epistemic uncertainty

Two kinds of uncertainty are routinely conflated, and conflating them causes real damage.

  • Aleatoric uncertainty is irreducible randomness in the environment. A dice roll, a slipping wheel, a stochastic reward channel. It does not shrink with more data, because there is nothing more to know.
  • Epistemic uncertainty is uncertainty about the model itself, arising from limited relevant evidence. It can shrink with informative data near the query point and often rises away from observed regions, though approximate or misspecified models need not show either trend reliably.

For a penalty meant to guard against model extrapolation, epistemic uncertainty is the quantity of interest: irreducible randomness alone does not show that the model lacks evidence. A risk-sensitive or safety objective may deliberately penalise stochastic outcomes for a different reason. Practical scores can mix these components. MOPO's published scale-based heuristic is one example, and its penalty coefficient must be tuned rather than read as a calibrated error bound.

Consider a robot driving on ice. If friction varies unpredictably, more examples may clarify the distribution of slides without making each slide deterministic. Penalising that randomness could make the robot avoid a well-understood route. By contrast, if the robot has never driven on a particular gravel surface, the uncertainty reflects missing evidence about that surface. A naive predictive-variance score can mix these two cases.

Quantifying epistemic uncertainty

We will compare five approaches to estimating or bounding uncertainty. They differ in cost, calibration, and how badly they can fail.

A schematic decomposition makes the distinction concrete before the individual estimator families are described.

In[7]:
Code
import numpy as np

# A schematic variance decomposition over a one-dimensional input. Data cover
# the interval [2, 8]; aleatoric noise is a constant 0.04 everywhere. The
# epistemic term grows with distance from the covered interval, so the total
# predictive variance is non-zero even where the environment is well sampled.
_x = np.linspace(0.0, 10.0, 401)
_data_low, _data_high = 2.0, 8.0
distance = np.where(
    _x < _data_low,
    _data_low - _x,
    np.where(_x > _data_high, _x - _data_high, 0.0),
)
aleatoric = np.full_like(_x, 0.04)
epistemic = 0.5 * (1.0 - np.exp(-0.5 * distance**2))
predictive = epistemic + aleatoric
Out[8]:
Visualization
Three curves showing total, epistemic, and aleatoric variance across a covered interval.
Schematic decomposition of predictive variance along a one-dimensional input. Epistemic variance grows outside the covered interval, while irreducible aleatoric variance stays constant. Total variance is therefore not a pure score for missing model evidence; a separate risk-sensitive objective may still care about the aleatoric component.

No single family dominates: the choice depends on whether you need speed, a defensible statistical statement, or robustness to shared model bias.

Ensembles and disagreement. Train KK dynamics models {P^θ1,…,P^θK}\{\hat P_{\theta_1}, \ldots, \hat P_{\theta_K}\} with different random initialisations and different bootstrapped data subsets. Each predicts a next-state mean μk(s,a)\mu_k(s,a). Then

u(s,a)=max⁡k∥μθk(s,a)−μˉ(s,a)∥2,μˉ(s,a)=1K∑k=1Kμθk(s,a).u(s,a) = \max_k \Big\| \mu_{\theta_k}(s,a) - \bar\mu(s,a)\Big\|_2, \qquad \bar\mu(s,a) = \frac{1}{K}\sum_{k=1}^{K}\mu_{\theta_k}(s,a).

where:

  • u(s,a)u(s,a): the disagreement-based uncertainty score at (s,a)(s,a)
  • μθk(s,a)\mu_{\theta_k}(s,a): the next-state mean predicted by ensemble member kk with parameters θk\theta_k
  • μˉ(s,a)\bar\mu(s,a): the mean prediction averaged over the KK members
  • ∥⋅∥2\|\cdot\|_2: the Euclidean norm

This is one possible disagreement score, not MOPO's practical penalty. MOPO uses the maximum Frobenius norm of a member's predicted Gaussian scale matrix across its probabilistic model ensemble, which its authors describe as a maximum standard-deviation score. Ensemble disagreement can be useful, but it is not a calibrated posterior uncertainty; members with shared architecture and data may agree even when they are all wrong. Nor does predictive variance automatically isolate epistemic uncertainty from environmental randomness. The disagreement captures how much the members differ from each other, not how far the ensemble as a whole is from the truth. If all members are biased in the same way, the disagreement stays small while the error stays large.

Count-based uncertainty. In discrete or discretisable spaces, uncertainty is a function of the visit count n(s,a)n(s,a). A common form is

u(s,a)=αα+n(s,a),oru(s,a)=1n(s,a)+1.u(s,a) = \frac{\alpha}{\alpha + n(s,a)}, \qquad \text{or} \qquad u(s,a) = \frac{1}{\sqrt{n(s,a)+1}}.

where:

  • u(s,a)u(s,a): the uncertainty score at (s,a)(s,a)
  • n(s,a)n(s,a): the number of times (s,a)(s,a) was visited in the dataset
  • α\alpha: a smoothing prior (set to 1 in our implementation below)

The second has the 1/n1/\sqrt n scaling of many concentration bounds. It is not, without a specified prior and observation-noise variance, an exact posterior standard deviation. Count-based scores are cheap and interpretable when counts are meaningful; we use one below to isolate the penalty mechanism from neural-network training noise. Their weakness is that "count" requires a definition of "the same event," and in continuous state spaces that definition is a discretisation, which introduces its own choices.

Density and likelihood models. Fit ρ^β(s,a)\hat\rho_\beta(s,a) with a normalising flow or a VAE, then use a quantity such as −log⁡ρ^β(s,a)-\log \hat\rho_\beta(s,a) as a proxy for low data density. It is not automatically a calibrated bound on transition-model error or a reliable test of support. BEAR instead uses behavior-action samples and an MMD constraint in policy improvement. FisherBRC learns a behavior density for its critic decomposition Qθ(s,a)=Oθ(s,a)+log⁡μ^(a∣s)Q_\theta(s,a)=O_\theta(s,a)+\log\hat\mu(a\mid s) and regularizes action gradients of the offset; neither method defines this density score as a world-model reward penalty. Likelihood is not the same as support: a generative model can give high likelihood to out-of-distribution inputs, as likelihood-based OOD studies demonstrate.

Bayesian neural networks and variational inference. Place a prior over model parameters and use a posterior predictive distribution. These methods introduce inference and calibration choices that differ from those of a bootstrap ensemble; no single family wins across tasks and model classes. PILCO's Gaussian-process dynamics, discussed in Part VIII: Decision-Centric Research Lineages, provide posterior uncertainty under the GP model assumptions. Naive exact GP fitting has cubic time cost in the number of training points, so that particular implementation does not scale cheaply to large datasets.

Spectral and Lipschitz-based bounds. A smooth learned model alone does not guarantee a small error outside the data. If the error function is known to be Lipschitz, or both true and learned dynamics have compatible known Lipschitz bounds, an error measured near a query point can be extended with a distance term. Such a result is deterministic under its assumptions, though its constants may be too loose to guide planning. The needed quantity is a bound on how quickly error can change, not merely on how quickly the learned prediction changes.

Practical ensemble hygiene

Ensembles are easy to get subtly wrong. Three details matter.

  1. Diversity source. Varying only the initialisation often produces models that agree too much. Bootstrapping the data (sample NN transitions with replacement per member) or using different data orderings helps.
  2. Noise handling. If each member samples an independent stochastic successor during a rollout, raw sample spread mixes environmental randomness with model disagreement. Compare predictions under a controlled noise coupling, or separate disagreement among predicted means from each model's predictive variance. The two components need different interpretations; MOPO's published practical penalty uses a norm of predicted Gaussian scale (standard deviation), not the mean-disagreement formula above.
  3. Ensemble size. A small ensemble is a common computational compromise, but there is no universal member count at which disagreement becomes calibrated or further models cease to help. Check size against held-out model error and compute cost in the task at hand.

A gridworld where the failure is visible

Now we build something concrete. The goal is not to reproduce a benchmark. It is to construct the smallest system in which we can see the mechanism with our own eyes, compute everything exactly, and check the answer against ground truth. Minimality matters here. Offline RL failures on real benchmarks are opaque: you see a poor return, but you cannot easily separate model error from coverage. In a small grid we can isolate those two factors directly.

The environment is a 7×607 \times 60 grid. Row 3 is a corridor that runs the full width; the goal sits at the far right end of that row. The grid edge acts as a wall. The goal is absorbing with zero reward, and reaching it pays 1.0. The discount factor is γ=0.99\gamma = 0.99.

The behaviour policy is a random walk along row 3, moving right with probability 0.7 and left with probability 0.3. It never leaves the corridor. This is the crucial design choice: the dataset covers one row out of seven, and covers it densely. Every other cell in the grid has never been seen.

The leftward moves matter. With a purely rightward policy, each of the 59 pre-goal corridor cells would be visited once per episode, while the absorbing goal would still receive 20 logged tail visits. The 30% leftward action instead makes the walk revisit earlier cells, producing a non-uniform pre-goal visit distribution.

The left edge clips a left action back to the same cell, while the rightward bias eventually carries most episodes to the goal. Because episodes reset at the start and stop after a short tail of goal self-transitions, the logged counts are not a stationary distribution of an endless reflecting walk. They need not be uniform along the corridor; the heatmap below shows the actual finite-episode coverage.

The optimal policy is trivially "walk right 59 times," so the return from the start state is ∑t=058γt⋅rt+1=γ58≈0.9958≈0.558\sum_{t=0}^{58} \gamma^t \cdot r_{t+1} = \gamma^{58} \approx 0.99^{58} \approx 0.558. (The reward is paid on arrival at the goal, after 59 steps, which is step index 58 from the start.) So we can compute every quantity exactly and ask whether each method finds it.

In[9]:
Code
import numpy as np

# A 7 x 60 grid. Row 3 is the corridor that the behaviour policy walks along;
# every other cell is never visited. The goal sits at the far right of row 3.
H, W = 7, 60
N_STATES = H * W
N_ACTIONS = 4
CORRIDOR_ROW = 3
START = CORRIDOR_ROW * W
GOAL = CORRIDOR_ROW * W + (W - 1)
GAMMA = 0.99
DELTAS = ((-1, 0), (1, 0), (0, -1), (0, 1))
ACTION_NAMES = ("up", "down", "left", "right")


def true_step(state, action):
    """Ground-truth transition. The goal is absorbing with zero reward."""
    if state == GOAL:
        return GOAL, 0.0
    row, col = divmod(state, W)
    d_row, d_col = DELTAS[action]
    row = min(max(row + d_row, 0), H - 1)  # the grid edge acts as a wall
    col = min(max(col + d_col, 0), W - 1)
    nxt = row * W + col
    return nxt, 1.0 if nxt == GOAL else 0.0

Next we roll out the behaviour policy and record the transitions. Each episode runs until the goal is reached, followed by a short tail of goal self-transitions. The logged tail includes only left and right actions at the goal, so the empirical model learns those two self-loops; up and down remain unobserved and receive the same uniform fallback as other unsupported pairs. The true goal is absorbing under every action. This distinction matters: samples of some terminal actions do not tell a purely empirical model that all terminal actions must self-loop unless that structure is supplied explicitly.

In[10]:
Code
def behaviour_action(rng):
    """A random walk confined to row 3: 70% right, 30% left."""
    return 3 if rng.random() < 0.7 else 2


rng = np.random.default_rng(0)
transitions = []
for _ in range(300):
    state = START
    for _ in range(400):
        action = behaviour_action(rng)
        nxt, reward = true_step(state, action)
        transitions.append((state, action, reward, nxt))
        state = nxt
        if state == GOAL:
            break
    assert state == GOAL, (
        "This fixed-seed episode must reach the goal before its tail"
    )
    for _ in range(20):  # a short tail of goal self-transitions
        action = behaviour_action(rng)
        nxt, reward = true_step(state, action)
        transitions.append((state, action, reward, nxt))
        state = nxt

dataset = np.asarray(transitions, dtype=np.int64)

We now tabulate the data. counts records how often each state-action pair was taken, reward_sums accumulates the observed reward, and next_counts records how often each particular successor followed. Everything downstream is built from these three arrays: the empirical model is a normalisation of them, and the uncertainty scores are functions of counts.

In[11]:
Code
counts = np.zeros((N_STATES, N_ACTIONS))
reward_sums = np.zeros((N_STATES, N_ACTIONS))
next_counts = np.zeros((N_STATES, N_ACTIONS, N_STATES))

for state, action, reward, nxt in dataset:
    counts[state, action] += 1
    reward_sums[state, action] += reward
    next_counts[state, action, nxt] += 1
Out[12]:
Console
transitions collected        : 49,937
state-action pairs observed  : 120 of 1,680
grid cells visited           : 60 of 420

The printed counts show 120 of 1,680 state-action pairs observed, about 7.1%. The 60 visited cells are exactly the ones in row 3. Every other cell is unobserved, and every method below has to decide what to do there. The 7.1% figure is worth pausing on: it is not an extreme number, and yet it is enough to make a naive planner fail completely, as we are about to see.

Fitting the empirical model

The model is deliberately simple: a maximum-likelihood empirical transition model with a uniform fallback for unobserved pairs. This is the tabular analogue of a neural dynamics model with no regularisation, and it has the property we want to study: off-support, it returns something arbitrary, determined by our fallback choice rather than by evidence.

P^(s′∣s,a)={n(s,a,s′)n(s,a),n(s,a)>0,1∣S∣,n(s,a)=0,r^(s,a)={1n(s,a)∑i:(si,ai)=(s,a)ri,n(s,a)>0,0,n(s,a)=0.\hat P(s'\mid s,a) = \begin{cases} \dfrac{n(s,a,s')}{n(s,a)}, & n(s,a) > 0,\\[2mm] \dfrac{1}{|\mathcal{S}|}, & n(s,a) = 0, \end{cases} \qquad \hat r(s,a) = \begin{cases} \dfrac{1}{n(s,a)}\displaystyle\sum_{i:(s_i,a_i)=(s,a)} r_i, & n(s,a) > 0,\\[2mm] 0, & n(s,a) = 0. \end{cases}

where:

  • P^(s′∣s,a)\hat P(s'\mid s,a): the estimated transition probability from (s,a)(s,a) to s′s'
  • r^(s,a)\hat r(s,a): the estimated expected reward at (s,a)(s,a)
  • n(s,a)n(s,a): the number of times (s,a)(s,a) appears in the dataset
  • n(s,a,s′)n(s,a,s'): the number of times s′s' followed (s,a)(s,a)
  • ∣S∣|\mathcal{S}|: the number of states, used as the uniform fallback denominator

The uniform fallback is a stand-in for extrapolation, not a claim that other approximators behave identically. A neural network may continue a learned trend; a random forest prediction reflects its selected leaves; a Gaussian-process posterior mean approaches its prior mean when covariance with the training inputs vanishes, as can happen far away under a decaying kernel. None is automatically correct off-support. The uniform choice makes the failure mode vivid: an unvisited transition can land anywhere in the grid, including directly on the goal.

In[13]:
Code
uniform = np.full(N_STATES, 1.0 / N_STATES)
P_hat = np.tile(uniform, (N_STATES, N_ACTIONS, 1))
R_hat = np.zeros((N_STATES, N_ACTIONS))

seen = counts > 0
P_hat[seen] = next_counts[seen] / counts[seen][:, None]
R_hat[seen] = reward_sums[seen] / counts[seen]

Counting the uncertainty

We use a count-based uncertainty proxy because the state-action visit counts are exact and interpretable in this tabular toy. In a larger continuous-state system, an ensemble is one common alternative; the count score isolates the penalty mechanism from neural-network training noise.

u(s,a)=αα+n(s,a),α=1.u(s,a) = \frac{\alpha}{\alpha + n(s,a)}, \qquad \alpha = 1.

where:

  • u(s,a)u(s,a): the count-based uncertainty score
  • n(s,a)n(s,a): the visit count from the dataset
  • α\alpha: the prior pseudo-count, fixed at 1 here

If a pair has never been observed, u=1u = 1. If it has been observed 999 times, u=1/1000=0.001u = 1/1000 = 0.001. This is monotone, bounded in (0,1](0,1], and it shrinks at roughly 1/n1/n in the well-sampled regime. The 1/n\sqrt{1/n} version has a different decay rate, closer to common concentration-bound scaling; here we use 1/n1/n only as a simple bounded heuristic. Choosing a bounded score also keeps the penalty comparable across state-action pairs: no single pair can swamp the sum with an arbitrarily large uncertainty value.

In[14]:
Code
# Count-based uncertainty proxy in (0, 1]: exactly 1 for an unobserved pair.
UNCERTAINTY_PRIOR = 1.0
u_hat = UNCERTAINTY_PRIOR / (UNCERTAINTY_PRIOR + counts)
visits = counts.sum(axis=1).reshape(H, W)
uncertainty_grid = u_hat[:, 3].reshape(
    H, W
)  # right action, sampled in the corridor
Out[15]:
Visualization
Heatmap of a seven by sixty grid with non-zero counts only in the middle row.
Visit counts per cell in the offline dataset, shown on a log scale. The behaviour policy never leaves row 3, so six of the seven rows are completely unobserved, and the 60 visited cells account for only 120 of the 1,680 state-action pairs, about 7 percent. The corridor is the only region the model has evidence about.
Heatmap of right-action uncertainty, with a lower-valued band in the logged middle row and maximum values elsewhere.
Count-based uncertainty for the right action across the grid. Uncertainty is lower along the logged corridor and equals 1 everywhere else, so even inside the corridor the vertical actions remain unsupported. A familiar state does not make every action at that state well supported.

Two things to notice. Right-action uncertainty is maximal off the corridor, because those pairs were never observed, and lower where rightward actions appear in the logs. The up and down actions remain maximally uncertain even on the corridor; that distinction is in the underlying state-action counts, though this heatmap shows only the right-action slice. A familiar state does not make every action at that state well supported.

A third thing to notice is that cell counts are not uniform even within the corridor. Fifty of the 59 nonterminal corridor cells have between 700 and 800 visits in this seeded run, but the two cells immediately before the goal have only 600 and 428. The absorbing goal has 6000 visits because every one of the 300 episodes adds a tail of 20 goal self-transitions. That bright goal cell should not be mistaken for stronger coverage of the approach states, and cell counts alone still say nothing about unsupported actions at those states.

Solving the model without pessimism

Now we plan. The generic recipe is value iteration on the learned model:

Q^(s,a)=r^(s,a)+γ∑s′P^(s′∣s,a)V^(s′),V^(s)=max⁡aQ^(s,a),\hat Q(s,a) = \hat r(s,a) + \gamma\sum_{s'}\hat P(s'\mid s,a)\hat V(s'), \qquad \hat V(s) = \max_a \hat Q(s,a),

where:

  • Q^(s,a)\hat Q(s,a): the estimated action-value at (s,a)(s,a) under the learned model
  • V^(s)\hat V(s): the estimated state value, the maximum of Q^\hat Q over actions
  • r^(s,a)\hat r(s,a): the estimated reward
  • P^(s′∣s,a)\hat P(s'\mid s,a): the estimated transition probability
  • γ\gamma: the discount factor

iterated to convergence. This is the same fixed-point computation that underlies every model-based planner in Part VII: Planning and Agency, with a learned model in place of the true one. The reason it is dangerous offline is precisely that value iteration trusts the model everywhere: the max⁡a\max_a operator steps through all four actions at every state, including states the model has never seen, and it takes whatever action looks best according to estimates that may be arbitrary.

In[16]:
Code
def value_iteration(P, R, gamma=GAMMA, tol=1e-10, max_sweeps=6000):
    """Value iteration for a finite MDP given as dense (S, A, S) and (S, A) arrays."""
    V = np.zeros(P.shape[0])
    for _ in range(max_sweeps):
        Q = R + gamma * (P @ V)  # (S, A, S) @ (S,) -> (S, A)
        V_next = Q.max(axis=1)
        if np.max(np.abs(V_next - V)) < tol:
            V = V_next
            break
        V = V_next
    return V, R + gamma * (P @ V)

To judge the result we also need the ground truth: the true MDP tensor, a direct policy evaluator, and the true optimal value. The evaluator solves the linear system Vπ=Rπ+γPπVπV^\pi = R^\pi + \gamma P^\pi V^\pi directly. The selected policies' true returns therefore come from a linear solve; planner scores and the optimal value come from value iteration to the stated tolerance. This is one of the luxuries of the gridworld setting: we can compare a method's belief about a policy against the policy's actual value in the true environment with negligible numerical error on the comparison.

In[17]:
Code
P_true = np.zeros((N_STATES, N_ACTIONS, N_STATES))
R_true = np.zeros((N_STATES, N_ACTIONS))
for state in range(N_STATES):
    for action in range(N_ACTIONS):
        nxt, reward = true_step(state, action)
        P_true[state, action, nxt] = 1.0
        R_true[state, action] = reward


def evaluate_policy(P, R, policy, gamma=GAMMA):
    """Exact evaluation of a fixed deterministic policy in the supplied MDP."""
    grid = np.arange(N_STATES)
    P_pi = P[grid, policy]
    R_pi = R[grid, policy]
    return np.linalg.solve(np.eye(N_STATES) - gamma * P_pi, R_pi)


def evaluate_in_true_mdp(policy):
    return evaluate_policy(P_true, R_true, policy)


V_star, _ = value_iteration(P_true, R_true)
In[18]:
Code
PENALTY = 0.5
V_naive, Q_naive = value_iteration(P_hat, R_hat)
V_pess, Q_pess = value_iteration(P_hat, R_hat - PENALTY * u_hat)
pi_naive = Q_naive.argmax(axis=1)
pi_pess = Q_pess.argmax(axis=1)
V_true_naive = evaluate_in_true_mdp(pi_naive)
V_true_pess = evaluate_in_true_mdp(pi_pess)
V_model_naive_policy = evaluate_policy(P_hat, R_hat, pi_naive)
V_model_pess_policy = evaluate_policy(P_hat, R_hat, pi_pess)
Out[19]:
Console
naive model planner        : planner score +3.082 | real -0.000
pessimistic model planner  : planner score +0.509 | real +0.558
true optimal policy        :               | real +0.558

There it is. The unpenalised planner estimates a return of about 3.08 from the start state. Under the true dynamics its greedy policy earns zero. The model's unsupported transition fallback has created a profitable imagined route that the policy exploits.

The reason is worth tracing. The supported route rightward along the corridor reaches the goal in 59 steps, worth γ58=0.9958≈0.558\gamma^{58}=0.99^{58}\approx0.558 in the true MDP. Moving vertically enters unsupported pairs whose fitted transition distribution is uniform over all 420 cells. Some of those random successors lie near the goal and have observed, rewarding actions. At the true absorbing goal, the logged left and right actions self-loop, but the unlogged up and down actions also get the model's uniform fallback, so imagined exits from the goal add another optimistic path. Repeated imagined opportunities make the unsupported route look better than the corridor, even though true vertical moves never teleport. Landing directly on the true goal earns no immediate reward in this implementation; the learned model can nevertheless imagine future reward after an unsupported goal action.

Once the policy steps into an unobserved cell, all four actions have identical estimated value, so the argmax picks the first one, "up." In the true grid it moves from (3,0)(3,0) through (2,0)(2,0) and (1,0)(1,0) to (0,0)(0,0). Further up actions are clipped by the top boundary, leaving it at (0,0)(0,0) forever. It never reaches the goal. The value map below shows the same failure structurally: the unpenalised value function is flat across the unvisited region, a plateau of imagined reward.

The flatness is the tell. In the true grid, a policy that moves toward the goal would have values that vary with distance. Instead, the unpenalised model assigns roughly the same value to every unsampled cell because it uses the same uniform successor distribution for each unsupported pair. This plateau exposes the uniform extrapolation built into the toy model; other model errors need not produce the same visual pattern.

In[20]:
Code
def greedy_trajectory(policy, steps=150):
    state = START
    path = [state]
    for _ in range(steps):
        state, _ = true_step(state, policy[state])
        path.append(state)
    return np.array(path)


path_naive = greedy_trajectory(pi_naive)
path_pess = greedy_trajectory(pi_pess)
rows_naive, cols_naive = np.divmod(path_naive, W)
rows_pess, cols_pess = np.divmod(path_pess, W)
Out[21]:
Visualization
Heatmap of a grid showing a high-value plateau confined to the unvisited off-support region, with the corridor row differing.
Value estimate of the unpenalised planner across the grid. Unsupported cells inherit value from uniform imagined transitions, so a route away from the logged corridor looks attractive; the value is exactly flat (3.08) across all 360 unvisited cells, while the visited corridor row carries the only gradient, rising to about 4.05.
Heatmap of the same grid showing negative values off the corridor and a gradient along it.
Value estimate of the pessimistic planner across the grid. The unvisited region collapses to strongly negative values, and the only positive gradient runs along the corridor toward the goal.
Out[22]:
Visualization
Grid map with a short vertical trajectory line at the left edge that never moves right.
Greedy trajectory of the unpenalised policy overlaid on the coverage heatmap. It leaves row 3 on the first move, reaches (0,0) after three moves, and remains there, earning zero in the true grid while the model estimates about 3.08.
Grid map with a horizontal trajectory line running the full width of the covered row.
Greedy trajectory of the pessimistic policy overlaid on the coverage heatmap. The same model with a count-based penalty walks the corridor to the goal in exactly 59 steps, matching the optimal return, so a penalty of half a unit per unobserved transition flips a policy worth zero into the optimal one.

The two trajectories show the mechanism. Without a penalty, the model's ignorance is read as opportunity and the policy moves off the data. With a penalty, the same model, the same dataset, and the same value iteration produce the optimal policy in this toy. A penalty of 0.5 per unobserved transition is enough to change the selected policy from zero return to the full optimum.

Pessimism and Conservative Objectives

The previous section diagnosed the disease: an unconstrained estimate is maximised by the optimiser, and off-support the estimate is arbitrary. This section describes the cure. The idea is almost embarrassingly simple, but its theoretical consequences are substantial, and the design choices around it are where most of the engineering lives.

The pessimistic principle

Let M\mathcal{M} be a set of MDPs consistent with the data, in the sense that no MDP in the set is ruled out by the observed transitions. Offline, the data cannot distinguish them, so the honest thing to do is to consider the worst case within the set:

Vpessπ(s)=min⁡M∈MVMπ(s).V^{\pi}_{\text{pess}}(s) = \min_{M \in \mathcal{M}} V^{\pi}_M(s).

This is pessimism. It says: if the true environment lies in the candidate set, I will plan for the candidate that treats me worst. Under that condition, the pessimistic value is a lower bound on the true value of a fixed candidate policy; how informative that bound is depends on the set. The corresponding minimax formulation has the agent choose a policy, "nature" choose an MDP from the specified set, and the agent receive the worst-case value. This is a planning objective, not a guarantee of risk-free outcomes.

In practice M\mathcal{M} is never enumerated. Instead, the worst case is approximated by subtracting a penalty from the reward:

r~(s,a)=r^(s,a)−λ u(s,a),\tilde r(s,a) = \hat r(s,a) - \lambda\, u(s,a),

where uu is the uncertainty function from the previous section and λ\lambda is the pessimism coefficient. The penalty is called "conservative" because it reduces the appeal of uncertain regions relative to their unpenalised estimates. A high predicted reward can still outweigh a high penalty; this objective does not impose a confidence threshold. The intuition is that the penalty proxies the gap between r^\hat r and the worst-case reward over M\mathcal{M}, and where uncertainty is large that proxy can be large.

Pessimism and optimism in bandits

The contrast with the classical exploration literature is sharp. In a bandit or an online MDP, a well-known family of algorithms is optimistic: add a bonus to the estimated reward, rUCB(s,a)=r^(s,a)+λu(s,a)r_{\text{UCB}}(s,a) = \hat r(s,a) + \lambda u(s,a), and the agent is driven toward the unfamiliar because the bonus makes it look attractive. The point is to gather information.

During offline training, an optimistic bonus can steer planning toward poorly supported actions without collecting any new evidence. This motivates pessimistic penalties in methods such as the one illustrated here. The policy must be chosen using the fixed dataset, not the prospect of gathering more information after a risky action.

The penalty must dominate the model error

The penalty is not automatically a guarantee. To obtain a lower bound, it must cover the model's possible overestimation at every relevant state-action pair. Here is a sufficient condition for a fixed deterministic policy.

Suppose the true and learned Bellman operators differ by at most ϵ\epsilon when applied to the true value function:

∣(TπVMπ)(s)−(T^πVMπ)(s)∣≤ϵfor all s.\big|(\mathcal{T}^\pi V_M^\pi)(s) - (\hat{\mathcal{T}}^\pi V_M^\pi)(s)\big| \le \epsilon \quad \text{for all } s.

The contraction property then gives ∥VMπ−VM^π∥∞≤ϵ/(1−γ)\|V^\pi_M-V^\pi_{\hat M}\|_\infty\leq\epsilon/(1-\gamma). More locally, let the penalised model Bellman operator be T^penπV(s)=r^(s,π(s))−λu(s,π(s))+γEP^V(s′)\hat{\mathcal T}_{\rm pen}^\pi V(s)=\hat r(s,\pi(s))-\lambda u(s,\pi(s))+\gamma\mathbb E_{\hat P}V(s'). Its value, averaged over the initial distribution, is

J^pen(π)=J^(π)−11−γ E(s,a)∼dM^π[λ u(s,a)].\hat J_{\text{pen}}(\pi) = \hat J(\pi) - \frac{1}{1-\gamma}\,\mathbb{E}_{(s,a)\sim d^\pi_{\hat M}}\big[\lambda\,u(s,a)\big].

For a pointwise lower bound, a sufficient condition is

λu(s,π(s))  ≥  (T^πVMπ)(s)−(TπVMπ)(s)for every s.\lambda u(s,\pi(s))\;\geq\; (\hat{\mathcal T}^\pi V_M^\pi)(s)-(\mathcal T^\pi V_M^\pi)(s)\quad\text{for every }s.

Under this condition, T^penπVMπ≤VMπ\hat{\mathcal T}_{\rm pen}^\pi V_M^\pi\leq V_M^\pi. Monotonicity and contraction of the discounted Bellman operator imply

V^penπ(s)≤VMπ(s)for all s.\hat V^\pi_{\text{pen}}(s) \le V^\pi_M(s) \quad \text{for all } s.

The pessimistic estimate is a certified lower bound only when the stated condition holds. Since VMπV_M^\pi and the true Bellman error are unknown offline, a count or ensemble score is not automatically a valid bound on that error. Policy-improvement guarantees require additional coverage, estimation, and optimisation assumptions; the displayed inequality alone does not give a suboptimality rate.

Two things follow. First, a penalty that is too small to cover overestimation lacks this lower-bound guarantee, whereas one that is too large can suppress useful policies. Second, a uniform penalty calibrated to a global worst case taxes even well-understood steps. Here is the deep problem: the condition depends on VMπV^\pi_M, which you do not have. A generic bound is ∣VMπ∣≤Rmax⁡/(1−γ)|V^\pi_M|\leq R_{\max}/(1-\gamma); tighter local bounds require additional assumptions. Using only the generic bound can produce the vacuous guarantee from the worked example earlier. More selective methods restrict where pessimism enters, and the rest of this section compares those choices.

The algorithm families

MOPO. Model-based Offline Policy Optimization learns an ensemble of probabilistic dynamics models. Its practical uncertainty score takes the largest Frobenius norm of a member's predicted Gaussian scale matrix (described in the paper as maximum standard deviation), u(s,a)=max⁡i∥Σi(s,a)∥Fu(s,a)=\max_i\|\Sigma_i(s,a)\|_F; it subtracts λu\lambda u from reward and trains a policy with MBPO-style short model rollouts. That score is heuristic rather than the admissible model-error upper bound required by the paper's theorem. The theoretical comparison also uses the true reward function in its penalized model; the practical implementation learns a reward model. Under the theorem's function-class, exact-reward, and admissible-dynamics-error assumptions, the result compares the learned policy with any candidate policy by

JM(π^)  ≥  sup⁡π{JM(π)−2λuˉ(π)},J_M(\hat\pi) \;\geq\; \sup_\pi\big\{J_M(\pi)-2\lambda\bar u(\pi)\big\},

where uˉ(π)\bar u(\pi) is the discounted sum of u(s,a)u(s,a) under policy π\pi in the learned model. This is a conditional guarantee, not a promise that the practical Gaussian scale (standard-deviation) heuristic upper-bounds real model error.

The gap between the theorem and the algorithm is instructive. The theorem assumes an admissible bound on relevant dynamics error and an exact reward; the practical algorithm estimates a Gaussian scale score and a reward model from data. That heuristic may work empirically, but it does not inherit the theorem's guarantee without additional calibration and reward-error assumptions. Read the proof as a statement about what the penalty would need to control, not a certification of the fitted ensemble.

MOReL. The Model-based Offline Reinforcement Learning algorithm of Kidambi and colleagues takes a harder line. It partitions the state-action space into known and unknown using ensemble disagreement, and then constructs a pessimistic MDP in which unknown pairs lead to an absorbing terminal state carrying a penalty:

Pp(shalt∣s,a)=1for unknown (s,a),rp(shalt,a)=−κ.P_p(s_{\text{halt}}\mid s,a)=1\quad\text{for unknown }(s,a),\qquad r_p(s_{\text{halt}},a)=-\kappa.

The halt state is absorbing and carries the negative reward at every subsequent step; known pairs use the fitted transition. Thus its value under the displayed convention is −κ/(1−γ)-\kappa/(1-\gamma), not a one-time −κ-\kappa. This hard unknown-pair rule prevents a model rollout from inventing a favorable continuation after the halt, whereas a soft penalty allows it to continue. MOReL's algorithm optionally uses a behavior-cloned policy to initialize its planner; that is not a required behavior-cloning regularizer in its objective.

COMBO. Conservative Offline Model-Based Policy Optimization learns a Q-function from a mixture of real transitions and model rollouts. Its conservative term pushes down Q-values on state-action pairs sampled under model rollouts and up Q-values on logged state-action pairs. Thus it does not apply pessimism only to real data. Unlike MOPO's practical reward penalty, this regularizer acts on the value function and does not require an explicit model-uncertainty estimate.

The COMBO idea is subtly different from the others. It does not need to know where the model is uncertain; it can simply compare where the Q-function is being trained on real data versus rollout data, and it penalises the difference. This makes COMBO less dependent on the quality of an uncertainty estimator, but more dependent on having a good critic architecture and optimiser.

RAMBO. Instead of subtracting a hand-designed uncertainty score from reward, RAMBO (Rigter and colleagues) trains the model adversarially against the value function while also fitting the data. Its model objective still has a coefficient λ\lambda controlling the adversarial term, so pessimism has not become hyperparameter-free. It is a value-aware model-learning approach rather than an uncertainty-penalty approach. The trade-off is that the model is trained both for data fit and for appropriately pessimistic values, which adds training complexity.

The model-free cousins, briefly

It is worth knowing that the same ideas appear on the model-free side, often with different vocabulary.

  • BCQ (Batch-Constrained Q-learning) trains a generative model Gω(s)G_\omega(s) over the dataset's actions and restricts the policy to perturbing samples from it. The constraint is on the support of the action distribution.
  • BEAR applies an MMD constraint between the learned policy's action distribution and the behaviour policy's during actor policy improvement; it does not add MMD to the critic's objective.
  • TD3+BC combines Q-maximisation with a plain action-regression term in the actor objective: max⁡π E(s,a)∼D[λQ(s,π(s))−∥π(s)−a∥2]\max_\pi\,\mathbb E_{(s,a)\sim\mathcal D}[\lambda Q(s,\pi(s))-\|\pi(s)-a\|^2], with λ=α/E(s,a)∼D∣Q(s,a)∣\lambda=\alpha/\mathbb E_{(s,a)\sim\mathcal D}|Q(s,a)|. Thus α\alpha scales the Q term through λ\lambda, not the regression term. It is startlingly simple and competitive.
  • AWAC and AWR perform advantage-weighted behaviour cloning: the policy update is a weighted maximum-likelihood fit where the weight is exp⁡(A(s,a)/β)\exp(A(s,a)/\beta). Actions with low advantage get near-zero weight, so the actor update moves toward better logged actions. That statement does not guarantee that every critic backup, including AWAC's, avoids querying actions outside the dataset.
  • IQL removes the OOD action query entirely by estimating the value function with expectile regression back-ups, which never need max⁡aQ(s,a)\max_a Q(s,a). It is arguably the cleanest statement of the underlying principle: do not ask the model a question the data cannot answer.

All of these address the same distribution-shift problem from different sides: pessimism discounts uncertain estimates, while a policy constraint limits which estimates are queried. They are complementary design choices, not generally a formal dual pair. The model-free family does not require a dynamics model, yet it still faces coverage problems when a critic is asked to value unsupported actions.

Implementing the penalty

The implementation is one line. Everything visible in the previous section came from adding - PENALTY * u_hat to the reward.

The choice of λ\lambda is the interesting part. Below some threshold, the penalty is too weak to overcome the model's optimism and the planner still walks off the data. Above the sampled threshold in this toy grid, the greedy policy stays on the corridor and reaches the goal; this is not a general safety guarantee. Push it much higher and you do not get a better policy, only a more pessimistic value estimate, which matters when you use that estimate to choose between candidate policies rather than to act. This is why tuning λ\lambda is subtler than it first appears: the optimal λ\lambda for shaping a policy may be different from the optimal λ\lambda for evaluating one.

To see the threshold, sweep the positive values of λ\lambda across about 3.3 orders of magnitude (with zero as an additional baseline), solve the penalised model at each value, extract the greedy policy, and evaluate that policy exactly in the true grid.

In[23]:
Code
PENALTY_GRID = np.array(
    [0.0, 0.005, 0.01, 0.02, 0.03, 0.05, 0.1, 0.5, 2.0, 10.0]
)
sweep_planner_score = np.zeros(len(PENALTY_GRID))
sweep_true = np.zeros(len(PENALTY_GRID))

for index, penalty in enumerate(PENALTY_GRID):
    V_p, Q_p = value_iteration(P_hat, R_hat - penalty * u_hat)
    sweep_planner_score[index] = V_p[START]
    sweep_true[index] = evaluate_in_true_mdp(Q_p.argmax(axis=1))[START]
Out[24]:
Visualization
Line chart over categorically spaced sampled penalty coefficients: the penalised planner score falls, while realised return jumps from zero to a plateau.
Penalised model-planning score and realised return from the start state as the positive pessimism coefficients span about 3.3 orders of magnitude, with zero as a baseline. Equal x-axis spacing denotes sampled coefficients, not linear or logarithmic spacing. The realised return jumps from zero to the optimum once the corridor beats the imagined off-support route; the planner score keeps falling as extra penalty depresses the objective without changing nonterminal behavior or realised return.

At λ=0\lambda = 0 the realised return is zero and the unpenalised model score is comfortably positive: the model is confidently wrong. Once the penalty makes the observed corridor more attractive than the imagined route, the realised return reaches the optimum and stays there. The penalised planner score keeps falling because model trajectories pay an additional tax. At the right of the plot that planning objective is negative while the policy remains optimal; it is not an unpenalised off-policy estimate of that policy's return.

That divergence between the penalised score and the policy is one warning sign of over-conservatism. It may not affect action quality in this grid, but it matters if penalised scores are compared across candidate policies or penalty settings. The lesson is not "pick the largest λ\lambda that keeps the policy optimal." The lesson is to distinguish a planning objective from a policy-value estimate and choose the penalty for the decision it must support.

Out[25]:
Visualization
Grouped bar chart comparing planner objective scores and realised returns for naive and pessimistic planners.
Comparison of each planner's optimisation score with the realised return of its selected policy. The naive planner scores 3.08 but earns zero in the true grid. The pessimistic planner's penalised score is 0.509 and its policy earns 0.558; its unpenalised learned-model policy evaluation is also 0.558 (not plotted), so the lower bar reflects the penalty rather than model-evaluation bias.

Over-conservatism and how it shows up

Over-conservatism deserves more than a sentence, because it is the failure mode practitioners hit most often and it does not look like failure.

In the clean gridworld above, over-conservatism is invisible in the realised return. The corridor is the only route that pays anything, so increasing the penalty leaves nonterminal behavior and realised return unchanged in this example. In a realistic environment the pattern can differ. A dataset might cover a mediocre but safe region densely and a rewarding but thinly covered region sparsely. As λ\lambda grows, the latter pays a larger penalty even if the model's uncertainty score is bounded. The agent can retreat to the safe region and produce a stable-looking policy that misses attainable reward.

The mechanism is a tax that rises as logged coverage falls, regardless of the region's actual reward. Well-sampled regions pay almost nothing; thinly sampled regions pay much more. This biases the policy toward the well-sampled region even when the well-sampled region is not where the reward lives. The result is not a broken policy. It is a policy that looks reasonable, that runs cleanly, and that quietly underperforms.

There are three symptoms worth watching for:

  • Return collapse relative to behaviour cloning. A policy that fails to beat a behaviour-cloning baseline may be over-constrained or over-penalised, though weak coverage, model error, or optimisation failure can produce the same symptom. The baseline is a useful diagnostic, not proof of the cause.
  • Sensitivity of the chosen policy to λ\lambda across the middle of the range. If the argmax policy changes shape when λ\lambda moves from 0.3 to 0.5, the penalty is not just taxing the estimate, it is deciding the policy. That is fragile.
  • A large gap between penalised and unpenalised model scores. In this fixed model, the optimal penalised score cannot increase as λ\lambda rises. A negative penalised score shows that the objective is heavily taxed; it does not imply the task's true return is negative or that the model believes the task is worthless.

The remedy depends on the cause. Reducing λ\lambda may recover a good policy but can also reintroduce model exploitation. Better data about the promising region, where collection is allowed, can reduce uncertainty; longer training on the same unsupported data cannot manufacture that missing evidence. A better-calibrated local penalty or constraint can distinguish well-supported and thinly supported regions. MOPO's penalty varies with (s,a)(s,a); MOReL halts on pairs classified as unknown. Neither automatically calibrates uncertainty to true model error.

An optimiser can seek out overestimated actions, so upper-tail value errors are especially dangerous. Underestimation also matters: it can discard a good policy or prevent improvement beyond the behaviour policy. Conservative methods trade some of that opportunity for protection against unsupported optimistic choices. Whether the trade is appropriate depends on coverage and the decision at hand.

Policy Constraints and Offline Evaluation

Pessimism and constraints are two responses to distribution shift, often used together rather than as formal duals. This section covers the constraint family, then turns to a separate problem: how can you judge a learned policy without running it?

Support constraints as an alternative to pessimism

Pessimism reduces estimated values where a penalty signals model risk, seeking to discourage unsupported choices. A support constraint instead restricts the policies considered. Neither automatically ensures that the resulting policy stays near the data: a heuristic penalty can miss model error, and an approximate constraint can admit poorly covered actions. They act on different parts of the optimization problem, and their effectiveness depends on their assumptions and implementation.

A support constraint restricts π\pi to the set of policies whose state-action distribution is covered by the dataset:

π∈Πsupp={π:dπ(s,a)>0⇒ρβ(s,a)>0}.\pi \in \Pi_{\text{supp}} = \Big\{ \pi : d^\pi(s,a) > 0 \Rightarrow \rho_\beta(s,a) > 0 \Big\}.

This is an idealized constraint on the full state-action occupancy, including states reached by the target policy. BCQ's conditional generative model and bounded action perturbation approximate the behavior-action support at queried states; they do not enforce this global occupancy condition, especially at unseen states. Under a continuous-density behaviour distribution, each individual sampled point has probability zero even when the distribution has full-dimensional support. That conditional support is hard to infer from a finite log, so practical methods use divergences or generative action models as approximations.

A behaviour constraint is softer: rather than requiring full support, it penalises divergence from the behaviour policy:

max⁡π  Es∼D[Ea∼π(⋅∣s)Q(s,a)  −  β D(π(⋅∣s) ∥ πβ(⋅∣s))],\max_\pi \; \mathbb{E}_{s\sim\mathcal{D}}\Big[\mathbb{E}_{a\sim\pi(\cdot\mid s)}Q(s,a) \;-\; \beta\,D\big(\pi(\cdot\mid s)\,\|\,\pi_\beta(\cdot\mid s)\big)\Big],

where DD denotes a discrepancy or penalty, such as KL, MMD, or squared action distance. BEAR constrains its actor with MMD, while TD3+BC uses an action-regression term. AWAC instead performs advantage-weighted behaviour cloning; its update is not literally this displayed divergence objective. The choice matters: KL is asymmetric, MMD compares distributions through a kernel, and squared distance ignores much of the action distribution's structure.

These mechanisms can be complementary, but no general equivalence follows. A support constraint may still allow optimistic errors within its permitted set; a pessimistic critic may still be poorly calibrated where data are scarce. CQL applies conservatism to critic values, MOPO penalises rewards in model rollouts, and MOReL changes unknown transitions to lead to HALT. BEAR and TD3+BC illustrate direct actor restrictions. Specific implementations may combine mechanisms, but the papers do not establish a universal constraint–penalty duality.

Constraint strength is a hyperparameter too

An exact support constraint is binary, but practical approximations usually use a divergence, a coefficient β\beta, or a generative action model. A very strong restriction can leave little room for policy improvement; a weak one can admit unsupported actions and critic overestimation. This resembles the conservatism trade-off in a penalty sweep, although constraint and penalty coefficients have different meanings and need not produce the same policies.

Evaluation is the unsolved part

Here is the part of offline RL that practitioners consistently underestimate. You have produced a policy. You cannot run it. You must decide whether to deploy it, and if you have several candidates, which one to pick.

The candidates are usually not just policies. They are (policy, model architecture, ensemble size, rollout length, penalty coefficient, learning rate) tuples, and the number of combinations can be large. Selecting among them is offline model selection, a separate problem that can be difficult with only logged data. A model-based estimate of a policy's value can inherit errors from the model used for learning. If that model overestimates a region, it may also overestimate policies that visit the region and favor them in selection. This is a related form of model exploitation on the evaluation side.

Three families of estimators exist.

Importance sampling and its variants. Given mm independent trajectories from a known behaviour policy πβ\pi_\beta, reweight each discounted trajectory return Gi=∑t=0Ti−1γtri,tG_i=\sum_{t=0}^{T_i-1}\gamma^t r_{i,t} by the ratio of action probabilities:

J^IS(π)=1m∑i=1m(∏t=0Ti−1π(ai,t∣si,t)πβ(ai,t∣si,t))Gi.\hat J_{IS}(\pi) = \frac{1}{m}\sum_{i=1}^{m} \left(\prod_{t=0}^{T_i-1} \frac{\pi(a_{i,t}\mid s_{i,t})}{\pi_\beta(a_{i,t}\mid s_{i,t})}\right) G_i .

Under matching initial-state and environment distributions, known behaviour propensities, target-policy support within behaviour support, and finite expectation, ordinary trajectory IS is unbiased for the return represented by each complete sampled trajectory. For the infinite-horizon J(π)J(\pi) defined above, the recorded trajectory must include all nonzero rewards (for example, by ending in a zero-reward absorbing state) or use a valid tail correction; otherwise this formula estimates a truncated return. Its variance can grow exponentially with horizon when the policies differ enough, but this is not inevitable. Per-decision IS (PDIS) applies prefix ratios to individual rewards and is unbiased for the same return target under the corresponding support, sampling, and trajectory-completeness conditions; weighted/self-normalised IS trades finite-sample bias for possible variance reduction. Doubly robust estimators combine an approximate value function with importance weights. Long horizons and continuous actions can make ratio-based estimators impractical when overlap is weak, but their suitability depends on the policies and data rather than horizon alone.

Fitted Q-evaluation. Learn QπQ^\pi by iterating fitted Bellman back-ups on the offline dataset, using actions from π\pi at the successor state rather than an argmax. The resulting value estimate depends on approximation, optimisation, and evaluation-policy coverage; a policy produced by an offline algorithm is not by construction covered by the data. FQE avoids multiplying whole-trajectory probability ratios, but fitting a critic can be computationally more expensive than computing IS weights. It is a useful OPE baseline when its assumptions and implementation are stated, not an automatic certificate for every D4RL or RL Unplugged task.

Model-based evaluation. Evaluate a fixed candidate policy in the learned model. Holding the policy fixed removes an additional maximisation within that evaluation calculation, but it does not erase how the policy was selected. A policy optimised on the same model may already exploit its weak spots, and choosing the highest of many model-based estimates can do so again. Model-based evaluation is useful when model accuracy is supported along that policy's occupancy, with uncertainty and coverage reported alongside the estimate.

It also has an important failure mode, and it is the same one. The evaluator shares the model's blind spots. If the model has a large optimistic region and the candidate policy goes there, the evaluator reports a high return for exactly the policy you should not deploy. Techniques that help include averaging over the ensemble, reporting the worst-case estimate across members, and using the uncertainty map to flag evaluations whose trajectories enter high-uu regions.

In[26]:
Code
# Unpenalised model-based evaluation of each fixed policy, distinct from the
# penalised objective used to select the pessimistic policy.
mb_estimate_naive = V_model_naive_policy[START]
mb_estimate_pess = V_model_pess_policy[START]
pessimistic_planner_score = V_pess[START]


# The fraction of each candidate's model rollout that visits unobserved states,
# computed from the occupancy implied by rolling the policy out in the model.
def model_occupancy(policy, sweeps=4000, tolerance=1e-12):
    occupancy = np.zeros(N_STATES)
    state_distribution = np.zeros(N_STATES)
    state_distribution[START] = 1.0
    for _ in range(sweeps):
        occupancy += state_distribution
        state_distribution = GAMMA * (
            P_hat[np.arange(N_STATES), policy].T @ state_distribution
        )
        if state_distribution.sum() < tolerance:
            break
    return occupancy / occupancy.sum()


occupancy_naive = model_occupancy(pi_naive)
occupancy_pess = model_occupancy(pi_pess)
unsupported_naive = float(occupancy_naive[(counts.sum(axis=1) == 0)].sum())
unsupported_pess = float(occupancy_pess[(counts.sum(axis=1) == 0)].sum())
Out[27]:
Console
model-based evaluation
  naive candidate       : estimated +3.082
  pessimistic candidate : estimated +0.558
  pessimistic planner score (with penalty): +0.509

discounted occupancy of unvisited states under each candidate
  naive candidate       :  45.2%
  pessimistic candidate :   0.0%

The discounted unvisited-state occupancy diagnostic shows that 45.2% of the naive candidate's model-rollout mass lies in cells absent from the dataset, versus 0.0% for the pessimistic candidate. The unpenalised learned-model evaluation of the pessimistic policy is about 0.558, while its penalised planner score is about 0.509. The naive policy's model estimate relies partly on unobserved dynamics and should not be read as a guarantee about its real return.

This is the practical lesson for offline evaluation: pair estimates with coverage diagnostics and sensitivity checks. This state-only fraction flags the naive policy's exposure to unvisited cells, but misses unsupported actions within visited states. Inspect discounted state-action occupancy or logged pair counts as well; neither check by itself certifies model accuracy or the value estimate.

Benchmarks and protocols

The standard offline RL benchmarks are:

  • D4RL. The benchmark includes MuJoCo locomotion, AntMaze, Franka kitchen, and Adroit. Locomotion datasets use labels such as random, medium, medium-replay, medium-expert, and expert; those labels do not apply uniformly to the other task families. Medium-quality locomotion data provide a useful improvement-from-imperfect-behaviour test, but no one quality label is universally the hardest across tasks and methods.
  • RL Unplugged. A suite spanning control and games, including continuous-control and Atari-style settings, with an emphasis on realistic logged data collection.
  • NeoRL. Datasets collected by policies of controlled quality, designed for studying near-realistic offline settings where the behaviour policy is neither random nor expert.

Two protocol points matter more than the choice of benchmark. First, report the benchmark's prescribed score normalization where it is defined; D4RL commonly normalizes scores against reference random and expert returns so raw returns across tasks are more comparable. Second, report the variation across seeds and the aggregation statistic. Mean and median can differ substantially, so a ranking that changes with aggregation deserves scrutiny.

For rollout-based methods such as MOPO and COMBO, one common protocol is to fix the dataset, fit a model per seed, generate short model rollouts, train a policy or critic using real and imagined data, and choose the final policy with an offline criterion. MOReL instead plans in its pessimistic learned MDP, and Trajectory Transformer plans over predicted sequences; the rollout-mixture recipe does not describe every model-based method. In rollout methods, horizon is a consequential hyperparameter: a short rollout limits direct propagation of model error but leaves more work to the critic, while a long rollout can accumulate unsupported predictions. Policy selection needs its own reported criterion because no online return is available to settle competing settings.

Limitations and Impact

Pessimism can reduce an optimiser's incentive to exploit model error where the penalty captures the relevant risk. The gridworld made the mechanism visible: the same fitted model, dataset, and value iteration produced a zero-return policy without a penalty and an optimal policy with a reward penalty of half a unit per under-supported transition. That is an instructive example, not a universal guarantee that pessimism resolves every ill-posed offline decision problem.

Pessimism also does not manufacture information about unvisited actions. If the log contains no evidence about an action and no justified structural assumption links it to observed actions, two environments can agree on every logged transition yet require different decisions there. That rules out a uniform worst-case guarantee from the log alone. Known physics, smoothness, or transferable pretraining may provide additional information, but their assumptions and target-domain validity must be made explicit. A warehouse fleet that has only driven down the middle of the aisles has no logged outcomes from the edges. A conservative policy may avoid an edge route that is actually useful; its caution is a response to uncertainty, not proof the route is bad. Where feasible, collect relevant data; otherwise state what prior structure supports any extrapolation and how it was checked.

The practical costs are real. Pessimism introduces at least one task-dependent hyperparameter. In our gridworld policy quality changes abruptly across a threshold in λ\lambda; real problems can change less cleanly, and penalties interact with rollout length, model capacity, ensemble design, and reward scale. Training multiple models increases modelling cost. Because the dataset is fixed, additional fitting or ensemble members cannot by themselves supply transition evidence in an unsupported region; new data would be required to resolve that uncertainty empirically.

There is also a subtler limitation: ensemble disagreement need not reflect shared model bias. Members trained on the same data and similar architectures may agree on a wrong extrapolation, leaving a small disagreement score where the actual transition error is large. Adding similar members does not reliably resolve that missing-data problem. Broader pretraining may help in some settings, a direction discussed in Part IX: Foundation and World-Action Models, but its uncertainty must still be checked against the target domain.

Under-dispersion limits a pessimistic method that relies on that uncertainty score, not offline RL as a whole. If the score is systematically small where model error matters, the penalty is small there too, so it cannot be treated as a certified error bound. More varied prior data or model assumptions might help, but neither certifies calibrated uncertainty without target-domain checks.

The offline setting forces coverage, model validity, and estimation bias into the open because there is no new interaction to check a candidate policy during learning. Online feedback can sometimes expose errors, but it does not automatically rescue a poor policy or make exploration safe. The same discipline matters beyond offline RL: in the evaluation protocols of Part XI: Evaluation and Understanding, in the safety arguments of Part XII: Reliable World Models, and wherever a learned model informs a decision before its consequences can be checked.

Summary

During its offline training phase, offline RL learns from a fixed dataset without further interaction. The removal of the training-time feedback loop creates two coupled problems: the log supplies no direct target where it is silent, and the optimiser can favour those places when unsupported model estimates are optimistic. Justified structural assumptions may narrow the uncertainty, but the log alone cannot test those predictions there.

  • Coverage or justified structure is fundamental. Without sufficient logged coverage or informative assumptions that constrain unobserved outcomes, uniform worst-case improvement is impossible. A finite concentrability coefficient formalizes one form of coverage; different algorithms use different structural and statistical assumptions.
  • The simulation lemma explains fixed-policy error. The value gap is an average of local errors weighted by the policy's learned-model rollout occupancy, which can differ from its true occupancy. Policy optimization can then favor optimistic model estimates and change where that learned-model occupancy lies. Error propagation and policy selection can reinforce each other, but the fixed-policy identity alone does not prove exploitation.
  • Worst-case bounds can be vacuous. In the illustrative example with γ=0.99\gamma=0.99 and a uniform 0.05 transition-TV bound, the transition-error term is 990 and the full bound, including ϵr=0.01\epsilon_r=0.01, is 991. Both exceed a range-based upper bound on the possible value difference under the stated reward range. Useful guarantees need stronger information about where the policy goes and where the model is accurate.
  • Uncertainty scores need interpretation. Counts and ensemble disagreement can flag missing evidence, but predictive variance may also include irreducible randomness. Shared model biases can make an ensemble agree while all its members are wrong.
  • Pessimism can lower-bound value under conditions. A penalty must cover possible model overestimation at relevant pairs; a heuristic score alone is not a certificate. Too little penalty permits exploitation, while too much can discard useful policies.
  • Constraints are complementary. Support constraints and behaviour regularisers restrict where a policy may go; pessimism reduces the estimated worth of uncertain choices. Some methods combine these mechanisms, but they are not generally formal duals.
  • Evaluation remains difficult. Trajectory and per-decision importance sampling can be unbiased under support, sampling, and complete-return or valid tail-correction conditions, but may have severe variance under policy mismatch. Weighted IS can introduce finite-sample bias; fitted Q-evaluation and model rollouts introduce approximation or model bias. Coverage diagnostics help identify risk but do not certify an estimate.
  • Model selection is a separate problem. Choosing among candidate policies and hyperparameters using offline data only needs its own evaluation and uncertainty checks; a favorable learned-model estimate is not enough.

The chapter's gridworld compressed the argument into a few hundred lines of NumPy: one row of a seven-row grid, one behaviour policy that never left it, and an additive reward penalty that changed the selected policy from zero return to the true optimum. The general mechanism is model exploitation: an optimiser can favor actions where estimates are unsupported. Beyond the toy grid, pessimism helps only to the extent that its uncertainty score tracks the relevant errors.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about offline and conservative model-based reinforcement learning.

Offline and Conservative Model-Based RL

Question 1 of 80 of 8 completed
What is the key distinction between offline RL and off-policy RL?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026offlineconservative, author = {Michael Brenndoerfer}, title = {Offline and Conservative Model-Based RL}, year = {2026}, url = {https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Offline and Conservative Model-Based RL. Retrieved from https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning
MLAAcademic
Michael Brenndoerfer. "Offline and Conservative Model-Based RL." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning>.
CHICAGOAcademic
Michael Brenndoerfer. "Offline and Conservative Model-Based RL." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Offline and Conservative Model-Based RL'. Available at: https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Offline and Conservative Model-Based RL. https://mbrenndoerfer.com/writing/offline-conservative-model-based-reinforcement-learning

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.