Perceptual and Generative Evaluation

Michael BrenndoerferAugust 1, 202662 min read

Part of World Models Handbook

Score predicted observations with MSE, PSNR, SSIM, LPIPS, FID, and FVD. Protocols, conditioning, best-of-K, and human judgment shape the results.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Perceptual and Generative Evaluation

An action can have more than one plausible future. In the small example below, a blob moves either a short or a long distance to the right. The realized future takes the short step. A sharp prediction of the long step puts the blob in the wrong place; averaging the two possible frames produces a faded, two-blob image with lower squared pixel error than that wrong sharp prediction. Yet the sharp prediction matching the realized short step has zero error and beats both. Which comparison should an evaluation report?

The example separates three questions: does an output look plausible, does it match a particular realized future, and does the predictive distribution match the possible futures and their probabilities? These properties can disagree. A score may provide evidence about more than one property, but no single score discussed here certifies all three. This chapter explains what each score measures and what remains unchecked.

We will compute the pixel scores, but we will not run a human-preference study or a learned perceptual network on this toy. Its faded average illustrates why an expected-error optimum need not resemble a sample. Whether people prefer either output requires a separate experiment. The useful task is to choose scores whose questions match the intended use.

This chapter sits inside Part XI: Evaluation and Understanding, after The World-Model Evaluation Ladder. That ladder distinguishes one-step prediction, fixed-action open-loop rollouts, interventions, closed-loop decisions, and transfer. Here we focus on scoring observations under declared prediction protocols. Observation scores can reveal useful failures, but they cannot by themselves certify state, physics, or causal fidelity. The next chapter develops those diagnostics.

The scope is observation evaluation, not a certificate of useful internal state or decision quality. Pixels, feature distances, likelihoods, and human judgments provide different kinds of evidence. A table of those scores becomes interpretable only when the evaluation protocol is stated.

We begin with a synthetic scene and score hand-designed predictions against its realized future. The example requires neither checkpoints nor network training. We then distinguish the questions answered by MSE, PSNR, and SSIM from those answered by LPIPS, FID, and FVD. For stochastic futures, we examine conditional scoring, future-reference best-of-KK selection, and an action reversal that pooling conceals. Human judgment protocols and learned evaluators complete the scorecard.

A 48 by 48 grayscale canvas lets us inspect the pixels and recompute the scores. Its two equally probable futures illustrate the difference between matching one realization and describing a distribution. The example does not reproduce the architecture, data, or failure modes of a large video model.

As in Chapter 61, keep three prediction regimes separate in your head:

  • Teacher-forced one-step prediction: the model receives the realized observation or history and the action needed to predict the next observation. Errors from its earlier predictions are not fed back as context.
  • Fixed-action open-loop rollout: the model receives hth_t and actions at:t+H−1a_{t:t+H-1} and predicts ot+1:t+Ho_{t+1:t+H} without seeing intermediate realized observations. Earlier prediction errors can affect later predictions, although error need not increase monotonically.
  • Closed-loop decision: an agent uses the model inside a planner or policy and receives environment feedback. Task return evaluates decision usefulness; image scores remain diagnostic. Planning, Control, and Policy Evaluation covers that regime.

An MSE of 0.003 needs context: which regime produced it, how many steps were rolled out, and what was averaged? Changing the regime can change the score substantially even with the same model. Keep the regimes separate when reading a prediction table.

The following curves are prescribed illustrations, not measurements from a trained predictor. We construct a gently increasing teacher-forced curve and an exponentially increasing open-loop curve to illustrate a possible difference between protocols. The teacher-forced curve is not constant, and neither shape is a universal law.

In[3]:
Code
import numpy as np

horizon_steps = np.arange(1, 17)
teacher_forced_mse = 0.004 + 0.0015 * horizon_steps
open_loop_mse = 0.004 * np.exp(0.22 * horizon_steps)
Out[4]:
Visualization
Prescribed linear and exponential illustrative MSE curves over sixteen steps.
Illustrative error curves defined directly by two analytic expressions: a linear teacher-forced curve and an exponential open-loop curve. They are not measured model errors and do not establish that open-loop error must grow exponentially or exceed one-step error at every horizon.

Everything we build below uses the convention pθ(ot+1:t+H∣ht,at:t+H−1)p_\theta(o_{t+1:t+H} \mid h_t, a_{t:t+H-1}), where:

  • ot+1:t+Ho_{t+1:t+H}: the sequence of future observations the model predicts, from time t+1t+1 through time t+Ht+H
  • ht=(o≤t,a<t)h_t = (o_{\le t}, a_{<t}): the history, consisting of observations up to and including time tt and actions strictly before time tt
  • at:t+H−1a_{t:t+H-1}: the sequence of actions from time tt to time t+H−1t+H-1, one action per predicted transition
  • HH: the prediction horizon, i.e. the number of future frames or steps predicted
  • θ\theta: the model parameters

History can supply predictive information missing from a single frame in partially observed settings. It is not, by itself, a definition of a world model or a guarantee that the history is sufficient. Supplying actions lets us ask conditional prediction questions; observational action conditioning alone does not identify the effects of causal interventions.

Use HimgH_{\mathrm{img}} for image height and reserve HH for the prediction horizon. Grayscale frames have shape (Himg,W)(H_{\mathrm{img}},W) and color frames (Himg,W,C)(H_{\mathrm{img}},W,C), with values in [0,1][0,1]. Here WW is width and CC is the channel count. Other layouts, such as channels-first, change the axes to average but not the metric's meaning. A rollout adds a time dimension whose length is HH.

Reconstruction and Predictive Accuracy

Reconstruction scoring and future-prediction scoring agree on the math and disagree on the story. In reconstruction the model sees oto_t, produces o^t\hat o_t, and we compare o^t\hat o_t to oto_t. In future prediction the model sees hth_t and at:t+H−1a_{t:t+H-1} and produces o^t+1:t+H\hat o_{t+1:t+H}, and we compare the predicted sequence to the realized sequence ot+1:t+Ho_{t+1:t+H}. Pixel-wise mean squared error (MSE) and peak signal-to-noise ratio (PSNR) are the two workhorses of both. Frame MSE averages squared pixel residuals; frame PSNR applies a logarithm to that spatially averaged MSE. For sequences, averaging frame PSNR is not the same as applying the logarithm to a pooled MSE. The averaging conventions matter and are the subject of the next formulas.

The reference changes meaning between tasks. Ordinary reconstruction compares against the fixed input observation; zero pixel MSE requires reproducing that array, but reconstruction quality alone does not establish useful internal state. Future prediction compares against a realized outcome, potentially one of several possible futures. The same distance therefore has a different interpretation.

The mean squared error (MSE) between a prediction x^\hat x and a target xx, both of shape (Himg,W)(H_{\mathrm{img}}, W) with pixel values in [0,1][0,1], is

MSE(x^,x)  =  1HimgW∑i=1Himg∑j=1W(x^ij−xij)2\mathrm{MSE}(\hat x, x) \;=\; \frac{1}{H_{\mathrm{img}} W} \sum_{i=1}^{H_{\mathrm{img}}} \sum_{j=1}^{W} \bigl(\hat x_{ij} - x_{ij}\bigr)^2

where:

  • x^ij\hat x_{ij}: the predicted pixel value at spatial position (i,j)(i, j)
  • xijx_{ij}: the target pixel value at the same position
  • Himg,WH_{\mathrm{img}}, W: the height and width of the frame in pixels
  • 1HimgW\frac{1}{H_{\mathrm{img}}W}: normalization by the pixel count, making the result an average rather than a sum

The subtraction x^ij−xij\hat x_{ij} - x_{ij} is the signed per-pixel error. Squaring removes its sign and heavily penalizes large deviations.

The square grows quadratically with error magnitude. Uniform error of 0.1 gives MSE 0.01. Error of 0.3 on 10 percent of pixels, with zero error elsewhere, gives MSE 0.09×0.1=0.0090.09\times0.1=0.009. These close averages do not describe where the errors occur. The perceptual importance of a localized error depends on the scene and viewing task; we have not measured it in this calculation.

The quadratic penalty is easiest to see as a curve. The same squared-error contribution can come from a small error everywhere or a larger error in a small region.

In[5]:
Code
import numpy as np

error_magnitude = np.linspace(0.0, 0.35, 200)
squared_error = error_magnitude**2
uniform_error = 0.1
localized_error = 0.3
uniform_contribution = uniform_error**2
localized_contribution = localized_error**2
# The prose claim is that a localized 0.3 error on 10% of pixels scores comparably to
# a uniform 0.1 error on all pixels: 0.09 * 0.1 = 0.009 vs 0.01. Verify the contest is
# actually close enough that the plot supports the prose rather than showing a trivial gap.
uniform_total = uniform_contribution
localized_total = localized_contribution * 0.1
assert abs(uniform_total - localized_total) / uniform_total < 0.2
Out[6]:
Visualization
Line plot of squared error versus error magnitude with uniform and localized examples.
Squared error versus per-pixel error magnitude. A 0.1 error contributes 0.01 per pixel; a 0.3 error contributes 0.09 per affected pixel. If only 10 percent of pixels have the latter error, total MSE is 0.009. The curve shows the arithmetic penalty, not a human-visibility experiment.

For a multi-step prediction, we first compute the per-frame MSE for each horizon step and then average those per-frame values over time. The sequence MSE is

MSEseq(o^t+1:t+H,ot+1:t+H)  =  1H∑k=1HMSE(o^t+k,ot+k)\mathrm{MSE}_{\text{seq}}(\hat o_{t+1:t+H}, o_{t+1:t+H}) \;=\; \frac{1}{H} \sum_{k=1}^{H} \mathrm{MSE}\bigl(\hat o_{t+k}, o_{t+k}\bigr)

where:

  • o^t+k\hat o_{t+k}: the predicted frame at horizon step kk
  • ot+ko_{t+k}: the realized frame at the same horizon step
  • HH: the number of future frames being averaged
  • MSE(o^t+k,ot+k)\mathrm{MSE}(\hat o_{t+k}, o_{t+k}): the per-frame MSE defined above

Notice the two averaging decisions: how to combine spatial and channel dimensions into a per-frame value, and how to combine frames into a per-rollout value. Do not compare two numbers from different conventions. Always specify the pixel range (here in [0,1][0,1]) and the tensor layout, and state whether frames of different length are padded and excluded via a valid mask mk∈{0,1}m_k \in \{0, 1\}:

MSEseq  =  ∑k=1Hmk MSE(o^t+k,ot+k)∑k=1Hmk\mathrm{MSE}_{\text{seq}} \;=\; \frac{\sum_{k=1}^{H} m_k\, \mathrm{MSE}(\hat o_{t+k}, o_{t+k})}{\sum_{k=1}^{H} m_k}

where:

  • mk∈{0,1}m_k \in \{0, 1\}: a validity indicator for frame kk, where mk=1m_k = 1 includes the frame and mk=0m_k = 0 excludes it (e.g., padding)
  • HH: the prediction horizon, as before
  • ∑k=1Hmk\sum_{k=1}^{H} m_k: the valid-frame count, which must be positive. An all-padding sequence has no defined score under this formula and must be excluded or handled by a declared convention.

The valid mask is not a cosmetic detail. Suppose a model is trained to predict up to 64 frames, but two datasets provide rollouts of length 32 and 48 respectively. If the evaluation pads all rollouts to length 64 and averages the padded MSE without masking, the shorter dataset will be artificially penalized or rewarded depending on what value padded frames take. A mask removes this artifact by restricting the average to the frames the data actually contains. The same averaging convention question also arises when sequences are grouped into batches: the natural convention is to compute a per-sequence score and then average over sequences, rather than pooling all frames across the batch, because the latter weights long sequences more heavily.

Peak signal-to-noise ratio (PSNR) is MSE expressed on a logarithmic scale. It answers the question: how large is the squared peak signal relative to the squared reconstruction error? Given a declared peak value LL (usually the maximum representable pixel value, so L=1L = 1 for a [0,1][0, 1] range and L=255L = 255 for 8-bit images), PSNR is

PSNR(x^,x)  =  10log⁡10L2MSE(x^,x)\mathrm{PSNR}(\hat x, x) \;=\; 10 \log_{10} \frac{L^2}{\mathrm{MSE}(\hat x, x)}

where:

  • LL: the declared peak value of the pixel range, i.e. the maximum representable pixel value
  • MSE(x^,x)\mathrm{MSE}(\hat x, x): the pixel MSE between prediction and target, defined above
  • 10log⁡10(⋅)10 \log_{10}(\cdot): a logarithmic rescaling of the ratio into decibels (dB). At fixed LL, lower MSE means higher PSNR, so the preference ordering is equivalent with opposite score directions.

When MSE=0\mathrm{MSE} = 0, PSNR is +∞+\infty in the mathematical definition. If a report clips exact matches to a finite cap, it must state that choice. Declare LL and keep it consistent with the pixel units. Rescaling both images by 255 multiplies MSE by 2552255^2 and changes LL from 1 to 255, leaving PSNR unchanged. Changing only LL while keeping the same numerical errors adds 20log⁡10(255)≈48.1320\log_{10}(255) \approx 48.13 dB, an inconsistent protocol rather than a model improvement. At fixed LL, PSNR is a strictly convex function of positive MSE: a constant minus 10log⁡10(MSE)10\log_{10}(\mathrm{MSE}).

For positive frame errors e1,…,ene_1,\ldots,e_n and the same LL, Jensen's inequality gives mean frame PSNR greater than or equal to PSNR of mean MSE, with equality when all errors agree. Their gap is 10log⁡10(AM(e)/GM(e))10\log_{10}(\mathrm{AM}(e)/\mathrm{GM}(e)), the logarithm of the arithmetic-to-geometric mean ratio. It increases under a mean-preserving spread of the errors; variance alone does not determine it. State which averaging convention the table uses. Exact-zero errors need an explicit infinity or clipping convention.

The next bar chart applies both averaging conventions to the same eight positive frame errors. Its difference measures the averaging gap for this fixed set, not a comparison between two predictors.

In[7]:
Code
import numpy as np

frame_mse = np.array([0.002, 0.003, 0.004, 0.006, 0.010, 0.018, 0.030, 0.050])
mean_psnr = float(np.mean(10.0 * np.log10(1.0 / frame_mse)))
psnr_of_mean = float(10.0 * np.log10(1.0 / np.mean(frame_mse)))
Out[8]:
Visualization
Bar chart comparing mean of per-frame PSNR with PSNR computed from mean MSE.
Mean per-frame PSNR exceeds PSNR computed from mean MSE for these eight unequal positive errors. At a common peak value, their gap is the logarithm of the arithmetic-to-geometric mean ratio. It increases under a mean-preserving spread, not necessarily whenever variance increases.

Expected MSE is minimized by the conditional mean. Suppose yy is a real-valued target with finite second moment, xx contains the conditioning information, and f(x)f(x) is an unrestricted measurable predictor with finite squared error. The expectation is over the joint distribution p(x,y)p(x,y). The squared error decomposes into an irreducible term plus the squared gap between ff and the conditional mean:

E[(y−f(x))2]  =  E[(y−E[y∣x])2]  +  E[(E[y∣x]−f(x))2]\mathbb{E}\bigl[(y - f(x))^2\bigr] \;=\; \mathbb{E}\bigl[(y - \mathbb{E}[y \mid x])^2\bigr] \;+\; \mathbb{E}\bigl[(\mathbb{E}[y \mid x] - f(x))^2\bigr]

Expanding the square and conditioning on xx eliminates the cross term: the conditional mean of y−E[y∣x]y-\mathbb{E}[y\mid x] is zero almost surely. The irreducible first term does not depend on ff; the second is nonnegative and is zero exactly when f(x)=E[y∣x]f(x)=\mathbb{E}[y\mid x] almost surely. Thus the conditional mean minimizes expected squared error under the stated integrability assumptions.

For a finite-dimensional image vector, apply the same decomposition to each coordinate and average. The conditional mean still minimizes expected pixel MSE. If possible images place an object at separated locations, averaging their pixels can produce a faded overlay rather than an image from either mode. This is a possible consequence of the geometry, not a claim that every multimodal distribution has an implausible mean.

In our separated-object example, an MSE-optimal point estimate need not be a plausible sample. That is a mismatch between the prediction target and the sample-quality question, not a contradiction in the theorem. Blur can also reflect limited decoder resolution or capacity; the image alone does not identify its cause.

Structural similarity (SSIM), introduced by Wang, Bovik, Sheikh, and Simoncelli (2004), compares local luminance, contrast, and structure rather than uniformly weighting pixel-coordinate squared errors. For each sliding-window position, it summarizes corresponding patches by weighted means, variances, and covariance. With normalized Gaussian window weights ww, these statistics are defined as:

μx=∑iwixi,σx2=∑iwi(xi−μx)2,σxy=∑iwi(xi−μx)(yi−μy)\mu_x = \sum_i w_i x_i, \qquad \sigma_x^2 = \sum_i w_i (x_i - \mu_x)^2, \qquad \sigma_{xy} = \sum_i w_i (x_i - \mu_x)(y_i - \mu_y)

where:

  • wiw_i: the window weight at position ii (Gaussian or uniform), summing to 1
  • xi,yix_i, y_i: pixel values of the two images within the window
  • μx,μy\mu_x, \mu_y: the weighted local means
  • σx2,σy2\sigma_x^2, \sigma_y^2: the weighted local variances
  • σxy\sigma_{xy}: the weighted local covariance between the two patches, computed as ∑iwi(xi−μx)(yi−μy)\sum_i w_i (x_i - \mu_x)(y_i - \mu_y)

The SSIM score for that window is

SSIM(x,y)  =  (2μxμy+C1)(2σxy+C2)(μx2+μy2+C1)(σx2+σy2+C2)\mathrm{SSIM}(x, y) \;=\; \frac{(2 \mu_x \mu_y + C_1)(2 \sigma_{xy} + C_2)}{(\mu_x^2 + \mu_y^2 + C_1)(\sigma_x^2 + \sigma_y^2 + C_2)}

where:

  • (2μxμy+C1)(2 \mu_x \mu_y + C_1): the numerator of the luminance term, comparing the two local means
  • (μx2+μy2+C1)(\mu_x^2 + \mu_y^2 + C_1): the denominator of the luminance term, normalizing that comparison
  • (2σxy+C2)(2 \sigma_{xy} + C_2): the numerator of the combined contrast-and-structure term
  • (σx2+σy2+C2)(\sigma_x^2 + \sigma_y^2 + C_2): its denominator; together the factors compare covariance and local variance, not a squared correlation alone
  • C1=(K1L)2C_1 = (K_1 L)^2, C2=(K2L)2C_2 = (K_2 L)^2: small stabilizing constants using the reference defaults K1=0.01K_1 = 0.01 and K2=0.03K_2 = 0.03, where LL is the declared dynamic range

The first ratio measures luminance agreement and the second measures structural (contrast and correlation) agreement; the constants C1,C2C_1, C_2 keep the ratio finite when the local variances or means are near zero. The per-window scores are averaged over all window positions, giving the mean SSIM (MSSIM) that is reported for image-quality benchmarking.

The second ratio is not squared Pearson correlation. Without stabilizers and with positive local standard deviations, it equals the correlation multiplied by 2σxσy/(σx2+σy2)2\sigma_x\sigma_y/(\sigma_x^2+\sigma_y^2), a contrast-agreement factor. A positive gain can preserve correlation while reducing contrast agreement; an offset can reduce luminance agreement. SSIM does not ignore brightness shifts. Rescaling both images and the declared range LL together preserves this formula, but that change of units is not invariance to a physical gain or offset applied to just one image.

SSIM depends on the window size, whether the window is Gaussian or uniform, the standard deviation of the Gaussian, border-handling, whether the input is color or grayscale, and the declared LL. Two papers calling their metric "SSIM" may be computing different numbers. And, more importantly, SSIM is a proxy for human similarity judgments on natural images; it is not a semantic or physical truth classifier. A prediction with high SSIM can still place an object on the wrong side of a wall.

SSIM's agreement with human ratings was evaluated on specified image-distortion datasets. Agreement on a new domain or task is not guaranteed, but domain change does not necessarily reduce it. Validate the score against the distinctions that matter for the target task.

Full-reference vs. no-reference

MSE, PSNR, and SSIM are full-reference metrics: they require a specified, aligned reference. That reference need not be physical ground truth; an evaluation should state its source. For stochastic prediction, the realized future is one draw from the true conditional, so an alternative valid mode may score poorly against it. Comparing sample distributions changes the evaluation unit but introduces other limitations.

Consider a grayscale 48×4848 \times 48 canvas with a Gaussian blob of width σ=2.5\sigma=2.5 pixels. We prescribe two equally likely future locations under the same action, which could represent unobserved friction or contact variation. The code renders the two endpoints; it does not simulate that physical mechanism. The realized reference is the short-step endpoint. We compare a sharp match, a sharp alternative, and their pixel-wise mean using MSE and PSNR. At a fixed range on a single pair, PSNR gives exactly the reverse numerical ordering of MSE, so their preferences agree. We discuss SSIM separately rather than pretending to have computed it here.

The soft rendering makes the pixel-wise average visible as a faded superposition. Gaussian blobs have overlapping tails, so they are separated rather than strictly disjoint. This particular average is neither sharp endpoint; the conditional-mean theorem does not assert that every mean of separated modes lies in a low-density region.

In[9]:
Code
import matplotlib.pyplot as plt
import numpy as np

SIZE = 48
BLOB_SIGMA = 2.5


def render_object(x, y=None, sigma=BLOB_SIGMA, size=SIZE):
    """Render a soft Gaussian blob centered at (x, y) on a size x size canvas."""
    if y is None:
        y = size / 2.0
    yy, xx = np.mgrid[0:size, 0:size].astype(float)
    return np.exp(-(((xx - x) ** 2 + (yy - y) ** 2) / (2.0 * sigma**2)))


def mse(a, b):
    return float(np.mean((a - b) ** 2))


def psnr(a, b, peak=1.0):
    m = mse(a, b)
    if m == 0.0:
        return float("inf")
    return 10.0 * np.log10(peak**2 / m)

Now we build a single realized future and three candidate predictions. The realized future is mode A. The candidates are mode A (Sharp-A), mode B (Sharp-B), and the pixel-wise conditional mean. The mean under a two-mode conditional with equal weights is the average of the two mode images.

In[10]:
Code
MODE_A_X = 20.0  # small step to the right under action a_t = +1
MODE_B_X = 28.0  # large step to the right under the same action
REALIZED_X = MODE_A_X  # the future that actually happened

observed = render_object(REALIZED_X)
pred_sharp_A = render_object(MODE_A_X)
pred_sharp_B = render_object(MODE_B_X)
pred_mean = 0.5 * (render_object(MODE_A_X) + render_object(MODE_B_X))

assert observed.shape == (SIZE, SIZE)
assert observed.min() >= 0.0 and observed.max() <= 1.0

Now score the predictions against the realized future. The mean averages two separated blobs with overlapping tails. Its maximum intensity is approximately 0.503 in this discretized example, close to half the peak of either endpoint, so it looks faded.

Out[11]:
Console
Sharp-A  MSE=0.000000  PSNR=inf
Sharp-B  MSE=0.015727  PSNR=18.03 dB
Mean     MSE=0.003932  PSNR=24.05 dB
Out[12]:
Visualization
Realized grayscale blob at horizontal position 20.
The realized future contains the blob at mode A.
Sharp-A grayscale prediction at position 20, matching the reference.
Sharp-A matches the realized mode exactly.
Sharp-B grayscale prediction at position 28, the alternative mode.
Sharp-B matches the other plausible mode, not this realization.
Two faded grayscale blobs formed by averaging both mode images.
The conditional mean overlays both modes at reduced intensity.
Out[13]:
Visualization
Bar chart of MSE for three predictions, ranked Sharp-A, Mean, Sharp-B.
Paired pixel MSE for the three predictions against the realized future. Sharp-A wins trivially because it happens to match the realized mode, Sharp-B is heavily penalized even though the data generator itself can produce it, and the conditional mean sits between them. PSNR gives the same ranking on a log scale.

The bar chart makes the first-class concern of this chapter concrete. Pixel error rewards the model that got lucky on this particular draw, not the model that captured the conditional. Sharp-B is a mode the ground-truth process itself can produce; if the realized future had been Sharp-B instead, the ranking would flip. And the conditional mean, which is MSE-optimal on average, looks neither like Sharp-A nor Sharp-B in this particular draw.

One paired score evaluates one prediction against one realized future. Repeated paired MSE estimates expected squared error, but even that expectation does not characterize the full conditional distribution: the conditional mean is optimal without representing its modes or probabilities. Distributional tests with explicit conditioning or proper probabilistic scores answer a different question.

MSE remains useful for measuring paired accuracy. To assess the forecast distribution, use a declared probabilistic scoring rule or conditional distributional diagnostic over enough held-out cases. Merely averaging more paired MSE values does not turn MSE into a complete distributional test.

Perceptual Quality and Distributional Metrics

Human observers are not pixel sensors. Two images that differ in MSE by a factor of four can look nearly identical, and two that agree numerically can look very different. Learned perceptual metrics were introduced to close that gap, and distributional metrics were introduced to handle the fact that a stochastic future has no single correct answer. The next two subsections define each family precisely and state the assumptions under which the scores are meaningful.

Perceptual distances such as LPIPS still compare a prediction with a reference, using a feature comparison calibrated against human judgments. They do not remove stochastic reference ambiguity. FID and FVD compare collections through feature moments rather than aligned pairs; that changes the evaluation unit but can discard conditioning information. Neither change solves every reference problem.

Learned perceptual metrics (LPIPS)

LPIPS, from Zhang et al. (2018), compares two images using pretrained convolutional features rather than pixel coordinates. The recipe is:

  • Feed both images through a pretrained backbone (typically AlexNet or VGG on ImageNet).
  • Extract feature maps at several layers.
  • Normalize the channel vector at each spatial location to unit L2L_2 norm, with the implementation's numerical stabilization.
  • Take the squared difference of the normalized features at each layer.
  • Apply learned channel weights to the differences, sum over channels, and average over spatial positions.
  • Sum the layer contributions. The calibration weights are fitted to human perceptual judgments on a specific dataset.

The intuition behind the recipe is that early layers in a supervised network respond to edges and textures, while later layers respond to semantic content. Comparing several layers can therefore use different kinds of visual information, but the calibrated distance is not a general guarantee of semantic correctness. Unit-normalizing the channel vector removes its overall scale, apart from numerical stabilization; it does not balance individual channels, and relatively large channels can still dominate its direction. The per-channel coefficients are fitted so that the combined distance better agrees with the calibration judgments. Those coefficients do not necessarily identify which individual features people notice or what causally drives their verdicts.

In the original formulation, the squared difference is taken on the unit-normalized feature map ϕ^(l)\hat\phi^{(l)}. For a raw feature map ϕ(l)(x)∈RCl×Hl×Wl\phi^{(l)}(x) \in \mathbb{R}^{C_l \times H_l \times W_l} at layer ll, the LPIPS distance between two images xx and yy is defined as:

LPIPS(x,y)=∑l1HlWl∑i,j∑cwlc 2(ϕ^cij(l)(x)−ϕ^cij(l)(y))2\mathrm{LPIPS}(x,y) = \sum_l \frac{1}{H_l W_l} \sum_{i,j}\sum_c w_{lc}^{\,2}\Bigl(\hat\phi^{(l)}_{cij}(x)-\hat\phi^{(l)}_{cij}(y)\Bigr)^2

where:

  • x,yx, y: the two input images being compared
  • ϕ(l)(x)\phi^{(l)}(x): the feature map at layer ll of the pretrained backbone applied to xx, of shape Cl×Hl×WlC_l \times H_l \times W_l
  • ϕ^cij(l)(x)\hat\phi^{(l)}_{cij}(x): the channel-normalized activation at channel cc, spatial position (i,j)(i,j), layer ll for image xx
  • Cl,Hl,WlC_l, H_l, W_l: number of channels, height, and width of the feature map at layer ll
  • wlcw_{lc}: the channel calibration weight for channel cc at layer ll in the paper's squared weighted-norm notation. The learned coefficient on a squared channel difference is wlc2w_{lc}^{2}.

The expression is a sum of squared, weighted feature differences, not an ordinary Euclidean distance with a guaranteed triangle inequality. Similar normalized feature maps give a small score even when raw pixels differ. Classification training and channel normalization do not guarantee invariance to input brightness or color changes. Its perceptual interpretation comes from the reported human calibration and evaluation, not a universal invariance theorem.

This formula describes the original calibrated comparison with a fixed feature backbone. Calibration learns channel coefficients; the backbone already contains parameters learned during its own training. Report the metric implementation and version rather than assuming all feature-distance variants use this exact convention.

It is easy to misread LPIPS as "a universal perceptual distance." It is not. Three sources of dependence matter:

  • The backbone. AlexNet and VGG can produce different numbers on the same pair. Scores can also differ between ImageNet-trained and out-of-domain backbones.
  • The preprocessing. Image resizing, normalization statistics, and channel order (RGB vs BGR) shift the numbers.
  • The training set of the linear weights. The weights are calibrated against human judgments on the specific datasets (e.g., BAPPS) used in the original paper. On distributions far from that calibration, the metric is being extrapolated.

LPIPS is a full-reference perceptual-similarity proxy, not a no-reference realism detector. A beautifully rendered but physically impossible future is not guaranteed a low LPIPS: the score also depends on its reference. Conversely, a low score does not certify physical correctness. Perceptual similarity and correct physics are separate questions.

LPIPS does not measure realism without a reference. An image compared with itself has zero feature difference whether it shows noise or a detailed scene. Two independently drawn images from the same output style need not have near-zero LPIPS. In stochastic prediction, a valid alternative future can also differ perceptually from the particular realized reference.

Distributional metrics: FID

FID, introduced by Heusel et al. (2017), compares two image sample populations without requiring paired reference images. It is the squared 2-Wasserstein distance between multivariate Gaussians fitted to their Inception feature distributions. Let μr,Σr\mu_r, \Sigma_r denote real-image feature moments and μg,Σg\mu_g, \Sigma_g generated-image feature moments. The Gaussian distance is:

FID  =  ∥μr−μg∥22  +  Tr ⁣(Σr+Σg−2(Σr1/2ΣgΣr1/2)1/2)\mathrm{FID} \;=\; \lVert \mu_r - \mu_g \rVert_2^2 \;+\; \mathrm{Tr}\!\left( \Sigma_r + \Sigma_g - 2\bigl(\Sigma_r^{1/2} \Sigma_g \Sigma_r^{1/2}\bigr)^{1/2} \right)

where:

  • μr,Σr\mu_r, \Sigma_r: the mean and covariance of features of real images under the Inception network
  • μg,Σg\mu_g, \Sigma_g: the mean and covariance of features of generated images under the same encoder
  • ∥μr−μg∥22\lVert \mu_r - \mu_g \rVert_2^2: the squared Euclidean distance between the two feature means, capturing a shift in average feature location
  • Tr(⋅)\mathrm{Tr}(\cdot): the matrix trace, summing the diagonal entries
  • Σr1/2\Sigma_r^{1/2}: the symmetric positive-semidefinite square root of Σr\Sigma_r
  • (Σr1/2ΣgΣr1/2)1/2\bigl(\Sigma_r^{1/2} \Sigma_g \Sigma_r^{1/2}\bigr)^{1/2}: the matrix square root appearing in the Bures distance between the two covariances, capturing a difference in feature spread and correlation

The trace term is the squared Bures distance on positive-semidefinite covariance matrices. The symmetric matrix inside the inner square root is positive-semidefinite, so its positive-semidefinite square root is well-defined. The full expression is the squared 2-Wasserstein distance between the two fitted Gaussians under a fixed encoder, and is zero if and only if those Gaussians have the same mean and covariance.

The first term measures a shift in the center of the feature cloud. The second term measures a mismatch in the shape of the cloud, including spread and correlation between feature dimensions. Both are computed on features, not on pixels, so FID never asks whether a specific generated frame looks like a specific real frame. It asks whether the statistics of the two feature clouds match.

FID is not a distance between the full distributions. It compares first and second moments of Inception features. Differences in higher moments, in tail behavior, or in the actual support of either distribution are invisible to FID. Two clearly different image distributions can have identical FID of zero if their Inception features share the same mean and covariance.

The moment-blindness claim is visible in one dimension. A single Gaussian and a two-mode mixture can share a mean and variance while looking completely different.

In[14]:
Code
import numpy as np


def normal_pdf(x, mu, sigma):
    return np.exp(-0.5 * ((x - mu) / sigma) ** 2) / (
        sigma * np.sqrt(2.0 * np.pi)
    )


x_grid = np.linspace(-3.5, 3.5, 400)
single_gaussian = normal_pdf(x_grid, 0.0, 1.0)
# Mixture of two Gaussians with equal weights. Variance = sigma^2 + a^2 = 0.36 + 0.64 = 1.
two_mode_mixture = 0.5 * normal_pdf(x_grid, -0.8, 0.6) + 0.5 * normal_pdf(
    x_grid, 0.8, 0.6
)
# Verify the moment-matching claim the figure caption and surrounding prose rely on.
_mix_mean = 0.5 * (-0.8) + 0.5 * (0.8)
_mix_var = 0.5 * (0.6**2 + (-0.8 - _mix_mean) ** 2) + 0.5 * (
    0.6**2 + (0.8 - _mix_mean) ** 2
)
assert abs(_mix_mean - 0.0) < 1e-9
assert abs(_mix_var - 1.0) < 1e-9
Out[15]:
Visualization
Two density curves with equal mean and variance but different shapes.
A single Gaussian and a two-mode mixture that share the same mean and variance. Because FID compares only the first and second moments of encoder features, two distributions this different in shape can receive the same score when their moments match.

FID uses Gaussian fits to feature means and covariances, not feature histograms or pixel-aligned reference pairs.

The phrase "identical FID of zero" is not a hypothetical edge case. Any two distributions with the same first two moments of Inception features score identically under FID, by construction. If one of them happens to place all its mass on a small set of atypical images while the other spreads mass broadly, FID still returns zero. Higher-order structure, all the way down to which specific images appear, is unconstrained.

Three practical issues come with FID.

  • Finite-sample estimation. In practice the moments are estimated from finite samples: μ^r,Σ^r\hat\mu_r, \hat\Sigma_r from real features and μ^g,Σ^g\hat\mu_g, \hat\Sigma_g from generated features. A centered covariance estimate from nn vectors in dd dimensions has rank at most min⁡(d,n−1)\min(d,n-1), so it is necessarily rank-deficient when n≤dn\le d. Chong and Forsyth (CVPR 2020), in "Effectively Unbiased FID and Inception Score and Where to Find Them", show finite-sample, model-dependent FID bias. Rank deficiency alone does not imply a universally downward bias, and using the same sample count does not eliminate differences in estimator bias between models. Report real and generated sample counts and the estimation protocol.
  • Encoder dependence. Conventional FID uses Inception-v3 features pretrained for ImageNet classification, not a feature distance fitted directly to human judgments. Domain changes can alter which distinctions the features retain. Validate their relevance before interpreting scores on a new image domain.
  • Preprocessing and protocol sensitivity. Image resizing, interpolation, JPEG compression, and whether the model produces the same number of samples as the reference all affect FID. The right discipline is to hold these constant when comparing models and to report them.

A small FID difference may be comparable to estimation noise or model-dependent bias. A bootstrap interval can describe sampling uncertainty, but it does not automatically remove that bias. Resample independent units: clips from the same trajectory may require trajectory-level grouping. Repeated high-dimensional covariance square roots can also be costly. Report the sample counts, estimator, resampling unit, and interval construction.

FVD (Fréchet Video Distance), from Unterthiner et al. (2018), applies the Gaussian feature-distance construction to video clips, using spatiotemporal features such as I3D embeddings. State clip duration, frame rate, spatial resolution, encoder checkpoint, and feature layer. Changes in these can change the score. FVD can detect some temporal failures, but matching its feature moments does not certify object persistence, identity, or causality.

What FID/FVD do and do not measure
  • They measure agreement between moments of features under a specified encoder, not agreement between full distributions.
  • They are not pixel-aligned and never compare individual generations to specific references.
  • Both depend on encoder choice, sample counts, preprocessing, and spatial resolution. FVD also depends on clip duration and frame rate.
  • They do not have a natural interpretation as a "physical correctness" or "causality" score.

The practical implication is that perceptual and distributional metrics are useful complements to paired-error metrics, not replacements. For either FID or FVD, state the encoder, sample counts, preprocessing, and spatial resolution. For FVD, also state clip duration and frame rate; for FID computed on video frames, state how frames were sampled. Still-image FID has no intrinsic clip-length parameter. When you report LPIPS, state the backbone and the weights. When you report PSNR, state the dynamic range. A single number without its protocol is not a measurement.

A signature failure: mass-matching without conditioning

A pooled marginal can hide action-conditioning errors. Suppose two actions occur with equal weight and a predictor swaps their distinct future distributions. The pooled mixture is unchanged, so a metric of that marginal cannot detect the swap. A conditional comparison can reveal it if its features and statistic distinguish the two conditionals. Equal pooling weights matter; mirror-image shapes alone do not guarantee equality with unequal action frequencies.

Pooling discards the association between a condition and its output. Stratifying by action or history is one remedy; testing the joint condition-output distribution or using a proper conditional probabilistic score are others. A stratified moment metric still has encoder and moment-matching blind spots.

The next section demonstrates this pooling failure with a one-dimensional toy Gaussian-feature distance, not a published metric. The same loss of conditioning can affect pooled FID or FVD. Paired LPIPS and SSIM retain their specific references, so this argument does not apply to them in the same way.

Diversity, Sharpness, and Multimodality

Unobserved quantities and measurement noise can make a future stochastic given the available history. An exposed point prediction does not communicate that uncertainty, even if the model maintains a distribution internally. To evaluate probabilistic predictions, inspect the declared conditional distribution as well as point-loss performance.

The vocabulary here is easy to get wrong, so let us separate what these words mean when we use them as metrics.

  • Visual sharpness: crispness or high-frequency structure in individual rendered samples. This is distinct from forecast sharpness, the concentration of a predictive distribution. A visually sharp sample can come from a broad multimodal forecast.
  • Mode coverage: whether the generator represents materially probable alternatives, evaluated with a specified feature-space or event-based diagnostic. Merely assigning nonzero density everywhere is not useful coverage.
  • Diversity: the spread of samples the model can generate. High diversity alone is not virtue either; a model that samples noise is diverse but useless.
  • Calibration: the agreement between the model's implied probabilities and the empirical frequencies of events it predicts.

These are different properties, and a single number cannot summarize all of them. Image-gradient or frequency diagnostics can describe visual sharpness but also reward noise. Entropy or interval widths can describe forecast concentration under a declared representation. Precision/recall-style diagnostics describe feature-space fidelity and coverage; variance or mode counts describe sampled diversity; event reliability describes calibration. NLL evaluates the probabilities or densities assigned to outcomes, not sharpness alone.

It is worth pausing on why these words are easy to conflate. Visual sharpness is a property of one rendered image. Diversity is a property of a set of rendered images. Mode coverage is a property of the set relative to the true conditional. Calibration is a property of the model's probability assignments, which may or may not be visible in any rendered image. A model can score high on any one of these and low on the others. A table that reports a single number under the heading "quality" is hiding at least three questions.

Best-of-KK: oracle selection inflates scores

Best-of-KK selection appears in stochastic video-prediction work: the SAVP project presents stochastic predictions selected for their best similarity to the reference among 100 samples. The metric samples KK predictions from the model's conditional, compares each to the realized future, and reports the best score. This number improves monotonically as KK grows, because you are taking a minimum over more samples; formally, for each realized future yy and predicted set {y^1,…,y^K}\{\hat y_1, \dots, \hat y_K\}, the best-of-KK error is min⁡k∈{1,…,K}d(y^k,y)\min_{k \in \{1,\dots,K\}} d(\hat y_k, y) for a chosen distance dd, and the expectation of a minimum over KK i.i.d. samples is non-increasing in KK. It is oracle selection: no real deployment gets to see the realized future and pick the closest of KK candidates. Best-of-KK only has interpretation if KK is held fixed across the models you compare and the full sampling protocol is reported.

Adding a sample to a nested pool cannot increase its minimum error. The expectation for i.i.d. samples follows the same ordering. A misspecified forecast can therefore obtain a small best-of-KK error if it assigns positive probability to a sufficiently close neighborhood of the reference: the chance of at least one hit is 1−(1−q)K1-(1-q)^K, where qq is that neighborhood's probability. If q=0q=0, no sample budget creates coverage. This improvement is selection, not evidence of calibration.

In the following experiment, each repetition draws one maximum-budget sample pool from a Gaussian centered slightly off the realized future. Each KK uses the first KK samples from that same pool. These nested sets guarantee non-increasing error for every repetition and for their mean. Independently resampled finite estimates for different KK need not be monotone, even though their expectations are non-increasing. Neither result establishes calibration.

In[16]:
Code
rng = np.random.default_rng(20240501)
SIGMA_POS = 1.2

K_values = np.array([1, 2, 4, 8, 16, 32, 64, 128])
repeats = 400
sample_pool = rng.normal(
    REALIZED_X + 1.5, SIGMA_POS, size=(repeats, int(K_values.max()))
)
prefix_best = np.minimum.accumulate(np.abs(sample_pool - REALIZED_X), axis=1)
best_of_k_mean = prefix_best[:, K_values - 1].mean(axis=0)

assert np.all(np.diff(best_of_k_mean) <= 1e-9), (
    "best-of-K must be non-increasing in K"
)
assert best_of_k_mean[0] > best_of_k_mean[-1]
Out[17]:
Visualization
Line chart of mean best-of-K error decreasing as K grows from 1 to 128 on a log scale.
Mean best-of-$K$ position error for a conditional that samples from a Gaussian offset from the realized future. The score improves monotonically as $K$ grows because a larger budget makes it more likely that some sample lands near the target. This is oracle selection, so the number is only interpretable when $K$ is fixed across models and the full sampling protocol is reported.

The figure is not evidence that the model is getting better. It is evidence that oracle selection is happening. When you see a table that reports best-of-KK numbers, check whether the same holdout data is used for all models, whether KK is constant, and whether the number is a mean or a minimum. If the sampling protocol is not in the paper, treat the number as uninterpretable.

A deployment may generate one sample or an ensemble. What it cannot generally do is select the sample closest to a future that has not occurred. Report the deployment sample budget and selection procedure separately from reference-based best-of-KK. Single-sample error and distributional or probabilistic scores are useful complementary diagnostics, not substitutes for that distinction.

Diversity through high variance is not coverage

High variance need not mean useful diversity. An increasingly wide Gaussian centered between two narrow modes has full real-line support, yet assigns vanishing probability to fixed neighborhoods of those modes as its scale grows. Useful coverage concerns probability mass near plausible alternatives, not merely nonzero density or spread. Feature-space sample diagnostics and event-based probability checks answer different parts of that question.

The reason spread is a weak proxy for coverage is that a distribution can spread out while simultaneously thinning out on the regions that matter.

The same distinction appears as a density plot. The target mixes two Gaussians centered at −1-1 and +1+1, each with standard deviation 0.3. The comparison forecast is a zero-mean Gaussian with standard deviation 2. Its variance is 4, compared with the target's 1+0.32=1.091+0.3^2=1.09, so these distributions are not variance-matched. The broad forecast has full support but less density at the target modes, about 26.5% of the target's mode density. We also compute probability mass in the two fixed intervals within 0.3 of each mode: density at a point and probability in a neighborhood are different quantities. The plotted window shows the central region, not the full tails.

In[18]:
Code
import numpy as np


def gaussian_density(x, mu, sigma):
    return np.exp(-0.5 * ((x - mu) / sigma) ** 2) / (
        sigma * np.sqrt(2.0 * np.pi)
    )


x_coverage = np.linspace(-4.0, 4.0, 400)
true_bimodal = 0.5 * gaussian_density(
    x_coverage, -1.0, 0.3
) + 0.5 * gaussian_density(x_coverage, 1.0, 0.3)
broad_gaussian = gaussian_density(x_coverage, 0.0, 2.0)
_target_at_mode = float(
    0.5 * gaussian_density(1.0, 1.0, 0.3)
    + 0.5 * gaussian_density(1.0, -1.0, 0.3)
)
_broad_at_mode = float(gaussian_density(1.0, 0.0, 2.0))
mode_density_ratio = _broad_at_mode / _target_at_mode
# Analytic density ratio, not an unsupported near-zero threshold.
expected_density_ratio = 0.3 * np.exp(-0.125) / (1 + np.exp(-2 / 0.09))
assert np.isclose(mode_density_ratio, expected_density_ratio)
target_variance, broad_variance = 1 + 0.3**2, 2.0**2
assert np.isclose(target_variance, 1.09) and broad_variance == 4.0

from math import erf


def interval_mass(lo, hi, mu, sigma):
    return 0.5 * (
        erf((hi - mu) / (sigma * np.sqrt(2.0)))
        - erf((lo - mu) / (sigma * np.sqrt(2.0)))
    )


mode_intervals = [(-1.3, -0.7), (0.7, 1.3)]
target_mode_mass = sum(
    0.5 * interval_mass(lo, hi, mu, 0.3)
    for lo, hi in mode_intervals
    for mu in (-1.0, 1.0)
)
broad_mode_mass = sum(
    interval_mass(lo, hi, 0.0, 2.0) for lo, hi in mode_intervals
)
assert 0 < broad_mode_mass < target_mode_mass < 1

moment_grid = np.linspace(-20.0, 20.0, 20001)
target_density = 0.5 * gaussian_density(
    moment_grid, -1.0, 0.3
) + 0.5 * gaussian_density(moment_grid, 1.0, 0.3)
broad_density = gaussian_density(moment_grid, 0.0, 2.0)
assert np.isclose(
    np.trapezoid(moment_grid**2 * target_density, moment_grid), target_variance
)
assert np.isclose(
    np.trapezoid(moment_grid**2 * broad_density, moment_grid), broad_variance
)
Out[19]:
Visualization
Central density curves of a narrow bimodal target and a broad zero-mean Gaussian with standard deviation 2.
A bimodal target and a zero-mean Gaussian with standard deviation 2. Their variances are 1.09 and 4, respectively, not matched. The broad forecast has full support but lower density near the target modes. Only the central feature window is displayed.

As this Gaussian's standard deviation tends to infinity, probability in any fixed bounded neighborhood tends to zero. At every finite positive scale, however, that probability remains positive: samples near the target modes are unlikely rather than impossible. Continuous point probabilities are zero for both distributions; the plotted heights are densities, not point masses.

Sajjadi et al. (2018) formulated generative precision and recall, and Kynkaanniemi et al. (2019) proposed an improved feature-manifold estimator. For that estimator, the two quantities separate:

  • Precision: the fraction of generated feature samples inside an estimated real-data support region.
  • Recall: the fraction of real feature samples inside an estimated generated-data support region. It is not a fraction of a manifold's geometric volume.

The decomposition is useful because the two failure modes have different fixes. Low precision suggests the generator needs to be constrained toward the real data distribution. Low recall suggests it needs to diversify, for example by using more modes or a less collapsed latent. Reporting a single score hides which of the two problems is present.

These diagnostics are useful because they separate coverage from fidelity, and they avoid the moment-matching limitations of FID. But their semantics depend on hyperparameters (typically kk-nearest-neighbor radii or manifold balls), on feature choice, and on sample counts. Two implementations with different conventions can rank models differently. Report the implementation you used.

A toy feature-distance illustration

The following example is deliberately simple. It uses one synthesized feature dimension and two discrete actions. We work in one feature dimension with two actions. Under action a1a_1 the true conditional feature is Gaussian with mean +1+1 and standard deviation 0.50.5, and under action a2a_2 it is Gaussian with mean −1-1 and standard deviation 0.50.5. We have two candidate generators:

  • G1 (matched) produces features with mean +1+1 under a1a_1 and mean −1-1 under a2a_2. It matches the ground-truth conditional.
  • G2 (reversed) produces features with mean −1-1 under a1a_1 and mean +1+1 under a2a_2. It reverses every action.

With equal action weights, G1 and G2 have identical population marginals: the same two Gaussians appear in the same proportions. We draw their finite samples independently, so the empirical pooled sets and scores will not be identical. Their small differences are sampling effects, not evidence that pooling detects the reversed action labels.

We compute a "toy feature Gaussian distance", defined as the squared 2-Wasserstein distance between two Gaussian fits of the empirical features. For one-dimensional features with means μx,μy\mu_x, \mu_y and standard deviations σx,σy\sigma_x, \sigma_y,

toy-dist2   =  (μx−μy)2+(σx−σy)2\text{toy-dist}^2\, \;=\; (\mu_x - \mu_y)^2 + (\sigma_x - \sigma_y)^2

where:

  • μx\mu_x: the sample mean of the first empirical feature set
  • μy\mu_y: the sample mean of the second empirical feature set
  • σx\sigma_x: the sample standard deviation of the first feature set
  • σy\sigma_y: the sample standard deviation of the second feature set

In one dimension, the 2-Wasserstein distance between two Gaussians N(μx,σx2)\mathcal{N}(\mu_x, \sigma_x^2) and N(μy,σy2)\mathcal{N}(\mu_y, \sigma_y^2) reduces to W22=(μx−μy)2+(σx−σy)2W_2^2 = (\mu_x - \mu_y)^2 + (\sigma_x - \sigma_y)^2. The first difference captures a location shift and the second captures a spread mismatch.

The reason we use a Wasserstein distance rather than an L2L_2 distance between densities or parameters is that it is the same construction that FID uses, just with matrices in place of scalars. Keeping the algebra identical while collapsing dimension is what makes the toy example informative: the failure mode you see here is exactly the failure mode a real FID exhibits, minus the high-dimensional bookkeeping.

We compute the distance separately within each action and average with equal weights, then compute it on the concatenated pools. Population pooled distances are both zero. Finite independent samples yield small nonzero pooled distances, while the conditional distance reveals G2's reversal because its action means differ from their references.

This is explicitly a toy feature Gaussian distance. It is not FID, not LPIPS, not SSIM, and not FVD. It is a pedagogical device for the conceptual failure. Real FID and FVD would use high-dimensional learned features and a matrix square root. The algebra is the same and the failure mode is identical.

In[20]:
Code
def toy_feature_gauss_dist(x, y):
    """Toy 1-D Gaussian Wasserstein distance squared between two empirical feature sets.

    For Gaussian features this equals $(\mu_x - \mu_y)^2 + (\sigma_x - \sigma_y)^2$.
    This is *not* FID, LPIPS, SSIM, or FVD; it is a hand-designed device to
    illustrate the action-conditioned vs pooled-marginal failure.
    """
    mu_x, mu_y = float(np.mean(x)), float(np.mean(y))
    sd_x, sd_y = float(np.std(x, ddof=1)), float(np.std(y, ddof=1))
    return (mu_x - mu_y) ** 2 + (sd_x - sd_y) ** 2
In[21]:
Code
rng = np.random.default_rng(7)
N_PER_ACTION = 400
MODE_MEAN_POS = 1.0
MODE_MEAN_NEG = -1.0
FEATURE_SD = 0.5
# Large separation relative to sampling noise is a design heuristic for this toy,
# not a necessary detection threshold.
assert (MODE_MEAN_POS - MODE_MEAN_NEG) > 4 * FEATURE_SD / np.sqrt(N_PER_ACTION)

# Ground-truth conditional features under each action.
gt_a1 = rng.normal(MODE_MEAN_POS, FEATURE_SD, N_PER_ACTION)
gt_a2 = rng.normal(MODE_MEAN_NEG, FEATURE_SD, N_PER_ACTION)

# G1: matched conditional.
g1_a1 = rng.normal(MODE_MEAN_POS, FEATURE_SD, N_PER_ACTION)
g1_a2 = rng.normal(MODE_MEAN_NEG, FEATURE_SD, N_PER_ACTION)

# G2: reversed conditional (a1 features look like a2 features and vice versa).
g2_a1 = rng.normal(MODE_MEAN_NEG, FEATURE_SD, N_PER_ACTION)
g2_a2 = rng.normal(MODE_MEAN_POS, FEATURE_SD, N_PER_ACTION)
Out[22]:
Console
G1 conditional   0.006109
G2 conditional   4.021280
G1 pooled        0.004088
G2 pooled        0.001985
Out[23]:
Visualization
Bar chart: conditional scores of about 0 (G1) and 4 (G2); pooled scores of about 0 for both.
Toy one-dimensional Gaussian-feature distances for matched and action-reversed generators. Conditional distances are approximately 0.0061 for G1 and 4.0213 for G2. Independent finite samples give pooled distances approximately 0.0041 and 0.0020, respectively, despite identical population marginals under equal action weights. This toy is not published FID or FVD.

A reported FID needs its conditioning protocol: a single headline number might be pooled, stratified, or computed for one condition. Pooling alone cannot rule out the balanced label-swap failure. One diagnostic computes FID or FVD within stated action or history strata, then aggregates with declared weights. Multiple covariance estimates and matrix square roots can increase cost, while smaller per-stratum sample sizes increase estimation uncertainty. Stratification does not remove feature-moment blind spots.

Each stratum needs enough samples for reliable statistics. Binning or clustering is one approach for continuous or high-dimensional conditions, but the chosen resolution can merge important differences. Conditional two-sample methods or joint condition-output tests are alternatives. Report the grouping, weights, sample counts, and remaining blind spots.

Finite-sample estimation bias

For two independent samples of size n≥2n\ge2 from the same nondegenerate Gaussian, the population toy distance is zero but the expected plug-in estimate is positive. Mean and standard-deviation errors are of order n−1/2n^{-1/2} here. In particular, the expected squared difference of sample means alone is 2σ2/n2\sigma^2/n. The additional squared standard-deviation difference is nonnegative, so this matched-sample expectation remains positive at finite nn and tends to zero as nn grows. Degenerate data or comparing a sample with itself are different cases.

The positive expectation arises from squaring estimator differences. It is a property of this plug-in estimator and sampling setup, not a theorem that every finite-sample metric must be biased or that every realization decreases monotonically with sample count.

In[24]:
Code
rng = np.random.default_rng(99)
n_list = np.array([50, 100, 200, 400, 800, 1600])
bias_estimates = []
for n in n_list:
    vals = []
    for _ in range(300):
        a = rng.normal(MODE_MEAN_POS, FEATURE_SD, int(n))
        b = rng.normal(MODE_MEAN_POS, FEATURE_SD, int(n))
        vals.append(toy_feature_gauss_dist(a, b))
    bias_estimates.append(float(np.mean(vals)))
bias_estimates = np.array(bias_estimates)
# The claim is that bias is positive at finite n and decays as n grows.
assert np.all(bias_estimates > 0.0)
assert bias_estimates[0] > bias_estimates[-1]
Out[25]:
Visualization
Line chart of toy feature distance decreasing as sample size per generator increases from 50 to 1600.
Mean plug-in toy Gaussian-feature distance over 300 independent matched-Gaussian sample pairs at each size. The population distance is zero; the estimates have positive expectation and decline with sample size in this experiment. This one-dimensional result does not establish the bias direction or magnitude of every FID/FVD comparison.

This is a finite-sample illustration for independent Gaussian samples. It does not establish the magnitude or direction of bias for every real FID/FVD protocol. Real feature covariance estimates additionally become rank-deficient when sample counts are too small for their dimension. Chong and Forsyth (2020) show that FID's finite-sample bias depends on the evaluated generator, so equal sample counts alone do not eliminate biased comparisons. A reported score depends on the model, encoder, sample counts, and estimator.

Using the same sample count for two models does not guarantee equal bias. Report the feature encoder, sampling protocol, estimator, and uncertainty. Confidence intervals describe sampling variation under that protocol; they do not remove estimator bias. Chong and Forsyth's analysis explains why the model-dependent bias also matters when comparing scores.

Sharpness metrics and their limits

Negative log-likelihood (NLL) evaluates outcome assignments, not visual sharpness alone. For a declared discrete probability mass function or continuous density pθp_\theta, the realized-future score is

NLL=−log⁡pθ(ot+1:t+H∣ht,at:t+H−1)\mathrm{NLL} = -\log p_\theta(o_{t+1:t+H}\mid h_t,a_{t:t+H-1})

where:

  • NLL\mathrm{NLL}: the negative log-likelihood of the realized future under the model's predictive distribution (the quantity being defined)
  • pθ(ot+1:t+H∣ht,at:t+H−1)p_\theta(o_{t+1:t+H}\mid h_t,a_{t:t+H-1}): the model's predictive probability (or density) of the realized future observations given history and actions
  • hth_t: the history at time tt, as defined earlier
  • at:t+H−1a_{t:t+H-1}: the sequence of actions over the prediction horizon

Here the logarithm is natural, so the score is in nats. Lower values indicate more assigned probability mass or density at the realized outcome. With a common reference measure and finite expectations, expected NLL is the true conditional entropy plus the conditional KL divergence, averaged over histories and actions. The true conditional distribution therefore minimizes expected NLL among unrestricted normalized forecasts. Calibration on selected events alone does not establish this optimum.

Zero assigned mass or density gives positive infinity. Discrete NLL is nonnegative; continuous densities can exceed one, giving negative NLL, and their numerical values depend on units and base measure. Clipping a log score is not necessarily scoring a normalized model: flooring the probabilities (0.9,0.1)(0.9,0.1) at 0.20.2 gives (0.9,0.2)(0.9,0.2), whose sum exceeds one. Renormalize a floored distribution, or use a declared mixture with a normalized contamination distribution, and disclose the change. A uniform contamination density on an unbounded Euclidean observation space cannot be normalized.

Near-optimal NLL and imperfect event reliability can coexist; exact optimal expected NLL under these assumptions is a stronger statement. For a binary event occurring with probability 0.90.9, forecasting 0.850.85 gives expected NLL about 0.3360.336 nats, versus the optimum of about 0.3250.325 nats, yet the forecast is off by 0.050.05. Reliability diagrams reveal that mismatch rather than replacing the log score.

The following reliability curves are stipulated illustrations, not estimates from labeled outcomes. The diagonal represents calibration. The second curve has event frequencies more extreme than the forecast probabilities, so it depicts underconfidence. No average accuracy or NLL is computed for either curve.

In[26]:
Code
import numpy as np

probability_bins = np.array([0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9])
calibrated_frequency = probability_bins
underconfident_frequency = probability_bins + 0.15 * (probability_bins - 0.5)
underconfident_frequency = np.clip(underconfident_frequency, 0.0, 1.0)
assert np.all(
    (underconfident_frequency - probability_bins) * (probability_bins - 0.5)
    >= 0
)
Out[27]:
Visualization
Analytic calibrated diagonal and underconfident reliability curve.
Illustrative reliability curves, specified analytically rather than measured from outcomes. The underconfident curve has event frequencies farther from 0.5 than the forecast probabilities: below the diagonal at low probabilities and above it at high probabilities. The figure does not report accuracy, bin counts, sampling uncertainty, or NLL.

For probabilistic forecasts, distinguish concentration of the forecast from its calibration and coverage. Where these diagnostics are defined, report NLL, a diversity diagnostic, and reliability on relevant events. State the interpretation and assumptions of any rank histogram. A sharpness claim alone leaves those other properties untested.

Human Judgment and Learned Evaluators

When the target is human perception, direct ratings provide evidence about what the sampled people judge under the tested presentation conditions. They are not an objective certificate of physical correctness. Human studies also have sampling uncertainty and presentation biases, so their protocol belongs beside their scores.

The essentials of a defensible protocol are:

  • Rater population and instructions. State whom the study represents, how raters were recruited and instructed, and any quality-screening or exclusion rules. Do not infer valid ratings from agreement alone.
  • Blinded, randomized presentation. Hide method identity and randomize candidate order independently of method. A consistency task may need an explicitly identified reference history or action; blinding does not mean hiding information needed to answer the question.
  • Separate questions for visual realism and conditional consistency. "Does this look like a real image?" and "Does this prediction correctly continue from the given history under the given action?" measure different things. Ask them separately. A model can be visually realistic but inconsistent, and vice versa.
  • Matched playback and sampling. Match horizon, frame rate, resolution, and sample-selection budget when comparing methods on those terms. If presentation is itself an experimental variable, vary it explicitly rather than confounding it with method identity.
  • Explicit handling of ties and uncertainty. Choose and report a response format appropriate to the task. A tie or "can't tell" option can express uncertainty; a two-alternative forced-choice design is also legitimate, but it does not directly measure indifference.
  • Uncertainty that respects clustering. Frames within a trajectory can be correlated. For inference over independent new trajectories with fixed raters, resample trajectory clusters when within-trajectory dependence is possible. To generalize over raters as well, account for crossed trajectory and rater variation with a suitable resampling scheme or statistical model.
  • Reported rater agreement. Use an agreement measure suited to the rating scale and design. Low agreement can reflect ambiguous instructions, heterogeneous preferences, or rater reliability; it does not identify the cause by itself.

These controls address different threats. Blinding reduces identity cues; it cannot guarantee that a method's visual signature is unrecognizable. Separate realism and consistency questions make a trade-off visible rather than merging it into one preference. Matched playback controls presentation differences, while cluster-aware uncertainty avoids treating dependent frames as independent evidence. Agreement statistics describe the ratings, not the physical truth of a continuation.

Report presentation context, including display, audio, color, and playback speed. These can affect judgments. Differences in protocol limit direct comparison between studies unless their effects have been controlled or separately assessed.

Learned evaluators

Collecting human ratings at scale can be expensive. One alternative is a learned evaluator: a model that predicts a human judgment from an image, a video clip, or a pair of clips. This includes image-aesthetic regressors, video-quality models, video-language models prompted to compare two candidate continuations, and multimodal systems that imitate pairwise comparisons. Within a validated task, such a model can approximate human judgments. Any claimed cost saving needs a separate accounting of repeated inference, training, adaptation, and held-out human validation; agreement alone does not establish a cost ratio. A learned evaluator can also produce scores that reflect the biases of its training set.

Three checks before using a learned evaluator for a reported comparison:

  • Held-out human agreement. The evaluator must be shown to agree with held-out human judgments on a dataset that is not part of its training. Training-set agreement alone is not independent validation.
  • Domain validation. A judge trained on generic internet video is not guaranteed to give useful scores on surgical footage, satellite imagery, or robot manipulation scenes. Validate on the target domain.
  • Specification of prompt and protocol. For a prompted evaluator, report the full prompt, candidate order, temperature, and sampling scheme. Test sensitivity to these choices for the actual judgment task rather than assuming prompt invariance.

There are systemic risks.

  • Evaluator overfitting. Training or selecting a generator against a fixed evaluator can exploit its errors. Keep an independent evaluation set and, where possible, independent judges; optimization does not guarantee that a blind spot will be found in every case.
  • Shared training bias. A shared pretraining corpus can create correlated blind spots in the generator and evaluator. Do not treat their agreement as independent verification without independent, task-specific validation.
  • Adversarial or gaming risk. Learned evaluators may be sensitive to input perturbations or out-of-distribution content. Test robustness relevant to the intended domain rather than assuming that every evaluator has the same vulnerability.

Use a learned evaluator as a measured proxy with a stated validation domain. Held-out human agreement, documented prompts, and matched sample budgets strengthen a comparison but do not guarantee a correct ranking. If the evaluator is used for training, keep that role separate from independent reporting and check whether optimization degrades agreement on held-out judgments.

Learned evaluators are a useful complement to human evaluation, especially for exploring large design spaces. They are not a substitute for it, and they are definitely not a substitute for the fact-based checks that come next. No language model judge is a substitute for verifying that predicted object positions match world coordinates, that velocities are consistent with physics, that interventions produce the right responses, or that the memory state preserves object identity across occlusions.

When to trust a learned evaluator

A learned evaluator can support a ranking within its validated scope. Report held-out human agreement, prompt and protocol, and sample budgets. Domain validation is necessary evidence for that use, not a guarantee. A training judge should not be presented as independent evaluation, and a preference score does not replace factual consistency checks.

Putting the pieces together: a complementary scorecard

No single number is a complete evaluation. A defensible scorecard for perceptual and generative performance addresses the following layers with diagnostics suited to the model's outputs and the evaluation task, each reported with its protocol:

  • Paired one-step scores. MSE and PSNR with declared dynamic range, on teacher-forced prediction. Report both spatial and per-frame averaging conventions.
  • Paired multi-step scores. Same metrics on fixed-action open-loop rollouts, with horizon curves (metric versus kk) instead of a single pooled number.
  • Perceptual metrics. LPIPS with a stated backbone and preprocessing, or SSIM with stated window, Gaussian parameters, and LL.
  • Conditional distribution checks. Where sample support allows it, report FID or FVD within declared action/history strata with explicit aggregation weights, or use another conditional test. Pooled scores can miss reversals; stratified feature moments can also miss conditional differences that preserve those moments.
  • Diversity, coverage, and calibration. Use precision/recall-style diagnostics and event reliability where defined. Report realized-future NLL when a normalized mass or density is evaluable under a declared common reference measure. For a sample-only predictor, use a suitable proper sample-based score where its assumptions hold rather than implying that samples expose a likelihood. Disclose normalized smoothing or other forecast modifications; if clipping does not define a normalized forecast, identify the result as a modified log score rather than its NLL.
  • Sample-selection protocol. Report KK and the selection rule. Future-reference best-of-KK is an oracle diagnostic, not deployable selection. A deployment may use one sample, an ensemble, or another rule; evaluate that rule separately.
  • Human or learned evaluation. Held-out human agreement for any learned evaluator; blinded randomized protocol for direct human studies.

Discuss score disagreements within their measurement scope. Better pooled FVD but worse action-stratified FVD suggests different aggregate and conditional feature-moment behavior; it does not establish a causal mechanism. Lower test MSE but higher LPIPS means a better empirical squared-error score and a worse feature-dissimilarity score on that protocol, not global optimality or a universal human preference.

Trace each row back to its measured quantity. Pooled FVD summarizes spatiotemporal feature moments over its sample population. Action-stratified FVD compares those moments within the selected action groups. Their disagreement can expose a weakness hidden by pooling, but agreement does not certify complete conditional correctness.

Observation scores can respond to visible consequences of physics or interventions, but do not by themselves certify internal state, persistent memory, or causal correctness. Those checks belong to State, Physics, Causality, and Memory Evaluation. Whether the model supports better decisions is a separate question in Planning, Control, and Policy Evaluation. Datasets, Benchmarks, and Experimental Design covers the broader protocol choices.

Limitations and Impact

Distributional feature metrics allow comparisons between unpaired real and generated sample sets. Learned perceptual metrics provide comparisons calibrated to particular human similarity judgments. Those capabilities complement paired pixel errors; none measures every property of a useful world model.

FID and FVD compare first and second feature moments under a fixed encoder. They miss differences that preserve those moments, including the balanced action reversal in this chapter. LPIPS measures feature similarity calibrated to particular human judgments, not realism in general. SSIM is a local similarity proxy whose window, weights, and dynamic range must be specified. MSE and PSNR can score a stochastic future, but closeness to one realized draw does not identify its conditional distribution. Human ratings measure the sampled judgments under the study conditions; learned judges require their own validation. Neither replaces physical or causal checks.

Optimizing a reported metric does not remove its blind spots. A model designed to minimize FID may match Inception feature moments while predicting the wrong dynamics. A model optimized against a learned evaluator may satisfy that evaluator while violating physics. Report complementary checks for the properties the application needs, and disclose their evaluation cost rather than assuming feature metrics are cheap.

The practical implication is a discipline of reporting. State the protocol, stratify by condition, use horizon curves, hold KK fixed, and publish the prompts, backbones, and encoders. Prefer complementary scores over a single headline number. A world model that scores well on every perceptual metric we have discussed and yet moves objects in the wrong direction under a well-specified action is not a good world model, no matter how good the numbers look. The next chapter starts from exactly that point.

Summary

Perceptual and generative evaluation answers a specific question: how closely do a model's predicted observations resemble the observations the world produces, individually and in distribution. The chapter developed this question in four layers.

  • Reconstruction and predictive accuracy. MSE, PSNR, and SSIM are paired full-reference metrics with different score directions. Expected squared error is minimized by the conditional mean under finite-second-moment assumptions; this theorem does not establish an optimum for expected PSNR or SSIM. In our separated two-mode example the mean looks faded, while a sharp prediction matching the realized mode wins on that draw. Declare dynamic range, averaging conventions, and SSIM preprocessing.
  • Perceptual and distributional metrics. LPIPS compares images in the feature space of a pretrained network with weights fit to human judgments, and inherits the backbone, preprocessing, and calibration-set dependencies of that construction. FID is the squared 2-Wasserstein distance between Gaussian fits of Inception features of real and generated distributions; FVD generalizes to spatiotemporal features. Both compare moments, not full distributions, and both depend on encoder choice, sample counts, preprocessing, and spatial resolution. FVD also depends on clip duration and frame rate; video-frame FID needs a declared frame-sampling protocol, not an intrinsic clip-length input.
  • Diversity, sharpness, and multimodality. Distributional checks complement point losses for stochastic futures. Nested best-of-KK errors cannot increase as KK grows, but this oracle rule does not test calibration or deployable selection. High variance is not coverage. Precision/recall depends on its feature-space estimator. Pooling can hide balanced action reversals; stratification exposes our separated-mean example but is not a complete conditional test. Conventional plug-in feature distances can have model-dependent finite-sample bias; sample size, estimator, and uncertainty all matter.
  • Human judgment and learned evaluators. Direct ratings require a stated question, presentation and sampling protocol, and uncertainty appropriate to trajectory and rater dependence. Learned evaluators need held-out agreement and domain validation. Their scores remain proxies, particularly when the same judge is used to train or select a generator.

Perceptual and generative scores provide evidence about specified observation properties. They do not certify state, physics, causality, memory, or decision quality. Use them alongside tests designed for those properties. The next chapter examines what the model preserves and predicts beyond visual similarity.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about perceptual and generative evaluation of world models.

Perceptual and Generative Evaluation

Question 1 of 80 of 8 completed
The chapter states that expected squared error is minimized by the conditional mean. What does this imply for a multimodal future where objects can occupy separated locations?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026perceptualgenerative, author = {Michael Brenndoerfer}, title = {Perceptual and Generative Evaluation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-10-11} }
APAAcademic
Michael Brenndoerfer (2026). Perceptual and Generative Evaluation. Retrieved from https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips
MLAAcademic
Michael Brenndoerfer. "Perceptual and Generative Evaluation." 2026. Web. October 11, 2026. <https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips>.
CHICAGOAcademic
Michael Brenndoerfer. "Perceptual and Generative Evaluation." Accessed October 11, 2026. https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Perceptual and Generative Evaluation'. Available at: https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips (Accessed: October 11, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Perceptual and Generative Evaluation. https://mbrenndoerfer.com/writing/perceptual-generative-evaluation-metrics-fid-fvd-lpips

About the author

Continue with the full handbook

This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore World Models Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.