Part of World Models Handbook
Model web, desktop, and code-execution state as partial observations, then verify actions, correct beliefs, and separate prediction error from task success.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Digital and Software Environments
Imagine building an agent that files bug reports in a locally hosted issue tracker. It predicts what a title entry, assignee selection, and Save click will do. In a hypothetical training set, most sessions are valid and most assignees belong to the team, so predicting that Save adds a row looks accurate. The interesting question is what happens when those conditions change.
The agent types "Login fails", selects "ada", and clicks Save. Suppose this application's client shows an optimistic "Saved" banner before the server responds, but the server rejects the write because the session has expired. The persistent issue list is unchanged. If the agent treats the banner as proof of persistence, its belief can drift from the store. This is an illustrative interface design, not a claim that every application handles Save this way.
The distinction is between a rendered observation and the state that determines an action's effects. Software workflows can have partial effects: a record may be committed while a notification fails, or one service may update while another does not. An atomic database transaction provides all-or-nothing behavior only within its stated boundary. Physical systems can also have thresholds and discrete failures, such as a latch that either engages or fails to engage. The useful software-specific opportunity is not a universal binary-versus-continuous divide; it is access to structured interfaces, executable checks, and controlled environments.
Digital environments can make some verification and reset operations inexpensive. A local test fixture may expose a record read, a file check, or a reproducible snapshot. Production is different: an authorized read can be slow or stale, and an email or payment may already have caused irreversible harm before a check detects failure. Cheap checking does not bound the cost of being wrong. We therefore need to model both the transition and the evidence that establishes its outcome.
This chapter develops that view for web pages, desktop applications, and code interpreters. We separate full state, observations, beliefs, actions, asynchronous effects, and goals. GUI and API interfaces offer different observation channels; contracts and tests constrain covered behavior without replacing a complete dynamics model. We then examine rollout errors, controlled replay, hierarchy, and verification. A deterministic local issue-tracker experiment makes the distinction between predictive fidelity, corrected belief agreement, task completion, rejected actions, and recovery cost explicit.
Web, Desktop, and Code Execution State
A world-model state contains the variables needed to predict the next transition, as introduced in Part II. For a software system, these may include processes, files, database records, pending events, and clock-dependent conditions. The agent usually sees only a projection and maintains a belief over relevant states from its observation-action history. An observation is evidence about state, not a replacement for state.
A digital environment is a tuple where:
- : the set of possible sufficient system configurations, including relevant process, file, database, timing, queue, and external-input variables
- : the set of observations an agent can collect (screen pixels, DOM trees, API responses)
- : the action set (GUI events, HTTP requests, code evaluations)
- : the transition kernel giving the distribution over next states given the current state and action
- : the observation model giving the distribution over observations given the state
- : a set of task goals
An agent usually observes rather than , so partial observability must be handled explicitly. This goal-oriented tuple is not a complete standard POMDP specification: a conventional POMDP also specifies rewards and other evaluation conventions. Nothing in partial observability requires hidden states eventually to become distinguishable. See Kaelbling, Littman, and Cassandra's POMDP treatment for belief-state modelling.
The state can mix discrete variables, continuous clocks, buffers, and external inputs. Observations are renderings or partial serializations of that state. Actions may change as an interface changes. Transitions can be deterministic in a controlled fixture yet uncertain to an agent that lacks the relevant state; real systems may also introduce stochastic scheduling and external events. Distinguish uncertainty about hidden state from randomness in the transition.
What typically lives in the state? We can sort the components into visible observations, hidden state, actions, asynchronous effects, and task goals.
Visible observations are what an agent can capture directly:
- Rendered pixels of the current screen.
- A DOM snapshot or accessibility tree: distinct structural and semantic projections that can include or omit different interface elements.
- HTTP API responses: status and headers, plus any structured or unstructured response payload (which may be absent).
- Console or log output from an application or a code interpreter.
Each channel exposes only part of the state. Pixels may conceal offscreen or masked values. A DOM snapshot need not describe a native dialog or all canvas content. API responses expose the endpoint's authorized projection, not necessarily the whole database. Printed output is not the interpreter's full memory. For a desktop editor, for example, the current document text may be visible while unsaved buffers, the selected file path, focus, and a pending background save affect the next action.
Hidden state is what visible observations cannot disambiguate:
- Session and authentication state: cookies, bearer tokens, expirations, refresh timers.
- Permission scopes: which rows a user is authorized to read or write.
- Feature flags and A/B assignments: which version of a page the user is currently seeing.
- Server-side persistent state: rows in a database, objects in a blob store, entries in a search index.
- Rate-limit counters, retry budgets, and quotas.
- Background jobs: scheduled tasks, asynchronous workers, webhook deliveries.
What is hidden depends on the available tools and permissions. Browser tooling may expose cookie expiry, but that alone does not establish server-side session validity or revocation status. Some interfaces reveal feature assignments or quota headers; others do not. A variable can remain unidentifiable even after many permitted observations. An agent should record what each tool can establish rather than assume a universal read-back channel.
Actions include GUI events, structured API requests, and code evaluations. A GUI event may initiate an asynchronous request; an API response may acknowledge queued work rather than completion; a code statement may spawn a process that outlives the statement. The action boundary used by a model therefore needs an explicit timing convention, such as "response received" or "job completed".
Asynchronous effects include queues, background saves, webhook deliveries, and client reconciliation. An optimistic client may later reconcile with a server response, but that reconciled display can still lag another writer or read replica. A full-state model includes pending work and time-relevant variables. A model built only from the immediate screen can miss those delayed effects.
Task goals are usually compositional. A checkout flow decomposes into "purchase item", "enter shipping information", and "pay", each with its own preconditions and postconditions. Filing a bug report decomposes into opening a form, filling in the fields, and confirming the submission. Composition matters because the semantics of an action depend on the milestone being executed: the same click is a submission in one context and a navigation in another. When we later consider hierarchical decomposition, the goal structure is what makes that decomposition natural.
Why the Markov Approximation Usually Fails
The Markov property says that a sufficient full state screens off earlier history:
The practical failure is usually the approximation that a single observation is such a state. Two identical screens can correspond to different session or permission states and therefore different outcomes of the same Save click.
A history can contain clues absent from the current screen. A belief represents the remaining ambiguity. If all available histories are indistinguishable, a calibrated model can predict a mixture of outcomes; it need not pretend that one outcome is certain.
Wall-clock time, pending queues, random-generator state, and relevant external inputs belong in the state or transition convention when they affect future behavior. Including them can restore a Markov description at the full-state level. Their omission from an observation-level representation does not prove that the complete system violates the property. Changes to a UI, schema, or server implementation introduce a separate distribution-shift problem; physical environments can change too.
Code Execution State
Code execution has several state boundaries. The interpreter contains variables, scopes, and stack frames; the surrounding execution environment contributes files, process resources, and external effects. Relevant variables can include:
- The value of every variable in every active scope.
- The current instruction and call stack.
- The contents of the file system reachable from the process.
- The working directory, environment variables, and open file descriptors.
- Any child processes spawned by the program.
- Network sockets and buffered I/O.
A debugger, instrumented interpreter, or local test harness may inspect some of this state. A printed value or exit code inspects much less. A predicate such as result == 42 gives an answer about that result under the executed conditions. A successful file-existence check provides evidence about a path at a particular time, not about its contents; a false result may also reflect an inability to inspect the path. Unit and property-based tests verify their assertions and exercised cases, not every future transition.
Controlled code execution can help separate predictive fidelity from representation quality and decision usefulness. Compare predictions against explicit observed values, test which hidden variables a representation preserves, and evaluate the decisions it supports. The comparison is only as complete as the instrumentation and assertions. This chapter executes a closed, deterministic fixture; it does not execute arbitrary agent-generated code or contact live services.
The issue-tracker example will use a local read to correct a predicted observation. That isolates what feedback changes in the agent's belief and what it leaves unchanged in the predictor.
GUI and API Transition Models
A full-state learned model approximates the environment's transition kernel:
Here denotes the next state. In deployment, the agent generally cannot supply the true hidden directly.
Under the observation convention in our definition, the ideal observation distribution conditional on a known state is
The sum weights each possible next state's observation distribution by the probability of reaching that state. For continuous state spaces, the sum becomes an integral. A deployable prediction instead conditions on history and marginalizes the uncertain current state:
Here is the observation-action history and is the belief weight assigned to current state . The outer sum averages over that uncertainty; it does not require access to the true hidden state. The learned observation model approximates this marginal, possibly through an estimated latent representation. Conditioning only on is an additional approximation, appropriate only when it retains the relevant predictive information.
The following GUI and API categories organize interfaces, not an exhaustive taxonomy or a ranking of model families. Both can be combined with learned dynamics, explicit simulators, contracts, and verifiers.
GUI Transition Models
A GUI world model predicts changes in interface observations. Here we distinguish pixel-level and structural inputs; hybrid representations can use both.
A pixel-level model predicts the next screenshot from observations or history, an action, and possibly an estimated latent state. Video-prediction architectures from Part V and Part IX offer relevant tools, but an architecture's ability to accept screenshots does not establish transfer to any application. Exact text can be important: "carol@example.com" and "carol@example.org" differ in three character positions. Downsampling or compression may make text harder to distinguish, but this chapter does not demonstrate a raster collision between those strings.
A structural model uses a DOM, accessibility tree, or serialized widgets. These channels can expose exact text and control identities without asking an image decoder to reproduce them. They may omit visual details such as canvas content or an unlabeled icon, and their coverage depends on the interface. Model size and comparative accuracy depend on the architecture, representation, training data, and evaluation; neither follows from "structural" alone.
Neither channel automatically reveals server-side persistence. The same current screenshot may appear with and without a saved row in the authoritative store. That screenshot alone cannot identify which latent state holds. Earlier responses or later reads may distinguish them, and a model can represent uncertainty rather than being forced to predict "saved". Additional scale cannot create information absent from every available input, but improved history use or a new observation tool can change what is identifiable.
API Transition Models
An API observation model predicts structured responses or later observations from a request and the available history. A response can expose identifiers and validation outcomes more directly than a screenshot, but it need not reveal all server state. A successful POST can return different statuses depending on endpoint semantics: creation may return 201, acceptance for later processing 202, or a completed request with no response content 204. The protocol status and application contract must be interpreted together; RFC 9110 defines these success statuses.
A payload validator can reject a prediction that violates a schema. It cannot show that the server executed the predicted side effect. Some APIs supply structured error details; others return sparse codes, free text, or transport failures without a response. A useful model separates response format, documented semantics, hidden permissions, and the evidence of eventual persistence. This chapter makes no comparative claim about how quickly API and GUI agents learn.
Use only interfaces and observations the agent is authorized to access. An undocumented endpoint observed in a client's traffic may change without a stable contract, and observing it does not grant permission to invoke it. If no authorized API exposes the relevant evidence, GUI observations, local instrumentation, or a human confirmation may be necessary.
Explicit Contracts as Transition Models
Contracts describe constraints on behavior. An OpenAPI document can describe operations and response formats; a JSON Schema can constrain payload structure. Types constrain representable values, migrations describe database changes, and tests assert behavior in covered cases. None automatically specifies the complete transition kernel, including session revocation, concurrent writes, queues, and external effects. The OpenAPI specification describes an interface, while JSON Schema validation defines structural validation keywords.
- Learned dynamics estimate outcomes from data and can be used where rules are incomplete, but generalization to a new field or redesigned UI must be measured.
- Contracts and executable checks provide explicit constraints for their stated scope, but can be incomplete, stale, or inconsistent with the implementation.
An executable specification that truly defines all relevant transitions can serve as a simulator . A schema alone is not such a simulator. When available, use contracts to reject malformed predictions and tests to challenge predicted effects; use learned or hand-built dynamics for effects not specified by those constraints. These roles complement one another without establishing that one family generally outperforms another.
Hybrid Framings
A hybrid design can use the DOM to locate a control, an authorized API to obtain an identified record, and a learned model to predict how navigation changes the screen. This is one possible design, not a claim about the architecture of a typical deployed agent. It connects to the hybrid models discussed in Part V.
Track provenance: which constraint is documented, which behavior was tested, which outcome was observed, and which prediction is an extrapolation. Predictive entropy can reflect both transition randomness and uncertainty about the model, depending on the formulation; a single entropy value does not reveal that provenance or guarantee calibration.
Long-Horizon Task Simulation
Our issue-tracker task uses three actions. Real workflows can contain many more, including branches and dependencies: create a ticket, identify a duplicate, update the assignee, then verify the final record. The exact action count depends on the interface and task definition. Long-horizon simulation asks how prediction errors affect that sequence, not just whether the final screen looks plausible.
The reason long horizons are hard is compounding error. Suppose a transition model has per-step accuracy , defined as the probability of predicting a single step correctly, and per-step error probability , defined as the complementary probability that the model predicts a single step incorrectly. These two quantities are related by:
where:
- : the probability that the model predicts a single step correctly
- : the per-step error probability, defined so that
If correctness events were independent with the same probability , the probability of tracing steps correctly would be
since each of the steps must succeed independently. Here is the rollout horizon (the number of steps in the trace). The intuition is simple: every step is an independent chance to be wrong, and the chance of getting every one of steps right is the product of their individual success probabilities. With : we can compute this quantity for ten-step and fifty-step rollouts:
where:
- : the per-step success probability of the transition model, equal to
- : the per-step error probability; for , the model is wrong about one in every twenty steps
- : the probability that a rollout of steps is traced correctly under the assumption that step errors are independent
- : the rollout horizon, i.e. the number of steps in the trace
Under this independent, constant-probability model, nineteen out of twenty correct predictions per step still yield a low chance of an entirely correct long trace. At fourteen steps the probability of at least one error is about fifty-one percent; at fifty it is about ninety-two percent. These are clean-prediction probabilities, not task-success measurements. A real model's errors can be correlated and concentrated on decisive actions, so a measured marginal one-step accuracy cannot generally be exponentiated. For example, if an entire episode is either wholly correct or wholly wrong, every step can have marginal accuracy while the all-correct episode probability remains , not .
The same product can be plotted directly instead of evaluated at two points. The next cell evaluates on a horizon grid and records the ten-step and fifty-step checkpoints that the chapter quotes. The chart cell that follows only renders those precomputed values.
import numpy as np
# Closed-form evaluation of the independent-error product (1 - eps)^k.
# No sampling is involved, so no random seed is required. The three error rates
# span an optimistic regime (0.01), the chapter's working value (0.05), and a
# pessimistic regime (0.20). The horizon grid runs from one to one hundred
# steps, and the checkpoints are the ten-step and fifty-step horizons quoted in
# the text.
error_rates = [0.01, 0.05, 0.20]
horizons = np.arange(1, 101)
success_curves = {eps: (1.0 - eps) ** horizons for eps in error_rates}
checkpoints = [10, 50]
Where the Errors Compound
The independent product is an illustration, not a bound on actual software tasks. Three mechanisms deserve separate attention.
First, an incorrect predicted observation can redirect planning. If a model treats an unfilled field as filled, later predictions may use an incorrect state estimate. This can propagate through an autoregressive rollout, but need not corrupt every later prediction: dynamics, new observations, and policies can attenuate or correct the error. Interleaving observations helps only when they expose the relevant discrepancy.
Second, some effects cannot be reversed by a simple inverse action. A completed order may require cancellation or a refund, which is a new workflow rather than a reset. For high-impact actions, verify preconditions, permissions, and user authorization before acting as well as checking effects afterward; post-hoc detection cannot undo all harm.
Third, queues, caches, and webhooks can delay effects beyond the immediate response. Include pending events and timing in the full state when modelling them. Delayed consequences can make a screen-only representation non-Markovian without contradicting a sufficient full-state Markov model. They also complicate credit assignment and the choice of when a verifier should check.
These mechanisms concern different failures: propagation of a mistaken estimate, costly recovery from an irreversible effect, and incomplete observation of delayed work. A single drift score cannot diagnose all three.
Snapshots, Reset, and Replay
In a controlled local environment, snapshots and replay can support inexpensive branching. A valid snapshot contains all relevant state within the controlled boundary, including pending work, time conventions, and random-generator state when needed. Restoring only a database is insufficient if an external queue or client buffer still differs. Chandy and Lamport's distributed snapshot algorithm illustrates why process and channel state both matter. With that qualification, a branch samples
where:
- : the next state sampled on the counterfactual branch
- : the true transition kernel
- : the fixed snapshot state
- : the alternative action replayed from the snapshot
This compares alternative actions from an equivalent controlled starting condition. As discussed earlier in Part VII, it can support search and belief-space planning. Snapshot storage and replay still have a cost, and experimental isolation is part of the design rather than an automatic property of software.
A production snapshot cannot retract an email or erase a payment's external consequences. Some internal changes are reversible, while others need compensating workflows, audit records, or approval. The reliable-world-model discussion in Part XII builds on this boundary between controlled experimentation and real intervention.
Hierarchical Decomposition
One way to organize a checkout flow is as a sequence of milestones with local actions beneath them. "Open the cart" and "confirm shipping address" can be separate milestones. The useful boundaries depend on the application and task; the analogy does not establish how every human or agent plans.
If a task has milestones with about actions each, a planner can make local rollouts of roughly actions rather than one rollout of . This reduces the horizon of an individual planning call, not the total number of actions or the probability of an entirely error-free task. Local repair works when milestone boundaries are meaningful and errors do not invalidate earlier dependent work. Milestone outcomes and high-level decisions can themselves be uncertain. See the earlier hierarchical planning discussion in Part VII.
Interface structure can suggest milestones, preconditions, and postconditions. Using that structure may reduce a planner's search space. Whether it reduces data requirements or improves task completion is an empirical question, especially when the documented flow omits exceptions.
Reset and Verification Together
Verification and local replanning address different parts of the loop. A reliable observation can correct the variables it reveals; replanning can change the actions chosen from that revised belief. Neither repairs hidden variables automatically. The verification schedule should reflect information quality, action risk, and cost.
An agent can alternate prediction, action, observation, correction, and replanning instead of committing to one long imagined trajectory. This can limit uncorrected drift in observed variables. It is not a guarantee of stability for arbitrarily long tasks: missed effects, unavailable evidence, repeated errors, and unrecoverable actions remain possible.
The next calculation isolates one limited benefit: under independent errors with probability 0.05 per step, a segment of steps is more likely to be predicted cleanly when is shorter. It does not calculate the success of a hundred-step task. If every one of those hundred predictions must be correct, then splitting them into twenty five-step segments still gives . Task success with correction additionally depends on detection, repair, cost, and the consequences of mistakes.
# Independent, constant-error illustration: these are clean segment
# probabilities, not whole-task success probabilities under correction.
task_horizon = 100
verified_epsilon = 0.05
verify_intervals = [1, 2, 5, 10, 25, 50, 100]
interval_success = {
v: (1.0 - verified_epsilon) ** min(v, task_horizon)
for v in verify_intervals
}
no_verification_success = (1.0 - verified_epsilon) ** task_horizon
segmented_full_clean = interval_success[5] ** (task_horizon // 5)
assert np.isclose(segmented_full_clean, no_verification_success)
Verifiable Actions and Environment Feedback
Many software effects can be checked through an additional observation, but a check is evidence only for the variables and conditions it covers. A useful read-back must come from an authoritative source, be authorized, have suitable consistency and freshness, and identify the result attributable to this action. Reading an optimistic client cache or a lagging replica may leave the original ambiguity unresolved.
A postcondition is a predicate expressing a desired property after an action. Checking requires an observation that reveals the variables the predicate needs. A true predicate is not automatically proof that this particular action caused the property: the record may already have existed or another writer may have created it.
After creating an issue, a practical verifier might read an authoritative record by the returned unique ID and check its fields, ownership, and expected version. A title appearing in a list is weaker evidence because titles need not be unique. An aggregate count can increase for unrelated reasons. In the controlled toy below, the fresh store, single writer, synchronous action, and complete local read make a count increment attributable; those assumptions must not be transferred unexamined to a live application.
Verification Is Not Reward
Reward and verification serve different roles, but their information content depends on their definitions. A reward is a scalar used by an objective; a Boolean postcondition records whether a specified property holds. If
then and the Boolean encode exactly the same single bit. Neither distinguishes "missing record" from "wrong title" when both make false.
Richer diagnostics require a separate structured observation, for example
Here the components describe what the verifier established; an unavailable check should not be reported as false. A scalar reward can also encode different failures or be shaped by the designer. The useful distinction is the objective's use of a signal versus the verifier's evidence about specific properties, not a universal rule that every verifier contains more information than every reward.
Verification can supply local evidence soon after an action, while a task reward may arrive much later. But asynchronous effects can require delayed verification, and rewards can also be immediate. Tie both signals to an explicit timing and attribution convention.
Types of Failures That Verification Catches
Read-back can reveal failures hidden by an immediate optimistic display, when the observation channel covers them.
- Silent rejection: the server rejects the request but the UI does not show the error. For example, the application might reject an expired session, an unsupported schema version, or an action disabled by a feature flag. The click alone does not establish its effect.
- Optimistic UI divergence: the client renders a success state before the server responds. The server may still fail. Verification must read the server state, not the client's optimistic rendering.
- Idempotency violations: retrying an action without an idempotency key can duplicate a purchase, a message, or a database row. Verification can detect the duplicate.
- Permission masks: an action appears available, but the server returns 403, indicating refusal. Check the relevant intended effect rather than assuming the entire system is unchanged; logging, concurrent actions, or earlier partial effects may still alter state.
- Delayed effects: the effect is real, but it has not yet occurred. Verification at step sees nothing, while verification at step , with where is the rollout horizon, sees the result. Timing matters when the effect is real but delayed.
The listed checks are conditional. A stale response may conceal a duplicate, a permission-limited list may omit an existing record, and an absent queued effect may mean "not yet" rather than "failed". A terminal status, version, idempotency identifier, or bounded polling convention may be needed. Verification does not grant access to otherwise unauthorized state.
The Verification Budget
Verification is not free. Local assertions and membership tests consume local resources but need not make a remote request; request-based checks can face latency, rate limits, or noisy responses. An agent that verifies after every action can spend more effort checking than acting, especially when each check requires reading a large list of records. Scanning thousands of rows to see whether one was added can be costly in a remote or metered collection; the cost depends on data size, pagination, consistency requirements, and the available budget.
Choose checks according to risk and information quality. High-impact actions require appropriate precondition and authorization checks even when a model is confident; low entropy is not permission to skip them. Read cost, rate limits, freshness, and the availability of a safe recovery all matter.
One possible heuristic triggers an extra observation when
Here is the available history and uses the same entropy units as : bits for logarithms base two, nats for natural logarithms. The strict inequality means equality alone does not trigger this rule. Predictive entropy can mix aleatoric and epistemic uncertainty, and a poorly calibrated model can be confidently wrong.
This is a heuristic, not a value-of-information calculation. A decision-theoretic alternative estimates how the proposed observation changes expected downstream utility, then compares that benefit with its cost and risks. An observation with little entropy reduction can still be crucial for an expensive decision; a high-entropy variable can be irrelevant. This connects to uncertainty in Part IV and the earlier active-perception and dual-control discussion in Part VII.
Task Success Is Not World-Model Fidelity
Task success is one useful evaluation metric for software agents: did the agent complete the goal? This is useful but reductive. A single task-success number folds together at least four distinct capabilities:
- The quality of the transition model, if one is used.
- The quality of the planning or policy on top of it.
- The ability to observe relevant state, including hidden parts.
- The ability to recover from failures.
Two agents can both achieve 60% task success for entirely different reasons. One has a great model and a bad recovery policy; the other has a bad model but verifies and corrects on every step. The first is a good model applied imperfectly, and the second is a weak model compensated by a strong loop. The same score tells us nothing about which of these is happening, and so it tells us nothing about where to invest effort next. To understand why an agent succeeds or fails, we have to measure the components separately. We will develop this idea systematically in Part XI: Evaluation and Understanding, especially in the chapter on the world-model evaluation ladder.
A related point is that task success mixes predictive fidelity, representation quality, uncertainty quality, and decision usefulness into one number. A model can predict well and represent badly, or represent well and predict no better than chance on the specific transitions the task requires. Disentangling these is a prerequisite for making sense of benchmark results. Instrumented software fixtures can support separate evaluation when they provide explicit reference measurements for those components; representation quality is not automatically observable in every software task.
A Worked Experiment: Optimistic vs Verified Predictors
Let us make this concrete. We will build a deterministic local issue tracker with two hidden facts, a session validity flag and a team membership roster, and then compare three agents:
- Optimistic: predicts transitions using only the visible form fields.
- Verified: uses the same model but performs a read-after-write correction after every action.
- Verified + recovery: additionally attempts a targeted recovery strategy when a save fails.
The three agents share an identical transition model, which isolates the contribution of verification and recovery from the contribution of the model itself. This is deliberate: the point of the experiment is not to demonstrate that a small state machine is inadequate, but to show that the same inadequate model can succeed or fail depending on the loop around it.
For each run we report these quantities separately:
- Raw one-step prediction accuracy: exact agreement of the prediction made before the action with the resulting observation, scored before any correction.
- Post-correction belief agreement: agreement after the optional read-back correction.
- Final observable drift: the number of mismatched components of the final observation triple; this is not full hidden-state drift.
- Task completion: whether the intended title-assignee tuple exists in the fresh, single-writer store.
- Rejected Save attempts: every actual Save click that appends no row, including retries.
- Successful recoveries and recovery-action cost: failed primary saves repaired by the retry policy, and additional actions it executes.
The environment exposes the observation
where is the issue count. Its full toy state includes the complete issue list, form fields, session-validity flag, and the roster set . A Boolean membership answer for the current assignee is not enough to predict all future assignee choices; the roster set is the relevant fixed environment parameter.
The local fixture has no queue, clock, concurrent writer, or network. Those omissions make its transitions deterministic. They are deliberate experimental boundaries, not claims about real trackers.
Let us define the environment and the optimistic model.
import numpy as np # noqa: E402
TEAM = {"ada", "bob"} # hidden server-side roster
class IssueTracker:
"""A tiny local issue tracker with hidden session and team state."""
def __init__(self, session_valid: bool):
self.session_valid = session_valid # hidden
self.title = None # visible form field
self.assignee = None # visible form field
self.issues = [] # persistent store
def observe(self):
return (self.title, self.assignee, len(self.issues))
def step(self, action):
kind, arg = action
if kind == "type_title":
self.title = arg
elif kind == "type_assignee":
self.assignee = arg
elif kind == "relogin":
self.session_valid = True
elif kind == "click_save":
ok = (
self.session_valid
and self.title
and self.assignee
and self.assignee in TEAM
)
if ok:
self.issues.append((self.title, self.assignee))
return self.observe()The class exposes only the visible triple through observe(). The click_save branch is where things get interesting: it is silently a no-op unless every hidden condition holds. In state-transition terms, the transition conditioned on appends to the issue list only when and ; otherwise the transition is the identity on the persistent store. This is a deterministic kernel, so the distribution places all its mass on a single next state.
The optimistic model is a transition function over the visible triple. Formally, it approximates
using only , ignoring the hidden components of . It reasons about required fields but has no access to session validity or team membership.
def optimistic_step(state, action):
"""Transition model over visible state only."""
title, assignee, n_issues = state
kind, arg = action
if kind == "type_title":
title = arg
elif kind == "type_assignee":
assignee = arg
elif kind == "click_save":
if title and assignee:
n_issues += 1
return (title, assignee, n_issues)This is a hand-written optimistic predictor, not a learned model. It models the required visible fields but ignores session validity and the roster. Its Save prediction is wrong in the two hidden-precondition failures. The comparison therefore illustrates feedback around an imperfect predictor; it does not estimate the performance of a trained GUI or API model.
The driver executes the same three primary actions in each mode against a fresh tracker. It logs a raw prediction before every real environment action, then stores the resulting observation and optionally corrected belief. Recovery adds relogin and one Save retry, each logged and scored separately. The policy tries a bounded repair without first identifying the cause; it can therefore repeat a failure when the assignee is invalid.
# NOTE: this driver mutates environment state and runs the simulation loop, so it is
# explicitly not theme-rerunnable; the plot cells below only read results from `results`.
def run(session_valid, actions, mode):
if mode not in ("optimistic", "verified", "verified_recovery"):
raise ValueError("Unknown policy mode")
if not actions:
raise ValueError("This experiment requires a nonempty action sequence")
verify = mode in ("verified", "verified_recovery")
recover = mode == "verified_recovery"
env = IssueTracker(session_valid)
belief = env.observe()
steps = []
recovered = 0
def execute(action, phase):
nonlocal belief
before = len(env.issues)
raw_prediction = optimistic_step(belief, action)
observation = env.step(action)
# Score the prediction before correction; keep both in the audit log.
raw_match = int(raw_prediction == observation)
belief = observation if verify else raw_prediction
record = dict(
action=action,
phase=phase,
raw_prediction=raw_prediction,
observation=observation,
belief=belief,
raw_match=raw_match,
belief_match=int(belief == observation),
rejected_save=int(
action[0] == "click_save" and len(env.issues) == before
),
)
steps.append(record)
return record
for action in actions:
record = execute(action, "primary")
if recover and record["rejected_save"]:
execute(("relogin", None), "recovery")
retry = execute(action, "recovery")
recovered += int(not retry["rejected_save"])
final_observation = env.observe()
drift = sum(int(a != b) for a, b in zip(belief, final_observation))
primary = [step for step in steps if step["phase"] == "primary"]
# The goal is derived from the requested form values, not a hardcoded result.
goal_title, goal_assignee = None, None
for kind, arg in actions:
if kind == "type_title":
goal_title = arg
elif kind == "type_assignee":
goal_assignee = arg
return dict(
steps=steps,
primary_raw_accuracy=sum(s["raw_match"] for s in primary)
/ len(primary),
raw_accuracy=sum(s["raw_match"] for s in steps) / len(steps),
belief_agreement=sum(s["belief_match"] for s in steps) / len(steps),
action_count=len(steps),
final_drift=drift,
task_complete=int((goal_title, goal_assignee) in env.issues),
save_attempts=sum(s["action"][0] == "click_save" for s in steps),
invalid_actions=sum(s["rejected_save"] for s in steps),
recovered=recovered,
recovery_actions=sum(s["phase"] == "recovery" for s in steps),
)Three scenarios cover the interesting cases: a happy path, an expired session, and an unknown assignee who is not on the team roster.
scenarios = {
"happy": (
True,
[
("type_title", "Login fails"),
("type_assignee", "ada"),
("click_save", None),
],
),
"expired_session": (
False,
[
("type_title", "Crash on save"),
("type_assignee", "bob"),
("click_save", None),
],
),
"unknown_assignee": (
True,
[
("type_title", "Flaky test"),
("type_assignee", "carol"),
("click_save", None),
],
),
}
modes = ["optimistic", "verified", "verified_recovery"]
results = {}
for name, (sv, acts) in scenarios.items():
for mode in modes:
results[(name, mode)] = run(sv, acts, mode)The table reports primary-action accuracy over the common three actions, all-action raw accuracy with its denominator, post-correction agreement, observable drift, task completion, every rejected Save, repaired saves, and recovery-action cost. Recovery rows can contain five actions; the other rows contain three. The all-action scores therefore have different denominators and should not be mistaken for identical test sets.
scenario mode primary raw N belief drift done saves reject repair extra --------------------------------------------------------------------------------------------------------- happy optimistic 1.00 1.00 3 1.00 0 1 1 0 0 0 happy verified 1.00 1.00 3 1.00 0 1 1 0 0 0 happy verified_recovery 1.00 1.00 3 1.00 0 1 1 0 0 0 expired_session optimistic 0.67 0.67 3 0.67 1 0 1 1 0 0 expired_session verified 0.67 0.67 3 1.00 0 0 1 1 0 0 expired_session verified_recovery 0.67 0.80 5 1.00 0 1 2 1 1 2 unknown_assignee optimistic 0.67 0.67 3 0.67 1 0 1 1 0 0 unknown_assignee verified 0.67 0.67 3 1.00 0 0 1 1 0 0 unknown_assignee verified_recovery 0.67 0.60 5 1.00 0 0 2 2 0 2
Reading the row for expired_session under optimistic: the model believes the save worked, but the environment silently rejected it. Drift is 1 and task completion is 0. The model is confidently wrong. The unknown_assignee scenario under optimistic fails the same way, because the model has no idea that carol is not on the team. Recall that drift counts the number of visible components on which the model's final belief and the true final observation disagree. Writing the three components as indexed entries of the observation vector, we have:
where:
- : the final step index of the rollout
- : the -th component of the model's final predicted observation
- : the -th component of the true final observation
- : the indicator function, equal to 1 when its argument is true and 0 otherwise
- the index ranges over the three visible components (title, assignee, and number of issues)
The verified rows have post-correction agreement 1.00 and final observable drift zero because this fixture's complete local observation is copied into the belief. Their raw primary-action accuracy remains 2/3 in the failing scenarios. Verification has not made the predictor perfect; it has corrected its exposed errors. Verification alone still leaves both tasks incomplete.
For an expired session, relogin plus retry repairs the failed Save. The initial rejection remains counted, so this run has two Save attempts, one rejection, and two extra recovery actions. For an unknown assignee, both the initial Save and the retry are rejected: two Save attempts, two rejections, two extra actions, and no completed task. Recovery is useful only when its actions address the actual precondition failure.
Let us visualize the drift comparison first.

For the common primary actions, raw accuracy is in both failing scenarios for all three modes. Let be the number of real actions included in a score and the prediction recorded before correction:
Post-correction agreement uses the corrected belief instead and measures a different property. The logged records retain both values so the verifier cannot erase prediction errors from the evaluation.
The primary traces end at the first incorrect prediction. They demonstrate that an aggregate over easy typing actions can hide the decisive Save miss, not that this miss propagated over a long rollout. In recovery mode the expired-session trajectory has four correct raw predictions among five actions; the unknown-assignee trajectory has three among five. These are trajectory-dependent scores: verified modes feed corrected beliefs into later predictions, and the recovery policy also changes the actions.
Now let us examine task completion and all rejected Save attempts under the recovery policy. These counts were computed in the experiment cell above (the verified_recovery rows).

The chart retains the initial failed Save even when recovery succeeds. Relogin changes session validity but cannot add carol to the roster, so the unknown-assignee retry is rejected too. Task completion and rejected attempts answer different questions: a successful task can still incur a failed action and a recovery cost.
What the Experiment Shows
The table supports three limited conclusions.
First, the optimistic predictor's common three-action accuracy is in each failing scenario while task completion is zero. Two correct typing predictions do not compensate for the decisive Save miss. The short experiment does not demonstrate long-horizon compounding.
Second, the fixture's observation channel corrects observable drift without changing raw prediction accuracy on the common actions. This illustrates belief correction under an exact, fresh local read. It does not establish that a real authorized read covers every hidden variable or that correction increases full-task success by a particular factor.
Third, recovery and detection differ. The bounded relogin-retry policy fixes the expired session but fails on the unknown assignee, adding two actions in either case. Recovery can use search, retry, a fallback, or an explicit causal hypothesis; successful repair does not always require a named diagnosis. More targeted repair of the assignee would require authorized roster information or another justified action. A supplementary observation channel could be written
where reveals allowed assignees under the channel's permissions. This is additional evidence, not something the original observation triple already contains.
Limitations and Impact
Real software is much messier than the small state machine this chapter has examined. A few limitations deserve careful spelling out.
Heterogeneity. A spreadsheet, desktop editor, and issue tracker have different observations, transition conventions, and task goals. Transfer must be tested rather than inferred from a common "software" label. Modular observation tools and workflow rules are one possible design; this chapter provides no evidence that such scaffolding generally dominates learned-model performance.
Learned dynamics and explicit contracts. Use a schema to constrain structure, a test to verify a covered assertion, and a dynamics model to predict effects. Only a complete executable specification justifies writing a specified transition kernel
That notation is not a claim that an OpenAPI document, type, or test suite already supplies the kernel. A hybrid can use formal constraints wherever their coverage is established and learned predictions where uncertainty remains. Its comparative performance needs evaluation.
Verification is neither free nor universal. A read can be unauthorized, stale, expensive, or rate-limited. Deletion does not always make evidence impossible: a tombstone or audit log may remain, if the agent may read it. Conversely, a visible record need not be attributable to the action under evaluation. Treat verification as an information-gathering action with explicit costs, as in the earlier Part VII discussion.
Task success conflates capabilities. The experiment made this numerically explicit: two policies can achieve different task-success scores for entirely different reasons. We can write task success as a composition, though the decomposition below is schematic rather than an exact identity:
where:
- : the overall task-success rate
- : a composite function bundling four capability factors into one scalar
- : the quality of the transition model
- : the quality of the planning or control policy
- : the coverage and reliability of state observations
- : the ability to detect and correct failures
A single scalar does not identify which capability caused a success or failure. Report predictive fidelity and calibration, observation coverage, policy performance, recovery outcomes, and costs separately. Part XI develops this evaluation perspective further.
Reset has a controlled boundary. Replays are useful when the fixture restores every relevant state component and external effects are isolated. Production may instead require compensation, audit-preserving corrections, or manual approval. A policy trained with low-cost reversibility should be reevaluated when those assumptions no longer hold.
Software fixtures let us examine prediction and feedback with explicit states and executable checks. The benefit comes from the chosen observation and control boundary, not a guarantee that software exposes truth cheaply. The next chapter turns to scientific, medical, and social environments, where the evidence available after an intervention can be delayed, incomplete, or confounded.
Summary
Digital and software environments combine structured state, interface-dependent observations, and workflows whose outcomes may be checked through authorized tools. The critical distinction is between a predicted observation, a corrected belief, and evidence that a task's intended effect occurred.
- A screenshot, DOM tree, or API response is a partial observation. Relevant full state can include sessions, permissions, files, database records, clocks, queues, and external inputs.
- GUI and API models predict through different channels. Contracts and tests constrain covered behavior; they are not automatically complete transition kernels.
- Independent-error products describe clean prediction traces under stated assumptions, not generic task success. Hierarchy shortens local planning calls without eliminating all-task risk.
- Verification can correct exposed discrepancies when its source, authorization, consistency, freshness, and attribution are adequate. A Boolean predicate and its indicator reward contain the same outcome bit; structured diagnostics require additional evidence.
- Score raw predictions before correction. Report observable drift, task completion, every rejected attempt, successful recovery, and added recovery actions separately.
- Controlled replay requires a complete reset boundary. Cheap checks or reversible local fixtures do not make harmful production actions safe.
The next chapter turns to scientific, medical, and social worlds. It examines how observation quality, intervention, and reversibility change the prediction-and-verification loop. Scientific simulations and instrumented experiments may offer precise checks too. What those checks establish depends on the measurements and controls available.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about digital and software environments.
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of World Models Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore World Models HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!