Planning: Task Decomposition and Goal-Directed LLM Agents

Michael BrenndoerferJanuary 1, 202661 min read

Part of Language AI Handbook

Explains how LLM agents plan, decompose tasks, execute multi-step goals, and recover from failures using architectures like ReWOO, Tree-of-Thought.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Planning

Language model agents can answer a question, call a tool, and synthesize a result in a single step. But what happens when the task requires ten steps, some of which depend on the outcomes of earlier ones, and the path to success is not obvious at the start? Single-step reasoning breaks down. The agent commits to an action too early, gets stuck when something unexpected happens, or loses track of the overall goal in a tangle of sub-problems.

Planning is the capacity to reason about sequences of actions before taking them. A planning agent does not immediately act. It decomposes the goal into subtasks, estimates what each subtask requires, identifies dependencies between them, and only then begins execution. When execution goes wrong, a good planner revises the plan rather than abandoning the goal entirely. This is not a minor engineering detail. It is the difference between an agent that can be delegated a complex goal and one that requires constant supervision to stay on track.

This chapter covers the core ideas behind LLM-based planning: what planning means in the context of language model agents, how task decomposition works, the main planning architectures available today, how plans get executed and monitored, and how agents recover when plans fail. We also implement a working planner from scratch that demonstrates these ideas in code. Building on the ReAct pattern and agent architectures covered in earlier parts of this book, planning adds a layer of deliberate foresight that separates simple tool-calling agents from goal-directed systems capable of multi-step autonomous work.

What Planning Means for LLM Agents

Planning in classical AI refers to the problem of finding a sequence of actions that turns an initial state into a goal state. In formal planning systems like STRIPS (Stanford Research Institute Problem Solver), actions have explicit preconditions and effects, and a planner searches through a state space to find a valid path from start to goal. This formalism works beautifully when the world is fully observable and the action space is small, but it becomes computationally intractable for real-world tasks where state spaces are enormous and actions are hard to enumerate. A robotics task with a few dozen states and ten actions is tractable. Planning how to complete a software engineering task involving hundreds of files and unbounded possible code changes is not.

LLM-based planning takes a different approach. Instead of explicit state-space search, the language model itself acts as the planner, using its broad knowledge to generate plausible action sequences in natural language. The planner does not search exhaustively through possibilities; it generates a reasonable plan based on its training distribution and then adapts as execution reveals new information. This is closer to how humans plan than how classical AI planners work. You do not enumerate all possible routes from your apartment to a meeting before leaving. You recall roughly how to get there, start walking, and adjust when you encounter a closed street or unexpected construction.

Plan

A plan is an ordered set of steps that, if executed correctly, achieves a specified goal. In the LLM agent context, each step typically maps to a tool call, a sub-query to another model, or a reasoning action. Plans can be fully determined upfront (static plans) or constructed incrementally as the agent learns more about the task (dynamic plans).

The key difference between a planning agent and a simple ReAct agent is the scope of deliberation. A ReAct agent thinks one step ahead: it reasons about what to do next, acts, observes the result, and repeats. A planning agent thinks many steps ahead: it considers the full structure of the task before committing to any action, identifies which steps can be parallelized and which must be sequential, and maintains a representation of the overall goal throughout execution. The ReAct loop is reactive; the planning loop is anticipatory.

This distinction matters for several reasons. First, some tasks have global constraints that only become visible when you look at the whole problem at once. Booking a trip that satisfies flight and hotel constraints within a calendar schedule is harder to handle step-by-step than with a plan that considers all three dimensions from the start. A reactive agent might book a flight, then discover no hotel is available at the destination on those dates, and have to restart entirely. A planning agent checks availability across all dimensions before committing to any booking, avoiding the wasted effort. Second, planning reduces redundant work by ordering steps to avoid re-doing computations. If the plan recognizes that three different research questions all require information from the same source document, it retrieves that document once and passes it to all three questions, rather than fetching it three times. Third, a plan provides a scaffold for recovery: when step four fails, the planner knows what steps three and five are and can reason about how to get back on track, whereas a reactive agent has no pre-existing structure to repair.

The Spectrum of Deliberation

It is useful to think of planning capability as a spectrum rather than a binary. At one end is pure reaction: the agent decides what to do based only on the immediate observation, with no representation of future steps. At the other end is full symbolic planning: the agent constructs a complete, verified action sequence before taking any action at all. LLM-based agents occupy a large middle region of this spectrum, and different architectures sit at different positions.

Most production systems combine some upfront deliberation with ongoing adaptation. A good planning agent commits enough upfront to gain the benefits of structured reasoning, while remaining flexible enough to adapt when the world does not cooperate. Finding this balance is one of the central design challenges in building planning agents. Commit too much upfront and the agent becomes brittle when the environment deviates from assumptions; commit too little and planning provides no benefit over pure reaction.

Understanding this spectrum also helps explain why planning capability is hard to benchmark reliably. A task that requires only a few predictable steps looks like a win for lightweight reactive agents. A task with ten steps, complex dependencies, and multiple potential failure modes strongly favors deliberate planning. The right architecture always depends on the task structure, and benchmark results from one domain do not necessarily transfer to another.

Why Planning Fails Without Structure

To appreciate what planning provides, it helps to examine the specific failure modes that arise when agents act without it. Consider a research agent tasked with writing a comparative analysis of two competing scientific theories. Without planning, the agent begins immediately: it searches for the first theory, gets some results, searches for the second theory, gets some results, and then attempts to write a comparison. But the comparison requires understanding what dimensions to compare across, which requires knowing both theories well before deciding what to show. A planning agent would recognize this upfront and structure the task differently: first identify the key comparison dimensions, then gather evidence for each dimension across both theories, then write the comparison section by section.

The structural flaw in reactive execution is that commitment and information arrive in the wrong order. The agent commits to an approach (searching for theory A first) before it has the information needed to know whether that approach is sensible (what dimensions matter for the comparison). Planning defers commitment until enough information has been specified to make good decisions. This is the same reason architects draw blueprints before construction workers begin building, and why programmers design APIs before implementing them.

Task Decomposition

The foundation of planning is decomposition: breaking a complex goal into simpler subgoals that are individually tractable. This is where most planning agents begin, and it is the step that most strongly determines whether the plan will succeed. A bad decomposition cannot be rescued by a good executor; it produces failures that look like execution problems but originate in planning.

Task Decomposition

Task decomposition is the process of taking a high-level goal and recursively subdividing it into smaller, more concrete subtasks until each subtask can be accomplished by a single action or tool call.

Decomposition makes individual steps easier and makes the structure of the problem explicit. An agent that has decomposed "write a research report on climate change" into "gather sources", "extract key findings", "organize by theme", and "write the report" has a shared vocabulary for discussing what needs to happen, a framework for monitoring progress, and a natural scaffold for recovery if any piece goes wrong. Without decomposition, failure is monolithic: either the whole task succeeds or it fails. With decomposition, failure is local: one subtask fails, and the rest of the structure remains intact.

Hierarchical Decomposition

Goals naturally form hierarchies. "Write a research report on climate change" decomposes into "gather sources", "extract key findings", "organize findings by theme", and "write the report". Each of these decomposes further: "gather sources" breaks into "search academic databases", "search news archives", and "identify expert opinions". At the bottom of the hierarchy are leaf tasks that correspond to concrete actions: run a search query, fetch a URL, extract text from a document.

Hierarchical decomposition has a critical property: the planner can reason at the appropriate level of abstraction for each decision. High-level sequencing decisions, such as what major phases to tackle first, do not need to know the details of low-level implementation. Low-level execution does not need to be aware of the strategic reasoning that led to the current subtask. This separation of concerns makes planning tractable. When you are writing the "gather sources" phase, you do not need to think about how you will format the final report. When you are formatting the final report, you do not need to remember exactly how you ran each database query.

In practice, LLM planners implement hierarchical decomposition through prompting. The model is asked to break the goal into top-level steps, and then each step is recursively expanded as needed. The depth of decomposition is controlled by whether a subtask is "atomic", meaning directly executable as a tool call, or "composite", meaning it still requires further breakdown. Atomic tasks are the leaves of the tree. Composite tasks are internal nodes that the planner must continue breaking down before they can be executed.

The key challenge in hierarchical decomposition is knowing when to stop. Decomposing too deeply produces a plan with hundreds of trivially small steps that creates coordination overhead without benefit. Decomposing too shallowly produces vague instructions that the executor cannot act on directly. A search step that says "search for relevant sources" is too vague; a search step that says "search Google Scholar for papers on climate feedback loops published after 2020" is appropriately specific. Calibrating this granularity requires the planner to have some knowledge of the executor's capabilities, which in practice means knowing what tools are available and what each tool can reliably do.

Sequential vs. Parallel Decomposition

Not all subtasks need to be completed in order. Understanding which tasks are independent enables parallelism that dramatically reduces wall-clock time on multi-agent systems, and also simplifies reasoning about the structure of the plan.

Sequential dependency occurs when the output of one task is the input to another. Searching for relevant documents must precede summarizing those documents. The search results are a required input to the summarization step. No amount of parallel execution can change this: you cannot summarize a document before you have it.

Independent tasks, by contrast, can be run in any order or simultaneously. Searching three different databases for background on a topic does not require the first search to complete before the second begins. The databases are separate systems with no shared state, and the results of one search are not inputs to the others. A parallelized planner can dispatch all three searches at once and wait for all results before proceeding, potentially completing three steps in the time it would take a sequential planner to complete one.

Identifying independence correctly is non-trivial for LLMs, and errors in either direction are costly. A planner that wrongly treats a sequential dependency as parallel will attempt to use outputs before they exist, generating errors or hallucinations. A planner that wrongly treats parallel tasks as sequential wastes time by running them one after another. In practice, LLMs tend to be conservative, defaulting to sequential orderings even when parallelism would be safe. Explicit prompting to reason about data dependencies before assigning step orderings significantly improves the quality of parallelism detection.

Parallelism is useful only in multi-agent or multi-threaded execution environments. A single-threaded planning agent running one step at a time gains nothing from identifying parallel steps during execution, though it can be confident it is not missing a dependency constraint when it runs them sequentially. In systems with multiple worker agents or parallel tool execution, correctly identifying independent steps translates directly into shorter task completion times. For long research tasks with many independent information-gathering steps, the wall-clock time reduction from parallelism can be substantial.

Out[3]:
Visualization
DAG diagram showing steps 1-4 in parallel feeding into step 5 with dependency arrows.
Dependency graph for a 5-step climate research plan. Steps 1 through 4 are independent searches that can execute in parallel (shown at the same level), while step 5 depends on all four and must wait until each search completes. Arrows indicate data flow. A parallel executor running this plan completes it in the time of two sequential steps rather than five.

Decomposition Strategies

The strategies available for task decomposition range from purely instructive (tell the model how to decompose) to structurally guided (ask the model to reason about dependencies explicitly). The right strategy depends on the complexity of the task and the quality of the model's knowledge about the domain. For familiar, well-structured domains, simple chain-of-thought decomposition usually suffices. For novel or loosely specified tasks, more structured approaches improve reliability by forcing the model to reason about what it does not yet know before committing to a path.

Several prompting-based decomposition strategies have emerged in practice:

Chain-of-thought decomposition instructs the model to think step-by-step before acting. The model writes out a sequence of steps, and this sequence becomes the plan. This is the simplest approach and works well for relatively straightforward tasks where the subtasks are obvious from the goal description. Its weakness is that it does not force the model to reason about dependencies or parallelism; steps are listed in the order the model thinks of them, which may not be the most efficient execution order.

Least-to-most prompting decomposes by asking the model first what the easiest sub-problem to solve is, solving it, then asking what the next easiest problem is given the previous solution. This builds up a plan from the bottom, which is particularly useful when earlier steps reveal knowledge that changes how later steps should be formulated. The technique was introduced specifically to address compositional reasoning tasks where simpler components must be understood before more complex ones can be tackled.

Divide-and-conquer decomposition splits the problem into roughly equal halves, solves each independently, and combines the results. This pattern appears naturally in tasks like comparing two large documents (compare the first halves, compare the second halves, then synthesize) or processing long inputs that exceed the context window. It is especially useful when the goal has natural symmetry or when partial results can be combined without requiring a single global pass over all information.

Goal decomposition prompting explicitly asks the model to list the preconditions of the goal, what must be true before the goal can be achieved, then identify what actions establish those preconditions, and recurse. This maps closely to backward-chaining search in classical planning and is effective for tasks with clear prerequisite structures. A goal like "publish a blog post" has preconditions including "have a written draft" and "have an approved title", each of which has its own preconditions, making backward-chaining a natural decomposition strategy.

Plan-sketch refinement asks the model to produce a rough outline first, then refine each outline item into an executable step. This two-pass approach trades one additional LLM call for substantially more coherent plans, because the first pass establishes the global structure and the second pass fills in details that are consistent with that structure. The risk of local incoherence, where individual steps are fine but do not fit together as a whole, is significantly lower with this approach.

What Makes a Good Decomposition

Beyond the specific strategy used, a good decomposition shares several characteristics that distinguish high-quality plans from poor ones. Understanding these properties helps both in evaluating model-generated plans and in designing prompts that elicit better decompositions.

A good decomposition is complete: every step needed to achieve the goal is present, with no implicit assumptions that would require the executor to figure out something the plan did not specify. Incompleteness is a common failure mode, particularly for steps near the end of a plan where the model implicitly assumes that earlier steps have generated everything needed.

A good decomposition is non-redundant: no step repeats work done by an earlier step. Redundancy often appears when a plan is generated without careful attention to what information each step produces. If step two already retrieves the text of a source document, step five should not retrieve it again just to check a different fact.

A good decomposition has well-specified interfaces: each step's output is defined clearly enough that the next step can use it without guessing. "Search for information about X" has an unclear output interface. "Search for information about X and return the three most relevant passages" has a clear one. Well-specified interfaces are especially important at the boundaries between parallel branches that must be combined into a synthesis step.

A good decomposition is appropriately granular: steps are neither so large that they require the executor to make sub-decisions that should be in the plan, nor so small that the coordination overhead dominates the work. Calibrating granularity requires the planner to have a model of executor capability, which in practice means knowing what tools are available and what each tool does reliably.

Goal-Directed Planning Architectures

Several architectures have been proposed for LLM-based goal-directed planning. They differ in how much reasoning is done upfront versus during execution, in how the plan is represented, and in how failures are handled. Choosing the right architecture for a given task type is one of the most consequential design decisions in building a planning agent.

The fundamental tradeoff that distinguishes these architectures is between commitment and flexibility. An architecture that commits fully to a plan before execution is efficient when the plan is correct but fragile when it is not. An architecture that reasons dynamically during execution can adapt to surprises but pays a higher per-step reasoning cost. No single architecture dominates across all task types, which is why practitioners need to understand when each one is appropriate.

A second dimension of variation is how parallelism is handled. Some architectures are inherently sequential: each step must complete before the next can begin. Others make dependencies explicit, enabling steps that do not depend on each other to run concurrently. For multi-agent systems where multiple specialized agents can act simultaneously, this distinction has a major impact on total task completion time.

Plan-and-Execute

The simplest architecture separates planning from execution into two distinct phases. In the planning phase, the model receives the goal and generates a complete plan: a numbered list of steps to be taken. In the execution phase, each step is executed in order by an agent with access to tools. The executing agent does not revise the plan; it simply follows it.

This separation has appealing properties. The planner can reason about the full structure of the task without being distracted by implementation details. The executor can focus on completing individual steps without worrying about the global goal. The phases can even use different models: a large, expensive model for planning and a smaller, cheaper model for execution. This model-size split is a practical optimization that some production systems use to reduce cost without sacrificing plan quality, since planning requires broad knowledge and multi-step reasoning while execution often requires only the ability to correctly invoke a tool with specified parameters.

The weakness of plan-and-execute is brittleness. Plans generated upfront cannot anticipate every contingency. When step three fails because a tool returns an unexpected error, or when the output of step two reveals that the approach planned in steps four through seven is no longer appropriate, the executor has no mechanism to adapt. The plan either fails or the system falls back to replanning from scratch, which is expensive. Plan-and-execute is therefore best suited for tasks with high predictability: the steps are known before execution begins, tool behavior is reliable, and the environment is stable.

ReWOO (Reasoning Without Observation)

ReWOO, introduced by Xu et al. (2023), extends the plan-and-execute idea by making dependencies between steps explicit. Instead of just listing steps, the planner writes variable names that represent the outputs of earlier steps, which later steps can reference. For example:

  • Step 1: Search for the population of France. Result: #E1.
  • Step 2: Search for the GDP of France. Result: #E2.
  • Step 3: Calculate GDP per capita using #E1 and #E2. Result: #E3.

The planner knows from the start that step three depends on steps one and two. The executor runs steps one and two (possibly in parallel, since they are independent), collects their results, substitutes them for the variable references in step three, and executes step three with the concrete values. This makes dependencies explicit and enables safe parallelism: steps that share no variable dependencies can be run concurrently.

The name "Reasoning Without Observation" refers to the fact that the planner generates the entire plan before any observations are made. This is efficient when the plan structure is predictable, because the executor never needs to call the planner again during execution. The planner's reasoning is concentrated in a single upfront call, and the executor's job is simple substitution and dispatch. The cost is the same as plan-and-execute: the planner cannot adapt to surprising intermediate results because it committed to the plan structure before seeing any observations.

ReWOO works particularly well for structured research and question-answering tasks where the information retrieval pattern is predictable and intermediate results flow clearly from one step to the next. It works less well when the plan structure itself depends on what intermediate results look like, for example when a search might return zero results and the appropriate follow-up query depends on what was or was not found.

LLM+P

LLM+P (Liu et al., 2023) hybridizes language model planning with classical symbolic planning. The language model translates a natural language goal description into a formal PDDL (Planning Domain Definition Language) problem, a classical planner like FastDownward solves the formal problem and produces an optimal or near-optimal plan, and the language model translates the formal plan back into natural language actions.

PDDL is the standard language for classical AI planning. It describes a planning domain in terms of types (objects that exist in the world), predicates (properties and relationships between objects), and actions (operators that have preconditions and effects). Given a domain description and a problem specification (initial state and goal state), a PDDL planner performs state-space search to find a valid action sequence. FastDownward, one of the most widely used PDDL planners, combines multiple heuristic search strategies to find solutions efficiently.

This hybrid approach inherits the strengths of classical planning, including provable correctness, optimality guarantees, and deterministic behavior, while using the LLM to bridge the gap between natural language and formal representation. The LLM's job is translation, not planning. The planning is done by a verified solver that cannot hallucinate. LLM+P works well for well-structured domains such as logistics and scheduling, including navigation problems, where a clean formal model can be constructed. It struggles when the task involves ambiguity, open-ended information retrieval, or actions that cannot be enumerated in a formal domain description. Most real-world LLM agent tasks fall into this difficult category, which limits LLM+P's practical applicability despite its theoretical appeal.

Tree-of-Thought Planning

Tree-of-Thought (ToT) planning, introduced by Yao et al. (2023), treats planning as a search problem over a tree of reasoning states. Rather than committing to a single linear plan, the model explores multiple branching possibilities simultaneously. At each decision point, the model generates several candidate next steps, evaluates each candidate (either with self-evaluation prompts or an external verifier), and expands the most promising branches.

This is significantly more expensive than linear planning because each node in the tree may require multiple model calls: one to generate candidates, one per candidate to evaluate it, and further calls to expand chosen branches. But it is also more resilient: dead ends are abandoned before they consume execution resources, and the best path through the problem space is found even when the first plausible path turns out to be wrong.

Tree-of-Thought Planning

Tree-of-Thought planning frames planning as a search over a tree where nodes are reasoning states and edges are candidate actions. The model generates multiple candidate continuations from each node, evaluates their promise, and expands the best ones. This is similar to Monte Carlo Tree Search in game-playing AI but implemented entirely through language model prompting.

The search strategy in ToT can be breadth-first (explore all nodes at depth dd before advancing to depth d+1d+1), depth-first (follow the most promising branch as far as possible before backtracking), or best-first (always expand the globally highest-scoring node). Each has different tradeoffs: breadth-first is more thorough but requires maintaining many candidate paths simultaneously; depth-first is efficient when good paths exist and poor paths are quickly identifiable; best-first is generally the best choice when a reliable evaluation function is available.

The evaluation function determines which branches ToT explores. If the model evaluates candidate steps poorly, ToT can waste enormous compute exploring bad branches while missing good ones. Effective evaluation prompts ask the model to reason about whether a candidate step makes progress toward the goal, whether it is logically consistent with prior steps, and whether it opens or closes future options. External verifiers, such as code execution for planning tasks that involve writing code, can provide ground-truth feedback that is more reliable than self-evaluation.

ToT is most valuable for tasks with a large search space of possible approaches where early choices significantly constrain what is possible later. Mathematical problem solving, creative writing with specific constraints, and code generation for complex algorithms are natural fits. It is overkill for straightforward retrieval and summarization tasks where the first plausible plan is usually correct.

ReAct with Planning (PlanReAct)

A practical middle ground between pure plan-and-execute and fully dynamic step-by-step reasoning is to combine an initial planning step with the ReAct loop. The agent first generates a high-level plan in natural language, and then executes the plan using the ReAct thought-action-observation cycle. Each iteration of the ReAct loop is guided by the current plan step, but the agent retains the ability to revise the plan if observations suggest it is no longer valid.

This architecture is more flexible than plan-and-execute because observations can trigger plan revisions, and simpler than full tree-of-thought planning because the search tree is collapsed to a linear sequence except when explicit replanning is needed. The plan functions as a soft constraint on behavior rather than a hard specification. The agent uses it as a guide but can deviate when the situation requires it.

PlanReAct is the most commonly used architecture in practical deployments as of 2024. Most production agent systems, including various implementations of LangGraph, AutoGPT-style systems, and the OpenAI Assistants API, use some version of this pattern. The combination of upfront deliberation with ongoing adaptation captures most of the benefit of pure planning while retaining most of the flexibility of pure reaction. Its weakness is that plan revisions can be expensive, since each revision requires a full model call with substantial context, and in adversarial or rapidly changing environments the agent may spend more time replanning than executing.

Graph-Based Planning

An extension of explicit dependency management that is gaining traction in production systems is representing plans as directed acyclic graphs (DAGs) rather than lists. In a list plan, steps are ordered linearly and parallelism must be inferred from the absence of dependencies. In a graph plan, nodes are tasks and directed edges represent data dependencies: an edge from node A to node B means B requires A's output as input.

Graph-based plans make the parallel structure immediately visible and machine-readable without requiring any inference. An orchestration engine reading the graph can trivially identify which nodes have all their dependencies satisfied and dispatch them simultaneously. When a node fails, the graph makes it clear exactly which downstream nodes are blocked, enabling targeted recovery reasoning rather than full replanning.

The LangGraph library implements this pattern directly. Workflows are defined as graphs where nodes are agent steps (model calls or tool invocations) and edges encode control flow and data dependencies. This representation also supports cycles, which pure DAG plans cannot express, enabling loops with exit conditions that are essential for tasks requiring iterative refinement.

Plan Execution and Monitoring

Generating a plan is only half the problem. Executing it correctly requires managing the flow of control, monitoring progress, handling intermediate results, and detecting when something has gone wrong. These engineering concerns are as important as the planning algorithms themselves, and in practice they often determine whether a planning agent works reliably in production.

Step Execution and Context Management

As each step of the plan is executed, its result must be stored and made available to subsequent steps. The planning agent needs to maintain a context object that records:

  • The original goal
  • The full plan with step statuses (pending, in progress, completed, failed)
  • The outputs of each completed step
  • Any observations or errors encountered during execution

This context object is the agent's working memory. It prevents the agent from forgetting what it was doing and ensures that later steps have access to all information generated by earlier steps. In practice, this context is often passed as part of the system prompt or injected into the model's context window on each step. For very long plans where the accumulated context would exceed the context window, selective summarization of completed steps keeps the total size manageable while preserving the information most relevant to current execution.

A significant engineering challenge is context window management. As plans grow longer and step outputs accumulate, the total number of tokens in context can exceed the model's limit. Strategies for managing this include summarizing the outputs of completed steps rather than preserving them verbatim, keeping only the most relevant recent context, and using an external memory store (as discussed in the Agent Memory chapter) to offload history that does not need to be in the immediate context. A common heuristic is to keep the last kk steps in full and maintain compressed summaries of everything older, where kk is chosen to be large enough to handle the typical dependency depth of the plan.

The choice of what to include in context at each step is not trivial. Including too little risks the executor making decisions without relevant information. Including too much fills the context window with irrelevant content that dilutes the signal the model can act on. The optimal context for a given step is the goal, the current step description, the outputs of direct predecessor steps, and any recent observations that might affect the current step. Everything else can be deferred to external memory.

Progress Monitoring

A well-designed planning agent monitors its own progress. After each step completes, the agent checks whether the step's output is consistent with what was expected, whether the plan's preconditions for the next step have been satisfied, and whether any new information suggests the plan should be revised.

This monitoring can be implemented as a simple check after each execution: the agent is prompted to assess whether the step output is valid and whether the plan should proceed, be modified, or be abandoned. More sophisticated implementations use a separate critic model to evaluate outputs, or define explicit success criteria for each step that can be checked programmatically. Programmatic success criteria, such as verifying that a search step returned at least one result or that a code execution step produced no errors, are preferable when they can be defined because they do not require an additional model call and cannot be fooled by plausible-sounding but incorrect model self-assessment.

Progress monitoring also is a form of early stopping. If a plan's first few steps produce unexpected results that make the goal unachievable or nonsensical, catching this early saves the cost of executing the remaining steps. An agent researching a factual question that discovers in step two that its initial premise was wrong should stop and reflect before replanning rather than continuing to execute a plan built on an incorrect assumption.

Handling Tool Failures

Tool failures are one of the most common sources of plan breakdown. APIs become unavailable, search queries return empty results, code execution raises exceptions, rate limits trigger. The plan executor must distinguish between three fundamentally different types of failure, because the correct response differs for each:

  • Recoverable failures: A tool call fails due to a transient error (network timeout, rate limit). The right response is to retry with backoff. These failures are not about the plan being wrong; they are about infrastructure instability. A well-designed executor handles them without involving the planner at all.
  • Fixable failures: The tool call fails because the input was malformed or the query was poorly specified. The right response is to revise the tool call. The plan structure is correct, but the specific parameters need adjustment. This is a local repair: the planner needs to modify the step's tool input but can leave the rest of the plan intact.
  • Plan-invalidating failures: The tool call reveals that a fundamental assumption in the plan is wrong. A search returns evidence that the factual premise of the plan does not hold, or a required API endpoint does not exist. The right response is replanning because the plan's structure is incorrect; changing a single step's parameters cannot fix it.

Distinguishing these failure types requires the agent to reason about why a failure occurred, not just that it occurred. This is a place where the language model's reasoning capability is especially valuable: given an error message and the context in which it occurred, the model can usually identify which category the failure falls into and propose an appropriate response. A 429 Too Many Requests error is clearly recoverable. A 404 Not Found error on an API endpoint the plan required is plan-invalidating. A malformed query parameter that produces no results is fixable.

Retry policies for recoverable failures should implement exponential backoff to avoid overwhelming already-stressed infrastructure. A simple policy retries after 1 second, then 2 seconds, then 4 seconds, then declares the step permanently failed. The maximum retry count is a hyperparameter that trades resilience for latency: higher values handle longer transient outages but increase wait time in the common case.

Detecting Semantic Failures

Beyond hard tool failures, planning agents must also detect semantic failures: cases where a tool call technically succeeds but returns content that does not satisfy the step's requirements. A search that returns results about a different topic than intended is a semantic failure. A code generation step that produces syntactically valid code that does not implement the specified algorithm is a semantic failure.

Semantic failure detection requires the agent to verify that step outputs satisfy their intended purpose, not just that the tool call returned without an error. This is harder than detecting hard failures because it requires understanding what the step was supposed to produce, which requires maintaining a clear specification of each step's expected output in the plan. When steps are specified only as vague descriptions like "gather information about X", semantic failure detection is nearly impossible because there is no clear criterion for what would count as success. Well-specified steps with explicit output requirements make semantic failure detection tractable.

Plan Revision and Recovery

The most practically important aspect of planning is handling the gap between the planned sequence and what happens during execution. Plans fail, observations surprise, new information changes what is optimal. A resilient planning agent does not give up when this happens. It revises.

Here planning systems earn their value. A system that can generate a plan but cannot recover from failures offers little improvement over a system that executes blindly step by step. Recovery capability justifies the overhead of upfront planning because it turns a failure from a dead end into a detour. The structured representation of the plan makes recovery reasoning tractable: the agent knows exactly what has been accomplished, what was expected next, and what the ultimate goal is, which is precisely the information needed to chart a new course.

Replanning Triggers

Three conditions should trigger replanning:

Execution failure: A step fails and cannot be completed despite retries and reformulation attempts. The plan must be revised to find an alternative path to the goal that does not depend on the failed step. This might mean finding a different tool that provides the same information, reformulating the goal in a way that does not require the unavailable resource, or acknowledging that the original goal cannot be fully achieved and proposing a partial completion.

Observation surprise: A step completes but its output is qualitatively different from what the plan assumed. For example, the plan assumed a document would contain information about a topic, but the retrieved document does not. Steps downstream that depend on this information need to be revised or replaced. Observation surprises are more subtle than execution failures because the step appears to succeed, but the success is hollow: the output is not what the plan needed.

Goal clarification: During execution, the agent realizes that the original goal specification was ambiguous and its interpretation was wrong. The plan needs to be rebuilt around the corrected understanding of what success looks like. Goal clarification is particularly important in multi-turn agent interactions where the human user provides feedback during execution that reveals the agent's initial understanding was incomplete.

Replanning Strategies

When a trigger is detected, the agent has several options for how to replan, ranging from minimally invasive local repair to complete plan reconstruction. The right choice depends on how severe the trigger is and how much of the existing plan remains valid.

Local repair modifies only the steps immediately affected by the failure, leaving the rest of the plan intact. If step four fails because a specific database is unavailable, a local repair replaces step four with a call to an alternative database. This is efficient because the planner does not need to regenerate the entire plan. The context window cost is low, the latency is minimal, and the risk of introducing new errors into the unchanged parts of the plan is zero. Local repair is the right first response to most execution failures and observation surprises.

Suffix replanning abandons the remaining steps of the plan and generates a new plan for the remaining subgoal, starting from the current state. The planner is given the original goal, the steps already completed and their outputs, and asked to generate a plan for what remains. This is more expensive than local repair but handles cases where the failure fundamentally changes the approach needed for subsequent steps. If step four's failure reveals that the entire approach to the second half of the plan is wrong, local repair to step four is insufficient; the whole suffix needs to be rebuilt.

Full replanning restarts the planning process from scratch given what has been learned. This is the most expensive option but is appropriate when the plan's underlying assumptions have been invalidated entirely. Full replanning benefits from the information gathered during the failed execution: the planner now knows what does not work and can avoid those paths. It should explicitly include a summary of prior failed attempts in the planning prompt so the new plan is informed by what was learned.

In practice, most implementations try local repair first, escalate to suffix replanning if local repair fails twice, and use full replanning only as a last resort. This escalation policy minimizes the average replanning cost across all executions while ensuring that severe failures eventually trigger a full response.

Self-Correction with Reflection

Reflection-based replanning (as in the Reflexion framework by Shinn et al., 2023) adds a systematic self-critique step before replanning. After a plan failure, the agent is prompted to reflect on what went wrong: what assumption failed, what information was missing, what strategy was flawed. This reflection is stored in a "working memory" of past failures and consulted at the start of the next planning attempt, helping the agent avoid repeating the same mistake.

Reflexion

Reflexion is a framework for LLM agents that uses verbal reinforcement learning: after each failed attempt at a task, the agent generates a natural language critique of its own performance and stores this critique in an episodic memory buffer. On the next attempt, this buffer is included in the planning prompt, giving the agent explicit information about what went wrong and how to avoid it.

The key insight behind Reflexion is that the gradient signal in traditional machine learning, which pushes model weights toward better performance, can be replaced for LLM agents by a verbal gradient: a natural language description of what went wrong and what to do differently. This requires no weight updates and can be implemented entirely through prompting, making it immediately applicable to any LLM without fine-tuning.

To make this concrete, consider an agent that tries and fails to answer a multi-hop research question. The standard approach is to restart from scratch. Reflexion instead instructs the agent to generate a short critique like "I assumed the primary source contained the answer, but it did not. Next time, I should verify the source topic before relying on it." This critique is prepended to the next attempt's planning prompt. The model reads its own past failure analysis and uses it as a prior when generating the new plan, systematically avoiding the strategy that failed.

Across multiple attempts, the agent accumulates a growing library of failure modes specific to this task, and the planning quality improves trial by trial without any change to model weights. The Reflexion paper demonstrated that this approach could solve tasks that pure execution loops could not, by systematically eliminating strategies that had already proven to fail. The limitation is that the failure memory accumulates only within a single task session; it does not generalize across different tasks or persist beyond the current conversation.

The Role of Uncertainty in Replanning

A subtle aspect of recovery that is often overlooked in simple implementations is the role of uncertainty. Not all plan failures are equally surprising. A step that depended on a source known to be occasionally unavailable should trigger a different response than a step that failed for a completely unexpected reason.

Well-designed planning agents maintain explicit uncertainty estimates about which steps are likely to succeed. Steps that depend on unreliable tools, ambiguous inputs, or volatile external state should be flagged as high-risk during planning, and contingency branches should be pre-planned for them. This is analogous to how a careful human planner identifies the riskiest parts of their plan and thinks through fallbacks before starting execution. If the database query might return empty results, plan for it now rather than discovering it mid-execution with no fallback ready.

Uncertainty-aware planning is harder to implement but significantly more resilient in practice. The simplest implementation asks the model to annotate each step in the plan with a confidence score and a fallback action. The executor checks the confidence score before each step and pre-allocates resources for the fallback if the confidence is below a threshold. More sophisticated implementations model uncertainty as probability distributions over step outcomes and use expected value calculations to choose between plan variants.

Implementation

We will build a working planning agent in Python that demonstrates task decomposition, plan representation, sequential execution with monitoring, and replanning on failure. The implementation uses simulated tools that occasionally fail to demonstrate recovery logic without requiring external API calls.

The architecture has four layers: data structures that represent plans and steps, tool implementations that do the actual work, an execution engine that dispatches steps in dependency order, and a recovery layer that handles failures. Each layer can be inspected and modified independently, which mirrors the separation of concerns in production planning agents.

Setup and Data Structures

In[4]:
Code
import random

random.seed(42)

We define the data structures that represent a plan and its execution state. The design prioritizes clarity over performance: every field is explicitly named, and status transitions are tracked as an enum rather than as boolean flags. This makes the execution logic easier to reason about and debug.

In[5]:
Code
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional


class StepStatus(Enum):
    PENDING = "pending"
    IN_PROGRESS = "in_progress"
    COMPLETED = "completed"
    FAILED = "failed"
    SKIPPED = "skipped"


@dataclass
class PlanStep:
    step_id: int
    description: str
    tool: str
    tool_input: str
    depends_on: list = field(default_factory=list)
    status: StepStatus = StepStatus.PENDING
    output: Optional[str] = None
    error: Optional[str] = None
    attempts: int = 0


@dataclass
class Plan:
    goal: str
    steps: list = field(default_factory=list)
    context: dict = field(default_factory=dict)

    def get_step(self, step_id: int) -> Optional[PlanStep]:
        for step in self.steps:
            if step.step_id == step_id:
                return step
        return None

    def pending_steps(self) -> list:
        return [s for s in self.steps if s.status == StepStatus.PENDING]

    def completed_steps(self) -> list:
        return [s for s in self.steps if s.status == StepStatus.COMPLETED]

    def failed_steps(self) -> list:
        return [s for s in self.steps if s.status == StepStatus.FAILED]

    def is_complete(self) -> bool:
        return all(
            s.status in (StepStatus.COMPLETED, StepStatus.SKIPPED)
            for s in self.steps
        )

The PlanStep dataclass encodes everything the executor needs for a single step: the tool to call, the input to pass it, the steps it depends on, its current status, and its output once it completes. The depends_on list is the key field for dependency management. The Plan class wraps a list of steps and provides convenience methods for querying subsets of them by status. Note that is_complete treats SKIPPED as a terminal state alongside COMPLETED, which allows the planner to mark unreachable downstream steps as skipped after a plan-invalidating failure rather than leaving them permanently pending.

Tool Definitions

Each tool has a small probability of failing to demonstrate recovery logic. In a real system, these would be actual API calls or code execution environments.

In[6]:
Code
# Simulated tool implementations.
# fail_prob controls how often a tool raises a transient error.


def search_tool(query: str, fail_prob: float = 0.15) -> dict:
    if random.random() < fail_prob:
        raise RuntimeError(f"Search API timeout for query: '{query}'")
    results = {
        "climate change temperature": "Global mean temperature has risen 1.1 C above pre-industrial levels.",
        "renewable energy capacity": "Global renewable capacity reached 3,372 GW in 2023, up 295 GW from 2022.",
        "carbon emissions 2023": "Global CO2 emissions hit 36.8 billion tonnes in 2023, a record high.",
        "sea level rise rate": "Sea levels are rising at 3.7 mm per year on average since 1993.",
    }
    for key, value in results.items():
        if any(word in query.lower() for word in key.split()):
            return {"query": query, "result": value, "source": "simulated_db"}
    return {
        "query": query,
        "result": f"No specific data found for: {query}",
        "source": "simulated_db",
    }


def compute_tool(expression: str, fail_prob: float = 0.05) -> dict:
    if random.random() < fail_prob:
        raise RuntimeError(f"Compute service unavailable for: '{expression}'")
    # Parse simple arithmetic expressions manually to avoid security concerns
    import ast
    import operator

    ops = {
        ast.Add: operator.add,
        ast.Sub: operator.sub,
        ast.Mult: operator.mul,
        ast.Div: operator.truediv,
    }

    def ast_compute(node):
        if isinstance(node, ast.Constant):
            return node.n
        if isinstance(node, ast.BinOp) and type(node.op) in ops:
            return ops[type(node.op)](
                ast_compute(node.left), ast_compute(node.right)
            )
        raise ValueError(f"Unsupported expression: {ast.dump(node)}")

    try:
        tree = ast.parse(expression, mode="eval")
        result = ast_compute(tree.body)
        return {"expression": expression, "result": result}
    except Exception as exc:
        raise RuntimeError(f"Computation failed: {exc}") from exc


def summarize_tool(text: str, fail_prob: float = 0.05) -> dict:
    if random.random() < fail_prob:
        raise RuntimeError("Summarization model unavailable")
    words = text.split()
    summary = " ".join(words[: min(20, len(words))]) + (
        "..." if len(words) > 20 else ""
    )
    return {"original_length": len(words), "summary": summary}


TOOLS = {
    "search": search_tool,
    "compute": compute_tool,
    "summarize": summarize_tool,
}

The tools dictionary maps string names to callable functions. The execution engine uses string lookup at runtime rather than direct function references, which mirrors how real tool-calling agents work: the planner specifies tools by name in the plan, and the executor resolves names to implementations when each step runs.

Plan Generation

The planner takes a goal and generates a structured plan. In a real system, this function calls an LLM with a planning prompt. Here we hardcode two example plans to focus on demonstrating execution and recovery logic without requiring an API connection.

In[7]:
Code
def generate_plan(goal: str) -> Plan:
    """
    Generate a structured plan for the given goal.
    Returns a Plan with PlanStep objects encoding dependencies.
    """
    plan = Plan(goal=goal)

    if "climate" in goal.lower():
        plan.steps = [
            PlanStep(
                step_id=1,
                description="Search for current global temperature anomaly data",
                tool="search",
                tool_input="climate change temperature",
                depends_on=[],
            ),
            PlanStep(
                step_id=2,
                description="Search for renewable energy capacity data",
                tool="search",
                tool_input="renewable energy capacity",
                depends_on=[],
            ),
            PlanStep(
                step_id=3,
                description="Search for current CO2 emissions data",
                tool="search",
                tool_input="carbon emissions 2023",
                depends_on=[],
            ),
            PlanStep(
                step_id=4,
                description="Search for sea level rise data",
                tool="search",
                tool_input="sea level rise rate",
                depends_on=[],
            ),
            PlanStep(
                step_id=5,
                description="Summarize all gathered climate data into a brief report",
                tool="summarize",
                tool_input="AGGREGATE",  # Filled at runtime from steps 1-4
                depends_on=[1, 2, 3, 4],
            ),
        ]
    else:
        plan.steps = [
            PlanStep(
                step_id=1,
                description=f"Search for information about: {goal}",
                tool="search",
                tool_input=goal[:50],
                depends_on=[],
            ),
            PlanStep(
                step_id=2,
                description="Summarize the gathered information",
                tool="summarize",
                tool_input="AGGREGATE",
                depends_on=[1],
            ),
        ]
    return plan

Execution Engine

The executor runs each step in dependency order, handling retries for transient failures and storing outputs for downstream steps.

In[8]:
Code
import time

MAX_RETRIES = 2
RETRY_DELAY = 0.1  # seconds


def can_execute(step: PlanStep, plan: Plan) -> bool:
    """Return True if all dependencies for this step have completed."""
    for dep_id in step.depends_on:
        dep = plan.get_step(dep_id)
        if dep is None or dep.status != StepStatus.COMPLETED:
            return False
    return True


def build_tool_input(step: PlanStep, plan: Plan) -> str:
    """
    For steps with tool_input == 'AGGREGATE', concatenate the outputs
    of all dependency steps as the tool input.
    """
    if step.tool_input != "AGGREGATE":
        return step.tool_input
    parts = []
    for dep_id in step.depends_on:
        dep = plan.get_step(dep_id)
        if dep and dep.output:
            parts.append(dep.output)
    return " | ".join(parts)


def execute_step(step: PlanStep, plan: Plan) -> bool:
    """
    Execute a single plan step. Retries up to MAX_RETRIES times on failure.
    Returns True on success, False if all retries are exhausted.
    """
    tool_fn = TOOLS.get(step.tool)
    if tool_fn is None:
        step.status = StepStatus.FAILED
        step.error = f"Unknown tool: {step.tool}"
        return False

    tool_input = build_tool_input(step, plan)
    step.status = StepStatus.IN_PROGRESS

    for attempt in range(1, MAX_RETRIES + 2):
        step.attempts = attempt
        try:
            result = tool_fn(tool_input)
            output_val = result.get("result") or result.get("summary", "")
            step.output = str(output_val)
            step.status = StepStatus.COMPLETED
            return True
        except RuntimeError as exc:
            step.error = str(exc)
            if attempt <= MAX_RETRIES:
                time.sleep(RETRY_DELAY)
            else:
                step.status = StepStatus.FAILED
                return False

    step.status = StepStatus.FAILED
    return False


def execute_plan(plan: Plan, verbose: bool = False) -> list:
    """
    Execute all steps in dependency order.
    Returns an execution log (list of dicts, one per step executed).
    """
    execution_log = []
    max_rounds = len(plan.steps) * 2  # Safety limit

    for _ in range(max_rounds):
        ready = [s for s in plan.pending_steps() if can_execute(s, plan)]
        if not ready:
            break

        for step in ready:
            success = execute_step(step, plan)
            execution_log.append(
                {
                    "step_id": step.step_id,
                    "description": step.description,
                    "tool": step.tool,
                    "status": step.status.value,
                    "attempts": step.attempts,
                    "output_preview": step.output[:80] if step.output else None,
                    "error": step.error,
                }
            )

            if verbose:
                status_label = "COMPLETED" if success else "FAILED"
                attempts_label = f" ({step.attempts} attempt{'s' if step.attempts > 1 else ''})"
                detail = step.output[:60] if success else step.error
                print(
                    f"  Step {step.step_id} [{status_label}{attempts_label}]: {detail}"
                )

        if plan.is_complete():
            break

    return execution_log

The execute_plan function implements a simple round-robin scheduler: at each iteration, it finds all pending steps whose dependencies have all completed and dispatches them. In a real parallel executor, each ready step would be dispatched to a separate worker and the loop would wait on futures. In this sequential implementation, ready steps within a single round are executed one after another, but across rounds the dependency structure is respected.

The build_tool_input function handles the AGGREGATE sentinel value, which signals that the step's input should be constructed by concatenating all dependency outputs. This is how the summarization step in the climate plan receives the combined results of the four preceding search steps.

Replanning Logic

When steps fail, the agent first attempts local repair by modifying the tool input. If that fails, it escalates to suffix replanning.

In[9]:
Code
def local_repair(failed_step: PlanStep, plan: Plan) -> bool:
    """
    Attempt to repair a single failed step by modifying its tool input.
    In a real system this calls the LLM with failure context.
    Returns True if a repair was applied.
    """
    if failed_step.tool == "search":
        # Broaden the query as a fallback
        failed_step.tool_input = failed_step.tool_input + " overview"
        failed_step.status = StepStatus.PENDING
        failed_step.error = None
        failed_step.attempts = 0
        return True
    return False


def suffix_replan(plan: Plan) -> Plan:
    """
    Suffix replanning: given completed steps and their outputs, generate
    a continuation plan for the remaining subgoal.
    Returns a new Plan that extends the completed steps.
    """
    completed_outputs = [
        f"Step {s.step_id}: {s.output}"
        for s in plan.completed_steps()
        if s.output
    ]

    new_plan = Plan(goal=plan.goal)
    new_plan.steps = [s for s in plan.steps if s.status == StepStatus.COMPLETED]

    max_existing_id = max((s.step_id for s in new_plan.steps), default=0)
    aggregated_input = (
        " | ".join(completed_outputs) if completed_outputs else plan.goal
    )

    new_plan.steps.append(
        PlanStep(
            step_id=max_existing_id + 1,
            description="Synthesize available data into partial report",
            tool="summarize",
            tool_input=aggregated_input,
            depends_on=[s.step_id for s in new_plan.steps],
        )
    )
    new_plan.context["replanned"] = True
    return new_plan


def execute_with_recovery(goal: str) -> dict:
    """
    Full pipeline: generate plan, execute, recover from failures.
    Returns a summary dict with counts and outputs.
    """
    plan = generate_plan(goal)
    all_logs = execute_plan(plan, verbose=False)

    # Local repair pass
    failed = plan.failed_steps()
    if failed:
        for step in failed:
            local_repair(step, plan)
        logs2 = execute_plan(plan, verbose=False)
        all_logs.extend(logs2)

    # Suffix replanning if still failing
    still_failed = plan.failed_steps()
    if still_failed:
        plan = suffix_replan(plan)
        logs3 = execute_plan(plan, verbose=False)
        all_logs.extend(logs3)

    return {
        "goal": goal,
        "plan_steps": len(plan.steps),
        "completed": len(plan.completed_steps()),
        "failed": len(plan.failed_steps()),
        "replanned": plan.context.get("replanned", False),
        "final_outputs": {
            s.step_id: s.output for s in plan.completed_steps() if s.output
        },
        "execution_log": all_logs,
    }

Running the Planner

Let's execute the planner on a climate research task and inspect the results.

In[10]:
Code
result = execute_with_recovery(
    goal="Research the current state of climate change and prepare a summary report"
)
Out[11]:
Console
Goal: Research the current state of climate change and prepare a s...
Plan steps:       5
Completed:        5
Failed:           0
Replanning used:  False

Final outputs (5 steps produced output):
  Step 1: Global mean temperature has risen 1.1 C above pre-industrial levels.
  Step 2: Global renewable capacity reached 3,372 GW in 2023, up 295 GW from 202...
  Step 3: Global CO2 emissions hit 36.8 billion tonnes in 2023, a record high.
  Step 4: Sea levels are rising at 3.7 mm per year on average since 1993.
  Step 5: Global mean temperature has risen 1.1 C above pre-industrial levels. |...

The output shows the plan's final state: how many steps completed, whether replanning was triggered, and what each step produced. The dependency structure ensures step 5 (summarization) only runs after all four search steps have produced their outputs, even though the four search steps run in sequence in this single-threaded implementation.

Tracking Planning Metrics

To understand aggregate planner performance, we run the pipeline across many trials and collect statistics.

In[12]:
Code
def run_planning_experiment(n_trials: int = 30) -> dict:
    """Run the planning pipeline n_trials times and return aggregate statistics."""
    first_pass_rates = []
    total_attempts_per_run = []
    failures_per_run = []

    for _ in range(n_trials):
        plan = generate_plan(
            "Research the current state of climate change and prepare a summary report"
        )
        execute_plan(plan, verbose=False)

        total_steps = len(plan.steps)
        first_pass = sum(
            1
            for s in plan.steps
            if s.status == StepStatus.COMPLETED and s.attempts == 1
        )
        total_attempts = sum(s.attempts for s in plan.steps)
        failed_count = sum(
            1 for s in plan.steps if s.status == StepStatus.FAILED
        )

        first_pass_rates.append(
            first_pass / total_steps if total_steps else 0.0
        )
        total_attempts_per_run.append(total_attempts)
        failures_per_run.append(failed_count)

    n = n_trials
    return {
        "n_trials": n,
        "avg_first_pass_rate": sum(first_pass_rates) / n,
        "avg_attempts_per_run": sum(total_attempts_per_run) / n,
        "avg_failures_per_run": sum(failures_per_run) / n,
        "raw_first_pass_rates": first_pass_rates,
        "raw_total_attempts": total_attempts_per_run,
    }


experiment_results = run_planning_experiment(n_trials=30)
Out[13]:
Console
Planning Experiment Results (30 trials, 5-step climate plan)
  Avg first-pass completion rate:  86.0%
  Avg tool call attempts per run:  5.7
  Avg failures per run:            0.03

Most runs complete with every step succeeding on the first attempt. Occasionally a search step times out and requires a retry, adding one or two extra attempts. The summarization step (step 5) very rarely fails because it only runs after the four search steps have completed and its input is always a non-empty concatenation of their outputs. This illustrates an important design principle: dependent steps tend to have higher success rates because their inputs are pre-validated by the successful completion of their predecessor steps. A well-structured dependency graph naturally concentrates failure risk in the early, independent steps and shields later synthesis steps from missing-input failures.

Key Parameters

The key parameters controlling planning agent behavior are:

  • MAX_RETRIES: Number of times to retry a failed tool call before marking the step as failed. Higher values improve resilience to transient failures but increase latency.
  • depends_on (step parameter): The list of step IDs that must complete before this step executes. Correctly specifying dependencies is the most important factor in plan correctness.
  • fail_prob (tool parameter): The probability that any individual tool call raises a transient error. In production this is determined by infrastructure reliability.
  • AGGREGATE sentinel: The special tool_input value that signals the execution engine to concatenate the outputs of all dependency steps before passing them to the tool.

Visualizations

Out[14]:
Visualization
Stacked bar chart of step outcomes across 30 planning trials showing mostly first-pass completions.
Step outcomes across 30 simulated planning trials. Each bar shows the number of steps that completed on the first attempt (green), required retries (orange), or failed permanently (red). The retry mechanism resolves most transient tool failures, leaving final failure rates close to zero across most trials.
Out[15]:
Visualization
CDF curve of total tool call attempts per trial showing median near 5 to 6 attempts.
Cumulative distribution of total tool call attempts per planning trial. Most trials require 5 to 6 attempts for the 5-step climate plan (one per step), with the right tail reflecting trials where one or more steps needed retries before succeeding.
Sorted bar chart of first-pass completion rates across 30 trials showing mostly high rates.
First-pass completion rates across 30 planning trials sorted from lowest to highest. The distribution shows that most trials achieve 80 to 100 percent first-pass completion, with a minority experiencing multiple concurrent failures that illustrate the tail risk in plan execution.

The step-outcomes chart reveals that the majority of plan executions complete cleanly in a single pass, with occasional retries accounting for a small fraction of additional calls. The CDF of attempts shows the typical plan requires between 5 and 7 total tool invocations for a 5-step plan, with the extra calls reflecting retry overhead.

Planning Approaches Compared

The different planning architectures represent a spectrum of tradeoffs between planning cost, execution flexibility, and recovery capability. No single architecture is best for all tasks. Understanding the tradeoffs allows you to choose the right tool for a given deployment context.

Comparison of LLM planning architectures by planning cost, recovery ability, parallelism support, and ideal use case.
ArchitecturePlanning CostRecovery AbilityParallelismBest For
Plan-and-ExecuteLowNoneManualSimple, predictable tasks
ReWOOLowLimitedExplicitTasks with clear variable dependencies
PlanReActMediumModerateNoneGeneral-purpose agents
Tree-of-ThoughtHighHighBranchingTasks with many valid approaches
LLM+PMediumHighDependsFormal, well-structured domains

"Planning cost" refers to the number of model calls required before execution begins. Plan-and-execute requires one planning call. Tree-of-Thought may require dozens of calls just to evaluate candidate branches at the first decision node. In latency-sensitive or cost-sensitive deployments, this cost matters significantly. A planning step that costs several cents in API calls and multiple seconds of latency is acceptable for a long research task but not for a real-time customer service interaction.

"Recovery ability" captures how well the architecture handles unexpected execution outcomes. Plan-and-execute has no native recovery; it either completes or fails. PlanReAct can revise the plan on any step based on observations. Tree-of-Thought inherently explores alternatives, so a failed branch does not terminate the search.

For most production deployments, PlanReAct offers the best balance. The initial plan provides structure and enables progress monitoring. The ReAct loop within each plan step provides adaptability. And the ability to trigger replanning when observations deviate significantly from expectations handles severe failures without requiring the overhead of full tree search.

Limitations and Impact

Planning is one of the most actively researched areas in LLM agents, and current systems have significant limitations alongside demonstrated capabilities. Understanding both is essential for setting appropriate expectations and designing systems that work reliably in practice.

Current Limitations

The most fundamental limitation is hallucinated planning. When a language model generates a plan, it draws on statistical associations in its training data. It does not simulate the execution of the plan or verify that each step's preconditions will be satisfied when that step is reached. Plans can appear coherent while containing logical errors: step four assumes that step two produces a certain output, but the actual execution of step two never produces that output because the assumption was wrong. Humans reviewing plans detect these errors through domain knowledge; the model often does not. The result is plans that look correct at first glance but fail at specific transition points between steps.

Long-horizon degradation is a practical challenge that limits the applicability of current planning systems to medium-length tasks. Plans longer than eight to ten steps degrade rapidly in quality. The model struggles to maintain a coherent representation of all dependencies, intermediate states, and global constraints over many steps. For truly long-horizon tasks, such as planning a multi-week software project or coordinating a multi-agent pipeline with dozens of steps, current models must be scaffolded with external state tracking and regular re-grounding prompts to avoid drift. Re-grounding means periodically summarizing the current state of the plan and re-stating the original goal to prevent the model from losing context of what it was trying to achieve.

Over-planning and under-planning are complementary failure modes that are hard to address with simple prompt engineering. Some models generate overly detailed plans with too many small steps, creating coordination overhead without benefit. Others generate under-specified plans where steps are so vague that the executor cannot act on them. Calibrating plan granularity to the task is itself a skill that current models handle inconsistently, and the right granularity varies substantially across domains and tool sets.

Replanning cost is non-trivial and can accumulate rapidly in adversarial or unpredictable environments. Each replanning call requires sending substantial context to the model: the original goal, all completed steps and their outputs, the failure description, and possibly the reflection summary from prior attempts. This can easily consume tens of thousands of tokens per replanning event, which is expensive in both latency and API cost. Systems that replan frequently are significantly more expensive than those that plan once and execute cleanly. This creates a tension between resilience (many replanning opportunities) and cost (each replanning event is expensive).

Evaluation is difficult. Measuring whether a planning agent improved task performance requires comparing against well-designed baselines on tasks where the "right" plan structure is known. Many published benchmarks evaluate only final task success rate, which does not distinguish between good planning and lucky reactive execution. A system that generates random plans and occasionally succeeds looks competitive with a careful planner if success rate is the only metric. Richer evaluation metrics that measure plan efficiency, recovery rate, and step accuracy are still being developed.

Practical Impact

Despite these limitations, planning-capable agents have demonstrated substantial improvements over step-by-step reasoning on multi-step tasks. On benchmarks like HotpotQA (multi-hop question answering), FEVER (fact verification), and WebArena (web navigation tasks), plan-based agents consistently outperform reactive agents on tasks requiring more than three steps. The improvement grows with task length: the longer the task, the larger the gap between agents that plan and agents that react.

The practical impact is most visible in software engineering agents. Systems like Devin, SWE-agent, and Claude Code use planning to decompose code changes into file modifications, test runs, and error fixes. These agents can complete software engineering tasks that require understanding multiple files, identifying the root cause of a bug, implementing a fix, and verifying it, all as a coherent planned sequence rather than a series of isolated actions. The planning capability is what allows these systems to maintain coherence across the many steps required to resolve a non-trivial bug in a real codebase.

Planning also enables a qualitatively new kind of delegation. When an agent can reason about multi-step goals, users can specify objectives at a higher level of abstraction, closer to "write a report on X" than "first search for X, then summarize the first result". This changes the user-agent interaction from supervised step-by-step guidance to goal delegation, with the agent handling the operational details autonomously. The cognitive load on the user decreases, and the tasks that can be delegated become substantially more complex.

The research trajectory for planning is toward more reliable upfront plan generation (better planning prompts, more capable models), more efficient recovery (local repair without full replanning), and longer-horizon execution (better state tracking and context management for plans with many steps). Each of these improvements makes planning agents applicable to more demanding tasks. The field is advancing rapidly, and systems that struggle with ten-step plans today will likely handle twenty-step plans within a few years.

The connection to upcoming topics in this book is direct. Multi-agent systems build on planning by distributing subtasks across specialized agents: a planner decomposes the goal and assigns subtasks, specialist agents execute the assigned subtasks, and a synthesizer aggregates results. Planning is what makes this orchestration coherent rather than chaotic. Without a plan, a collection of specialist agents has no shared understanding of what the overall goal is or how their individual contributions fit together. With a plan, each agent receives a clearly specified subtask with well-defined inputs and outputs, and the orchestrator can monitor progress and trigger recovery at the right level of granularity.

Summary

Planning gives language model agents the ability to reason about multi-step goals before committing to action. The core ideas are:

  • Task decomposition breaks goals into subtasks, identifying sequential dependencies and opportunities for parallelism. Good decompositions are complete, non-redundant, and appropriately granular.
  • Planning architectures range from simple plan-and-execute to search-based tree-of-thought, trading planning cost for execution flexibility and recovery capability. PlanReAct offers the best balance for most production deployments.
  • Plan execution requires managing context across steps, handling tool failures with retry logic, and monitoring progress against plan expectations. Efficient recovery depends on distinguishing recoverable failures from fixable and plan-invalidating ones.
  • Replanning triggers on execution failures, observation surprises, and goal clarification, with local repair, suffix replanning, and full replanning available as escalating responses. Start with local repair and escalate only when necessary.
  • Reflection-based recovery as in Reflexion stores verbal critiques of past failures to avoid repeating the same mistakes across planning attempts, implementing a form of verbal reinforcement learning without weight updates.

Planning is not a solved problem. Current systems hallucinate plans, degrade on long horizons, and pay high costs for recovery. But even imperfect planning significantly improves agent performance on complex multi-step tasks, and the field is advancing rapidly. The planning architectures covered here form the foundation for everything from software engineering agents to research assistants to autonomous workflow systems. Understanding them deeply is essential for anyone building or evaluating goal-directed AI systems.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about planning in LLM agents.

Planning in LLM Agents

Question 1 of 80 of 8 completed
What is the key difference between a ReAct agent and a planning agent?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026planningtask, author = {Michael Brenndoerfer}, title = {Planning: Task Decomposition and Goal-Directed LLM Agents}, year = {2026}, url = {https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Planning: Task Decomposition and Goal-Directed LLM Agents. Retrieved from https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents
MLAAcademic
Michael Brenndoerfer. "Planning: Task Decomposition and Goal-Directed LLM Agents." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents>.
CHICAGOAcademic
Michael Brenndoerfer. "Planning: Task Decomposition and Goal-Directed LLM Agents." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Planning: Task Decomposition and Goal-Directed LLM Agents'. Available at: https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Planning: Task Decomposition and Goal-Directed LLM Agents. https://mbrenndoerfer.com/writing/planning-task-decomposition-goal-directed-llm-agents

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.