Agent Architectures: Control Loops, State & Planning

Michael BrenndoerferFebruary 4, 202653 min read

Part of Language AI Handbook

Covers LLM agent architectures including control loops, state management strategies, planning mechanisms, and termination conditions for autonomous AI systems.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Agent Architectures: Loop, State, Planning, Termination

In our exploration of function calling, we saw how Large Language Models can invoke external tools to overcome their inherent limitations: accessing real-time data, performing calculations, and interacting with external systems. But function calling alone does not make an agent. It gives the capability to act, but not the architecture for sustained, goal-directed behavior.

An agent architecture is the structural framework that orchestrates how an LLM perceives its environment, maintains internal state, plans multi-step strategies, executes actions, and determines when its work is complete. While a simple function-calling system might answer a single query with a single tool call, an agent architecture lets complex workflows: researching a topic across multiple sources, iteratively debugging code, or managing a long-running task with conditional branches and error recovery.

The move from static prompt-response interactions to dynamic agent architectures is a basic change in how we build AI systems. Instead of treating the LLM as a stateless function that maps input to output, we treat it as a cognitive engine embedded within a control loop that can persist state across time, reflect on intermediate results, and adapt its strategy based on feedback. This transformation mirrors the evolution from simple reflex agents in classical artificial intelligence to advanced deliberative agents capable of maintaining beliefs and desires as well as intentions over extended periods.

This chapter examines the four core components of agent architectures: the agent loop that drives execution, state management strategies that maintain context, planning mechanisms that structure multi-step reasoning, and termination conditions that bound agent behavior. By understanding how these pieces fit together, you will be equipped to design agents that are reliable and efficient, with autonomy bounded appropriately.

The Agent Loop

At the heart of every agent architecture lies a control loop that governs the interaction between the language model and its environment. This pattern, often called the Observation-Thought-Action loop, gives the basic rhythm of agent execution. The concept draws inspiration from cognitive psychology and control theory, where intelligent behavior emerges not from static input-output mappings but from continuous interaction with an environment through feedback cycles.

The loop is the basic unit of agent cognition. Without it, an LLM is a advanced text completion engine that responds to a fixed input and halts. With it, the LLM becomes an actor situated in a dynamic world, capable of pursuing a goal across time by repeatedly sensing and reasoning before acting. The loop is what distinguishes reactive tools from purposeful agents.

The Basic Cycle

The agent loop operates through a continuous cycle of three basic operations.

Observation is the first stage of every iteration. The agent receives input from its environment: a user query, the result of a previous tool execution, an error message from a failed action, or a system event like a timeout. The observation grounds the agent in the current state of the world. This gives the raw material upon which decisions are made. Without accurate observation, the agent operates on false premises, much like a navigator trying to plot a course with an outdated map. Every piece of information the agent receives in this phase, no matter how small, shapes the subsequent reasoning.

Reasoning is where the cognitive work happens. The LLM processes the observation in the context of its goal and available tools, generating an internal representation of what needs to be done next. This might mean deciding which tool to call and with what arguments, recognizing that the goal is already achieved, or concluding that the current strategy is not working and should be replaced. The reasoning step turns raw observations into structured intentions. Critically, the quality of this step depends on everything that has come before it: the richness of the conversation history, the clarity of the current plan, and the relevance of information in working memory. A well-designed architecture ensures that the LLM has exactly what it needs for this reasoning step, no more and no less.

Action is the agent's attempt to change the state of the world based on its reasoning. Actions can take many forms: calling a tool and awaiting its result, generating a response for the user, updating internal state, or deciding to terminate. The action bridges intention and effect. It is the moment when the agent's internal deliberation becomes external consequence. Every action produces an outcome that becomes the next observation, closing the loop and beginning the next cycle.

This three-phase cycle repeats until a termination condition is met, which we discuss in detail later. Each iteration consumes the previous action's output as its observation input, creating a chain of dependent operations where the context evolves based on the history of interactions. This chaining creates temporal coherence, letting the agent to build upon previous work rather than starting fresh with each exchange. The agent accumulates understanding across iterations the way a researcher accumulates evidence across a literature review: each source adds a piece to a growing picture.

Loop Variants

Not all agent loops have the same structure. The complexity of the loop should match the demands of the task and the dynamics of the environment.

Single-step loops execute one thought-action pair per user query. The agent receives input, generates a plan, executes all necessary actions in a predetermined sequence, and returns. This works for simple tasks where the environment is predictable and actions are largely independent of each other. For example, an agent asked to calculate the total cost of items in a shopping cart might use a single-step loop: it reads the items, calls the calculator tool once, and returns the result. However, single-step loops struggle with dynamic environments where intermediate results should influence subsequent actions. If the agent needs to check inventory before calculating shipping, and shipping options depend on warehouse locations determined by inventory checks, a single-step approach cannot adapt its course based on what it discovers at each stage.

Multi-step loops allow the agent to iterate many times, using each action's result to inform the next decision. This makes reactive behavior possible: if a web search returns insufficient results, the agent can reformulate its query; if code execution fails, the agent can diagnose the error and retry with corrections. Multi-step loops create the possibility for true interaction with external systems, where the agent engages in a dialogue with the world, learning and adjusting as it goes. This resembles how a human researcher pursues a line of inquiry, following references, revising hypotheses, and adjusting the search strategy based on what each source reveals. The cost is complexity: multi-step loops require reliable state management and carefully designed termination conditions to prevent infinite cycles.

Hierarchical loops nest agent execution within higher-level planning. A top-level planner agent decomposes a complex goal into subtasks, then delegates each subtask to specialized worker agents, each running their own internal loops. The results from the workers feed back to the planner, which coordinates their efforts and resolves dependencies between outputs. This architecture mirrors organizational structures in human enterprises, where managers coordinate specialists who possess deep expertise in specific domains. A software development agent might have separate sub-agents for requirements analysis and code generation as well as testing and documentation, with a meta-agent sequencing their work and managing the handoffs between stages. Hierarchical loops introduce significant coordination overhead but enable the agent to tackle problems that are too large or complex for any single agent to handle.

The choice of loop structure depends on task complexity, the dynamics of the environment, and the cost constraints of the deployment. Simple information retrieval might use single-step loops, while autonomous research assistants require multi-step or hierarchical architectures to handle the uncertainty and branching paths inherent in exploratory tasks.

Why the Loop Matters

The loop embodies a philosophical shift in how we conceptualize language model capabilities. A single-shot LLM interaction is fundamentally limited by the quality of the input prompt: if the user does not frame the question correctly, or if the task requires information not present at query time, the model cannot compensate. The agent loop removes this limitation by letting the model to actively seek missing information, correct its own errors, and adapt its strategy to what it learns along the way.

This shift also changes how we think about failure. In a single-shot system, a poor answer is a dead end. In an agent loop, a poor action is a learning opportunity: the next observation reveals that the action failed, and the agent can reason about why and try something different. This error-recovery capacity is one of the most practically important properties of agent architectures, letting them to succeed at tasks where fixed sequences of actions would inevitably fail due to environmental variability.

State Management

State is the memory of an agent system: the accumulated information that persists across loop iterations and lets coherent, context-aware behavior. Without effective state management, each loop iteration would be an isolated, stateless exchange, incapable of maintaining continuity or learning from previous actions. The challenge is particularly acute in LLM-based agents because, unlike traditional software systems that maintain arbitrary data structures in memory, agents must compress their state into the limited context window of the language model or manage it externally through retrieval systems.

The problem of state management becomes visible when you trace what an agent needs to know at each iteration. It needs the original goal, the actions it has already taken, the results of those actions, facts it has discovered along the way, the current plan, and any error conditions encountered. All of this information must be available during the reasoning step, or the agent loses coherence. Managing these different streams of information and keeping them organized and prioritized within the bounds of what the model can process is the central challenge of state management.

Types of Agent State

Agent architectures typically manage several categories of state, each serving distinct cognitive functions.

Conversation history maintains the chronological sequence of observations and thoughts followed by actions. This gives the LLM with context about what has already been attempted and what has been discovered. As we discussed in Part XXVIII on GPT Architecture, transformers process context through attention mechanisms, meaning the entire conversation history influences each new token generation. However, raw conversation history can become lengthy and redundant, containing facts alongside the procedural noise of how those facts were obtained: failed attempts, dead ends, and repeated queries that eventually yielded results. The agent must balance the need for complete context against the attention costs and context window limits of processing lengthy histories.

Working memory holds structured information extracted from observations: facts retrieved from knowledge bases, intermediate calculation results, parsed user intents, and active hypotheses. Unlike raw conversation history, working memory is often summarized or structured to reduce token consumption and highlight salient information. Think of working memory as the agent's scratchpad, where it keeps the current hypotheses, active goals, and necessary facts that demand immediate attention. This is analogous to human working memory, which has limited capacity but fast access, holding the information currently needed for the task at hand. A research agent might maintain working memory as a dictionary mapping topic headings to bullet-point summaries of what has been learned about each.

Episodic memory stores past experiences for long-term retrieval, persisting beyond the boundaries of a single conversation. While conversation history is linear and ephemeral (cleared between sessions), episodic memory allows agents to recall relevant previous interactions when encountering similar situations. This might include remembering that a particular database query pattern failed in a past session, or that a specific user prefers concise answers without technical jargon. We will explore memory architectures in greater detail in the upcoming chapter on Agent Memory, but episodic memory is worth understanding here as the mechanism that allows agents to learn from experience across conversations rather than starting fresh each time.

Tool state tracks the status of external systems: database connections, API session tokens, file handles, or the current position in paginated search results. This state ensures that tools maintain consistency across multiple invocations within a session. For example, if an agent is iterating through a large dataset using cursor-based pagination, the cursor position is needed state that must persist between tool calls. If this state is lost, the agent would restart from the beginning of the dataset on every call, potentially missing records or processing duplicates. Similarly, maintaining authentication tokens across API calls prevents the overhead of re-authenticating with each iteration, which could be both slow and potentially trigger rate limits.

State Persistence Strategies

The finite context window of LLMs creates a basic tension: agents need complete state to make informed decisions, but they can only attend to limited amounts of text. Several strategies address this constraint, each with different tradeoffs between completeness and efficiency.

Truncation with summarization discards older conversation turns while preserving a condensed summary. The agent might maintain the full context of recent interactions but compress distant history into key facts and decisions. For example, after ten iterations of web searching, the agent might summarize the first five iterations as a brief paragraph capturing the main sources found and their key claims, while keeping the last five iterations in full detail. This approach risks losing nuance (a seemingly minor detail from iteration two might become relevant at iteration eleven) but preserves the most immediately relevant information within a manageable token budget.

Hierarchical state organizes information at different granularities. Immediate context (the last few turns) is preserved verbatim; short-term summaries capture the current task phase; long-term memory stores extracted knowledge in structured form. This mirrors human cognitive architectures where working memory and long-term memory serve different functions, with information flowing between them through attention and consolidation. In practice, this might mean the agent keeps the last three tool calls in full detail, maintains a paragraph summary of the current subtask progress, and stores key facts in a structured knowledge graph. Each tier serves a different temporal horizon of reasoning: the immediate context supports fine-grained decision making, while the long-term store supports goal-level planning.

External state stores move state outside the LLM's context window into databases, vector stores, or structured files. The agent then uses retrieval mechanisms to pull relevant state into context on demand, building on the dense retrieval concepts from our discussion of retrieval-augmented generation. This approach allows agents to maintain effectively unlimited state while keeping context windows manageable. A legal research agent might store all discovered case law in a vector database, querying it for relevant precedents only when drafting specific arguments, rather than keeping every case in the active context. The tradeoff is retrieval quality: if the retrieval mechanism fails to retrieve a relevant piece of state, the agent proceeds without it, potentially making decisions inconsistent with prior work.

Out[4]:
Visualization
Scatter plot comparing four state management strategies on computational efficiency versus context retention axes, with bubble size representing implementation complexity.
Comparison of four state management strategies by computational efficiency and context retention. Full History reaches the highest retention (95%) but the lowest efficiency (20%), while External Store balances both dimensions at 85% retention and 80% efficiency. Bubble size indicates relative implementation complexity, with external stores requiring the most advanced infrastructure but delivering the most favorable tradeoff profile.

State Consistency

Maintaining consistent state across operations presents significant challenges, particularly when agents execute long-running tools or parallel subtasks where state updates may arrive out of order or fail entirely. Reliable architectures implement several strategies to ensure coherence.

Atomic updates ensure that state changes occur as complete transactions, preventing partial updates that leave the agent in inconsistent configurations. If an agent is updating both a task list and a knowledge base simultaneously, both updates should succeed or both should fail. Allowing one to succeed while the other fails creates an inconsistency where completed tasks are not reflected in the knowledge store, or conversely, where knowledge is updated for tasks that were never marked complete.

Conflict resolution gives rules for reconciling conflicting changes when concurrent processes modify shared state. If two sub-agents simultaneously update the same working memory key with different values, the architecture needs a clear policy for determining which update takes precedence. Common strategies include timestamp-based ordering (most recent write wins), priority-based resolution (the higher-priority agent wins), and merge operations (both values are combined, for example by appending to a list rather than overwriting).

Checkpointing creates periodic snapshots of agent state that let recovery from failures without losing all progress. If an agent crashes midway through a complex task, checkpointing allows it to resume from the most recent checkpoint rather than starting over from the beginning. This is especially important for long-running agents where the cost of restarting from scratch, in both time and API tokens, is prohibitive.

State management directly affects an agent's ability to reason coherently. An agent operating on inconsistent state may generate contradictory plans, repeat actions it has already taken, or fail to recognize when a goal has been achieved. Investing in reliable state management is investing in the agent's cognitive reliability.

Planning Strategies

Planning turns high-level goals into executable sequences of actions. While simple agents might react opportunistically to each observation, advanced architectures employ explicit planning mechanisms to structure their approach to complex tasks. Planning serves several necessary functions: it reduces cognitive load by breaking complex problems into manageable pieces, it lets the agent to anticipate future needs and lay groundwork for later steps, and it gives a framework for monitoring progress and detecting when the current approach is not working.

The difference between a planning agent and a reactive agent is the difference between a surgeon who reviews the patient's file, rehearses the procedure, and assembles the necessary instruments before making the first incision, versus one who improvises at each step based only on what is immediately in front of them. For routine, well-understood tasks, the reactive approach may be adequate. For novel, complex tasks with many interdependencies, explicit planning is needed for achieving a coherent outcome.

Planning Taxonomy

Planning strategies fall along a spectrum from reactive to deliberative, with different approaches suited to different problem structures.

Single-step planning generates one action at a time based on the current state. The LLM examines the current observation and immediately decides the next tool to call, with no explicit plan persisting between iterations. This approach is computationally efficient and naturally responsive to environmental changes: there is no outdated plan to discard when circumstances shift, because the agent reacts fresh each time. However, single-step planning may lack global coherence for multi-step tasks. It resembles the strategy of a chess player who considers only the current board position without thinking ahead several moves. Early decisions can close off better paths that would have been available with foresight, and without a plan to refer to, the agent may lose track of the overall goal when absorbed in solving immediate subproblems.

Chain-of-thought planning explicitly generates a reasoning sequence before selecting actions. As we explored in our discussion of chain-of-thought emergence, intermediate reasoning steps improve performance on complex problems by decomposing them into manageable sub-tasks. Rather than immediately jumping to a tool call, the agent first articulates its understanding of the problem and its intended approach. This externalized reasoning serves multiple purposes: it forces the agent to be explicit about its assumptions, it creates an audit trail for debugging when things go wrong, and it often produces better results by letting the model to work through the logic step by step rather than jumping straight to a conclusion. The cost is additional tokens per iteration, which can add up over many steps.

Hierarchical task networks decompose goals into sub-goals, which further decompose into primitive actions. A research task might break down into: identify key sources, extract relevant information from each source, and synthesize findings into a coherent narrative. Each sub-goal generates its own planning context, reducing cognitive load and letting modular reasoning. This approach mirrors project management methodologies where complex deliverables are recursively divided into smaller, actionable work packages. The hierarchical structure also facilitates parallelization, as independent sub-goals can be delegated to separate worker processes running simultaneously. The challenge is correctly identifying the right decomposition: a poor decomposition can create sub-goals that are internally coherent but fail to combine into a correct overall solution.

Tree search planning explores multiple possible action sequences simultaneously, evaluating each branch's likely outcome before committing to one. This approach, common in classical AI planning and game playing, allows agents to backtrack from dead ends and select optimal paths through combinatorial search spaces. For example, when debugging code, an agent might explore several hypotheses about the bug's cause simultaneously, pursuing each line of investigation until evidence supports or refutes it, then consolidating the findings into a diagnosis. While computationally expensive, tree search protects against local optima and unexpected obstacles. Current LLM-based agents rarely implement full tree search due to cost, but prompting techniques like "let's consider several approaches" approximate it by generating multiple candidate plans before selecting one.

Out[5]:
Visualization
Grouped bar chart comparing Single-step, Chain-of-thought, Hierarchical, and Tree search planning strategies on computational cost, adaptability, and global coherence dimensions rated from 1 to 5.
Performance profiles of four planning strategies across computational cost, adaptability, and global coherence dimensions rated from 1 to 5. Tree search reaches the highest global coherence (5.0) but at maximum computational cost (4.5), which makes it suitable for offline planning but impractical for real-time agents. Chain-of-thought offers a favorable intermediate position with moderate cost (1.5), strong adaptability (3.5), and meaningfully improved coherence (2.5) over single-step baselines.

Plan Representation

Plans can be represented in several formats, each with tradeoffs in expressiveness and interpretability as well as ease of execution.

Linear sequences specify an ordered list of actions such as [Search, Extract, Summarize]. This works well for straightforward workflows where each step follows naturally from the last, but struggles with conditional logic. Linear plans are easy to visualize and validate, and they give a clear sense of progress as items are checked off. However, they assume the environment will cooperate with the predetermined sequence. If the Search step returns empty results, a linear plan has no mechanism for branching to a different information source: it simply proceeds to Extract, operating on nothing.

State machines define transitions between states based on observations. The agent moves from "Researching" to "Synthesizing" only when sufficient information has been gathered, handling branching and looping naturally. State machines excel at representing workflows with clear phase transitions and validation gates. A data processing pipeline might use a state machine where records must pass quality checks before advancing to analysis stages, and failed records branch to a review queue rather than proceeding. The state machine gives explicit structure for the agent's workflow phases, which makes it easy to reason about which state the agent should be in at any point and what conditions govern transitions.

Programmatic structures stand for plans as code or pseudocode, allowing complex control flow with loops and conditionals as well as function calls. This approach treats planning as code generation, using the LLM's training on programming languages to produce structured execution plans. A programmatic plan might look like a function with conditional branches and loops expressed in natural pseudocode. This representation bridges the gap between high-level planning and executable instructions, though it requires careful sandboxing to prevent unsafe code execution when plans are translated into actual runnable code.

Replanning and Adaptation

Static plans fail when environments change or actions produce unexpected results. A plan formulated at the outset of a task is based on assumptions about the world that may prove false as the agent gathers information. Effective agent architectures incorporate replanning triggers: conditions that invalidate current plans and initiate new planning cycles.

Common replanning triggers include tool execution failures or unexpected return values that indicate the current approach is unworkable, new information that invalidates previous assumptions about the task structure, user interventions or goal modifications that shift the objective mid-stream, and timeout conditions showing stalled progress or potential infinite loops.

Replanning frequency presents a real design tradeoff. Replanning too frequently wastes computation and creates instability, as the agent constantly second-guesses itself without making sustained progress toward the goal. But replanning too infrequently produces rigid behavior that cannot adapt to changing circumstances, like a navigation app that insists on turning left onto a road that was closed after the route was calculated. Many architectures implement opportunistic replanning: maintaining stable plans while remaining alert to significant environmental changes that warrant a strategic revision, much like a human project manager who follows a project plan but calls a team meeting when a key dependency unexpectedly fails.

The decision about when to replan can itself be part of the agent's reasoning. At each iteration, the agent can evaluate whether its current plan remains valid given the latest observation, and trigger replanning only when the gap between expected and actual outcomes exceeds a threshold. This meta-level reasoning about plans distinguishes advanced agents from brittle ones.

Termination Conditions

Without explicit termination conditions, agents might loop indefinitely, consume excessive resources, or pursue goals long after they have been achieved. Defining when an agent should stop is as necessary as defining how it should act. Termination conditions serve as guardrails that ensure agents remain practical tools rather than runaway processes, and they give the closure that allows users to receive results and move on.

Termination is philosophically interesting because it requires the agent to judge its own work. A human completing a task has an intuitive sense of "doneness": the report is written, the dish is cooked, the equation is solved. An LLM-based agent must make this judgment computationally, often without perfect information. Designing termination conditions that are neither too eager (cutting off before the task is complete) nor too reluctant (continuing long after the goal is achieved) is one of the more fine-grained aspects of agent architecture design.

Goal-Based Termination

The most intuitive termination condition occurs when the agent reaches its stated objective. For question-answering agents, this means giving a satisfactory answer; for task-completion agents, it means executing the requested action successfully. However, determining "success" requires careful definition upfront, because ambiguous success criteria produce agents that cannot reliably know when to stop.

Explicit success signals occur when external systems confirm task completion: a file is successfully written, a database transaction commits, or an API returns a success status code. These give unambiguous termination criteria. When a file write returns a success code, the agent can confidently conclude that the persistence task is complete. These signals are objective and verifiable, eliminating the need for the agent to make a subjective judgment about whether the work is done. Whenever possible, designing tasks around explicit success signals makes termination reliable and predictable.

Implicit success detection requires the agent to recognize when its goal is satisfied without an external signal. A research agent might terminate when it has gathered sufficient information to answer the original question, requiring self-assessment of coverage and confidence. This is akin to knowing when to stop reading and start writing an essay: there is no external signal that you have read enough, only an internal judgment that you have achieved sufficient understanding. Implementing implicit success detection often requires heuristics such as information saturation (new sources are not adding facts not already in working memory), confidence thresholds in generated answers, or a coverage checklist that the agent maintains and checks against its accumulated knowledge.

Resource Limits

Practical deployments impose hard constraints on agent execution to control costs and prevent system degradation. These limits are not signs of poor design but prudent engineering: they acknowledge that agents operate in resource-constrained environments and must be accountable for what they consume.

Token budgets limit the total number of tokens processed or generated. Given the cost and latency implications of LLM inference, agents might terminate when approaching token limits, returning partial results or requesting user guidance. This is especially important in commercial deployments where API costs scale directly with token usage. An agent configured with a budget of 100,000 tokens must prioritize its remaining actions carefully once it has consumed 80,000, perhaps switching from exploratory search to synthesis even if the exploration feels incomplete. Token budgets force agents to make economically rational tradeoffs between thoroughness and cost.

Time limits cap wall-clock execution duration. Long-running agents might hit timeouts during extended tool executions or during reasoning phases. Time limits protect against hung processes and ensure that users receive responses within acceptable windows. They also force agents to prioritize speed versus thoroughness in ways that should be calibrated to the use case: a customer service chatbot needs responses in under a second, while a research assistant running as a background task might be allowed minutes or even hours to compile a complete report. The time limit communicates to the agent what kind of task it is performing and what level of thoroughness is expected.

Iteration limits restrict the number of loop cycles. This prevents infinite loops in misconfigured agents and guards against tasks that expand beyond their original scope through uncontrolled replanning or goal drift. Without iteration limits, an agent stuck in a cycle of "search for more information" might never conclude that it has enough data. Iteration limits are typically set between 5 and 50 depending on task complexity: simple fact retrieval might need only 3 iterations, while complex debugging might require 20 or more before a fix is found and validated.

Safety Termination

Beyond resource constraints, safety considerations demand termination conditions that protect users and systems from harm. These conditions must fire unconditionally, regardless of task completion status, because the risk of proceeding outweighs any benefit.

Error thresholds trigger termination when consecutive tool failures or reasoning errors exceed a set limit, preventing agents from spiraling into unproductive error states. If an agent fails to connect to a database five times consecutively, continuing to retry is unlikely to help and may lock resources, trigger security alerts, or violate rate limits. Error thresholds recognize when the agent has lost its ability to make progress and redirect control to human operators or graceful failure handlers.

Content policies halt execution when agents generate prohibited content or attempt harmful actions. An agent instructed to research sensitive topics might encounter content that violates privacy policies, generates legally problematic material, or produces outputs that could facilitate harm. Content policy checks should run after every action and terminate execution immediately when violations are detected, with clear logging to support incident review.

User intervention allows human operators to terminate or redirect agents during execution, maintaining human oversight over autonomous systems. This might take the form of a "stop button" in a user interface, or periodic checkpoints where the agent pauses to request confirmation before proceeding with irreversible actions like deleting files, sending emails, or committing database transactions. User intervention preserves human agency in necessary decision loops. This keeps automated systems remain tools serving human intent rather than autonomous actors pursuing misaligned objectives. As agent autonomy increases, the design of intervention mechanisms becomes more important, not less: the more capable an agent is, the more consequential it is to get these mechanisms right.

Worked Example: Research Assistant Agent

To solidify these concepts, let us design a research assistant agent that helps users investigate topics by searching the web, reading documents, and synthesizing findings. This example illustrates how the agent loop and state management work with planning and termination to produce coherent, goal-directed behavior across multiple iterations.

Consider the user query: "What are the latest developments in efficient transformer attention mechanisms from 2024?"

Decomposing the Task

Upon receiving this query, the agent enters the planning phase. It must decompose a broad, open-ended request into concrete, executable subtasks. The agent recognizes several implicit requirements: "latest developments" implies a need for recency filtering, "efficient transformer attention" suggests specific technical keywords such as FlashAttention, linear attention, and sparse attention patterns, and "developments" (plural) implies the agent should gather multiple sources rather than stopping at the first relevant result.

The agent formulates an initial plan with three stages. First, search for recent papers and blog posts on efficient attention mechanisms from 2024 to build a list of candidate sources. Second, prioritize and read the most promising sources, extracting key claims and performance benchmarks. Third, synthesize the findings into a coherent summary that covers the main innovations, their claimed benefits, and their practical status.

This initial plan is deliberately high-level. The agent does not yet know which search queries will be most fruitful, how many sources will be relevant, or what unexpected angles will emerge from the literature. The plan gives structure and direction, but the agent expects to refine it as it learns more about the subject.

Iterative Execution

Iteration 1: Initial search. The agent executes a search for "efficient attention transformers 2024" on arXiv. The search returns 15 papers with varying levels of relevance. Rather than immediately reading all of them, which would be inefficient, the agent updates its working memory with the paper list and metadata (titles, abstracts, citation counts) and applies a prioritization heuristic: papers with more citations or titles explicitly mentioning FlashAttention, linear attention, or sparse patterns receive higher priority. The agent selects the top 5 papers for deeper reading. This shows a key planning adaptation: the agent expected a small number of highly relevant papers, found instead a large set of varying relevance, and adjusted its strategy to triage before committing to deep reading.

Iteration 2: Abstract extraction. The agent fetches the abstracts and introduction sections of the top 5 papers. It extracts key claims into structured working memory entries: FlashAttention-3 reaches roughly 1.5x speedup on H100 GPUs through warp-specialization, a new linear attention method reduces complexity from O(n2)O(n^2) to O(n)O(n) while maintaining 95% of standard attention quality, and a sparse attention approach reduces memory usage by 60% on sequences longer than 8,192 tokens. These facts are now in working memory as structured data, not just raw text, making them easy to reference in later reasoning steps without re-reading the abstracts.

Iteration 3: Targeted deep dive. During its reasoning phase, the agent identifies a gap: it has abstracts claiming performance improvements, but lacks implementation details or independent evaluations for FlashAttention-3 that would help a reader assess practical applicability. It updates its plan to include a targeted search for FlashAttention-3 GitHub repositories, technical blog posts, and benchmarking results. This replanning shows the flexibility of the architecture: new information revealed a gap in the initial plan, triggering a targeted addition rather than a wholesale plan replacement.

Iteration 4: Synthesis preparation. The agent now has sufficient raw material and begins organizing its findings into synthesis categories: hardware-aware algorithmic optimizations (FlashAttention-3), mathematical approximations that change the complexity class (linear attention), and structural sparsity methods (sparse patterns). It drafts an outline for the final response and identifies any remaining gaps, noting that it has limited information about practical adoption in production systems.

Iteration 5: Termination and synthesis. The agent applies implicit success detection. It checks whether recent queries are returning redundant information already captured in working memory: yes, the last search returned papers already in the list. It checks whether its synthesis outline covers all major themes from the collected sources: yes. It determines that additional iterations would yield diminishing returns and generates the final synthesis, citing specific papers and benchmarks, and terminates.

What the Example Illustrates

Throughout this process, the agent maintained conversation history (all search queries and their full results), working memory (extracted facts and priorities plus synthesis outlines), and plan state (current subtask, identified gaps, and completed stages). It replanned when searches returned unexpected volumes or exposed knowledge gaps, and it terminated based on an internally derived coverage criterion rather than a fixed iteration count. The four architectural components worked together smoothly: the loop drove execution, state enabled context-awareness across iterations, planning provided strategic coherence, and termination ensured the agent concluded appropriately.

Out[6]:
Visualization
Horizontal bar timeline showing five sequential agent iterations labeled Planning, Search, Reading, Analysis, and Synthesis across iteration numbers 1 through 5.
Research assistant agent execution across five iterations showing the progression from planning through information gathering to final synthesis. Early iterations focus on information acquisition (Search, Reading), while later iterations shift to knowledge processing (Analysis, Synthesis), which shows the natural arc of a research task.

Code Implementation

Let's implement a simplified agent architecture that shows the core loop, state management, and planning components. Our agent will solve mathematical word problems by using a calculator tool, showing how the architectural principles translate into working code.

The goal is not production-ready software but rather a clear illustration of how each architectural component is expressed in code. By tracing the execution of this simplified agent, you can see exactly how the loop drives iterations, how working memory accumulates results, and how termination conditions fire.

We start by defining the data structures that stand for the agent's state:

In[7]:
Code
from dataclasses import dataclass, field
from typing import Any, Dict, List, Optional


@dataclass
class Observation:
    """Input received from the environment at each iteration."""

    content: str
    source: str  # 'user', 'tool', or 'system'
    step: int


@dataclass
class Action:
    """The agent's decision about what to do next."""

    tool_name: Optional[str]  # None signals a final answer (termination)
    tool_input: Dict[str, Any]
    reasoning: str


@dataclass
class AgentStep:
    """A complete record of one agent loop iteration."""

    observation: Observation
    thought: str
    action: Action
    result: Optional[str] = None


@dataclass
class AgentState:
    """The full state of the agent across all iterations."""

    steps: List[AgentStep] = field(default_factory=list)
    working_memory: Dict[str, Any] = field(default_factory=dict)
    current_plan: List[str] = field(default_factory=list)
    iteration_count: int = 0

These four dataclasses capture the key concepts from our earlier discussion. Observation is input from the environment. Action is the agent's output decision, with tool_name=None serving as the implicit termination signal. AgentStep records a complete iteration for later inspection. AgentState aggregates everything the agent knows and has done, representing its persistent memory.

Now we implement the calculator tool and the tool registry using safe arithmetic (avoiding dynamic code execution by parsing operations explicitly):

In[8]:
Code
import operator


def calculator(expression: str) -> str:
    """
    Safe calculator that evaluates simple arithmetic expressions.
    Supports addition, subtraction, multiplication, and division.
    Parses operations explicitly to avoid any code execution concerns.
    """
    # Remove all whitespace
    expression = expression.replace(" ", "")

    # Try to match a simple two-operand expression
    pattern = re.compile(r"^(-?\d+(?:\.\d+)?)([\+\-\*/])(-?\d+(?:\.\d+)?)$")
    match = pattern.match(expression)
    if not match:
        return "Error: Expression must be in the form 'a op b' (e.g. '100 / 4')"

    left = float(match.group(1))
    op_char = match.group(2)
    right = float(match.group(3))

    ops = {
        "+": operator.add,
        "-": operator.sub,
        "*": operator.mul,
        "/": operator.truediv,
    }

    if op_char not in ops:
        return f"Error: Unsupported operator '{op_char}'"

    if op_char == "/" and right == 0:
        return "Error: Division by zero"

    result = ops[op_char](left, right)
    # Return integer if result is whole number
    if result == int(result):
        return str(int(result))
    return str(round(result, 6))


# Tool registry: maps tool names to their functions and schemas
TOOLS = {
    "calculator": {
        "function": calculator,
        "description": "Calculate simple two-operand arithmetic. Input: expression like '150 - 45'.",
        "parameters": {
            "type": "object",
            "properties": {
                "expression": {
                    "type": "string",
                    "description": "Arithmetic expression to evaluate, e.g. '150 - 45'",
                }
            },
            "required": ["expression"],
        },
    }
}

The core of our agent is the planning and execution logic. Since this example runs without a live LLM API, we simulate the reasoning step with deterministic heuristics that capture the needed decision patterns:

In[9]:
Code
import re


class SimpleAgent:
    """
    Agent architecture showing loop, state, planning, and termination.
    Uses rule-based reasoning to simulate LLM decision making.
    """

    def __init__(self, max_iterations: int = 5, tools: Dict = None):
        self.max_iterations = max_iterations
        self.tools = tools or {}
        self.state = AgentState()

    def plan(self, observation: Observation) -> List[str]:
        """
        Generate a task plan based on the observation content.
        A real agent would call an LLM here.
        """
        content = observation.content.lower()
        if "calculate" in content or any(
            op in content for op in ["+", "-", "*", "/"]
        ):
            return ["extract_expression", "calculate", "verify"]
        return ["analyze", "respond"]

    def _extract_expression(self, content: str) -> Optional[str]:
        """Extract a simple arithmetic expression from the content string."""
        # Strip leading instruction words
        text = re.sub(r"calculate\s*", "", content, flags=re.IGNORECASE).strip()

        # Match patterns like "150 - 45" or "100/4"
        pattern = re.compile(
            r"(-?\d+(?:\.\d+)?)\s*([\+\-\*/])\s*(-?\d+(?:\.\d+)?)"
        )
        match = pattern.search(text)
        if match:
            return f"{match.group(1)}{match.group(2)}{match.group(3)}"

        # Detect word-form operations
        numbers = re.findall(r"\d+(?:\.\d+)?", content)
        if len(numbers) >= 2:
            if "minus" in content.lower() or "subtract" in content.lower():
                return f"{numbers[0]}-{numbers[1]}"
            if (
                "plus" in content.lower()
                or "total" in content.lower()
                or "sum" in content.lower()
            ):
                return f"{numbers[0]}+{numbers[1]}"
            if "divided" in content.lower():
                return f"{numbers[0]}/{numbers[1]}"
            if "times" in content.lower() or "multiplied" in content.lower():
                return f"{numbers[0]}*{numbers[1]}"

        return None

    def decide_action(self, observation: Observation) -> Action:
        """
        Choose the next action based on observation and current plan.
        A real agent would call an LLM here.
        """
        content = observation.content

        # Hard termination: iteration limit exceeded
        if self.state.iteration_count >= self.max_iterations:
            return Action(
                tool_name=None,
                tool_input={"answer": "Maximum iterations reached"},
                reasoning="Terminating due to iteration limit",
            )

        # Try to extract and calculate an arithmetic expression
        expr = self._extract_expression(content)
        if expr:
            return Action(
                tool_name="calculator",
                tool_input={"expression": expr},
                reasoning=f"Arithmetic expression detected: {expr}",
            )

        # Default: answer directly without a tool
        return Action(
            tool_name=None,
            tool_input={"answer": f"Processed: {content}"},
            reasoning="No arithmetic detected, giving direct response",
        )

    def should_terminate(self, action: Action) -> bool:
        """Determine whether the agent should stop after this action."""
        if action.tool_name is None:
            return True
        if self.state.iteration_count >= self.max_iterations:
            return True
        return False

    def run(self, query: str) -> Dict:
        """Execute the agent loop until a termination condition is met."""
        print(f"Starting agent with query: '{query}'\n")

        observation = Observation(content=query, source="user", step=0)

        while True:
            self.state.iteration_count += 1
            step_num = self.state.iteration_count

            print(f"Iteration {step_num}")
            print(f"  Observation: {observation.content[:70]}")

            # Planning phase: generate plan on first iteration
            if not self.state.current_plan:
                self.state.current_plan = self.plan(observation)
                print(f"  Plan: {' -> '.join(self.state.current_plan)}")

            # Reasoning phase: decide next action
            action = self.decide_action(observation)
            print(f"  Reasoning: {action.reasoning}")

            if action.tool_name:
                print(f"  Action: {action.tool_name}({action.tool_input})")
            else:
                print("  Action: Final answer")

            # Termination check
            if self.should_terminate(action):
                step = AgentStep(
                    observation=observation,
                    thought=action.reasoning,
                    action=action,
                    result=action.tool_input.get("answer", "Done"),
                )
                self.state.steps.append(step)
                print(f"\nTerminated after {step_num} iteration(s)")
                return {
                    "answer": action.tool_input.get("answer"),
                    "iterations": step_num,
                    "history": self.state.steps,
                }

            # Execution phase: call the tool
            tool_func = self.tools.get(action.tool_name, {}).get("function")
            if tool_func:
                try:
                    result = tool_func(**action.tool_input)
                    print(f"  Result: {result}")
                except Exception as e:
                    result = f"Error: {str(e)}"
                    print(f"  Result: {result}")
            else:
                result = f"Unknown tool: {action.tool_name}"
                print(f"  Result: {result}")

            # Update working memory with tool result
            self.state.working_memory[f"step_{step_num}_result"] = result

            # Record the completed step
            step = AgentStep(
                observation=observation,
                thought=action.reasoning,
                action=action,
                result=result,
            )
            self.state.steps.append(step)

            # The tool result becomes the next observation
            observation = Observation(
                content=result, source="tool", step=step_num
            )

            # Advance the plan
            if self.state.current_plan:
                self.state.current_plan.pop(0)

            print()

The run method is where all four architectural components converge. The while True loop with an explicit return is our agent loop. The self.state object is our state management. The self.plan() call is our planning phase. The self.should_terminate() check implements our termination conditions. Notice how naturally each component maps to a distinct section of code.

Let's run the agent on a word problem that requires calculation:

In[10]:
Code
agent = SimpleAgent(max_iterations=3, tools=TOOLS)
result = agent.run(
    "If I have 150 apples and give away 45, how many remain? Calculate 150 - 45"
)
Out[11]:
Console
Final Answer: Processed: 105
Total iterations: 2
Working memory keys: ['step_1_result']

The agent recognized the arithmetic expression, invoked the calculator tool with 150-45, received the result 105, and terminated with a final answer. Working memory contains one entry recording the tool result, which would be available for subsequent reasoning if the task required further computation.

Now let's examine state accumulation more explicitly with a stateful extension:

In[12]:
Code
class StatefulAgent(SimpleAgent):
    """Extended agent with explicit state reporting capabilities."""

    def get_state_summary(self) -> str:
        """Generate a human-readable summary of the agent's current state."""
        lines = [
            f"Iterations completed: {self.state.iteration_count}",
            f"Working memory entries: {len(self.state.working_memory)}",
            f"Conversation length: {len(self.state.steps)} step(s)",
        ]

        if self.state.steps:
            last_step = self.state.steps[-1]
            lines.append(
                f"Last action: {last_step.action.tool_name or 'final_answer'}"
            )
            lines.append(f"Last result: {str(last_step.result)[:60]}")

        if self.state.working_memory:
            lines.append("Working memory contents:")
            for key, val in self.state.working_memory.items():
                lines.append(f"  {key}: {val}")

        return "\n".join(lines)


stateful_agent = StatefulAgent(max_iterations=3, tools=TOOLS)
stateful_result = stateful_agent.run("Calculate 100 / 4")
Out[13]:
Console
Final Answer: Processed: 25
Total iterations: 2

Final State Summary:
Iterations completed: 2
Working memory entries: 1
Conversation length: 2 step(s)
Last action: final_answer
Last result: Processed: 25
Working memory contents:
  step_1_result: 25

The state summary reveals what the agent retained after execution: the number of iterations, the working memory entries with their calculated values, and the action history. In a more complex agent, working memory might contain dozens of entries representing facts gathered across many iterations, all available for synthesis at the final step.

Let's visualize the agent control loop across iterations:

Out[14]:
Visualization
Flowchart showing two agent loop iterations with observation, reasoning, action, and termination check stages connected by arrows.
Agent control loop execution flow across two iterations. The first iteration processes the user query, invokes the calculator tool, and records the result. The second iteration receives the tool output as its observation, determines the goal is complete, and terminates. The Agent State box shows the working memory that persists across iterations, letting context-aware decision making.

The diagram captures the needed control flow: two iterations connected by an arrow representing the feedback loop where tool results become the next iteration's observations. The state box on the right accumulates working memory across iterations, representing the agent's persistent context.

Key Design Parameters

The key parameters for our agent architecture are:

  • max_iterations: Maximum number of loop cycles before forced termination. This acts as a necessary safety guardrail, preventing infinite loops and controlling computation costs. Setting it too low may cause the agent to terminate prematurely before solving complex problems; too high risks excessive costs if the agent enters an unproductive cycle. Typical values range from 3 for simple tasks to 20 or more for complex research or debugging workflows.

  • tools: Dictionary of available functions the agent can invoke. Each entry specifies the function and description as well as the parameter schema. The quality of tool descriptions materially impacts agent performance, because these descriptions form part of the prompt that guides the LLM's reasoning about which tool to use. Poorly described tools lead to confusion and incorrect selections.

  • working_memory: Dictionary storing intermediate results and extracted facts across iterations, letting stateful behavior. The structure and content of working memory determine what information the agent retains for future reasoning. Well-designed schemas organize information by type and priority, while poor schemas allow important details to be buried under irrelevant data.

  • current_plan: Ordered list of planned actions guiding the agent's execution strategy, updated as tasks progress. Dynamic plan updating allows the agent to adapt to new information while maintaining strategic coherence.

  • steps: Complete history of observations and thoughts followed by actions and results, giving conversation context. This chronological record is the agent's episodic memory for the current session. This makes it possible to avoid redundant actions and justify its decisions based on prior reasoning.

Limitations and Impact

Agent architectures, while powerful, introduce significant challenges that practitioners must manage carefully. The flexibility that lets complex problem-solving also creates vectors for failure modes that simpler systems avoid entirely. Understanding these limitations is needed for deploying agents responsibly and effectively in production environments.

Reliability and Error Propagation

Perhaps the most pressing practical concern is that errors compound across iterations. In a single-step function-calling system, an error is localized and easy to attribute: either the tool call succeeded or it failed, and the failure is immediately visible. In a multi-step agent loop, an incorrect inference in step 3 might not manifest obviously until step 7, by which time the agent has built elaborate reasoning on false premises and taken several actions based on that reasoning. Without explicit verification steps or self-consistency checks, agents can confidently pursue entirely wrong trajectories while appearing to make progress.

This problem is compounded by the LLM's tendency toward confident generation even when uncertain. A language model that is unsure whether a web search returned reliable results may still generate a confident synthesis, masking the uncertainty from the planning logic. Building in explicit uncertainty tracking, where the agent maintains confidence scores for key claims and triggers human review when confidence drops below a threshold, is one mitigation strategy, but it adds complexity and is itself susceptible to miscalibration.

Out[15]:
Visualization
Line chart showing cumulative error probability from iteration 1 to 20, with a red solid line for no validation rising steeply and a green dashed line for validation every 5 steps staying below 25 percent.
Cumulative error probability across 20 agent loop iterations comparing unvalidated execution against a periodic validation strategy. Without validation, error probability reaches 64% by iteration 20 due to compounding effects at a 5% per-step base error rate. Validation checkpoints every 5 iterations reset accumulation and maintain error probability below 23%, showing the value of intermediate verification even at moderate cost.

The figure illustrates why validation checkpoints matter even when they add cost. At a modest 5% per-step error rate, unvalidated agents face a 64% chance of accumulated error by iteration 20. Periodic validation resets this accumulation, keeping cumulative error probability manageable across long-running tasks. The tradeoff is the additional tokens and latency of the validation steps themselves.

The Exploration-Exploitation Tradeoff

Agent architectures face a basic optimization problem in deciding how to allocate their iteration budget between exploration and exploitation. Too much exploration, trying different approaches, searching for additional information, considering alternative plans, wastes tokens and time without converging on a solution. Too much exploitation, committing early to the first promising approach, leads to local optima where agents miss better solutions because they failed to consider alternatives.

Current architectures rely heavily on heuristics to balance this tradeoff: fixed iteration limits, simple termination conditions based on the absence of new information, and greedy action selection that always picks the highest-confidence next step. These heuristics work reasonably well for well-structured tasks but can fail badly on tasks with deceptive local optima, where the most obvious approach turns out to be a dead end that only becomes apparent after significant investment.

More advanced approaches might employ tree search, Monte Carlo rollouts, or learned value functions to guide exploration, similar to the techniques used in reinforcement learning. However, these methods increase computational cost substantially and require either a simulator of the environment (for tree search) or extensive training data (for learned value functions). The practical reality is that most deployed agents use greedy, heuristic-based planning because it is the only approach that remains tractable at inference time with current hardware.

State Management at Scale

Our implementation demonstrated simple working memory, but real-world agents managing database transactions, file systems, and API states face consistency challenges that echo those in distributed systems. When an agent executes a sequence of operations, reading a file, processing it, and writing results, what happens if the write fails? Should the agent roll back previous operations? Should it retry? Should it escalate to human operators? These are the distributed systems problems that database engineers have addressed with ACID transactions and two-phase commit protocols, but those solutions assume a controlled database environment. Agent architectures operate across heterogeneous external systems that may not give transactional guarantees.

Building reliable state management requires importing lessons from distributed computing: idempotency keys to prevent duplicate operations when retrying failed steps, circuit breakers to fail fast when dependent services are unavailable, and saga patterns to manage distributed transactions that span multiple systems. Most current agent frameworks address these challenges incompletely, placing the burden of state consistency on the application developer rather than giving reliable infrastructure support.

Latency and Cost

Each loop iteration requires at least one LLM inference call. An agent that iterates 10 times to solve a problem costs 10 times more in API fees and generates 10 times the latency compared to a single-shot solution. This creates real pressure to minimize iterations, sometimes at the expense of solution quality. For consumer applications with sub-second response expectations, agents requiring multiple reasoning steps may simply not be deployable without aggressive caching, batching, or distillation into smaller, faster models fine-tuned for specific tasks.

The cost issue is particularly acute for hierarchical agent architectures, where a top-level planner spawns multiple specialized sub-agents, each running their own internal loops. A task that decomposes into five parallel sub-tasks, each requiring five iterations, generates 25 LLM calls plus the planner's own calls, potentially making the overall system an order of magnitude more expensive than a direct single-model approach. Cost-effectiveness analysis is a necessary part of agent architecture design, not an afterthought.

Safety and Alignment

Safety concerns intensify with agent autonomy in ways that differ qualitatively from single-shot systems. A function-calling system does exactly what it is asked to do. An agent, given a goal and left to pursue it through multiple iterations with broad tool access, might develop instrumental subgoals that conflict with operator intent. The agent might determine that to complete its task efficiently, it should bypass rate limits, access data it was not explicitly authorized to use, or take irreversible actions without seeking confirmation.

This phenomenon, sometimes called instrumental convergence in the AI safety literature, suggests that agents with broad capabilities may default to certain problematic behaviors regardless of their terminal goals, unless those behaviors are explicitly constrained. The concern is not that agents are malicious but that goal-directedness combined with insufficient specification of constraints can produce unexpected solutions that violate constraints. Ensuring that agent architectures remain controllable and aligned with operator intent requires reliable sandboxing, minimal-privilege tool access, explicit constraint sets, and human review checkpoints for consequential actions. We will explore these considerations in greater depth in the upcoming chapter on Agent Evaluation and Safety.

Despite these challenges, agent architectures stand for a important evolution in language AI capability. They transform LLMs from passive responders into active problem-solvers capable of sustained effort, error recovery, and adaptation. The architectural patterns we have examined here, loops for sustained action, state for context maintenance, planning for strategic coherence, and termination for bounded behavior, give the scaffolding upon which increasingly capable autonomous systems are being built. In the next chapter on Agent Memory, we will see how extending these architectures with long-term, cross-session memory capabilities further expands what language model-based agents can reach.

Summary

Agent architectures give the structural framework that turns Large Language Models from stateless text generators into goal-directed systems able to sustained, adaptive behavior. This chapter examined four pillars of agent design, illustrating how each component contributes to coherent and reliable autonomous operation.

The agent loop establishes the basic control flow: Observation and Reasoning followed by Action, cycling continuously until task completion. Whether implemented as simple single-step loops or complex hierarchical structures, the loop is what turns static inference into dynamic interaction with the world. Each iteration produces an observation that informs the next. This creates a chain of context-aware decisions that can accumulate toward goals too complex for any single model call.

State management addresses the challenge of maintaining context across iterations. Through conversation history, working memory, episodic memory, and tool state, agents accumulate information that informs future decisions. The tension between complete state and limited context windows drives architectural innovations in summarization, hierarchical organization, and external retrieval. Effective state management distinguishes agents that build coherently on prior work from those that repeat the same mistakes across iterations.

Planning strategies determine how agents approach complex tasks. From reactive single-step planning to deliberative hierarchical task networks, the planning component structures agent behavior toward goal achievement. Effective architectures balance plan stability with adaptive replanning, letting agents to pursue coherent long-term strategies while remaining responsive to unexpected findings. The representation of plans as linear sequences, state machines, or programmatic structures shapes the flexibility and reliability of the agent's approach to novel situations.

Termination conditions bound agent execution, preventing infinite loops and resource exhaustion. Whether triggered by goal achievement, resource limits, or safety constraints, termination mechanisms ensure that agents remain practical tools rather than uncontrolled processes. Defining success criteria, whether through explicit external signals or implicit coverage heuristics, requires careful consideration of both user needs and operational constraints.

Together, these components form the blueprint for autonomous language AI systems. While significant challenges remain in reliability and efficiency, with safety requiring equal attention, the architectural patterns established in this chapter are the foundation on which increasingly capable agents are being built. The next chapter extends this foundation by examining how long-term memory systems allow agents to accumulate knowledge and learn from experience across many conversations, not just within a single session.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about agent architectures, state management, and planning strategies.

Agent Architectures Quiz

Question 1 of 60 of 6 completed
What are the three fundamental operations in the basic agent loop cycle?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026agentarchitectures, author = {Michael Brenndoerfer}, title = {Agent Architectures: Control Loops, State & Planning}, year = {2026}, url = {https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Agent Architectures: Control Loops, State & Planning. Retrieved from https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination
MLAAcademic
Michael Brenndoerfer. "Agent Architectures: Control Loops, State & Planning." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination>.
CHICAGOAcademic
Michael Brenndoerfer. "Agent Architectures: Control Loops, State & Planning." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Agent Architectures: Control Loops, State & Planning'. Available at: https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Agent Architectures: Control Loops, State & Planning. https://mbrenndoerfer.com/writing/agent-architectures-loop-state-planning-termination

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.