ReAct: Reasoning and Action for LLM Agents

Michael BrenndoerferFebruary 3, 202650 min read

Part of Language AI Handbook

Explains how the ReAct pattern enables LLM agents to interleave reasoning with tool execution.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

ReAct Pattern

Large language models generate coherent text, analyze complex ideas, and synthesize information from their training data. However, they face a basic constraint: they cannot reliably access external information, verify facts against current data sources, or interact with computational tools beyond their static parameter knowledge. In the previous chapter, we explored Function Calling, which provides models with the syntactic structures necessary to invoke external tools and APIs. While this capability represents a significant advancement, knowing how to call a function differs fundamentally from knowing when to invoke it, which specific function to choose among alternatives, and what to do with the results once they arrive.

Consider the contrast between knowing a hammer exists and knowing how to build a house. Function calling gives the model a hammer. ReAct teaches it to build. The difference lies not in the tool itself, but in the reasoning that governs when and how to swing it. A model that knows how to call a web search API still needs to decide what to search for, interpret the result it receives, recognize when the result answers a sub-question, and determine what to search next. Each of those decisions requires the model to maintain a mental model of where it stands in the problem-solving process and what still needs to be resolved.

ReAct, which stands for Reasoning + Acting, bridges this necessary gap by interleaving chain-of-thought reasoning with tool execution in a tight, iterative loop. Instead of generating a complete, monolithic plan upfront or making blind, uncontextualized tool calls, the model thinks step-by-step, acts on those thoughts, observes the results of those actions, and adapts its next thought based on the new information received. This creates a dynamic, interactive problem-solving process that closely mirrors how humans approach complex tasks in the real world: we think, act, observe the consequences, and reconsider our approach based on what we learn.

The ReAct pattern turns a static language model into an interactive agent able to multi-step reasoning, dynamic information gathering, and sophisticated error recovery. It shifts from simple prompt-response interactions to persistent, stateful computation where the model maintains working memory of its reasoning trajectory. This allows it to solve problems beyond the complexity of single-turn queries.

This design emerged from a well-observed limitation in pure Chain-of-Thought prompting (covered in Part XXX, Chapter 3), where models reason beautifully over problems they already know the answer to but hallucinate fluently when they lack the necessary knowledge. ReAct solves this by making knowledge retrieval an explicit, auditable step rather than an implicit, invisible one. The model can no longer silently confabulate: it must fetch the information it claims to need.

The ReAct Formulation

ReAct, introduced by Yao et al. (2022), formalizes the interaction between an LLM and an environment as a structured sequence of thought-action-observation triples. This mathematical framework characterizes how the agent maintains state and progresses toward goals through iterative refinement.

The core insight behind the formulation is that an agent solving a task should not be treated as a single forward pass through a model, but as a dynamic process unfolding over time. Each moment in that process consists of three things: what the agent currently thinks (thought), what it chooses to do (action), and what it learns from doing it (observation). Concatenate enough of these moments and you have a complete trace of how the agent reasoned its way to an answer.

The model maintains a trajectory composed of these sequential triples, represented mathematically as:

τ=(t1,a1,o1,t2,a2,o2,…,tn,an,on)\tau = (t_1, a_1, o_1, t_2, a_2, o_2, \ldots, t_n, a_n, o_n)

where:

  • τ\tau: the complete trajectory of the ReAct session, representing the full history of the agent's reasoning and actions
  • tit_i: the thought generated at step ii, containing the model's internal reasoning about its current state
  • aia_i: the action taken at step ii, representing a specific tool invocation from the available action space
  • oio_i: the observation received after executing action aia_i, giving environmental feedback from the external system

Each step in this sequence consists of three distinct components working in concert:

  • Thought (tit_i): The model's internal reasoning about the current state, including what it currently knows, what information it still needs to acquire, and what specific action to take next to progress toward the goal. Thoughts are free-form natural language, though they should be focused and task-relevant.
  • Action (aia_i): A specific tool invocation with parameters, drawn from a predefined action space available to the agent. Actions must be parseable by the environment, so their format is more constrained than thoughts.
  • Observation (oio_i): The factual result returned by the environment after executing the action. This provides ground truth for subsequent reasoning. Observations are not generated by the model; they come from external systems.

This formulation differs from Chain-of-Thought prompting, where reasoning occurs entirely within the model's internal monologue without environmental interaction. In ReAct, thoughts drive actions that change the external state or retrieve new information, and observations feed back to inform subsequent reasoning steps. The environment becomes an integral part of the computational graph rather than a passive recipient of outputs.

The formal trajectory also is the agent's working memory. Because the entire history τ\tau is included in the context window at each step, the model at step nn can reason over everything it has learned in steps 11 through n−1n-1. This is not memory in the neural sense, but memory through explicit context: a written record the model can attend to. The limitation is that this memory is bounded by the context window, and we will examine that constraint closely in the limitations section.

The Action Space

The action space A\mathcal{A} defines what operations the agent can perform. Different tasks call for different action spaces, and defining the right set of tools is one of the most consequential design decisions in building a ReAct agent. Too few tools and the agent cannot gather the information it needs. Too many tools and the model struggles to select the right one or generates actions that match no tool in the registry.

Action spaces typically fall into four categories:

  • Information retrieval: Search queries, database lookups, document retrieval, API calls to external knowledge sources, or Wikipedia lookups. These expand what the agent knows beyond its training data.
  • Computation: Calculator operations, code execution environments, mathematical solvers, symbolic reasoning engines, or spreadsheet functions. These offload precise numerical or logical work to deterministic systems.
  • Interaction: Sending messages, filling forms, controlling external systems, or manipulating digital interfaces. These apply to agents that must affect the world, not just query it.
  • Termination: A finish or answer action that signals task completion and carries the final answer. Without this, the agent has no way to declare it is done.

For a question-answering agent, the minimal action space is often just search[query] and finish[answer]. For a coding agent, it might include write_file[path, content], run_code[code], read_file[path], and search_docs[query]. The key principle is that the action space should be complete enough to solve the intended class of problems, but minimal enough that the model can reliably select and format the correct action.

Connecting to the Policy Formulation

If you think of the agent as a policy in the reinforcement learning sense, each thought and action can be understood as the policy making a decision conditioned on the full trajectory seen so far:

ai∼πθ(a∣t1,a1,o1,…,ti−1,ai−1,oi−1,ti)a_i \sim \pi_\theta(a \mid t_1, a_1, o_1, \ldots, t_{i-1}, a_{i-1}, o_{i-1}, t_i)

where πθ\pi_\theta is the language model with parameters θ\theta. The thought tit_i is generated first, conditioning on all prior context, and then the action aia_i is generated conditioned on all prior context including the new thought. This two-stage generation is important: the thought gives the model space to reason before committing to an action, which empirically reduces the rate of nonsensical or irrelevant tool calls.

The Reasoning-Action Interleaving

We might naturally ask: why interleave reasoning with action rather than separating these phases into distinct stages? Consider the alternative approaches and their inherent limitations.

The first alternative, Reason-then-Act, generates a complete, complete plan before executing any actions, then follows that plan rigidly. This strategy fails when intermediate steps reveal unexpected information that invalidates subsequent parts of the plan. Imagine planning a multi-day road trip in detail, then discovering the first highway is closed. Your entire plan, including every turn and overnight stop, may need to be rewritten from scratch. For knowledge-intensive tasks, the plan made before any retrieval is essentially a plan made in ignorance, and the more complex the task, the more likely some intermediate discovery will invalidate the original plan.

The second alternative, Act-without-Reason, calls tools reactively based on pattern matching or simple heuristics without strategic planning. This approach often leads the agent into fruitless search loops, redundant queries, or missed connections between disparate pieces of information. Without explicit reasoning to guide tool selection, the agent lacks the strategic overview necessary for efficient problem-solving. It may retrieve the right facts in the wrong order, fail to recognize when a retrieved fact answers a key question, or repeat the same failed search query hoping for a different result.

The third approach, ReAct itself, tightly couples each action to explicit reasoning, and allows each observation to immediately inform the next reasoning step. The agent maintains flexibility while preserving strategic coherence. When the first search returns an unexpected result, the thought phase immediately processes that result, updates the agent's understanding of the problem, and formulates a better next step. The tight coupling creates a self-correcting system that approximates the scientific method by generating a hypothesis, testing it, observing the result, and revising the hypothesis.

This iterative refinement allows the model to recover from errors dynamically. If a search returns irrelevant results, the thought process acknowledges this mismatch and reformulates the query with better keywords. If a calculation yields an unexpected value, the model can verify its assumptions or check for unit conversion errors. If an API returns an error, the model can try an alternative tool or reformulate the request. None of these recoveries are possible in a purely plan-first or purely reactive system.

Out[3]:
Visualization
Grouped bar chart comparing three agent architectures across flexibility, error recovery, strategic planning, and efficiency.
Comparison of Reason-then-Act, Act-without-Reason, and ReAct architectures across four capability metrics. ReAct achieves superior flexibility and error recovery by interleaving reasoning with environmental feedback, while maintaining comparable strategic planning performance. Scores represent qualitative estimates based on the architectural properties of each approach.

The Thought-Action-Observation Loop

The ReAct loop operates as a finite state machine with three distinct phases per iteration. Understanding each phase's precise role helps clarify why the pattern succeeds where simpler, linear approaches fail to maintain coherence across complex tasks. But before we drill into each phase, it helps to see the entire loop from the outside.

At the start of each iteration, the model receives the full trajectory so far (system prompt, the original question, all prior thoughts and actions plus their observations) as a single long prompt. It then generates a thought, followed immediately by an action. The system intercepts the action, executes it using the appropriate tool, appends the observation, and feeds the whole updated prompt back to the model. This continues until the agent calls finish or hits the maximum iteration limit.

The simplicity is deliberate. There is no separate planning module, no memory retrieval system, no explicit belief state. Everything the agent knows is in the context window, written out as text. This makes ReAct straightforward to implement and debug, and it uses the model's strongest capability: reasoning over text.

Thought: Explicit Reasoning

The thought phase generates a natural language string that serves multiple cognitive functions simultaneously. A well-constructed thought does several things at once:

  • State tracking: The model explicitly records what it has determined so far. For example: "I have confirmed that Ernest Hemingway wrote 'The Old Man and the Sea.'"
  • Gap identification: The model identifies what remains unknown. For example: "I still need to find his birthplace."
  • Action justification: The model commits to a next step and explains why. For example: "I should search for 'Ernest Hemingway birthplace' to find this."
  • Error acknowledgment: When something goes wrong, the thought registers it and proposes a correction. For example: "The previous search returned no relevant results; I should try a more specific query."

By forcing the model to articulate its reasoning explicitly in natural language, we create an audit trail that human operators can inspect and verify. This explicitness reduces hallucinated actions because the model must justify each tool invocation before executing it. If a thought does not justify the action that follows it, the generated trajectory looks incoherent compared to the few-shot examples, which creates implicit pressure toward sensible reasoning.

The thought also acts as a persistent scratchpad that accumulates across iterations, maintaining context and intermediate conclusions that might otherwise be diluted by attention mechanisms across a very long context. When the model generates thought t3t_3, it can look back at t1t_1, o1o_1, t2t_2, o2o_2, and everything else to reconstruct its current understanding. Writing things down is not just useful for humans.

Prompt Engineering for Thoughts

ReAct prompting typically uses few-shot examples showing the exact format expected:

Thought: [reasoning about current state]
Action: [tool_name]([parameters])
Observation: [result from environment]

The model learns to generate thoughts that logically precede and justify actions, creating a coherent narrative flow. This structured prompting is needed for reliable parsing: if the model generates a thought that runs into the action line without a clear separator, the parser breaks. Good few-shot examples prevent this by showing the expected structure consistently.

Action: Structured Tool Use

The action phase translates abstract thought into concrete, executable commands. As we covered in Function Calling, modern LLMs can generate structured JSON, domain-specific languages, or simple command syntax for tool invocation. In ReAct, the action must be parseable and executable by the environment without ambiguity.

The specific action space is task-dependent and must be clearly defined in the system prompt. For question-answering tasks, the action space might include:

  • search[query]: Web search or document retrieval to find factual information
  • lookup[keyword]: Search within previously retrieved documents for specific terms
  • calculate[expression]: Mathematical evaluation for numerical reasoning
  • finish[answer]: Terminate the session with the final response

The syntax must be rigid enough for reliable parsing but intuitive enough for the model to generate correctly. A useful heuristic: if you would have to read the format specification twice to understand it, the model probably will too. Simple bracket notation (tool[params]) works well for most tasks. JSON objects are better when parameters are complex or multiple arguments are needed.

One subtlety that matters enormously in practice: the action format must be consistent across all few-shot examples. Models are exquisitely sensitive to format inconsistency. If one example uses search["query in quotes"] and another uses search[query without quotes], the model will sometimes generate one format and sometimes the other, causing parsing failures. Pick one format and stick to it throughout all examples and the system prompt.

Observation: Environmental Feedback

The observation phase closes the loop by giving factual grounding from external systems. Unlike the thought and action phases, which are generated by the LLM, observations come from external systems: search APIs, databases, calculators, sensors, or software interfaces. The model reads observations but does not generate them.

Observations serve as ground truth anchors that prevent the model from drifting into hallucination or maintaining false assumptions. If the model mistakenly thinks "The capital of France is London" based on some confused reasoning, a search observation returning "Paris" creates immediate corrective pressure on the next thought phase. The model must reconcile its internal belief with the external fact, and in practice, models do update their beliefs when observations contradict their assumptions.

This grounding is the key feature that distinguishes ReAct from pure reasoning approaches. By tethering the reasoning process to verifiable external facts, we create a system that can correct its own misconceptions and build conclusions on solid foundations rather than parametric memory alone. The model's parametric knowledge becomes a heuristic for guiding retrieval, not a substitute for it.

There is an important engineering detail here: the observation should be presented to the model as coming from the environment, not from the model itself. The prompt format typically wraps observations in a way that makes this clear:

Thought: I should search for the population of Paris. Action: search[population of Paris] Observation: The population of Paris is approximately 2.1 million. Thought: I now have the population. I can provide the final answer.

The model sees this entire string in its context window. The Observation: prefix signals that the following text is authoritative environmental data, not model-generated content. This framing makes a meaningful difference: models tend to trust observations more than their own prior assertions, which is exactly the behavior we want.

ReAct Prompting Strategies

Effective ReAct implementation requires careful prompt engineering that teaches the model the available tools and the pattern of reasoning that connects thoughts to actions and incorporates observations into an evolving understanding. The prompt is more than instructions: it is a behavioral specification that shapes every aspect of how the agent reasons.

Few-Shot Prompting

The standard approach provides two to three examples of successful ReAct trajectories before presenting the actual task. Each example should demonstrate multiple things:

  1. How to format thoughts (free-form but task-focused reasoning, neither too brief nor too verbose)
  2. How to format actions (strict syntax for reliable parsing)
  3. How to incorporate observations into subsequent reasoning steps
  4. How to recognize when you have enough information to call finish
  5. Ideally, at least one example where an initial approach fails and the agent recovers
Example 1: Question: What is the elevation of the mountain range where Mount Everest is located? Thought: I need to find which mountain range Mount Everest belongs to, then find the elevation of that range. Action: search[Mount Everest mountain range] Observation: Mount Everest is in the Himalayas. Thought: Now I need to find the elevation of the Himalayas. Action: search[elevation of Himalayas] Observation: The Himalayas have an average elevation of 20,000 feet (6,100 meters). Thought: I have found the answer. Action: finish[20,000 feet (6,100 meters)] Example 2: [...]

These examples serve as behavioral templates, showing the model how to decompose complex queries into manageable steps, how to recover from unexpected observations, and how to know when sufficient information has been gathered to provide a final answer. The quality of these examples has a large effect on agent performance. Poorly constructed examples that skip reasoning steps or use inconsistent formatting produce agents that make the same mistakes.

Choosing the Right Number of Examples

More examples are not always better. For highly capable models like GPT-4 or Claude 3, two or three carefully chosen examples often suffice. For smaller or older models, four or five examples may be needed to establish the pattern reliably. Beyond about six examples, the law of diminishing returns sets in, and the additional tokens are better spent on a longer system prompt with clearer tool descriptions.

The diversity of examples matters as much as their number. If all your examples show simple two-step retrievals, the agent will struggle with tasks requiring computation, backtracking, or multi-step dependency chains. Include at least one example that involves a calculation, one that requires multiple searches on different topics, and one where a search returns partial or ambiguous information.

Zero-Shot ReAct

While few-shot prompting works well for highly capable models, some situations call for zero-shot prompting: explicit instruction templates that describe the pattern without examples. This approach is useful when the action space is novel (and therefore no prior examples exist), when context window limits prevent including examples, or when the model is large enough to generalize from instructions alone.

You are an AI assistant that helps users by thinking step-by-step and using tools when necessary. For each step, you must provide: 1. Thought: Your reasoning about what to do next 2. Action: One of [search[query], calculate[expr], finish[answer]] 3. Observation: (provided by system after Action) Current Task: {task} Thought:

The zero-shot approach relies on the model's pre-trained understanding of task decomposition and tool use patterns, which makes it more flexible but potentially less reliable than few-shot examples for complex reasoning chains. In practice, the gap between few-shot and zero-shot narrows as model capability increases: a sufficiently large and capable model can infer the pattern from the instruction alone.

System Prompt Design

Beyond the examples themselves, the system prompt must convey several things clearly:

  • The available tools and exactly what each one does
  • The expected output format for each action
  • Any constraints (for example, maximum number of iterations, policies on when to use which tool)
  • The stopping condition (when should the agent call finish)

The tool descriptions deserve special attention. A vague tool description like "search: searches for things" is far less effective than "search[query]: Performs a keyword search and returns the top 3 most relevant results. Use this for factual questions, current events, or when you need information not in your training data." Detailed descriptions reduce the rate at which models misuse tools or fail to use them when appropriate.

Worked Example: Multi-Hop Question Answering

Let's trace through a concrete ReAct execution for the query: "What is the population of the city where the author of 'The Old Man and the Sea' was born?"

This question requires three distinct information retrieval steps: finding the author, then their birthplace, then the population of that city. Each step depends on the result of the previous one, which makes this a canonical multi-hop reasoning problem. A model without tool access would need to rely entirely on its parametric memory for all three facts, which introduces significant hallucination risk, especially for the specific population figure.

The dependency structure is:

answer←population(city)←birthplace(author)←author(novel)\text{answer} \leftarrow \text{population}(\text{city}) \leftarrow \text{birthplace}(\text{author}) \leftarrow \text{author}(\text{novel})

You cannot resolve the left side without resolving the right side first, which makes this unsolvable with a single tool call. The agent must execute the chain from right to left, resolving each dependency before proceeding.

Step 1: Initial Thought and Action

The model receives the question and generates an initial plan:

Thought: I need to identify the author of 'The Old Man and the Sea' first. I believe this is Ernest Hemingway, but I should verify with a search. Action: search[author of The Old Man and the Sea]

Notice how the thought explicitly states the immediate goal. The model acknowledges it already has a hypothesis (Hemingway) but chooses to verify rather than assume. This metacognitive pattern, being aware that parametric knowledge can be wrong and deferring to retrieval, is exactly what we want.

Step 2: Observation and Reformulation

The environment returns factual data:

Observation: Ernest Hemingway wrote The Old Man and the Sea in 1952.

The model now incorporates this observation and updates its plan:

Thought: The author is Ernest Hemingway. Now I need to find where he was born. Action: search[Ernest Hemingway birthplace]

The thought explicitly names the newly confirmed fact ("The author is Ernest Hemingway") and then states the next open question. This structure, state what you now know and then state what you still need, is a reliable template for well-structured thoughts.

Step 3: Second Information Retrieval

The environment provides the next fact:

Observation: Ernest Hemingway was born in Oak Park, Illinois.

The model updates its understanding and identifies the final missing piece:

Thought: Hemingway was born in Oak Park, Illinois. Now I need the population of Oak Park. Action: search[population of Oak Park Illinois]

Step 4: Final Retrieval and Termination

The observation provides the final required fact:

Observation: Oak Park, Illinois has a population of 54,583 as of the 2020 census.

The final thought demonstrates task completion recognition:

Thought: I now have all the information needed. Oak Park, Illinois is where Hemingway was born, and its population is 54,583. Action: finish[54,583]

Notice how each thought explicitly references the previous observation, creating a coherent chain of evidence that can be audited. The trajectory is a complete written record of how the agent reasoned from question to answer, which makes it possible to diagnose exactly where things went wrong when they do.

If the second search had returned "Ernest Hemingway was born in Key West, Florida" (which is wrong but plausible), the agent would have searched for Key West's population and returned an incorrect answer. This illustrates the retrieval quality dependence that we will discuss in the limitations section.

Code Implementation

Let's implement a minimal but complete ReAct system. We'll create a simple question-answering agent with access to a search tool and a calculator. This shows the core loop mechanics. The goal here is not to build a production agent, but to make the mechanics of the loop concrete and observable.

The implementation has five parts: the tools themselves, the agent class that manages the loop, a simulated LLM for demonstration, a function to run the agent, and the visualization of its trajectory.

Setting Up the Tools

First, we define the tools available to the agent. In production, these would be real API calls; here they use a hardcoded knowledge base for reproducibility:

In[4]:
Code
import ast
import operator
from typing import Callable, Dict


def search(query: str) -> str:
    """Simulated search tool with hardcoded knowledge for demo purposes."""
    knowledge_base = {
        "capital of france": "Paris",
        "population of paris": "2.1 million (2020 estimate)",
        "area of paris": "105 square kilometers",
        "eiffel tower height": "330 meters",
        "eiffel tower year built": "1889",
        "author old man sea": "Ernest Hemingway wrote 'The Old Man and the Sea' in 1952",
        "ernest hemingway birthplace": "Oak Park, Illinois",
        "population oak park illinois": "54,583",
        "author the old man and the sea": "Ernest Hemingway wrote The Old Man and the Sea in 1952",
    }
    query_lower = query.lower()
    for key, value in knowledge_base.items():
        if all(word in query_lower for word in key.split()):
            return value
    return f"No results found for: {query}"


# Safe arithmetic evaluator using ast module
_SAFE_OPS = {
    ast.Add: operator.add,
    ast.Sub: operator.sub,
    ast.Mult: operator.mul,
    ast.Div: operator.truediv,
    ast.Pow: operator.pow,
    ast.USub: operator.neg,
}


def _safe_eval(node):
    """Recursively evaluate a parsed arithmetic AST node."""
    if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
        return node.value
    elif isinstance(node, ast.BinOp) and type(node.op) in _SAFE_OPS:
        left = _safe_eval(node.left)
        right = _safe_eval(node.right)
        return _SAFE_OPS[type(node.op)](left, right)
    elif isinstance(node, ast.UnaryOp) and type(node.op) in _SAFE_OPS:
        return _SAFE_OPS[type(node.op)](_safe_eval(node.operand))
    else:
        raise ValueError(f"Unsupported expression type: {type(node)}")


def calculate(expression: str) -> str:
    """Safe arithmetic calculator that parses and evaluates mathematical expressions."""
    try:
        tree = ast.parse(expression.strip(), mode="eval")
        result = _safe_eval(tree.body)
        return str(result)
    except (ValueError, ZeroDivisionError, SyntaxError) as e:
        return f"Error: {str(e)}"


def finish(answer: str) -> str:
    """Terminal action that returns the final answer."""
    return f"FINAL_ANSWER: {answer}"


# Tool registry
TOOLS: Dict[str, Callable] = {
    "search": search,
    "calculate": calculate,
    "finish": finish,
}

The tool registry is a dictionary mapping tool names to callable functions. This design makes it easy to add or remove tools without changing the agent's core logic. In production, you might also store metadata about each tool (description, parameter schema, usage examples) alongside the callable, which is essentially what the OpenAI function-calling API does.

The calculator uses Python's ast module to parse the expression into a syntax tree and then evaluate only the numeric operations explicitly listed in _SAFE_OPS. This is safer than passing arbitrary strings to eval because it only accepts a restricted grammar of arithmetic operations.

The ReAct Agent Class

Now we implement the agent class that manages the thought-action-observation loop:

In[5]:
Code
import re
from typing import Any, Callable, Dict, List, Tuple


class ReActAgent:
    def __init__(self, tools: Dict[str, Callable], max_iterations: int = 10):
        self.tools = tools
        self.max_iterations = max_iterations
        self.trajectory: List[Dict[str, str]] = []

    def parse_action(self, text: str) -> Tuple[str, str]:
        """Parse 'Action: tool[params]' format from model output."""
        action_match = re.search(r"Action:\s*(\w+)\[(.*?)\]", text, re.DOTALL)
        if not action_match:
            return None, None
        tool_name = action_match.group(1)
        params = action_match.group(2)
        return tool_name, params

    def format_prompt(self, question: str) -> str:
        """Format the full prompt with examples and current trajectory."""
        few_shot_examples = """You are a helpful assistant that answers questions by thinking step-by-step and using tools.
Available tools: search[query], calculate[expression], finish[answer]

Example 1:
Question: What is 15 multiplied by 7?
Thought: I need to calculate 15 * 7.
Action: calculate[15 * 7]
Observation: 105
Thought: I have the answer.
Action: finish[105]

Example 2:
Question: What is the capital of France?
Thought: I need to search for the capital of France.
Action: search[capital of France]
Observation: Paris
Thought: I have found the capital.
Action: finish[Paris]

Now answer the following:"""

        # Build trajectory context
        context = ""
        for step in self.trajectory:
            context += f"\nThought: {step['thought']}"
            context += f"\nAction: {step['action']}"
            context += f"\nObservation: {step['observation']}"

        prompt = f"{few_shot_examples}\nQuestion: {question}{context}\nThought:"
        return prompt

    def step(self, question: str, llm_response: str) -> Tuple[str, bool]:
        """Execute one step of the ReAct loop.

        Returns: (observation, is_finished)
        """
        # Parse thought (everything before Action)
        thought_match = re.match(r"(.*?)Action:", llm_response, re.DOTALL)
        if thought_match:
            thought = thought_match.group(1).strip()
            if thought.startswith("Thought:"):
                thought = thought[8:].strip()
        else:
            thought = llm_response.strip()

        # Parse action
        tool_name, params = self.parse_action(llm_response)

        if not tool_name:
            return (
                "Error: Could not parse action. Please use format 'Action: tool[params]'",
                False,
            )

        if tool_name == "finish":
            return self.tools[tool_name](params), True

        if tool_name not in self.tools:
            return f"Error: Unknown tool '{tool_name}'", False

        # Execute tool
        observation = self.tools[tool_name](params)

        # Record in trajectory
        self.trajectory.append(
            {
                "thought": thought,
                "action": f"{tool_name}[{params}]",
                "observation": observation,
            }
        )

        return observation, False

    def run(
        self, question: str, llm_simulator: Callable[[str], str]
    ) -> Dict[str, Any]:
        """Run the full ReAct loop.

        Args:
            question: The question to answer
            llm_simulator: Function that takes prompt and returns LLM response
        """
        self.trajectory = []

        for i in range(self.max_iterations):
            prompt = self.format_prompt(question)
            response = llm_simulator(prompt)

            # Add the generated thought/action to prompt for parsing
            full_response = "Thought: " + response

            observation, is_finished = self.step(question, full_response)

            if is_finished:
                return {
                    "success": True,
                    "answer": observation.replace("FINAL_ANSWER: ", ""),
                    "trajectory": self.trajectory,
                    "steps": i + 1,
                }

        return {
            "success": False,
            "answer": None,
            "trajectory": self.trajectory,
            "steps": self.max_iterations,
            "reason": "Max iterations reached",
        }

The parse_action method uses a regular expression to extract the tool name and parameters from the model's output. The regex (\w+)\[(.*?)\] matches patterns like search[query text] or calculate[15 * 7]. This is deliberately simple. Production systems often need more reliable parsers that handle multi-line parameters, special characters, or JSON-formatted arguments.

The format_prompt method constructs the full context window contents at each step. Notice that it always starts from the full few-shot examples and then appends the complete trajectory. This means the model sees the same examples every time, which is important for maintaining format consistency. The trajectory context appended at the end grows with each iteration, which is the source of the linear token growth discussed in the limitations section.

The run method implements the outer loop with a maximum iteration guard. The guard is needed: without it, a confused or stuck agent would consume API credits indefinitely. Hitting the maximum iteration limit without a finish action is a detectable failure mode that the calling system can handle, for example by returning a "could not answer" response or escalating to a human.

Simulating LLM Responses

For demonstration purposes, we simulate the LLM responses to show exactly what a correct trajectory looks like:

In[6]:
Code
def simulate_llm(
    prompt: str, step_responses: List[str], step_counter: List[int]
) -> str:
    """Simulate LLM responses for demo purposes."""
    current_step = step_counter[0]
    step_counter[0] += 1
    if current_step < len(step_responses):
        return step_responses[current_step]
    return "finish[Unable to complete]"


# Define the response sequence for our multi-hop question
responses = [
    # Step 1: Search for author
    "I need to find who wrote 'The Old Man and the Sea'.\nAction: search[author The Old Man and the Sea]",
    # Step 2: Search for birthplace (after seeing observation)
    "The author is Ernest Hemingway. Now I need to find where he was born.\nAction: search[Ernest Hemingway birthplace]",
    # Step 3: Search for population
    "He was born in Oak Park, Illinois. Now I need the population of Oak Park.\nAction: search[population Oak Park Illinois]",
    # Step 4: Finish
    "I have found that Oak Park, Illinois has a population of 54,583. I can now provide the final answer.\nAction: finish[54,583]",
]

Running the Agent

Now let's run the agent and observe its trajectory:

In[7]:
Code
# Initialize agent and run
agent = ReActAgent(TOOLS, max_iterations=10)
step_counter = [0]  # Use list to allow mutation in closure

result = agent.run(
    "What is the population of the city where the author of 'The Old Man and the Sea' was born?",
    lambda p: simulate_llm(p, responses, step_counter),
)
Out[8]:
Console
Success: True
Answer: 54,583
Steps taken: 4

Full Trajectory:

Step 1:
  Thought: I need to find who wrote 'The Old Man and the Sea'.
  Action: search[author The Old Man and the Sea]
  Observation: Ernest Hemingway wrote 'The Old Man and the Sea' in 1952

Step 2:
  Thought: The author is Ernest Hemingway. Now I need to find where he was born.
  Action: search[Ernest Hemingway birthplace]
  Observation: Oak Park, Illinois

Step 3:
  Thought: He was born in Oak Park, Illinois. Now I need the population of Oak Park.
  Action: search[population Oak Park Illinois]
  Observation: 54,583

The agent successfully completes the multi-hop reasoning task in four steps. This shows how the ReAct pattern chains together disparate pieces of information into a coherent solution. Each thought explicitly references the previous observation. This creates an auditable reasoning chain that traces the path from question to answer.

Visualizing the Trajectory

The trajectory diagram below makes the iterative structure concrete. Each column represents a phase (Thought, Action, Observation), and each row represents one iteration:

Out[9]:
Visualization
Flowchart showing three rows of thought, action, observation boxes connected by arrows, leading to a final answer box.
Three iterations of the ReAct thought-action-observation loop for the multi-hop question about Hemingway's birthplace. Each row shows one iteration, with observations feeding directly into the next thought. The dashed lines on the left indicate that cumulative context from all prior iterations is available at each step.

The diagram makes one important feature visible that is easy to overlook in the code: the cumulative context arrows on the left side. Each new thought has access to the immediate previous observation and to the entire trajectory of prior thoughts and actions plus their observations. This accumulation is what allows the agent to refer back to earlier findings even several steps later.

Context Length and Token Growth

One of the most practically important aspects of ReAct is that context grows with each iteration. The prompt at step nn contains:

∣Pn∣=∣Psystem∣+∑i=1n−1(∣ti∣+∣ai∣+∣oi∣)|P_n| = |P_{\text{system}}| + \sum_{i=1}^{n-1} (|t_i| + |a_i| + |o_i|)

where ∣Psystem∣|P_{\text{system}}| is the length of the system prompt (including few-shot examples), and each iteration adds tokens for a thought and action plus the resulting observation. For typical ReAct setups with roughly 350 tokens for the system prompt and roughly 280 tokens per iteration (about 80 for the thought, 40 for the action, and 160 for the observation), the context length at step nn is approximately:

∣Pn∣≈350+280⋅(n−1)|P_n| \approx 350 + 280 \cdot (n - 1)

This linear growth has real consequences. Modern models support context windows of 8k, 32k, 128k, or even 1M tokens, but most API calls are billed per token, and latency increases with context length. An agent running 20 iterations with verbose observations could easily consume 10,000 tokens per LLM call, and it makes 20 such calls, for a total of 200,000 tokens for a single user question. At typical API pricing, this adds up quickly.

Out[10]:
Visualization
Line chart showing cumulative token count rising linearly across 12 ReAct iterations, with a dashed red line at 4,096 tokens.
Cumulative context length growth across ReAct iterations, assuming 350 tokens for the system prompt and 280 tokens per iteration. The red dashed line marks a 4,096-token context window limit, showing that agents running more than about 13 iterations will exceed this boundary. Modern models with larger context windows shift this limit right but do not eliminate the linear growth problem.

The chart shows that even with a modest 4k context window, the agent starts hitting limits around iteration 13. With modern 128k context windows, this is less immediately problematic, but the cost and latency implications remain. Several strategies address this:

  • Trajectory truncation: Drop the oldest few iterations when the context approaches a threshold. This loses some historical context but keeps the agent running.
  • Observation summarization: Instead of storing the full observation, store a one-sentence summary. This is lossy but dramatically reduces token usage.
  • Selective retention: Keep all thoughts and actions, but only keep observations when they contributed new information. Redundant or failed observations are discarded.
  • External memory: For very long sessions, store the trajectory in a database and retrieve only the most relevant past steps using semantic similarity search.

The right strategy depends on the task. For short research queries (5-10 steps), the raw trajectory works fine. For complex research or coding tasks that might run dozens of iterations, some form of compression becomes necessary.

Key Agent Parameters

The key parameters for the ReActAgent are:

  • tools: A dictionary mapping tool names to callable functions. These define the action space available to the agent and determine what capabilities the system can use to solve problems. Choosing the right tools is the most important architectural decision in building a ReAct agent.
  • max_iterations: The maximum number of thought-action-observation cycles allowed before the agent terminates. This safety parameter prevents infinite loops if the model cannot reach a conclusion or enters a repetitive cycle. Ten iterations is a reasonable default for most question-answering tasks, but complex coding or research tasks may need 20-50.

Integration with Real LLMs

In production systems, the llm_simulator function would be replaced with actual API calls to models like GPT-4, Claude, or open-source alternatives such as Llama or Mistral. Modern models trained with instruction tuning (covered in Part XXXVI) and RLHF (covered in Part XXXVII) typically require minimal few-shot examples to adopt the ReAct pattern effectively, though the specific formatting may need adjustment based on the model's training.

The following example integrates the pattern with the OpenAI API:

In[19]:
Code
import openai


def llm_call(prompt: str, model: str = "gpt-4") -> str:
    """Call OpenAI API with ReAct-formatted prompt."""
    client = openai.OpenAI()
    response = client.chat.completions.create(
        model=model,
        messages=[
            {
                "role": "system",
                "content": "You are a helpful assistant that uses tools by thinking step-by-step.",
            },
            {"role": "user", "content": prompt},
        ],
        temperature=0,  # Low temperature for deterministic reasoning
        stop=["\nObservation:"],  # Stop before generating observation
    )
    return response.choices[0].message.content

The stop parameter is important: it prevents the model from hallucinating observations. By specifying that generation should halt before creating an observation line, we ensure that observations come exclusively from actual tool execution rather than the model's imagination. This maintains the integrity of the grounding mechanism that makes ReAct reliable. Without this stop sequence, a sufficiently confident model might generate both the action and its supposed result in a single pass, completely bypassing the external tool call.

The temperature setting also matters. Low temperatures (0.0-0.2) produce more deterministic reasoning that follows the expected format reliably. Higher temperatures can cause format drift, where the model starts omitting the "Thought:" prefix or using slightly different action syntax. For most ReAct applications, temperature 0 is the right choice.

Structured Output Approach

An alternative to parsing free-form text is to use structured outputs or function calling to extract actions. Instead of asking the model to generate Action: search[query] in text and parsing it with regex, you ask the model to generate a JSON object with a defined schema:

In[21]:
Code
import openai

tools_spec = [
    {
        "type": "function",
        "function": {
            "name": "search",
            "description": "Search for information on a topic",
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {
                        "type": "string",
                        "description": "The search query",
                    }
                },
                "required": ["query"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "calculate",
            "description": "Evaluate a mathematical expression",
            "parameters": {
                "type": "object",
                "properties": {
                    "expression": {
                        "type": "string",
                        "description": "Mathematical expression to evaluate",
                    }
                },
                "required": ["expression"],
            },
        },
    },
]


def llm_call_structured(messages: list, tools: list) -> dict:
    """Call OpenAI API with structured tool definitions."""
    client = openai.OpenAI()
    response = client.chat.completions.create(
        model="gpt-4",
        messages=messages,
        tools=tools,
        tool_choice="auto",
    )
    return response.choices[0].message

The structured approach eliminates the need for regex parsing and reduces format errors significantly. The tradeoff is that the model no longer generates explicit thought text before each action, which reduces the interpretability of its reasoning. Some implementations address this by asking the model to include its reasoning in the tool call arguments or in a separate "thinking" step before calling tools.

Error Handling and Reliability

Real-world ReAct implementations must handle several failure modes that can disrupt the reasoning loop or lead to incorrect conclusions. These are not edge cases: they occur frequently in production, and an agent that cannot handle them gracefully will fail on a substantial fraction of queries.

Invalid Action Parsing

When the model generates malformed actions, for example search[query with a missing closing bracket, the system should return a helpful error observation that prompts the model to correct the syntax in the next iteration. The error message should be explicit and directive:

Observation: Error: Could not parse action. Expected format: search[your query here]. Please retry with correct bracket syntax.

This phrasing tells the model exactly what went wrong and how to fix it. A vague error like "invalid input" often leads the model to try increasingly divergent formats before eventually giving up. Specific error messages lead to faster recovery.

Tool Execution Failures

External APIs timeout, databases return connection errors, and calculations occasionally overflow or divide by zero. These exceptions should be caught gracefully and returned as observations so the model can adapt its strategy. The error handling should distinguish between transient failures (network timeout, rate limit) and permanent ones (invalid query syntax, resource not found). For transient failures, the observation can suggest a retry. For permanent failures, it should suggest an alternative approach.

Reasoning Loops

The model might enter cycles, repeatedly searching for the same information or alternating between two equivalent states without making progress. You can detect this by tracking the action history: if the same action appears twice in the trajectory, note the repetition in the next prompt.

A simple loop-detection mechanism:

In[11]:
Code
def detect_loop(trajectory: List[Dict[str, str]], window: int = 3) -> bool:
    """Check if the agent is repeating actions within the last `window` steps."""
    if len(trajectory) < window:
        return False
    recent_actions = [step["action"] for step in trajectory[-window:]]
    return len(set(recent_actions)) < len(recent_actions)

When a loop is detected, the observation for the repeated action can include a nudge: "This search was already performed in step 2. Consider trying a different search term or checking whether you already have enough information to answer."

Hallucinated Facts in Thoughts

While observations provide grounding for the final answer, the thoughts themselves can contain reasoning errors or false assumptions. A model might write "Hemingway lived in Paris for many years, so his birthplace is likely France" as a thought, then search based on that incorrect assumption. The observation corrects it, but only if the model pays attention.

Advanced systems address this by using self-consistency sampling: running the ReAct loop multiple times with different random seeds and checking whether the trajectories agree on intermediate facts. When trajectories diverge on a sub-question, that sub-question gets flagged for extra verification. This is expensive but useful for high-stakes applications.

Evaluating ReAct Agents

Evaluating ReAct agents is more complex than evaluating a single forward pass, because there are multiple dimensions of quality to assess. The final answer is the most obvious metric, but it is not the only thing that matters.

Answer Accuracy

The most direct metric is whether the final answer is correct. For factual question-answering benchmarks like HotpotQA (multi-hop questions) and FEVER (fact verification), ReAct was evaluated against pure chain-of-thought baselines. The original Yao et al. paper showed that ReAct improved accuracy by 3-8 percentage points on HotpotQA and substantially more on tasks requiring current information that was absent from the model's training data.

Trajectory Efficiency

A correct answer reached in 3 steps is better than the same correct answer reached in 8 steps, because fewer steps means fewer tokens, lower latency, and lower cost. Efficiency is measured as the number of tool calls required to reach the correct answer, averaged across a test set. An agent that makes redundant or irrelevant searches has low trajectory efficiency even if it eventually reaches the right answer.

Error Recovery Rate

When the agent encounters an obstacle (a failed search, an ambiguous result, a calculation error), does it recover? You can measure this by deliberately introducing failures in the tool responses and observing whether the agent adapts. A well-designed agent should recover from a majority of single-step failures. Recovery from multi-step errors (where an early incorrect observation cascades into wrong subsequent actions) is harder and depends heavily on the model's ability to revisit assumptions.

Faithfulness

The thought trajectory should accurately reflect the reasoning behind each action. An unfaithful trajectory might show a thought saying "I need to search for Hemingway's birthplace" while generating a completely unrelated action. Faithfulness can be measured by checking whether the action logically follows from the thought. While this is partially qualitative, it is important for interpretability and for building user trust.

Out[12]:
Visualization
Radar chart with four axes showing evaluation scores for small and large ReAct agents across accuracy, efficiency, error recovery, and faithfulness.
Radar chart showing four evaluation dimensions for ReAct agents at two model scales. Larger models score higher across all dimensions, with the most pronounced gap in error recovery and faithfulness. Answer accuracy and trajectory efficiency also improve substantially with scale, which reflects the benefit of stronger reasoning and world knowledge.

The radar chart illustrates a consistent pattern: larger models score higher across all four evaluation dimensions, with the greatest gains in error recovery and faithfulness. This makes intuitive sense. Error recovery requires recognizing that a previous step was wrong and formulating a correction, which demands more sophisticated reasoning than simply following a familiar pattern. Faithfulness requires the thought and action to be internally consistent, which benefits from the model's ability to maintain coherent multi-step reasoning.

Limitations and Impact

The ReAct pattern has significant limitations that constrain its use in production environments, and understanding them is needed for building reliable agents rather than ones that work in demos but fail in production.

Token Efficiency

The most immediate constraint is token cost. Each iteration appends the thought and action plus the resulting observation to the prompt, causing linear growth in context length. For complex tasks requiring dozens of steps, the context window fills rapidly, increasing latency and cost while potentially pushing important early reasoning out of the model's effective attention span. In practice, models attending to a 100k-token context may not effectively utilize information from the first few thousand tokens, effectively losing early reasoning steps even when they technically fit in the window.

A concrete example: if a user asks an agent to research a company for an investment report, the agent might need 30-50 tool calls to gather complete information. Each call costs latency and money. If each iteration averages 500 tokens and you run 40 iterations, that is 20,000 tokens per query just in trajectory context, plus the 40 individual API calls. At typical API pricing, you are spending a meaningful amount per research query just on context, plus API call overhead.

Error Propagation

Error propagation presents another necessary challenge. In our multi-hop example, if the first search incorrectly identified the author as "John Steinbeck" due to ambiguous query results, all subsequent steps would compound this error. Unlike human researchers who might question suspicious premises or cross-verify facts, language models tend to accept previous trajectory steps as ground truth. This commitment bias means early mistakes often prove fatal to the final answer, and the system lacks built-in mechanisms for backtracking or revisiting earlier assumptions.

The figure below illustrates how error probability compounds across hops. Even a modest per-step error rate of 5% leads to a 40% probability of at least one error by step ten. Chain-of-thought without grounding, which uses the model's parametric memory exclusively, compounds errors faster because it lacks the corrective signal that observations provide.

Out[13]:
Visualization
Line chart comparing cumulative error probability versus reasoning depth for Chain-of-Thought and ReAct, with ReAct showing a lower error curve.
Cumulative error probability across reasoning depth for Chain-of-Thought (10% per-step error rate) versus ReAct (5% per-step error rate). Both curves grow toward certainty at sufficient depth, but ReAct's lower per-step error rate from observation grounding provides meaningful protection for chains of up to 10 hops. Beyond that depth, both approaches require additional verification strategies.

Action Space Rigidity

The action space limitation also constrains ReAct's flexibility. The model can only invoke tools explicitly provided in the prompt. It cannot create new tools on the fly, combine existing ones in novel ways not anticipated by the designer, or recognize when a problem requires a capability not in its repertoire. This brittleness contrasts with human problem-solving, where we dynamically adapt our tool use or seek new instruments when existing ones prove inadequate.

This limitation has a practical consequence for deployment: the tool set must be carefully designed upfront for the expected class of tasks. An agent given only search and calculate cannot write or execute code, no matter how clearly the task demands it. Getting the tool set right requires understanding the nominal task and all the edge cases and unexpected sub-problems that real queries trigger.

Retrieval Quality Dependence

ReAct's grounding mechanism only works as well as the tools it grounds against. If the search tool returns poor results, the agent builds its reasoning on a shaky foundation. In academic benchmarks, the knowledge base is often curated and reliable. In production, search results can be noisy, outdated, incomplete, or misleading. An agent that trusts its observations absolutely becomes vulnerable to anyone who can manipulate the data sources it queries, which is a real security concern for agents that search the open web.

Planning Depth

ReAct is fundamentally reactive: each thought responds to the most recent observation, and the agent has no mechanism for lookahead or planning multiple steps ahead. For problems with complex dependency structures, where completing step 5 reveals that step 3 was unnecessary but step 7 depends on a sub-question that should have been asked at step 1, a purely reactive approach is inefficient. More sophisticated agent architectures, such as Tree of Thoughts or MCTS-based planners, address this by maintaining multiple hypothesis trees and exploring them in parallel, but these come with additional complexity and cost.

Impact and Legacy

Despite these limitations, ReAct represents a basic advance in language model capabilities and a important step toward autonomous agent systems. By externalizing working memory into a traceable trajectory, it enables complex reasoning that exceeds the model's single-forward-pass capacity. By grounding each reasoning step in environmental feedback, it dramatically reduces the rate of confident hallucination on factual queries.

The pattern has become foundational for production agent systems. Browser-use agents that browse websites, coding assistants that write and execute code iteratively, research tools that synthesize information from multiple sources, and customer service bots that look up account information all implement some form of the ReAct loop. The architecture is simple enough to implement in an afternoon, interpretable enough to debug without specialized tools, and reliable enough to handle a wide range of real-world tasks.

ReAct also influenced the design of more sophisticated agent frameworks that followed. LangChain's agent abstraction, LlamaIndex's query engines, AutoGPT's planning loop, and OpenAI's Assistants API all build on the same basic insight: language models reason best when they can interleave thinking with information retrieval rather than doing everything in a single forward pass. As we explore in upcoming chapters on agent memory and multi-agent coordination, ReAct provides the basic iterative loop within which more sophisticated strategies for planning and reflection during learning can operate.

Summary

ReAct (Reasoning + Acting) interleaves chain-of-thought reasoning with tool execution in an iterative loop. The pattern consists of three phases per iteration: Thought, which contains explicit natural language reasoning about the current state and next steps; Action, which is a structured tool invocation with specific parameters; and Observation, which is environmental feedback that grounds reasoning in facts from external systems. The complete sequence of thought-action-observation triples forms the agent's trajectory, which is its working memory and audit trail.

The formal trajectory τ=(t1,a1,o1,t2,a2,o2,…)\tau = (t_1, a_1, o_1, t_2, a_2, o_2, \ldots) captures the agent's full reasoning history. At each step, the model conditions on this trajectory to generate the next thought and action. The key insight is that observations provide factual grounding that corrects the model's internal assumptions, reducing hallucination rates substantially compared to pure chain-of-thought prompting.

Key implementation considerations include:

  • Few-shot prompting to teach the expected format through concrete examples. Example quality and diversity matter more than count.
  • Careful action space design to match the tools available to the class of problems the agent will face.
  • Reliable action parsing with helpful error messages that guide the model toward format compliance.
  • Stop sequences when using real LLMs to prevent the model from hallucinating observations.
  • Loop detection and graceful termination when the agent is not making progress.
  • Context management strategies (truncation, summarization, selective retention) for tasks requiring many iterations.

The limitations are real: linear token growth, error propagation in multi-hop chains, action space rigidity, and retrieval quality dependence all require thoughtful mitigation in production systems. But the pattern is simple and interpretable while remaining effective, which has made it the dominant approach to building tool-using LLM agents. It turns a static language model into an interactive agent able to multi-step reasoning, dynamic information gathering, and structured problem-solving. This provides the reasoning foundation for modern autonomous AI systems.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the ReAct pattern, thought-action-observation loops, and agent architectures.

ReAct Pattern Fundamentals

Question 1 of 80 of 8 completed
What does the acronym ReAct represent in the context of LLM agent architectures?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026reactreasoning, author = {Michael Brenndoerfer}, title = {ReAct: Reasoning and Action for LLM Agents}, year = {2026}, url = {https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). ReAct: Reasoning and Action for LLM Agents. Retrieved from https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents
MLAAcademic
Michael Brenndoerfer. "ReAct: Reasoning and Action for LLM Agents." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents>.
CHICAGOAcademic
Michael Brenndoerfer. "ReAct: Reasoning and Action for LLM Agents." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents.
HARVARDAcademic
Michael Brenndoerfer (2026) 'ReAct: Reasoning and Action for LLM Agents'. Available at: https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). ReAct: Reasoning and Action for LLM Agents. https://mbrenndoerfer.com/writing/react-pattern-llm-reasoning-action-agents

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.