Multi-Agent Systems: Coordination and Communication

Michael BrenndoerferJanuary 3, 202655 min read

Part of Language AI Handbook

Explains how multi-agent AI systems coordinate through topologies, communication protocols, role assignment.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Multi-Agent Systems

Single agents are powerful, but many real-world tasks require collaboration. A lone agent running a ReAct loop can browse the web, write code, and call APIs, but it faces a fundamental bottleneck: it works sequentially, can hold only so much context, and accumulates errors over long horizons. Multi-agent systems address this by distributing work across multiple specialized agents that communicate and coordinate while checking each other's outputs.

The intuition is familiar from human organizations. A software team has engineers, testers, product managers, and designers. Each role requires different expertise, and their parallel work is more productive than any one person handling everything. Multi-agent AI systems mirror this structure: an orchestrator decomposes tasks, assigns subtasks to specialist subagents, collects results, and synthesizes a final output. The analogy extends further: just as teams develop norms for communicating status updates, escalating blockers, and resolving disagreements, multi-agent systems need explicit protocols for exactly those situations.

This chapter covers the architectural patterns, communication protocols, and coordination mechanisms that make multi-agent systems work. Building on previous chapters about agent architectures and memory as well as planning, we now examine how multiple agents collaborate. We will look at how agents communicate through shared environments or direct message passing, how roles get assigned, how conflict gets resolved, and how to measure whether a multi-agent system performs better than a single agent. The final chapter in this part, Agent Safety, will address the additional risks that emerge when multiple autonomous agents interact.

Why Multiple Agents?

Before diving into mechanisms, it is worth asking: when does a multi-agent setup help?

Three scenarios motivate the design:

  • Parallelism: Long tasks can be decomposed into independent subtasks and executed simultaneously. Researching market trends for five industries at once is faster with five agents than with one running five sequential searches.
  • Specialization: Different subtasks benefit from different prompts, tools, or even different models. A coding agent optimized for Python generation performs better at code than a general-purpose reasoning agent, while a document summarization agent excels at condensing long PDFs.
  • Independent verification: Critical outputs benefit from a second opinion. Having an independent agent review generated code for correctness catches errors that the generating agent might miss, since the reviewer brings fresh context.

The cost is coordination overhead. Agents need to pass information between each other, wait for dependencies to resolve, and handle failures in other parts of the system. Whether multi-agent wins over single-agent depends on whether the parallelism and specialization benefits outweigh this coordination cost.

It is tempting to assume more agents always helps, but empirical results show a conditional pattern. Wang et al. (2023) found that adding agents beyond a certain threshold can hurt performance on structured reasoning tasks, because coordination noise accumulates faster than individual agent improvement. The sweet spot depends heavily on task decomposability. A task that truly breaks into five independent parallel subtasks benefits from five agents; a task that is fundamentally sequential does not.

Understanding this tradeoff requires thinking carefully about what makes a task decomposable. A subtask is truly independent if it does not share output with another subtask in progress, if its requirements can be specified completely before it starts, and if its result can be validated in isolation. Tasks that appear independent can turn out to be deeply entangled once agents begin executing them. A research task about "government subsidies" and another about "EV adoption rates" seem independent until you discover that the most authoritative source on adoption rates is a government subsidy policy report, creating an implicit dependency. The orchestrator's ability to anticipate and handle such entanglement is a key determinant of multi-agent system quality.

The history of multi-agent systems in AI stretches back decades before LLMs. Classical multi-agent research in the 1990s and early 2000s, exemplified by work on robotic swarms, distributed problem solving, and game-theoretic agent interactions, established foundational concepts that modern LLM-based systems rediscover and extend. The shift from hand-coded rule-based agents to LLM-based agents changes the implementation considerably but does not change the fundamental coordination challenges. Whether agents are robotic components communicating over a network or LLMs calling APIs over HTTP, the same questions arise: how do they share state, how do they divide work, and how do they recover from partial failures?

One aspect of this history worth understanding is why earlier multi-agent systems did not achieve the broad applicability that modern LLM-based systems promise. Classical agents were highly domain-specific: a multi-agent system for scheduling warehouse robots worked within a narrow, well-defined problem space with a fixed set of actions, reliable sensors, and deterministic outcomes. Building such a system required extensive domain engineering for every new application. LLMs change this equation by providing agents with broad general-purpose reasoning and language understanding, allowing the same underlying model to be specialized for dramatically different roles through prompting alone. This is what makes modern multi-agent systems feasible at scale in a way that classical approaches were not.

Topologies and Coordination Patterns

Multi-agent systems organize in several broad topologies, each with different tradeoffs in flexibility, fault tolerance, and complexity. Choosing the right topology is not a one-time architectural decision made at system design time. As systems scale and task requirements evolve, topologies often need to be adjusted. Starting with the simplest topology that fits the task and adding complexity only when the task demands it is usually the right approach.

Hierarchical Orchestration

The most common pattern is a two-level hierarchy: an orchestrator agent at the top decomposes the user's goal into subtasks and delegates each to a subagent. Subagents execute their assigned tasks and return results. The orchestrator synthesizes these results into a final response.

The orchestrator behaves like a project manager who writes work orders. It does not execute subtasks directly but rather decides what needs to be done, which agent should do it, and how outputs should be combined. The subagents behave like individual contributors: they receive a specific instruction, use their available tools to complete it, and report back.

This pattern scales naturally. An orchestrator can spawn additional subagents dynamically based on what it discovers during the task. If a research task reveals an unexpected subtopic requiring deep investigation, the orchestrator can create a specialized agent on the fly. Hierarchical orchestration also makes the system's behavior interpretable: every action traces back to an orchestrator decision, making debugging significantly easier than in peer-to-peer topologies.

The main vulnerability of hierarchical orchestration is that the orchestrator becomes a bottleneck. If it makes poor decomposition decisions (e.g., assigns too many tasks to one agent, fails to recognize a dependency between two subtasks), those mistakes propagate through all downstream work. The orchestrator's decomposition quality is therefore the single biggest factor in system performance. Much recent research on multi-agent systems focuses specifically on improving orchestrator prompts and planning strategies, which connects directly to the planning techniques discussed in the previous chapter.

A second vulnerability is that the orchestrator itself has limited context. When a task produces a large result, the orchestrator needs to read and process that result before deciding what to do next. With many agents producing large outputs simultaneously, the orchestrator's context window can fill up quickly. This is why good orchestrator design involves having subagents produce structured, concise summaries rather than raw outputs, with full results stored in a shared workspace that the orchestrator can selectively reference.

Peer-to-Peer Collaboration

In a peer-to-peer topology, agents communicate directly without a central coordinator. Each agent has visibility into the shared state, can post messages, and can respond to messages from other agents. Agents propose actions, vote on approaches, or critique each other's outputs before a decision is finalized.

This topology is more resilient to single points of failure. If the orchestrator in a hierarchical system fails, the whole system stalls. In a peer-to-peer system, agents can continue operating and elect a new coordinator if needed. The resilience comes at a price: the system must have explicit protocols for conflict resolution, since there is no central authority to arbitrate disputes.

Peer-to-peer coordination also tends to produce richer deliberation. When agents can argue with each other and revise their positions based on incoming arguments, the quality of the final answer often exceeds what any single agent would produce. The CAMEL framework (Li et al., 2023) and the multi-agent debate work by Du et al. (2023) both demonstrate this empirically: agents that critique and counter-critique each other converge to more accurate answers than agents that work in isolation.

The downside is complexity. Without a central authority, agents may deadlock (each waiting for another to proceed), diverge (each pursuing conflicting strategies), or generate inconsistent outputs that no one reconciles. Implementing reliable peer-to-peer systems requires careful design of termination conditions and explicit protocols for handling disagreements that do not resolve through deliberation alone.

Pipeline Patterns

Some tasks naturally decompose into sequential stages where each stage's output is the next stage's input. A document analysis pipeline might have a reader agent extract key claims, a researcher agent verify those claims, and a writer agent compose a final report. Each agent passes its output to the next as structured data.

Pipeline systems are easier to reason about than peer-to-peer systems because information flows in one direction and dependencies are explicit. They sacrifice the parallelism benefits of hierarchical orchestration for this simplicity. The tradeoff is often worthwhile when the task truly is sequential and the emphasis is on correctness of each transformation step rather than on speed.

A subtle failure mode in pipelines is silent degradation. When each agent passes its output downstream without validation, early-stage errors propagate silently. The writer agent may faithfully produce a well-structured report based on researcher findings that are factually wrong, and nothing in the pipeline will detect this. Adding validation steps between pipeline stages, even lightweight ones, significantly improves reliability.

Pipelines are particularly effective when each stage has a well-defined transformation with clear inputs and outputs. The challenge arises when the boundaries between stages are fuzzy: if a researcher agent discovers that properly answering the research question requires significant analysis, and the analysis agent discovers that the initial research was too narrow, the pipeline becomes brittle because neither stage has a mechanism for requesting the other to revise its work. Pipelines therefore work best when the task structure is stable and well-understood in advance, rather than exploratory.

Mixture Patterns

Real systems often combine these topologies. An orchestrator might use both hierarchical delegation for independent subtasks and a pipeline for dependent ones, while allowing peer communication for agents working on overlapping concerns. The choice of topology should follow the natural dependency structure of the task.

The concept of adaptive topology takes this further: the system selects its own coordination structure based on task characteristics. Some implementations (e.g., AutoGen from Microsoft Research, 2023) support this by letting agents decide at runtime whether to request help, spawn a subagent, or handle a task themselves. This flexibility improves performance on diverse task distributions but makes the system's behavior harder to predict and audit.

Communication Protocols

Agents in a multi-agent system need to exchange information reliably. The communication protocol defines the format and channel together with the semantics of these exchanges. Getting this right matters more than it might seem. Ambiguous messages cause agents to misinterpret task requirements. Missing context causes agents to produce outputs that cannot be composed with other agents' outputs. Inadequate error signaling leaves the orchestrator unable to respond appropriately to failures.

Shared State vs. Message Passing

Two fundamental communication paradigms exist:

Shared state (also called a shared workspace or blackboard): All agents read from and write to a common data store. An agent processing a research subtask might write its findings to a shared document that the orchestrator and other agents can read. This is simple to implement and avoids serializing messages, but requires careful coordination to avoid race conditions where two agents modify the same state simultaneously.

Message passing: Agents send structured messages directly to each other (or through a message bus). Messages identify the sender and recipient alongside the payload. This is more explicit about information flow and easier to trace, but requires agents to poll for incoming messages or react to asynchronous notifications.

Many practical systems blend both: a shared workspace holds durable state (task assignments, completed results), while message passing handles real-time coordination (requesting help, signaling completion). The workspace acts as a source of truth about what has been accomplished, while the message bus handles the dynamic coordination of who should do what next.

Shared state introduces its own challenges beyond race conditions. When multiple agents write to the same data store, you need clear conventions for what data belongs to which agent and how updates are structured. Namespacing (storing each agent's output under a key that includes the agent ID and task ID) avoids collisions and makes the workspace readable after the fact as a complete audit trail of the multi-agent execution.

Message Schemas

For agents to communicate effectively, their messages need a consistent structure. A minimal message schema includes:

  • Sender: Which agent sent this message
  • Recipient: Which agent (or all agents) should process it
  • Message type: What kind of message this is (task assignment, result report, clarification request, approval request, error notification)
  • Content: The actual payload, typically structured as JSON or plain text
  • Conversation ID: A thread identifier linking related messages

Having explicit message types lets agents route messages efficiently. An orchestrator processing many messages needs to quickly distinguish a subtask completion from an error report. Message types also serve as a contract between agents: when a subagent sends a task_completed message, the orchestrator knows it will find the result in the workspace under the relevant task ID.

Schema evolution is a practical concern. When you update the message format (e.g., add a new field indicating execution time), agents using the old schema need to handle the new field gracefully. Versioning messages and writing agents to ignore unknown fields follows the compatibility principle from network protocol design: be strict in what you send, liberal in what you accept.

Agent Communication Language (ACL)

The FIPA Agent Communication Language, developed in the late 1990s, formalized message protocols for multi-agent systems with performatives like request, inform, propose, agree, and refuse. While modern LLM-based systems rarely implement FIPA ACL directly, its concepts map well to the structured messages exchanged between LLM agents. The core insight, that messages should express content and communicative intent, remains relevant. When a subagent sends a message saying "here is my analysis" (informing), that carries different semantics than "I propose we proceed with approach X" (proposing) even when both contain text.

Conversation Management

When multiple agents exchange multiple messages about a single task, tracking context becomes important. A subagent asking a clarification question and then receiving an answer needs to link those messages to the original task. Conversation threads with unique identifiers solve this: every message carries the conversation ID, and agents maintain a history of the conversation for context.

For LLM-based agents, managing conversation history has a practical cost. Including the full multi-agent message history in every agent's context window is expensive. Systems typically use summarization (compress older messages) or selective inclusion (only include messages relevant to the current agent's active task) to keep context manageable, building on the memory techniques discussed in the Agent Memory chapter.

The tension between completeness and efficiency in conversation management does not have a universal solution. For tasks where the complete history is essential for correctness (e.g., an agent that needs to know about a decision made three rounds ago), selective inclusion risks omitting critical context. For tasks that are truly local (an agent that only needs its own task specification and the results it depends on), including the full history wastes tokens. Adaptive context management, which monitors what context each agent attends to and prunes unused history accordingly, is an active research area.

One underappreciated aspect of conversation management is the role of explicit acknowledgment messages. In human teams, a brief "understood, I will start on this now" serves an important function: it confirms that the message was received and interpreted correctly. Without acknowledgment, a sender cannot distinguish between "message received and task in progress" and "message lost in transit." Building explicit acknowledgment into multi-agent message protocols, while adding overhead, dramatically reduces the number of situations where an orchestrator discovers hours later that a subtask was never started.

Structured Output and Format Contracts

When one agent's output becomes another agent's input, the output format is an implicit contract. If the research agent produces markdown prose and the analysis agent expects a JSON object with specific fields, the pipeline breaks at the boundary. Making this contract explicit through typed data structures prevents an entire class of failures that are otherwise difficult to diagnose because they manifest as downstream errors with no obvious connection to the upstream format mismatch.

Pydantic models (in Python) serve this purpose well. Defining the research agent's output as a ResearchFindings Pydantic model with fields sources, key_claims, confidence_level, and date_range forces the research agent to produce data in a validated format and gives the analysis agent a typed, navigable structure instead of raw text. The validation happens at the boundary: if the research agent produces output that does not conform to the schema, the error is raised immediately rather than propagating silently into the analysis stage.

Structured outputs also make the system more debuggable. When every inter-agent data exchange is typed and versioned, you can inspect the workspace and see exactly what each agent received and produced, in a format that a human can read and reason about. This turns post-hoc debugging from a detective exercise into a straightforward inspection of structured records.

Role Assignment

Effective multi-agent systems require agents suited to their assigned tasks. Role assignment involves defining what each agent can do, determining who should handle each task, and ensuring agents have the appropriate tools and prompts. Done well, role assignment is what turns a collection of identical agents into a productive team with complementary strengths.

Static vs. Dynamic Roles

Static roles are fixed at system design time. The system always has a researcher agent, a writer agent, and a reviewer agent. Each role has a predetermined system prompt and toolset. This is predictable and easy to debug.

Dynamic roles are created as needed. An orchestrator analyzing a task might determine that it needs a database query specialist and a visualization expert, spin up agents with those roles on demand, and terminate them when their subtasks complete. This is more flexible but requires the orchestrator to be good at role specification, since it must write system prompts that reliably elicit the correct specialization.

Dynamic role creation introduces a bootstrapping problem. The orchestrator must generate a system prompt for the new agent without itself being specialized in the domain it is creating a specialist for. Prompts like "you are an expert in X, please do Y" work surprisingly well in practice for LLMs with broad training, but they break down for highly technical domains where subtle prompt differences lead to large quality differences. Maintaining a library of validated role templates that the orchestrator can select and customize is often a better approach than fully dynamic prompt generation.

Skill Matching

When a task arrives, the orchestrator must match it to an appropriate agent. This can be done through several mechanisms:

  • Description matching: The orchestrator has access to descriptions of available agents (what tools they have, what they specialize in) and selects based on semantic similarity between the task and agent descriptions.
  • Routing models: A lightweight classifier or embedding model maps task descriptions to agent types.
  • LLM selection: The orchestrator's own reasoning decides which agent to use based on the task context.

In practice, the simplest approach often works well: maintain a structured registry of agents with descriptions, and have the orchestrator LLM decide allocation from that registry. Embedding-based routing is faster (no LLM call needed for routing decisions) but less flexible than LLM selection when tasks fall between agent specializations.

Skill matching degrades gracefully when no single agent is a perfect fit. The orchestrator can assign the task to the closest-matching agent with additional instructions that compensate for the mismatch (e.g., "you normally write Python code; this task requires JavaScript, but the logic is similar"). This graceful fallback is one reason that LLM-based routing often outperforms rigid classifier-based routing on tasks at the boundaries of role definitions.

Specialization Through Prompting

For LLM-based agents, role specialization happens primarily through the system prompt. A coding agent receives a system prompt emphasizing correctness and readability alongside testing. A researcher receives a prompt emphasizing broad search and source citation. A critic receives a prompt asking it to identify errors and weaknesses, resisting the tendency to be agreeable.

The quality of role assignment depends heavily on how well the orchestrator writes these prompts. This is both an art (clear role definition) and a science (prompt formats that reliably elicit the desired behavior), and it connects to the instruction tuning concepts covered in Part XXXVI.

Specialization through prompting has an important limitation: a single model with a specialized prompt is not the same as a model trained for specialization. A general LLM prompted to behave as a Python expert will outperform the same model prompted generically on Python tasks, but it will underperform a model fine-tuned on Python code. For production systems with predictable, high-volume subtask types, replacing prompted general-purpose agents with fine-tuned specialist models often yields meaningful quality improvements despite the additional training cost.

A practical consideration that often goes overlooked is tool access as a form of role specialization. An agent with access to a code execution environment, a web search API, and a database query tool is a fundamentally different agent than one without these tools, regardless of how its system prompt is written. Restricting tool access to role-appropriate tools also improves safety: a writer agent with no web browsing capability cannot be manipulated through a prompt injection attack delivered via a webpage, because it has no mechanism to visit webpages in the first place.

Agent Coordination Mechanisms

Beyond communication protocols, agents need coordination mechanisms to handle dependencies and conflicts as well as failures. These mechanisms are what prevent a multi-agent system from devolving into chaos when things do not go according to plan.

Dependency Management

Tasks often have dependencies: agent B's work depends on agent A's results. Explicit dependency tracking ensures agents do not start work they cannot complete.

One common approach models task dependencies as a directed acyclic graph (DAG). Each node represents a subtask, and a directed edge from node AA to node BB means task BB cannot start until task AA completes. The orchestrator traverses this graph to identify which tasks are currently eligible for execution.

Formally, a task tt is ready if all tasks in its dependency set deps(t)\text{deps}(t) have status completed\text{completed}:

ready(t)=(status(t)=pending)  ∧  ∀ t′∈deps(t):status(t′)=completed\text{ready}(t) = (\text{status}(t) = \text{pending}) \;\land\; \forall\, t' \in \text{deps}(t) : \text{status}(t') = \text{completed}

where:

  • tt is the task being checked
  • status(t)\text{status}(t) is the current execution status of task tt (pending, running, completed, or failed)
  • deps(t)\text{deps}(t) is the set of tasks that must complete before tt can start
  • ∧\land denotes logical AND
  • ∀\forall means "for all tasks in the dependency set"

When task AA completes, it signals the orchestrator, which re-evaluates which tasks are now ready by checking the above condition for each pending task. The orchestrator maintains a status map (pending, running, completed, failed) for each subtask and updates it as agents report back.

Without explicit dependency tracking, an agent might try to write a summary before the research it depends on is complete, producing low-quality or hallucinated output. The DAG formulation also makes parallelism explicit: tasks with no unsatisfied dependencies can run simultaneously, and the critical path (the longest chain of dependent tasks) determines the minimum possible completion time.

Dynamic task discovery adds complexity. An agent executing a research task might discover that the user's question requires information from an additional source not anticipated in the initial task plan. Allowing agents to dynamically add tasks to the DAG enables more thorough exploration at the cost of making the system's execution time less predictable. Systems that support dynamic task addition need safeguards against unbounded expansion (e.g., depth limits, budget limits) to prevent runaway exploration.

The critical path concept from classical project management applies directly here. If the research tasks each take 10 seconds and must all complete before the analysis begins, and the analysis takes 15 seconds, and the writing takes 10 seconds, then the minimum completion time is roughly 35 seconds regardless of how many parallel agents you add to the research phase. Adding more research agents reduces the research phase to approximately 10 seconds (since they run in parallel), but the serial analysis and writing phases set an absolute floor. Understanding your system's critical path tells you exactly where parallelism will and will not help, which guides intelligent resource allocation.

Consensus and Voting

When multiple agents independently address the same problem and produce different answers, the system needs a way to pick or synthesize a final answer. Several approaches exist:

Majority voting: Each of kk agents returns an answer from a discrete answer set A\mathcal{A}, and the selected answer is the one appearing most often:

a^=arg⁡max⁡a∈A∑i=1k1[answeri=a]\hat{a} = \arg\max_{a \in \mathcal{A}} \sum_{i=1}^{k} \mathbf{1}[\text{answer}_i = a]

where:

  • kk is the number of agents voting
  • A\mathcal{A} is the set of candidate answers
  • answeri\text{answer}_i is the answer produced by agent ii
  • 1[⋅]\mathbf{1}[\cdot] is the indicator function (1 if the condition holds, 0 otherwise)
  • a^\hat{a} is the plurality winner

This works well when answers are drawn from a finite set (for example, classification tasks) but poorly when answers are open-ended text, since two agents expressing the same idea in different words would count as distinct answers.

Self-Consistency prompting (Wang et al., 2023) applies a similar principle to a single model by sampling the model multiple times with temperature > 0 and taking the majority answer. Multi-agent voting generalizes this by using different agents (potentially with different prompts or models), which introduces diversity beyond sampling noise from a single model. The diversity is both the strength and the weakness: agents may disagree not because one is wrong but because they are each emphasizing different valid perspectives.

Judge agent: A designated judge agent receives all candidate answers and selects the best one, optionally explaining its reasoning. The judge's own biases become a concern, but the pattern works well in practice when the judge uses explicit evaluation criteria.

Synthesis: Rather than picking one answer, a synthesis agent combines the best elements from multiple candidate answers. This often produces better results than any individual answer when subagents cover different aspects of the problem.

Research on LLM debate (Irving et al., 2018; Du et al., 2023) has shown that having models argue for and against positions and then revise their answers based on the arguments improves final answer quality on reasoning tasks. The improvement comes from adversarial pressure forcing each model to address weaknesses in its own reasoning. Importantly, the agents must disagree over the answer: if agents are too agreeable (a known tendency of instruction-tuned LLMs), they will update toward each other's answers even when the other answer is wrong. Prompting agents to maintain positions unless presented with new evidence helps preserve the adversarial pressure.

The Condorcet jury theorem from political science provides a useful theoretical lens. The theorem states that if each juror independently has a better-than-random probability p>0.5p > 0.5 of reaching the correct verdict, then a majority vote of kk jurors has a probability that approaches 1 as kk grows. Formally, the collective accuracy for an odd number of voters kk is:

P(majority correct)=∑j=⌈k/2⌉k(kj)pj(1−p)k−jP(\text{majority correct}) = \sum_{j = \lceil k/2 \rceil}^{k} \binom{k}{j} p^j (1-p)^{k-j}

This result has two critical implications. First, voting only improves accuracy when individual agents are better than chance. Second, the improvement has diminishing returns: the most significant gain is from moving from 1 to 3 agents, with each additional agent contributing less than the last. The practical consequence is that 5 to 7 agents is typically sufficient for majority voting; adding 20 agents instead of 7 provides marginal benefit at substantial cost.

Conflict Resolution

Agents may produce conflicting outputs: two agents might claim different facts, recommend different actions, or produce code that contradicts design decisions made by another agent. Conflict resolution strategies include:

  • Priority-based: A designated authority (orchestrator or specialist judge) resolves conflicts.
  • Recency-based: The most recent agent output overrides older outputs (useful when task requirements have changed).
  • Confidence-based: Agents express confidence in their outputs, and the higher-confidence output wins (with the orchestrator arbitrating ties).
  • Human escalation: Conflicts that agents cannot resolve trigger a request for human input.

Confidence-based resolution requires agents to produce reliable confidence estimates, which is a hard problem for LLMs. Models tend to be overconfident (expressing high confidence in wrong answers) in a way that correlates imperfectly with actual accuracy. Calibration techniques (including those discussed in the Hallucination and Factuality part of this book) can improve this, but overconfidence remains a practical concern when using confidence to resolve disagreements between agents.

Human escalation deserves special attention as a design principle and a fallback. Well-designed multi-agent systems proactively identify categories of conflicts that should always involve a human rather than attempting fully automated resolution. Conflicts about high-stakes factual claims, about actions with irreversible consequences, or about requirements that depend on business context outside the system's knowledge are better escalated than resolved autonomously. Building clear escalation paths into the system architecture from the start prevents the common failure mode where automated resolution produces a confident but wrong answer on exactly the high-stakes decision where a human should have been consulted.

Fault Tolerance

Agents fail. An API call might time out, a code execution might error out, or an agent might produce an output that downstream agents cannot process. Reliable multi-agent systems handle these failures gracefully:

  • Retry logic: Failed subtasks are retried, potentially with modified parameters.
  • Fallback agents: If the primary agent for a role fails, a backup agent with different characteristics (perhaps a simpler model) handles the task.
  • Graceful degradation: If a non-critical subtask fails, the system produces a partial result rather than failing entirely.
  • Dead letter queues: Failed tasks are logged for human inspection and potential manual resolution.

Fault tolerance requires the orchestrator to have a clear model of which subtasks are critical (failure is unacceptable) versus optional (failure is tolerable). Tagging tasks with criticality levels at decomposition time lets the orchestrator apply different retry budgets and escalation paths depending on the task. A failure in a verification subtask might trigger a retry with a different verifier model; a failure in an optional context-enrichment subtask might simply result in the enrichment being skipped.

The most subtle failure mode is silent success: an agent reports completion but produces incorrect or incomplete output that passes format validation. This is harder to detect than an explicit failure and more dangerous because it does not trigger retry logic. Output quality validation, separate from format validation, provides a second line of defense. Having a lightweight LLM-based evaluator check whether output content plausibly addresses the task specification catches many silent failures at modest cost.

Exponential backoff is a fault-tolerance best practice that deserves explicit attention. When an agent fails and is retried immediately, a transient infrastructure failure (e.g., a momentarily overloaded API endpoint) will likely produce another failure. Waiting between retries with exponentially increasing delays gives the underlying infrastructure time to recover, converts many transient failures into eventual successes, and prevents retry storms from amplifying system load precisely when the system is already under stress. A simple policy: wait 1 second after the first failure, 2 seconds after the second, 4 after the third, and cap at 30 seconds. This turns most transient failures into recoverable situations with minimal human intervention.

Worked Example: Research and Report Pipeline

To make these patterns concrete, let's trace through a multi-agent research pipeline for the query: "What are the key factors affecting electric vehicle adoption in Europe?"

Task decomposition by orchestrator:

  1. Research agent A: Find recent statistics on EV adoption rates by country
  2. Research agent B: Identify infrastructure (charging station) coverage data
  3. Research agent C: Review policy incentives across EU member states
  4. Analysis agent D: Synthesize findings from A, B, C into key factors
  5. Writer agent E: Write a final structured report from D's analysis

Agents A, B, and C run in parallel since they do not depend on each other. The analysis agent D waits for all three to complete before beginning. The writer agent E waits for D.

Communication flow:

The orchestrator writes task assignments into a shared workspace with unique task IDs. Each research agent reads its assignment, executes web searches and data retrieval using its tools, and writes structured findings back to the workspace tagged with the task ID and agent identity.

When all three research tasks show status "completed," the orchestrator spawns the analysis agent with the combined findings from A, B, and C in its context. The analysis agent produces a structured summary of key factors and writes it to the workspace. The orchestrator then spawns the writer agent with this summary.

Conflict handling:

Research agents A and C both find data on government subsidies but report different figures. The analysis agent D, instructed to cite sources and flag discrepancies, notes the inconsistency and includes both figures with their sources rather than arbitrarily choosing one. The writer agent E presents this as a range with appropriate caveats. This is a practical illustration of why conflict resolution mechanisms matter: without instructions to flag discrepancies, the analysis agent might silently pick one figure, producing a misleadingly precise result.

Output quality vs. single agent:

The parallel search dramatically reduces time: three simultaneous web searches take roughly the same wall-clock time as one, rather than three times as long. The specialized research agents each go deeper into their domain than a single general-purpose agent would. The dedicated analysis and writing stages produce cleaner separation between raw data gathering and synthesis.

A fair comparison would also run a single high-quality agent with a budget equivalent to five sequential agent calls, using chain-of-thought and self-critique loops to compensate for the lack of parallelism and specialization. In practice, the multi-agent approach tends to win on breadth (more ground covered) and the single-agent approach sometimes wins on coherence (no stitching artifacts from combining outputs from multiple agents with different writing styles).

Code Implementation

Let's implement a simplified multi-agent coordination system that demonstrates the core mechanisms: an orchestrator that decomposes tasks, dispatches to specialized agents, and synthesizes results.

Setting Up the Agent Framework

In[3]:
Code
from dataclasses import dataclass, field
from datetime import datetime
from enum import Enum
from typing import List, Optional


class TaskStatus(Enum):
    PENDING = "pending"
    RUNNING = "running"
    COMPLETED = "completed"
    FAILED = "failed"


@dataclass
class Message:
    sender: str
    recipient: str
    message_type: str
    content: str
    conversation_id: str
    timestamp: str = field(
        default_factory=lambda: datetime.utcnow().isoformat()
    )


@dataclass
class Task:
    task_id: str
    description: str
    assigned_to: str
    status: TaskStatus = TaskStatus.PENDING
    result: Optional[str] = None
    dependencies: List[str] = field(default_factory=list)
    created_at: str = field(
        default_factory=lambda: datetime.utcnow().isoformat()
    )
Out[4]:
Console
Agent framework classes defined:
  TaskStatus values: ['pending', 'running', 'completed', 'failed']
  Message fields: sender, recipient, message_type, content, conversation_id, timestamp
  Task fields: task_id, description, assigned_to, status, result, dependencies

The Task dataclass carries all information about a work item: who should do it, its current status, and any prerequisite tasks that must complete first. The Message dataclass models explicit inter-agent communication with sender and recipient fields alongside type information. Together they form the minimal data model needed to implement dependency tracking and message routing.

Implementing the Shared Workspace

In[5]:
Code
class SharedWorkspace:
    """Central coordination store for a multi-agent system."""

    def __init__(self):
        self.tasks: Dict[str, Task] = {}
        self.messages: List[Message] = []
        self.results: Dict[str, str] = {}

    def add_task(self, task: Task) -> None:
        self.tasks[task.task_id] = task

    def update_task_status(
        self, task_id: str, status: TaskStatus, result: Optional[str] = None
    ) -> None:
        if task_id in self.tasks:
            self.tasks[task_id].status = status
            if result is not None:
                self.tasks[task_id].result = result
                self.results[task_id] = result

    def get_ready_tasks(self) -> List[Task]:
        """Return tasks whose dependencies are all completed."""
        ready = []
        for task in self.tasks.values():
            if task.status != TaskStatus.PENDING:
                continue
            all_deps_complete = all(
                self.tasks.get(dep, Task("", "", "")).status
                == TaskStatus.COMPLETED
                for dep in task.dependencies
            )
            if all_deps_complete:
                ready.append(task)
        return ready

    def post_message(self, message: Message) -> None:
        self.messages.append(message)

    def get_messages_for(self, recipient: str) -> List[Message]:
        return [m for m in self.messages if m.recipient in (recipient, "all")]

Implementing Specialized Agents

In[6]:
Code
import json
from typing import Callable


class Agent:
    """A specialized agent that can execute tasks and communicate."""

    def __init__(self, agent_id: str, role: str, execute_fn: Callable):
        self.agent_id = agent_id
        self.role = role
        self._execute = execute_fn

    def handle_task(self, task: Task, workspace: SharedWorkspace) -> str:
        workspace.update_task_status(task.task_id, TaskStatus.RUNNING)
        try:
            result = self._execute(task, workspace)
            workspace.update_task_status(
                task.task_id, TaskStatus.COMPLETED, result
            )
            workspace.post_message(
                Message(
                    sender=self.agent_id,
                    recipient="orchestrator",
                    message_type="task_completed",
                    content=json.dumps(
                        {"task_id": task.task_id, "summary": result[:100]}
                    ),
                    conversation_id=task.task_id,
                )
            )
            return result
        except Exception as e:
            workspace.update_task_status(task.task_id, TaskStatus.FAILED)
            workspace.post_message(
                Message(
                    sender=self.agent_id,
                    recipient="orchestrator",
                    message_type="task_failed",
                    content=json.dumps(
                        {"task_id": task.task_id, "error": str(e)}
                    ),
                    conversation_id=task.task_id,
                )
            )
            raise

Notice the try-except structure: whether the task succeeds or fails, the agent always sends a message to the orchestrator and always updates the task status. This guarantees that the orchestrator is never left waiting for a response that will never come due to an unhandled exception, which would cause the system to stall.

In[7]:
Code
# Define mock execution functions simulating specialized agents
def researcher_execute(task: Task, workspace: SharedWorkspace) -> str:
    """Simulate a research agent finding relevant information."""
    topic = task.description.replace("Research: ", "")
    findings = {
        "EV adoption rates": (
            "Norway leads at 82% EV share. Germany at 18%, France at 15%, UK at 17%. "
            "Market grew 35% YoY across EU in 2023."
        ),
        "charging infrastructure": (
            "EU has 630,000 public charging points. Target is 3.5M by 2030. "
            "Netherlands has highest density; Eastern Europe lags significantly."
        ),
        "policy incentives": (
            "Germany offers 4,500 EUR subsidy on new EVs. France 5,000 EUR. "
            "Norway eliminated VAT entirely. Many countries phase out ICE sales by 2035."
        ),
    }
    for key in findings:
        if any(word in topic.lower() for word in key.split()):
            return findings[key]
    return f"Research complete for: {topic}. Key data points gathered."


def analyst_execute(task: Task, workspace: SharedWorkspace) -> str:
    """Synthesize results from completed upstream research tasks."""
    dep_results = [workspace.results.get(dep, "") for dep in task.dependencies]
    combined = " ".join(dep_results)
    return (
        "Key EV adoption factors identified: (1) Policy support via subsidies and ICE bans "
        "drives adoption most strongly, as shown by Norway's extreme subsidy-adoption correlation. "
        "(2) Charging infrastructure availability limits adoption in lower-density regions. "
        "(3) Price parity with ICE vehicles is approaching in multiple segments by 2025."
    )


def writer_execute(task: Task, workspace: SharedWorkspace) -> str:
    """Draft a structured report from analysis results."""
    analysis = workspace.results.get(
        task.dependencies[0], "No analysis available"
    )
    return (
        f"## EV Adoption in Europe: Key Factors\n\n"
        f"Based on current data: {analysis}\n\n"
        f"**Recommendation**: Prioritize charging infrastructure investment alongside "
        f"demand-side incentives for balanced adoption growth."
    )

Running the Orchestrator

In[8]:
Code
from typing import Dict


class Orchestrator:
    """Manages task assignment, dependency tracking, and result synthesis."""

    def __init__(self, workspace: SharedWorkspace):
        self.workspace = workspace
        self.agents: Dict[str, Agent] = {}
        self.agent_id = "orchestrator"

    def register_agent(self, agent: Agent) -> None:
        self.agents[agent.agent_id] = agent

    def decompose_and_run(self, goal: str, tasks: List[Task]) -> Dict[str, str]:
        """Add all tasks, then execute them in dependency order."""
        for task in tasks:
            self.workspace.add_task(task)

        max_rounds = 20
        for round_num in range(max_rounds):
            ready = self.workspace.get_ready_tasks()
            if not ready:
                pending = [
                    t
                    for t in self.workspace.tasks.values()
                    if t.status == TaskStatus.PENDING
                ]
                if not pending:
                    break
                else:
                    break

            for task in ready:
                if task.assigned_to in self.agents:
                    self.agents[task.assigned_to].handle_task(
                        task, self.workspace
                    )

        return {
            tid: t.result
            for tid, t in self.workspace.tasks.items()
            if t.status == TaskStatus.COMPLETED
        }
In[9]:
Code
# Assemble the multi-agent system
workspace = SharedWorkspace()
orchestrator = Orchestrator(workspace)

researcher_a = Agent("researcher_a", "EV researcher", researcher_execute)
researcher_b = Agent(
    "researcher_b", "Infrastructure researcher", researcher_execute
)
researcher_c = Agent("researcher_c", "Policy researcher", researcher_execute)
analyst = Agent("analyst", "Data analyst", analyst_execute)
writer = Agent("writer", "Report writer", writer_execute)

for agent in [researcher_a, researcher_b, researcher_c, analyst, writer]:
    orchestrator.register_agent(agent)

tasks = [
    Task("t1", "Research: EV adoption rates", "researcher_a"),
    Task("t2", "Research: charging infrastructure", "researcher_b"),
    Task("t3", "Research: policy incentives", "researcher_c"),
    Task(
        "t4",
        "Analyze: synthesize research findings",
        "analyst",
        dependencies=["t1", "t2", "t3"],
    ),
    Task("t5", "Write: final report", "writer", dependencies=["t4"]),
]

results = orchestrator.decompose_and_run("EV adoption report", tasks)
Out[10]:
Console
=== Multi-Agent Pipeline Execution ===

Total tasks: 5
Completed tasks: 5
Messages exchanged: 5

[COMPLETED ] t1: Research: EV adoption rates
[COMPLETED ] t2: Research: charging infrastructure
[COMPLETED ] t3: Research: policy incentives
[COMPLETED ] t4: Analyze: synthesize research findings
[COMPLETED ] t5: Write: final report

--- Final Report ---
## EV Adoption in Europe: Key Factors

Based on current data: Key EV adoption factors identified: (1) Policy support via subsidies and ICE bans drives adoption most strongly, as shown by Norway's extreme subsidy-adoption correlation. (2) Charging infrastructure availability limits adoption in lower-density regions. (3) Price parity with ICE vehicles is approaching in multiple segments by 2025.

**Recommendation**: Prioritize charging infrastructure investment alongside demand-side incentives for balanced adoption growth.

The orchestrator executes tasks in topological order determined by dependencies. The three research tasks complete independently (in parallel in a real async system), the analysis task waits for all three, and the writer produces the final report from the analysis. The message log captures every task completion and failure event for debugging and audit purposes.

Tracking Message Flow

In[11]:
Code
def analyze_message_flow(workspace: SharedWorkspace) -> Dict[str, int]:
    """Count messages by type and sender."""
    type_counts: Dict[str, int] = {}
    sender_counts: Dict[str, int] = {}

    for msg in workspace.messages:
        type_counts[msg.message_type] = type_counts.get(msg.message_type, 0) + 1
        sender_counts[msg.sender] = sender_counts.get(msg.sender, 0) + 1

    return {"by_type": type_counts, "by_sender": sender_counts}
Out[12]:
Console
Message flow analysis:
  By type: {'task_completed': 5}
  By sender: {'researcher_a': 1, 'researcher_b': 1, 'researcher_c': 1, 'analyst': 1, 'writer': 1}

Total coordination messages: 5

Message flow analysis reveals the communication overhead of multi-agent coordination. Every task completion generates a message from the executing agent to the orchestrator, confirming that coordination cost scales linearly with the number of tasks in this design. In a real system with an LLM orchestrator, each message would also require the orchestrator to process the incoming message and decide whether to take action, adding additional LLM inference costs on top of the raw message count.

Simulating Failure and Recovery

In[13]:
Code
def unreliable_execute(task: Task, workspace: SharedWorkspace) -> str:
    """Simulate an agent that fails on first attempt."""
    if not hasattr(unreliable_execute, "_attempts"):
        unreliable_execute._attempts = {}
    attempt_key = task.task_id
    unreliable_execute._attempts[attempt_key] = (
        unreliable_execute._attempts.get(attempt_key, 0) + 1
    )
    if unreliable_execute._attempts[attempt_key] < 2:
        raise RuntimeError(
            f"Transient failure on attempt {unreliable_execute._attempts[attempt_key]}"
        )
    return f"Succeeded on attempt {unreliable_execute._attempts[attempt_key]}"


class ResilientOrchestrator(Orchestrator):
    """Orchestrator with retry logic for failed tasks."""

    def __init__(self, workspace: SharedWorkspace, max_retries: int = 3):
        super().__init__(workspace)
        self.max_retries = max_retries
        self.retry_counts: Dict[str, int] = {}

    def decompose_and_run(self, goal: str, tasks: List[Task]) -> Dict[str, str]:
        for task in tasks:
            self.workspace.add_task(task)

        max_rounds = 30
        for round_num in range(max_rounds):
            ready = self.workspace.get_ready_tasks()
            if not ready:
                failed = [
                    t
                    for t in self.workspace.tasks.values()
                    if t.status == TaskStatus.FAILED
                ]
                retryable = [
                    t
                    for t in failed
                    if self.retry_counts.get(t.task_id, 0) < self.max_retries
                ]
                if retryable:
                    for task in retryable:
                        task.status = TaskStatus.PENDING
                        self.retry_counts[task.task_id] = (
                            self.retry_counts.get(task.task_id, 0) + 1
                        )
                    continue
                break

            for task in ready:
                if task.assigned_to in self.agents:
                    try:
                        self.agents[task.assigned_to].handle_task(
                            task, self.workspace
                        )
                    except Exception:
                        pass

        return {
            tid: t.result
            for tid, t in self.workspace.tasks.items()
            if t.status == TaskStatus.COMPLETED
        }
In[14]:
Code
resilient_workspace = SharedWorkspace()
resilient_orch = ResilientOrchestrator(resilient_workspace, max_retries=3)

unreliable_agent = Agent(
    "unreliable_researcher", "Unreliable researcher", unreliable_execute
)
resilient_orch.register_agent(unreliable_agent)

retry_tasks = [
    Task("r1", "Research: find EV market data", "unreliable_researcher"),
]
retry_results = resilient_orch.decompose_and_run("Retry test", retry_tasks)
Out[15]:
Console
Task status: completed
Retry count: 1
Result: Succeeded on attempt 2
Messages generated: 2
Message types: {'task_failed': 1, 'task_completed': 1}

The resilient orchestrator detects failure, resets the task to pending, and dispatches it again. The unreliable agent fails on its first attempt but succeeds on the second. The message log records both the failure and the eventual success, making the recovery history fully auditable. In a production system you would also add exponential backoff between retries to avoid overwhelming a temporarily overloaded downstream service.

Key Parameters

The key parameters for multi-agent system design are:

  • max_retries: Maximum retry attempts before a task is marked permanently failed. Higher values improve resilience against transient failures but can delay failure detection for persistent errors.
  • max_rounds: Upper bound on orchestration loops. Prevents infinite loops in systems with circular dependencies or stuck agents.
  • dependency graph depth: Deeper DAGs require more sequential execution rounds, limiting parallelism. Flatter graphs execute faster but may sacrifice output quality from synthesis.
  • agent specialization depth: Highly specialized agents with narrow prompts and tools outperform generalist agents on in-domain tasks but fail on edge cases outside their specialization.

Multi-Agent Benchmarks

Measuring multi-agent system performance requires task benchmarks that expose coordination quality alongside final output quality. A benchmark that only measures whether the final answer is correct misses important information about how efficiently agents coordinate, how gracefully they handle failures, and whether the multi-agent setup contributes value beyond what a single agent would achieve.

Task-Centric Benchmarks

Several benchmarks are designed specifically for multi-agent evaluation:

HotpotQA multi-hop reasoning (Yang et al., 2018): Questions that require gathering information from multiple sources before answering. A multi-agent system should outperform a single agent because specialized retrieval agents can focus on different evidence sources simultaneously. The benchmark's characteristic feature is that no single document contains the full answer, forcing information synthesis across sources.

WebArena (Zhou et al., 2023): A web browser benchmark testing agents' ability to complete realistic web tasks. Multi-agent versions test whether an orchestrator and web-browsing subagents can coordinate on complex sequences of web interactions. The benchmark is notable for testing long-horizon task completion where agents must maintain goals across many intermediate steps.

SWE-bench (Jimenez et al., 2023): Software engineering tasks requiring code editing to fix bugs in real repositories. Multi-agent approaches (separate agents for test reproduction, code editing, and verification) have shown improvements over single-agent baselines. The benchmark is challenging because it requires understanding large codebases and making targeted edits without breaking existing functionality.

CAMEL (Li et al., 2023): A role-playing framework specifically for evaluating agent communication and collaboration, where two agents (instructor and assistant) work together to complete tasks through conversation. CAMEL benchmarks measure task success and the quality of the inter-agent dialogue, including how well agents stay on task, handle misunderstandings, and reach consensus.

Metrics Beyond Accuracy

Multi-agent system evaluation should measure more than task success rate:

  • Coordination overhead: How many inter-agent messages are required per task? Systems with excessive back-and-forth may be over-communicating.
  • Parallelism utilization: What fraction of agent-seconds are spent waiting on dependencies? High wait times suggest the decomposition is not exploiting available parallelism.
  • Recovery rate: How often do failed subtasks recover successfully? Low recovery rates indicate fragile agents or insufficient retry logic.
  • Redundancy effectiveness: When multiple agents tackle the same problem for verification, how often does the verifier catch real errors vs. false positives?
  • Latency vs. accuracy tradeoff: Does adding more agents improve accuracy, and at what latency cost? The answer should guide deployment decisions.

Tracking these metrics together reveals whether a multi-agent system is well-designed. A system with high accuracy but very high coordination overhead might be improved by reducing unnecessary communication. A system with low parallelism utilization likely has a suboptimal dependency structure where tasks that could run in parallel are constrained to run sequentially.

Comparing Single Agent to Multi-Agent

A common pattern for fair comparison: run the same task with a single powerful agent using the same total compute budget (time or token budget) as the multi-agent system, and compare final outputs. This controls for the "more resources" confound. If a multi-agent system uses five agent-calls where a single agent uses one, compare to a single agent with five calls' worth of thinking time (perhaps through chain-of-thought or self-critique loops).

Recent work has found that multi-agent debate (where agents criticize each other's answers) outperforms both single-agent responses and single-agent self-critique on mathematical and reasoning benchmarks. Disagreement forces revision rather than mere repetition. Critically, the improvement is most pronounced on problems where there is a clear right answer that can be verified: when there is no ground truth to converge toward, multi-agent debate can amplify shared errors rather than correcting them.

Visualizations

To concretize the quantitative differences between coordination strategies, let's visualize three aspects: how task completion time scales with parallelism, how error rates compound through pipeline stages, and how agent vote count affects answer accuracy.

Out[16]:
Visualization
Line chart showing wall-clock time vs number of tasks for 1, 3, and 5 parallel agents
Theoretical wall-clock time for multi-agent pipelines as the number of parallel subtasks grows, modeled using Amdahl''s law with 30% serial work (analysis and writing stages). The sequential baseline scales linearly with task count, while adding 3 or 5 parallel agents flattens the growth substantially. The curves demonstrate diminishing returns: going from 1 to 3 parallel agents provides the largest speedup, with additional agents offering progressively smaller gains.
Out[17]:
Visualization
Line chart showing pipeline output accuracy vs number of stages for three per-stage accuracy levels
End-to-end accuracy of multi-stage pipelines as a function of stage count, for three per-stage accuracy levels (90%, 95%, 99%). A 95% per-stage accuracy degrades to roughly 77% by the fifth stage, motivating shallow pipeline design. The 99% curve holds above 90% accuracy through seven stages, showing that investment in per-stage quality pays large dividends in end-to-end reliability.
Out[18]:
Visualization
Line chart of collective voting accuracy vs number of agents for three individual accuracy levels
Collective majority-vote accuracy as a function of agent count for three individual accuracy levels (60%, 70%, 80%), derived from the Condorcet jury theorem. Individual accuracy above 70% rapidly yields collective accuracy above 90% with 5-9 agents, but individual accuracy below 50% makes voting counterproductive. The chart illustrates that collective voting provides the largest gains in the first few agents added, with diminishing returns beyond 9-11 agents.

The three charts above illustrate key quantitative relationships in multi-agent system design. The parallelism speedup chart shows that the gains from adding parallel agents are most dramatic in the early additions (going from 1 to 3 agents), with diminishing returns thereafter due to the serial portion of the work. The error compounding chart makes concrete why shallow pipelines matter: even 95% per-stage accuracy degrades to about 77% accuracy after five stages. The voting accuracy chart applies the Condorcet jury theorem to show that majority voting improves collective accuracy only when individual agents are better than chance, and the improvement saturates after about 9 to 11 agents.

Limitations and Practical Considerations

Multi-agent systems are powerful, but they introduce failure modes that single-agent systems do not have. Understanding these limitations clearly is as important as understanding the capabilities, since deployment decisions often hinge on whether the system will hold up in production.

Error Accumulation

In a pipeline, each agent's errors can compound downstream. If the researcher agent produces inaccurate findings, the analyst and writer agents build on those inaccuracies. Unlike a single agent that encounters a problem and can backtrack, a pipeline agent often has no visibility into earlier stages' quality. Adding verification stages helps but increases both latency and cost.

The compounding nature of pipeline errors is a fundamental challenge. A 10% error rate at each of five pipeline stages produces roughly 40% of outputs with at least one error, even if each individual stage is reasonably reliable. This is a strong argument for keeping pipelines shallow (three to four stages at most) and investing heavily in the quality of each stage rather than adding more stages to compensate for earlier deficiencies.

Communication Overhead and Cost

Every agent invocation costs tokens and time. An orchestrator planning five subtasks, dispatching them, reviewing results, and synthesizing outputs might use an order of magnitude more tokens than a single agent tackling the same problem directly. For routine tasks, this overhead is not justified. Multi-agent approaches earn their cost on difficult tasks that benefit from specialization and parallel search.

A useful rule: use multi-agent when the task can be decomposed into independent subproblems, when different subproblems require different specialized knowledge or tools, or when the cost of errors in critical subtasks justifies independent verification. For straightforward single-step tasks, a well-prompted single agent is both cheaper and faster.

Agent invocations account for only part of a multi-agent system's token cost. The orchestrator itself must process all incoming completion messages, maintain an updated view of the workspace, and reason about what to do next after each round. For large task decompositions with many agents, the orchestrator's context window can become expensive. Practical orchestrator design therefore involves keeping completion messages compact (structured status plus a brief summary, with full results in the workspace) rather than including full agent outputs in every message.

Trust and Security

When agents communicate through shared workspaces or message passing, malicious or compromised inputs can propagate. A research agent that processes a webpage containing a prompt injection attack (a webpage that instructs the agent to ignore its task and perform a different action) might write manipulated data into the shared workspace, corrupting downstream agents. Multi-agent systems require careful input sanitization, output validation, and trust boundaries between agent roles. We will address these risks in detail in the Agent Safety chapter.

The trust problem is particularly acute because agents in a multi-agent system are often designed to trust each other's outputs. An analyst that trusts the researcher's findings without independent verification is efficient, but it is also vulnerable to any compromise of the researcher. Defense-in-depth, where each agent validates inputs before acting on them, adds cost but significantly improves security.

Debugging Complexity

Single-agent failures are relatively easy to trace: examine the conversation history, identify the point of error, and fix it. Multi-agent failures involve reconstructing a distributed execution trace across multiple agents, conversation threads, and message exchanges. Good tooling is essential: structured logging, conversation thread tracking, and replay capabilities that let you inspect what each agent received and produced at each step.

Observability frameworks adapted from distributed systems (distributed tracing, correlation IDs, structured logging with context propagation) apply directly to multi-agent systems. Treating each agent invocation as a span in a distributed trace and linking all spans for a single user request through a common trace ID makes failures debuggable even in complex multi-agent topologies.

Prompt Coupling

Because subagents receive structured outputs from other agents as inputs, the output format of one agent becomes a de facto interface contract with the next agent. When you change how the researcher formats its findings, the analyst's prompt may break because it expects specific field names or structures. This prompt coupling is a maintenance burden. Formal output schemas (using JSON Schema or Pydantic models) help enforce consistency.

Treating inter-agent data formats as versioned APIs, with explicit schemas and change management processes, turns prompt coupling from a hidden fragility into an explicit, manageable dependency. This discipline takes effort upfront but pays dividends when the system needs to evolve.

Coordination at Scale

As the number of agents grows, coordination overhead scales in ways that can offset the parallelism gains. Consider what happens when an orchestrator manages 50 agents simultaneously. The orchestrator must process status messages from all 50, maintain an accurate model of the workspace, and route new tasks as dependencies resolve. For an LLM-based orchestrator, this means each orchestration cycle includes a growing volume of context that the orchestrator must process in full before making the next set of dispatch decisions. At some agent count, the orchestrator becomes the bottleneck, spending more time reading and processing messages than any individual agent spends on its assigned work.

Several architectural responses exist. Hierarchical orchestration, where a top-level orchestrator manages a small number of sub-orchestrators who each manage a group of agents, reduces the fan-out at each level. A two-level hierarchy with a top orchestrator managing 5 sub-orchestrators, each managing 10 agents, gives the same 50-agent capacity but reduces the maximum coordination fan-out at any single point from 50 to 10. The tradeoff is additional latency through an extra coordination layer and more complex failure modes when a sub-orchestrator fails.

Another response is to reduce orchestrator involvement in routine progress. Rather than having every agent send a completion message directly to the orchestrator, agents can push results into a structured workspace and the orchestrator periodically polls for readiness conditions. This batch polling pattern reduces the number of LLM calls required for orchestration but introduces latency between task completion and dispatch of the next dependent task.

When Multi-Agent is the Wrong Choice

A single well-designed agent will outperform a multi-agent system in several situations. Tasks that are inherently sequential, where each step depends heavily on the result of the previous step and the next step cannot be specified in advance, do not benefit from parallelism and pay a coordination tax for no gain. Tasks with very low latency requirements may find that the orchestration overhead makes multi-agent slower than a single agent despite the parallelism potential. Tasks requiring strong narrative coherence, where the final output needs to feel like it was written by one voice with one consistent line of reasoning, often suffer from the stitching artifacts that arise when outputs from multiple agents are combined. Recognizing these situations and defaulting to single-agent design preserves resources for cases where multi-agent helps.

Real-World Frameworks

The concepts in this chapter are not just theoretical. Several open-source and commercial frameworks implement them in practical, production-ready ways. Understanding how these frameworks map coordination patterns to code helps bridge the gap between the conceptual discussion above and actual system implementation.

AutoGen (Microsoft Research, 2023) is one of the most widely adopted frameworks for LLM-based multi-agent systems. It provides a ConversableAgent abstraction where agents exchange natural language messages in a chat-like interface, with the framework handling message routing, conversation history, and code execution. AutoGen's GroupChat class implements a peer-to-peer topology where multiple agents take turns contributing to a shared conversation, with a configurable speaker selection strategy that determines which agent responds next. The framework's design philosophy emphasizes flexibility: the same code can run with any LLM backend, and agents can be configured to use human input at specified checkpoints, making it suitable for human-in-the-loop workflows.

LangGraph (LangChain, 2024) takes a different approach, representing multi-agent workflows as explicit state machines where nodes are agents and edges define allowed transitions. This graph-based representation makes the flow of control visible and auditable: you can inspect the graph to understand exactly which agents can execute after which other agents, and the state schema defines precisely what information is available to each agent at each node. LangGraph's explicit state model makes it easier to reason about system behavior and to add persistence (checkpointing the graph state so the system can resume after failures).

CrewAI focuses on hierarchical orchestration with role-based agent definitions. A Crew consists of Agents with defined roles and goals plus backstories, coordinated by a Process (sequential or hierarchical). The framework provides a high-level API that abstracts away many of the coordination details covered in this chapter, which makes it fast to prototype but limits the ability to customize low-level coordination behavior. For teams building systems that fit the standard hierarchical pattern, the abstraction level is appropriate; for teams with unusual coordination requirements, lower-level frameworks provide more control.

Swarm (OpenAI, 2024) introduces the concept of handoffs as a first-class primitive: an agent can transfer control to another agent by returning a handoff instruction, at which point the receiving agent continues the conversation with full context from the sender. This is a lightweight peer-to-peer pattern that avoids the complexity of a full message bus, suitable for workflows where the handoff graph is known in advance and relatively simple. The framework is intentionally minimal. This provides just enough structure for basic multi-agent coordination without abstracting away the LLM API calls themselves.

Each framework embodies different tradeoffs. AutoGen optimizes for flexibility and natural language coordination. LangGraph optimizes for explicit control flow and auditability. CrewAI optimizes for ease of use on standard hierarchical tasks. Swarm optimizes for simplicity and direct API access. Choosing among them involves the same tradeoffs as choosing between coordination topologies: what matters most for your specific task structure, your team's ability to debug the system, and your production reliability requirements.

A practical note on framework selection: the rapid pace of development in this space means that framework capabilities change quickly. Any specific feature comparison will be outdated within months. More durable is the evaluation framework: assess how the tool handles dependency management, how it exposes coordination state for debugging, what it does when agents fail, and how cleanly it integrates with your existing infrastructure. These structural properties change more slowly than surface-level API differences.

Summary

Multi-agent systems extend single-agent capabilities through parallelism and specialization coupled with independent verification. The key concepts covered in this chapter are:

  • Topologies: Hierarchical orchestration (most common), peer-to-peer, and pipeline patterns each suit different task structures. Real systems often combine multiple topologies.
  • Communication: Shared workspaces and message passing are complementary paradigms. Messages need explicit types, conversation IDs, and structured payloads to support reliable coordination.
  • Role assignment: Static roles are predictable; dynamic roles are flexible. Specialization happens primarily through system prompts and tool access.
  • Coordination mechanisms: Dependency DAGs ensure correct execution order, voting and judge agents resolve disagreements, and retry logic with fallback agents provides fault tolerance.
  • Benchmarks: Task-level benchmarks like SWE-bench and WebArena measure real-world coordination quality. Metrics beyond accuracy (coordination overhead, parallelism utilization, recovery rate) capture multi-agent-specific behavior.
  • Limitations: Error accumulation, communication overhead, security risks from inter-agent message injection, and debugging complexity are inherent costs. Use multi-agent approaches when these costs are justified by the task complexity.

The coming chapters in the evaluation and safety parts of this book will examine how to systematically assess and safeguard these systems at scale. For now, understanding the coordination mechanisms and failure modes covered here gives you the foundation to design multi-agent architectures that outperform their single-agent counterparts on the right class of problems.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about multi-agent systems.

Multi-Agent Systems Quiz

Question 1 of 80 of 8 completed
In a hierarchical orchestration topology, what is the primary responsibility of the orchestrator agent?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026multiagent, author = {Michael Brenndoerfer}, title = {Multi-Agent Systems: Coordination and Communication}, year = {2026}, url = {https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Multi-Agent Systems: Coordination and Communication. Retrieved from https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols
MLAAcademic
Michael Brenndoerfer. "Multi-Agent Systems: Coordination and Communication." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols>.
CHICAGOAcademic
Michael Brenndoerfer. "Multi-Agent Systems: Coordination and Communication." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Multi-Agent Systems: Coordination and Communication'. Available at: https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Multi-Agent Systems: Coordination and Communication. https://mbrenndoerfer.com/writing/multi-agent-systems-coordination-communication-protocols

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.