Tool Selection for LLM Agents: Routing Strategies

Michael BrenndoerferFebruary 4, 202658 min read

Part of Language AI Handbook

Covers LLM tool selection through embedding-based routing, hybrid strategies, and semantic interfaces. Build scalable multi-tool agent systems.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Tool Selection and LLM Routing Strategies

In Chapter 2: Function Calling, we explored how large language models can generate structured outputs to invoke external capabilities. However, knowing how to call a tool is only half the battle. The more basic challenge lies in determining which tool to call, or whether to use a tool at all. This decision process, known as tool selection, sits at the heart of effective agent systems.

When you provide an LLM with access to a calculator, a search engine, a code interpreter, and a database, the model faces a routing problem analogous to a network packet finding its destination. But unlike network routing, which relies on fixed lookup tables and deterministic protocols, tool selection requires deep semantic understanding. The model must parse your intent, map it against available capabilities, and decide on the most appropriate action sequence. This semantic routing challenge is complicated by the ambiguity inherent in natural language: a request to "check my balance" might refer to a bank account, a calorie counter, or a chemical equilibrium depending on context. The model must resolve these ambiguities without explicit programmatic logic, relying instead on pattern recognition and contextual inference.

Think about what happens when you ask a well-organized professional for help. They do not blindly grab the first tool on their workbench. Instead, they quickly assess what you need, mentally scan their available capabilities, and select the most appropriate instrument. If your request is ambiguous, they ask clarifying questions. If no tool fits, they say so. This is precisely what a good tool selection system does, and building one requires careful attention to description quality, routing architecture, and failure handling.

This chapter examines the mechanisms that enable accurate tool selection. We will explore how tool descriptions function as the interface between models and capabilities, investigate routing strategies ranging from embedding-based retrieval to end-to-end learning, and address the complexities that arise when multiple tools must work together. By understanding these selection mechanisms, you will be able to design tool-augmented systems that route requests efficiently and accurately.

Tool Descriptions as Interfaces

The relationship between an LLM and its available tools is mediated entirely through descriptions. Unlike software APIs where function signatures provide strict type contracts and compilation enforces correctness, LLMs interpret tools through natural language and structured schemas. This creates both extraordinary flexibility and inherent fragility: descriptions must be precise enough to guide correct selection, yet general enough to cover diverse use cases. The model cannot execute a tool to see what it does; it must infer functionality from text alone, much like a human reading a manual without ever touching the device.

This textual interface creates a basic design tension. If descriptions are too vague, the model cannot distinguish between similar tools. If they are too specific, they may fail to cover legitimate use cases that fall outside the narrow examples provided. Crafting effective descriptions requires understanding how language models interpret text: they rely on distributional semantics, where meaning emerges from patterns in training data rather than grounded experience.

The quality of tool descriptions has an outsized impact on routing accuracy. A study on GPT-4 function calling found that improving description clarity alone increased correct tool selection by 15-20 percentage points, without any changes to the underlying model. This makes description engineering one of the highest-impact activities when building tool-augmented systems. Developers who treat descriptions as afterthoughts, writing one-line summaries like "Weather tool" or "Get data," find that their systems make systematic routing errors that no amount of prompt tuning elsewhere can fix.

Schema-Based Descriptions

The most common approach to tool description follows the JSON Schema specification, which we encountered in our discussion of function calling. This standard provides a structured yet flexible way to define tool interfaces. Each tool is defined by three needed components:

  • Name: A unique identifier that the model uses to reference the tool in its output. The name should be descriptive but concise, serving as a memorable handle for the capability.
  • Description: A natural language explanation of what the tool does, when to use it, and when not to use it. This field is the primary semantic signal for selection.
  • Parameters: A structured schema defining required and optional inputs, including types and constraints plus descriptions for each argument.

The description field is the primary signal for selection because it bridges the gap between user intent and tool capability. Research has shown that models are highly sensitive to phrasing variations, a phenomenon rooted in how attention mechanisms weight token relationships. A description like "Search the web for current information" elicits different behavior than "Retrieve up-to-date facts from the internet," even when referring to the same underlying capability. The former emphasizes the search action and recency, while the latter suggests fact retrieval, potentially triggering different patterns in the model's associative memory.

Tool Schema Example
{
  "name": "weather_lookup",
  "description": "Retrieve current weather conditions and forecasts for a specific location. Use this when the user asks about temperature, precipitation, or weather conditions.",
  "parameters": {
    "type": "object",
    "properties": {
      "location": {
        "type": "string",
        "description": "City and country, e.g., 'Paris, France'"
      },
      "days": {
        "type": "integer",
        "description": "Number of days for forecast (1-7)",
        "default": 1
      }
    },
    "required": ["location"]
  }
}

The parameter descriptions are equally necessary to successful tool selection. When a model hesitates between two similar tools, subtle differences in parameter definitions often break ties. For instance, if one calculator tool accepts "expression" (mathematical notation) while another accepts "query" (natural language math), the parameter description guides the model toward the appropriate choice based on input format. This distinction helps the model route "2+2" to the expression-based calculator and "what is two plus two" to the natural language variant, even when both tools perform arithmetic.

Naming conventions matter more than developers typically expect. A tool named get_current_weather is more likely to be invoked for weather queries than one named atmospheric_conditions_retriever, even if they are functionally identical, because the former's name pattern is more commonly associated with weather-related contexts in training data. This means you should favor common, domain-appropriate naming over clever or technically precise identifiers. The model's internal associations are built on statistical patterns from human-written code and documentation, not from first principles.

Semantic Richness and Examples

Beyond basic schemas, effective tool descriptions often include few-shot examples showing proper usage. These examples anchor the model's understanding through concrete instances, showing what the tool does and the boundary conditions of when to invoke it. This approach uses the few-shot learning capabilities of transformer models, where performance improves dramatically with task-specific examples in the context.

Consider the distinction between these two description strategies for a calendar tool:

Minimal description:

Add events to the user's calendar.

Rich description with examples:

Schedule events, meetings, and reminders on the user's calendar. Use this when the user mentions specific dates, times, or scheduling requests. Examples of when to use: - "Schedule a meeting tomorrow at 3pm" - "Remind me to call mom on Sunday" - "Block 2 hours for deep work on Friday" Do NOT use for: - General time questions (use get_current_time instead) - Checking availability without scheduling (use check_availability instead)

The rich description provides negative examples (when not to use the tool), which helps reduce false positives by defining decision boundaries. It also establishes relationships between tools, implicitly defining the routing policy through description content rather than explicit logic. This contextual steering is particularly important in crowded tool ecosystems where multiple tools might plausibly handle a request.

Negative examples deserve particular emphasis because they address one of the most common failure modes in production systems: false positives. Without guidance about when not to invoke a tool, models tend to over-trigger on tools whose names or descriptions partially match the query. A calendar tool might get invoked when the user simply asks "what day is it?" because both the query and the description contain temporal concepts. Explicit negative examples prevent this kind of shallow matching by forcing the model to distinguish between related but distinct capabilities.

The examples you choose should reflect the actual distribution of queries your system will receive, not idealized textbook cases. If your users frequently ask questions in informal language ("whats the weather tmrw?"), your description examples should include informal phrasings, not just grammatically pristine sentences. This alignment between training distribution and deployment distribution is a core principle of machine learning, and it applies equally to the "micro-training" that few-shot examples provide within the context window.

Description Embeddings and Retrieval

When the number of available tools grows large, including all descriptions in the context window becomes impractical or impossible. As we discussed in Part XLIV: Retrieval-Augmented Generation, we can treat tool selection as a retrieval problem, applying the same principles used for document retrieval to the domain of capability discovery.

In this paradigm, each tool description is embedded into a vector space using a sentence encoder. When a user query arrives, we embed the query and retrieve the kk most similar tool descriptions via vector similarity. Only these top-kk candidates are presented to the LLM for final selection. This approach turns an intractable NN-way classification problem into a manageable kk-way decision, where k≪Nk \ll N.

This approach, sometimes called embedding-based routing, reduces context length and focuses the model's attention on relevant capabilities. However, it introduces a retrieval bottleneck: if the embedding model fails to retrieve the correct tool, the LLM never sees it, regardless of its reasoning capabilities. This creates a hard ceiling on system performance determined by the retriever's recall, making the choice of embedding model and similarity metric necessary design decisions.

The retrieval bottleneck is particularly problematic for rare or specialized tools. A tool that gets invoked infrequently will have few query examples in the embedding space, which makes it harder to surface through similarity search. This is a form of the long-tail problem: common tools become even more dominant because they are more easily retrieved, while niche but potentially valuable tools are systematically underexposed. Addressing this requires careful attention to description diversity and, in some cases, data augmentation techniques that generate synthetic query examples for underrepresented tools.

Designing Descriptions for Scale

When you anticipate having dozens or hundreds of tools, description design requires a systematic approach rather than ad-hoc writing. Several principles improve description quality at scale.

Consistent vocabulary: Choose a consistent vocabulary for describing tool capabilities across your entire ecosystem. If some tools use "retrieve" and others use "fetch," "get," and "obtain" to describe the same operation, embedding similarity becomes noisy. Standardizing on a small set of action verbs that map to distinct capability categories helps embeddings cluster correctly.

Scope clarity: Each description should make the tool's scope unmistakably clear. Rather than "Handles user data," write "Read and write user profile fields including name and email plus preferences. Does not handle authentication or payment information." The more precisely you define what a tool covers, the sharper the decision boundary becomes.

Hierarchical structure: For large tool ecosystems, consider organizing tools into namespaces or categories. A tool named finance.get_balance is more clearly scoped than get_balance, and the namespace prefix helps the model reason about tool groups before selecting specific tools. This hierarchical organization also enables two-stage routing, where the first stage selects a category and the second selects a specific tool within that category.

Version and deprecation notes: In evolving systems, descriptions should note when tools are preferred over others. "Use this instead of the deprecated old_search_tool" directly steers the model away from legacy endpoints that might still appear in its training data.

Tool Routing Strategies

Routing determines how a system maps incoming requests to tool invocations. The strategy you choose depends on latency requirements, accuracy needs, and the complexity of your tool ecosystem. Simple systems might delegate all decisions to the LLM, while high-throughput applications require sophisticated multi-stage pipelines. Understanding these trade-offs enables architects to match routing mechanisms to operational constraints.

The routing problem can be framed as a classification task: given an input query qq and a set of available tools T={t1,t2,…,tN}\mathcal{T} = \{t_1, t_2, \ldots, t_N\}, we want to find the tool t∗t^* that maximizes some measure of appropriateness:

t∗=arg⁡max⁡t∈T∪{none} P(t∣q,T)t^* = \underset{t \in \mathcal{T} \cup \{\text{none}\}}{\arg\max} \ P(t \mid q, \mathcal{T})

where the "none" option allows the system to decline tool use when no tool is appropriate, preventing hallucinated invocations on unrelated queries. Different routing strategies represent different approaches to estimating this probability, trading off computational cost against estimation quality.

LLM-Based Routing

The most straightforward approach uses the LLM itself as the router. The model receives the user query along with all available tool descriptions, and generates a structured output showing which tool (if any) to invoke. This is the standard function calling mechanism we covered previously, using the model's inherent reasoning capabilities to perform semantic matching.

The advantage of LLM-based routing is flexibility. The model can handle fine-grained distinctions between tools, resolve ambiguities through implicit world knowledge, and chain multiple selections in a single reasoning trace. It can recognize that "find my keys" requires a different tool than "find information about keys" based on subtle linguistic cues. However, this flexibility comes at a cost: it incurs the full latency and computational expense of LLM inference for every routing decision, which makes it unsuitable for high-frequency applications or environments with strict latency budgets.

Another important benefit of LLM-based routing is its ability to use conversational context. When routing a query in isolation, the model might ambiguously select between a database lookup and a web search. But given the preceding conversation, the model can infer that the user is working within a company's internal system, ruling out public web search as the appropriate choice. This context-sensitivity is difficult to replicate in simpler routing approaches that treat each query independently.

In[3]:
Code
import json

# Define available tools
tools = [
    {
        "name": "calculate",
        "description": "Evaluate mathematical expressions. Use for arithmetic, algebra, and numeric computations.",
        "parameters": {
            "expression": {
                "type": "string",
                "description": "Math expression like '2 + 2' or 'sqrt(16)'",
            }
        },
    },
    {
        "name": "search",
        "description": "Find current information on the web. Use for recent events, facts, or information not in training data.",
        "parameters": {
            "query": {"type": "string", "description": "Search query"}
        },
    },
    {
        "name": "calendar",
        "description": "Check schedule and availability. Use when users ask about free time or existing commitments.",
        "parameters": {
            "date": {
                "type": "string",
                "description": "Date to check, e.g., '2024-01-15'",
            }
        },
    },
]


def format_tools_for_prompt(tool_list):
    """Format tools into a system prompt section."""
    return json.dumps(tool_list, indent=2)
Out[4]:
Console
Available tools formatted for LLM context:
[
  {
    "name": "calculate",
    "description": "Evaluate mathematical expressions. Use for arithmetic, algebra, and numeric computations.",
    "parameters": {
      "expression": {
        "type": "string",
        "description": "Math expression like '2 + 2' or 'sqrt(16)'"
      }
    }
  },
  {
    "name": "search",
    "description": "Find current information on the web. Use for recent events, facts, or information not in training data.",
    "parameters": {
      "query": {
        "type": "string",
        "description": "Search query"
      }
    }
  },
  {
    "name": "calendar",
    "description": "Check schedule and availability. Use when users ask about free time or existing commitments.",
    "parameters": {
      "date": {
        "type": "string",
        "description": "Date to check, e.g., '2024-01-15'"
      }
    }
  }
]

This JSON representation shows how tool schemas are structured for the LLM context, with each tool's description and parameters clearly delineated to guide the model's selection process.

In this setup, the LLM evaluates the semantic similarity between the user intent and each tool description, effectively performing classification over the tool set. The model's attention mechanisms compute relevance scores implicitly during the forward pass, weighing the semantic overlap between query tokens and description tokens to determine the most appropriate tool.

The practical limitation of pure LLM routing becomes apparent at scale. With 50 tools, each with a 200-word description, you are pushing roughly 10,000 tokens of tool context into every request. At GPT-4 pricing, that adds real cost. Beyond cost, there is a cognitive load problem: models perform worse at discriminating between the correct tool and distractors when the context window contains many irrelevant options. This is a manifestation of the attention dilution problem, where relevant signals get buried in noise.

Similarity-Based Routing

For applications requiring lower latency or operating at scale, we can decouple routing from the LLM entirely. As mentioned earlier, this approach uses dense retrieval to pre-filter tools. Building on our understanding of Dense Retrieval from Part XLIV, we encode tool descriptions and user queries into the same vector space using bi-encoder architectures.

The routing decision becomes a nearest neighbor search in this embedding space:

toolselected=arg⁡max⁡t∈Tools sim(q,desc(t))\text{tool}_{\text{selected}} = \underset{t \in \text{Tools}}{\arg\max} \ \text{sim}(q, \text{desc}(t))

where:

  • toolselected\text{tool}_{\text{selected}}: the tool chosen for execution
  • tt: a candidate tool from the available set Tools\text{Tools}
  • qq: the embedding vector representing the user's query
  • desc(t)\text{desc}(t): the embedding vector of tool tt's description
  • sim(⋅,⋅)\text{sim}(\cdot, \cdot): a similarity function (typically cosine similarity or dot product) measuring alignment between two embeddings

This approach works because embeddings capture semantic meaning through the distributional hypothesis: words and phrases with similar meanings appear in similar contexts during training, causing them to map to nearby points in the vector space. Thus, geometric distance approximates semantic relevance, letting us to find tools whose descriptions match the query intent without expensive neural inference at decision time.

The key question is which embedding model to use. General-purpose models like all-MiniLM-L6-v2 or text-embedding-3-small work well for everyday language but can struggle with domain-specific terminology or unusual tool names. A query about "reconciling journal entries" will embed close to "accounting" and "finance" concepts in general-purpose space, but only a domain-aware model will correctly distance it from "journalism" and "diary writing." For specialized applications, fine-tuning a bi-encoder on domain-specific query-tool pairs consistently outperforms zero-shot retrieval.

In[5]:
Code
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "all-MiniLM-L6-v2"
)  # Run once locally, then comment out

# Create embeddings for tool descriptions
tool_descriptions = [
    "Evaluate mathematical expressions and perform calculations",
    "Search the web for current information and facts",
    "Check calendar schedule and availability",
    "Send emails and messages to contacts",
    "Translate text between different languages",
]

tool_embeddings = model.encode(tool_descriptions)

# Define tool metadata
tool_registry = [
    {"name": "calculate", "embedding_idx": 0},
    {"name": "search", "embedding_idx": 1},
    {"name": "calendar", "embedding_idx": 2},
    {"name": "email", "embedding_idx": 3},
    {"name": "translate", "embedding_idx": 4},
]
In[6]:
Code
import numpy as np


def route_query(query, threshold=0.3):
    """Route query to most similar tool using embeddings."""
    query_embedding = model.encode([query])[0]

    # Compute cosine similarities
    similarities = np.dot(tool_embeddings, query_embedding) / (
        np.linalg.norm(tool_embeddings, axis=1)
        * np.linalg.norm(query_embedding)
    )

    best_idx = np.argmax(similarities)
    best_score = similarities[best_idx]

    if best_score < threshold:
        return None, best_score  # No suitable tool found

    return tool_registry[best_idx]["name"], best_score


# Test routing
test_queries = [
    "What is 15 times 23?",
    "Who won the World Series last year?",
    "Do I have meetings on Friday?",
    "What's the weather like today?",  # No matching tool
]
Out[7]:
Console
Similarity-based routing results:

Query: 'What is 15 times 23?'
  Confidence: 0.122 → NO MATCH

Query: 'Who won the World Series last year?'
  Confidence: 0.097 → NO MATCH

Query: 'Do I have meetings on Friday?'
  Confidence: 0.341 → calendar

Query: 'What's the weather like today?'
  Confidence: 0.088 → NO MATCH

The results demonstrate how embedding-based routing assigns confidence scores based on semantic similarity. High scores (above 0.5) indicate strong matches between query intent and tool descriptions, while scores below the 0.3 threshold correctly trigger the "NO MATCH" fallback for irrelevant queries like "What's the weather like today?" which has no corresponding tool in our registry.

Out[8]:
Visualization
Scatter plot of tool descriptions as blue squares and user queries as orange circles in 2D PCA space, with accepted query Q3 connected by a dashed line to the calendar tool.
Two-dimensional PCA projection of tool descriptions and user queries in embedding space. Geometric proximity corresponds to semantic similarity, and a dashed line connects a query to its nearest tool only when similarity clears the 0.3 acceptance threshold, making accepted and unmatched intents visible.

Similarity-based routing is computationally efficient but lacks the fine-grained reasoning of LLM-based approaches. It cannot easily handle cases where tool selection depends on subtle constraints, requires multi-hop reasoning, or depends on conversational context not captured in the current query. It also struggles with out-of-distribution queries that are semantically distant from any tool description in the embedding space.

The threshold parameter is one of the most important design decisions in embedding-based routing. Set it too high, and many valid queries get dropped; set it too low, and the system invokes tools for irrelevant queries. Rather than using a single global threshold, calibrated per-tool thresholds tend to work better: a highly specific tool like convert_currency can tolerate a lower threshold because false positives are less costly, while a destructive tool like delete_records should require a higher threshold to prevent accidental invocations.

Hybrid Routing

Production systems often combine both approaches in a cascade that uses the strengths of each while mitigating their weaknesses:

  1. Retrieval stage: Use fast embedding similarity to filter from NN tools to a candidate set of k≪Nk \ll N tools. This stage acts as a coarse filter, eliminating obviously irrelevant options at minimal computational cost.
  2. Ranking stage: Use the LLM to select from the kk candidates, or determine that none are appropriate. This stage provides fine-grained discrimination, handling nuances and dependencies that embeddings cannot capture.

This hybrid approach balances efficiency with accuracy. The embedding step acts as a coarse filter, reducing the cognitive load on the LLM by eliminating distractors, while the LLM provides the final fine-grained discrimination necessary for high-stakes decisions. The approach is analogous to human decision-making: we first retrieve relevant options from memory, then deliberate carefully among the candidates.

The choice of kk governs the basic trade-off. A small kk (e.g., 3) minimizes LLM context length and latency but risks missing the correct tool if the retriever makes an error. A large kk (e.g., 10) provides more safety margin but adds more noise for the LLM to reason through. Empirically, values between 3 and 7 tend to work well for ecosystems with up to a few hundred tools, but the right value depends on how semantically similar your tools are to each other. Ecosystems with many closely related tools benefit from larger kk; ecosystems with distinct, non-overlapping tools can use smaller kk.

Out[9]:
Visualization
Scatter plot of four routing strategies comparing accuracy versus latency. LLM-Only is highest accuracy and highest latency; Embedding-Only is fastest but least accurate; Hybrid variants are in between.
Accuracy versus latency trade-offs for four routing strategies. The hybrid approaches (k=3 and k=5) occupy the sweet spot between the speed of embedding-only routing and the precision of LLM-only routing, showing how the two-stage cascade captures most of the accuracy gain at a fraction of the latency cost.

The diagram illustrates an important pattern: the marginal accuracy gain from moving from Hybrid (k=5) to LLM-Only is small (about 1%), but the latency cost is 2.5 times higher. This diminishing return is typical: the last few percentage points of accuracy in routing systems require disproportionately more compute.

Learned Routers

For systems with hundreds or thousands of tools, even presenting kk candidates to the LLM becomes expensive or intractable. In these scenarios, we can train a dedicated routing model, typically a smaller classifier or a mixture-of-experts system as discussed in Part XXXIII: Mixture of Experts.

The router is trained on historical query-tool pairs, learning to map directly from query text to tool indices. This can be formulated as a multi-class classification problem or, more scalably, as a retrieval task using contrastive learning where similar query-tool pairs are pulled together in embedding space while dissimilar pairs are pushed apart.

The routing function becomes:

P(t∣q)=softmax(Wr⋅Encoder(q)+br)P(t \mid q) = \text{softmax}(W_r \cdot \text{Encoder}(q) + b_r)

where:

  • P(t∣q)P(t \mid q): the probability of selecting tool tt given query qq
  • WrW_r: the learned weight matrix of the router layer
  • brb_r: the learned bias vector of the router layer
  • Encoder(⋅)\text{Encoder}(\cdot): a neural network (typically a transformer or sentence encoder) that converts text into dense vector representations
  • softmax(⋅)\text{softmax}(\cdot): a function that converts raw scores into a probability distribution over tools

The softmax ensures that all selection probabilities sum to 1.0, letting the model to express uncertainty when multiple tools might be appropriate. The learned parameters WrW_r and brb_r adapt the generic encoder outputs to the specific tool selection task through training data. This learned projection can capture domain-specific patterns that generic embeddings miss, such as industry jargon or company-specific terminology.

The training objective for a learned router is typically cross-entropy loss over tool selections:

Lrouter=−∑(q,t∗)∈Dlog⁡P(t∗∣q)\mathcal{L}_{\text{router}} = -\sum_{(q, t^*) \in \mathcal{D}} \log P(t^* \mid q)

where D\mathcal{D} is a dataset of query-tool pairs and t∗t^* is the correct tool for each query. This objective pushes the model to assign high probability to correct tools and low probability to incorrect ones, directly optimizing the decision boundary in the embedding space. With enough training data, the learned router can significantly outperform even the embedding-based retrieval approach, because it is optimized end-to-end for the specific tool set rather than relying on a general-purpose encoder.

The trade-off with learned routers is maintenance cost. Adding a new tool requires retraining the classifier, unlike embedding-based approaches where you simply add a new entry to the vector index. This makes learned routers more appropriate for stable, mature tool ecosystems than for rapidly evolving ones. A common hybrid is to use a learned router for the core set of frequently used tools and fall back to embedding retrieval for the long tail of specialized tools that are added less frequently.

Confidence Calibration and Abstention

All routing approaches should support abstention: the ability to decline routing when confidence is insufficient. Without this capability, the system will always select some tool, even for queries that have nothing to do with available capabilities. This creates a frustrating user experience where the system attempts irrelevant tool calls rather than simply acknowledging it cannot help.

For embedding-based routing, abstention is implemented via the similarity threshold we discussed earlier. For LLM-based routing, it requires explicitly including a "none" option in the tool list and instructing the model to choose it when appropriate:

If none of the available tools are appropriate for this request, respond with: {"name": "none", "reason": "The request cannot be handled by available tools"}

For learned routers, abstention can be modeled as a special class or implemented by thresholding on the maximum softmax probability. The key insight is that a well-calibrated router should be able to express uncertainty: when the maximum P(t∣q)P(t \mid q) across all tools is low, the router should abstain rather than commit to a low-confidence selection.

Multi-Tool Scenarios

Real-world tasks rarely require just one tool. Complex workflows demand sequential tool use, parallel invocations, or conditional branching based on intermediate results. These multi-tool scenarios introduce combinatorial complexity: the model must select appropriate tools, determine their execution order, handle data dependencies, and manage failure modes across the chain.

The basic shift in multi-tool scenarios is from selection to planning. Single-tool routing is a one-shot classification problem: given a query, output a tool. Multi-tool routing is closer to program synthesis: given a query, output a sequence of tool calls that collectively accomplish the goal. The model must reason about what information each step produces and what the next step requires, maintaining a mental model of state as execution proceeds.

Sequential Tool Use

When tools have dependencies, they must execute in sequence. For example, to answer "What's the weather like at the location of my 3pm meeting?", the system must:

  1. Query the calendar to find the meeting location
  2. Query the weather service for that specific location

This creates a tool chain where the output of tool t1t_1 becomes input to tool t2t_2. The LLM must recognize this dependency pattern during the initial selection phase and plan accordingly, understanding that the weather tool cannot be called until the calendar tool provides the location parameter.

This dependency structure is often invisible in natural language. The user does not say "first get the location, then get the weather"; they expect the system to infer this ordering from the semantic structure of the request. Models trained on sufficient examples of chained tool use develop a kind of procedural understanding: they learn that certain types of queries require information gathering before information computation, and that certain parameters cannot be filled until upstream tools have executed.

In[10]:
Code
class ToolChain:
    def __init__(self):
        self.memory = {}  # Store intermediate results

    def execute_sequential(self, steps):
        """
        Execute a sequence of tool calls where later steps
        depend on earlier results.
        """
        results = []
        for step in steps:
            tool_name = step["tool"]
            # Allow referencing previous results via template
            params = self._resolve_params(step["params"])
            result = self._call_tool(tool_name, params)
            self.memory[step["id"]] = result
            results.append(result)
        return results

    def _resolve_params(self, params):
        """Replace template variables with actual values from memory."""
        resolved = {}
        for key, value in params.items():
            if isinstance(value, str) and value.startswith("$"):
                var_name = value[1:]
                resolved[key] = self.memory.get(var_name, value)
            else:
                resolved[key] = value
        return resolved

    def _call_tool(self, tool, params):
        # Mock implementation
        return f"{tool}({params})"

The _resolve_params method implements a simple template substitution: when a parameter value starts with "$", it is treated as a reference to a previous result stored in memory. This allows the LLM to express data dependencies declaratively in its output, rather than requiring it to embed actual values before they are known.

In more sophisticated implementations, the memory store becomes a scratchpad where intermediate results are indexed by step name. The LLM can then reference specific fields extracted from a previous tool output: "$calendar_lookup.location" might reference only the location field of the calendar result rather than the entire JSON object. This fine-grained referencing reduces the amount of context passed to each tool and makes the data flow explicit.

Parallel Tool Selection

Some tasks benefit from invoking multiple tools simultaneously. If a user asks "Compare the weather in New York and London," the system should recognize that two independent weather queries can execute in parallel, reducing total latency.

Modern LLMs support parallel function calling, generating multiple tool invocations in a single response. This requires the selection mechanism to identify independence: tools that do not share input dependencies and do not have side effects that interfere with each other. Determining independence requires analyzing data flow graphs and understanding which tools read versus write shared resources.

The independence check is subtler than it appears. Two weather tool calls for different cities are clearly independent. But two calls to a "search" tool might not be if they are drawing from a rate-limited API that throttles concurrent requests. And two calls to a "database_write" tool are never independent because they might contend for the same record. A reliable parallelism framework requires identifying data dependencies along with resource contention and side-effect semantics, information that must be encoded in tool metadata beyond the basic schema.

Parallel tool calling also affects error handling. When sequential tools fail, recovery is straightforward: you know exactly which step failed and what state was accumulated before the failure. When parallel tools fail, you must handle partial results: some tools may have succeeded while others failed, and the appropriate recovery strategy depends on which combination succeeded. Systems that support parallel tool calling need explicit failure policies: "all-or-nothing" (retry everything if any tool fails), "best-effort" (use whatever succeeded), or "conditional" (retry only failed steps with their dependencies).

Conditional and Branching Selection

Advanced scenarios require conditional logic: "If the meeting room is booked, find an alternative time." This introduces tool selection as planning, where the model must generate a program or plan tree rather than a flat sequence. The selection mechanism must evaluate runtime conditions and dynamically choose branches based on intermediate results.

The ReAct pattern (Reasoning + Acting), which we explore in the next chapter, addresses this by interleaving reasoning traces with tool invocations. Each reasoning step evaluates the current state and determines the next tool selection dynamically, rather than planning the entire sequence upfront. This reactive approach handles uncertainty and failure gracefully, letting the system to backtrack or pivot when tools return unexpected results.

The key insight behind ReAct is that planning and execution should be interleaved, not separated. A system that generates a complete plan before executing any tool is betting that all assumptions about tool outputs are correct. When they are wrong, the entire plan may need to be discarded. A system that reasons between each tool call can adapt: if the calendar tool returns "room unavailable," the next reasoning step can decide to call the "find_alternative_room" tool, rather than trying to anticipate this possibility in the original plan.

This adaptive behavior requires the LLM to maintain coherent state across multiple turns of tool invocation and reasoning. The model must keep track of what has been tried, what succeeded, what failed, and what remains to be done. This is a form of working memory management, and it represents one of the harder challenges in building reliable multi-tool agents. Models with larger context windows and stronger instruction-following capabilities tend to perform better at this kind of stateful planning.

Tool Selection Training

While general-purpose LLMs demonstrate zero-shot tool selection capabilities, specialized performance requires targeted training. Teaching models to reliably choose between tools involves curating demonstration data and optimizing for selection accuracy. The gap between zero-shot and fine-tuned performance is particularly pronounced in domains with specialized terminology or fine-grained distinctions between similar tools.

Zero-shot selection works because modern LLMs have seen many examples of API calls, command-line tool usage, and programmatic interfaces in their pre-training data. They have internalized a general pattern: when a task requires external information or computation, look for a named capability and invoke it with appropriate arguments. But this general pattern does not always translate to correct selection in specific tool ecosystems, particularly when tools have overlapping capabilities or use unconventional naming.

Fine-Tuning for Tool Selection

The standard approach involves supervised fine-tuning on examples of correct tool usage. Each training example consists of:

  • Context: User query and available tool descriptions
  • Target: The correct tool invocation (name and parameters)

The training objective maximizes the log-likelihood of the correct tool selection:

L=−∑(q,t∗)∈Dlog⁡P(t∗∣q,Tools)\mathcal{L} = -\sum_{(q, t^*) \in \mathcal{D}} \log P(t^* \mid q, \text{Tools})

where:

  • L\mathcal{L}: the training loss (negative log-likelihood) to be minimized
  • (q,t∗)∈D(q, t^*) \in \mathcal{D}: a query-ground-truth-tool pair from the training dataset D\mathcal{D}
  • qq: the user query input to the model
  • t∗t^*: the ground-truth (correct) tool for query qq
  • P(t∗∣q,Tools)P(t^* \mid q, \text{Tools}): the probability the model assigns to selecting the correct tool t∗t^* given the query qq and set of available tools

This mirrors the instruction tuning approach covered in Part XXXVI: Instruction Tuning, but with structured outputs constrained to valid tool schemas. The loss function penalizes the model when it assigns low probability to the correct tool, gradually adjusting its internal representations to favor appropriate selections.

Data curation is necessary for effective fine-tuning. The training set must include:

  • Positive examples: Clear cases where the tool is appropriate, establishing the core usage patterns
  • Negative examples: Borderline cases where the tool should not be used, teaching the model decision boundaries
  • Ambiguous cases: Queries that could match multiple tools, teaching the model to distinguish fine-grained differences through subtle cues

The ratio of positive to negative examples matters. Too many positive examples produces a model that over-triggers, invoking tools for marginally related queries. Too many negative examples produces a model that under-triggers, declining tool use even when appropriate. A balanced dataset with roughly equal positive and negative examples, plus a smaller set of hard negatives (queries that are very close to positive cases but require a different tool), tends to produce the most reliable selectors.

Out[11]:
Visualization
Dual-axis chart showing training loss decreasing from 2.3 to 0.48 and validation accuracy rising from 0.45 to 0.945 over 10 epochs.
Training loss and validation accuracy over 10 epochs of tool selection fine-tuning. The model achieves most of its improvement in the first four epochs, showing rapid acquisition of selection patterns, with diminishing returns thereafter as it refines edge cases.

The rapid improvement in early epochs reflects the model quickly learning the basic selection patterns from the training examples. The slower improvement in later epochs corresponds to refinement of edge cases and hard negatives, which require more subtle adjustments to the decision boundary.

Synthetic Data Generation

Collecting real-world tool usage data is expensive and time-consuming. Synthetic generation offers an alternative, using larger teacher models to generate training examples for smaller student models. This distillation approach transfers reasoning capabilities from powerful but expensive models to efficient production models.

The process typically involves:

  1. Tool documentation analysis: Parse API schemas and documentation to understand capabilities and constraints
  2. Query generation: Prompt a teacher LLM to generate diverse queries that would require each tool, covering edge cases and variations in phrasing
  3. Verification: Use the teacher to verify that generated queries correctly map to intended tools, filtering out hallucinated or inconsistent examples
  4. Diversity filtering: Ensure coverage of edge cases and negative examples, using clustering or embedding-based deduplication to maximize informational diversity
In[12]:
Code
def generate_synthetic_training_data(tool_schema, num_examples=100):
    """
    Generate synthetic training examples for tool selection.
    In practice, this would call a teacher LLM API.
    """
    examples = []

    # Generate positive examples
    for i in range(num_examples // 2):
        prompt = f"""
        Tool: {tool_schema["name"]}
        Description: {tool_schema["description"]}
        Generate a user query that would require using this tool.
        Query:"""
        # Simulated LLM response
        query = f"Example query {i} for {tool_schema['name']}"
        examples.append(
            {
                "query": query,
                "tool": tool_schema["name"],
                "parameters": {},  # Would be filled by LLM
            }
        )

    # Generate negative examples (queries for other tools)
    other_tools = ["search", "calendar", "email"]
    for i in range(num_examples // 2):
        tool = other_tools[i % len(other_tools)]
        query = (
            f"Example query that should use {tool}, not {tool_schema['name']}"
        )
        examples.append(
            {
                "query": query,
                "tool": tool,  # Correct tool is different
                "negative_for": tool_schema["name"],
            }
        )

    return examples

Synthetic data generation works particularly well for establishing coverage of rare query patterns. If your tool has a seldom-used parameter combination, you can generate many synthetic examples exercising that combination, even though it almost never appears in real usage. This is impossible with purely organic data collection, which naturally concentrates examples around common patterns and under-represents edge cases.

The quality of synthetic data depends heavily on the teacher model's understanding of the tool ecosystem. A teacher that does not accurately understand the semantics of each tool will generate misleading examples. This is why the verification step is necessary: after generation, each example should be checked to confirm that the teacher would route the generated query to the correct tool. Examples where the teacher disagrees with the intended label should be discarded or reviewed by humans before inclusion in the training set.

Diversity filtering addresses a different problem: even with verification, synthetic generation can produce highly redundant examples that all look like slight variations of "what is [number] times [number]?" for a calculator tool. Redundant examples waste training budget without improving coverage. Embedding-based deduplication, where examples that are too similar in embedding space are removed, ensures that each training example contributes unique coverage of the query distribution.

Reinforcement Learning from Tool Feedback

Beyond imitation learning, we can optimize tool selection using reinforcement learning. The reward function captures task success: did the selected tool accomplish the user's goal? This aligns with the RLHF approaches discussed in Part XXXVII: Alignment and RLHF.

However, tool selection presents unique challenges for RL:

  • Sparse rewards: Success is often only observable after a full sequence of tool calls, which makes it difficult to attribute success or failure to individual selection decisions
  • Credit assignment: Determining which selection decision caused failure in a multi-step workflow requires sophisticated attribution mechanisms
  • Safety constraints: Exploring random tool selections may have real-world consequences (sending emails, making purchases, modifying databases), limiting the exploration strategies available during training

These challenges favor conservative exploration strategies and heavy use of simulation environments during training, where the agent can learn from mistakes without causing real harm. Building faithful simulation environments is itself a significant engineering challenge: the simulator must accurately model tool behavior, including error conditions, rate limits, and side effects, or the policy learned in simulation will fail to transfer to production.

The credit assignment problem deserves particular attention. In a five-step tool chain, if the final step fails because a parameter was incorrectly extracted two steps earlier, the reinforcement signal must somehow propagate back to that earlier decision. Standard discount-factor approaches from RL can accomplish this mathematically, but they require many samples to estimate reliably. Techniques like hindsight experience replay, where failed trajectories are relabeled with modified goals that make them "successes," can improve sample efficiency significantly.

Rejection Sampling and Iteration

A practical middle ground between SFT and RL is rejection sampling (also called Best-of-N sampling). The model generates NN candidate tool selections for a given query, and we filter or rank them based on:

  • Execution success: Does the tool accept the parameters without schema validation errors?
  • Result relevance: Does the output help answer the query?
  • Consistency checks: Would a human expert have chosen this tool?

The successful trajectories become training data for the next iteration, creating a self-improving loop similar to the iterative alignment methods covered in Part XXXVII. This approach uses the model's own generated diversity to discover better solutions without requiring explicit reward modeling.

Rejection sampling has an important practical advantage: it requires no reward model. Instead, execution results serve as direct feedback. If the tool accepts the parameters and returns a useful result, the selection was likely correct. If the tool throws a validation error or returns an empty result, the selection was likely wrong. This execution-based feedback is more reliable than a learned reward model, which can be misspecified or exploited through reward hacking.

The iteration loop is particularly powerful for improving parameter extraction alongside tool selection. Many tool selection failures are not outright wrong tool choices, but rather correct tool choices with incorrectly formatted parameters. By filtering on execution success (which includes parameter validation), the rejection sampling loop trains the model to simultaneously improve both the selection decision and the parameter extraction.

Worked Example: Multi-Tool Selection

Let's walk through a concrete example showing how an LLM works within a complex tool ecosystem. Consider a personal assistant with access to five tools:

  1. Calculator: Mathematical computations
  2. Calendar: Schedule checking and event creation
  3. Email: Sending messages
  4. Search: Web information retrieval
  5. Notes: Saving and retrieving personal notes

User query: "I need to prepare a presentation for next week's team meeting. Can you find the sales figures from Q3, calculate the growth rate from Q2, and schedule 3 hours this Friday to work on slides? Also, remind me to email the team about it."

This query requires decomposing intent across multiple tools, recognizing dependencies between data retrieval and computation, and scheduling actions in the correct temporal order. The LLM must simultaneously parse several sub-goals from a single, complex sentence and construct an execution plan that respects their interdependencies.

In[13]:
Code
class ToolSelector:
    def __init__(self):
        self.tools = {
            "calculator": {
                "desc": "Perform mathematical calculations",
                "triggers": [
                    "calculate",
                    "compute",
                    "growth rate",
                    "percentage",
                    "sum",
                    "average",
                ],
            },
            "calendar": {
                "desc": "Schedule events and check availability",
                "triggers": [
                    "schedule",
                    "meeting",
                    "remind",
                    "block time",
                    "available",
                    "friday",
                    "next week",
                ],
            },
            "email": {
                "desc": "Send emails and messages",
                "triggers": ["email", "send", "message", "write to", "notify"],
            },
            "search": {
                "desc": "Find information and facts",
                "triggers": [
                    "find",
                    "look up",
                    "search",
                    "get",
                    "figures",
                    "data",
                    "q3",
                    "sales",
                ],
            },
            "notes": {
                "desc": "Save and retrieve notes",
                "triggers": ["save", "remember", "note", "write down", "store"],
            },
        }

    def analyze_intent(self, query):
        """Simple keyword-based analysis for demonstration."""
        query_lower = query.lower()
        detected_tools = []

        for tool_name, metadata in self.tools.items():
            score = sum(
                1 for trigger in metadata["triggers"] if trigger in query_lower
            )
            if score > 0:
                detected_tools.append((tool_name, score))

        # Sort by relevance score
        detected_tools.sort(key=lambda x: x[1], reverse=True)
        return detected_tools

    def create_execution_plan(self, query):
        """Create ordered plan based on dependencies."""
        tools_needed = self.analyze_intent(query)

        # Define dependencies: some tools must run before others
        dependencies = {
            "email": ["calendar"],  # Need meeting scheduled before emailing
            "calculator": ["search"],  # Need Q3 data before calculating growth
        }

        plan = []
        completed = set()

        while len(plan) < len(tools_needed):
            added_this_round = False
            for tool_name, score in tools_needed:
                if tool_name in completed:
                    continue

                # Check if dependencies satisfied
                deps = dependencies.get(tool_name, [])
                if all(d in completed for d in deps):
                    plan.append(
                        {
                            "tool": tool_name,
                            "reason": f"Detected intent: '{self.tools[tool_name]['desc']}'",
                            "params": self._extract_params(tool_name, query),
                        }
                    )
                    completed.add(tool_name)
                    added_this_round = True

            if not added_this_round and len(plan) < len(tools_needed):
                # Circular dependency or missing tool
                break

        return plan

    def _extract_params(self, tool_name, query):
        # Simplified parameter extraction
        params = {}
        if tool_name == "calendar":
            if "friday" in query.lower():
                params["day"] = "Friday"
            if "3 hours" in query:
                params["duration"] = "3 hours"
        elif tool_name == "search":
            if "Q3" in query and "sales" in query:
                params["query"] = "Q3 sales figures"
        elif tool_name == "calculator":
            params["task"] = "growth rate Q2 to Q3"
        elif tool_name == "email":
            params["recipient"] = "team"
            params["subject"] = "Presentation preparation"

        return params
In[14]:
Code
selector = ToolSelector()
complex_query = "I need to prepare a presentation for next week's team meeting. Can you find the sales figures from Q3, calculate the growth rate from Q2, and schedule 3 hours this Friday to work on slides? Also, remind me to email the team about it."
intent_analysis = selector.analyze_intent(complex_query)
plan = selector.create_execution_plan(complex_query)
Out[15]:
Console
Query Analysis:
============================================================
User: 'I need to prepare a presentation for next week's team meeting. Can you find the sales figures from Q3, calculate the growth rate from Q2, and schedule 3 hours this Friday to work on slides? Also, remind me to email the team about it.'

Detected tool intents (by keyword matching):
  - calendar: relevance 5
  - search: relevance 4
  - calculator: relevance 2
  - email: relevance 1

============================================================
Execution Plan:

Step 1: CALENDAR
  Reason: Detected intent: 'Schedule events and check availability'
  Parameters: {'day': 'Friday', 'duration': '3 hours'}

Step 2: SEARCH
  Reason: Detected intent: 'Find information and facts'
  Parameters: {'query': 'Q3 sales figures'}

Step 3: CALCULATOR
  Reason: Detected intent: 'Perform mathematical calculations'
  Parameters: {'task': 'growth rate Q2 to Q3'}

Step 4: EMAIL
  Reason: Detected intent: 'Send emails and messages'
  Parameters: {'recipient': 'team', 'subject': 'Presentation preparation'}

The execution plan demonstrates how the system decomposes a complex, multi-part request into discrete tool invocations ordered by their data dependencies. The keyword-based detection successfully identified four relevant tools, with the dependency graph so that data-retrieval steps (search) precede computation steps (calculator). Notice that the email step is deferred to last because it depends on the calendar event being scheduled first.

Out[16]:
Visualization
Horizontal flow diagram showing calendar, search, calculator, and email as colored rectangles connected by arrows in execution order.
Sequential execution plan for the multi-step presentation query. Tools are ordered left to right by their data dependencies: search must complete before calculator can compute the growth rate, and calendar must complete before email can reference the scheduled session.

The analysis reveals four distinct tool invocations. The implicit dependencies (calculator needs search results; email needs calendar event) are automatically resolved by the dependency-aware planner. A sophisticated tool selector recognizes these constraints and orders execution accordingly, or makes parallel calls where independence allows.

In a real LLM-based system, this dependency reasoning happens inside the model's forward pass rather than through explicit dependency graph construction. The model has learned from training data that certain sequences of tool calls commonly appear together, and that certain parameter values are derived from previous tool outputs. This learned co-occurrence pattern effectively captures the dependency structure without requiring the developer to manually enumerate it.

Limitations and Practical Challenges

Tool selection systems face several inherent limitations that impact reliability and scalability. Understanding these constraints is needed for designing reliable production systems and setting appropriate expectations for performance.

Description Sensitivity and Brittleness

Tool selection remains surprisingly sensitive to description phrasing. Minor edits, such as changing "Retrieve current weather" to "Get current weather," can shift selection probabilities enough to alter routing decisions. This brittleness stems from the fact that LLMs lack true understanding of tool capabilities; they rely on pattern matching between query semantics and description semantics, making them vulnerable to superficial variations in wording.

In production systems, this necessitates extensive A/B testing of description variants and continuous monitoring of selection distributions. A sudden shift in the frequency of calculator invocations might indicate that a description change inadvertently altered the decision boundary between mathematical and factual queries. Maintaining stable tool ecosystems requires treating descriptions as necessary code, with version control and regression testing.

The brittleness problem is compounded by the fact that LLMs themselves change over time. A model update that improves performance on one task might subtly shift the token probability distributions used for tool selection, causing previously stable routing decisions to become unreliable. This means that tool descriptions need to be re-validated whenever the underlying model is updated, not just when the tool ecosystem changes. Organizations that run large tool ecosystems often maintain automated regression test suites specifically for routing behavior.

Another dimension of brittleness is language variation. Users phrase requests differently based on their background, expertise level, and communication style. A tool description tuned to match formal, precise queries might fail on casual or colloquial phrasings. Diverse test sets should include informal and abbreviated queries as well as multilingual ones to assess reliability across real-world language variation.

Scaling Challenges

As tool ecosystems grow, selection accuracy degrades due to the curse of dimensionality in semantic space. With two tools, discrimination is easy. With two hundred, the problem becomes significantly harder, suffering from the "long tail" of rare tools that receive insufficient training signal and are often overshadowed by more frequently mentioned tools.

Research suggests that LLM-based routing struggles beyond approximately twenty tools without specialized training. Solutions include hierarchical routing (tools organized into categories, with a two-stage selection process first choosing the category then the specific tool) and dynamic tool retrieval (only presenting relevant tools based on conversation context or user preferences).

Out[17]:
Visualization
Line chart on log-scale x-axis comparing zero-shot LLM and fine-tuned router accuracy versus number of tools, with a dashed vertical line at 20 tools marking the necessary threshold.
Tool selection accuracy as a function of the number of available tools for zero-shot LLM routing and a fine-tuned router. Both approaches degrade as tool count grows, with zero-shot performance declining sharply beyond the 20-tool threshold. Fine-tuning extends reliable performance but cannot fully overcome the scaling challenge.

The scaling curve reveals an important design implication: if your tool ecosystem will eventually grow to hundreds of tools, plan for hierarchical routing from the beginning. Retrofitting a flat routing architecture with hierarchical organization requires changes throughout the system, including tool descriptions, routing logic, and evaluation pipelines. Starting with a two-level hierarchy (categories, then tools) even when you have only twenty tools makes it much easier to add tools later without disrupting existing routing behavior.

Tool Hallucination

Just as LLMs hallucinate facts, they can hallucinate tools: generating invocations for non-existent capabilities or parameters outside defined schemas. This is particularly problematic when the model invents tools that "sound right" for a query but do not exist in the registry, or hallucinates parameters that seem plausible but violate the schema.

Mitigation strategies include:

  • Constrained decoding: Using grammars or logits processors to force valid tool names and parameter structures, as discussed in Constrained Decoding
  • Validation layers: Checking tool names against the registry before execution, rejecting any unknown tools immediately
  • Fallback prompts: Explicitly listing available tools and asking the model to choose from that closed set, reducing the space of possible hallucinations

Constrained decoding is the most reliable defense against hallucination because it operates at the token level: the model physically cannot generate invalid tool names because those token sequences are assigned zero probability. The trade-off is that constrained decoding requires integrating with the inference infrastructure at a low level, which is not always possible with third-party API providers. For hosted model APIs, validation layers and fallback prompts provide the next best defense.

It is worth distinguishing between two types of tool hallucination. The first type invents entirely non-existent tools ("use the financial_forecast_pro tool" when no such tool exists). This is relatively easy to catch with a registry validation check. The second type invokes real tools with hallucinated parameters (calling weather_lookup with {"location": "the moon"}). This is harder to catch because the parameter value passes schema validation but would fail at execution. Downstream validation, where tool outputs are checked for reasonableness before being used, is needed to catch this type of error.

Latency and Cost Trade-offs

Every tool selection strategy involves trade-offs between accuracy and speed. LLM-based routing provides the best accuracy but adds hundreds of milliseconds to response time. Embedding-based routing is faster but less accurate. Hybrid approaches add system complexity and potential failure points at each stage.

For high-frequency applications, the cost of inference becomes significant. Routing a query through GPT-4 to select between local tools may cost more than simply attempting multiple tool calls in parallel and filtering results. The optimal strategy depends on the relative costs of latency and compute along with error rates, requiring careful economic modeling of the trade-off space.

In practice, the right routing architecture is often determined by the failure cost of incorrect selections. For a system that can only send one email per request and cannot undo sends, the cost of a wrong tool selection is high, justifying expensive LLM-based routing. For a system that runs read-only database queries, the cost of a wrong selection is low (just a wasted API call), making fast embedding routing the better choice. Align your routing investment with the cost structure of your tool ecosystem.

Code Implementation: Building a Tool Router

Let's implement a complete tool selection system that combines embedding-based retrieval with LLM-based ranking. This hybrid approach scales to many tools while maintaining accuracy through the two-stage cascade.

In[18]:
Code
from dataclasses import dataclass, field
from typing import Dict, List, Tuple

import numpy as np
from sentence_transformers import SentenceTransformer


@dataclass
class Tool:
    name: str
    description: str
    parameters: Dict
    examples: List[str] = field(default_factory=list)


class HybridToolRouter:
    def __init__(self, embedding_model: str = "all-MiniLM-L6-v2"):
        self.tools: List[Tool] = []
        self.embedder = SentenceTransformer(
            embedding_model
        )  # Run once locally, then comment out
        self.tool_embeddings = None
        self.tool_texts = []

    def register_tool(self, tool: Tool):
        """Add a tool to the registry."""
        self.tools.append(tool)
        # Create rich text representation for embedding
        text = f"{tool.name}: {tool.description}"
        if tool.examples:
            text += f" Examples: {'; '.join(tool.examples)}"
        self.tool_texts.append(text)

    def build_index(self):
        """Compute embeddings for all tools."""
        if not self.tool_texts:
            raise ValueError("No tools registered")
        self.tool_embeddings = self.embedder.encode(self.tool_texts)

    def retrieve_candidates(
        self, query: str, top_k: int = 3
    ) -> List[Tuple[Tool, float]]:
        """Retrieve top-k candidate tools using embeddings."""
        if self.tool_embeddings is None:
            self.build_index()

        query_embedding = self.embedder.encode([query])[0]

        # Compute cosine similarities
        similarities = np.dot(self.tool_embeddings, query_embedding) / (
            np.linalg.norm(self.tool_embeddings, axis=1)
            * np.linalg.norm(query_embedding)
        )

        # Get top-k indices
        top_indices = np.argsort(similarities)[-top_k:][::-1]

        candidates = []
        for idx in top_indices:
            candidates.append((self.tools[idx], float(similarities[idx])))

        return candidates
In[19]:
Code
# Initialize router and register tools
router = HybridToolRouter()

# Register our tool ecosystem
router.register_tool(
    Tool(
        name="calculator",
        description="Evaluate mathematical expressions and perform calculations",
        parameters={"expression": "string"},
        examples=["What is 15 times 23?", "Calculate the square root of 144"],
    )
)

router.register_tool(
    Tool(
        name="weather",
        description="Get current weather conditions and forecasts for locations",
        parameters={"location": "string", "days": "integer"},
        examples=[
            "What's the weather in Paris?",
            "Will it rain tomorrow in London?",
        ],
    )
)

router.register_tool(
    Tool(
        name="search",
        description="Search for current information on the internet",
        parameters={"query": "string"},
        examples=["Who is the current president?", "Latest news about AI"],
    )
)

router.register_tool(
    Tool(
        name="calendar",
        description="Check schedule and create calendar events",
        parameters={"date": "string", "event": "string"},
        examples=["Do I have meetings tomorrow?", "Schedule a call for 3pm"],
    )
)

router.register_tool(
    Tool(
        name="translate",
        description="Translate text between languages",
        parameters={"text": "string", "target_language": "string"},
        examples=[
            "Translate 'hello' to French",
            "How do you say 'thank you' in Japanese?",
        ],
    )
)

router.build_index()
Out[20]:
Console
Registered 5 tools
Embeddings indexed: True

The output confirms that all five tools are registered and embeddings are successfully indexed, showing the vector search infrastructure is ready. The router is now initialized with five diverse tools spanning calculation and weather as well as search and scheduling plus translation capabilities, ready to process incoming queries through the two-stage retrieval and ranking pipeline.

In[21]:
Code
# Test retrieval on ambiguous and clear queries
test_cases = [
    "What's 25 divided by 5?",  # Clear: calculator
    "Will it rain in Seattle tomorrow?",  # Clear: weather
    "How do you say 'goodbye' in Spanish?",  # Clear: translate
    "Find me the latest stock prices",  # Ambiguous: search vs calculator
    "Schedule a meeting for 2pm",  # Clear: calendar
    "Calculate the weather forecast for next week",  # Nonsensical: should fail
]
Out[22]:
Console
Retrieval Results (Top-3 Candidates):
======================================================================

Query: 'What's 25 divided by 5?'
  1. calculator   (score: 0.335) <-- BEST
  2. weather      (score: 0.055) 
  3. calendar     (score: 0.042) 

Query: 'Will it rain in Seattle tomorrow?'
  1. weather      (score: 0.370) <-- BEST
  2. calendar     (score: 0.212) 
  3. calculator   (score: 0.018) 

Query: 'How do you say 'goodbye' in Spanish?'
  1. translate    (score: 0.477) <-- BEST
  2. calendar     (score: 0.070) 
  3. calculator   (score: 0.053) 

Query: 'Find me the latest stock prices'
  1. search       (score: 0.343) <-- BEST
  2. weather      (score: 0.239) 
  3. calculator   (score: 0.101) 

Query: 'Schedule a meeting for 2pm'
  1. calendar     (score: 0.677) <-- BEST
  2. weather      (score: 0.116) 
  3. calculator   (score: 0.073) 

Query: 'Calculate the weather forecast for next week'
  1. weather      (score: 0.412) <-- BEST
  2. calendar     (score: 0.253) 
  3. calculator   (score: 0.183)

The retrieval results show clear semantic clustering: mathematical queries align with the calculator (scores around 0.6-0.7), weather queries match the weather tool, and nonsensical queries like "Calculate the weather forecast" produce low scores across all candidates (below 0.4), signaling that no appropriate tool exists.

Out[23]:
Visualization
Histogram of cosine similarity scores with bars color-coded red below 0.3 and green above 0.5, with dashed threshold lines.
Distribution of retrieval confidence scores across all test queries and top-3 candidates. The bimodal distribution reveals a natural separation between high-confidence matches (above 0.5, shown in green) and low-confidence mismatches (below 0.3, shown in red), with the threshold region in between capturing ambiguous cases that benefit from LLM disambiguation.
In[24]:
Code
from typing import Optional


def simulate_llm_selection(
    query: str, candidates: List[Tuple[Tool, float]]
) -> Optional[Tool]:
    """
    Simulate LLM-based selection from candidates.
    In production, this would be an actual LLM API call.
    """
    if not candidates:
        return None

    best_tool, best_score = candidates[0]

    # Decline if best match is below threshold
    if best_score < 0.3:
        return None

    # Check for parameter compatibility (simulated)
    if best_tool.name == "calculator" and not any(c.isdigit() for c in query):
        # No numbers in query, likely not a calculation
        if len(candidates) > 1 and candidates[1][1] > 0.25:
            return candidates[1][0]  # Select second best

    return best_tool


def generate_selection_prompt(
    query: str, candidates: List[Tuple[Tool, float]]
) -> str:
    """Generate the prompt that would be sent to an LLM for final selection."""
    prompt = f'User query: "{query}"\n\nAvailable tools:\n'
    for i, (tool, score) in enumerate(candidates, 1):
        prompt += f"\n{i}. {tool.name}: {tool.description}\n"
        prompt += f"   Parameters: {json.dumps(tool.parameters)}\n"
        prompt += f"   Examples: {tool.examples}\n"
    prompt += "\nSelect the most appropriate tool by name, or respond 'none' if no tool is suitable."
    return prompt
Out[25]:
Console
LLM Selection Simulation:
======================================================================

Query: 'What's 25 divided by 5?'
Prompt that would be sent to LLM:
--------------------------------------------------
User query: "What's 25 divided by 5?"

Available tools:

1. calculator: Evaluate mathematical expressions and perform calculations
   Parameters: {"expression": "string"}
   Examples: ['What is 15 times 23?', 'Calculate the square root of 144']

2. weather: Get current weather conditions and forecasts for locations
   Parameters: {"location": "string", "days": "integer"}
   Examples: ["What's the ...
--------------------------------------------------
Selected tool: calculator


Query: 'Will it rain in Seattle tomorrow?'
Prompt that would be sent to LLM:
--------------------------------------------------
User query: "Will it rain in Seattle tomorrow?"

Available tools:

1. weather: Get current weather conditions and forecasts for locations
   Parameters: {"location": "string", "days": "integer"}
   Examples: ["What's the weather in Paris?", 'Will it rain tomorrow in London?']

2. calendar: Check schedule and create calendar events
   Parameters: {"date": "string", "event": "string"}
   Examples: [...
--------------------------------------------------
Selected tool: weather


Query: 'How do you say 'goodbye' in Spanish?'
Prompt that would be sent to LLM:
--------------------------------------------------
User query: "How do you say 'goodbye' in Spanish?"

Available tools:

1. translate: Translate text between languages
   Parameters: {"text": "string", "target_language": "string"}
   Examples: ["Translate 'hello' to French", "How do you say 'thank you' in Japanese?"]

2. calendar: Check schedule and create calendar events
   Parameters: {"date": "string", "event": "string"}
   Examples: ['Do I hav...
--------------------------------------------------
Selected tool: translate


Query: 'Find me the latest stock prices'
Prompt that would be sent to LLM:
--------------------------------------------------
User query: "Find me the latest stock prices"

Available tools:

1. search: Search for current information on the internet
   Parameters: {"query": "string"}
   Examples: ['Who is the current president?', 'Latest news about AI']

2. weather: Get current weather conditions and forecasts for locations
   Parameters: {"location": "string", "days": "integer"}
   Examples: ["What's the weather in Paris...
--------------------------------------------------
Selected tool: search

The simulation shows how the LLM would receive a focused prompt containing only the top-3 candidate tools with their schemas and examples, letting it to make a fine-grained selection based on fine-grained criteria that pure similarity matching cannot capture, such as detecting when a query lacks required parameters for the top candidate.

This implementation demonstrates the two-stage hybrid approach: fast embedding retrieval narrows the field, then the LLM performs fine-grained discrimination. The prompt construction shows how tool descriptions are formatted for the LLM context, including examples that anchor the model's understanding of appropriate usage boundaries. The generate_selection_prompt function is the necessary bridge between the retrieval stage and the LLM stage: its output quality directly affects the final selection decision.

Key Design Parameters

The key parameters for the hybrid tool router interact in non-obvious ways, and understanding them helps you tune the system for your specific requirements:

  • embedding_model: The sentence transformer model used to encode tool descriptions and queries. all-MiniLM-L6-v2 provides a good balance of speed and accuracy, though domain-specific models may offer better performance for specialized tool ecosystems. Larger models like all-mpnet-base-v2 provide better accuracy at higher latency.
  • top_k: Number of candidate tools to retrieve via embedding similarity. Higher values give the LLM more options but increase context length and latency. Typical values range from 3 to 10, depending on the semantic similarity between tools in your ecosystem. For ecosystems with many similar tools, larger kk reduces the risk of the retriever missing the correct tool.
  • threshold: Minimum similarity score for routing. Below this threshold, no tool is selected, triggering a fallback response. This parameter acts as a confidence gate, preventing the system from invoking tools when user intent is unclear or unrelated to available capabilities. Per-tool thresholds often work better than a single global value.
  • tool_text_construction: The method used to build the text representation for each tool before embedding. Including examples in the embedded text ("calculator: Evaluate math expressions. Examples: What is 2+2? Calculate sqrt(16)") consistently outperforms embedding the description alone, because the examples introduce the query-like vocabulary that will appear in actual routing queries.

Summary

Tool selection turns language models from passive text generators into active agents able to use external capabilities. The mechanisms we explored reveal that effective selection depends on three pillars: rich descriptions that bridge semantic gaps between user intent and tool functionality, routing strategies that balance latency against accuracy, and training methodologies that refine selection instincts through demonstration and feedback.

Key takeaways from this chapter include:

  • Descriptions serve as interfaces: Rich descriptions with examples, negative constraints, and clear boundaries outperform minimal schemas by giving the contextual cues necessary for fine-grained discrimination. Treat descriptions as first-class engineering artifacts, not documentation afterthoughts.

  • Hybrid routing optimizes efficiency: Combining fast embedding-based retrieval with precise LLM-based ranking allows systems to scale to large tool ecosystems without sacrificing accuracy or incurring prohibitive latency. This coarse-to-fine approach mirrors human cognitive strategies for working through complex choice spaces, and typically achieves 90% of LLM-only accuracy at 25-30% of the latency cost.

  • Multi-tool scenarios require planning: Complex tasks demand sequential or parallel tool orchestration, introducing dependencies that selection mechanisms must recognize and respect. Understanding these dependency patterns is important for building systems that can execute workflows rather than isolated actions. The shift from selection to planning is one of the defining challenges of multi-step agent systems.

  • Training improves reliability: While zero-shot selection works for simple cases, fine-tuning on curated examples, synthetic data generation, or reinforcement learning from tool feedback significantly enhances selection reliability, particularly in specialized domains with fine-grained distinctions between similar tools. Systematic data curation, including hard negatives and coverage of edge cases, is as important as the choice of training algorithm.

  • Abstention is needed: A routing system that always selects some tool is worse than one that can confidently decline. Proper confidence thresholding and explicit "none" handling prevent frustrating user experiences caused by irrelevant tool invocations.

  • Scaling requires architecture: Flat routing architectures break down beyond roughly twenty tools without specialized training. Plan for hierarchical routing, dynamic retrieval, and learned routers when building systems that will grow over time.

As we move toward more sophisticated agent architectures in the next chapter, tool selection is the foundational decision layer. The ability to correctly route intent to capability determines whether an AI system acts as a helpful assistant or a frustrating automaton that repeatedly reaches for the wrong instrument. Mastering these selection mechanisms enables the construction of reliable, scalable systems that effectively extend LLM capabilities through external tools.

We will explore how these selection capabilities integrate into broader agent architectures that can reason and plan before executing complex multi-step workflows, building on the routing foundations established here.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about tool selection.

Tool Selection Quiz

Question 1 of 60 of 6 completed
What is the primary function of tool descriptions in LLM-based systems?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026toolselection, author = {Michael Brenndoerfer}, title = {Tool Selection for LLM Agents: Routing Strategies}, year = {2026}, url = {https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Tool Selection for LLM Agents: Routing Strategies. Retrieved from https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies
MLAAcademic
Michael Brenndoerfer. "Tool Selection for LLM Agents: Routing Strategies." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies>.
CHICAGOAcademic
Michael Brenndoerfer. "Tool Selection for LLM Agents: Routing Strategies." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Tool Selection for LLM Agents: Routing Strategies'. Available at: https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Tool Selection for LLM Agents: Routing Strategies. https://mbrenndoerfer.com/writing/tool-selection-llm-agents-routing-strategies

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.