Function Calling: Structured Tool Use for LLMs

Michael BrenndoerferFebruary 2, 202656 min read

Part of Language AI Handbook

Explains how function calling enables LLMs to invoke external tools and APIs through structured JSON schemas, bridging natural language and executable code.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Function Calling

Large language models excel at reasoning and writing as well as analysis, yet they remain confined to the text they were trained on. They cannot check the weather, calculate precise mathematical expressions, query databases, or interact with external APIs. Function calling bridges this gap by letting models to express intent to use external tools in a structured, machine-readable format. Rather than simply describing what a calculator might do, the model outputs a precise JSON object specifying which function to invoke and with what arguments. This structured output turns the model from a passive text generator into an active agent that can retrieve information, perform computations, and affect external systems.

The shift from natural language tool descriptions to structured function calling represents an important architectural decision. Early approaches to tool use relied on in-context learning, where you provided examples of tool usage within the prompt yourself. While effective for simple cases, this approach consumed valuable context window space and suffered from inconsistency as the complexity of available tools increased. Native function calling capabilities, by contrast, bake the understanding of tool schemas directly into the model's weights through specialized fine-tuning. This creates a more reliable, scalable interface, where the model learns to treat tool schemas as first-class citizens in its output space, much like it learns grammar or factual associations.

As we discussed in Tool Use Motivation, the shift from in-context tool demonstrations to native function calling capabilities represents a significant evolution in how language models interface with the world. In this chapter, we examine the mechanics of function calling: how you define tool schemas, how models generate structured calls, how execution results feed back into the generation process, and how to fine-tune models for reliable tool-use capabilities.

Function Calling vs. Tool Use

The terms "function calling" and "tool use" are often used interchangeably, but they carry slightly different connotations. Function calling emphasizes the mechanism: the model emits structured calls with named parameters. Tool use is the broader concept: the model employs external capabilities to answer queries. Every function call is a form of tool use, but tool use includes other patterns like web search, code execution, and memory retrieval that may not follow a strict function-call schema.

Understanding why function calling works requires thinking about the information-theoretic problem it solves. A language model generating free text must simultaneously decide what information to communicate and how to format it for the recipient. When the recipient is a human, natural language is ideal. When the recipient is a software system, natural language is ambiguous and brittle. Function calling resolves this by separating the decision of what to do (natural language understanding) from the specification of how to do it (structured output generation). The model uses its linguistic training to understand intent and its fine-tuned schema knowledge to express that intent unambiguously.

Function Schema Definition

For a language model to invoke external tools, it must first understand what tools are available, what parameters they accept, and what types of values those parameters require. This metadata is provided through function schemas, typically expressed in JSON Schema format, which is a contract between the model and the external environment. The schema acts as a Rosetta Stone translating between the unstructured world of human language and the rigid requirements of software APIs.

Think of a function schema as a documentation standard that a model can read and act on. When a human developer reads API documentation, they build a mental model of what the function does, what inputs it needs, and what outputs to expect. They can then write code that calls the function correctly. The model does something analogous at inference time: it reads the schema, develops a representation of the tool's purpose, and generates calls that conform to the specification. The quality of the schema directly determines how reliably this process works.

JSON Schema Structure

A function schema describes the interface of a callable function using a structured dictionary. At minimum, it specifies the function name, a natural language description of its purpose, and the expected parameters.

In[3]:
Code
weather_schema = {
    "name": "get_weather",
    "description": "Retrieve current weather conditions for a specific location",
    "parameters": {
        "type": "object",
        "properties": {
            "location": {
                "type": "string",
                "description": "The city and state/country, e.g., 'Boston, MA' or 'Paris, France'",
            },
            "unit": {
                "type": "string",
                "enum": ["celsius", "fahrenheit"],
                "description": "Temperature unit to use for the forecast",
            },
        },
        "required": ["location"],
    },
}

The key components of a function schema are:

  • name: A unique identifier for the function using snake_case conventions. This becomes the token the model emits to signal which tool to invoke. The choice of name matters semantically, as models often infer function purpose partially from the name itself. A function named get_current_temperature will likely be invoked for different queries than one named get_historical_weather_data, even if their schemas are similar.
  • description: A clear, imperative explanation of what the function does and when to use it. This text is semantically embedded into the model's context, and heavily influences whether the model chooses to invoke this particular tool. Well-written descriptions act as soft classifiers, helping the model distinguish between similar tools. For instance, if you have both a search_products and get_product_details function, the description should clarify that the former is for finding items matching criteria while the latter requires a specific product ID.
  • parameters: A JSON Schema object defining the function's arguments, including type constraints, valid ranges, and nested object structures. This section effectively constrains the model's output space, limiting what it can generate to syntactically valid structures.
  • required: A list of parameter names that must be provided for the function call to be valid. This creates a hard constraint: if the model attempts to call the function without these parameters, the call fails validation. This forces the model to either extract the necessary information from your query or ask for clarification, rather than hallucinating missing values.

The description field deserves special attention because it operates on two levels simultaneously. At the syntactic level, it tells the model the signature of the function. At the semantic level, it guides the model's judgment about when to invoke the function at all. A description like "Retrieve current weather conditions" implicitly communicates that this function is appropriate for present-tense weather queries but not for historical weather data or general weather science questions. Writing descriptions that calibrate this boundary accurately is one of the most important skills in function calling design.

Type System and Constraints

JSON Schema provides rich typing capabilities that help constrain model outputs and prevent hallucinated parameters. These constraints serve as guardrails that reduce the cognitive load on the model during generation. When a parameter is constrained to an integer type with a minimum value of zero, the model need not consider negative numbers or fractional values, effectively narrowing the search space during token sampling.

The available type primitives and constraint mechanisms include:

  • Primitive types: string, number, integer, boolean, array, object, and null. These basic types align with most programming language type systems, making translation to actual function arguments straightforward.
  • Enum constraints: Restrict string values to a predefined set of options, reducing the probability of invalid values. For example, a color parameter with enum ["red", "green", "blue"] prevents the model from inventing colors like "turquoise" when only primary colors are supported by the underlying API.
  • Array schemas: Define homogeneous arrays with items specifications or heterogeneous tuples with prefixItems. This allows for complex inputs like lists of coordinates or structured records.
  • Nested objects: Support complex parameter structures through recursive properties definitions. This is needed for APIs that require structured data, such as shipping addresses with nested fields for street and city plus a postal code.
  • Validation keywords: minimum, maximum, minLength, maxLength, pattern (regex), and format (email, URI, date-time). These constraints serve dual purposes. During inference, they guide the model toward valid outputs through the semantic cues in parameter descriptions. During structured decoding (as discussed in Constrained Decoding), they can be enforced grammatically to guarantee syntactic validity.

The interplay between these constraints creates a reliable interface. When a model sees a parameter defined as a string with pattern ^[0-9]{5}$, it understands that a zip code is required and that the value must consist of exactly five digits. This semantic signal helps the model extract the correct information from your queries or recognize when it lacks necessary information.

A well-designed schema reduces ambiguity at every level. Numeric ranges communicate expected magnitudes: a page_number parameter with minimum: 1 implicitly tells the model that pages are one-indexed, not zero-indexed. String format constraints communicate the expected structure of textual inputs. Even the choice between number and integer provides semantic information: using integer for a quantity parameter signals that fractional quantities are not meaningful. Every constraint you add is information the model can use to generate more accurate calls.

Schema Design Best Practices

Writing effective schemas is an iterative process that benefits from understanding how the model interprets them. The descriptions you write are not just documentation for human readers; they are the primary signal the model uses to decide whether and how to invoke the tool. Treat every word in a description as meaningful input to the model's decision process.

Several principles consistently produce more reliable schemas. First, be specific about what the function does rather than what it is. "Search products" is weaker than "Search the catalog for products matching keywords, returning up to 20 results sorted by relevance." The latter communicates the function's behavior, expected output format, and appropriate use case. Second, explicitly describe what the function does not do when there is potential for confusion with similar tools. If you have both a search_products and a lookup_product_by_id function, adding "Do not use this function when you have a specific product ID; use lookup_product_by_id instead" to the search function's description prevents the model from using the wrong tool.

Third, use consistent naming conventions across all functions in a toolkit. If some functions use user_id as a parameter name and others use userId or uid, the model may struggle to recognize that they refer to the same concept. Consistency reduces the cognitive overhead the model faces when reading schemas and makes parameter extraction more reliable. Fourth, keep required parameters minimal. Every required parameter is a point of potential failure: if the model cannot extract or infer the value, the call fails. When a parameter can be given a sensible default, make it optional with a default value documented in the description.

Finally, test your schemas with adversarial queries: questions that are adjacent to the function's domain but should not trigger it, questions that require the function but provide information in unexpected formats, and questions that are ambiguous between multiple tools. The cases where the model makes wrong decisions reveal gaps in your schema descriptions that can be addressed iteratively.

Multiple Function Definitions

Real-world applications rarely expose a single tool. Instead, they present the model with a toolkit containing multiple functions. The schemas are provided as a list, and the model must learn to select the appropriate function based on your intent. This selection process mirrors the routing logic in traditional software systems, but happens dynamically based on natural language understanding.

In[4]:
Code
tools = [
    {
        "name": "search_database",
        "description": "Query the product database for items matching criteria",
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string", "description": "Search terms"},
                "category": {
                    "type": "string",
                    "enum": ["electronics", "clothing", "books"],
                },
                "max_price": {
                    "type": "number",
                    "description": "Maximum price filter",
                },
            },
            "required": ["query"],
        },
    },
    {
        "name": "calculate_shipping",
        "description": "Calculate shipping cost based on weight and destination",
        "parameters": {
            "type": "object",
            "properties": {
                "weight_kg": {"type": "number", "minimum": 0},
                "destination_zip": {"type": "string", "pattern": "^[0-9]{5}$"},
                "express": {"type": "boolean", "default": False},
            },
            "required": ["weight_kg", "destination_zip"],
        },
    },
]

When multiple functions are available, the model performs implicit tool selection by emitting the name field corresponding to the appropriate schema. This selection mechanism relies on the semantic alignment between your query, function descriptions, and parameter descriptions. The model essentially performs a form of nearest-neighbor search in embedding space, matching your intent against the semantic content of available tool descriptions.

Consider asking, "How much to ship a 5kg package to Boston?" The model must recognize that while both tools are available, only calculate_shipping addresses the specific question. It extracts "5kg" as the weight parameter, infers the zip code for Boston (or recognizes it needs this specific information), and understands that the express parameter is optional. This discrimination capability requires the model to develop a sophisticated understanding of tool capabilities through training on diverse examples. We'll explore more sophisticated selection strategies in the upcoming chapter on Tool Selection.

The challenge of multi-tool selection scales non-linearly with the number of tools available. With two tools, the selection problem is straightforward. With ten tools, the model must maintain a richer understanding of tool boundaries. With hundreds of tools (as in enterprise environments with dozens of integrated APIs), the selection problem can overwhelm the context window and tax the model's attention. Practical systems often solve this by implementing a two-stage retrieval: first retrieving the most relevant tools using embedding similarity, then presenting only the top-k candidates to the model. This mirrors how human experts work through large toolsets: they first recall which category of tool applies, then select the specific tool within that category.

Out[5]:
Visualization
Line chart showing tool selection accuracy declining from near 100 percent at 1 tool to around 65 percent at 50 tools, with a second line for two-stage retrieval maintaining around 85 percent accuracy across all tool counts.
Illustrative tool selection accuracy as the number of available tools grows. The conceptual values show direct selection degrading faster than a two-stage retrieval approach that first narrows the candidate set. They demonstrate the expected scaling behavior rather than report a measured benchmark.

Function Call Generation

Given a set of function schemas and your query, the model must determine whether to answer directly, request clarification, or emit a structured function call. This decision emerges from the model's training on large corpora of tool-use demonstrations. The generation process represents a specialized form of structured output prediction, where the output must conform to a specific grammar defined by the function schema.

The key insight is that function call generation is not a fundamentally different capability from ordinary text generation: it is the same next-token prediction mechanism applied to a different output distribution. What changes is the target distribution. During pre-training, the model learns to predict the next token in human-written text. During function calling fine-tuning, the model learns to predict the next token in structured, schema-conforming JSON. The architecture is identical; only the training data and target format differ.

Training for Tool Use

Function calling capability is typically instilled through supervised fine-tuning on specialized datasets. These datasets consist of conversation traces where:

  1. The system prompt includes available function definitions
  2. User queries require tool invocation to answer correctly
  3. Assistant responses alternate between function call objects and final answers based on observation results

The training objective remains next-token prediction, but the target distribution now includes structured JSON snippets interleaved with natural language. This multimodal output space requires the model to learn special transition dynamics: when to switch from natural language to JSON, how to maintain syntactic validity across token boundaries, and when to terminate a function call versus continuing with explanation.

Models learn to recognize patterns such as:

  • Queries containing temporal references ("current weather," "latest stock price") likely require tool calls because the model's training data has a cutoff date and cannot provide real-time information.
  • Mathematical precision requirements often necessitate calculator tools, especially for complex expressions where the model might otherwise hallucinate incorrect calculations.
  • Questions about private or real-time data require retrieval functions, as these fall outside the model's parametric knowledge.

The training process also instills tool awareness: the ability to recognize when a query falls within the domain of available tools versus when it requires general knowledge. A poorly trained model might call a weather API for "What's the capital of France?", while a well-calibrated model recognizes this as factual knowledge requiring no external tool.

One subtle aspect of this training is teaching the model to be appropriately uncertain. When a query is ambiguous, the right behavior is often to ask for clarification rather than guess. "What's the weather?" lacks a location, so a well-trained model should respond with a clarifying question rather than inventing a location. Achieving this calibration requires training examples that explicitly demonstrate the clarification behavior, not just tool-use examples.

Generation Formats

Different model families employ varying output formats for function calls, though all share the common thread of structured, parseable text. The choice of format involves trade-offs between human readability, parsing complexity, and semantic clarity.

OpenAI-style function calling uses a specific message role and JSON structure:

{
  "role": "assistant",
  "tool_calls": [
    {
      "id": "call_abc123",
      "type": "function",
      "function": {
        "name": "get_weather",
        "arguments": "{\"location\": \"Boston, MA\", \"unit\": \"celsius\"}"
      }
    }
  ]
}

This format treats function calls as first-class objects in the message hierarchy, with explicit typing and unique identifiers. The nested structure separates the function metadata (name) from the payload (arguments), which makes it easy for downstream parsers to route calls to appropriate handlers.

Anthropic-style tool use embeds XML tags within the text:

<function_calls>
<invoke name="get_weather">
<parameter name="location">Boston, MA</parameter>
<parameter name="unit">celsius</parameter>
</invoke>
</function_calls>

XML formats offer advantages in streaming scenarios, where partial JSON might be invalid but partial XML maintains structure. They also allow for easier human inspection and debugging, as the hierarchical structure is visually apparent through indentation and tag matching.

Generalized JSON mode simply instructs the model to output valid JSON matching the schema, often with special tokens delimiting the start and end of tool calls. This approach maximizes flexibility but requires careful prompt engineering to ensure the model distinguishes between explanatory text and executable code.

Regardless of format, the underlying mechanism relies on the decoder's ability to generate structured text that conforms to a grammar. Building on our discussion of Autoregressive Generation, the model samples tokens conditioned on the function schema context, with the schema effectively biasing the output distribution toward valid JSON structures. The attention mechanism must learn to attend to specific parts of the schema when generating corresponding parameters. This keeps the location value in the output aligns with the location parameter description in the input.

The format choice also has implications for error recovery. JSON formats are all-or-nothing: a single missing bracket or unescaped quote invalidates the entire structure. XML formats are more forgiving in streaming contexts because each field is independently delimited. Some production systems implement constrained decoding, as we discussed in Constrained Decoding, where the decoder's sampling distribution is grammatically restricted to guarantee syntactic validity. This eliminates parse errors entirely but adds implementation complexity and may slightly restrict the model's expressive range.

Parallel Function Calling

Modern implementations support parallel function calling, where the model identifies multiple independent operations that can be executed simultaneously. For example, given the query "Compare the weather in Boston and San Francisco," the model might emit two separate function calls in a single response:

{
  "tool_calls": [
    {"id": "call_1", "function": {"name": "get_weather", "arguments": "{\"location\": \"Boston, MA\"}"}},
    {"id": "call_2", "function": {"name": "get_weather", "arguments": "{\"location\": \"San Francisco, CA\"}"}}
  ]
}

This capability requires the model to recognize functional dependencies and independence between requested operations. In the weather comparison example, the two calls are independent: neither requires the output of the other. However, for a query like "What's the weather in the capital of California?", the model must recognize the dependency chain: first determine the capital (Sacramento), then get its weather. Attempting to parallelize these calls would fail because the second call requires information only available after the first completes.

Parallel calling significantly reduces latency in multi-tool scenarios but requires the execution framework to handle concurrent API calls and aggregate results. The framework must track which calls belong to which logical operation, handle partial failures (where one call succeeds and another fails), and manage race conditions in stateful operations. This pattern essentially turns the linear request-response cycle into a graph execution problem, where the model defines the nodes (function calls) and the system manages the edges (dependencies and data flow).

Out[6]:
Visualization
Grouped bar chart comparing latency in seconds for sequential versus parallel execution across 1 to 5 function calls. Sequential latency grows linearly while parallel latency grows much more slowly.
Illustrative latency comparison for independent function calls under a simple timing model. Sequential execution adds 1.2 seconds per call, while parallel execution adds a smaller coordination overhead; with five calls, the constructed example grows from 1.2 to 6.0 seconds sequentially but remains below 2.0 seconds in parallel. These values are assumptions, not benchmark measurements.

Function Output Handling

Function calling establishes a request-response cycle between the language model and external tools. Once a function executes, its return value must be formatted and presented back to the model to enable the final response generation. This cycle creates a feedback loop where the model can react to real-world data, correcting misconceptions or filling knowledge gaps dynamically.

The design of this feedback loop is more subtle than it first appears. The model must receive the function output and interpret it in the context of the original query. A raw JSON response from a weather API contains temperature, humidity, wind speed, and conditions. The model must select which of these fields are relevant to your question, convert units if requested, and frame the information in natural language that answers the original intent. This interpretation step requires the model to maintain awareness of the original query throughout the tool-execution cycle.

The Observation Pattern

The standard execution loop follows an observe-act pattern with distinct stages. Your query enters the conversation and the model processes it against available schemas. If a tool is warranted, the model emits a structured function call. The application executes the function with the provided arguments and formats the result as an observation message. This observation is appended to the conversation history, and the model generates a final natural language response incorporating the new information.

This pattern mirrors the perception-action cycles found in cognitive architectures and robotics. The model acts (generates a function call), observes the result (receives the tool output), and then acts again (generates a response) based on the updated state. This loop can iterate multiple times for complex queries requiring sequential tool use.

The observation is typically formatted as a tool message (or function result message) containing the function output, often JSON-serialized:

{
  "role": "tool",
  "tool_call_id": "call_abc123",
  "name": "get_weather",
  "content": "{\"temperature\": 22, \"conditions\": \"Partly cloudy\", \"humidity\": 65}"
}

This message structure is necessary because it maintains the conversation state necessary for the model to generate coherent multi-turn interactions. The tool_call_id links the observation back to the specific request, letting the model to handle parallel calls correctly. Without this linkage, the model might confuse which result corresponds to which query, especially when multiple similar tools are invoked simultaneously.

The content of the observation should be structured to maximize the model's comprehension. Raw API responses often contain extraneous metadata, status codes, and internal identifiers. Best practice involves transforming these into clean, semantic representations that highlight the information relevant to your query. For weather data, this might mean extracting just temperature and conditions, while for database queries, it might involve formatting records as readable text or markdown tables.

There is an important design choice here: how much preprocessing to do before returning tool output to the model. Returning raw API responses preserves all information but may confuse the model with irrelevant fields. Returning heavily summarized results is more efficient but risks discarding information the model might need. A practical middle ground is to return the full relevant payload while using clear field names and removing only truly irrelevant metadata (like internal request IDs, rate limit headers, and server timestamps).

Context Window Implications

Each function call and observation consumes tokens in the context window. In scenarios involving multiple tool invocations or verbose API responses, this can quickly exhaust available context length. A complex database query might return hundreds of rows, each consuming dozens of tokens when serialized as JSON. As we explored in Context Length Challenges, long tool outputs may need summarization or selective filtering before being presented to the model.

Strategies for managing context in tool-heavy conversations include:

  • Result summarization: Using smaller models or heuristics to compress verbose API outputs into key points before presenting them to the main model.
  • Pagination: Breaking large result sets into chunks and letting the model to request specific pages or aggregates.
  • Selective retention: Keeping only the most recent tool results in context while archiving older ones to a separate memory store.

The recursive nature of tool use, where observations trigger additional function calls, creates deep conversation histories. A complex research task might involve ten or more tool calls, each adding multiple messages to the thread. Efficient management of this history, potentially through KV Cache Compression or selective context pruning, becomes needed for production deployments. Some systems implement conversation summarization at fixed intervals, condensing older tool interactions into high-level summaries that preserve needed information while freeing up tokens for new operations.

The token cost of function calling is frequently underestimated. A system prompt listing ten detailed tool schemas might consume 2,000 tokens before a single word of your query is processed. This front-loading of context means that function calling applications effectively have shorter usable context windows than their nominal context limit suggests. Designing compact but informative schemas, and using retrieval-based schema selection to present only relevant tools, helps recover this overhead.

Out[7]:
Visualization
Line chart with filled area showing cumulative token count growing from the initial stage through four tool-use iterations to the final response stage, with a dashed red horizontal line marking a typical context limit.
Illustrative cumulative token consumption across function calling iterations. The constructed example starts with 920 tokens, adds 600 tokens per tool-call and observation cycle, and compares the total with an example 4,096-token limit. The values show the accumulation mechanism rather than a universal context-window budget.

Error Handling

Not all function calls succeed. Networks fail, APIs return errors, and arguments may be invalid. The error handling strategy significantly affects reliability and user experience. A brittle system that crashes on API timeouts provides little value, while a resilient system that gracefully degrades maintains utility even under adverse conditions.

Common error handling patterns include:

  • Retry with correction: Present the error to the model and allow it to generate a corrected call. For example, if a weather API returns "Location not found" for "Bostn, MA", the model might infer the typo and retry with "Boston, MA". This requires the error message to be descriptive enough to enable diagnostic reasoning.
  • Fallback to knowledge: If the tool fails, the model falls back to its parametric knowledge with appropriate uncertainty qualifiers. For instance, "I was unable to check the live weather, but based on my training data, Boston in January is typically cold, often below freezing." This maintains utility while signaling uncertainty.
  • Escalation to you: For necessary failures, the system asks you for clarification or manual input. This is appropriate when the model lacks sufficient information to recover autonomously, such as when required authentication tokens expire.

Error messages should be structured to help the model diagnose issues. Rather than generic "Error 500" messages, provide specific feedback: "Invalid location format: expected 'City, State' but received 'Boston'". This specificity enables the model to adjust its approach, perhaps by asking you for clarification or by reformatting the parameter according to the API's expectations. The error format should mirror the success format (JSON with consistent fields) to ensure the model can parse it reliably.

There is also a category of logical errors that are harder to detect: cases where the function executes successfully but returns data that is semantically incorrect for the query. If your query asks about tomorrow's weather and the API returns today's weather, the function call technically succeeded. The model must reason about whether the returned data answers the question. Designing observation formats that include metadata about the query parameters (such as the date range returned by a weather API) helps the model perform this validation.

Worked Example: Multi-Step Calculation

Let's walk through a concrete example involving a calculator tool to illustrate the full function calling cycle.

Scenario: You ask, "If I have 150 apples and give away 30%, then buy 2 dozen more, how many do I have?"

While modern LLMs can solve this arithmetically, we'll assume our model has been instructed to use a calculator for precision. This scenario demonstrates how function calling handles multi-step reasoning where natural language must be translated into mathematical expressions.

Step 1: Schema Definition

In[8]:
Code
calculator_schema = {
    "name": "calculate",
    "description": "Evaluate a mathematical expression safely",
    "parameters": {
        "type": "object",
        "properties": {
            "expression": {
                "type": "string",
                "description": "Mathematical expression to evaluate, e.g., '(150 * 0.7) + 24'",
            }
        },
        "required": ["expression"],
    },
}

Step 2: Initial Query

Your message enters the conversation context along with the available function definitions. The model processes this input through its attention layers, comparing the query against the calculator description to determine that mathematical evaluation is required.

Step 3: Call Generation

The model recognizes the arithmetic nature of the query and generates:

{
  "name": "calculate",
  "arguments": "{\"expression\": \"150 * 0.7 + 24\"}"
}

Note that the model has translated "give away 30%" into the multiplicative factor 0.7 and "2 dozen" into 24. This shows its ability to perform semantic translation into executable parameters. This translation step shows the model's role as an interface layer between human communication patterns and machine-executable instructions. The model must understand that "giving away 30%" leaves 70%, and that a "dozen" equals twelve. This makes two dozen equal twenty-four.

The translation is non-trivial. The model must correctly chain two operations: first apply the percentage reduction, then add the new quantity. Expressing this as a single expression 150 * 0.7 + 24 requires understanding operator precedence (multiplication before addition) and recognizing that the 30% reduction and the subsequent purchase are independent operations on the running count.

Step 4: Execution

The application layer parses the JSON, validates that the expression contains only safe mathematical operations (preventing code injection), and evaluates:

In[9]:
Code
import ast
import operator


def safe_calculate(expression):
    """Safely evaluate mathematical expression using AST parsing.
    All nodes are validated against an allowlist before any evaluation occurs."""
    allowed_operators = {
        ast.Add: operator.add,
        ast.Sub: operator.sub,
        ast.Mult: operator.mul,
        ast.Div: operator.truediv,
        ast.USub: operator.neg,
    }

    def eval_node(node):
        if isinstance(node, ast.Num):  # Python 3.7
            return node.n
        elif isinstance(node, ast.Constant):  # Python 3.8+
            return node.value
        elif isinstance(node, ast.BinOp):
            op_type = type(node.op)
            if op_type in allowed_operators:
                return allowed_operators[op_type](
                    eval_node(node.left), eval_node(node.right)
                )
        elif isinstance(node, ast.UnaryOp):
            if isinstance(node.op, ast.USub):
                return -eval_node(node.operand)
        raise ValueError(f"Unsupported operation: {type(node)}")

    tree = ast.parse(expression, mode="eval")
    result = eval_node(tree.body)
    return result


result = safe_calculate("150 * 0.7 + 24")
Out[10]:
Console
Calculation result: 129.0

The calculation yields 129.0 apples, confirming that after giving away 30% (45 apples) from the original 150 and adding 24 more, you have 129 apples remaining.

Step 5: Observation Injection

The result 129.0 is formatted as a tool message and appended to the context. This injection step is transparent to you but important for the model's reasoning chain. The observation provides grounding. This keeps the final answer is based on actual computation rather than estimation.

Step 6: Final Response

The model now generates: "You would have 129 apples."

This example demonstrates the complete loop: intent recognition, parameter extraction, safe execution, and natural language synthesis based on structured data. It illustrates how function calling creates a division of labor: the model handles language understanding and translation, while specialized tools handle precise computation. This separation of concerns allows each component to excel at what it does best, combining the linguistic flexibility of LLMs with the computational accuracy of traditional software.

The example also illustrates why safe execution matters. The safe_calculate function uses Python's AST module to parse the expression before evaluating it, explicitly checking that only arithmetic operations are present. The entire expression tree is walked and validated against an allowlist before any computation occurs. This architecture is a standard pattern for secure expression evaluation in function calling contexts.

Code Implementation: Building a Function Calling System

Let's implement a complete function calling pipeline using Python. We'll create a mock LLM interface that simulates generation, then build the execution framework around it. This implementation demonstrates the architectural patterns used in production systems, abstracted for clarity.

The architecture we build here separates three distinct concerns. The tool registry manages the available capabilities: it stores schemas (the contract with the LLM) and implementations (the actual executable code). The mock LLM encapsulates the decision logic: which tool to call and with what arguments. The agent orchestrates the interaction: it maintains conversation state, sends queries to the LLM, executes tool calls, and feeds results back into the conversation. In a production system, the mock LLM would be replaced by real API calls, but the registry and agent patterns remain unchanged. This separation of concerns is what makes function calling systems maintainable as they grow: you can swap out the LLM backend without touching the tool implementations, or add new tools without modifying the orchestration logic.

In[11]:
Code
from dataclasses import dataclass
from typing import Any, Dict, List


@dataclass
class FunctionCall:
    name: str
    arguments: Dict[str, Any]
    call_id: str


@dataclass
class Message:
    role: str
    content: str = None
    tool_calls: List[FunctionCall] = None
    tool_call_id: str = None
    name: str = None

First, we define a registry for available tools. This registry maps function names to implementations and maintains their schemas. The registry pattern centralizes tool management. This provides a single source of truth for both the schemas (consumed by the LLM) and the implementations (consumed by the execution environment).

In[12]:
Code
from typing import Callable


class ToolRegistry:
    def __init__(self):
        self.tools: Dict[str, Dict] = {}
        self.implementations: Dict[str, Callable] = {}

    def register(self, name: str, schema: Dict, implementation: Callable):
        self.tools[name] = schema
        self.implementations[name] = implementation

    def get_schemas(self) -> List[Dict]:
        return [
            {"type": "function", "function": schema}
            for schema in self.tools.values()
        ]

    def execute(self, call: FunctionCall) -> Any:
        if call.name not in self.implementations:
            raise ValueError(f"Unknown function: {call.name}")

        func = self.implementations[call.name]
        try:
            result = func(**call.arguments)
            return {"status": "success", "result": result}
        except Exception as e:
            return {"status": "error", "error": str(e)}

The execute method wraps the function call in a try-except block and returns a structured result with a status field. This pattern ensures that even when tools fail, the return value is parseable JSON that the model can reason about. A failed call that returns {"status": "error", "error": "Location not found"} gives the model actionable information; a failed call that raises an unhandled exception gives it nothing.

Now we implement the tools. We'll create a weather lookup and a calculator, simulating the weather API while implementing the calculator. This mixed approach is typical in development environments, where some tools connect to real services while others use mocks or local implementations.

In[13]:
Code
def mock_weather_api(location: str, unit: str = "celsius") -> dict:
    """Simulated weather API for demonstration"""
    weather_db = {
        "boston, ma": {
            "temp": 22,
            "condition": "Partly cloudy",
            "humidity": 65,
        },
        "san francisco, ca": {"temp": 18, "condition": "Foggy", "humidity": 80},
        "london, uk": {"temp": 15, "condition": "Rainy", "humidity": 90},
    }

    key = location.lower().strip()
    if key not in weather_db:
        return {"error": f"No data available for {location}"}

    data = weather_db[key]
    if unit == "fahrenheit":
        data = {**data, "temp": data["temp"] * 9 / 5 + 32}

    return data


def calculator_tool(expression: str) -> float:
    """Safe calculator using AST-based expression evaluation.
    All nodes are validated against an allowlist before any evaluation occurs."""
    allowed_ops = {
        ast.Add: operator.add,
        ast.Sub: operator.sub,
        ast.Mult: operator.mul,
        ast.Div: operator.truediv,
        ast.USub: operator.neg,
    }

    def eval_node(node):
        if isinstance(node, ast.Constant):
            if not isinstance(node.value, (int, float)):
                raise ValueError("Only numeric constants allowed")
            return node.value
        elif isinstance(node, ast.BinOp):
            op_type = type(node.op)
            if op_type not in allowed_ops:
                raise ValueError(f"Disallowed operator: {op_type.__name__}")
            return allowed_ops[op_type](
                eval_node(node.left), eval_node(node.right)
            )
        elif isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.USub):
            return -eval_node(node.operand)
        raise ValueError(f"Disallowed node type: {type(node).__name__}")

    tree = ast.parse(expression, mode="eval")
    return eval_node(tree.body)

We instantiate the registry and register our tools:

In[14]:
Code
registry = ToolRegistry()

weather_schema = {
    "name": "get_weather",
    "description": "Get current weather for a location",
    "parameters": {
        "type": "object",
        "properties": {
            "location": {
                "type": "string",
                "description": "City and state/country",
            },
            "unit": {
                "type": "string",
                "enum": ["celsius", "fahrenheit"],
                "default": "celsius",
            },
        },
        "required": ["location"],
    },
}

calc_schema = {
    "name": "calculate",
    "description": "Calculate mathematical expressions",
    "parameters": {
        "type": "object",
        "properties": {
            "expression": {
                "type": "string",
                "description": "Math expression to evaluate",
            }
        },
        "required": ["expression"],
    },
}

registry.register("get_weather", weather_schema, mock_weather_api)
registry.register("calculate", calc_schema, calculator_tool)

For demonstration purposes, we'll simulate the LLM's function calling behavior with rule-based responses. In practice, this would be replaced with actual API calls to GPT-4, Claude, or open-source models with function calling capabilities. This mock implementation illustrates the expected interface: the LLM receives conversation history and returns either a text response or a structured tool call.

In[15]:
Code
import re


class MockLLM:
    """Simulates an LLM with function calling capabilities"""

    def __init__(self, registry: ToolRegistry):
        self.registry = registry
        self.call_count = 0

    def generate(self, messages: List[Message]) -> Message:
        """Simulate generation based on message history"""
        last_message = messages[-1]
        content = last_message.content.lower() if last_message.content else ""

        # Simple pattern matching to simulate tool use decisions
        if "weather" in content:
            self.call_count += 1
            # Extract location with simple regex
            match = re.search(r"weather\s+(?:in|for)?\s+(.+?)(?:\?|$)", content)
            if match:
                location = match.group(1).strip()
                return Message(
                    role="assistant",
                    tool_calls=[
                        FunctionCall(
                            name="get_weather",
                            arguments={"location": location},
                            call_id=f"call_{self.call_count}",
                        )
                    ],
                )

        elif any(
            word in content
            for word in ["calculate", "compute", "sum", "product", "how many"]
        ):
            self.call_count += 1
            # Look for math expressions
            numbers = re.findall(r"\d+", content)
            if len(numbers) >= 2 and "sum" in content:
                return Message(
                    role="assistant",
                    tool_calls=[
                        FunctionCall(
                            name="calculate",
                            arguments={
                                "expression": f"{numbers[0]} + {numbers[1]}"
                            },
                            call_id=f"call_{self.call_count}",
                        )
                    ],
                )

        # Default response if no tool needed
        return Message(
            role="assistant", content="I don't need any tools to answer this."
        )

Now we implement the execution loop that orchestrates the interaction between the LLM and tools. This agent class encapsulates the state management and control flow, maintaining the conversation history and managing the iterative cycle of generation and execution.

In[16]:
Code
import json


class FunctionCallingAgent:
    def __init__(
        self, llm: MockLLM, registry: ToolRegistry, max_iterations: int = 5
    ):
        self.llm = llm
        self.registry = registry
        self.max_iterations = max_iterations
        self.conversation: List[Message] = []

    def run(self, user_query: str) -> str:
        """Execute the full function calling loop"""
        # Add system message with tool descriptions
        system_msg = f"You have access to the following tools: {json.dumps(self.registry.get_schemas())}"
        self.conversation = [Message(role="system", content=system_msg)]
        self.conversation.append(Message(role="user", content=user_query))

        for iteration in range(self.max_iterations):
            # Generate response
            response = self.llm.generate(self.conversation)

            # Check if tool calls were made
            if response.tool_calls:
                self.conversation.append(response)

                # Execute all tool calls (parallel execution)
                for call in response.tool_calls:
                    result = self.registry.execute(call)

                    # Add observation to conversation
                    obs_msg = Message(
                        role="tool",
                        content=json.dumps(result),
                        tool_call_id=call.call_id,
                        name=call.name,
                    )
                    self.conversation.append(obs_msg)

                # Continue loop for final generation
                continue
            else:
                # No tool calls, return final answer
                return response.content

        return "Maximum iterations reached without final answer."
In[17]:
Code
agent = FunctionCallingAgent(MockLLM(registry), registry)
agent_result = agent.run("What's the weather in Boston, MA?")
Out[18]:
Console

Final result: I don't need any tools to answer this.

The execution trace shows the complete cycle: the agent recognizes the weather query, extracts the location parameter, executes the mock API, and would typically generate a final response (in this simulation, the mock LLM returns a placeholder). This architecture separates concerns cleanly: the agent manages state and orchestration, the registry handles tool management, and the LLM handles decision-making. Such separation simplifies extension and testing as well as maintenance of production function calling systems.

The max_iterations guard deserves particular attention. Without it, a malfunctioning model that repeatedly generates tool calls without ever creating a final answer would loop forever, consuming API credits and blocking other requests. The iteration limit is a safety valve, not an expected boundary condition. Well-calibrated models almost never reach it. When they do, it is often a signal that the query is ambiguous, the tool outputs are confusing the model, or the model is stuck in a pattern matching loop. Logging iteration counts in production helps identify these pathological cases.

Key Parameters

The key parameters governing a function calling system's behavior are:

  • max_iterations: Maximum number of tool invocation cycles allowed before terminating. This prevents infinite loops in cases where the model repeatedly invokes tools without reaching a final answer. Setting this parameter requires balancing completeness against resource constraints. A value of 3 to 5 is typically sufficient for most use cases, while complex research tasks might require 10 or more. Exceeding this limit usually indicates either a poorly calibrated model stuck in a loop or a query that requires more steps than the system allows.
  • tool_call_id: Unique identifier linking each observation back to its corresponding function call, needed for handling parallel function calls correctly. In production systems, these IDs must be globally unique and persistent across the conversation lifecycle to ensure proper attribution of results. UUIDs or timestamped identifiers are commonly used.
  • required: List of parameter names in the function schema that must be provided for the call to be valid. This keeps necessary arguments are not omitted. This list acts as a contract enforcement mechanism. When the model attempts to call a function without a required parameter, the system should reject the call and either prompt the model for the missing information or return a validation error that the model can use to correct its approach.

Function Calling Fine-Tuning

While general-purpose models like GPT-4 come with reliable function calling capabilities, open-source models often require fine-tuning to reliably emit structured tool calls. This process involves curating datasets of tool-use conversations and optimizing the model for schema adherence. Fine-tuning turns a base model from a general text generator into a specialized tool-calling agent.

The need for fine-tuning is about more than format compliance. A base model might generate syntactically valid JSON that is semantically wrong: it might call the right function but with hallucinated parameter values, or invoke a tool when the answer is already in its parametric knowledge. Fine-tuning teaches the model the output format and the judgment of when and how to invoke tools correctly. This judgment is fundamentally a classification problem: for every query, the model must decide among (a) answer directly, (b) request clarification, (c) call one or more tools.

Dataset Construction

Fine-tuning data for function calling follows a conversation format where each example consists of:

  1. System Prompt: Defines available functions and general behavior guidelines
  2. User Messages: Queries that require tool invocation
  3. Assistant Messages: Either function calls (in the structured format) or final answers
  4. Tool Results: Observations returned from executed functions
  5. Final Answers: Natural language responses incorporating tool outputs

A single training example might look like this in ShareGPT format:

{
  "conversations": [
    {"from": "system", "value": "You are a helpful assistant with access to tools..."},
    {"from": "human", "value": "What's the weather in Paris?"},
    {"from": "gpt", "value": "<tool_call>{\"name\": \"get_weather\", \"arguments\": {\"location\": \"Paris\"}}</tool_call>"},
    {"from": "tool", "value": "{\"temperature\": 20, \"condition\": \"Sunny\"}"},
    {"from": "gpt", "value": "The weather in Paris is sunny with a temperature of 20 degrees Celsius."}
  ]
}

The necessary aspect is maintaining the exact output format expected at inference time. If the model is expected to wrap tool calls in <tool_call> tags, the training data must include these tags consistently. Inconsistency between training and inference formats leads to parsing errors and failed tool invocations. Additionally, training datasets should include diverse examples covering edge cases: tool calls with nested parameters, parallel invocations, error recoveries where the model must retry with corrected parameters, and clarifications where the model asks for missing required fields.

Negative examples are equally important. The dataset should include instances where the model should not call a tool, instead answering from its parametric knowledge. This prevents over-reliance on external tools and reduces unnecessary API calls and latency.

The diversity of the dataset matters as much as its size. A dataset with 10,000 examples covering only simple single-tool queries will produce a model that fails on multi-tool queries, ambiguous queries, and error recovery scenarios. A well-constructed dataset of 2,000 examples spanning all behavioral categories will generalize better. Think of the fine-tuning data as a specification of the system's behavioral contract: every pattern you want the model to handle correctly must appear somewhere in the training data.

Constructing high-quality fine-tuning data is expensive because it requires human annotation of correct tool-use behavior. Some teams have found success with synthetic data generation: using a powerful model (like GPT-4) to generate diverse queries and then annotating the correct tool calls programmatically, then using a weaker model to filter out low-quality examples. This approach can produce large, diverse datasets at reasonable cost but requires careful quality control to prevent the fine-tuned model from inheriting the generator model's failure modes.

Out[19]:
Visualization
Bar chart showing the percentage breakdown of four conversation types in a fine-tuning dataset: tool-required queries at 45 percent, direct answer queries at 30 percent, clarification requests at 15 percent, and error recovery at 10 percent.
Illustrative composition of a function calling fine-tuning dataset. The conceptual mix assigns 45 percent to tool-required queries while retaining direct-answer, clarification, and error-recovery examples to teach calibrated tool use. It is a design example rather than a measured dataset distribution.

Training Objectives

The fine-tuning process uses standard supervised learning with cross-entropy loss, but with attention to special tokens that demarcate tool boundaries. As discussed in Instruction Tuning Training, the learning rate is typically lower than pre-training to preserve general capabilities while adapting to the specific tool-calling format. The model must learn the format and the semantics of when to transition between natural language and structured calls.

Key hyperparameters for function calling fine-tuning include:

  • Learning rate: Often 1e-5 to 2e-5, smaller than general fine-tuning to maintain stability. Higher rates might cause catastrophic forgetting of conversational abilities or overfitting to specific tool formats.
  • Sequence length: Must accommodate long tool descriptions and multi-turn conversations. Function schemas can be verbose, especially with nested objects, requiring context windows of 4k tokens or more.
  • Masking strategy: Typically, only assistant messages (including tool calls) are used for loss computation, while user and system prompts are masked. This focuses the model's learning on generating appropriate responses rather than predicting your inputs.
  • Data mixture: Combining tool-use data with general instruction data prevents catastrophic forgetting of conversational abilities, as explored in Catastrophic Forgetting. A typical mixture might be 70% tool-use data and 30% general instruction following data, though this varies by base model and target application.

The masking strategy deserves a closer look. When computing the cross-entropy loss, we want the model to learn to generate correct tool calls and final responses. We do not want it to expend capacity predicting the user's messages or the system prompt. By masking these tokens out of the loss computation (setting their weights to zero), we ensure the gradient signal comes entirely from the model's own outputs. This is the same technique used in instruction tuning more broadly and is important for efficient learning.

The training process must also account for the observation phase. Some implementations include simulated tool results in the training data, while others train only on the call generation phase and handle observations as context during inference. The former approach creates a more complete understanding of the tool-use cycle but requires a dataset of realistic tool outputs. Including observations in training is generally preferred because it teaches the model how to synthesize tool results into natural language responses, which is the final step that makes function calling useful to users.

Format-Specific Considerations

Different model families require different fine-tuning approaches based on their native chat templates and special token vocabularies.

ChatML format (used by many open-source models) structures tool calls as separate message types:

<|im_start|>system You have access to the following functions... <|im_end|> <|im_start|>user What's the weather? <|im_end|> <|im_start|>assistant <tool_call>{"name": "get_weather", "arguments": {"location": "Boston"}}</tool_call> <|im_end|> <|im_start|>tool {"temperature": 22} <|im_end|>

Llama-2/3 chat formats use specific header tokens to distinguish between tool planning and final responses, often requiring the model to generate thinking tokens or plan steps before emitting the final JSON. These formats may require the model to first output reasoning in special tags, then the tool call, creating a chain-of-thought style approach that improves accuracy on complex multi-step problems.

The fine-tuning objective trains the model to recognize when its internal knowledge is insufficient (triggering a tool call) versus when it can answer directly. This calibration requires diverse training examples including:

  • Tool-required queries: Questions that cannot be answered without external data, such as current weather or stock prices
  • Direct answer queries: Questions within the model's parametric knowledge to prevent over-reliance on tools, such as historical facts or general knowledge
  • Clarification requests: Ambiguous queries where the model should ask for parameter specification rather than hallucinate values, training the model to recognize uncertainty and request your input

The boundary between these categories is not always sharp, which is precisely what makes calibration difficult. "What's the population of Tokyo?" can be answered from parametric knowledge with reasonable accuracy (the model knows it's around 14 million in the city proper, 37 million in the greater metropolitan area) but might warrant a database lookup for an application requiring precise current figures. The right behavior depends on the application's accuracy requirements, which should be communicated through the system prompt.

Evaluation Metrics

Validating function calling models requires specialized metrics beyond standard perplexity or BLEU scores. These metrics assess both the syntactic correctness and semantic appropriateness of tool use:

  • Schema adherence rate: Percentage of generated calls that parse as valid JSON and match the defined schema. This is a hard constraint; even minor syntax errors render the call unusable.
  • Parameter accuracy: Correctness of extracted parameters (e.g., did the model correctly identify "Boston" as the location). This measures the model's ability to map natural language entities to structured fields.
  • Tool selection accuracy: Frequency with which the model chooses the correct function from multiple options. This assesses the model's semantic understanding of tool purposes and boundaries.
  • False positive rate: How often the model invokes tools unnecessarily for answerable questions. High false positive rates indicate poor calibration and lead to wasted computational resources.
  • Hallucination rate: Frequency of invented parameter values or non-existent function calls. This measures the model's tendency to confabulate when faced with uncertainty.

These metrics are typically evaluated on held-out test sets with gold-standard annotations showing the correct tool sequence for each query. Evaluation should include adversarial examples designed to trick the model into incorrect tool selection or parameter hallucination. This tests resilience against edge cases.

The interaction between these metrics creates interesting trade-offs. Aggressive training on tool-required examples reduces the false positive rate (the model learns that direct questions should be answered directly), but may reduce schema adherence if the training data lacks sufficient format diversity. Evaluating on all metrics simultaneously and tracking their correlations helps identify which aspects of the training data need reinforcement.

Out[20]:
Visualization
Grouped bar chart comparing base model and fine-tuned model accuracy scores across four evaluation metrics: schema adherence, parameter accuracy, tool selection, and no-hallucination rate.
Illustrative function calling evaluation metrics for a base and fine-tuned model. The conceptual scores show improvement across schema adherence, parameter accuracy, tool selection, and no-hallucination rate; they demonstrate how a comparison can be reported and are not results from a named empirical benchmark.

Structured Output Generation and Schema Enforcement

Function calling and structured output generation are closely related but distinct capabilities. Function calling specifically refers to the model emitting a call to a named function with arguments. Structured output generation is the broader capability of emitting text that conforms to any specified schema, which includes function calls but also encompasses JSON objects, XML documents, CSV records, and other structured formats.

Understanding this relationship matters because the same techniques that enable function calling also underlie many other LLM capabilities. When a model generates a structured resume from a job description, or populates a database template from unstructured text, or extracts named entities in a specified JSON format, it is using the same structural generation machinery as function calling, just pointed at different schemas.

Token-Level Schema Enforcement

At the token generation level, two broad approaches enforce schema conformance. The first is training-time enforcement, where the fine-tuning data exclusively contains valid schema-conforming examples, and the model learns to generate valid structures probabilistically. The second is inference-time enforcement, where the decoder's sampling distribution is dynamically constrained to only allow tokens that maintain schema validity at each step.

Training-time enforcement is simpler to implement but produces stochastic compliance: the model usually generates valid structures but occasionally produces invalid ones, especially on unusual inputs or when under-specified schemas leave ambiguity. The failure rate depends on the quality and diversity of training data and typically improves with more training.

Inference-time enforcement, as covered in Constrained Decoding, guarantees syntactic validity by construction. The system maintains a parser state that tracks which tokens are currently valid according to the schema grammar, and sets the logit of all invalid tokens to −∞-\infty before sampling. This approach is computationally more expensive but eliminates parsing errors entirely. For production systems where unparseable outputs represent hard failures, the overhead is generally worth it.

The trade-off between these approaches also affects the model's reasoning quality. Inference-time constraints can sometimes prevent the model from generating the contextually best argument value if that value happens to violate a syntactic constraint mid-generation. Training-time enforcement, while noisier, allows the model more freedom to express its full linguistic capability and may produce more semantically accurate outputs even if some are syntactically invalid. Hybrid approaches train the model for high baseline compliance and apply selective inference-time constraints only for the most necessary schema fields (like function names and required parameters).

Structured Output Beyond Function Calling

The same principles that govern function calling apply when you need LLMs to produce structured data in other contexts. Consider information extraction: given a long article about a company's earnings report, you want to extract a structured record with fields for revenue, profit, year-over-year growth, and key executive commentary. You define a JSON schema for the record, provide it to the model alongside the article, and prompt the model to populate the schema.

The schema serves the same roles it does in function calling: it constrains the output space, provides semantic guidance about expected field types, and gives the model a structured template to fill. The difference is that instead of generating a function invocation, the model is generating a data record. The underlying mechanism is identical, which is why models with strong function calling capabilities tend to excel at structured extraction tasks as well.

This generalization suggests a productive mental model: think of JSON schemas not as function call specifications but as output templates that guide what the model generates. Any time you want the model to produce structured information, defining a schema is likely to improve reliability and consistency. The model has been trained to treat schemas as authoritative specifications of its expected output format.

Limitations and Impact

Function calling represents a significant advance in language model capabilities, yet it introduces specific constraints and failure modes that practitioners must manage. Understanding these limitations is important for deploying reliable systems and setting appropriate user expectations.

Schema Rigidity and Generalization

Models trained on function calling often struggle with schema variations not seen during fine-tuning. If a training set always presents weather queries with "city, state" format, the model may fail when encountering "city, country" or coordinate-based locations. This brittleness contrasts with the flexibility of natural language, requiring careful prompt engineering or few-shot examples to handle edge cases. The model essentially overfits to the specific parameter patterns in its training data, lacking the compositional generalization that humans exhibit when encountering novel but semantically equivalent inputs.

Models may also exhibit parameter hallucination when facing ambiguous queries. Given "What's the weather like?" without a specified location, a poorly calibrated model might guess a location rather than ask for clarification. This necessitates reliable validation layers that check required parameters before execution and implement retry loops with explicit error feedback. Some systems implement "guardian" models that pre-screen function calls for hallucinated parameters, adding a layer of safety at the cost of additional latency.

The hallucination problem is particularly challenging for string parameters that are not constrained by enums or patterns. A model generating a customer_id parameter has no schema-level guardrail preventing it from inventing a plausible-looking but non-existent ID. The only defense is application-layer validation that checks whether the generated ID exists in the system before executing the function. This validation step adds latency but prevents silent errors where the model's output looks valid but refers to non-existent entities.

Latency and Cost Considerations

The function calling loop introduces multiple round-trips to the language model: initial call generation, observation injection, and final response generation. Each trip incurs latency and API costs. In high-throughput applications, these costs compound rapidly. A query requiring three tool calls might take three times as long and cost three times as much as a single-turn response.

Parallel function calling mitigates some latency concerns by batching independent operations, but sequential dependencies (where one tool's output is needed for the next call) require serial execution. Planning strategies, which we'll explore in the ReAct Pattern chapter, help optimize these dependency chains by letting the model to reason about tool dependencies before execution. Caching frequent tool results and using smaller, faster models for simple parameter extraction can also reduce costs in production environments.

The cost model for function calling applications also includes the token overhead of schema definitions in the system prompt. For applications with dozens of available tools, the system prompt alone might cost hundreds of tokens per request. At scale, these schema tokens represent a substantial fraction of total API costs. Techniques like schema compression (removing verbose descriptions while preserving needed type information), retrieval-based schema selection (presenting only the most relevant tools per query), and schema caching (reusing parsed schema representations across requests) help manage this overhead.

Security Implications

Function calling creates a direct channel between unconstrained natural language and executable code. This presents significant security risks that must be addressed through defense in depth:

  • Prompt injection: Malicious users might craft queries that invoke tools with dangerous arguments. This is analogous to SQL injection in web applications, where user input is executed as code.
  • Tool confusion: Attackers may describe tools in ways that trick the model into invoking privileged functions with attacker-controlled parameters. Renaming a sensitive function to resemble a benign one might fool the model into using it inappropriately.
  • Information leakage: Tool outputs might contain sensitive data that gets exposed in subsequent generations. A database query tool might return private customer information that the model then reveals to unauthorized parties.

Mitigation strategies include strict input validation against schemas, allowlisting permitted operations, executing tools in sandboxed environments, and implementing human-in-the-loop confirmation for destructive operations. The principle of least privilege should govern tool design: functions should expose minimal capabilities necessary for their specific purpose. Additionally, output filtering can prevent sensitive data from being returned to you, and rate limiting can prevent abuse of expensive or sensitive tools.

The prompt injection threat is particularly insidious because it can be embedded in tool outputs, not just user inputs. If a web search tool returns a result containing text that resembles instructions, the model might treat this as a legitimate system command. Sanitizing tool outputs to remove instruction-like patterns and maintaining clear separation between data and instructions in the conversation format are important defenses.

The Path to Agency

Despite these limitations, function calling turns language models from passive responders into active agents able to affect the world. By giving a structured interface to external capabilities, it enables the construction of complex workflows where models plan and execute, then observe and adapt. This capability represents the foundation of autonomous agent architectures, where the LLM is the reasoning engine coordinating multiple tools to achieve high-level goals.

This capability enables applications ranging from automated research assistants that query databases and calculate statistics to customer service bots that modify orders and check shipping status. As we move toward more sophisticated agent architectures in subsequent chapters, function calling is the basic building block, the "motor cortex" that translates high-level intentions into concrete actions. The ReAct Pattern builds directly on this foundation, adding explicit reasoning steps between observations and actions to create more reliable, traceable agent behaviors. Understanding function calling is therefore needed for anyone building the next generation of AI applications that interact with the real world.

Summary

Function calling enables language models to bridge the gap between natural language understanding and actionable computation. By defining tool schemas in JSON format, models learn to emit structured function calls with appropriate parameters, execute external code safely, and synthesize results into coherent responses. This capability turns static text generators into dynamic agents able to retrieving information, performing calculations, and interacting with external systems.

Key takeaways include:

  • Schema definition uses JSON Schema to specify function interfaces, parameter types, and constraints. This provides the model with the metadata necessary to generate valid calls. Well-designed schemas act as both technical specifications and semantic guides, helping the model understand when and how to use each tool.
  • Call generation uses supervised fine-tuning on tool-use corpora, teaching models to recognize when external data is required and how to format requests appropriately. This training instills tool awareness and the ability to map natural language queries to structured parameters.
  • Execution loops follow an observe-act pattern where function results are injected back into the context window as observations, letting multi-step reasoning. This cycle allows models to react to real-world data and correct course based on external feedback.
  • Fine-tuning requires carefully curated conversation datasets that demonstrate proper tool selection, parameter extraction, and natural language synthesis from structured data. Success requires attention to format consistency, diverse examples, and proper evaluation metrics.
  • Practical deployment demands attention to error handling, security validation, latency optimization, and context window management. Production systems must guard against injection attacks, handle API failures gracefully, and manage the computational costs of multi-turn interactions.

As the interface between language models and the digital world, function calling represents the key primitive upon which autonomous agents are constructed. Mastering its mechanics, both the mathematical foundations of structured generation and the engineering challenges of safe execution, prepares practitioners to build systems that move beyond text generation into purposeful action. The skills developed here, designing schemas, managing execution loops, and handling edge cases, form the bedrock of modern AI application development.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about function calling in large language models.

Function Calling Fundamentals

Question 1 of 70 of 7 completed
What is the primary purpose of the 'required' field in a function schema?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026functioncalling, author = {Michael Brenndoerfer}, title = {Function Calling: Structured Tool Use for LLMs}, year = {2026}, url = {https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Function Calling: Structured Tool Use for LLMs. Retrieved from https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools
MLAAcademic
Michael Brenndoerfer. "Function Calling: Structured Tool Use for LLMs." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools>.
CHICAGOAcademic
Michael Brenndoerfer. "Function Calling: Structured Tool Use for LLMs." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Function Calling: Structured Tool Use for LLMs'. Available at: https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Function Calling: Structured Tool Use for LLMs. https://mbrenndoerfer.com/writing/function-calling-llm-structured-tools

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.