Instruction Format: Chat Templates and Roles

Michael BrenndoerferDecember 18, 202556 min read

Part of Language AI Handbook

Explains how chat templates, prompt formats, and role definitions structure conversations for language model instruction tuning and reliable inference.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Instruction Format

When you fine-tune a language model to follow instructions, the instruction format you use to structure those instructions matters enormously. A model trained on one format will struggle to understand inputs presented in a different format at inference time. This chapter explains how templates and their special-token conventions transform raw text into structured conversations that language models can reliably interpret.

The core challenge is deceptively simple: language models process sequences of tokens, but we want them to understand the distinction between user instructions and assistant responses. We need consistent delimiters that tell the model "this is an instruction" versus "this is the expected response." These delimiters must be unambiguous, consistent during training and inference, and reliable enough to handle complex multi-turn conversations. Consistent delimiters are a prerequisite for instruction tuning to work.

Think of an instruction format as the grammar of a conversation protocol. Just as HTTP headers must follow a precise format for a web server to parse a request correctly, language model conversations must follow a precise structural convention for the model to parse roles and content reliably. The model learns this grammar from tens of thousands of training examples. If you present data in a slightly different format at inference time, you are effectively speaking a dialect the model has never studied, and the resulting degradation in output quality can be large.

As we discussed in the Instruction Following chapter, instruction-tuned models learn to recognize patterns in how instructions are presented. The format is a structural scaffold that helps the model parse inputs correctly and generate outputs in the expected positions. Without consistent formatting, the model cannot reliably distinguish between instruction and response, between the user and the assistant, or between system context and conversation content. Every structural irregularity introduces ambiguity that the model must resolve using context alone, and context alone is often insufficient.

Format design also has deeper implications for what capabilities emerge. The presence of a system-message role, for instance, makes it possible to deploy one base model across thousands of applications, each with its own persona and behavioral constraints. The way loss masking interacts with multi-turn format determines whether the model learns to generate coherent conversations or merely to predict isolated responses. These design choices shape the character of the resulting model in ways that go far beyond superficial presentation.

Historical Context

Early instruction-tuning efforts like FLAN (2022) simply prepended task descriptions to inputs without any structural delimiter framework. The Alpaca paper (2023) introduced the three-part ### Instruction: / ### Input: / ### Response: template that became the de facto standard for early open-source efforts. OpenAI then introduced ChatML internally, which leaked into the open-source community through models like Mistral and the Qwen family. By 2024, the Hugging Face apply_chat_template API had standardized how practitioners interact with these divergent formats. This gives a single abstraction layer over the entire ecosystem's menagerie of templates.

Prompt Templates

A prompt template defines the structure that wraps each instruction-response pair during training. Think of it as a contract between you and the model: you promise to always present instructions in a certain way, and in return, the model learns to recognize and respond to that presentation reliably. At its simplest, a template specifies where the instruction goes, where the response goes, and what separates them. However, template design has a significant impact on model behavior, affecting everything from how confidently the model generates responses to how gracefully it handles edge cases in multi-turn dialogue.

Language models are pattern recognition engines. During pre-training, they learn statistical regularities in text. During instruction tuning, we want them to learn a new regularity: the pattern that connects instructions to appropriate responses. The template creates this pattern by giving consistent structural cues that the model can latch onto. Without these cues, the model would have no way to determine where an instruction ends and where it should begin generating. The template is the primary mechanism through which the model learns the concept of "following an instruction."

Consider the most basic template structure:

### Instruction: {instruction} ### Response: {response}

This template uses natural language markers to delineate sections. The model learns that text after ### Instruction: contains what it should respond to, and text after ### Response: contains the expected output. During training, both sections are present. This allows the model to learn the mapping from instruction to response. During inference, you give everything up through ### Response: and let the model generate the continuation. The blank space after ### Response: is an invitation for the model to fill in what should come next.

Why do these markers work? The model learns through thousands of training examples that whenever it sees ### Instruction: followed by some text and then ### Response:, the text that follows ### Response: should be a helpful, relevant answer to whatever appeared after ### Instruction:. The consistency of this pattern allows the model to generalize: when it encounters the same structure with a new instruction at inference time, it knows exactly what role it should play. This generalization is the foundation of instruction tuning, and it depends entirely on the template being applied with perfect consistency throughout the training process.

The choice of marker matters more than it might initially appear. Natural language markers like ### Instruction: have the advantage of being human-readable and easy to work with during development and debugging. However, they carry a risk: the text of the marker might appear in normal content. Imagine a training example that asks the model to explain what the Alpaca ### Instruction: format is. The model would encounter ### Instruction: embedded within what is supposed to be instruction content, creating an ambiguous signal about where one section ends and another begins. This is why the community progressively shifted toward special tokens.

The Alpaca dataset, which pioneered many instruction-tuning practices, used a slightly more elaborate template that included an optional input field:

Below is an instruction that describes a task, paired with an input that gives further context. Write a response that appropriately completes the request. ### Instruction: {instruction} ### Input: {input} ### Response: {response}

This three-part structure separates the task description from any context needed to complete it. For a translation task, the instruction might be "Translate the following text to French" while the input contains the actual text to translate. This separation helps the model distinguish between instructions and context. The preamble at the beginning also serves an important function: it primes the model with context about what kind of task this is, setting expectations before the specific instruction appears. The model processes that preamble and adjusts its attention accordingly, arriving at the specific instruction already calibrated to produce a "response that appropriately completes the request."

When the input field is empty, the Alpaca format omitted it entirely and used a different preamble: "Below is an instruction that describes a task. Write a response that appropriately completes the request." This variability in the preamble means the model learns that the preamble is informative but not a hard structural signal; the real markers are ### Instruction: and ### Response:. Recognizing these as the primary anchors helps the model generalize across both preamble variants.

Template Consistency

The exact template you use matters less than using it consistently. A model trained with ### Instruction: markers will expect those same markers at inference time. Switching to [INST] delimiters without retraining will degrade performance because the model has never learned to recognize those markers as instruction boundaries. Even minor variations, such as extra whitespace or different capitalization, can cause measurable degradation when they are applied inconsistently. When building a training pipeline, encode your template as a single function and call it from every data preparation step.

Modern templates typically use special tokens rather than natural language markers. Special tokens are added to the tokenizer's vocabulary specifically to serve as structural delimiters. They are less likely to appear in normal text, making them unambiguous markers of format structure. Consider the difference: if you use ### Response: as a marker, and you ask the model to "explain what ### Response: means in the Alpaca format," the model might become confused about where the actual response should begin. Special tokens like <|assistant|> avoid this ambiguity because they exist outside the normal vocabulary of human-written text. The tokenizer always maps them to a single, dedicated token ID that cannot be produced by composing subword pieces from ordinary text.

The transition from natural-language markers to special tokens also changed how models stand for format structure internally. A natural-language marker like ### Response: is tokenized into a sequence of tokens that each carry semantic meaning from pre-training. The hash characters, the word "Response," the colon, all carry prior statistical associations. A special token like <|im_end|> starts as a blank slate during fine-tuning; its entire learned meaning derives from its structural role in the conversation format. This gives the model a cleaner signal: the special token says "structural boundary" without any confounding semantic content.

Worked Example: Tracing a Single Example Through a Template

Let's trace a concrete instruction-response pair through the Alpaca template step by step, tracking exactly what the model sees and what it learns to predict.

Raw training pair:

  • Instruction: "Summarize the following paragraph in one sentence."
  • Input: "The Amazon rainforest, often called the 'lungs of the Earth,' produces 20% of the world's oxygen and is home to 10% of all species."
  • Response: "The Amazon rainforest produces a fifth of global oxygen and harbors a tenth of Earth's species."

After template application, the model receives the following token sequence during training:

Below is an instruction that describes a task, paired with an input that gives further context. Write a response that appropriately completes the request. ### Instruction: Summarize the following paragraph in one sentence. ### Input: The Amazon rainforest, often called the 'lungs of the Earth,' produces 20% of the world's oxygen and is home to 10% of all species. ### Response: The Amazon rainforest produces a fifth of global oxygen and harbors a tenth of Earth's species.

What the model learns to predict: During causal language modeling, the model must predict each token from all previous tokens. But we only care about its ability to predict the response tokens. The tokens up through and including the newline after ### Response: are treated as context (their losses are masked to -100). The model is penalized only for mispredicting the response tokens. After many such examples, the model internalizes: "when I see ### Response: after an instruction and some input, I should produce a concise summary of the input that addresses the instruction."

At inference time: You give exactly:

Below is an instruction that describes a task, paired with an input... ### Instruction: Summarize the following paragraph in one sentence. ### Input: {new text} ### Response:

The model sees this partial sequence and generates the next tokens, which become the response. The final ### Response: header with no text after it is the model's cue to begin generating.

The key insight is that the template creates a deterministic boundary between "context to be processed" and "generation target." Everything before the response header is consumed as context. Everything after it is generated. The consistency of this boundary across all training examples is what makes the model reliable at inference time.

Role Definitions

Structured conversations require clear identification of who is speaking. This might seem obvious when you think about human conversation, where we always know who said what, but for a language model processing a stream of tokens, speaker identity is not inherent in the text. The model must learn to recognize speaker boundaries and adjust its behavior accordingly. Without role markers, the model has no way to distinguish a user's question from its own previous answer, making coherent multi-turn conversation impossible.

Instruction-tuned models typically recognize three basic roles, each serving a distinct purpose in the conversation architecture:

  • System: Provides context and persona definitions. The system role defines the behavior and persona for the conversation, persisting as a foundational constraint across all subsequent exchanges.

  • User: Represents the human interacting with the model. User messages contain the questions or instructions that the model must address.

  • Assistant: Represents the model itself. This role defines the identity and style of the generated responses, and tokens in this role are the primary targets of the training loss.

These roles emerged from practical necessity rather than theoretical elegance. Early chatbots needed to distinguish human input from bot output to maintain coherent conversations. As models became more advanced, practitioners discovered that adding a system role allowed behavior customization without cluttering the conversation with meta-instructions. Instead of asking the user to include "remember to be concise" in every message, you could establish that expectation once in the system prompt and have it apply throughout. This separation of concerns between behavioral configuration and user interaction is one of the most practically powerful aspects of the three-role design.

Role separation serves several purposes beyond mere bookkeeping. First, it lets training on multi-turn conversations where the model must track who said what and respond appropriately to the conversational context. Second, it allows system prompts that establish consistent behavior across interactions, meaning the same base model can behave differently for different applications without any additional fine-tuning. Third, it gives clear boundaries for computing the training loss: typically, we only compute loss on assistant tokens, not on user or system tokens. This choice ensures the model learns to generate appropriate responses rather than predicting user inputs. The model is not trying to learn what humans ask; it is learning how to answer.

The asymmetry between roles also has a strong effect on how the model processes conversation context internally. Because loss is computed only on assistant tokens, the model's gradient updates are driven entirely by its ability to produce coherent assistant turns. User and system tokens flow through the network as pure context, shaping attention patterns and hidden state representations, but they do not directly shape the model's weights in the same way. The model becomes specialized at "given this system context and this user query, produce this assistant response," which is precisely the capability we want.

Think of the three roles as a theatrical script where the system role is the director's notes (setting the stage and defining character), the user role is the other actor's lines (the inputs the model must respond to), and the assistant role is the model's own lines (what it learns to produce). Only the model's own lines are rehearsed during training; the director's notes and other actors' lines are fixed context that the model adapts to.

Out[4]:
Visualization
Diagram showing five stacked colored boxes representing a conversation sequence: System, User, Assistant, User, Assistant.
Sequence of conversation messages showing roles across multiple turns. The system message establishes foundational constraints, while subsequent user and assistant messages alternate to maintain state. This structure helps the model distinguish between context-setting instructions and the dialogue itself.

How Roles Shape Model Behavior

The way a model responds to role boundaries extends well beyond simple parsing. When the model encounters the assistant role marker during generation, it has already processed all prior context: the system instructions, the user's current query, and any previous conversational turns. Each of those context segments has already modified the model's hidden states. The role marker signals that it is now time to synthesize that context into a coherent generation, and the model's learned behavior for the assistant role is to produce helpful, on-topic, well-formed text.

This synthesis is especially powerful when the system message and user message interact. If the system message says "Always respond in the style of a Shakespearean playwright" and the user asks "What time is it?", the model must reconcile these two inputs. Its learned behavior for the assistant role includes respecting system constraints, so it might generate something like "The hour doth approach the second of the afternoon." The three-role structure is what makes this kind of behavioral conditioning possible, because the model has learned to treat the system role as a persistent constraint filter that all generations must pass through.

The distinction between user and assistant tokens also affects few-shot prompting dynamics. When you include example question-answer pairs in the conversation history (using user and assistant roles respectively), you are giving demonstrations in the format the model was trained on. The model recognizes these as examples of the instruction-following task and adjusts its generation accordingly. If you were to include those same examples without role markers, the model would see them as undifferentiated context and would generalize from them less reliably. The role structure gives a meta-signal that says "this is a demonstration," which triggers different inference behavior than plain context would.

Extended Roles: Tool Use and Reasoning

The basic three-role system works well for simple instruction following, but modern applications frequently require more fine-grained role distinctions. Tool-use models, for example, need a way to stand for the output of external functions called by the model. Many frameworks introduce a tool or function role for this purpose. When the model calls a search function, the results are injected into the conversation under the tool role, and the model then generates its final answer under the assistant role using those results as context.

Chain-of-thought reasoning introduces another pattern: some formats include a thinking or reasoning role that contains the model's internal reasoning steps, followed by an assistant role with the final answer. Models like DeepSeek-R1 and OpenAI's o1 series use this structure. Training on these extended roles teaches the model to explicitly work through problems before committing to an answer, which measurably improves performance on complex reasoning tasks. The role structure makes the separation between private reasoning and public output tractable from a training perspective.

System Messages

System messages establish the behavioral context for an entire conversation. Unlike user messages that stand for turn-by-turn interaction, system messages set persistent instructions that influence all subsequent responses. They act as a kind of meta-instruction that shapes how the model interprets and responds to everything that follows. Think of the system message as the briefing a professional receives before beginning an assignment: it sets expectations, establishes constraints, and defines the persona that should be maintained throughout.

System messages define the model's personality and constraints for a conversation. The model learns during training that system content should be treated as foundational truth that shapes all its responses. It learns to subordinate other considerations to system instructions, meaning a well-crafted system message can override default model behaviors and establish new conventions for a particular deployment.

A system message might specify:

You are a helpful assistant that gives concise, accurate answers. You always cite sources when making factual claims. If you're unsure about something, you acknowledge your uncertainty rather than guessing.

This tells the model how to behave across the entire conversation. Every response should be concise, should cite sources, and should acknowledge uncertainty. These constraints persist regardless of what the user asks. If the user later asks for a long, detailed explanation, the model must balance that request against the system instruction for conciseness. Through training on examples that contain similar tensions between system and user instructions, the model learns to resolve these situations, typically by honoring system constraints while still being as responsive as possible to user requests.

System messages became prominent with ChatGPT and similar conversational AI systems. They allow the same base model to behave differently in different contexts without any additional fine-tuning. A customer service application might use a system prompt stressing politeness and company policy, while a coding assistant might use one stressing technical accuracy and code quality. A creative writing assistant might receive a system prompt encouraging imagination and stylistic flourish. A medical information assistant might use a system prompt reminding the model to always recommend consulting a healthcare professional. This flexibility means a single trained model can power many different applications, each with its own personality and constraints, sharply reducing the cost and complexity of deploying specialized AI systems.

The depth of behavioral control that system messages give depends heavily on how much system-message diversity was included in the training data. A model trained almost entirely on a single system prompt will have learned to follow that one prompt expertly but may respond unpredictably to novel system prompts at inference time. Models trained on system prompts with varied writing styles and constraints generalize much better. This is why datasets like OpenAssistant and Dolly deliberately varied their system prompts, and why commercial models are typically trained on enormous pools of human-written system prompts.

The positioning of system messages matters for how the model processes them. Most formats place the system message at the very beginning of the conversation, before any user turns. This ensures the model processes instructions before encountering user requests, letting those instructions to condition all subsequent attention computations. Some formats allow system messages to appear mid-conversation, but this is less common and can confuse models not trained to expect it. The model has learned that system messages come first and set the stage; encountering one mid-conversation violates that learned pattern and may produce inconsistent behavior.

During training, system messages are typically included in the input but excluded from the loss calculation. The model sees the system message and learns to condition its responses on it, but it is not penalized for failing to predict the system message tokens themselves. This makes sense: we do not want the model to learn to generate system messages; we want it to learn to follow them. The asymmetry in how loss is applied creates an asymmetry in what the model learns to do: it becomes expert at being conditioned by system messages, not at creating them.

System Message Composition and Effectiveness

Writing effective system messages is both a science and an art. The model responds to system instructions because it has seen thousands of training examples where similar system instructions preceded certain kinds of behavior. The closer your system message is to the kinds of instructions the model saw during training, the more reliably it will follow them. Vague instructions like "be helpful" work poorly because helpfulness is already the default; the system message adds no specificity. Precise instructions like "always structure your response with a brief summary followed by numbered steps" work much better because they give the model a concrete behavioral pattern to match.

Some practitioners decompose complex behavioral requirements into multiple distinct instructions within a single system message, covering tone, format, knowledge boundaries, and persona separately. Others find that very long system messages exceed what the model can reliably attend to across a long conversation, since early tokens in a long context receive less attention weight than recent tokens. Finding the right balance between completeness and conciseness in system messages is an empirical question that depends on the model and the application.

Multi-Turn Conversation Format

Real conversations involve multiple exchanges between the user and the assistant. A single question and answer might suffice for simple tasks, but complex interactions require context to build over time. A multi-turn format must track the full conversation history while clearly marking each turn's boundaries and roles. This creates a cumulative context where each response can draw on everything that came before. This makes possible the model to resolve pronouns, follow up on previous statements, and maintain a coherent narrative across many exchanges.

Understanding multi-turn format is necessary because it determines how the model encodes conversational memory. The model does not have explicit memory; it has context. All prior turns are re-encoded on every generation step, and the model's responses are conditioned on the entire visible history. The format's job is to make that history unambiguously parseable so the model can correctly identify which content came from which speaker at each turn.

A simple multi-turn structure looks like:

User: What is the capital of France? Assistant: The capital of France is Paris. User: What's the population? Assistant: Paris has a population of approximately 2.2 million in the city proper, or about 12 million in the greater metropolitan area.

The model processes the entire history when generating each new response. For the second question, it sees both the original question about France's capital, the assistant's response, and the follow-up question about population. This context allows it to resolve the ambiguous reference "What's the population?" to mean Paris's population. Without the conversation history, the model would have no way to know what "the population" refers to. Multi-turn formats let natural, contextual conversation instead of isolated question-answering.

Multi-turn training requires careful attention to what receives gradient updates. Typically, you only compute loss on the assistant turns, not user turns or the system message. Standard causal attention masking ensures only assistant tokens contribute to the loss, preventing the model from learning to predict user inputs. User turns give context that shapes the assistant's response, but they are not targets for prediction. This selective loss computation is what makes the model specialized for response generation rather than conversation modeling in the general sense.

The training loss objective for a multi-turn example can be written as follows. Let x=(x1,x2,…,xT)\mathbf{x} = (x_1, x_2, \ldots, x_T) denote the full token sequence including system, user, and assistant tokens, and let A⊂{1,…,T}\mathcal{A} \subset \{1, \ldots, T\} denote the set of positions corresponding to assistant response tokens. The instruction-tuning loss is:

L=−∑t∈Alog⁡P(xt∣x1,…,xt−1;θ)\mathcal{L} = -\sum_{t \in \mathcal{A}} \log P(x_t \mid x_1, \ldots, x_{t-1}; \theta)

where:

  • xtx_t is the token at position tt
  • x1,…,xt−1x_1, \ldots, x_{t-1} is the full causal context up to position tt
  • θ\theta are the model parameters
  • A\mathcal{A} is the set of assistant-turn token positions

The key insight is that A\mathcal{A} does not include all positions, only those belonging to assistant turns. Non-assistant positions are still present in the input and influence the attention computation through the conditioning context x1,…,xt−1x_1, \ldots, x_{t-1}, but they do not contribute to the gradient signal. This means the model learns to generate good assistant responses while freely using system and user context to condition those responses.

The conversation format must handle several edge cases that arise in real-world applications. What happens if a user message is empty? What if there are consecutive user messages without assistant responses? How should the model handle extremely long conversation histories that exceed context length? Different formats and training procedures handle these cases differently, and your choice depends on your application's requirements. Some applications might truncate long histories, while others might summarize earlier turns, and still others might refuse to continue conversations that exceed certain lengths. None of these strategies is universally correct; each involves tradeoffs between context richness and computational feasibility.

Out[5]:
Visualization
Stacked area chart showing cumulative token counts over ten conversation turns with a red dashed horizontal line showing the context limit.
Cumulative context growth over ten conversation turns illustrating token consumption. Total token count increases linearly as exchanges accumulate, with assistant responses typically consuming more context than user prompts. The red dashed line marks the context limit, showing how rapidly long conversations approach token boundaries.

Packing Multiple Examples Into a Single Sequence

An important efficiency technique for multi-turn training is sequence packing, where multiple independent conversations are concatenated into a single long sequence to maximize GPU utilization. Instead of padding short conversations to the maximum sequence length (wasting compute on padding tokens), you fit as many conversations as possible into each sequence up to the context limit.

Packing requires careful attention masking. Each packed conversation must attend only to its own tokens, not to tokens from other conversations in the same sequence. This is achieved through block-diagonal attention masks, where each conversation forms a diagonal block that sees only its own tokens. Without these masks, the model would "leak" context from one conversation into another, corrupting the training signal. Most modern training frameworks support packing with block-diagonal attention, and the speedup over naive padding can be substantial, often two to four times on datasets with variable-length conversations.

Chat Templates

Chat templates standardize how conversations are converted to token sequences. Standardization ensures the training and inference formats match. Different model families use different templates. This shows different design choices and historical developments. Using the wrong template for a given model produces poor results because the model encounters structural cues it has never seen before. Understanding the major template formats helps you work effectively with the open-source model ecosystem and make informed decisions when designing your own instruction-tuning pipelines.

Think of chat templates as dialects of a common conversational language. All dialects convey the same basic information (who said what, when), but they do so using different vocabulary and syntax. A speaker of one dialect can be understood only by listeners who know that dialect. Similarly, a model trained on ChatML can be "understood" only by inference code that speaks ChatML.

ChatML Format

ChatML, which stands for Chat Markup Language, was introduced by OpenAI and has become one of the most common formats in the open-source ecosystem. It uses special tokens to mark message boundaries, giving clear and unambiguous structure:

<|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user What is 2 + 2?<|im_end|> <|im_start|>assistant 2 + 2 equals 4.<|im_end|>

The key elements are designed for clarity and consistency:

  • <|im_start|> marks the beginning of a message
  • <|im_end|> marks the end of a message
  • The role immediately follows <|im_start|> on the same line
  • Message content follows the role on subsequent lines

ChatML's structure makes role boundaries unambiguous. The special tokens are unlikely to appear in normal text, and the consistent structure simplifies parsing both for the model and for any code that needs to process these conversations. Many open-source models, including various Qwen versions, use ChatML or close variants. The widespread adoption of ChatML makes it easier to share training data and inference code across different projects.

The name "im" in <|im_start|> stands for "imaginary monologue," which shows the original design intent that each message is one participant's contribution to a dialogue. This naming is a historical artifact; in practice, everyone uses it simply as a structural marker. The single-token nature of <|im_start|> and <|im_end|> is important: the tokenizer maps each entire string to a single vocabulary entry, meaning the model always processes the boundary signal with a single attention step rather than needing to integrate information across multiple subword tokens.

Llama Chat Format

Meta's Llama models use a different format that evolved across model generations. This shows lessons learned and changing design priorities. Llama 2's chat format uses:

<s>[INST] <<SYS>> You are a helpful assistant. <</SYS>> What is 2 + 2? [/INST] 2 + 2 equals 4. </s><s>[INST] What about 3 + 3? [/INST]

The format nests the system message within special <<SYS>> tags, places instructions within [INST] and [/INST] markers, and uses <s> and </s> as sequence boundaries. For multi-turn conversations, each turn is wrapped with these markers. This nesting structure shows a hierarchical view of the conversation where system content is embedded within the first instruction turn rather than standing alone. Critics of this approach noted that it makes the system message structurally different from how it is conceptually: it is supposed to be persistent and separate, but it appears as a sub-component of the first user turn.

Llama 3 simplified and modified this format materially, introducing new special tokens like <|begin_of_text|>, <|start_header_id|>, and <|end_header_id|>:

<|begin_of_text|><|start_header_id|>system<|end_header_id|> You are a helpful assistant.<|eot_id|><|start_header_id|>user<|end_header_id|> What is 2 + 2?<|eot_id|><|start_header_id|>assistant<|end_header_id|> 2 + 2 equals 4.<|eot_id|>

This format gives more explicit structure with dedicated tokens for headers, end-of-turn markers, and text boundaries. The evolution from Llama 2 to Llama 3 format shows how template design improves as practitioners gain experience. The Llama 3 format separates concerns more cleanly: there are distinct tokens for marking role headers versus marking content boundaries, making the structure more regular and easier to parse programmatically. The system message now stands on equal footing with user and assistant turns structurally, which better shows its conceptual role.

Mistral Format

Mistral models use a simpler format that minimizes token overhead:

<s>[INST] You are a helpful assistant. What is 2 + 2? [/INST]2 + 2 equals 4.</s>[INST] What about 3 + 3? [/INST]

Mistral's format combines the system message with the first user message and uses minimal special tokens. This design choice prioritizes simplicity and efficiency. Notice there is no space after [/INST] before the assistant response. This detail matters for exact token alignment. During training, the model learns that the assistant response begins immediately after the closing instruction marker. Introducing a space at inference time would create a token sequence the model has never encountered, potentially affecting output quality. The Mistral format is unforgiving of whitespace errors in a way that ChatML is not, because ChatML uses newlines as natural separators.

Mistral's decision to embed the system message in the first user turn is a pragmatic simplification. It reduces the number of distinct roles the model must learn to handle, at the cost of making the system message structurally indistinguishable from the first user query. This is a reasonable tradeoff for models deployed in contexts where system messages are predictable and stable, but it limits fine-grained behavioral control in complex deployments.

Zephyr Format

Zephyr models, built on the Mistral architecture but fine-tuned for alignment, use a ChatML-like format with different tokens:

<|system|> You are a helpful assistant.</s> <|user|> What is 2 + 2?</s> <|assistant|> 2 + 2 equals 4.</s>

The structure resembles ChatML but uses <|system|>, <|user|>, and <|assistant|> as role markers with </s> as the end-of-message delimiter. This hybrid approach combines the clarity of ChatML-style role markers with the familiarity of the standard end-of-sequence token. The choice of format shows the Zephyr team's design preferences and the training data they used, which was sourced largely from UltraChat and other datasets that used this structure. The use of </s> as the end-of-message token rather than a dedicated <|im_end|> means the format shares structural tokens with the model's natural language generation, which can occasionally cause the model to end turns prematurely if the end-of-sequence token appears in normal content.

Phi-3 Format

Microsoft's Phi-3 family uses a format that blends elements of both ChatML and Llama 3:

<|system|> You are a helpful assistant.<|end|> <|user|> What is 2 + 2?<|end|> <|assistant|> 2 + 2 equals 4.<|end|>

The <|end|> token marks message boundaries, while <|system|>, <|user|>, and <|assistant|> mark role starts. The format is clean and easy to parse, with role markers on their own lines and content following on the next line. Phi-3 models were designed to reach strong performance with small parameter counts, and the clean template structure contributes to efficient training signal extraction. The regular pattern of "role marker, newline, content, end marker" is easy for the model to internalize and reliable to variations in content length.

Working with Chat Templates in Practice

The Hugging Face transformers library gives a standardized way to apply chat templates through the tokenizer's apply_chat_template method. This ensures you use the correct format for each model without manually constructing the template strings, which would be error-prone and tedious. The library abstracts away the differences between formats, letting you to write code that works across multiple model families.

The abstraction layer that apply_chat_template gives is more useful than it might first appear. Without it, you would need to maintain model-specific formatting code for every model you work with, update that code whenever a model's template changes, and carefully test that your code produces bit-identical output to the model's training pipeline. With it, you can write a single piece of code that formats conversations correctly for any model that has a template registered in its tokenizer. The templates are stored as Jinja2 strings in the tokenizer's tokenizer_config.json file, so they travel with the model weights.

Let's explore how chat templates work with a concrete example:

In[6]:
Code
from transformers import AutoTokenizer

# Load a tokenizer that has a chat template defined
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct")

# Define a conversation as a list of message dictionaries
messages = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {
        "role": "user",
        "content": "Write a Python function to calculate factorial.",
    },
    {
        "role": "assistant",
        "content": "def factorial(n):\n    if n <= 1:\n        return 1\n    return n * factorial(n - 1)",
    },
    {"role": "user", "content": "Can you add input validation?"},
]

The conversation is represented as a list of dictionaries, each with a role and content key. This standardized format works across different models regardless of their underlying template structure. The tokenizer handles converting it to the model-specific template, so you can write model-agnostic code that processes conversations uniformly. The dictionary representation is also the natural format for APIs like OpenAI's Chat Completions endpoint, so this same conversation structure often flows from API response logging directly into training data preparation.

In[7]:
Code
# Apply the chat template to get the formatted string
formatted_prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,  # Return string, not token IDs
    add_generation_prompt=True,  # Add the assistant turn start
)
Out[8]:
Console
Formatted conversation:
<|system|>
You are a helpful coding assistant.<|end|>
<|user|>
Write a Python function to calculate factorial.<|end|>
<|assistant|>
def factorial(n):
    if n <= 1:
        return 1
    return n * factorial(n - 1)<|end|>
<|user|>
Can you add input validation?<|end|>
<|assistant|>

The output displays the conversation formatted with the model's control tokens. The add_generation_prompt=True parameter is important for inference. It adds whatever tokens signal the start of an assistant turn, so the model knows to generate a response. Without it, you would get the conversation history but no prompt for the model to continue. The model would see a complete conversation and have no indication that it should add anything. This is one of the most common bugs in inference pipelines: forgetting add_generation_prompt=True and then wondering why the model either refuses to generate or generates continuation text rather than a fresh response.

Let's examine how the same conversation looks with different model templates:

In[9]:
Code
# Compare templates across different models
model_names = [
    "microsoft/Phi-3-mini-4k-instruct",
    "Qwen/Qwen2-0.5B-Instruct",
    "mistralai/Mistral-7B-Instruct-v0.1",
]

formatted_outputs = {}
for model_name in model_names:
    tok = AutoTokenizer.from_pretrained(model_name)

    # Use a simpler conversation for comparison
    simple_messages = [
        {"role": "user", "content": "Hello!"},
        {"role": "assistant", "content": "Hi there!"},
        {"role": "user", "content": "How are you?"},
    ]

    formatted = tok.apply_chat_template(
        simple_messages, tokenize=False, add_generation_prompt=True
    )
    formatted_outputs[model_name] = formatted
Out[10]:
Console

============================================================
Model: microsoft/Phi-3-mini-4k-instruct
============================================================
<|user|>
Hello!<|end|>
<|assistant|>
Hi there!<|end|>
<|user|>
How are you?<|end|>
<|assistant|>


============================================================
Model: Qwen/Qwen2-0.5B-Instruct
============================================================
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
Hi there!<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant


============================================================
Model: mistralai/Mistral-7B-Instruct-v0.1
============================================================
<s> [INST] Hello! [/INST] Hi there!</s> [INST] How are you? [/INST]

Notice how each model uses different delimiters and structures. The Phi model uses its specific template markers. Qwen uses ChatML-style <|im_start|> and <|im_end|> tokens. Mistral uses [INST] and [/INST] markers. Using the wrong template for a model would produce nonsensical outputs because the model has only been trained to recognize its specific format. This comparison illustrates why the apply_chat_template abstraction is so useful: it shields you from needing to know these details for every model you work with. When you load a model and call apply_chat_template, you are guaranteed to get exactly the format that model was trained on.

Out[11]:
Visualization
Stacked bar chart comparing token overhead across three chat template formats: Phi-3, Qwen2, and Mistral. Each bar shows content tokens in green and overhead tokens in purple.
Token overhead for Phi-3, Qwen2, and Mistral formats during a three-message exchange. Phi-3 uses the fewest structural tokens in this example, Mistral is intermediate, and the ChatML-based Qwen2 format uses the most, illustrating the cost of explicit role marking in long sequences.

Tokenizing for Training

During instruction tuning, you need to format the conversation and create labels that mask out non-assistant tokens from the loss calculation. This loss masking is basic to instruction tuning because it focuses the learning signal on the intended target: teaching the model to generate good responses. Without loss masking, the model would spend half its capacity learning to predict user messages, system prompts, and formatting tokens, diluting the gradient signal for what matters.

Loss masking works by replacing the labels for non-assistant tokens with -100, a value that PyTorch's CrossEntropyLoss function treats as "ignore this position." The loss at each position is computed as the cross-entropy between the model's predicted probability distribution and the true next token. When the label is -100, that position contributes zero to the total loss and zero to the gradient. Only positions where the label is a valid token ID contribute to learning.

Let's see how to prepare training data end to end:

In[12]:
Code
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct")

# Ensure we have a pad token
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

# Training conversation with complete assistant response
training_messages = [
    {"role": "system", "content": "You are a math tutor."},
    {"role": "user", "content": "What is 15 + 27?"},
    {"role": "assistant", "content": "15 + 27 = 42"},
]

# Tokenize the full conversation
import torch


def _to_tensor(ids):
    """Convert apply_chat_template output to a 2D tensor regardless of version."""
    if isinstance(ids, torch.Tensor):
        return ids.unsqueeze(0) if ids.dim() == 1 else ids
    # BatchEncoding (dict-like) or list
    if hasattr(ids, "input_ids"):
        ids = ids.input_ids
    if isinstance(ids, torch.Tensor):
        return ids.unsqueeze(0) if ids.dim() == 1 else ids
    # plain Python list of ints
    flat = (
        ids[0]
        if (isinstance(ids, list) and ids and isinstance(ids[0], list))
        else ids
    )
    return torch.tensor([flat])


_full_ids = tokenizer.apply_chat_template(
    training_messages,
    tokenize=True,
    add_generation_prompt=False,  # Don't add prompt since we have the response
)
full_tokens = _to_tensor(_full_ids)
Out[13]:
Console
Full sequence length: 35 tokens
Full sequence tokens: <|system|> You are a math tutor.<|end|><|user|> What is 15 + 27?<|end|><|assistant|> 15 + 27 = 42<|end|><|endoftext|>

The decoded sequence shows the conversation history merged into a single string with appropriate control tokens. Notice that we set add_generation_prompt=False because this is training data with a complete assistant response already included. We do not need to prompt for generation; we need to present the complete example for the model to learn from. If we accidentally set it to True, we would append the start-of-assistant marker a second time, corrupting the token sequence.

For training, we need to identify which tokens are from the assistant response, so we only compute loss on those:

In[14]:
Code
# Get everything up to the assistant response
messages_without_response = training_messages[:-1]

_prompt_ids = tokenizer.apply_chat_template(
    messages_without_response,
    tokenize=True,
    add_generation_prompt=True,  # Add assistant turn start
)
prompt_tokens = _to_tensor(_prompt_ids)

# The assistant response tokens are the difference
prompt_length = prompt_tokens.shape[1]
full_length = full_tokens.shape[1]
response_length = full_length - prompt_length
Out[15]:
Console
Prompt length: 22 tokens
Response length: 13 tokens

Prompt: <|system|> You are a math tutor.<|end|><|user|> What is 15 + 27?<|end|><|assistant|>

Response tokens only: 15 + 27 = 42<|end|><|endoftext|>

These outputs confirm the separation between the instruction context and the target response. By tokenizing the conversation with and without the assistant response, we can precisely identify which tokens constitute the response. This identification is important because those are the only tokens that should contribute to the training loss. The technique of comparing two tokenizations is called the prompt-continuation split, and it is the most reliable way to find the boundary between context and generation target.

Now we can create labels with masking for the prompt tokens:

In[16]:
Code
# Create labels for training
# -100 is the ignore index for CrossEntropyLoss
labels = full_tokens.clone()
labels[0, :prompt_length] = -100  # Mask out prompt tokens from loss
Out[17]:
Console
Token-by-token label visualization:
------------------------------------------------------------
  0: '<|system|>'         -> masked
  1: 'You'                -> masked
  2: 'are'                -> masked
  3: 'a'                  -> masked
  4: 'math'               -> masked
  5: 't'                  -> masked
  6: 'utor'               -> masked
  7: '.'                  -> masked
  8: '<|end|>'            -> masked
  9: '<|user|>'           -> masked
 10: 'What'               -> masked
 11: 'is'                 -> masked
 12: ''                   -> masked
 13: '1'                  -> masked
 14: '5'                  -> masked
 15: '+'                  -> masked
 16: ''                   -> masked
 17: '2'                  -> masked
 18: '7'                  -> masked
 19: '?'                  -> masked
 20: '<|end|>'            -> masked
 21: '<|assistant|>'      -> masked
 22: ''                   -> loss
 23: '1'                  -> loss
 24: '5'                  -> loss
 25: '+'                  -> loss
 26: ''                   -> loss
 27: '2'                  -> loss
 28: '7'                  -> loss
 29: '='                  -> loss
 30: ''                   -> loss
 31: '4'                  -> loss
 32: '2'                  -> loss
 33: '<|end|>'            -> loss
 34: '<|endoftext|>'      -> loss

The tokens marked as "loss" are from the assistant response and contribute to training. The masked tokens are from the system message, the user message, and format tokens. During backpropagation, gradients only flow from the assistant tokens, teaching the model to generate appropriate responses while preserving its understanding of how to parse the format structure. The special value -100 is recognized by PyTorch's CrossEntropyLoss as an ignore index, causing those positions to be excluded from the loss calculation entirely. Any attempt to include padding tokens in the loss would corrupt training with meaningless gradient signal, so the same masking technique also applies to padding positions when batching variable-length sequences.

Out[18]:
Visualization
Horizontal bar chart showing the first 30 tokens of an instruction-tuned training example, colored gray for masked tokens and green for response tokens that contribute to loss.
Token-level loss mask for instruction tuning across a conversational sequence. Assistant response tokens (green) contribute to the training loss, while system, user, and structural tokens (gray) are masked with -100 to be ignored. This ensures the model only learns to generate responses rather than the input context.

The Loss Fraction Problem

An important but often overlooked issue with loss masking is what practitioners call the loss fraction problem: in a heavily masked training example, only a small fraction of tokens contribute to the loss. If a training example has 200 context tokens and 20 response tokens, 90% of the tokens are masked. The model processes the entire 220-token sequence but only receives gradient updates from 20 positions.

This creates two related problems. First, short responses are implicitly downweighted relative to long responses: a dataset where some examples have 5-token responses and others have 500-token responses will train more effectively on the long-response examples by a factor of 100 in terms of gradient signal per example. Second, conversations with very short responses waste most of the compute spent processing the context.

Common solutions include filtering the training dataset to ensure responses meet a minimum length, upweighting examples with high response-to-context ratios, normalizing the loss by the number of non-masked tokens per example rather than per sequence, and using sequence packing to maximize response tokens per batch. The normalization strategy is particularly important: a naive implementation that averages loss over all token positions will systematically undervalue short, dense-information responses compared to long, verbose ones.

Custom Chat Templates

Sometimes you need to define a custom template for models that do not have one, or modify an existing template for specific use cases. This might occur when you are working with a base model that was never instruction-tuned, when you want to add new roles beyond the standard three, or when you need to match a specific data format that differs from standard templates. Hugging Face tokenizers use Jinja2 templating syntax, which gives the flexibility to define arbitrary formatting logic.

Jinja2 is a Python templating engine commonly used for web development. In the context of chat templates, it allows you to write conditional logic, iterate over messages, and format content based on roles. The template receives the list of message dictionaries and produces a formatted string. Think of Jinja2 templates as a tiny programming language embedded in a string, giving you the expressive power to handle any conversation structure your use case requires.

In[19]:
Code
# Load a base tokenizer without a chat template
base_tokenizer = AutoTokenizer.from_pretrained("gpt2")

# Define a custom chat template using Jinja2 syntax
custom_template = """{% for message in messages %}
{% if message['role'] == 'system' %}
[SYSTEM] {{ message['content'] }}
{% elif message['role'] == 'user' %}
[USER] {{ message['content'] }}
{% elif message['role'] == 'assistant' %}
[ASSISTANT] {{ message['content'] }}
{% endif %}
{% endfor %}
{% if add_generation_prompt %}
[ASSISTANT]{% endif %}"""

# Apply the custom template to the tokenizer
base_tokenizer.chat_template = custom_template

# Test it
test_messages = [
    {"role": "system", "content": "Be helpful."},
    {"role": "user", "content": "Hello!"},
]

custom_formatted = base_tokenizer.apply_chat_template(
    test_messages, tokenize=False, add_generation_prompt=True
)
Out[20]:
Console
Custom template output:
[SYSTEM] Be helpful.
[USER] Hello!
[ASSISTANT]

The Jinja2 template iterates over messages, applies role-specific formatting, and optionally adds a generation prompt. This flexibility allows you to design templates that match your training data format exactly. The template syntax supports conditionals, loops, and variable interpolation, giving you full control over how conversations are rendered to text.

When creating custom templates, ensure your special tokens (like [SYSTEM], [USER], [ASSISTANT]) are added to the tokenizer's vocabulary as special tokens. Otherwise, they will be split into subword pieces, making them harder for the model to recognize as format boundaries. A [USER] token that gets tokenized as [, USER, ] gives the model three separate tokens to attend to instead of one, and the model must learn that this particular three-token sequence is always a structural marker, which is harder than learning a single unique token:

In[21]:
Code
# Add special tokens to vocabulary
special_tokens = {
    "additional_special_tokens": ["[SYSTEM]", "[USER]", "[ASSISTANT]"]
}
num_added = base_tokenizer.add_special_tokens(special_tokens)
Out[22]:
Console
Added 3 special tokens
[USER] token ID: 50258
[ASSISTANT] token ID: 50259

The output verifies that the new tokens are now part of the vocabulary with assigned IDs. Each special token receives its own unique ID. This keeps it will always be tokenized as a single token rather than being broken into pieces. After adding special tokens, you must resize the model's embedding layer to accommodate the new vocabulary entries. This is typically done with model.resize_token_embeddings(len(tokenizer)). The new token embeddings will be randomly initialized and will learn appropriate representations during fine-tuning. This initialization is an important detail: the initial random values should be drawn from the same distribution as the existing embeddings to avoid instability during early training steps.

Advanced Jinja2 Template Features

The Jinja2 syntax used by Hugging Face supports several useful features beyond simple conditionals. You can use {% set ns = namespace(found=false) %} to track state across iterations, letting you to handle patterns like "add a separator after every turn except the last." You can use filters like {{ message['content'] | trim }} to clean whitespace. You can use {% if loop.last %} to detect the final message and apply special formatting.

Real-world templates for models like Llama 3 use these features to handle edge cases: what to do when there is no system message, how to format the generation prompt differently for different contexts, and how to handle tool calls and tool results as special message types. Reading the actual Jinja2 template stored in a model's tokenizer_config.json is one of the best ways to deeply understand how that model's format works, because the template is the canonical specification of the format from the model authors' perspective.

Key Parameters

The key parameters for chat templating and tokenization are:

  • messages: List of dictionaries containing the conversation history (roles and content). Each dictionary must have role and content keys.

  • tokenize: If True, returns token IDs; if False, returns a formatted string. Use False for debugging and inspecting the template output; use True when preparing inputs for model forward passes.

  • add_generation_prompt: If True, adds tokens to signal the start of an assistant response. Set to True at inference time and False when the complete assistant response is already included in the messages.

  • return_tensors: Specifies the format for returned data (e.g., 'pt' for PyTorch tensors). Useful when you want the tokenizer to return tensors directly rather than lists of integers.

  • special_tokens: A dictionary of tokens to add to the tokenizer's vocabulary. Use this to add domain-specific structural tokens when designing a custom format.

  • truncation and max_length: Control whether and how the template output is truncated to fit within a maximum sequence length. Essential for batch training where all sequences must be padded to the same length.

Format Considerations for Training Data

The format you choose affects several aspects of instruction tuning, from computational efficiency to model quality. Understanding these tradeoffs helps you make principled decisions when designing your training pipeline rather than simply defaulting to whatever the most popular current model uses.

Token efficiency matters for training speed and cost. Some formats use verbose markers that consume many tokens per message. A format that adds 20 tokens of overhead per turn might seem negligible, but across millions of training examples and dozens of turns per conversation, that overhead becomes substantial. ChatML's <|im_start|> and <|im_end|> are single tokens in models trained with them, but if you are adding them to a base model that has never seen them, they might be tokenized into multiple pieces initially. Compare the token counts of different formats on your data to understand the overhead and make informed decisions about format selection.

Consistency with pre-training improves results. If you are fine-tuning a model that was pre-trained with a specific format, continuing to use that format builds on what the model already learned. The model has already developed internal representations for those format tokens and patterns. This is why using the official chat template for instruct models is important: they were trained expecting that exact format, and any deviation forces the model to generalize in ways it was not prepared for. Even small inconsistencies, like using a slightly different whitespace pattern, can cause measurable degradation.

Edge cases need explicit handling. What happens when a message is empty? When the user sends multiple messages before the assistant responds? When the conversation exceeds context length and must be truncated? Your template and data processing pipeline should handle these cases gracefully. Consider whether empty messages should be included with empty content or skipped entirely. Define clear truncation strategies that preserve the most important context, typically the system message and the most recent exchanges.

System messages are not always supported. Some formats and models do not support system messages, or support them only in limited ways. Mistral's original chat format had limited system message support, requiring the system content to be prepended to the first user message. If your application relies on system prompts for behavior control, verify your chosen model and format support them properly before committing to a particular approach. Deploying a system-message-dependent application on a model that was trained without dedicated system message support will result in unreliable adherence to system instructions.

Limitations and Practical Considerations

Specific formats can make deployed systems fragile. A model trained on ChatML format will produce degraded outputs if you accidentally include text that looks like format markers. If your message contains <|im_end|>, the model might interpret it as an end-of-message signal and behave unexpectedly, potentially truncating its response or switching to a confused state. Reliable deployment requires input sanitization to escape or remove potential format tokens from your content before passing it to the model. This sanitization must be complete, covering all special tokens the model recognizes as structural. In practice, this means maintaining an allowlist of known special tokens for each model you deploy and stripping or escaping them in user inputs.

Format lock-in is another concern that affects long-term system design. Once you have trained a model on a specific format, changing formats requires retraining. This limits your ability to adopt new, potentially better formats as they emerge. Some practitioners address this by training on multiple formats simultaneously, teaching the model to recognize several template styles, though this increases training complexity and data requirements. The multi-format approach also requires careful attention to data balancing to ensure the model learns all formats equally well. Without balancing, the model will develop stronger adherence to whichever format is most represented in the training data.

The separation of roles, while useful, is an abstraction that does not perfectly map to all use cases. Some applications need more fine-grained roles than system, user, and assistant: perhaps a tool role for function calling results, or an observation role for chain-of-thought reasoning steps. Extended formats like those used in function-calling models introduce additional roles, but each extension adds complexity and requires format-specific training data. The more roles you introduce, the more advanced your training data must be to teach the model when and how each role should be used. A model that has seen the tool role in only 1% of its training examples will follow tool-format instructions much less reliably than one trained on a balanced mixture.

Multi-turn context management also presents challenges that become more pronounced as conversations grow longer. As conversations accumulate tokens, they eventually exceed the model's context window. Truncation strategies must preserve enough history for coherent responses while respecting token budgets. Simply truncating from the beginning can remove the system prompt or early context that later messages depend on. If a conversation began with establishing user preferences and later references those preferences implicitly, truncating the early turns will break that coherence. More advanced approaches keep the system message and recent turns while summarizing or dropping middle content, but these strategies require careful implementation and may introduce their own artifacts when the summary fails to capture all relevant details.

The training loss masking strategy, where loss is computed only on assistant tokens, means the model never directly learns to generate user messages or system prompts. This is intentional: we want the model to respond to instructions, not to generate them. However, it means the model's understanding of these roles comes only from their influence on its own outputs, not from explicit prediction of those tokens. The model learns that certain patterns precede assistant turns and uses that context, but it does not learn to produce those patterns itself. This creates an interesting asymmetry: the model is very good at following system instructions it has seen patterns of during training but may behave unpredictably when given system instructions that are structurally unlike anything in its training distribution.

The relationship between format and capability is not neutral. Some research suggests that certain format choices create subtle biases in model behavior. Models trained with very short system messages, for example, may respond less reliably to long, complex system prompts at inference time because the distribution of system-message lengths does not match their training experience. Similarly, models trained predominantly on single-turn conversations may struggle with coherent multi-turn reasoning even when multi-turn conversation format is applied correctly, because carrying context across turns requires targeted multi-turn training data. The format structure alone is insufficient.

Summary

Instruction formats give the structure that allows language models to parse conversations and generate correct responses. The design choices made in a template affect how the model parses inputs, what behavioral capabilities emerge from training, how reliable the deployed system is to adversarial or unusual inputs, and how flexibly the model can be deployed across different applications.

The key concepts we covered include:

  • Prompt templates wrap instruction-response pairs with consistent delimiters that the model learns to recognize. Whether using natural language markers like ### Instruction: or special tokens, consistency between training and inference is needed, and even minor formatting deviations can cause measurable degradation.

  • Role definitions separate system context from user input and assistant output. This three-role structure lets multi-turn conversations, behavioral customization through system prompts, and appropriate loss masking during training. Extended roles like tool and thinking support more advanced capabilities in modern models.

  • System messages establish persistent behavioral context that influences all subsequent responses. They allow the same base model to behave differently across applications without retraining, and their effectiveness depends on training data that presents diverse system-message styles and constraints.

  • Chat templates standardize how conversations become token sequences. Different model families use different formats: ChatML for many OpenAI-style models, [INST] markers for Llama 2 and Mistral, dedicated header tokens for Llama 3, and various others. Using the tokenizer's apply_chat_template method ensures correct formatting for each model and protects against manual formatting errors.

  • Training data preparation requires identifying assistant tokens for loss computation while masking out prompts, system messages, and format tokens. The prompt-continuation split technique reliably locates this boundary. The loss fraction problem, where short responses contribute disproportionately little gradient signal, deserves careful attention when designing the training data pipeline.

  • Custom templates using Jinja2 syntax allow full control over format design for novel use cases, but they require careful special-token management to ensure structural markers are single tokens rather than multi-token sequences.

The next chapter on Instruction Tuning Training will cover how to train models using these formatted datasets, including learning rate schedules, batch construction strategies, and convergence monitoring.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about instruction format and chat templates.

Instruction Format

Question 1 of 70 of 7 completed
Why is loss masking applied to non-assistant tokens during instruction tuning?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025instructionformat, author = {Michael Brenndoerfer}, title = {Instruction Format: Chat Templates and Roles}, year = {2025}, url = {https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Instruction Format: Chat Templates and Roles. Retrieved from https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm
MLAAcademic
Michael Brenndoerfer. "Instruction Format: Chat Templates and Roles." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm>.
CHICAGOAcademic
Michael Brenndoerfer. "Instruction Format: Chat Templates and Roles." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Instruction Format: Chat Templates and Roles'. Available at: https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Instruction Format: Chat Templates and Roles. https://mbrenndoerfer.com/writing/instruction-format-chat-templates-role-definitions-llm

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.