Synthetic Data: Generation, Quality, Diversity, Distillation

Michael BrenndoerferJanuary 13, 202665 min read

Part of Language AI Handbook

Explains how LLMs are trained on synthetic data, from Self-Instruct and Evol-Instruct to quality verification, diversity control, and knowledge distillation.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Synthetic Data

Real-world text data is often scarce, with uneven coverage across topics, languages, and tasks. A model trained only on naturally occurring web text will learn to reflect whatever is abundant online: English dominates, common topics crowd out rare ones, and instruction-following examples are nearly absent from raw crawls. Synthetic data generation solves this problem directly. Instead of scraping more of the same, you instruct a capable model to write the data you need, verify its quality, diversify it to cover the gaps, and distill the capability of a large teacher into a smaller student model.

Synthetic data has become one of the central tools in modern LLM training. Models like Phi-1, Phi-2, WizardLM, Orca, and Mixtral Instruct were all trained heavily on synthetic or model-generated content, often outperforming models trained on far larger natural corpora. Understanding how synthetic data is generated, why it works, and when it fails is essential for anyone working on language model training pipelines.

This chapter covers the full lifecycle of synthetic data for LLM training: how to generate diverse and high-quality examples, how to verify their quality automatically, how to ensure diversity across topics and styles, and how to use knowledge distillation to transfer capability from large teacher models to smaller student models.

Why Synthetic Data Works

Before diving into mechanics, it is worth understanding why using model-generated text to train models is not circular. The intuition might seem suspicious: if a model generates the training data, are you not just teaching the model to imitate itself?

The key insight is that generation is easier than learning from scratch. A capable teacher model already contains knowledge and reasoning capability accumulated during its own pretraining. Generating synthetic data is a process of eliciting and rendering that knowledge into structured, labeled examples that a student model can learn from efficiently. The student does not need to rediscover reasoning patterns from unstructured text; it gets curated demonstrations of those patterns, with the structure and framing optimized for learning.

Think of it this way: if you want to teach someone to write formal proofs, you could hand them a thousand pages of raw mathematical literature and let them figure it out, or you could show them five hundred carefully constructed proof examples that isolate exactly the reasoning moves they need to internalize. Both contain roughly the same underlying knowledge, but the second form is dramatically more sample-efficient because it delivers that knowledge at the right level of abstraction for the learner. Synthetic data is the second approach applied at scale.

A second reason synthetic data works is coverage. Natural corpora underrepresent many desirable behaviors: careful step-by-step reasoning, refusals for harmful requests, low-resource language responses, domain-specific technical writing, and structured output formats like JSON or code with specific function signatures. You can generate targeted examples for any of these at scale, filling the gaps that natural data cannot cover at any realistic crawl budget.

A third reason is label quality. In natural datasets, labels are often noisy or require expensive human annotation. A capable teacher model can generate both the input and the ideal output, effectively providing its own supervision signal. This is particularly powerful for instruction-following datasets, where the task is to train a model to respond helpfully to user instructions. Collecting 50,000 human-annotated instruction-response pairs requires hundreds of annotator hours. Generating 50,000 synthetic instruction-response pairs from a capable teacher requires a few API calls and a modest budget.

A fourth reason, sometimes overlooked, is controllability. Natural data reflects what happened to be written and published; synthetic data reflects what you deliberately construct. If a model is weak at a particular reasoning pattern, you can build a seed set that generates thousands of examples targeting that exact weakness. If a model overuses hedging language, you can generate training examples that demonstrate confident, direct answers. This degree of intentional shaping is impossible with natural data.

Building on the data quality pipeline assembled across Part XX, synthetic data is the final tool in the data curation toolkit: after deduplication (MinHash), quality filtering, toxicity removal, PII scrubbing, and careful data mixing, synthetic data fills the remaining gaps with purpose-built examples.

Synthetic Data Generation

The most common approach to synthetic data generation is prompted generation: you write a template that describes the kind of data you want, fill in variables from a seed set, and call a teacher model to produce the output. The architecture of a generation pipeline has three components that determine the quality of what comes out: the template, the seed data, and the teacher model.

Prompt Templates and Seed Data

A prompt template defines the structure of the generation request. For an instruction-following dataset, a typical template might look like:

You are an expert in {domain}. Write a detailed question about {topic} that a curious student might ask. Then provide a thorough, accurate answer. Format your response as: Question: <question here> Answer: <answer here>

The template is a contract: it specifies the output format, the desired style, the expected length, and the constraints the generation must satisfy. A well-written template is more than a fill-in-the-blank form; it is a specification of quality. Templates that describe desired qualities explicitly ("thorough", "accurate", "step-by-step") produce better outputs than vague templates ("write something about {topic}"), because language models attend to every word in the prompt and use it as guidance.

The seed data provides the values that fill the template variables: a list of domains, topics, or example inputs that drive diversity. Seeds can come from:

  • Existing natural corpora (Wikipedia titles, academic paper abstracts, Stack Overflow question titles)
  • Manually curated topic lists organized by domain and difficulty
  • Outputs from a previous generation round (self-play / iterative generation)
  • Human-written examples for few-shot prompting that anchor the style

The quality and breadth of your seed set directly determines the quality and breadth of the generated dataset. A narrow seed set produces a narrow dataset, regardless of how capable the teacher model is. If 80% of your seeds are about software engineering, 80% of your generated data will be about software engineering, and models trained on it will be disproportionately strong at software tasks. Seed curation is therefore not a secondary concern; it is the first place to invest effort in a synthetic data pipeline.

One underappreciated aspect of template design is the use of few-shot examples within the template. Instead of describing what you want in abstract terms, showing the model two or three examples of desired input-output pairs immediately constrains the generation toward the right format and style. The few-shot examples serve as in-context demonstrations of the quality bar, and the model tends to produce outputs that match that quality level.

Self-Instruct

Self-Instruct is a seminal synthetic data approach that bootstraps instruction-following data from a seed set of 175 human-written instruction examples. The paper introduced the idea that a model can expand its own training data by generating new instructions, classifying them, and producing input-output pairs, with only minimal human curation at the seed stage. The process works in four stages.

First, the model generates new instructions by sampling from a prompt that shows a few of the seed instructions and asks the model to create more. These are diverse because the prompt explicitly asks for variety, and because the few-shot examples are sampled randomly from the growing pool, so the model sees different combinations each time. Second, the model classifies each instruction as either a classification task or a generation task, because these two types require different input-output structures. Third, the model generates input-output pairs for each instruction. Fourth, low-quality examples are filtered using simple heuristics.

The heuristics in the original Self-Instruct paper catch obvious failures: instructions that are shorter than three words, instructions that begin with an imperative not appropriate for the task type, instructions that too closely match one of the seed examples (measured by ROUGE-L similarity), and input-output pairs where the model simply copied the instruction into the output field. These lightweight filters remove perhaps 30-40% of generated examples without requiring any additional model calls.

Self-Instruct demonstrated that a small human-curated seed set can bootstrap thousands of diverse instruction examples automatically. Alpaca, one of the first widely-used open instruction-tuned models, used a variant of Self-Instruct to generate 52,000 instruction-following examples from GPT-3.5 using a single calling batch. The entire dataset cost under $500 to generate at the time, making it a remarkably cost-effective approach to building instruction-following capability. The resulting model, despite being only 7B parameters, matched or exceeded GPT-3 on many instruction-following benchmarks.

One subtle but important design choice in Self-Instruct is that the few-shot prompt always shows examples from the growing pool of generated instructions, not just from the original seed. As the pool grows, the model sees increasingly diverse examples in its context, which reinforces diversity in what it generates. This feedback mechanism means the dataset becomes more varied over time, not less, provided the prompt selection samples broadly from the existing pool rather than repeatedly selecting the same popular examples.

The ROUGE-L similarity filter that removes near-duplicate instructions is worth examining more closely. ROUGE-L measures the longest common subsequence between two text sequences, normalized by the length of the sequences. If a newly generated instruction has a ROUGE-L score above 0.7 with any existing instruction in the pool, it is treated as a near-duplicate and discarded. This filter prevents the pool from being dominated by paraphrases of the same underlying instruction, which would waste generation budget without adding diversity. The threshold of 0.7 is a reasonable default, but you may want to lower it (to 0.5) for domains where natural paraphrase is common, or raise it (to 0.8) for domains with a naturally small vocabulary of instruction verbs.

Evol-Instruct

Evol-Instruct (from WizardLM) takes a different approach. Instead of generating new instructions from scratch, it starts with a base instruction and evolves it to be more complex through a series of mutations. The core insight is that flat instruction datasets have a difficulty ceiling: they tend to produce simple, direct questions and answers because the model generating them defaults to moderate complexity. By explicitly prompting the model to make an instruction harder, you can systematically populate the high-difficulty end of the training distribution.

The mutations include:

  • In-depth evolution: Adding constraints, reasoning steps, or requirements ("In addition, explain why the approach fails when the input contains null values...")
  • In-breadth evolution: Creating new, related instructions in adjacent areas of the topic
  • Difficulty increase: Making the task harder ("Instead of a simple list, provide a comparative analysis with quantitative trade-offs...")
  • Concreteness increase: Making vague instructions more specific by adding details about the desired format, length, or audience
  • Deepen a concept: Asking for a fuller explanation of a concept that was previously treated as assumed background knowledge

Each evolution step produces a harder version of the original instruction. After multiple rounds of evolution, you have a dataset that spans difficulty levels from simple to highly complex. The WizardLM models trained on Evol-Instruct data showed significantly better performance on complex reasoning benchmarks compared to models trained on flat instruction sets, because the training distribution included the complex end of the difficulty spectrum rather than truncating it.

The reason Evol-Instruct works is that it addresses a known failure mode of flat instruction datasets: they underrepresent complex, multi-step reasoning. When a model is only trained on simple instructions ("Summarize this passage", "Translate this sentence"), it learns to follow simple patterns but struggles with tasks that require chaining multiple steps, applying multiple constraints simultaneously, or synthesizing information from multiple angles. By systematically constructing harder variants of every seed instruction, Evol-Instruct ensures the training distribution includes the complex end of the difficulty spectrum.

One practical consideration is that evolved instructions can become incoherent if the evolution prompt is too aggressive. Adding too many constraints at once produces instructions that are technically longer but logically contradictory or impossible to satisfy. A student asked to "write a Python function that sorts a list in O(n) time using only built-in Python operations without using any list methods or built-in sort" is receiving an impossible instruction. Good Evol-Instruct pipelines include a coherence filter: after evolution, a judge model is asked whether the evolved instruction is a sensible, self-consistent task. Instructions flagged as incoherent are discarded rather than evolved further.

The difficulty of an evolved instruction is not the same as the quality of the training example it produces. A very hard instruction might have no good answer, in which case training on the generated response teaches the model to produce plausible-sounding non-answers for hard questions. Pairing Evol-Instruct with rejection sampling, where possible, provides the strongest combination: evolve instructions to be hard, then filter generated responses to keep only those that are verifiably correct or highly rated by a judge.

Backtranslation

Backtranslation is an elegant technique borrowed from machine translation. In translation, backtranslation generates synthetic source-language text by translating target-language monolingual data backward. In the instruction-tuning context, the analogous process runs in reverse from the typical generation direction:

  1. Start with high-quality response text (for example, a well-written Wikipedia paragraph, a well-reasoned essay, a clear technical tutorial)
  2. Ask the teacher model: "What question or instruction would elicit this response?"
  3. Use the generated instruction as the input and the original text as the output

This is powerful because you can use any large corpus of high-quality text as a source of responses, then synthetically create the corresponding instructions. The LIMA paper ("Less Is More for Alignment") used 1,000 carefully selected examples from this approach to fine-tune a strong model, showing that curated quality matters more than raw quantity. The LIMA model, fine-tuned on only 1,000 examples, was competitive with models fine-tuned on 52,000 Alpaca examples, illustrating how dramatically sample quality affects learning efficiency.

The backtranslation approach offers a natural quality guarantee: you start with text you know is high quality (a well-written encyclopedia article, a carefully argued essay, a detailed technical tutorial) and generate the instruction to match it. This is the reverse of the typical flow where you worry about whether the model's response is good. The response is already guaranteed to be good; the only uncertainty is whether the generated instruction is a natural and sensible request for that response. That is a much easier quality judgment to make, and it can often be verified by asking whether a reasonable person would ask such a question.

One extension of backtranslation is to use it for stylistic alignment. If you have a corpus of responses in a particular style (formal academic writing, friendly conversational explanations, concise technical documentation), backtranslation can generate corresponding instructions that target that style. The resulting training data teaches the model what to say and how to say it in response to particular kinds of requests.

The limitation of backtranslation is that the instructions it generates tend to be more abstract and less specific than instructions in natural use. "What is the main argument of this passage?" is a natural backtranslation instruction for an expository paragraph, but real users often ask more specific, contextual questions like "Why did the author reject the classical interpretation?". The fix is to use diverse high-quality source texts rather than encyclopedic survey articles, which tend to elicit more specific and varied instructions.

Rejection Sampling

Rejection sampling is used when you have a verifiable correctness criterion, such as math problems, code tasks, or logical puzzles. The approach is straightforward in principle:

  1. Generate many candidate solutions for a given problem
  2. Execute or verify each solution (run the code, check the math answer, evaluate logical validity)
  3. Keep only the solutions that pass verification
  4. Use the passing solutions as training examples

Rejection sampling is the backbone of math-focused synthetic datasets like MetaMath and OpenMathInstruct. Given a math problem, the teacher generates dozens of candidate solutions. Each solution is verified by checking whether the final numerical answer matches the known ground truth. Only solutions that reach the correct answer through valid reasoning steps are kept.

The main advantage of rejection sampling is that it provides a hard correctness signal, not a soft quality judgment. A solution either produces the right answer or it does not. This binary signal is especially useful for training reasoning capability. A soft quality score from a judge model might rate a plausible-sounding but incorrect proof highly, because the reasoning looks confident and well-structured. Rejection sampling bypasses this failure mode entirely: you never train on incorrect reasoning, regardless of how well it reads.

The MetaMath dataset, constructed by augmenting GSM8K and MATH problems with rejection-sampled solutions from a large model, produced training data where every example is verifiably correct, and models fine-tuned on it showed substantially improved mathematical reasoning compared to models trained on unfiltered outputs. The improvement was particularly pronounced on harder problems, where incorrect reasoning is most likely to pass superficial quality checks.

The limitation is that rejection sampling requires a verifier. For math, you can check numerical answers. For code, you can run tests. For open-ended creative writing or general conversation, there is no oracle to consult. This is why rejection sampling is primarily used for structured domains where ground truth is computable, and why the math and code domains have benefited disproportionately from synthetic data approaches.

A useful property of rejection sampling is that the acceptance rate is a proxy for problem difficulty. Easy problems have high acceptance rates (the model usually gets them right); hard problems have low acceptance rates. You can use this signal to curate datasets at specific difficulty levels, focusing on problems where the model gets it right between 20% and 80% of the time. Problems in this range tend to produce the most informative training examples: the model can solve them sometimes, which means it is operating near its current capability boundary, and learning to solve them reliably produces the most improvement.

Out[3]:
Visualization
Three overlapping histograms showing acceptance rate distributions for easy, medium, and hard problems
Simulated acceptance rate distributions for 200 problems in each of three difficulty tiers (easy, medium, and hard math problems). Easy problems cluster at high acceptance rates (mean around 0.80), while hard problems cluster at low rates (mean around 0.18). The shaded band marks the 20-80% sweet spot where problems are most informative: difficult enough to require effort but solvable enough that correct solutions exist and can be learned from.

Quality Verification

Generating synthetic data is cheap. Generating good synthetic data requires systematic quality verification. Models hallucinate, produce vague responses, copy instructions verbatim into answers, generate structurally malformed outputs, and occasionally produce fluent but completely wrong answers. Without quality filters, these examples enter the training set and degrade the student model. The quality verification layer is where a synthetic data pipeline earns its value.

Quality verification is a multi-stage funnel. Each stage catches a different class of failures. The stages are ordered by cost: fast rule-based filters first, slower model-based scoring second, and expensive consistency verification last. This ordering ensures that only examples that pass cheap filters proceed to expensive evaluation, keeping the overall pipeline cost manageable.

Rule-Based Filters

The first line of defense is a set of fast, cheap rule-based filters. These catch obvious quality failures without requiring model calls:

  • Length filters: Remove responses that are too short (a one-word answer to a detailed question) or too long (runaway generation exceeding a reasonable token budget for the task type)
  • Repetition filters: Detect n-gram repetition within a response, since degenerate generation often produces repeated phrases or entire paragraphs as the model loops
  • Format filters: Verify that structured outputs (JSON, code, markdown tables) parse correctly and conform to the required schema
  • Keyword filters: Flag responses containing phrases that indicate refusal ("As an AI language model, I cannot...") when the task does not require refusal
  • Language ID filters: Verify the response is in the target language, using a lightweight language identification library
  • Instruction-response length ratio filters: Catch cases where the response is much shorter than the instruction would imply necessary, which often indicates a non-answer

These filters are not perfect but they are fast. Running them on every generated example before any model-based scoring saves significant compute. In practice, a well-tuned rule-based filter removes 15-30% of generated examples at essentially no cost per example, compared to the per-token cost of a model-based scoring call.

One important implementation detail is to make rule-based filters configurable per task type. For a creative writing task, you want to allow longer responses. For a code completion task, the length ratio between instruction and response may be inverted. A general pipeline that applies the same filters uniformly to all task types will either be too restrictive for some tasks or too permissive for others.

Model-Based Quality Scoring

After rule-based filtering, model-based scoring evaluates the semantic quality of responses. The most common approach is to prompt a strong judge model (often the same teacher that generated the data, or a more capable model) to score each example on a numerical scale.

A typical scoring prompt asks the judge to evaluate the response along several dimensions:

  • Accuracy: Does the answer correctly address the question?
  • Completeness: Does the response cover all aspects of the request?
  • Relevance: Is the content directly relevant to the instruction, without unnecessary tangents?
  • Coherence: Is the response logically consistent and well-structured?
  • Helpfulness: Would a person asking this question be satisfied with this response?

The judge returns a score (often on a 1-5 or 1-10 scale) along with a brief explanation. Examples below a threshold are discarded. The explanation helps with debugging and threshold calibration: if the judge consistently explains that 4/5 scores correspond to minor omissions while 3/5 scores correspond to factual errors, you can use that information to set a threshold that matches your tolerance for different failure types.

A practical consideration when using model-based scoring is position bias: judge models tend to rate responses that use more confident, structured language more highly, independent of factual accuracy. If you generate ten responses to the same question and the one rated highest is the most verbose and well-formatted but not the most accurate, you are training the student to imitate a style rather than to be correct. Cross-referencing model-based scores with rejection sampling results, where available, helps detect this bias.

The choice of judge model also matters significantly. Using the same model as the teacher introduces correlated errors: the teacher's systematic biases will be rated favorably by itself. Using a stronger or independently trained model as judge reduces this correlation, at higher per-call cost. A practical middle ground is to use the teacher model for initial filtering and bring in a stronger judge for borderline examples near the threshold.

IFEval-Style Format Verification

When generating instruction-following data with explicit format constraints ("respond in JSON", "use bullet points", "keep your answer under 100 words"), you can verify format compliance programmatically. The IFEval benchmark evaluates models on instructions with verifiable format constraints, and the same verification functions can be used to filter training data.

For example, if a generated example has the instruction "Respond only with a numbered list", you can check:

  • Does the response contain numbered items?
  • Does the response start with a number, not a sentence?
  • Does it avoid prose paragraphs entirely?

Format verification provides a hard filter that complements the soft scores from model-based judging. It is particularly valuable when you want to train a model to follow explicit format instructions reliably, since format compliance is fully verifiable without any model call.

The same principle extends to many other structural constraints: word count limits (count words), response structure (check for required sections), code syntax (parse with AST), output in a specific language (language ID), presence of citations (pattern match for URLs or reference formats), and so on. Every constraint in your instruction set that is mechanically verifiable should have a corresponding filter function in your pipeline.

Reward Model Scoring

A trained reward model provides a learned quality signal rather than a prompted one. Reward models are trained on human preference data: pairs of responses where a human has labeled which is better. The reward model learns to predict human preference, and this learned function can score individual examples without pairwise comparison.

Using a reward model for synthetic data filtering has two advantages. First, it is faster than calling a large judge model, since the reward model is typically much smaller and requires no elaborate prompting. Second, it captures quality dimensions that are hard to articulate in a prompt, because they were learned implicitly from thousands of human judgments rather than specified by a single prompt author.

The limitation is that a reward model can overfit to surface features that humans associate with quality but that do not reflect correctness. Responses that are verbose, politely formatted, and carefully structured may score highly even when they are factually wrong. This is the reward model version of the same position bias problem described above.

Reward Model Hacking

When synthetic data is filtered using a reward model, and the student model is then trained on the filtered data, the student learns to produce outputs that score well on the reward model, not necessarily outputs that are correct. This is a subtle form of reward hacking. Mitigation strategies include using multiple diverse reward models trained on different human preference datasets, periodically replacing the reward model as the student improves and begins to exploit it, and combining reward model scores with ground-truth verification wherever possible.

Consistency Filtering

For factual questions, you can filter for consistency by generating multiple responses to the same question and checking agreement. If most responses agree on the answer, that consensus provides evidence for correctness. Responses that contradict the consensus are discarded.

This approach is analogous to self-consistency decoding (majority voting at inference time) applied to data generation. If you generate eight responses and six agree that the answer to a historical question is a particular year while two give different years, the majority response has higher evidence for correctness than any single response would. The consistency signal is particularly useful for math and factual knowledge, where there is a deterministic correct answer, but it works surprisingly well for open-ended questions too: consistent responses tend to reflect the model's well-grounded beliefs rather than hallucinated one-offs.

Consistency filtering requires multiple generation passes per example, which multiplies the API cost. A practical approach is to reserve consistency filtering for high-stakes subsets: factual knowledge about specific entities, dates, quantities, and causal claims, where a hallucinated answer is most damaging if it enters the training data.

Synthetic Data Diversity

A common failure mode of synthetic data pipelines is topic collapse: the generated data looks superficially varied but is dominated by a few frequent topics. If your seed set contains many software engineering examples and few biology examples, the generated data will reflect that imbalance. Models trained on such data will be strong at programming questions and weak at biology questions, not because biology is inherently harder to teach with synthetic data, but because the pipeline never tried.

Topic collapse is particularly insidious because it is invisible to standard quality metrics. Each individual example might score highly on accuracy and helpfulness, and the aggregate dataset looks large and varied at a glance. The collapse only becomes apparent when you analyze the topic distribution quantitatively or evaluate the trained model on a diverse benchmark.

Diversity is therefore a first-class concern in synthetic data design, not an afterthought.

Measuring Diversity

Before you can improve diversity, you need to measure it. Several metrics are useful:

Embedding-based diversity measures the spread of examples in embedding space. You embed all generated examples using a sentence encoder, then compute the average pairwise cosine distance across the dataset. A high average distance indicates diverse coverage; a low distance indicates that examples cluster together, which corresponds to topic collapse. The mean pairwise distance is sensitive to outliers but provides a useful aggregate signal.

N-gram diversity measures the proportion of unique n-grams across the dataset. Self-BLEU is a standard metric: for each example, compute its BLEU score against the rest of the dataset. A low average self-BLEU score indicates low redundancy, meaning each example contains mostly novel n-grams not present in other examples. High self-BLEU indicates that examples share a lot of phrasing, which often means topic collapse or template overuse.

Topic distribution analysis uses a topic model (LDA or BERTopic) to identify the latent topics in the dataset and measure how evenly they are distributed. An ideal dataset has roughly equal coverage across topics rather than a long tail of rare topics and a few dominant ones. The entropy of the topic distribution is a useful summary statistic: maximum entropy corresponds to uniform coverage, and lower entropy corresponds to concentration in fewer topics.

Vocabulary coverage measures the unique types (distinct words) divided by total tokens across the dataset. A dataset dominated by a few topics will reuse the same technical vocabulary repeatedly, producing a lower type-token ratio than a truly diverse dataset.

To make this concrete, consider what topic collapse looks like when visualized. A biased seed set produces a training distribution heavily skewed toward a few categories, while a well-curated pipeline produces near-uniform coverage. The figure below shows the difference between a collapsed distribution and a diverse one.

Out[4]:
Visualization
Bar chart showing skewed topic distribution with programming dominating
Topic distribution from a biased seed set, simulating topic collapse. A handful of categories (programming, machine learning) dominate the dataset while others receive almost no representation. A model trained on this distribution would underperform on underrepresented topics despite having seen large total training volumes.
Bar chart showing balanced topic distribution across domains
Topic distribution after applying topic-guided generation with a balanced seed list. Each domain receives roughly equal coverage, giving the trained model consistent exposure across all topics and producing more balanced benchmark performance.

Topic-Guided Generation

The most direct approach to diversity is to explicitly specify the topic during generation. Instead of letting the model choose what to write about, you provide a diverse list of topics and generate a fixed number of examples per topic. This hard allocation ensures that rare topics appear in the dataset regardless of how rarely they appear in the teacher's spontaneous outputs.

The challenge is building a sufficiently broad topic list. For general instruction-following, you can start with a broad taxonomy (science, history, math, code, creative writing, practical advice, and so on) and then expand each category into subcategories. Structured knowledge bases like Wikipedia's category tree, academic taxonomy lists from arXiv, or standard educational curriculum frameworks all provide good starting points for a coverage map.

Topic-guided generation ensures that rare topics appear in the dataset, but it requires manual work to maintain and expand the topic list as the model's capability scope evolves. A practical workflow is to start with a coarse topic list, generate an initial dataset, evaluate the model on a benchmark, identify the weakest domains, and expand the topic list with finer-grained categories in those domains for the next generation round. This feedback loop between evaluation and generation makes topic-guided generation an iterative process rather than a one-time setup.

One practical subtlety is that topics specified in the prompt are not the only determinant of what the model generates. If you specify "biology" but your template asks for a "technical question", the model will generate technical biology questions, not introductory biology questions. The combination of topic and framing in the template jointly determines the difficulty and style distribution within a topic. Varying both dimensions independently gives you a finer-grained control over the training distribution.

Persona-Based Generation

Persona-based generation is a complementary diversity strategy that varies the perspective rather than the topic. Instead of specifying the subject matter, you specify who is asking the question or what kind of responder is answering:

  • A high school student struggling with algebra for the first time
  • A software engineer debugging a production system under time pressure
  • A non-native English speaker asking for writing feedback on a job application
  • A medical professional researching drug interactions for a patient case

Different personas produce different question styles, vocabulary levels, and implicit context assumptions. A high school student's algebra question will be phrased differently from a university professor's algebra question, even if both concern the same underlying mathematical concept. The resulting dataset captures a wider range of the actual distribution of real users, which generalizes better than a dataset generated from a single implicit perspective (typically: educated, technical, native English speaker).

Persona-based generation is also useful for controlling response style. If you specify that the responder is a "patient, encouraging tutor who explains concepts from first principles", the generated responses will adopt that pedagogical style. This lets you deliberately shape the assistant persona encoded in the training data rather than accepting whatever implicit style the teacher model defaults to.

The Microsoft Persona Hub paper demonstrated that a diverse set of synthetic personas, generated from a model trained to represent a wide demographic and professional spread, produced significantly more diverse training data than generic prompt templates. The key insight is that personas encode implicit assumptions about vocabulary, domain familiarity, and communication style that naturally vary the generated content in ways that topic lists alone cannot capture.

Iterative Feedback and Self-Play

In iterative self-play, the model generates data, trains on it, and the improved model generates the next round of data. This bootstrapping process can progressively expand coverage if each round is designed to target the model's current weaknesses. The dynamic is appealing: as the model improves, the data generated in the next round can be harder and more demanding, because the improved model can produce and evaluate more sophisticated examples.

The Reinforcement Learning from AI Feedback (RLAIF) paradigm extends this idea by adding an explicit evaluation loop: an AI judge evaluates the model's outputs, provides feedback, and the model learns from that feedback through reinforcement learning. This creates a self-improvement loop that can operate at scale without human annotators. Constitutional AI, developed by Anthropic, is one example of this approach: a model trained to follow a set of principles uses those principles to critique and revise its own outputs, generating training data for further refinement.

The risk of iterative self-play is mode collapse: the model converges to a narrow set of high-reward outputs and loses diversity over successive rounds. A model might discover that verbose, well-structured responses always score highly regardless of content, and begin producing exclusively that style. Mitigation requires explicit diversity constraints in the generation process, such as requiring each batch to cover a minimum number of distinct topics, or using a reward model that explicitly penalizes stylistic uniformity alongside correctness.

A more subtle failure mode is capability stagnation: if the model can only generate data at its current capability level, iterative self-play may not push it beyond that ceiling. Each round produces data similar to the previous round, and the model plateaus. The most effective mitigation is to inject harder problems from external sources (benchmark datasets, human-designed challenges) into each round alongside the self-generated data, making sure that the training distribution always includes examples at and slightly above the model's current frontier.

Deduplication and Near-Duplicate Removal

Even with diversity controls in place, synthetic generation tends to produce near-duplicates. If the same seed appears multiple times in the generation queue, or if the model produces two responses to different seeds that happen to cover the same ground with slightly different wording, these near-duplicates add bulk without adding coverage.

As we discussed in the MinHash chapter, locality-sensitive hashing efficiently identifies near-duplicates across large datasets. Applying MinHash deduplication to the synthetic data after generation removes redundant examples and keeps the most diverse subset. The choice of Jaccard similarity threshold matters here: a very high threshold (0.95) only removes near-identical examples, while a lower threshold (0.7) removes paraphrases that cover the same ground. For instruction-following datasets, a threshold of 0.7-0.8 on the instruction text (ignoring the response) is typically appropriate.

An alternative to MinHash for small-to-medium datasets is embedding-based deduplication: embed all examples, then use a greedy selection algorithm to pick the largest subset where every pair has cosine similarity below a threshold. This approach is more semantically sensitive than MinHash but scales poorly to very large datasets. For datasets above 100,000 examples, MinHash or similar approximate methods are more practical.

Knowledge Distillation

Knowledge distillation is the process of transferring the capability of a large, expensive teacher model into a smaller, efficient student model. The fundamental motivation is economic: large models are expensive to run at inference time. If a 7B student model can achieve 90% of a 70B teacher's performance on your tasks, deploying the student saves roughly 10x in inference compute. Synthetic data is the primary mechanism for distillation in modern LLM training: the teacher generates training examples, and the student trains on those examples.

The connection between synthetic data generation and distillation is direct. When you call GPT-4 to generate instruction-following examples and train a smaller model on them, you are doing response distillation. The student is learning to produce responses in the style and at the quality level of the teacher, compressing the teacher's learned behavior into a smaller set of parameters.

Response Distillation

The simplest form of distillation is response distillation: the teacher generates complete responses to input prompts, and the student learns to mimic those responses via standard supervised fine-tuning.

Given a dataset of prompts D={x1,x2,…,xN}\mathcal{D} = \{x_1, x_2, \ldots, x_N\}, the teacher model TT generates a response yi=T(xi)y_i = T(x_i) for each prompt. The student model SS is then trained to minimize the negative log-likelihood of the teacher's response tokens, given the prompt and all preceding tokens. This is the standard next-token prediction loss, applied to teacher-generated text:

Lresponse=−∑i=1N∑t=1∣yi∣log⁡PS(yit∣xi,yi<t)\mathcal{L}_{\text{response}} = -\sum_{i=1}^{N} \sum_{t=1}^{|y_i|} \log P_S(y_i^t \mid x_i, y_i^{<t})

where:

  • NN is the number of training examples in D\mathcal{D}
  • xix_i is the input prompt for example ii
  • yiy_i is the teacher-generated response for prompt ii, consisting of ∣yi∣|y_i| tokens
  • yity_i^t is the tt-th token of response yiy_i, serving as the prediction target at position tt
  • yi<ty_i^{<t} is the sequence of response tokens at positions before tt (the left context)
  • PS(yit∣xi,yi<t)P_S(y_i^t \mid x_i, y_i^{<t}) is the student's predicted probability for the correct token at position tt

The outer sum over ii accumulates the loss across all examples. The inner sum over tt accumulates the token-level loss within a single response. Minimizing this objective encourages the student to assign high probability to exactly the sequence of tokens the teacher produced.

This is identical to standard language model fine-tuning, but the labels come from the teacher rather than human annotators. Response distillation is what most instruction-tuning pipelines do when they call GPT-4 to generate training data. The simplicity is its virtue: it requires no special training infrastructure, no access to the teacher's internal activations, and no modification to the standard fine-tuning pipeline.

Reasoning Chain Distillation (Chain-of-Thought)

A more powerful form of distillation transfers the reasoning process along with the final answer. Chain-of-thought (CoT) distillation generates step-by-step reasoning traces from the teacher and trains the student to reproduce those traces.

The Orca paper demonstrated that training a 13B parameter model on detailed reasoning traces from GPT-4 produced a student that substantially outperformed models trained on final answers alone. The reasoning traces provide explicit supervision for the intermediate steps that the model needs to solve complex problems, rather than supervising only the endpoint. The student learns the answer and the path for arriving at it.

Generating reasoning chains requires prompting the teacher carefully. System messages like "Think step by step and explain your reasoning in detail before giving the final answer" elicit chain-of-thought outputs. The generated chains then serve as training targets. The key insight is that chain-of-thought prompting was already known to improve reasoning at inference time; distillation makes that reasoning capability a trained property of the student rather than something that requires special prompting at inference time.

The quality of distilled reasoning chains depends on the quality of the teacher's chains. A teacher that produces crisp, well-structured reasoning traces produces better training data than a teacher whose reasoning is scattered or contains errors. For this reason, reasoning chain distillation is most effective with the strongest available teacher models, and the quality investment in the teacher pays off disproportionately in the student's downstream performance.

One important practical decision is whether to train the student to produce reasoning traces at inference time, or to use the reasoning traces only as intermediate supervision during training. The first approach (having the student reason aloud at inference time) produces more interpretable models and generally achieves better performance on complex tasks. The second approach produces a more compact model that does not need to generate reasoning tokens, which reduces inference latency. The choice depends on whether the deployment context values interpretability or speed more.

Logit Distillation

Classical knowledge distillation (from Hinton et al., 2015) transfers the teacher's answers and the teacher's soft probability distribution over the next token. Instead of training the student on the argmax output of the teacher, the student is trained to match the teacher's full output distribution.

The key idea is that a teacher model's probability distribution over tokens carries much more information than its single best prediction. When a teacher assigns 60% probability to token A, 30% to token B, and 10% to token C, it is communicating that A and B are both plausible, and that B is more plausible than C. Training on this soft distribution teaches the student the full structure of the teacher's knowledge about token plausibility, not just which token happened to be the argmax.

The distillation loss combines the student's cross-entropy loss on the correct labels with a divergence loss against the teacher's distribution:

Ldistill=(1−α)LCE(y,PS)+α⋅T2⋅KL(PT(T)∥PS(T))\mathcal{L}_{\text{distill}} = (1 - \alpha) \mathcal{L}_{\text{CE}}(y, P_S) + \alpha \cdot T^2 \cdot \text{KL}(P_T^{(T)} \parallel P_S^{(T)})

where:

  • LCE\mathcal{L}_{\text{CE}} is the standard cross-entropy loss against the ground truth label yy
  • PT(T)P_T^{(T)} is the teacher's softmax distribution at temperature TT
  • PS(T)P_S^{(T)} is the student's softmax distribution at temperature TT
  • α\alpha is a mixing coefficient controlling the balance between hard labels and soft targets
  • TT is the temperature parameter that softens both distributions

The temperature T>1T > 1 softens the distributions, spreading probability mass from the argmax token to non-argmax tokens. This reveals the teacher's "dark knowledge": its beliefs about which wrong answers are more plausible than others. A teacher that assigns 40% probability to token A and 35% to token B is conveying that A and B are nearly interchangeable in this context, which is much more information than a teacher that assigns 100% to A and nothing to B.

The T2T^2 scaling in the loss formula compensates for a mathematical artifact of temperature scaling. When you divide logits by TT before softmax, the gradients of the resulting KL divergence are also scaled by 1/T21/T^2. The T2T^2 multiplier in the loss cancels this scaling, making sure that the soft loss and the hard loss contribute comparable gradient magnitudes regardless of the temperature choice.

Logit distillation requires access to the teacher model's logit outputs, which is not available when using API-only models like GPT-4. It is most applicable when you own or have white-box access to the teacher model. For open-source teacher models (Llama, Mistral, and so on), logit distillation is straightforward to implement. For proprietary API teachers, response distillation is the only available option.

Speculative Decoding as Distillation

Speculative decoding is an inference efficiency technique that has a natural connection to distillation. A small draft model proposes token sequences, and the large verifier model either accepts or rejects them in parallel. Tokens accepted by the verifier match the verifier's distribution; tokens rejected are replaced.

The data generated during speculative decoding contains rich alignment signals: each accepted draft token is one where the small model's prediction aligned with the large model's. Collecting these accepted tokens across many inference runs produces a training dataset for improving the draft model, making it gradually more aligned with the verifier. This is an organic form of online distillation, where the distillation signal comes from the natural operation of the inference system rather than from a separate training data generation stage.

The appeal of this approach is that it turns inference compute into training data without requiring any additional API calls or prompt construction. Every deployment of a speculative decoding system automatically generates a stream of alignment data that can be used to improve the draft model, creating a flywheel where deployment improves the system over time.

Capacity Gaps and When Distillation Fails

Distillation works best when the teacher-student capacity gap is moderate. If the teacher is GPT-4 and the student is a 70B model, distillation can transfer substantial capability because the student has enough representational capacity to capture most of the teacher's behavior. If the student is a 1B model trying to learn from GPT-4's outputs, the gap may be too large: the student's representational capacity cannot fit the full range of teacher behaviors.

The mathematical intuition is that a neural network can only represent functions within its function class, and smaller networks have smaller function classes. No matter how many examples you provide from a teacher whose behavior lies outside that function class, the student cannot represent the teacher's full behavior. The student will find the best approximation it can, but there is a hard upper bound on how closely it can approximate the teacher.

In practice, this manifests as the student producing grammatically correct but factually shallow responses. The student has learned the style of the teacher's outputs (confident, well-structured, appropriately hedged) without acquiring the underlying knowledge. This is sometimes called style distillation as opposed to capability distillation. A student that has absorbed the teacher's stylistic patterns but not its knowledge will sound authoritative while being factually unreliable, which can be worse than a student that sounds uncertain, because users may trust the confident-sounding incorrect answer.

Mitigation strategies include:

  • Using smaller teacher models that are closer in capability to the student, so the capacity gap is more bridgeable
  • Curriculum learning: start with simpler examples from the teacher and gradually increase complexity as the student improves
  • Filtering the teacher's outputs to keep only examples within the student's current competence range, avoiding examples so far above the student's level that they produce only noise in the gradient signal
  • Combining distillation with pretraining on natural data, which provides the foundational knowledge the student needs before it can make sense of the teacher's outputs

Implementation

Let's build a complete synthetic data pipeline in Python, covering generation, quality filtering, and diversity analysis.

Setup and Dependencies

We start by installing the required libraries and configuring the pipeline.

In[5]:
Code
# Install required packages
# uv pip install openai sentence-transformers scikit-learn numpy matplotlib

Simulating Synthetic Generation

For demonstration purposes, we simulate a teacher model's outputs rather than calling an actual API. The patterns here mirror exactly what you would do with a real API call.

In[6]:
Code
import random
from typing import Optional

# Simulated teacher model outputs (in production, replace with actual API calls)
SIMULATED_RESPONSES = {
    "instruction": [
        {
            "instruction": "Explain the concept of gradient descent in machine learning.",
            "response": "Gradient descent is an optimization algorithm that iteratively adjusts model parameters to minimize a loss function. At each step, it computes the gradient of the loss with respect to the parameters and moves in the opposite direction of the gradient by a small amount called the learning rate. The process repeats until the loss converges to a minimum. Variants like stochastic gradient descent (SGD) and Adam use subsets of data or adaptive learning rates to improve convergence speed.",
            "domain": "machine_learning",
            "difficulty": "intermediate",
        },
        {
            "instruction": "What is the difference between precision and recall?",
            "response": "Precision measures what fraction of the model's positive predictions are actually correct: TP / (TP + FP). Recall measures what fraction of actual positives the model correctly identified: TP / (TP + FN). High precision means few false alarms; high recall means few missed detections. The F1 score combines both into a single metric as their harmonic mean. In medical diagnosis, recall is usually prioritized to avoid missing true cases, while in spam filtering, precision may matter more to avoid blocking legitimate email.",
            "domain": "machine_learning",
            "difficulty": "intermediate",
        },
        {
            "instruction": "Describe the water cycle.",
            "response": "The water cycle describes the continuous movement of water through Earth's systems. Water evaporates from oceans, lakes, and land surfaces when heated by the sun. Water vapor rises into the atmosphere, cools, and condenses into clouds through a process called condensation. When enough water accumulates in clouds, it falls as precipitation: rain, snow, sleet, or hail. Precipitation either runs off into rivers and eventually back to the ocean, soaks into the ground to replenish groundwater, or is taken up by plants and returned to the atmosphere through transpiration.",
            "domain": "earth_science",
            "difficulty": "beginner",
        },
        {
            "instruction": "How does HTTPS encryption protect data in transit?",
            "response": "HTTPS uses TLS (Transport Layer Security) to encrypt data between your browser and the server. The process starts with a TLS handshake: the server presents its certificate to prove its identity, and the client and server negotiate an encryption cipher. They then use asymmetric cryptography to securely exchange a symmetric session key. All subsequent data is encrypted using that symmetric key, which is fast and efficient. This prevents eavesdroppers from reading traffic and ensures the server you're talking to is who it claims to be.",
            "domain": "computer_security",
            "difficulty": "intermediate",
        },
        {
            "instruction": "Write a Python function to check if a string is a palindrome.",
            "response": "def is_palindrome(s: str) -> bool:\n    # Normalize: lowercase and remove non-alphanumeric characters\n    cleaned = re.sub(r'[^a-z0-9]', '', s.lower())\n    return cleaned == cleaned[::-1]\n\n# Examples\nprint(is_palindrome('racecar'))   # True\nprint(is_palindrome('A man a plan a canal Panama'))  # True\nprint(is_palindrome('hello'))     # False",
            "domain": "programming",
            "difficulty": "beginner",
        },
        {
            "instruction": "What is photosynthesis and why is it important?",
            "response": "Photosynthesis is the process by which plants, algae, and some bacteria convert light energy, carbon dioxide, and water into glucose and oxygen. The reaction takes place in chloroplasts using the pigment chlorophyll. It matters because it is the primary source of oxygen in Earth's atmosphere and the foundation of almost all food chains on land. Without photosynthesis, the oxygen levels in the atmosphere would decline, and the chemical energy that supports most living organisms would disappear.",
            "domain": "biology",
            "difficulty": "beginner",
        },
        {
            "instruction": "Explain transformer attention in one paragraph.",
            "response": "Transformer attention allows each token in a sequence to selectively gather information from all other tokens. For each position, the model computes a query vector, and each other position contributes a key vector and a value vector. The attention score between two positions is the dot product of their query and key vectors, scaled by the square root of the dimension to prevent magnitude blowup, then passed through softmax to produce weights. These weights are used to compute a weighted sum of the value vectors, producing a context-aware representation for each position.",
            "domain": "machine_learning",
            "difficulty": "advanced",
        },
        {
            "instruction": "What caused the French Revolution?",
            "response": "The French Revolution (1789-1799) arose from a convergence of financial crisis, social inequality, and Enlightenment ideals. France was effectively bankrupt after funding the American Revolution and years of expensive wars. The tax system was deeply unfair: nobles and clergy were largely exempt while commoners bore most of the burden. A series of bad harvests in the late 1780s caused widespread food shortages and rising bread prices, pushing the urban poor to crisis. Enlightenment philosophers like Rousseau and Voltaire had spread ideas about individual rights and the illegitimacy of absolute monarchy, providing intellectual fuel for the uprising.",
            "domain": "history",
            "difficulty": "intermediate",
        },
    ]
}


# Simulate generating synthetic examples from a teacher model
def simulate_teacher_generation(
    instructions: list[str], domain_filter: Optional[str] = None
) -> list[dict]:
    """
    In production, this function would call an API.
    Here we return pre-defined examples filtered by domain.
    """
    examples = SIMULATED_RESPONSES["instruction"]
    if domain_filter:
        examples = [e for e in examples if e["domain"] == domain_filter]
    # Shuffle to simulate non-determinism
    random.seed(42)
    result = random.choices(examples, k=len(instructions))
    return result

Rule-Based Quality Filtering

The rule-based filter checks each example against length constraints, repetition detection, and refusal phrase detection. These are the cheapest checks to run and should always come first in the pipeline.

In[7]:
Code
def check_length(
    response: str, min_words: int = 20, max_words: int = 500
) -> bool:
    words = len(response.split())
    return min_words <= words <= max_words


def check_repetition(
    response: str, ngram_size: int = 5, threshold: float = 0.3
) -> bool:
    """Returns True if repetition is below threshold (response is NOT repetitive)."""
    words = response.lower().split()
    if len(words) < ngram_size:
        return True
    ngrams = [
        tuple(words[i : i + ngram_size])
        for i in range(len(words) - ngram_size + 1)
    ]
    unique_ratio = len(set(ngrams)) / len(ngrams)
    return unique_ratio >= (1.0 - threshold)


def check_no_refusal(response: str) -> bool:
    """Returns True if the response does not contain refusal phrases."""
    refusal_phrases = [
        "as an ai language model",
        "i cannot provide",
        "i'm not able to",
        "i don't have the ability",
        "i am unable to assist",
    ]
    response_lower = response.lower()
    return not any(phrase in response_lower for phrase in refusal_phrases)


def rule_based_filter(examples: list[dict]) -> tuple[list[dict], dict]:
    """Apply all rule-based filters and return passing examples + stats."""
    passed = []
    stats = {
        "total": len(examples),
        "failed_length": 0,
        "failed_repetition": 0,
        "failed_refusal": 0,
    }

    for ex in examples:
        response = ex["response"]
        if not check_length(response):
            stats["failed_length"] += 1
            continue
        if not check_repetition(response):
            stats["failed_repetition"] += 1
            continue
        if not check_no_refusal(response):
            stats["failed_refusal"] += 1
            continue
        passed.append(ex)

    stats["passed"] = len(passed)
    return passed, stats

Model-Based Quality Scoring

The heuristic scorer below approximates what a judge model would evaluate: whether the response is appropriately detailed relative to the instruction's complexity, whether it contains structured content signals, and whether its vocabulary is rich.

In[8]:
Code
import re


def score_response_heuristic(instruction: str, response: str) -> float:
    """
    Heuristic quality scorer (simulates model-based scoring).
    Returns a score from 0.0 to 1.0.
    In production, replace with a judge model API call.
    """
    score = 0.0
    response_words = len(response.split())
    instr_words = len(instruction.split())

    # Length relative to instruction complexity
    if response_words > instr_words * 3:
        score += 0.25
    elif response_words > instr_words:
        score += 0.15

    # Presence of substantive content signals
    if any(c.isdigit() for c in response):
        score += 0.10  # Contains specific numbers/data
    if ":" in response:
        score += 0.10  # Contains structured content
    if response[0].isupper() and response[-1] in ".!?":
        score += 0.10  # Proper sentence structure

    # Vocabulary richness
    words = re.sub(r"[^a-z\s]", "", response.lower()).split()
    if len(words) > 0:
        vocab_richness = len(set(words)) / len(words)
        score += min(0.35, vocab_richness * 0.5)

    return min(1.0, score)


def model_based_filter(
    examples: list[dict], threshold: float = 0.5
) -> tuple[list[dict], list[float]]:
    """Score examples and filter by threshold."""
    scored = []
    scores = []
    for ex in examples:
        score = score_response_heuristic(ex["instruction"], ex["response"])
        scores.append(score)
        if score >= threshold:
            scored.append(ex)
    return scored, scores

Diversity Analysis

The diversity module computes three metrics: average pairwise TF-IDF distance (how spread-out the examples are in term-space), vocabulary diversity (type-token ratio), and domain entropy (how evenly domains are represented).

In[9]:
Code
from collections import Counter

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity


def compute_diversity_metrics(examples: list[dict]) -> dict:
    """
    Compute diversity metrics for a set of examples.
    Returns average pairwise distance and self-BLEU proxy.
    """
    if len(examples) < 2:
        return {
            "avg_pairwise_distance": 0.0,
            "vocabulary_diversity": 0.0,
            "domain_entropy": 0.0,
        }

    texts = [ex["instruction"] + " " + ex["response"] for ex in examples]

    # TF-IDF-based embedding diversity
    vectorizer = TfidfVectorizer(max_features=500, stop_words="english")
    try:
        tfidf_matrix = vectorizer.fit_transform(texts)
        similarity_matrix = cosine_similarity(tfidf_matrix)
        # Average pairwise distance (1 - similarity) for upper triangle
        n = len(texts)
        distances = []
        for i in range(n):
            for j in range(i + 1, n):
                distances.append(1.0 - similarity_matrix[i, j])
        avg_distance = np.mean(distances) if distances else 0.0
    except Exception:
        avg_distance = 0.0

    # Vocabulary diversity: unique words / total words across all texts
    all_words = " ".join(texts).lower().split()
    vocab_diversity = len(set(all_words)) / len(all_words) if all_words else 0.0

    # Domain entropy
    domains = [ex.get("domain", "unknown") for ex in examples]
    domain_counts = Counter(domains)
    total = sum(domain_counts.values())
    entropy = -sum(
        (c / total) * np.log2(c / total + 1e-9) for c in domain_counts.values()
    )
    max_entropy = np.log2(len(domain_counts)) if len(domain_counts) > 1 else 1.0
    domain_entropy = entropy / max_entropy if max_entropy > 0 else 0.0

    return {
        "avg_pairwise_distance": avg_distance,
        "vocabulary_diversity": vocab_diversity,
        "domain_entropy": domain_entropy,
        "n_domains": len(domain_counts),
        "domain_distribution": dict(domain_counts),
    }

Running the Full Pipeline

With all components in place, we run the complete pipeline end to end: generate, rule-filter, score, and measure diversity.

In[10]:
Code
# Step 1: Simulate generation from a teacher model
seed_instructions = [
    "Explain the concept of gradient descent in machine learning.",
    "What is the difference between precision and recall?",
    "Describe the water cycle.",
    "How does HTTPS encryption protect data in transit?",
    "Write a Python function to check if a string is a palindrome.",
    "What is photosynthesis and why is it important?",
    "Explain transformer attention in one paragraph.",
    "What caused the French Revolution?",
]

generated_examples = simulate_teacher_generation(seed_instructions)

# Step 2: Rule-based filtering
rule_passed, rule_stats = rule_based_filter(generated_examples)

# Step 3: Model-based quality scoring and filtering
quality_passed, quality_scores = model_based_filter(rule_passed, threshold=0.5)

# Step 4: Diversity analysis
diversity_before = compute_diversity_metrics(generated_examples)
diversity_after = compute_diversity_metrics(quality_passed)

# Step 5: Summary statistics
pipeline_summary = {
    "generated": len(generated_examples),
    "after_rule_filter": len(rule_passed),
    "after_quality_filter": len(quality_passed),
    "rule_filter_stats": rule_stats,
    "diversity_before": diversity_before,
    "diversity_after": diversity_after,
    "avg_quality_score": np.mean(quality_scores) if quality_scores else 0.0,
    "min_quality_score": np.min(quality_scores) if quality_scores else 0.0,
    "max_quality_score": np.max(quality_scores) if quality_scores else 0.0,
}
Out[11]:
Console
=== Synthetic Data Pipeline Results ===

Generated examples:         8
After rule-based filtering: 8 (100% retained)
After quality filtering:    8 (100% retained)

Rule filter breakdown:
  total: 8
  failed_length: 0
  failed_repetition: 0
  failed_refusal: 0
  passed: 8

Quality scores (on rule-passed examples):
  Mean:  0.762
  Min:   0.700
  Max:   0.900

Diversity metrics:
  Before filtering:
    avg_pairwise_distance: 0.838
    vocabulary_diversity: 0.397
    domain_entropy: 0.906
    n_domains: 4
    domain_distribution: {'biology': 3, 'machine_learning': 3, 'earth_science': 1, 'history': 1}
  After filtering:
    avg_pairwise_distance: 0.838
    vocabulary_diversity: 0.397
    domain_entropy: 0.906
    n_domains: 4
    domain_distribution: {'biology': 3, 'machine_learning': 3, 'earth_science': 1, 'history': 1}

The pipeline retains examples that pass both rule-based and quality filters. The diversity metrics show the domain distribution before and after filtering, which helps identify whether quality filtering accidentally removes certain domains disproportionately. A quality filter that is calibrated for one domain (say, technical writing) may penalize short responses from other domains (say, mathematics, where concise answers are appropriate). Checking domain-stratified retention rates as part of the pipeline diagnostic catches this kind of systematic bias early.

Out[12]:
Visualization
Histogram of quality scores from 0 to 1 with a threshold marker at 0.5
Simulated distribution of quality scores assigned to synthetic examples by a heuristic scorer. The vertical dashed line marks the filtering threshold at 0.5. The distribution is bimodal: a cluster of lower-quality examples (shown in red) that the scorer correctly identifies as substandard, and a cluster of higher-quality examples (shown in green) that are retained. Rule-based filters alone would miss many of the lower-quality examples since they pass basic length and format checks; the additional model-based scoring step is needed to surface their semantic weaknesses.

Knowledge Distillation Loss

The implementation below computes the combined distillation loss from Hinton et al., showing how temperature scaling spreads probability mass to non-argmax tokens and how the T2T^2 factor restores the gradient scale.

In[13]:
Code
import numpy as np


def softmax(logits: np.ndarray, temperature: float = 1.0) -> np.ndarray:
    """Compute softmax with temperature scaling."""
    scaled = logits / temperature
    exp_logits = np.exp(
        scaled - np.max(scaled)
    )  # Subtract max for numerical stability
    return exp_logits / exp_logits.sum()


def kl_divergence(p: np.ndarray, q: np.ndarray, eps: float = 1e-9) -> float:
    """KL divergence KL(p || q). Measures how much p differs from q."""
    p = p + eps
    q = q + eps
    return float(np.sum(p * np.log(p / q)))


def distillation_loss(
    student_logits: np.ndarray,
    teacher_logits: np.ndarray,
    true_label_idx: int,
    temperature: float = 2.0,
    alpha: float = 0.5,
) -> dict:
    """
    Compute the combined distillation loss.

    Args:
        student_logits: Student model's raw output scores (before softmax)
        teacher_logits: Teacher model's raw output scores
        true_label_idx: Index of the correct label
        temperature: Softening temperature (T > 1 reveals dark knowledge)
        alpha: Weight for distillation loss vs hard label loss

    Returns:
        Dictionary with loss components
    """
    vocab_size = len(student_logits)

    # Hard label cross-entropy loss (standard supervised learning)
    student_probs = softmax(student_logits, temperature=1.0)
    hard_ce_loss = -np.log(student_probs[true_label_idx] + 1e-9)

    # Soft targets from teacher at temperature T
    teacher_soft = softmax(teacher_logits, temperature=temperature)
    student_soft = softmax(student_logits, temperature=temperature)

    # KL divergence between teacher and student soft distributions
    # Multiplied by T^2 to compensate for the temperature scaling of gradients
    soft_kl_loss = (temperature**2) * kl_divergence(teacher_soft, student_soft)

    # Combined loss
    combined_loss = (1 - alpha) * hard_ce_loss + alpha * soft_kl_loss

    return {
        "hard_ce_loss": hard_ce_loss,
        "soft_kl_loss": soft_kl_loss,
        "combined_loss": combined_loss,
        "teacher_soft_probs": teacher_soft,
        "student_soft_probs": student_soft,
    }


# Example: 4-token vocabulary, teacher is more confident than student
np.random.seed(42)
vocab_size = 4
true_label = 0  # Correct token index

# Teacher: very confident about token 0, some mass on token 1
teacher_logits = np.array([3.5, 1.2, -0.5, -1.0])

# Student: less confident (smaller model, earlier in training)
student_logits = np.array([1.8, 1.5, 0.3, -0.2])

# Compute losses at different temperatures
results_by_temp = {}
for temp in [1.0, 2.0, 4.0]:
    results_by_temp[temp] = distillation_loss(
        student_logits, teacher_logits, true_label, temperature=temp, alpha=0.5
    )
Out[14]:
Console
=== Knowledge Distillation Loss Components ===

Vocabulary size: 4
True label index: 0

Teacher logits:  [ 3.5  1.2 -0.5 -1. ]
Student logits:  [ 1.8  1.5  0.3 -0.2]

Teacher argmax prediction: token 0 (correct: True)
Student argmax prediction: token 0 (correct: True)

 Temperature |  Hard CE Loss |  Soft KL Loss |  Combined Loss
------------------------------------------------------------
         1.0 |        0.7416 |        0.3770 |         0.5593
         2.0 |        0.7416 |        0.6163 |         0.6789
         4.0 |        0.7416 |        0.6389 |         0.6903

Teacher's soft distribution at T=2.0:
  Token 0: 0.6421  ████████████
  Token 1: 0.2033  ████
  Token 2: 0.0869  █
  Token 3: 0.0677  █

Student's soft distribution at T=2.0:
  Token 0: 0.3702  ███████
  Token 1: 0.3187  ██████
  Token 2: 0.1749  ███
  Token 3: 0.1362  ██

At higher temperatures, the soft distributions reveal more about the teacher's relative preferences among tokens, giving the student richer training signal. The combined loss balances learning from ground truth labels (hard CE loss) and from the teacher's full distribution (soft KL loss). Notice that at T=1T=1, the teacher's distribution is highly peaked, and the KL divergence is dominated by the difference in the top token's probability. At T=4T=4, the distribution is spread more evenly, and the student gets substantial gradient signal from tokens 1 and 2, learning that the teacher considers them plausible alternatives.

The effect of temperature on the soft distribution is worth visualizing directly. At T=1T=1, the teacher's distribution is dominated by the correct token. As temperature increases, probability mass spreads to non-argmax tokens, revealing relative plausibility rankings.

Out[15]:
Visualization
Bar charts at 4 temperatures showing probability spreading as T increases
Teacher soft probability distributions at four different temperatures (T=1, 2, 4, 8) over a 6-token vocabulary. At T=1, nearly all probability mass concentrates on token 0 (the correct answer). As temperature increases, the distribution flattens and spreads to neighboring tokens, revealing the teacher's relative preferences among alternatives. This 'dark knowledge' provides richer gradient signal than a hard one-hot label.
Line plot showing scaled KL loss peaking near T=2.3 and then declining, while CE loss remains flat as temperature increases
Distillation loss components as a function of temperature, with alpha=0.5. The hard cross-entropy loss (dashed red) stays constant since it always uses T=1. The T-squared-scaled soft KL loss (solid blue) rises to a peak near T=2.3 and then declines as the teacher and student distributions become more similar at higher temperatures. The combined loss (dotted green) follows the weighted average of both signals.

Key Parameters

The key parameters for the synthetic data pipeline are:

  • Temperature (generation): Controls diversity of teacher outputs. Higher temperatures produce more varied responses but also more errors. Typical range 0.7-1.0.
  • Distillation temperature TT: Controls how much soft knowledge is revealed. T=1T=1 is standard softmax; T=2T=2 to T=4T=4 reveals dark knowledge effectively.
  • Alpha α\alpha: Mixing weight between hard CE loss and soft KL loss. α=0\alpha=0 is pure supervised learning; α=1\alpha=1 is pure distillation.
  • Quality score threshold: Fraction of examples retained after model-based scoring. Typical values 0.5-0.7 depending on teacher quality and task type.
  • Seed diversity: Number of distinct topics, domains, and personas in the seed set. More seeds directly increases output diversity.

Limitations and Practical Considerations

Synthetic data is powerful but not without significant failure modes that practitioners need to understand and actively mitigate.

Hallucination propagation is the most serious concern. When a teacher model generates factually incorrect content, that error enters the training data. A student trained on such data learns the error as fact, and the error compounds across training steps as the student's own generations reinforce the incorrect belief. This is especially dangerous for factual knowledge (historical dates, scientific constants, drug dosages, legal facts) where the teacher may be confidently wrong. Hallucinations in a teacher model's output look exactly like correct outputs: they are fluent, confidently stated, and well-integrated into the response. Standard quality filters are nearly blind to them. Mitigation requires grounding verification: checking generated facts against authoritative sources wherever possible, and using rejection sampling with external validators for tasks that have verifiable ground truth.

Capability ceiling limits what distillation can achieve. A student model cannot systematically exceed the teacher's performance on the tasks the synthetic data covers. The ceiling is the teacher's capability level. To raise the ceiling, you need either a better teacher, a different training signal (human feedback, external verification), or tasks where the student can self-improve (mathematical reasoning with verifiable answers, for instance). This ceiling effect has significant practical implications: the widespread availability of GPT-4 as a teacher model has made 7B-70B student models much more capable, but it also means that all students distilled from GPT-4 will plateau at or below GPT-4's capability level.

Distribution mismatch occurs when synthetic data covers the training distribution well but differs from the deployment distribution. If users ask questions in styles, languages, or domains that the synthetic generation pipeline did not anticipate, the trained model may fail. Real user data collected during deployment provides a valuable complement to synthetic data, even in small amounts. A few thousand real user interactions can reveal distribution gaps that no amount of synthetic data generation would have surfaced, because the synthetic data reflects what the pipeline designers anticipated rather than what users do.

Repetition and template fatigue degrade model quality when too high a fraction of training data comes from a single generation pipeline with a single prompt template. The model learns to mimic the template's patterns, producing outputs that feel formulaic. If every training example begins with "Certainly! Here is a detailed explanation of..." and ends with "I hope this explanation was helpful!", the student learns to produce exactly that framing regardless of context. Rotating templates, mixing generation strategies (Self-Instruct alongside Evol-Instruct alongside backtranslation), and blending synthetic data with natural data all mitigate this effect. Analyzing the most common n-grams in your synthetic dataset is a simple diagnostic: if a handful of phrases appear in more than 5% of examples, your templates are leaving too much of a fingerprint.

Evaluation data contamination is a subtle risk: if the same teacher model that generates training data is also used to evaluate the student, the evaluation may be biased. The teacher's responses set the implicit style standard, and the student's outputs may look good simply because they mimic that style, even when they are wrong in ways the teacher-judge cannot detect. Using held-out human evaluators or independent evaluation benchmarks that the teacher model has not seen is essential for unbiased assessment.

Data mixing balance becomes critical when synthetic data constitutes a large fraction of the training mixture. Models trained on 100% synthetic data tend to have peculiar failure modes: they over-represent the teacher's writing style, they perform well on instruction-following benchmarks (which are themselves based on similar distribution assumptions) but poorly on open-ended real-world tasks, and they sometimes exhibit exaggerated confidence without appropriate uncertainty. Blending synthetic data with a substantial fraction of natural pretraining text maintains the breadth and naturalness that large-scale web crawls provide while adding the targeted capability that synthetic data delivers.

Despite these limitations, synthetic data remains one of the most cost-effective tools for improving model capability. The Phi family of models demonstrated that a carefully curated synthetic dataset of a few billion tokens could produce a model competitive with much larger models trained on much larger natural corpora. Phi-1 achieved state-of-the-art performance on Python coding benchmarks with only 1.3B parameters, trained almost entirely on synthetic "textbook quality" data generated by GPT-4. Phi-2, at 2.7B parameters, matched or exceeded models 5-10x its size on reasoning benchmarks, again largely through synthetic training data. As generation quality and verification pipelines improve, synthetic data will continue to grow in importance across the full spectrum of model training.

With the data curation pipeline complete, Part XXXI turns to training infrastructure: the GPU architecture, memory management, and distributed training strategies needed to train large models on the data assembled across Part XX.

Summary

Synthetic data generation is the practice of using capable teacher models to produce labeled training examples at scale. It works because generation is easier than learning from scratch, because it provides coverage for underrepresented behaviors, because it delivers high-quality labeled examples at low cost, and because it gives practitioners deliberate control over the training distribution. The key ideas covered in this chapter are:

  • Generation methods: Self-Instruct bootstraps from a small seed set using iterative pool expansion; Evol-Instruct evolves instructions to higher complexity through systematic mutations; backtranslation converts high-quality text into instruction-response pairs by generating instructions for existing responses; rejection sampling filters by verifiable correctness, giving a hard quality signal for structured domains
  • Quality verification: Rule-based filters catch obvious failures cheaply before any model calls; model-based scoring with a judge model evaluates semantic quality on a numerical scale; IFEval-style format verification checks compliance with explicit constraints programmatically; reward model scoring provides a learned preference signal; consistency filtering uses agreement across multiple generations as evidence for correctness
  • Diversity: Measured through embedding-based pairwise distance, vocabulary richness, and topic entropy; topic collapse is the most common failure mode; improved through explicit topic-guided generation, persona-based generation for perspective diversity, and iterative self-play with diversity constraints; near-duplicate removal via MinHash prevents redundant examples from dominating the dataset
  • Knowledge distillation: Response distillation trains students on teacher outputs via standard cross-entropy; chain-of-thought distillation transfers reasoning traces rather than just final answers; logit distillation transfers soft probability distributions including dark knowledge; the T2T^2-scaled KL loss combines hard and soft targets; distillation performance degrades when the teacher-student capacity gap is too large, producing style distillation rather than capability distillation

Synthetic data is not a replacement for real data but a complement to it, filling gaps that natural corpora cannot cover and providing targeted training signals for specific capabilities. The most effective training pipelines combine large-scale natural pretraining with carefully curated synthetic fine-tuning data, using each source for what it does best.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about synthetic data generation, quality verification, and knowledge distillation.

Synthetic Data Quiz

Question 1 of 80 of 8 completed
What is the primary purpose of Self-Instruct in synthetic data generation?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026syntheticdata, author = {Michael Brenndoerfer}, title = {Synthetic Data: Generation, Quality, Diversity, Distillation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Synthetic Data: Generation, Quality, Diversity, Distillation. Retrieved from https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm
MLAAcademic
Michael Brenndoerfer. "Synthetic Data: Generation, Quality, Diversity, Distillation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm>.
CHICAGOAcademic
Michael Brenndoerfer. "Synthetic Data: Generation, Quality, Diversity, Distillation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Synthetic Data: Generation, Quality, Diversity, Distillation'. Available at: https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Synthetic Data: Generation, Quality, Diversity, Distillation. https://mbrenndoerfer.com/writing/synthetic-data-generation-quality-diversity-distillation-llm

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.