Part of Language AI Handbook
Construct effective pretraining data recipes by setting domain proportions, applying quality weighting, running proxy experiments.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Data Mixing
Training a large language model from scratch involves more than assembling a massive corpus and pressing go. One of the most consequential decisions you will make is how to mix data from different sources: how much web text versus code, how much scientific writing versus social media, whether to oversample high-quality sources and downsample noisy ones. The mixture you choose shapes what your model knows, how it reasons, and where it fails. Get it wrong, and no amount of compute will fix the resulting gaps.
Data mixing combines empirical science with engineering judgment. Unlike many machine learning hyperparameters, there is no clean analytical formula that tells you the optimal proportions. The problem space is vast, the feedback loop is expensive (you need to train models to see results), and the right answer depends on what you want the model to do. Yet researchers have accumulated enough evidence over the past few years that we now have principled frameworks for thinking about the problem, and a set of algorithmic tools for searching it systematically.
Think of the challenge this way: the internet is not a balanced, curated library. It contains orders-of-magnitude more casual conversation than rigorous technical writing. More marketing copy than scientific analysis. More low-effort content farms than peer-reviewed research. A language model trained on this raw distribution learns to be fluent at what the internet does most: produce casual, chatty, moderately-accurate prose. That may not be what you need. Data mixing is the mechanism for correcting this imbalance.
The stakes are higher than they might first appear. Decisions made at the data mixing stage lock in capability trade-offs that are difficult or impossible to reverse through later fine-tuning. A model that spent most of its pretraining compute reading web crawl will have deeply embedded representations shaped by that distribution. Subsequent instruction tuning can adapt the model's behavior on the surface, but it cannot easily create capabilities that were never built during pretraining. Code generation ability, rigorous multi-step mathematical reasoning, scientific factual grounding, formal writing style: all of these emerge from pretraining, not fine-tuning. Data mixing is where you decide which of these capabilities your model will develop.
This chapter builds directly on the curation pipeline covered in earlier parts of this handbook. By this point, you have raw text that has been crawled from the web (Part XX, chapters 1 through 2), deduplicated using MinHash (chapters 4 through 5), filtered for quality and toxicity (chapters 6 through 7), and scrubbed for PII (chapter 8). What remains is a collection of domain-specific pools, each with different sizes, quality distributions, and relevance profiles. The question now is: how do you sample from these pools when constructing training batches?
Data mixing refers to the process of determining what fraction of each training batch should come from each data source or domain. A mixing recipe specifies proportions such as "30% web text, 20% code, 15% books, 10% scientific papers, and so on." These proportions can remain fixed throughout training or can change dynamically based on model performance.
Domain Proportions
The most fundamental mixing decision is how much weight to assign each broad domain category. Domains are typically coarse groupings like "web," "books," "code," "academic," and "conversational," though practitioners often define finer-grained categories depending on what data they have available. The granularity of domain definition is itself a design choice: coarser categories are easier to reason about and require fewer ablations, while finer categories give you more control but multiply the search space.
The word "proportions" is key here. At any given training step, the model processes a batch of tokens. Those tokens are sampled from the training corpus according to your mixing recipe. If your recipe says 30% web and 20% code, then roughly 30% of tokens in each batch come from web sources and 20% come from code repositories, on average across many batches. The mixing recipe is not about the total size of each corpus but about the probability of sampling from it. This distinction matters enormously: a corpus with 3 trillion tokens does not automatically deserve 3x more weight than a corpus with 1 trillion tokens. Weight reflects intended contribution to the model's learned behavior, not raw data availability.
Why Proportions Matter So Much
A naive approach is to sample uniformly from the combined corpus: each document is equally likely to appear in each training batch regardless of its source. This approach sounds principled, but it produces models that reflect the raw composition of the internet, which is heavily skewed. The web contains vastly more casual social media text than scientific publications. If you sample proportionally to corpus size, your model will have seen millions of Reddit comments for every medical journal article.
Representation is only one part of the problem. Domain composition also affects generalization. A model trained mostly on web text will struggle to follow the precise, structured reasoning expected in code or mathematics, even if it has technically seen some examples of those domains. Conversely, a model trained with an unrealistically high proportion of academic text may generate formal, stilted prose when asked for a conversational reply. The domain mixture shapes the model's default mode of expression, its tolerance for ambiguity, and its sensitivity to precise technical language.
The practical manifestation of this is well documented. The original GPT models, trained heavily on web text, were noticeably better at producing fluent prose than at solving multi-step math problems or writing correct code. Early code generation systems required dedicated code-heavy training data. The progression from GPT-3 to Codex, a code-specialized variant, involved dramatically increasing the fraction of Python and other programming language data, showing just how directly domain composition maps to capability. Adding code data does more than improve code generation: it also improves multi-step reasoning in general, because code imposes a discipline of formal, step-by-step logic that transfers to other domains.
The insight that makes domain mixing work is that you can systematically boost or suppress domains to achieve the capability profile you want. This is analogous to class reweighting in supervised learning, except you are reweighting entire source distributions rather than individual labels. And just as class reweighting in supervised learning has principled solutions (inverse frequency weighting, cost-sensitive learning), data mixing has developed its own set of principled approaches over the past few years.
Understanding the Token Budget
Before going further, it helps to have a clear mental model of what a training run consumes. Modern large language models train on between one trillion and ten trillion tokens. Each training step processes a batch of several thousand tokens. The total number of training steps times the batch size gives you the total token budget. When you specify a mixing recipe, you are specifying how that budget gets allocated across your data pools.
For a concrete example: suppose you are training a 7 billion parameter model on 2 trillion tokens, with a batch size of 4 million tokens per step. You will take 500,000 training steps. If your recipe assigns 20% to code, then across those 500,000 steps, roughly 400 billion tokens will come from code sources. Whether your code corpus has enough distinct tokens to sustain that without excessive repetition depends on its size. If your code pool has only 100 billion distinct tokens, you will see each token an average of 4 times during training. Whether that is a problem depends on the domain and the repetition tolerance, which we discuss in the oversampling section below.
This framing helps you reason about the practical feasibility of your recipe. Before running any ablations, you should calculate the implied oversampling ratio for each domain under your proposed recipe and flag any that look problematic. A recipe that works on paper may be impossible to execute without unacceptable repetition.
Empirical Starting Points
Several large-scale training efforts have published their domain compositions, giving a useful starting point for practitioners.
The Llama 2 training dataset used approximately 89% web text (CommonCrawl and related sources), 8% code from GitHub, and smaller fractions of books, Wikipedia, and other high-quality sources. The model's strong general-purpose performance despite this web-heavy mixture reflects the scale and quality of the filtered web data used. GPT-3 weighted its mixture toward filtered web data (CommonCrawl), with books at about 16%, WebText at 22%, and Wikipedia at 3%. These choices reflected the team's judgment about what capabilities mattered most for general-purpose text generation. The Chinchilla paper from DeepMind, while primarily focused on compute-optimal scaling laws, noted the importance of domain composition but did not publish full mixing weights.
The Falcon models, from the Technology Innovation Institute, trained on RefinedWeb, a carefully filtered subset of CommonCrawl, and found that a high-quality web corpus alone could match models trained on more diverse mixtures at equivalent scale. This finding is important: it suggests that quality filtering within a domain can sometimes substitute for domain diversification. A very clean, carefully filtered web corpus can deliver similar capabilities to a more diverse mixture, as long as the filtering is aggressive enough.
Dolma, the open dataset used to train OLMo, published detailed breakdowns: approximately 38% Common Crawl (post-filtering), 26% The Stack (code), 11% C4, 11% Reddit (filtered), 7% peS2o (scientific papers), and other sources making up the remainder. These proportions reflect deliberate choices about what capabilities the OLMo team wanted the model to emphasize. The 26% code allocation is high for a general-purpose model, showing the team's view that code data contributes broadly to reasoning abilities.
The Mistral and Mixtral models from Mistral AI have published limited information about their training data but are widely believed to use a heavy code fraction, consistent with their strong performance on coding and reasoning benchmarks. Phi-1 and its successors from Microsoft went even further, showing that small models trained almost entirely on high-quality synthetic and curated data can outperform much larger models trained on raw web text. These results challenge the assumption that larger corpora are always better and point toward the importance of data quality over quantity.
These published recipes should not be treated as ground truth. They reflect the goals, compute budgets, and data quality decisions of specific projects. A model intended primarily for coding assistance would weight code much more heavily. A model intended for biomedical applications would boost scientific literature. Use published recipes as baselines to understand the range of reasonable choices, not as universal answers. The right recipe for your model depends on your intended capabilities, your available data, and your compute budget.
The Oversampling Ceiling
When you have a small but high-quality domain (say, books or peer-reviewed papers), you will often want to oversample it. If your total corpus is 2 trillion tokens and your book data is 20 billion tokens, proportional sampling means the model sees book text only 1% of the time. To give books a 10% effective weight, you need to repeat the book data roughly 10 times over training.
Oversampling has diminishing returns. Repeating data exposes the model to the same examples many times, which can cause it to memorize rather than generalize. Practical experience across multiple large training runs suggests that repeating high-quality data up to 2-4 times produces meaningful gains, while oversampling beyond 10x tends to hurt. This creates a ceiling on how much you can compensate for data scarcity by upweighting.
Why exactly does high oversampling hurt? The explanation involves what the model's loss is measuring. Early in training, seeing the same text again provides useful gradient signal because the model has not yet learned to predict it well. But once the model has seen a passage several times, its loss on that passage approaches zero: it has effectively memorized the token sequence. The gradient signal from memorized examples becomes very small, and the model gains little from seeing them again. Worse, those memorized patterns may begin to act as spurious shortcuts that hurt generalization to new text.
To see where the ceiling comes from more concretely, consider how many times each token in domain will be seen during training. If domain has available tokens and you want it to contribute fraction of a total training budget of tokens, the expected number of times each token is seen is:
where:
- : the target weight for domain , a value between 0 and 1 with all domain weights summing to 1
- : the total training token budget for the full run
- : the number of distinct tokens available in domain
When , each token is seen more than once on average, meaning the data is oversampled. When , the model sees only a fraction of the available tokens, meaning the data is undersampled. Keeping below 4 for most high-quality domains is a reasonable empirical heuristic, though more data-scarce domains may tolerate higher repetition if the quality is very high and the corpus is short enough to be largely memorized without harm.
The ceiling creates a practical upper bound on how much you can compensate for small domain size through upweighting. If your books corpus has 20 billion tokens and you want to train for 2 trillion total tokens with books at 15% weight, you are oversampling books roughly 15x. Most practitioners would reduce the books weight to something compatible with 3-4x oversampling, or accept that books will be underweighted relative to their quality because the corpus is simply too small to sustain higher weight without problematic repetition.
There is an important nuance here. Not all oversampling is equally harmful. Data diversity within a domain modulates how much repetition hurts. A books corpus with 80 billion tokens drawn from hundreds of thousands of distinct books is more tolerant of oversampling than a corpus where the same few thousand books dominate. Within-domain diversity means the model sees different examples each time even if aggregate corpus size stays the same. This is one reason why practitioners care about deduplication within domains: removing near-duplicate documents effectively increases diversity and raises the oversampling ceiling.
Cross-Domain Interaction Effects
Domain proportions interact with each other in non-trivial ways. Adding more code data does not just improve coding performance: it typically also improves mathematical reasoning, logical deduction, and instruction following. Adding more high-quality web text does not just improve general language: it can also improve factual grounding because web text contains references to real-world entities and events. These spillover effects mean that the optimal mixing recipe is not a simple sum of domain-specific contributions.
Understanding these interactions requires controlled ablations. The standard approach is to hold all domains constant except one, vary that domain's weight across a range of values, and measure the impact on all capability dimensions. Repeating this for each domain gives you a sensitivity profile: how much does each capability change when you add or remove a given domain? Domains with large spillover effects are worth weighting more heavily than their primary-capability contribution alone would suggest.
Code is perhaps the best-documented case of cross-domain spillover. Multiple studies have shown that training on code, even for models that will not be used for coding tasks, improves performance on multi-step reasoning benchmarks, mathematical problem solving, and structured generation tasks. The hypothesis is that code imposes constraints that other domains do not: programs must be syntactically valid, variables must be defined before use, and function calls must match their signatures. Learning to model this kind of structured, rule-governed text trains the model to apply similar discipline in other contexts.
Quality Weighting
Domain labels are coarse. Within a single domain like "web crawl," you will find everything from carefully written encyclopedia articles to spam-filled link farms. Quality weighting addresses this heterogeneity by assigning individual documents or source sites different sampling probabilities based on estimated quality. While domain proportions determine how much weight each broad category receives, quality weighting determines which specific documents within each category get sampled most frequently.
The motivation here is efficiency. If 80% of your web crawl is low-quality noise, sampling uniformly from it means most of your training budget goes toward learning patterns that are not worth learning. Quality weighting lets you concentrate training on the informative 20%. In the limit, sufficiently aggressive quality filtering can produce a smaller, denser corpus that outperforms a much larger unfiltered corpus at the same training compute. The FineWeb dataset from Hugging Face demonstrated this trade-off: careful quality filtering of CommonCrawl produced a dataset where training on the filtered subset outperformed training on the full original despite having far fewer tokens.
Quality is not binary. The useful question is how much to sample from a document, rather than whether to include or exclude it outright. A document that scores 0.6 on a quality scale probably provides some learning signal; discarding it entirely wastes potentially useful data. Treating quality as a continuous weighting factor and adjusting sampling probability smoothly across the quality spectrum is generally more efficient than hard inclusion/exclusion thresholds.
Classifier-Based Quality Scores
One approach to quality weighting uses a classifier trained to distinguish high-quality text from low-quality text. The classifier assigns each document a score, and documents with higher scores are sampled more frequently.
The training data for such classifiers is typically constructed from contrasting examples. Wikipedia and books are treated as positive examples of high-quality text. Raw web crawl data is treated as negative. A simple binary classifier trained on these examples learns to identify the stylistic and structural features that distinguish clean, informative writing from noise: well-formed sentences, consistent punctuation, appropriate paragraph structure, factual density, absence of navigation menus and boilerplate.
GPT-3 used this approach: the OpenAI team trained a classifier to identify text similar to WebText (a curated collection of web pages linked from Reddit posts that received substantial karma, serving as a proxy for human-judged quality), and then used the classifier scores to filter and weight their CommonCrawl corpus. This single step dramatically improved the effective quality of their web data. The classifier learned to separate the long tail of low-quality web text from the informative content that made up the high-karma-link-target pages.
FineWeb, released by Hugging Face in 2024, takes a more sophisticated approach. Rather than a binary classifier, it applies a cascade of heuristics and learned filters, then evaluates the resulting corpus by training small proxy models and measuring performance on benchmarks. This evaluation-driven approach lets you see the impact of quality filtering choices without committing to a full training run. The FineWeb work also highlighted an important finding: different quality filters are not always additive. Some filters that look good in isolation degrade performance when combined, because they remove documents that carry specific types of useful signal even if those documents score low on some quality dimensions.
One important limitation of classifier-based approaches is that they are bounded by the quality of your positive examples. If your "high quality" training signal comes from Wikipedia and English-language books, your classifier will favor text that resembles Wikipedia and English books. It will down-score technically rigorous but stylistically unusual text, multilingual content, and anything that follows conventions different from your reference corpus. A data science tutorial written in informal, conversational style may score low even though it contains valuable technical content. A legal document in precise but arcane language may be penalized. This is not always a problem, but it is worth keeping in mind when working with diverse data sources, especially when building models intended for non-English languages or specialized technical domains.
URL-Based and Source-Based Weighting
An even simpler form of quality weighting assigns weights based on the source of a document rather than its content. Documents from known high-quality sources (academic publishers, major newspapers, authoritative encyclopedias) receive higher weights, while documents from low-quality domains (spam sites, link farms, autogenerated content) receive lower weights or are discarded entirely.
This approach is computationally cheap because it requires only a domain-name lookup, not content analysis. When you are working with a corpus of hundreds of billions of documents, the ability to apply quality weighting without processing each document individually is operationally important. Source-based weighting can be applied as a preprocessing step that drastically reduces the amount of content requiring per-document scoring.
The limitation of source-based weighting is that it applies a single quality score to all content from a given source, even when that source produces variable-quality content. A news site publishes both in-depth investigative journalism and listicle fluff. A university website hosts rigorous research papers alongside undergraduate course syllabi. A programming forum contains expert answers alongside incorrect, confused questions. Source-based weighting cannot distinguish between these. You are averaging quality across a source's entire output, which may wash out the best content from heterogeneous sources.
The C4 dataset used a set of heuristic filters including discarding pages from a blocklist of known low-quality sites. RedPajama constructed per-domain quality scores by estimating the fraction of content from each domain that survived multiple quality filters. This creates a hybrid: the quality score of a domain is derived from content-level filtering, but applied at the domain level for efficiency. Documents from a domain where 90% of content passes filters receive higher base weights than documents from a domain where only 30% of content passes, even before per-document scoring is applied.
Perplexity-Based Filtering
A third quality signal uses a small, high-quality language model to score documents. Documents that the model assigns low perplexity to (meaning they are coherent and predictable by that model) are treated as higher quality. Documents that receive high perplexity from the reference model are considered less coherent and may be downweighted or removed.
Perplexity measures how surprised a model is by a given piece of text. A well-formed, coherent sentence follows predictable linguistic patterns, and a language model trained on good text will assign it low perplexity. Fragmented, noisy, or incoherent text violates those patterns and receives high perplexity. The intuition is that text which is surprising even to a model trained on good data is probably not good data itself.
To make this concrete, consider a document containing sentences like: "The algorithm computes optimal solutions using dynamic programming techniques developed by Bellman in the 1950s." A language model trained on scientific and technical writing will assign this low perplexity because the vocabulary, syntax, and topic are all familiar. Now consider a spam document: "Click here!!! Buy now!!! Limited time!!! Amazing deals!!!" A good language model will assign this high perplexity not because the words are unusual but because the pattern is erratic and does not follow the coherent structure the model has learned to expect.
This approach has a subtle bias: it favors text that resembles whatever was used to train the scoring model. If your scoring model was trained on English news articles, it will assign low perplexity to similar text and high perplexity to code, non-English text, or unusual writing styles that might be valuable. Perplexity-based filtering is most effective when combined with other signals and when the scoring model is specifically calibrated for the domains you care about. Using a code-specialized model to score code documents is more appropriate than using a general-purpose English model.
There is also a practical consideration about the size of the scoring model. A scoring model needs to be small enough to process your entire corpus efficiently. Scoring hundreds of billions of documents with a 7-billion-parameter model is feasible but requires careful infrastructure design. Most practitioners use 100-million to 1-billion-parameter models for perplexity-based scoring, accepting some loss of scoring accuracy in exchange for computational tractability.
Combining Quality Signals
In practice, effective quality weighting combines multiple signals. A document's final sampling probability might be a product of factors including its domain weight, its classifier score, its source URL reputation, and any hard quality thresholds it survived during the earlier filtering pipeline described in prior chapters.
The combination function matters. Multiplicative combinations, where you multiply all factors together, tend to aggressively discard documents that score poorly on any single dimension, which can over-filter. A document that scores 0.4 on the classifier but 0.9 on URL reputation and 0.8 on perplexity might be acceptable, but multiplicative combination would assign it a combined score of around 0.29. The document gets heavily penalized for one weak signal even though two other signals rate it positively. Additive combinations or learned combination functions are sometimes more robust because they allow high scores on some dimensions to compensate for lower scores on others.
Threshold-based approaches (keep everything above 0.5, discard the rest) are simpler but discard potentially useful intermediate-quality content. A document at 0.49 and a document at 0.5 are nearly identical in quality, but a hard threshold treats them very differently. A continuous weighting scheme that smoothly adjusts sampling probability across the quality spectrum is generally preferable to hard thresholds, because it allows the model to learn from borderline content at reduced frequency rather than wasting that content entirely.
One important practical consideration: quality signals are expensive to compute, and you typically want to compute them once and store them as metadata attached to each document in your corpus. The data mixing system then uses these precomputed scores at training time. This separation between the scoring pipeline and the training pipeline simplifies both: scoring can run asynchronously on whatever hardware is convenient, and training just needs to look up the stored scores and apply the sampling probabilities.
Data Mixing Experiments
Determining the optimal mixing recipe empirically requires a principled experimental methodology. You cannot afford to train full-scale models for every combination you want to test. Even at the scale of a few hundred experiments, the compute cost of evaluating each configuration with a full training run would exceed the budget of most projects. Instead, the field has converged on a workflow centered around proxy evaluations at reduced scale.
The scale of the search space makes careful experimental design essential. If you have 10 domain categories and want to test mixing proportions at 5% increments, you have combinatorially many candidate mixtures. Even with proxy models, evaluating more than a few dozen configurations is expensive. Researchers have developed several strategies for searching this space efficiently, from hand-crafted ablation designs to automated search algorithms.
The Proxy Model Paradigm
The main insight allowing practical data mixing research is that the relative ranking of data mixtures tends to be consistent across model scales. A mixture that produces a better model at 1 billion parameters also tends to produce a better model at 70 billion parameters, even if the absolute performance numbers differ. This transfer property is what makes proxy experiments useful: you can identify good mixtures cheaply at small scale and trust that they will remain competitive at full scale.
This insight enables a much cheaper experimental workflow. You train small proxy models (typically 100 million to 1 billion parameters) on candidate data mixtures, then evaluate them on a benchmark suite covering the capabilities you care about, then rank the mixtures by proxy model performance, and finally use the top-ranked mixture for the full-scale training run. This workflow reduces the cost of each experiment by orders of magnitude compared to full-scale evaluation.
The evaluation suite should cover the capabilities you care about. If your target application is coding, your benchmarks should include coding tasks like HumanEval or MBPP. If you care about factual accuracy, include knowledge-intensive QA datasets like Natural Questions or TriviaQA. If reasoning matters, include mathematical benchmarks like GSM8K or MATH. Using a diverse benchmark suite reduces the risk of overfitting your mixture to a narrow set of tasks. A mixture that maximizes one benchmark while hurting others may not be the one you want.
The proxy model approach has important limits. The scale transfer assumption holds on average across many experiments, but individual exceptions exist. Some capabilities, particularly those associated with emergent abilities that appear suddenly at certain scale thresholds, simply do not manifest at proxy scale. At 100 million parameters, no mixing recipe triggers strong chain-of-thought reasoning or sophisticated multi-step problem solving. This makes it impossible to optimize for these capabilities through proxy experimentation alone. For emergent capabilities, you are essentially flying blind on mixing decisions and must rely on intuitions carried over from other design choices.
Sample-Efficient Ablations
Even proxy model training has a cost. To test many mixture configurations efficiently, researchers typically run very short training runs, sometimes just 1 to 5% of the full compute budget, and use the resulting loss curves as an early signal of mixture quality.
The loss at intermediate training steps reflects how quickly the model is learning from each domain. A mixture where all domains improve steadily typically indicates a well-balanced recipe. A mixture where some domains stagnate while others improve rapidly may indicate imbalance: the model is spending too little time on the stagnating domains to make progress there. Short training runs can surface these differences early enough to be actionable without committing to the full compute cost of a complete proxy run.
The risk with extremely short runs is that some data mixing effects only manifest after extended training. A source that is helpful in the long run might not show its benefit within the first few billion tokens, either because the benefit compounds over many exposures or because the capability in question requires a complex integration of skills that only develops after sustained training. This is one reason practitioners supplement short-horizon ablations with a small number of longer validation runs to check that short-horizon rankings hold, particularly for their top few candidate recipes.
Another practical technique for sample-efficient exploration is variance reduction through ensembling. Instead of running a single proxy training run for each candidate recipe, running 3 to 5 independent runs with different random seeds and averaging their performance gives you lower-variance estimates of each mixture's quality. This allows you to distinguish better recipes from ones that just happened to get lucky in one run. The extra cost is usually worth it when you are comparing recipes that perform similarly in your initial screening.
DoReMi: Distributionally Robust Mixing
DoReMi (Domain Reweighting with Minimax Optimization) is a more principled algorithm for data mixing that frames the problem as a distributionally robust optimization. Rather than manually testing mixing ratios and relying on human judgment to interpret the results, DoReMi uses a two-model approach that automatically finds mixing weights optimized against a formal objective.
The algorithm's goal is to minimize the worst-case excess loss across all domains. The intuition is similar to how a student should allocate study time: spend more effort on weak subjects, less on subjects already mastered. Domains where the proxy model performs poorly relative to the reference model receive higher weights, and domains where it already performs well receive lower weights. This continues iteratively until the proxy model's performance is as good as possible across all domains simultaneously, rather than being excellent on easy domains while struggling on hard ones.
The algorithm begins by training a reference model with uniform domain weights for some number of steps. This reference model establishes a baseline performance level for each domain under the null hypothesis of equal domain importance. Then a proxy model is trained with adaptive weights that are updated based on where the proxy is falling behind the reference.
The weight update at each step works by computing an excess loss signal for each domain. For domain , the excess loss measures how much worse the proxy model is compared to the reference:
where:
- : the average token-level cross-entropy loss of the proxy model on held-out data from domain
- : the average token-level cross-entropy loss of the reference model on the same held-out data
A positive means the proxy is struggling on domain relative to the reference. A negative value means the proxy has caught up or surpassed the reference on that domain. When the proxy has been trained with too little weight on a domain, its loss on that domain will remain high relative to the reference, generating a large positive and triggering a weight increase.
The weight update then uses an exponentiated gradient step to shift weight toward domains with higher excess loss:
where:
- : the weight assigned to domain at iteration
- : the step size controlling how aggressively weights shift based on the excess loss signal
- : the exponential function, which ensures weights remain positive throughout the update
After the update, weights are renormalized to sum to 1 by dividing each by the total:
This normalization step projects the updated weights back onto the probability simplex, the set of non-negative vectors that sum to 1, making sure the mixing recipe remains a valid probability distribution after each update step. Without this normalization, repeated exponentiation would cause weights to grow without bound.
The exponentiated gradient update has a nice property: it never sets any weight to exactly zero, so no domain is ever completely excluded. Domains that are consistently easy for the proxy model receive very small but nonzero weights, making sure the model retains at least some exposure to those domains throughout training. This softmax-like behavior is more desirable than an update rule that could collapse weights to zero, because it maintains the diversity of the training distribution even as it concentrates weight on the hardest domains.
The exponentiated gradient family of updates is well-studied in the online learning literature under the name "multiplicative weights" or "Hedge." DoReMi inherits theoretical guarantees about convergence toward the minimax-optimal weights under appropriate conditions. Specifically, the regret of the algorithm (how much worse it performs compared to the optimal static weight vector) grows as where is the number of update steps and is the number of domains. This guarantee holds regardless of how the domain losses evolve, making DoReMi robust to adversarial or unexpected loss dynamics.
The published DoReMi results showed meaningful improvements over manually designed mixing recipes on several benchmarks, particularly for domains that were systematically under-represented in human-designed recipes. The automated nature of the approach is also operationally appealing: it reduces the number of manual decisions you need to make and makes the mixing recipe more reproducible across different practitioners and projects.
DoReMi-2 and Successive Halving
Subsequent work extended the DoReMi framework and combined it with techniques from hyperparameter search. DoReMi-2 incorporates successive halving: rather than evaluating all domain candidates throughout training, the approach progressively eliminates the worst-performing configurations and reallocates compute to more promising ones. This is analogous to Hyperband or similar adaptive hyperparameter search methods. By discarding clearly inferior configurations early, successive halving concentrates the experimental budget on the part of the search space where the best answer is likely to live.
These techniques collectively shift data mixing from a manual art toward an automated optimization problem. You specify your domains, your evaluation benchmarks, and your compute budget, and the algorithm finds a near-optimal mixing recipe without requiring dozens of hand-designed experiments. The automation is not perfect: the proxy-to-full-scale transfer assumption still needs to hold, and the benchmarks you choose still encode your implicit preferences. But the process is substantially more systematic and less reliant on individual practitioner intuition than earlier approaches.
Optimal Mixing
"Optimal" mixing depends on what you are optimizing for, and that target is not always obvious. A model optimized for coding performance will sacrifice some general language quality. A model optimized for factual accuracy will see less conversational text and may produce stiffer responses. The optimal mixing recipe is never optimal in all dimensions simultaneously; it represents a set of tradeoffs chosen by the practitioners who design it. Framing the question correctly, as "optimal for what?" rather than "optimal in general," is the first step toward making principled mixing decisions.
Capability-Proportional Mixing
One philosophy is to mix data in proportion to how much each source contributes to the capabilities you want. This sounds circular but is operationalized through benchmark decomposition: run ablations that train models with one domain removed at a time, measure the capability drop on each benchmark, and use those drops to set domain weights. Domains that cause large capability drops when removed should receive more weight; domains with negligible impact can be reduced.
For example, if removing code data causes a large drop on programming benchmarks and a smaller drop on general language understanding, you know that code data is important for coding but has spillover benefits for general reasoning. You can use these marginal contribution estimates to find a weight that delivers acceptable coding performance without sacrificing too much on other tasks. The relative magnitudes of the drops across domains give you a gradient that points toward a better recipe.
The ablation approach reveals direct contributions and spillover effects. Code data often improves mathematical and logical reasoning beyond coding tasks, because code requires formal, step-by-step reasoning that transfers to other domains. Scientific papers improve factual grounding beyond scientific question answering because the writing style trains the model to make precise, verifiable claims. Understanding these transfer effects helps you make more informed weight decisions. A domain that has high spillover is often worth weighting more heavily than its benchmark-specific contribution alone would suggest.
One challenge with leave-one-out ablations is that they capture marginal effects at zero weight (what happens when you go from some weight to zero), which may not accurately predict the effect of smaller weight changes around your current recipe. A domain that shows large effect when completely removed might show a smaller marginal effect when you just reduce its weight by 5%. More granular ablation designs that vary each domain's weight over several levels give you a richer picture of the model's sensitivity, at the cost of running more experiments.
Task-Specific Weighting
When you have a specific target task or application in mind, you can use direct performance on that task as your optimization objective. If you are building a coding assistant, you care primarily about performance on coding benchmarks. If you are building a medical question-answering system, you care primarily about performance on medical factual accuracy tasks. You can design your mixing recipe explicitly to maximize performance on the tasks you care about, using proxy model ablations to evaluate each candidate recipe on those specific tasks.
This approach naturally incorporates the intended use case into the data curation process. It is the most direct path from "what do I want the model to do" to "what data mixture achieves that." The risk is that optimizing too narrowly for your target tasks may degrade performance on tasks you did not explicitly include in your evaluation suite. A model tuned entirely for medical question answering may lose the general language flexibility needed for the instruction-following and explanation tasks that come up in real medical deployments. Using a target-task-focused objective works best when you are building a highly specialized model where the narrow capability range is a feature rather than a limitation.
Diversity-Aware Mixing
A complementary view is that diversity itself is a goal. Models trained on homogeneous data, even if that data is very high quality, tend to fail in predictable ways on out-of-distribution inputs. Maintaining a minimum floor of coverage across diverse domains guards against blind spots that are difficult to anticipate in advance.
This manifests practically as constraints on the mixing optimization: no domain should fall below some minimum weight, regardless of what the quality-maximizing or capability-maximizing solution would suggest. The minimum floor values are usually set by heuristic, typically 1 to 5% of total training tokens, rather than formal optimization.
The intuition behind diversity floors is related to the model's ability to handle unexpected inputs at deployment. A model that has never seen certain kinds of text during training will have degraded performance on those inputs in a way that is difficult to predict in advance. By ensuring minimum exposure to each domain, you hedge against unknown unknowns. You do not know exactly what users will ask at deployment. Diversity floors ensure the model has at least encountered a wide range of topics, genres, styles, and registers, even if it has not deeply mastered all of them.
There is also a subtler benefit to diversity: it improves robustness to distribution shift. When training distribution and deployment distribution differ, which they almost always do, models trained on more diverse data tend to generalize better to the deployment distribution because they have been exposed to a wider range of input patterns. A model trained on only web text and code may struggle with the kinds of formal instructions, edge-case phrasings, and unusual query structures that real users produce. Including diverse source types, even at low weight, improves the model's overall robustness.
Curriculum Effects
Both the order in which domains appear during training and the overall proportions matter. A curriculum that starts with high-quality, structured text (books, Wikipedia) and then gradually introduces more varied sources often outperforms random shuffling of the same mixture.
The intuition is that language models benefit from first establishing strong grammatical and factual foundations before being exposed to noisy or stylistically unusual text. Early in training, the model is learning basic language patterns: how sentences are structured, how paragraphs flow, what vocabulary means in context. High-quality text teaches these patterns more efficiently. Later in training, after the basic representations are established, the model can extract useful signal even from noisier sources, because it can filter that signal through the representations it has already built.
If the first billion tokens a model sees are spam or grammatically mangled text, subsequent learning may be harmed more than if those same tokens appear later in training when the model has already formed basic representations. The model's early gradient steps are particularly influential in shaping the feature representations that all subsequent learning builds on. A poorly initialized representation structure is hard to correct later in training, because the model's accumulated knowledge is encoded in terms of those early representations.
The analogy to human education is imperfect but instructive. Children do not learn to read by being handed a dictionary and a noisy social media feed simultaneously. They start with structured, clear text and encounter complexity gradually. There may be a similar principle at work in neural network training, where early gradient steps are particularly influential in shaping the learned representations that all subsequent learning builds on. The analogy breaks down at some point (children are not language models), but the basic intuition that early training has disproportionate influence on final representations is consistent with empirical findings.
This is not universally accepted. Some practitioners argue that shuffling reduces overfitting to the early-seen distribution and that curriculum effects are overstated or not reproducible across different training setups. The evidence is mixed, and the right choice likely depends on the specific domains involved and the scale of training. Curriculum effects may be more significant at smaller scales where the model has less capacity to learn representations from noisy data, and may diminish at very large scale where the model eventually sees enough data to correct early impressions.
Compute-Optimal Mixing
The Chinchilla paper established that for a given compute budget, there exists an optimal balance between model size and training tokens. Data mixing adds another dimension to this optimization: for a given compute budget and model size, there is also an optimal allocation of those training tokens across domains.
Recent work by researchers at EleutherAI and elsewhere has explored whether Chinchilla-like scaling laws can be extended to cover domain mixing. The challenge is that the loss decomposition across domains is coupled in complex ways: improving performance on one domain may change the model's representations in ways that help or hurt other domains. The gradient updates from code training interact with the representations built from book training. A joint scaling law that accounts for both compute budget and domain composition would be extremely useful, but has not yet been derived in a fully satisfactory form. The problem is harder than single-domain scaling because the interaction effects are hard to characterize analytically.
A pragmatic approach is to treat the compute-optimal model size and training token count derived from Chinchilla scaling as fixed constraints, and then use DoReMi or proxy model ablations to find the best mixing ratio within that envelope. This decouples the two optimization problems and makes both more tractable. You use scaling laws to determine how big your model should be and how long to train, then use data mixing experiments to determine the best composition of that training budget.
Dynamic Mixing
All strategies discussed so far assume a fixed mixing recipe throughout training. Dynamic mixing adjusts proportions during the training run, typically based on how the model's performance evolves. The flexibility to change proportions during training opens up strategies that are impossible with fixed recipes, but it also introduces new risks and engineering complexity.
One simple variant is two-phase training: use a diversified mixture for most of training, then switch to a quality-focused mixture for a final fine-tuning phase. This approach appears in several large model training recipes, where a quality-heavy final phase (sometimes called "cooldown" or "annealing") is applied to improve instruction-following and factual accuracy. The idea is that broad diversity is valuable early in training when the model is building general representations, while quality focus is more valuable late in training when the model is refining its knowledge and aligning its outputs with human preferences. The final phase effectively specializes the model's behavior without discarding the broad capabilities built earlier.
A more sophisticated variant continuously monitors benchmark performance on held-out data and adjusts domain weights in response, essentially running a real-time version of the DoReMi algorithm throughout training. As the model improves on easy domains, weights on those domains are automatically reduced. As the model stagnates on hard domains, weights on those domains increase. This requires infrastructure to evaluate the model periodically during training, which adds engineering overhead but can improve final model quality by keeping the training distribution well-matched to the model's current state.
Dynamic mixing raises an important technical challenge: stability. When you shift domain weights during training, the data distribution the model is learning from changes suddenly. This distributional shift can cause training instability, particularly if the shift is large or abrupt. The optimizer's internal state (momentum in Adam, for example) reflects the gradient statistics of the previous distribution, and sudden distribution changes can make that accumulated state misleading. Practitioners typically apply gradual transitions rather than hard switches, linearly interpolating between the old and new mixing recipe over some number of training steps to give the optimizer time to adapt.
The learning rate schedule interacts with dynamic mixing in ways that are not fully understood. Many training recipes use cosine decay or linear decay of the learning rate, calibrated under the assumption of a fixed data distribution. When the data distribution changes partway through training, the previously calibrated schedule may be suboptimal for the new distribution. This is an active area of research, and current practice relies more on engineering intuition than on principled guidance.
Code Implementation
Let us work through a simulation that captures the key ideas in data mixing. We will implement a simplified mixing framework that demonstrates domain proportion effects, quality weighting, and the proxy model evaluation paradigm.
This simulation uses synthetic capability scores rather than real model training, which makes it fast enough to run interactively. The core dynamics, including the oversampling ceiling, the advantage of curated recipes over proportional sampling, and the DoReMi weight update, are accurately represented. The synthetic model captures the most important qualitative behaviors while abstracting away the engineering complexity of actual large-scale training.
Setting Up Domain Profiles
We start by defining synthetic domain profiles that represent a realistic pretraining corpus.
import numpy as np
# Define domain profiles with realistic characteristics
# Each domain has: corpus size (tokens), base quality score, capability contributions
domains = {
"web_common_crawl": {
"size_billions": 3000, # Very large pool
"quality": 0.4, # Mixed quality
"coding": 0.05,
"reasoning": 0.3,
"facts": 0.4,
"language": 0.5,
},
"web_filtered": {
"size_billions": 500,
"quality": 0.7,
"coding": 0.1,
"reasoning": 0.5,
"facts": 0.6,
"language": 0.7,
},
"books": {
"size_billions": 80,
"quality": 0.9,
"coding": 0.05,
"reasoning": 0.7,
"facts": 0.7,
"language": 0.95,
},
"wikipedia": {
"size_billions": 20,
"quality": 0.95,
"coding": 0.1,
"reasoning": 0.6,
"facts": 0.95,
"language": 0.85,
},
"code_github": {
"size_billions": 300,
"quality": 0.75,
"coding": 0.98,
"reasoning": 0.6,
"facts": 0.2,
"language": 0.3,
},
"scientific_papers": {
"size_billions": 60,
"quality": 0.88,
"coding": 0.2,
"reasoning": 0.85,
"facts": 0.9,
"language": 0.75,
},
"reddit_filtered": {
"size_billions": 150,
"quality": 0.55,
"coding": 0.15,
"reasoning": 0.35,
"facts": 0.3,
"language": 0.6,
},
}
total_corpus = sum(d["size_billions"] for d in domains.values())
proportional_weights = {
k: v["size_billions"] / total_corpus for k, v in domains.items()
}Total corpus size: 4,110B tokens Proportional weights (sampling uniformly by size): web_common_crawl : 0.730 (73.0%) web_filtered : 0.122 (12.2%) books : 0.019 (1.9%) wikipedia : 0.005 (0.5%) code_github : 0.073 (7.3%) scientific_papers : 0.015 (1.5%) reddit_filtered : 0.036 (3.6%)
When sampling proportionally to corpus size, the 3 trillion token CommonCrawl pool dominates the mixture at over 72% by default. Every other domain, including books and Wikipedia, receives a tiny fraction. If training used these natural proportions, the model would rarely encounter the highest-quality content in the corpus. Books would account for less than 2% of all training tokens, meaning the model spends the vast majority of its budget learning patterns from content of much lower average quality.


Simulating Capability Scores
Now we simulate how different mixing recipes affect model capabilities using a simplified proxy model. The simulation captures the most important dynamics: capability improves with exposure to relevant domains, quality modulates how much each exposure is worth, and oversampling has diminishing returns after a threshold.
def simulate_capability(mixing_weights, noise_std=0.02, rng=None):
"""
Simulate capability scores under a given mixing recipe.
This is a toy model: real capability depends on many more factors.
We capture the core dynamic: each domain contributes in proportion to its
sampling weight, quality, and relevance to the capability being measured.
"""
if rng is None:
rng = np.random
capability_scores = {
"coding": 0.0,
"reasoning": 0.0,
"facts": 0.0,
"language": 0.0,
}
for domain, weight in mixing_weights.items():
quality_adjusted_weight = weight * domains[domain]["quality"]
for cap in capability_scores:
capability_scores[cap] += (
quality_adjusted_weight * domains[domain][cap]
)
# Map weighted contributions into a plausible proxy-score range, then add
# small run-to-run variation.
capability_scores = {
k: float(np.clip(0.45 + 0.90 * v + rng.normal(0, noise_std), 0, 1))
for k, v in capability_scores.items()
}
return capability_scores
def overall_score(caps):
"""Weighted aggregate capability score."""
weights = {
"coding": 0.25,
"reasoning": 0.30,
"facts": 0.25,
"language": 0.20,
}
return sum(caps[k] * weights[k] for k in caps)Comparing Mixing Recipes
Let us compare three mixing strategies: proportional (sampling by corpus size), curated (a practitioner's informed guess), and quality-weighted (sampling proportionally to domain quality scores).
# Strategy 1: Proportional (baseline)
recipe_proportional = proportional_weights.copy()
# Strategy 2: Curated (typical LLM recipe, broadly similar to Llama 2)
recipe_curated = {
"web_common_crawl": 0.10,
"web_filtered": 0.35,
"books": 0.15,
"wikipedia": 0.05,
"code_github": 0.20,
"scientific_papers": 0.10,
"reddit_filtered": 0.05,
}
# Strategy 3: Quality-weighted (boost quality, penalize noise)
quality_scores = np.array([domains[d]["quality"] for d in domains])
quality_weights = quality_scores / quality_scores.sum()
recipe_quality = dict(zip(domains.keys(), quality_weights))
# Evaluate each recipe
recipe_rng = np.random.RandomState(0)
runs_per_recipe = 20
results = {}
for recipe_name, recipe in [
("Proportional", recipe_proportional),
("Curated", recipe_curated),
("Quality-Weighted", recipe_quality),
]:
scores = [
simulate_capability(recipe, rng=recipe_rng)
for _ in range(runs_per_recipe)
]
aggregate = [overall_score(s) for s in scores]
cap_means = {
k: np.mean([s[k] for s in scores])
for k in ["coding", "reasoning", "facts", "language"]
}
results[recipe_name] = {
"mean": np.mean(aggregate),
"std": np.std(aggregate),
"caps": cap_means,
}Recipe Comparison (proxy model simulation): Recipe Overall Coding Reasoning Facts Language ----------------------------------------------------------------- Proportional 0.616+-0.011 0.528 0.621 0.639 0.692 Curated 0.800+-0.011 0.638 0.847 0.827 0.898 Quality-Weighted 0.841+-0.011 0.614 0.885 0.920 0.957
The proportional recipe underperforms on all capability dimensions because the model spends most of its training budget on lower-quality, lower-diversity CommonCrawl text. The curated recipe shows the biggest gain in coding because it deliberately includes much more code. The quality-weighted recipe leads on reasoning and factual accuracy because it favors books, Wikipedia, and scientific papers. Notice that no single recipe dominates all capability dimensions: curated mixing wins on coding, quality weighting wins on reasoning and facts, and both beat proportional on language quality.

Simulating the Oversampling Ceiling
We now look at what happens when we push oversampling to its limits for the books domain. This experiment varies the weight assigned to books from 1% to 40% and measures how overall capability and language quality respond as the oversampling ratio increases.
# Vary how much we oversample books while keeping other weights proportional
book_weights = [0.01, 0.05, 0.10, 0.15, 0.20, 0.30, 0.40]
oversample_results = []
oversample_rng = np.random.RandomState(1)
for book_weight in book_weights:
# Distribute remaining weight proportionally among other domains
remaining = 1.0 - book_weight
other_domains = {
k: v for k, v in proportional_weights.items() if k != "books"
}
other_sum = sum(other_domains.values())
recipe = {k: (v / other_sum) * remaining for k, v in other_domains.items()}
recipe["books"] = book_weight
# Calculate actual oversample ratio for books domain
# (books has 80B tokens, we're training on 1000B total)
oversample_ratio = (book_weight * 1000) / domains["books"]["size_billions"]
scores = [
simulate_capability(recipe, rng=oversample_rng) for _ in range(30)
]
# A compact toy penalty represents memorization displacing generalization
# once the same book tokens have been repeated more than about three times.
repeat_penalty = 0.015 * max(oversample_ratio - 3.0, 0.0) ** 2
mean_score = np.mean([overall_score(s) for s in scores]) - repeat_penalty
language_score = np.mean([s["language"] for s in scores]) - repeat_penalty
oversample_results.append(
{
"book_weight": book_weight,
"oversample_ratio": oversample_ratio,
"mean_score": mean_score,
"language_score": language_score,
}
)Effect of oversampling books:
Book Weight Oversample Ratio Overall Score Language Score
--------------------------------------------------------------
1% 0.1x 0.615 0.683
5% 0.6x 0.630 0.705
10% 1.2x 0.643 0.729
15% 1.9x 0.659 0.763
20% 2.5x 0.676 0.792 <-- sweet spot
30% 3.8x 0.696 0.834 <-- sweet spot
40% 5.0x 0.676 0.834The results confirm the oversampling ceiling: language scores improve as you oversample books up to around 3 to 4 times, but gains plateau and may slightly reverse at very high ratios. The mechanism is that repeated exposure to the same book tokens eventually causes the model to memorize rather than learn generalizable patterns. The gradient signal from well-memorized text approaches zero, and those training steps contribute nothing to the model's ability to generalize.

Implementing a Simple DoReMi Step
We simulate one step of the DoReMi weight update to illustrate the algorithm's behavior. The key thing to observe is how the update naturally directs weight toward domains where the proxy model is struggling.
def doremi_weight_update(domain_weights, proxy_losses, ref_losses, eta=0.1):
"""
One step of the DoReMi exponentiated gradient update.
Increases weight for domains where proxy model struggles (high excess loss),
decreases weight for domains where proxy model has caught up.
"""
weights = np.array([domain_weights[d] for d in domains])
proxy_l = np.array([proxy_losses[d] for d in domains])
ref_l = np.array([ref_losses[d] for d in domains])
# Excess loss: positive means proxy is worse than reference
excess_loss = proxy_l - ref_l
# Exponentiated gradient update
updated_weights = weights * np.exp(eta * excess_loss)
# Project back to simplex (normalize)
updated_weights = updated_weights / updated_weights.sum()
return dict(zip(domains.keys(), updated_weights))
# Simulate a reference model with uniform weights trained for some steps
# Simulate proxy model losses: code and reasoning are harder domains
loss_rng = np.random.RandomState(7)
ref_losses = {
"web_common_crawl": 2.1 + loss_rng.normal(0, 0.05),
"web_filtered": 2.0 + loss_rng.normal(0, 0.05),
"books": 2.4 + loss_rng.normal(0, 0.05),
"wikipedia": 2.2 + loss_rng.normal(0, 0.05),
"code_github": 2.8 + loss_rng.normal(0, 0.05),
"scientific_papers": 2.6 + loss_rng.normal(0, 0.05),
"reddit_filtered": 1.9 + loss_rng.normal(0, 0.05),
}
# Proxy model with uniform weights is better on easy domains, worse on hard ones
proxy_losses = {
"web_common_crawl": ref_losses["web_common_crawl"] - 0.20,
"web_filtered": ref_losses["web_filtered"] - 0.15,
"books": ref_losses["books"] + 0.10, # Struggling
"wikipedia": ref_losses["wikipedia"] - 0.05,
"code_github": ref_losses["code_github"] + 0.25, # Struggling most
"scientific_papers": ref_losses["scientific_papers"] + 0.12, # Struggling
"reddit_filtered": ref_losses["reddit_filtered"] - 0.18,
}
# Initial uniform weights
initial_weights = {d: 1.0 / len(domains) for d in domains}
# Run several DoReMi update steps
weights = initial_weights.copy()
weight_history = [weights.copy()]
for step in range(20):
weights = doremi_weight_update(weights, proxy_losses, ref_losses, eta=0.15)
weight_history.append(weights.copy())
final_weights = weight_history[-1]DoReMi weight evolution (uniform start -> adaptive end): Domain Initial Final Change ---------------------------------------------------------- web_common_crawl 0.143 0.073 -0.070 web_filtered 0.143 0.085 -0.058 books 0.143 0.179 +0.036 wikipedia 0.143 0.114 -0.029 code_github 0.143 0.281 +0.138 scientific_papers 0.143 0.190 +0.047 reddit_filtered 0.143 0.077 -0.065
The DoReMi update naturally increases weights for code and scientific papers (where the proxy model struggles relative to the reference), while reducing weights for Reddit and common web crawl (where the proxy model performs well). This mirrors what human practitioners would do manually when they look at per-domain loss curves: boost the hard domains, suppress the easy ones. The difference is that DoReMi does this automatically and continuously throughout training, rather than requiring a human to review curves and adjust manually.

Key Parameters
The key parameters for a data mixing configuration are:
- domain_weights: The probability vector specifying how frequently each domain is sampled. Must sum to 1. This is the primary lever you control.
- training_budget: The total number of tokens to train on. Combined with domain weights, this determines the effective oversampling ratio for each domain.
- quality_threshold: If applying quality filtering before mixing, the minimum quality score for a document to be included. Raising this threshold reduces corpus size but increases average quality.
- eta: The learning rate for DoReMi-style adaptive mixing. Higher values cause more aggressive weight shifts; lower values produce smoother, more conservative updates.
- minimum_floor: A lower bound on any domain's weight, making sure no domain is completely excluded from training. Typical values are in the range of 1 to 5%.
Limitations and Practical Considerations
Data mixing can strongly affect a model, but it comes with practical limitations that matter for real training runs. Understanding where the approach is reliable and where it can mislead you is as important as knowing how to implement it.
The proxy model assumption is the most significant limitation. The claim that mixing rankings transfer across scales holds in aggregate across many published experiments, but individual exceptions exist. A mixture that looks good at 1 billion parameters may behave differently at 70 billion if certain domains interact with scale in non-obvious ways. Some capabilities, particularly the emergent ones that appear suddenly at larger scales, simply cannot be observed at proxy scale at all. The safest practice is to validate proxy results with at least one mid-scale run, say 10 to 20 billion parameters, before committing to a full training budget. This intermediate validation step costs more than proxy-only experiments but substantially reduces the risk of arriving at full scale with a suboptimal recipe.
Benchmark overfitting is another real risk. When you tune your data mixture to maximize scores on a specific set of benchmarks, you are implicitly optimizing for the capabilities those benchmarks measure. If your benchmark suite has blind spots, and all suites do, you may be trading unknown capabilities for better scores on the tests you are running. This is especially concerning if the benchmarks are publicly available and your training data includes documents that discuss or reference the benchmark tasks. A model that has seen discussions of BIG-Bench tasks during pretraining may perform better on those tasks not because of general capability but because of specific exposure. Using a diverse, held-out evaluation suite that is not publicly associated with your project helps reduce but cannot eliminate this risk.
Data mixing decisions made early in a project are expensive to revise. Changing your mixing recipe after training has started is disruptive: the model has developed internal representations tuned to the original distribution, and abruptly changing that distribution can cause instability or unlearning. This creates strong pressure to get the recipe right before large-scale training begins, which in turn justifies the investment in proxy experiments and ablation studies. The cost of a few extra proxy runs early in a project is tiny compared to the cost of a failed full-scale run caused by a preventable mixing error.
The interaction between data mixing and other training decisions is underexplored. Learning rate schedules, batch size, gradient clipping thresholds, and weight decay all interact with the data distribution in ways that are not well understood. Optimizer hyperparameters tuned for one mixing recipe may be suboptimal for another. Running mixing ablations under the exact hyperparameter configuration you plan to use for full training is important but expensive. In practice, most teams run mixing ablations under a standard configuration and then tune optimizer hyperparameters separately, accepting that the optimal mixing recipe and the optimal optimizer hyperparameters may not be exactly co-optimized.
The relationship between pretraining data mixing and subsequent fine-tuning is another important practical consideration. Data mixing during pretraining creates a foundation that all subsequent stages build on. A model with weak code understanding from pretraining will struggle to improve at coding through instruction tuning alone, because instruction tuning cannot easily create representations that were never built during pretraining. The same applies to mathematical reasoning, scientific factual grounding, and multilingual ability. Each of these capabilities needs to be developed during pretraining, before instruction tuning can refine and direct them. Practitioners sometimes underestimate this dependency and are surprised when instruction tuning cannot fix capability gaps left by poor pretraining data mixing.
Finally, the practical difficulty of data mixing is compounded by data availability uncertainty. You do not always know in advance exactly how much data you will have in each domain, or what quality distribution that data will have after filtering. Your recipe may be designed for 500 billion tokens of filtered web text, but if the filtering pipeline removes more content than expected, you may end up with 300 billion. Having contingency recipes for different corpus sizes, and monitoring corpus composition as the data pipeline runs, is important for maintaining control over the final training distribution.
Summary
Data mixing turns a raw collection of training corpora into a structured learning curriculum. The decisions made here propagate forward through the entire model development process, shaping every capability the final model will have. Getting these decisions right requires understanding both the principled frameworks developed by the research community and the practical constraints of real data pipelines.
The key ideas from this chapter are:
-
Domain proportions determine how often the model encounters each type of content. Proportional sampling by corpus size is almost never the right answer because corpus sizes do not reflect the relative value of different domains. Carefully curated recipes that deliberately over-represent high-quality, capability-relevant sources consistently outperform naive proportional baselines.
-
Quality weighting allows you to boost high-quality sources within a domain, either through trained classifiers, source-based heuristics, or perplexity scores from small reference models. Combining multiple quality signals with a continuous weighting scheme is more robust than applying hard thresholds.
-
Oversampling can compensate for data scarcity in high-value domains, but has diminishing returns beyond roughly 2 to 4 times repetition. Beyond that threshold, memorization begins to outweigh generalization, and gradient signal from repeated examples degrades toward zero.
-
Proxy model ablations make it feasible to explore the mixing space without committing full compute. The key assumption is that relative performance rankings transfer across model scales, which holds in aggregate but has individual exceptions. Intermediate-scale validation reduces the risk of scale-transfer failures.
-
DoReMi formalizes the mixing optimization as minimax distributionally robust optimization, automatically finding weights that minimize worst-case performance across domains through an exponentiated gradient update. It shifts mixing from a manual art toward a principled algorithmic process.
-
Dynamic mixing and curriculum effects can further improve over fixed recipes by adapting domain weights to the model's evolving performance profile, at the cost of additional infrastructure and the risk of training instability during distribution transitions.
-
Cross-domain spillover means that some domains contribute capabilities beyond their primary domain. Code data improves general reasoning. Scientific writing improves factual grounding. Understanding these interactions helps you build more efficient recipes that use spillover effects rather than treating each domain independently.
The next chapter covers synthetic data generation, which adds a new dimension to the mixing problem: rather than just mixing existing sources, you can generate new training examples to fill capability gaps that no existing corpus adequately covers.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about data mixing for language model pretraining.
Data Mixing Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!