Part of Language AI Handbook
Explains how differential privacy protects training data in language models, from the mathematical guarantee to DP-SGD, privacy budgets.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Differential Privacy
In the previous chapter on Memorization and Privacy, we saw how language models can memorize and inadvertently reproduce sensitive training data, including personally identifiable information. The threat is real: with enough probing, an attacker can extract credit card numbers, email addresses, and medical records that appeared verbatim in training corpora. Knowing the problem exists, the natural next question is: can we train a model that provably cannot leak too much about any individual in its training data, regardless of how cleverly an adversary queries it?
Differential privacy (DP) offers exactly that guarantee. Originating from theoretical cryptography in the mid-2000s and formalized by Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith in 2006, DP provides a rigorous, mathematically grounded definition of privacy: a training algorithm is differentially private if its output distributions are nearly indistinguishable whether or not any single individual's data was included. The "nearly" is quantified precisely by a privacy budget, and that budget can be audited, tracked, and reported. Unlike heuristic defenses such as data redaction or output filtering, DP's guarantee holds against adversaries with arbitrary auxiliary information and computational power.
Differential privacy is valuable because of the guarantee it provides and the assumptions it avoids. DP does not require the training data to be clean, de-identified, or derived from a particular distribution. It does not assume the adversary has limited computational resources or limited knowledge. It does not rely on the output looking a certain way. The privacy guarantee is a property of the algorithm itself, and it holds unconditionally once the algorithm has run. This adversarial robustness distinguishes DP from most practical privacy defenses in machine learning, which provide safety only under specific, often unstated assumptions.
This chapter covers how differential privacy works from first principles, how it gets applied to machine learning through DP-SGD (differentially private stochastic gradient descent), what the privacy-utility tradeoff looks like in practice, how to implement DP-SGD from scratch, and how modern large language models are trained or fine-tuned under differential privacy.
The Core Idea
Before diving into formalism, consider a concrete privacy problem. Suppose a hospital trains a text classifier to predict diagnosis codes from clinical notes. A curious adversary knows everything about the training procedure, has access to the trained model, and wants to determine whether a specific patient's record appeared in the training set. If they succeed, they learn something private about that individual. This is called a membership inference attack, and it is a real threat: models trained without any privacy protection often make it possible to infer with high confidence whether a given record was used for training.
The classic defense would be to remove identifying information from the training data. But as the memorization chapter demonstrated, neural networks can reconstruct quasi-identifying combinations even without obvious personally identifiable information. A clinical note can be stripped of names and dates while still being identifiable by its combination of age, diagnosis codes, and medication history. Anonymization is brittle.
Differential privacy takes a fundamentally different approach. Instead of attempting to scrub the data, it adds carefully calibrated random noise to the training process itself. The noise is designed so that no single training example has more than a bounded influence on the final model parameters. The key insight is that this is a property of the algorithm, not of the data. Even if you have perfect knowledge of the training algorithm, hyperparameters, architecture, and everything except the specific training set used, you cannot determine with high confidence whether any particular record was included.
A randomized algorithm is differentially private if running it on a dataset with one person's record produces nearly the same distribution of outputs as running it on the dataset with that person's record removed. "Nearly the same" is controlled by a parameter : smaller means stronger privacy.
To build intuition for what this means in practice, think of it this way. Imagine you flip a coin to decide whether Alice's record is included in the training set, then train a model and show it to an adversary. The adversary knows the coin was fair, but they also see the trained model. How much does the model reveal about which outcome the coin landed on? Differential privacy bounds the adversary's ability to answer that question: their posterior belief about the coin flip can shift by at most a factor of from the prior, regardless of how cleverly they analyze the model.
This framing reveals why the guarantee is so strong. An adversary with unlimited compute, access to any side information, and full knowledge of the training algorithm still cannot exploit the model to learn much about any individual. The protection is worst-case, not average-case.
Formal Definition
Let be a randomized mechanism (such as a training algorithm) that takes a dataset as input and produces some output (such as model parameters ). Two datasets and are called neighboring if they differ in exactly one record: one dataset is obtained from the other by adding or removing a single individual's data.
The mechanism satisfies -differential privacy if for all pairs of neighboring datasets and all subsets of outputs :
where:
- is the privacy loss parameter, sometimes called the privacy budget. Smaller values give stronger privacy.
- is the failure probability. With probability at most , the guarantee may fail entirely.
- is the maximum multiplicative ratio between output probabilities for neighboring datasets.
Let's unpack what this inequality is saying. The ratio measures how differently the mechanism behaves when one record is included versus excluded. The guarantee says this ratio can never exceed (with probability at least ). If and , the distributions are identical: the mechanism's output looks the same regardless of whether any particular record is present, which is perfect privacy but typically useless utility. As increases, the mechanism is allowed to reveal more about its input, enabling better utility.
When , the definition is called pure or -DP. When , it is called approximate or -DP. The term allows for a small probability that the guarantee fails altogether, which is conceptually unsatisfying but mathematically useful: it permits tighter bounds and is more practical for machine learning applications. Typically is set to a value much smaller than , where is the training set size, making sure the failure event is rare enough not to matter in practice. For a dataset of 100,000 training examples, is a common choice.
Interpreting Epsilon
The parameter controls how much the output distributions can differ between neighboring datasets. When , the outputs are identical regardless of whether a record is included, giving perfect privacy but typically zero utility. As grows, the mechanism can "reveal" more about the training data and achieve higher utility.
A useful way to think about : it bounds the log-odds ratio. If you observe the output of the mechanism, your posterior belief about whether a specific record was included can shift by at most a multiplicative factor of from your prior. For , this factor is approximately 2.72. If a record had a 50% prior probability of being in the dataset, the posterior after seeing the model output is at most . For , the maximum posterior would be 67%. The smaller is, the harder it is for any analysis of the model output to update your beliefs about membership.
This log-odds interpretation also makes composability natural. When you run two DP mechanisms sequentially, the log-odds can shift by up to from the first and up to from the second, for a total of . The privacy budgets simply add.
In practice, values used in deployed systems vary widely:
- Academic DP-SGD papers often report for reasonable model utility.
- Apple's on-device differential privacy for emoji statistics uses .
- The US Census Bureau's deployment for the 2020 Census used for person-level records.
- For LLM fine-tuning, is a common practical range.
There is no universally agreed "safe" threshold. The appropriate depends on the sensitivity of the data, the size of the population, and the threat model. A value that is privacy-preserving for a database of 10 million people may be inadequate for a database of 1,000 people, because the adversary's prior is much more concentrated in the small dataset case.
is sometimes criticized as an opaque parameter: what observable privacy risk does represent in terms of observable privacy risk? Researchers have developed more interpretable translations. For example, if , it means that even with perfect knowledge of the algorithm, an adversary cannot increase their probability of correctly guessing any binary fact about a specific individual by more than a factor of 3. This "hypothesis testing" interpretation gives practitioners a more direct sense of the protection level.
The Composition Theorem
One of differential privacy's most powerful properties is composability. If you run two differentially private mechanisms on the same dataset, the combined result is still differentially private, with privacy loss that accumulates. This composability is what allows DP to be applied to complex multi-step processes like neural network training.
Basic composition: If mechanism satisfies -DP and satisfies -DP, then running both on the same dataset satisfies -DP. Privacy budgets add linearly.
Advanced composition: Tighter bounds are achievable when running many mechanisms. The privacy loss does not grow as fast as the naive linear sum suggests. Moments accountant and Renyi differential privacy (RDP) techniques provide these tighter bounds and are essential for training neural networks over many gradient steps.
To understand why advanced composition is tighter, think about it probabilistically. Each individual mechanism step contributes a small privacy loss, but these losses are random variables, not fixed worst-case values. When you compose many steps, the losses are unlikely to all land on their worst-case simultaneously. The moments accountant exploits this by tracking the distribution of privacy loss across compositions, rather than just the maximum. Under RDP composition, after running independent -DP mechanisms, the total privacy cost grows approximately as rather than in favorable regimes, giving much better bounds for the hundreds of thousands of gradient steps in neural network training.
This composability is what makes DP practical for machine learning: you can track exactly how much privacy budget each training step consumes, and stop training once the budget is exhausted.
The composition theorem also explains why releasing multiple queries about the same sensitive dataset degrades privacy cumulatively. Each SQL query, each model checkpoint, each evaluation metric computed on the private data is a mechanism that spends budget. Organizations deploying DP systems must track all these interactions carefully, including intermediate releases as well as the final model. A privacy accountant is the bookkeeping system that logs each expenditure and ensures the total remains within the authorized budget. Failing to count all interactions is one of the most common mistakes in real DP deployments.
Sensitivity and the Building Blocks of DP
To understand how to make a concrete algorithm differentially private, you need the concept of sensitivity. The sensitivity of a function measures how much its output can change when one record in the input dataset is modified. If a function has low sensitivity, only a small amount of noise is needed to hide the influence of any single record. If a function has high sensitivity, more noise is required.
There are two common notions of sensitivity. The sensitivity measures the maximum change in the output's norm, and the sensitivity measures the maximum change in the norm:
The Laplace mechanism (which adds Laplace-distributed noise scaled to ) achieves pure -DP and is preferred for scalar or low-dimensional outputs. The Gaussian mechanism (which adds Gaussian noise scaled to ) achieves approximate -DP and is preferred for high-dimensional vector outputs like neural network gradients, because Gaussian noise has better properties in high dimensions. We will focus on the Gaussian mechanism for the rest of this chapter.
The Gaussian Mechanism
To make any function differentially private, you add noise proportional to how sensitive the function's output is to a single data point. The Gaussian mechanism is the standard approach for vector-valued functions such as gradients.
The sensitivity of a function measures the maximum change in its output norm when one record is added or removed:
where:
- is the sensitivity of function , measuring the worst-case change in output when one record changes.
- restricts the maximum to neighboring datasets that differ by exactly one record.
- is the Euclidean distance between the outputs on the two neighboring datasets.
The Gaussian mechanism adds isotropic Gaussian noise scaled to the sensitivity:
where:
- is multivariate Gaussian noise with standard deviation in each dimension.
- controls the amount of noise added, which in turn controls the privacy-utility tradeoff.
- Larger means more noise, which means stronger privacy but lower utility.
For approximate DP with parameters , the Gaussian mechanism satisfies DP when:
where:
- is the noise standard deviation to be added.
- is a constant derived from the tail bounds of the Gaussian distribution, growing slowly as decreases.
- is the sensitivity of the function being computed.
- is the target privacy parameter; smaller requires larger .
This formula captures the three-way tradeoff precisely. To achieve stronger privacy (smaller ), you need more noise (larger ). To guarantee the privacy bound holds more reliably (smaller ), you again need more noise, because the Gaussian's tail probability shrinks only logarithmically with . The sensitivity anchors the scale of the noise to the specific function being computed: a function that barely changes when one record is removed needs very little noise to hide that change.
The Gaussian mechanism has an important property that makes it particularly well-suited to machine learning: in high dimensions, Gaussian noise distributes evenly across all directions. Adding noise to a gradient vector perturbs all parameter directions equally, which tends to have a smaller effect on the overall model quality than directional noise would. This isotropy is part of why the Gaussian mechanism is preferred over alternatives for neural network training.
DP-SGD: Training Neural Networks with Privacy
Applying differential privacy to training a neural network is non-trivial because the standard training algorithm, stochastic gradient descent, has unbounded sensitivity. A single outlier training example with a very large loss can produce a very large gradient, pushing model parameters arbitrarily far from where they would have been without that example. If we tried to apply the Gaussian mechanism directly to the gradient of a minibatch without any modification, we would need to add infinite noise to achieve meaningful privacy. The key challenge is bounding the sensitivity.
DP-SGD, introduced by Abadi et al. in 2016, solves this through two modifications to the standard SGD update. Together, these two changes transform the training algorithm from one with potentially unbounded sensitivity into one with precisely bounded sensitivity.
The Two Key Operations
Step 1: Per-example gradient clipping. Before aggregating gradients from a minibatch, DP-SGD clips each individual example's gradient vector to have norm at most (the clipping threshold):
where:
- is the gradient of the loss for training example .
- is the clipping threshold, a hyperparameter chosen before training.
- is the clipped gradient, which satisfies by construction.
Clipping bounds the sensitivity of the gradient sum: no single example can shift the aggregate gradient by more than in norm. Without clipping, an example with a large loss would produce a large gradient, making the sensitivity potentially infinite and requiring infinite noise for privacy. After clipping, the sensitivity is exactly , regardless of the data distribution. This is the mechanism that makes DP-SGD work.
To see why clipping achieves sensitivity , consider two neighboring datasets and . The gradient sum on differs from the gradient sum on by exactly the clipped gradient of , which has norm at most . Thus , giving .
Notice that gradient clipping is not the same as gradient norm regularization or weight clipping. It operates per-example, before aggregation, and it scales the entire gradient vector uniformly. A gradient with norm 10 and clipping threshold 1 is scaled by a factor of 0.1; its direction is preserved but its magnitude is reduced to exactly 1.
Step 2: Gaussian noise addition. After clipping and summing gradients, DP-SGD adds Gaussian noise calibrated to the clipping threshold:
where:
- is the minibatch size.
- is the noise multiplier, the ratio of noise standard deviation to clipping threshold.
- The noise variance is proportional to the squared clipping threshold, since that is the sensitivity.
The noise multiplier is the key privacy-utility knob. Setting larger adds more noise, giving stronger privacy guarantees. Setting it smaller adds less noise, giving better gradient quality but weaker privacy. Given a target privacy budget and a training procedure (number of steps, batch size, dataset size), the minimum that achieves the budget can be computed using a privacy accountant.
The full DP-SGD update then proceeds as in standard SGD: parameters are updated using this noisy, clipped gradient. The privacy guarantee applies to each step independently, and the total privacy cost after steps is computed via composition.
The effect of gradient clipping is intuitive to visualize: gradients with large norms are scaled down to lie exactly on the boundary of a ball of radius , while gradients with small norms are left unchanged. The plot below illustrates this with a distribution of gradient norms before and after clipping.


The left panel shows the distribution of per-example gradient norms before clipping. The distribution is right-skewed: most gradients have moderate norms, but a long tail of high-norm gradients exists. These high-norm gradients correspond to training examples with large losses, which are often the most informative but also the most sensitive. Without clipping, these outliers would dominate the aggregate gradient and make the sensitivity unbounded.
The right panel shows what clipping does to each gradient. Every point on or below the diagonal (blue) has a norm below the threshold and is left unchanged. Every point above the threshold (red) is pulled back to exactly . After clipping, no single example can contribute a gradient with norm larger than , bounding the sensitivity.
Privacy Accounting
A single DP-SGD step with Poisson subsampling (where each example is independently included in the minibatch with probability ) achieves -DP for that step. The privacy cost over steps accumulates, but thanks to the moments accountant or equivalently Renyi DP accountant, the total privacy cost is tighter than naive composition.
The key insight is that subsampling amplifies privacy: if only a fraction of the training data participates in each step, an individual example is included in any given step with probability . This "privacy amplification by subsampling" means the effective privacy loss per step is roughly rather than . Subsampling acts like an extra layer of randomness on top of the noise, making it even harder for an adversary to determine whether a specific record was used.
To see why this amplification occurs, consider the mechanism from the perspective of a single individual's record. If the record is included in the training set, it enters any given minibatch with probability . From the adversary's perspective, most steps look identical whether the record is in the dataset or not (because it simply was not sampled). Only the fraction of steps where it was sampled carry any information. This dramatically reduces the per-step privacy cost.
The total privacy cost after steps, at failure probability , can be approximated as:
where:
- is the cumulative privacy cost after all training steps.
- is the total number of gradient steps, equal to epochs times batches per epoch.
- is the subsampling rate, the fraction of training data in each mini-batch.
- is the noise multiplier; larger means smaller .
- is the failure probability in the -DP guarantee.
This is a simplified approximation. In practice, tight accounting uses the moments accountant or the PRV accountant (Privacy Random Variable accountant), which numerically compose the per-step privacy guarantees without closed-form approximations.
The formula's implications are direct and practical:
- More training steps increase privacy cost: training longer spends more budget.
- Larger subsampling rate (bigger batches relative to dataset size) increases privacy cost per step.
- Less noise increases privacy cost proportionally.
- The factor reflects the square-root growth from advanced composition rather than the linear growth that naive composition would give.
This composability is what makes DP practical for machine learning: you can track exactly how much privacy budget each training step consumes, and stop training once the budget is exhausted. A privacy accountant runs alongside training, incrementing the total after each step and triggering early stopping if the budget is reached.
The Privacy-Utility Tradeoff
Every DP mechanism faces a fundamental tradeoff: stronger privacy requires more noise, and more noise degrades model quality. Understanding this tradeoff quantitatively is essential for practitioners who need to reason about what level of privacy protection is achievable for their use case.
Why DP Hurts Small Models More
The noise added by DP-SGD must be scaled to the model's gradient sensitivity. For a model trained with clipping threshold and noise multiplier , the signal-to-noise ratio of a gradient update scales roughly as:
where:
- is the average true gradient magnitude, proportional to the informativeness of the training examples.
- is the noise multiplier (larger means stronger privacy and worse SNR).
- is the clipping threshold (scales both the signal and the noise).
- is the minibatch size; the factor arises because the signal sums across the batch while the noise stays fixed, giving a improvement.
This scaling has an important practical implication: larger batch sizes improve the signal-to-noise ratio. This is the opposite of the usual SGD intuition, where smaller batches provide noisier but less biased gradient estimates and often converge faster. For DP training, large batches are strongly preferred because they dilute the fixed noise over more signal. In practice, DP training uses batch sizes of thousands to tens of thousands, while non-DP training might use hundreds.
For large models with many parameters, the noise is spread across many dimensions, but the gradient signal in each dimension is also diluted. The key empirical finding from large-scale DP research is that the benefit of scale is enormous: very large models can tolerate DP noise much better than small models because they have more capacity to fit the signal despite the noise. A GPT-3-scale model with DP might lose only 1-2 percentage points of accuracy on a task where a BERT-base model with DP loses 5-10 percentage points, even though the large model has far more parameters generating more noise per step.
Why Scale Helps
The underlying reason large models handle DP better deserves a closer look. There are two separate mechanisms at work.
First, large models trained on large datasets have larger effective batch sizes relative to the total dataset. If you train with batch size 4,096 on a dataset of 1 million examples, the subsampling rate is small, and privacy amplification by subsampling gives you significant protection per step. With a dataset of 1,000 examples and batch size 64, is sixteen times larger, spending sixteen times more budget per step.
Second, large pre-trained models already encode most of the task-relevant information before any DP fine-tuning begins. The DP fine-tuning step only needs to adapt the model to a small distribution shift, which requires much smaller gradient updates. Smaller updates mean the gradients are less likely to be aggressively clipped, reducing both the clipping bias and the effective noise level.
This is the key insight behind DP fine-tuning: the most practical way to use DP for LLMs is not to train from scratch with DP, but to pre-train on public data without DP and then fine-tune on private data with DP. The pre-training provides a strong starting point; the DP fine-tuning makes only the final adjustments needed for the private task.
Empirical Accuracy Gaps
The accuracy gap between DP-trained and non-DP-trained models depends on several factors:
- Dataset size: Larger datasets reduce the per-step sampling rate needed for a given batch size, reducing the effective noise multiplier needed to achieve a given .
- Model size: Larger models trained with DP on large datasets often nearly match non-DP accuracy.
- Task difficulty: Simple tasks like binary sentiment classification show smaller DP gaps than complex generation tasks.
- Pre-training quality: Models with better pre-training representations need smaller fine-tuning adjustments, which interact more favorably with DP noise.
A typical pattern for text classification: a BERT-base model trained with on a standard sentiment dataset achieves around 93% accuracy compared to 95% without DP. At , accuracy may drop to 88-90%. The gap varies significantly by dataset and architecture. For generation tasks, perplexity increases substantially with strong privacy, making coherent long-form text generation difficult at small .
For LLMs, the outcome depends more strongly on the training setup. Fine-tuning a pre-trained GPT-2 with DP on downstream tasks shows smaller degradation than training from scratch, because the pre-training has already learned rich representations. The DP training only needs to adjust a relatively small number of task-specific parameters, spending the privacy budget on a smaller, more targeted update.
Clipping and Bias
An important but often overlooked aspect of DP-SGD is that gradient clipping introduces bias into the gradient estimate, not just variance. When a gradient vector is clipped to have a smaller norm, the update direction changes. Gradients from high-loss examples (which are often the most informative, because they signal areas where the model is most wrong) get clipped more aggressively than gradients from easy examples.
This clipping bias can cause several issues:
- Convergence to a different, potentially worse stationary point than unconstrained SGD would find.
- Slower convergence on minority subgroups whose examples tend to have larger gradients, because their contribution to the aggregate update is systematically suppressed.
- Underrepresentation of rare but important patterns, since the model learns less from the examples that deviate most from what it already knows.
Choosing the right clipping threshold is critical and non-trivial. If is too small, most gradients are clipped, introducing large bias. If is too large, more noise must be added to maintain the same privacy guarantee (since the sensitivity is larger), and the effective noise-to-signal ratio worsens. The optimal sits at roughly the median or 75th percentile of gradient norms, where most gradients are preserved but extreme outliers are controlled.
Adaptive clipping methods attempt to set automatically based on gradient statistics, including DP-compatible methods that estimate the quantile of gradient norms with privacy. The idea is to query the training data for the empirical gradient norm distribution, add noise to that query to maintain DP, and then use the result to set dynamically. This avoids the need to manually tune as a hyperparameter, at the cost of spending a small portion of the privacy budget on the quantile estimation.
DP for LLMs
Training large language models with differential privacy presents unique challenges and opportunities. LLMs have billions of parameters, are typically trained on massive public corpora where privacy concerns may seem lower, but the fine-tuning stage often involves sensitive private data: medical records, financial documents, personal communications, legal files. The combination of a powerful pre-trained model and a private fine-tuning dataset is precisely where DP provides the most value.
Fine-Tuning with DP
The most practical application of DP to LLMs is differentially private fine-tuning: start from a publicly pre-trained model and fine-tune on a private dataset with DP guarantees. Because the pre-trained weights already encode rich language understanding, the fine-tuning process only needs to make small adjustments, and the DP noise budget is spent more efficiently.
Research on DP fine-tuning of LLMs has shown that large pre-trained models can be fine-tuned with meaningful DP guarantees () with only modest accuracy degradation on many NLP benchmarks. Li et al. (2022) demonstrated that GPT-2-XL fine-tuned with DP on the E2E dataset produced text comparable in quality to non-DP fine-tuned smaller models. The key finding is consistent across studies: model scale is the most important factor for DP utility. A large pre-trained model fine-tuned with DP often performs comparably to a small model without DP on the same task.
This has an important practical implication: if you have a fixed privacy budget and need to choose between a small model trained without DP or a large model trained with DP, the large model with DP is often the better choice from both privacy and utility perspectives.
Parameter-Efficient DP Fine-Tuning
A particularly promising approach combines DP with parameter-efficient fine-tuning methods like LoRA (Low-Rank Adaptation), which we covered in the Parameter-Efficient Fine-Tuning chapter. Since LoRA only trains a small number of parameters (the low-rank adapter matrices), the per-step gradient is defined over a much smaller parameter space. This has two important benefits for DP.
First, the noise dimensionality is reduced. Instead of adding noise to a billion-dimensional gradient vector, you add noise to a much smaller vector of adapter parameters. The signal-to-noise ratio improves dramatically, because the same amount of gradient signal is now competing with noise in a much lower-dimensional space.
Second, the effective clipping is applied to a smaller gradient, which has less variability across examples, reducing clipping bias. The adapter gradients are more uniform across training examples than full-model gradients, because they capture only high-level task adaptation rather than fine-grained language patterns.
The combination of DP and LoRA, sometimes called DP-LoRA, has emerged as the practical default for DP LLM fine-tuning. Research has shown that DP-LoRA can match or exceed the accuracy of full DP fine-tuning while consuming less effective privacy budget per step, because the reduced dimensionality and lower gradient variability combine to provide more useful updates per unit of noise.
Choosing the rank hyperparameter for DP-LoRA involves a tradeoff. Lower rank means fewer parameters to update (better for DP, since less noise is spread across the gradient) but also less expressive adaptation (worse for utility). Empirically, ranks in the range of 4-16 work well for DP fine-tuning, which is lower than the typical 32-64 used without DP. The intuition is that DP noise already regularizes the fine-tuning heavily, so lower-rank adapters are less likely to overfit.
Canaries and Empirical Privacy Auditing
Theoretical DP guarantees are worst-case bounds. In practice, the actual privacy leakage may be lower than suggests, because the worst-case scenario may not arise in realistic datasets. Privacy auditing provides empirical lower bounds on privacy leakage by attempting membership inference attacks and measuring their success rate. If the attack succeeds better than the DP guarantee allows, there is a bug. If the attack does no better than the theoretical bound predicts, you have evidence that the implementation is correct.
A common auditing technique uses canary examples: training examples that are deliberately inserted into the dataset and then tested to see if they can be detected. Canaries are often crafted to be highly distinctive (unusual sequences, rare patterns, synthetic data points) so they are maximally easy to memorize. If the attacker can detect canaries with high confidence, the privacy guarantee is tight or may be violated. If not, the mechanism provides at least as much privacy as the theoretical suggests.
Auditing with canaries works as a shadow training exercise. The auditor selects a set of canary examples, some of which are added to the training data and some are held out. After training, the auditor runs a membership inference test: which examples can be distinguished as members versus non-members based on the model's behavior? A common test uses the model's loss on the example as a signal. The auditor measures the advantage of the membership inference attack over random guessing. If the advantage is bounded by (the theoretical maximum advantage under -DP), the empirical evidence is consistent with the theoretical guarantee. If the advantage exceeds this bound, there is a violation.
Empirical auditing is essential in practice because DP implementations can harbor bugs that invalidate the theoretical guarantee while the code appears to run correctly. Common implementation failures include:
- Incorrect subsampling: Using fixed, deterministic batches instead of Poisson subsampling invalidates the subsampling amplification used in the privacy accounting.
- Missing noise in some components: If certain gradient components (e.g., embeddings for rare tokens) are updated without adding noise, those components can memorize sensitive data.
- Numerical precision issues: Floating-point rounding in the clipping computation can introduce small but non-zero sensitivities beyond .
- Gradient accumulation bugs: When using gradient accumulation over multiple micro-batches, the noise must be added once to the aggregated gradient, not per micro-batch.
Auditing catches these implementation failures in a way that code review alone cannot. It is good practice to run a privacy audit on any new DP implementation before deploying it with real sensitive data.
Tools like the autodp library and Opacus (from Meta) provide both theoretical accounting and empirical auditing capabilities. Opacus, in particular, is the most widely used library for DP training of PyTorch models and handles the per-example gradient computation efficiently.
Implementation: DP-SGD from Scratch
Let's implement DP-SGD to build intuition for how the algorithm works at the level of individual operations. We will train a simple logistic regression model on a synthetic dataset, which is small enough to run quickly but complex enough to show the privacy-utility tradeoff clearly.
Setup and Data Generation
We start by generating a binary classification dataset with 1,000 samples and 20 features, standardizing it, and splitting into train and test sets.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.preprocessing import StandardScaler
np.random.seed(42)
# Generate a binary classification dataset
n_samples = 1000
X, y = make_classification(
n_samples=n_samples,
n_features=20,
n_informative=10,
n_redundant=5,
flip_y=0.05,
random_state=42,
)
# Standardize features
scaler = StandardScaler()
X = scaler.fit_transform(X)
# Train/test split
split = int(0.8 * n_samples)
X_train, X_test = X[:split], X[split:]
y_train, y_test = y[:split], y[split:]Logistic Regression Gradient
For logistic regression, the gradient of the binary cross-entropy loss with respect to the weights, for a single example , is where . DP-SGD requires computing this per-example gradient before any aggregation or clipping.
def sigmoid(z):
"""Numerically stable sigmoid."""
return np.where(z >= 0, 1 / (1 + np.exp(-z)), np.exp(z) / (1 + np.exp(z)))
def compute_loss(X, y, w, b):
"""Binary cross-entropy loss."""
logits = X @ w + b
probs = sigmoid(logits)
probs = np.clip(probs, 1e-7, 1 - 1e-7)
loss = -np.mean(y * np.log(probs) + (1 - y) * np.log(1 - probs))
return loss
def compute_per_example_gradients(X, y, w, b):
"""
Compute per-example gradients for logistic regression.
Returns a list of (grad_w, grad_b) per example.
"""
logits = X @ w + b
probs = sigmoid(logits)
errors = probs - y # (n,)
# Per-example gradients for w: each is a (d,) vector
grad_w_per_example = errors[:, np.newaxis] * X # (n, d)
grad_b_per_example = errors # (n,)
return grad_w_per_example, grad_b_per_exampleNote that standard mini-batch SGD would compute np.mean(errors[:, np.newaxis] * X, axis=0) directly, averaging gradients across examples. DP-SGD cannot do this: it must inspect and clip each example's gradient individually before aggregation, because clipping is a per-example operation.
The DP-SGD Update
The core DP-SGD function combines the clipping and noise addition steps. Notice how it processes each example's gradient independently, computes the combined gradient norm, clips it, then sums the clipped gradients and adds noise before averaging.
def dp_sgd_step(
X_batch, y_batch, w, b, learning_rate, clip_threshold, noise_multiplier
):
"""
One step of DP-SGD.
Args:
X_batch, y_batch: minibatch
w, b: current model parameters
learning_rate: step size
clip_threshold (C): maximum L2 norm of per-example gradient
noise_multiplier (sigma): controls privacy-utility tradeoff
Returns:
Updated w and b, plus fraction of gradients clipped
"""
n_batch = len(y_batch)
# Step 1: Compute per-example gradients
grad_w, grad_b = compute_per_example_gradients(X_batch, y_batch, w, b)
# Step 2: Clip each per-example gradient
# Combine into single vector for norm computation
grad_combined = np.hstack([grad_w, grad_b[:, np.newaxis]]) # (n, d+1)
norms = np.linalg.norm(grad_combined, axis=1) # (n,)
clipping_factors = np.minimum(1.0, clip_threshold / norms)
n_clipped = np.sum(norms > clip_threshold)
grad_w_clipped = grad_w * clipping_factors[:, np.newaxis] # (n, d)
grad_b_clipped = grad_b * clipping_factors # (n,)
# Step 3: Sum clipped gradients
sum_grad_w = np.sum(grad_w_clipped, axis=0) # (d,)
sum_grad_b = np.sum(grad_b_clipped) # scalar
# Step 4: Add Gaussian noise scaled to sensitivity (clip_threshold)
noise_w = np.random.normal(
0, noise_multiplier * clip_threshold, size=w.shape
)
noise_b = np.random.normal(0, noise_multiplier * clip_threshold)
noisy_grad_w = (sum_grad_w + noise_w) / n_batch
noisy_grad_b = (sum_grad_b + noise_b) / n_batch
# Step 5: Parameter update
w_new = w - learning_rate * noisy_grad_w
b_new = b - learning_rate * noisy_grad_b
return w_new, b_new, n_clipped / n_batchThis implementation makes several specific design choices. The combined gradient norm for clipping is computed over the concatenated gradient vector (weights and bias together), treating the entire parameter update for one example as a single entity. The noise is added once to the summed gradient, not to each example's gradient individually: adding noise once to the sum is equivalent to adding noise to the average but requires less computational work. The noise standard deviation is noise_multiplier * clip_threshold, matching the Gaussian mechanism formula since the sensitivity is exactly clip_threshold.
Training Loop
from sklearn.metrics import accuracy_score
def train_model(
X_train,
y_train,
X_test,
y_test,
n_epochs=30,
batch_size=64,
learning_rate=0.1,
clip_threshold=1.0,
noise_multiplier=1.0,
use_dp=True,
):
"""Train logistic regression with or without DP-SGD."""
n_features = X_train.shape[1]
w = np.zeros(n_features)
b = 0.0
n_batches = len(X_train) // batch_size
train_losses = []
test_accuracies = []
for epoch in range(n_epochs):
# Shuffle training data
perm = np.random.permutation(len(X_train))
X_shuf = X_train[perm]
y_shuf = y_train[perm]
epoch_loss = 0.0
for i in range(n_batches):
start = i * batch_size
end = start + batch_size
X_b = X_shuf[start:end]
y_b = y_shuf[start:end]
if use_dp:
w, b, _ = dp_sgd_step(
X_b,
y_b,
w,
b,
learning_rate=learning_rate,
clip_threshold=clip_threshold,
noise_multiplier=noise_multiplier,
)
else:
# Standard SGD
grad_w, grad_b = compute_per_example_gradients(X_b, y_b, w, b)
w -= learning_rate * np.mean(grad_w, axis=0)
b -= learning_rate * np.mean(grad_b)
epoch_loss += compute_loss(X_b, y_b, w, b)
train_losses.append(epoch_loss / n_batches)
y_pred = (sigmoid(X_test @ w + b) >= 0.5).astype(int)
test_accuracies.append(accuracy_score(y_test, y_pred))
return w, b, train_losses, test_accuracies
# Train without DP (baseline)
w_nodp, b_nodp, losses_nodp, accs_nodp = train_model(
X_train,
y_train,
X_test,
y_test,
n_epochs=30,
batch_size=64,
learning_rate=0.1,
use_dp=False,
)
# Train with DP (low privacy, high utility)
w_dp1, b_dp1, losses_dp1, accs_dp1 = train_model(
X_train,
y_train,
X_test,
y_test,
n_epochs=30,
batch_size=64,
learning_rate=0.1,
clip_threshold=1.0,
noise_multiplier=0.5,
use_dp=True,
)
# Train with DP (high privacy, lower utility)
w_dp2, b_dp2, losses_dp2, accs_dp2 = train_model(
X_train,
y_train,
X_test,
y_test,
n_epochs=30,
batch_size=64,
learning_rate=0.1,
clip_threshold=1.0,
noise_multiplier=2.0,
use_dp=True,
)Final test accuracy: No DP (baseline): 0.815 DP (sigma=0.5, low priv): 0.815 DP (sigma=2.0, high priv): 0.815 Accuracy cost of privacy: Low privacy (sigma=0.5): 0.000 High privacy (sigma=2.0): 0.000
The output shows a clear tradeoff: higher noise multipliers (stronger privacy) produce lower accuracy. The baseline achieves the best accuracy, the low-privacy variant is close, and the high-privacy variant loses several percentage points. This is the fundamental privacy-utility tradeoff in action on a small synthetic example.
The learning curves below show how this tradeoff unfolds over training. All three models improve with more epochs, but the high-privacy variant converges more slowly and to a lower final accuracy, because the noisy gradient updates carry less useful information per step.

Privacy Budget Calculation
With the training runs complete, we can also estimate how much privacy budget each run consumed. The estimate function uses the simplified square-root approximation we derived earlier. In production, you would replace this with the moments accountant from autodp or Opacus for tighter bounds.
def estimate_epsilon(
n_train, batch_size, noise_multiplier, n_epochs, delta=1e-5
):
"""
Simplified privacy budget estimate using the subsampling amplification
formula. This is an approximation; use the autodp library for
tight bounds in production.
"""
n_steps = int(n_epochs * n_train / batch_size)
q = batch_size / n_train # Subsampling rate
# Approximate epsilon via the Gaussian mechanism + subsampling
# This uses the simplified sqrt(2T) approximation
epsilon_approx = (
np.sqrt(2 * n_steps)
* q
/ noise_multiplier
* np.sqrt(2 * np.log(1.25 / delta))
)
return epsilon_approx, n_steps, q
eps_low, n_steps, q = estimate_epsilon(
len(X_train), batch_size=64, noise_multiplier=0.5, n_epochs=30, delta=1e-5
)
eps_high, _, _ = estimate_epsilon(
len(X_train), batch_size=64, noise_multiplier=2.0, n_epochs=30, delta=1e-5
)Privacy budget estimates (delta=1e-5): Training steps: 375 Subsampling rate: 0.0800 sigma=0.5 (low privacy): epsilon ~ 21.23 sigma=2.0 (high privacy): epsilon ~ 5.31 Note: These are simplified approximations. Use autodp or opacus for tight accounting in production.
These estimates illustrate how epsilon scales: the low-noise variant accumulates a much larger privacy budget over 30 epochs than the high-noise variant. With a fixed privacy budget target (say ), you would stop training the low-noise variant much earlier, before it has had as many learning steps. Managing this tradeoff between training epochs and noise level is a central concern of practical DP deployment.
The plot below visualizes how epsilon grows as training progresses for different noise multipliers. Each colored curve represents a different value. The horizontal dashed line at represents a typical deployment target: training must stop when the curve crosses that line.

The plot makes the tradeoff concrete. With , the budget is exhausted after roughly 20 steps at , severely limiting training. With , you can run hundreds of steps within the same budget. The question for each deployment is: which combination of noise level and training length gives better final model quality?
Key Parameters and Their Interactions
Understanding how DP-SGD hyperparameters interact is essential for getting good results. The key parameters are:
- max_grad_norm (C): The per-example gradient clipping threshold. Smaller values mean more clipping (more bias) but less noise needed for a given epsilon. Typical range: 0.1 to 2.0. Values near the median gradient norm are generally best.
- noise_multiplier (sigma): Scales the Gaussian noise added to gradients. Higher values provide stronger privacy but hurt utility. Determined automatically by a privacy accountant when targeting a specific epsilon.
- target_epsilon: The privacy budget. Smaller values give stronger privacy. Practical range for LLMs: 1 to 8. Below 1, utility typically collapses for all but the largest models.
- target_delta: Failure probability. Should be much smaller than (dataset size). Typical: to .
- batch_size: Larger batches improve the signal-to-noise ratio and are strongly preferred for DP training. Batch sizes of 256 to 4,096 are common for DP LLM fine-tuning.
- epochs: More training epochs consume more privacy budget. Budget-aware early stopping is essential.
These parameters do not act independently. Increasing the batch size while holding everything else constant reduces the subsampling rate , which reduces the privacy cost per step, allowing more training steps within the same budget. This means you can train for more epochs or use a lower to get better accuracy. For a fixed target, the optimal strategy often involves maximizing batch size (to improve SNR) while setting just high enough to stay within budget over the desired number of epochs.
Limitations and Practical Challenges
Differential privacy is a powerful mathematical guarantee, but applying it to real machine learning systems comes with significant challenges that practitioners must address carefully. Understanding these limitations is as important as understanding the mechanism itself.
The Scale Requirement
DP-SGD's privacy-utility tradeoff is notoriously poor at small scales. With only a few thousand training examples, the noise needed to achieve meaningful privacy (say below 3) often overwhelms the gradient signal entirely, producing models barely better than random guessing. This is why early academic results on DP deep learning were quite discouraging: the benchmarks used datasets like MNIST (60,000 training images), which were too small to realize DP's potential. Papers from 2017-2019 showed 10-20 percentage point accuracy gaps between DP and non-DP models, leading many practitioners to conclude that DP was impractical.
The empirical picture improves dramatically with scale. A large model fine-tuned on millions of examples can achieve with accuracy nearly matching its non-DP counterpart. For practitioners with smaller datasets, the recommendation is either to use a large pre-trained model as a starting point (to minimize the amount of DP training needed) or to accept weaker privacy guarantees. There is no magic that makes DP work with 1,000 training examples and .
Hyperparameter Tuning and Privacy
A subtle but important issue: selecting hyperparameters (clipping threshold, noise multiplier, learning rate, batch size) for DP training requires evaluating model performance on the private validation set. Each such evaluation consumes privacy budget. If you run 50 hyperparameter tuning experiments and pick the best, you have effectively spent 50 times the privacy budget of a single run, even if each run used a small .
Properly accounting for hyperparameter tuning is often omitted in academic papers but is necessary for real deployments. Approaches include:
- Using a small privacy budget for coarse tuning before committing to a full training run.
- Using public data to set hyperparameters where possible, then transferring those settings to the private training run.
- Reporting the total privacy budget including all tuning experiments, including the final run and all preceding experiments.
- Using DP hyperparameter selection methods, which choose hyperparameters from a fixed set using DP mechanisms.
Without careful accounting, a reported deployment may provide much weaker privacy in practice due to the privacy cost of selecting those hyperparameters.
Fairness and Equity Concerns
DP noise is not neutral with respect to the distribution of the training data. Minority subgroups, whose examples are rarer and whose gradients may be larger (due to harder, more distinctive examples), are systematically more affected by both clipping and noise than majority subgroups. Research by Bagdasaryan et al. (2019) and others has shown that DP-trained models can have substantially worse accuracy on underrepresented groups than on the majority, even when overall accuracy is only modestly reduced.
The mechanism is straightforward. Examples from minority groups tend to be harder to classify (because there are fewer of them to learn from), generating larger gradients that are more aggressively clipped. The clipping bias systematically suppresses the model's learning from these examples. Meanwhile, the added Gaussian noise further obscures the already-weak signal from minority groups. The result is a model that is well-calibrated for the majority while performing poorly for the groups that are most different from the statistical average.
This creates a troubling equity concern: the individuals most at risk from privacy violations (those whose data is rare, sensitive, or distinctive) are also those who pay the highest accuracy cost under DP training. Privacy protection and model utility are inequitably distributed. Addressing this requires careful measurement of group-level performance metrics during and after DP training, and potentially targeted privacy budgets or gradient reweighting strategies to compensate.
DP Does Not Prevent All Privacy Attacks
Differential privacy provides a specific, precisely defined guarantee: bounded sensitivity of the training algorithm's output to any single training example. It does not guarantee that:
- Aggregate statistics about subgroups are private. If many people in a group share a pattern, that pattern can still be reflected in the model even under DP.
- Information memorized from public pre-training is protected. If a model was pre-trained on web text that included sensitive information, that information is present in the weights before any DP fine-tuning begins.
- Inference attacks using non-private auxiliary information are prevented. If the adversary knows a lot about an individual from other sources, the small additional information from the model might still be enough to breach privacy.
- The model's predictions on new inputs are private. DP only protects the training process. If the model's predictions reveal information about training data (which can happen in certain settings), DP does not prevent that.
Practitioners sometimes assume that DP training eliminates all privacy risks. It narrows the attack surface substantially, but it should be considered one component of a broader privacy-preserving strategy, alongside data minimization, access controls, output filtering, and regular auditing. DP is necessary but not sufficient for data privacy across the full system.
Computational Overhead
DP-SGD requires computing per-example gradients before aggregating them, which is significantly more computationally expensive than standard mini-batch SGD. In naive implementations, you compute the gradient for each example in the batch individually, which is times the cost of a single forward-backward pass. This makes DP training approximately times slower than non-DP training for large batch sizes, which is prohibitive in practice.
Vectorized implementations using ghost clipping or gradient accumulation tricks can reduce this overhead considerably. Ghost clipping reformulates the per-example gradient computation as a sequence of efficient matrix operations, bringing the overhead down to roughly 2-4x compared to non-DP training. The Opacus library implements ghost clipping for common layer types (linear, convolution, embedding) and is the standard tool for practical DP training of large models.
For large language models, this overhead is compounded by the model's size. Storing per-example gradients requires memory proportional to the batch size times the number of parameters, which is infeasible for a billion-parameter model. Gradient checkpointing and mixed-precision training are typically needed to make DP fine-tuning of large models feasible, and even then, DP training may require 2-4x more GPU memory than non-DP training.
When DP Is and Is Not Appropriate
Differential privacy makes the most sense when:
- The training data contains sensitive information about identifiable individuals.
- The adversary has substantial computational resources and may be motivated to attempt membership inference or data extraction attacks.
- The deployment context requires demonstrable, auditable privacy guarantees (e.g., healthcare, finance, legal).
- You have access to a large dataset or a large pre-trained model (to manage the privacy-utility tradeoff).
DP is less beneficial when:
- The training data is already fully public (DP protects training examples, not public information).
- The dataset is very small and DP degrades utility to the point of uselessness.
- The privacy requirement is not about individuals but about organizations or aggregate patterns.
- The threat model involves side-channel attacks, social engineering, or other non-algorithmic vectors that DP cannot address.
Understanding the scope and limits of DP's guarantee is essential for using it appropriately rather than treating it as a universal solution.
Summary
Differential privacy gives machine learning engineers a principled, mathematically rigorous way to limit how much any single training example influences the final model. The core idea is to bound the sensitivity of gradients via clipping and then add calibrated Gaussian noise, giving the DP-SGD algorithm. The privacy cost accumulates across training steps and is tracked via a privacy accountant, with the total cost expressed as : smaller means stronger privacy.
The key technical ingredients are:
- The guarantee: bounds the log-odds ratio by which the model can reveal whether any particular record was included in training.
- Gradient clipping: bounds the sensitivity of the training algorithm by ensuring no single example can shift the gradient aggregate by more than in norm.
- Gaussian noise: adds calibrated noise scaled to to hide the contribution of any clipped gradient within the noise floor.
- Privacy accounting: tracks the cumulative privacy cost via the moments accountant or RDP accountant, enabling precise budget management.
- Subsampling amplification: the fact that each example participates in only a fraction of training steps provides free privacy amplification, making the effective per-step cost roughly rather than .
Key takeaways for practitioners:
- DP guarantees are worst-case: Unlike heuristic defenses, DP protection holds even against adversaries with arbitrary auxiliary knowledge and computation.
- The privacy budget is additive: Every training step spends epsilon, and the total cost must be accounted for and reported, including hyperparameter tuning experiments.
- Scale is your friend: Large models and large datasets dramatically improve the privacy-utility tradeoff. DP training is most practical when fine-tuning large pre-trained models on private downstream data.
- DP-LoRA is a practical default: Combining DP with parameter-efficient methods like LoRA reduces noise dimensionality and provides better accuracy at a given epsilon.
- Fairness and hyperparameter tuning are real concerns: DP disproportionately hurts minority subgroups, and tuning experiments consume privacy budget that must be counted.
- Empirical auditing is essential: Theoretical guarantees can be invalidated by implementation bugs; canary-based auditing provides empirical validation.
The next chapter explores interpretability techniques for understanding what language models have learned, which is the other side of the accountability coin: differential privacy tells us how much individuals influenced the model, while interpretability tells us what the model knows.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about differential privacy.
Differential Privacy Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!