Part of Language AI Handbook
Covers diffusion models from the DDPM objective to latent diffusion, classifier-free guidance, image editing techniques, ControlNet, and FID evaluation metrics.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Image Generation: Diffusion Models and Text-to-Image Synthesis
For most of computing history, creating images required human artists. Generating a novel photorealistic scene, a painting in a specific style, or an illustration from a text description was something only people could do. That changed dramatically in the early 2020s, when diffusion models transformed image synthesis from a niche research curiosity into a technology capable of producing images indistinguishable from photographs and paintings. Today, you can type a sentence and receive a detailed, coherent image in seconds. Understanding how this works, at the level of math and mechanism, is the goal of this chapter.
The story of image generation is a story about probability. Generating an image means sampling from a distribution over pixels. The challenge is that this distribution is astronomically high-dimensional (a RGB image has nearly 800,000 dimensions), highly structured, and deeply non-Gaussian. Getting samples that look like photographs, not random noise, requires a model that has learned the complex correlations between pixels at every scale, from fine texture to global composition.
This chapter covers the full arc of modern image generation. We start with the historical context, then examine the mathematical foundations of diffusion: how adding and removing noise creates a tractable generative process. We then examine how language is connected to image synthesis, enabling text-to-image generation. We cover image editing techniques that build on generation, and we discuss how to evaluate the quality of generated images. By the end, you will understand how these systems work, why diffusion became the dominant paradigm, and what its real limitations are.
From GANs to Diffusion: A Brief History
Before diffusion models, the dominant paradigm for image generation was the Generative Adversarial Network (GAN), introduced by Goodfellow et al. in 2014. GANs set up a game between two neural networks: a generator that produces fake images and a discriminator that tries to distinguish real from fake. The generator learns by fooling the discriminator; the discriminator improves by catching the generator. In equilibrium, the generator produces realistic images.
GANs were powerful but notoriously difficult to train. They suffered from mode collapse, a failure mode where the generator produces only a narrow variety of images despite being trained on diverse data. A generator that outputs a single convincing image of a golden retriever technically fools the discriminator, but it has not learned the full distribution. Training instability was equally problematic: the generator and discriminator could fall into unproductive cycles, with the generator oscillating between certain artifact patterns rather than converging. Researchers spent years developing techniques like progressive growing (training first at low resolution and gradually increasing), spectral normalization, gradient penalties, and careful architectural choices just to stabilize training.
Variational Autoencoders (VAEs) offered a different approach. A VAE encodes images into a compact latent distribution, parameterized by a mean and variance, then samples from that distribution to decode new images. VAEs were stable to train and provided an explicit latent space suitable for interpolation and manipulation. However, they produced notoriously blurry outputs. The reason is their reconstruction loss: optimizing pixel-level mean squared error encourages the decoder to average over plausible reconstructions, producing a blurry mean rather than any specific sharp image. The latent bottleneck adds to this averaging pressure by compressing multiple plausible outputs into a single latent point.
Autoregressive models like PixelCNN and VQ-VAE-based approaches (DALL-E v1, ImageGPT) took a third path: model the joint pixel distribution as a product of conditionals, generating one pixel (or patch, or token) at a time. These models can generate very high-quality images but are extremely slow at inference time because generation is sequential. VQ-VAE with a transformer prior (used in the original DALL-E) combined discrete tokenization of images with autoregressive modeling, allowing text-conditioned generation, but still required hundreds to thousands of forward passes.
Diffusion models emerged as an alternative that combined the stability of VAE training with quality approaching or exceeding GANs. Instead of a generator-discriminator game or an encoder-decoder compression, diffusion models learn to reverse a gradual noise process. This framing turns image generation into a sequence of denoising steps, each more tractable than generating an image from scratch. The main point is elegantly simple: if you can reliably remove a small amount of noise, you can chain those removals together to generate images from nothing.
Diffusion Models: The Core Idea
The fundamental insight of diffusion models is this: if you have a process that gradually destroys structure, you can learn its reverse to build structure back.
Consider what happens when you gradually add Gaussian noise to an image. At first, the image is recognizable but slightly grainy. After more noise is added, the details blur and only global structure remains. Continue adding noise, and eventually you have something indistinguishable from pure random static, regardless of what the original image was. This forward process destroys information in a principled, mathematically tractable way.
Now flip the perspective. If you could learn to undo each small noise step, you could start with pure random noise and gradually recover a coherent image. You do not need to generate a complete image in one shot; you only need to denoise slightly at each step. This is dramatically more tractable because each denoising step is a small, local change to a partially formed image, rather than a jump from noise to image.
The Forward Process: Adding Noise
The forward process (also called the diffusion process) takes a real image and gradually adds Gaussian noise over timesteps:
where:
- : the noisy image at timestep
- : the image at the previous (less noisy) timestep
- : the noise schedule at timestep , controlling how much noise is added
- : a Gaussian distribution with mean and variance
Each step scales down the signal by (to preserve approximate unit variance) and adds Gaussian noise with variance . The scaling factor ensures that the overall variance of remains roughly constant over time: as signal energy decreases, noise energy increases by the same amount. With a carefully designed noise schedule, after steps (typically ), is approximately pure Gaussian noise, regardless of what was.
A mathematical property makes diffusion computationally tractable: you can compute the noisy image at any arbitrary timestep directly from , without simulating each intermediate step. This is possible because Gaussian distributions have the property that the composition of Gaussians is itself Gaussian. Define and (the cumulative product of all values up to step ). Then:
which means we can sample directly as:
where is pure Gaussian noise. This "reparameterization" has major practical implications. During training, you do not need to simulate the full forward chain from to for every training example. Instead, for each training image you pick a random timestep , sample noise , compute the corresponding using the formula above, and train the model on that single noisy image. This makes training embarrassingly parallel across timesteps.
The Noise Schedule
The noise schedule controls how quickly structure is destroyed as noise is added. Its design has a significant impact on the quality of the learned model.
A natural first choice is a linear schedule: increases linearly from a small value (like ) to a larger value (like ). This is easy to understand, but it has a practical flaw: it destroys too much structure too quickly at low noise levels, where most of the perceptual information lives, and too little at high noise levels. The model ends up spending many timesteps learning to denoise nearly-Gaussian noise, which carries little useful gradient signal for learning image structure.
The cosine schedule, proposed by Nichol and Dhariwal, addresses this by defining directly rather than working through :
where:
- : the total number of timesteps
- : a small offset (typically 0.008) that prevents from being too large near
- The cosine function ensures decreases smoothly from 1 to near 0 following a cosine curve
The cosine schedule keeps near 1 for a longer initial period before dropping more smoothly to near zero at the end. This allocates more timesteps to informative intermediate noise levels, where the model can learn meaningful structure. Empirically, cosine schedules produce better image quality than linear ones, especially for higher-resolution images.
More recent work has explored learned noise schedules (optimizing as parameters) and continuous-time diffusion (defining the noise process as a stochastic differential equation rather than discrete steps), but linear and cosine schedules remain widely used because of their simplicity and reliable performance.
The Reverse Process: Learning to Denoise
The reverse process aims to invert the forward process: starting from pure noise , progressively denoise to recover a clean image . The true reverse conditional distribution exists but is intractable to compute directly, because it requires integrating over all possible clean images consistent with the noisy image . However, if we condition on the clean image, the reverse conditional becomes tractable:
where is the posterior variance. This formula gives the optimal denoising direction when the clean image is known. In practice, we do not know during generation, so we learn a neural network to approximate the reverse conditional:
where is a neural network parameterized by that predicts the mean of the reverse distribution.
Rather than directly predicting the mean, most implementations train the network to predict the noise that was added. This is the -prediction formulation. Given the predicted noise , the reverse step mean becomes:
where:
- : un-scales the signal by reversing the scaling applied in the forward step
- : subtracts the predicted noise contribution, pushing the estimate toward the clean image
The -prediction formulation is used in practice because it produces stable training dynamics: the loss is well-scaled across all timesteps, and the noise prediction task is a well-defined, bounded regression problem. An alternative is -prediction (predicting the clean image directly), which is sometimes more stable at low noise levels, and -prediction (predicting a linear combination of noise and signal), which was introduced for improved training at high resolutions. Modern models often blend these formulations.
The Training Objective
The full diffusion training objective is derived from a variational lower bound on the log-likelihood of the data. Maximizing directly is intractable, but we can instead maximize a lower bound that decomposes into per-timestep KL divergences between the learned reverse process and the true posterior. Ho et al. (2020) showed that a simplified version of this objective works just as well in practice:
where:
- : a uniformly randomly sampled timestep
- : a real training image sampled from the training dataset
- : noise sampled from a standard Gaussian
- : the noisy image at timestep using the reparameterization trick
- : the neural network predicting the noise from the noisy image and timestep
In words: pick a random training image, pick a random timestep, add the corresponding amount of noise, and train the network to predict what noise was added. This is a denoising task framed differently at each timestep. At high , the image is mostly noise and the model must infer global structure. At low , the image is mostly clean and the model must predict the residual fine-grained noise.
The simplification from the full variational lower bound to drops per-timestep weighting that appears in the variational objective, trading theoretical optimality for empirical training stability. In practice, models trained with generate better images than those trained with the full variational bound, likely because the uniform weighting prevents some timesteps from dominating the gradient signal.
Diffusion models are closely related to score-based generative models. The "score" of a distribution is the gradient of the log probability with respect to the data: . If you know the score function everywhere, you can generate samples by starting from any point and following the gradient toward regions of higher probability, a procedure called Langevin dynamics. Song and Ermon (2019) showed how to estimate score functions at multiple noise levels through denoising score matching, and Song et al. (2021) unified this with diffusion models by showing that both processes are solutions to stochastic differential equations. Training a diffusion model to predict noise is mathematically equivalent to learning the score of a noisy data distribution at each noise level. This perspective has been productive for theoretical analysis and for developing improved samplers.
The U-Net Architecture for Noise Prediction
The neural network needs to accept a noisy image and a timestep, and output a same-sized noise prediction. The dominant architecture is a U-Net, originally developed by Ronneberger et al. for biomedical image segmentation, adapted for the diffusion setting.
The U-Net's key structural property is its encoder-decoder architecture with skip connections. The encoder path progressively downsamples the image using strided convolutions or pooling, building a hierarchy of features from fine-grained local patterns to coarse global structure. The decoder path progressively upsamples back to the original resolution using transposed convolutions or bilinear upsampling. Skip connections copy feature maps directly from each encoder stage to the corresponding decoder stage, allowing fine-grained spatial details to flow directly to the decoder. Without skip connections, the encoder bottleneck would force all spatial information through the narrow representation, causing loss of fine detail.
For diffusion, the U-Net is adapted with several modifications:
- Timestep embeddings: the scalar timestep is embedded via sinusoidal positional encodings (similar to those we discussed in Part XIV on positional encoding) and projected through small feed-forward networks. These embeddings are added to the feature maps at each resolution level, allowing the network to know at which noise level it is operating.
- Residual blocks: each U-Net block wraps its core convolutions in residual connections, improving gradient flow during training and allowing deeper networks.
- Self-attention layers: attention layers are added at lower resolutions where the spatial dimensions are small, making attention computationally feasible. These allow the network to model long-range dependencies between distant spatial locations. A global composition constraint, for example a sky being blue and grass being green, requires correlating distant pixels.
- Group normalization: normalization is applied within feature groups rather than across batch dimensions, which is more stable for the small batch sizes used in diffusion training.
The result is a flexible model that can process images at their full resolution while conditioning on the noise level. Because the same network handles all timesteps, it learns a range of denoising behaviors: at high noise (early in generation), it focuses on global structure and low-frequency signals; at low noise (late in generation), it refines fine details and textures. The model implicitly learns a curriculum from coarse to fine.
Latent Diffusion Models
A key limitation of pixel-space diffusion is computational cost. Running a U-Net over full-resolution images for hundreds of timesteps is expensive. A image has 786,432 pixels; processing this at each of 1000 timesteps during training requires enormous GPU memory and compute. Pixel-space diffusion models like the one used in DALL-E 2's decoder were trained on thousands of GPU-days.
Latent Diffusion Models (LDMs), introduced by Rombach et al. (2022) and forming the basis of Stable Diffusion, solve this by moving the diffusion process into a compressed latent space.
The approach uses a separate variational autoencoder (VAE) that:
- Encodes images to a compact latent representation: , where has spatial dimensions roughly smaller than
- Decodes latent representations back to images:
The VAE encoder-decoder is trained separately with a combination of losses. A reconstruction loss (typically L1 or L2 in pixel space) ensures the decoded image resembles the original. A perceptual loss (computing similarity in the feature space of a pre-trained VGG network rather than raw pixel space) ensures the reconstruction captures high-level semantic content, not just low-level pixel patterns. A GAN loss discriminates between real and decoded images, encouraging sharp textures and preventing blurry averaging. Additionally, a KL regularization term keeps the latent distribution close to a standard Gaussian, enabling sampling.
Once trained, the encoder-decoder is frozen. The diffusion model then operates exclusively on latent codes rather than on pixels. Because latent codes are much smaller (typically with 4 channels for a image), the U-Net processes a element feature map rather than a pixel image. This is a 64x reduction in spatial resolution, reducing the computational cost of each forward pass by roughly two orders of magnitude.
The practical consequence is significant. Models like Stable Diffusion can be trained and run on consumer GPUs with 8-24 GB of VRAM, whereas pixel-space models of comparable quality required data center hardware. This democratization of image generation through latent compression was arguably as important as the diffusion algorithm itself.
The latent space also has a useful structural property: it is smoother and more semantically organized than pixel space. Nearby points in latent space correspond to semantically similar images, which means the diffusion model can learn a smoother, more coherent distribution compared to the raw pixel distribution that contains abrupt transitions between adjacent pixel values.
Text-to-Image Generation
Generating images from text descriptions requires connecting language to the visual generation process. This is done through conditioning: the noise-prediction network receives the noisy image, the timestep, and a text representation that guides the generation.
The general principle is that the model learns a conditional distribution where is the text condition, rather than the unconditional . Every U-Net forward pass receives both the noisy image and the text embedding, allowing the noise prediction to be guided by the semantic content of the prompt.
Classifier-Free Guidance
The key technique for controlling text-conditioned generation is classifier-free guidance (CFG), introduced by Ho and Salimans. The motivating problem is that while a conditional diffusion model generates images corresponding to the text, the correspondence can be weak: the generated image might include elements from the prompt but not strongly emphasize the distinctive aspects that make the text description specific.
The idea is to train the model in two modes simultaneously: conditioned on text, and unconditioned (with the text randomly dropped out, typically 10-20% of the time by replacing the text conditioning with a null embedding). At inference, you run both modes and interpolate their noise predictions:
where:
- : the text conditioning (encoded prompt)
- : the null conditioning (empty or dropped text)
- : the guidance scale, controlling how strongly the generation follows the text
- : noise prediction conditioned on text
- : unconditional noise prediction
When , this reduces to the standard conditioned prediction. When (typical values are 7-12), the conditional direction is amplified beyond the raw conditional prediction: the model is pushed more strongly toward the text-conditioned manifold. When , the result is fully unconditional.
Why does this work? The term captures how much the denoising direction changes when conditioned on text versus not. This difference is the direction in noise prediction space that corresponds to "move toward samples that fit the text condition." Scaling this up with large amplifies the text-driven signal, pushing the generation more aggressively toward images that match the prompt.
The tradeoff is that higher guidance scales reduce diversity. With , the model produces images that very closely match the prompt but all tend to look similar; with , there is more variation at the cost of sometimes weaker prompt adherence. The guidance scale is one of the most important inference-time hyperparameters.
Classifier-free guidance was preceded by classifier guidance, introduced by Dhariwal and Nichol (2021). Their approach used a separately trained image classifier to guide the diffusion process. At each denoising step, the classifier computes , the gradient of the log probability that the noisy image belongs to class . This gradient is added to the noise prediction, pushing generation toward the target class. Classifier guidance demonstrated that guided diffusion could surpass GANs on image quality metrics. Classifier-free guidance achieves similar or better effects without requiring a separate classifier, making it applicable to any conditioning modality: text, images, poses, style references. It is not limited to predefined class labels.
Text Encoding
The text conditioning is typically derived from a pre-trained language model that maps the prompt to a sequence of dense embeddings.
CLIP text encoder: Used in DALL-E 2 and Stable Diffusion 1.x, the CLIP text encoder was trained to align with visual representations through contrastive learning on image-text pairs, as we discussed in the CLIP chapter (Part L). This makes CLIP text embeddings naturally suited for conditioning image generation: they encode language in a space that is semantically linked to visual concepts. The embedding captures semantic meaning in a vision-aware way, which helps the diffusion model bridge language and image.
Large language model encoders: Imagen (Google, 2022) used T5-XXL as its text encoder, a 4.6B parameter language model not specifically trained for visual alignment. Despite lacking CLIP's vision-language grounding, T5 embeddings proved highly effective because T5 handles linguistic structure and negation alongside composition and rare concepts better than CLIP's text encoder. The T5 embeddings carry richer linguistic context, which improves generation of complex, compositional prompts.
Multiple encoder combination: Stable Diffusion XL (SDXL, 2023) and Stable Diffusion 3 (SD3, 2024) use multiple text encoders simultaneously, combining OpenCLIP's ViT-bigG text encoder with the original CLIP ViT-L encoder (in SDXL) or with T5-XXL (in SD3). The outputs are concatenated or pooled together before being fed to the diffusion backbone. This combination captures both vision-aligned CLIP semantics and linguistically rich T5 representations. This provides broader coverage of both visual and textual concepts.
The text is tokenized and passed through the text encoder to produce a sequence of embeddings, one per token. These embeddings are injected into the U-Net through cross-attention layers: the image features at each spatial location act as queries, and the text embeddings serve as keys and values. Each spatial location in the image can thus attend to the text tokens most relevant to what should appear at that location, creating fine-grained text-image alignment. A text token for "red" might be strongly attended to by the image region where the red object should appear.
DALL-E 2 and the CLIP Prior
DALL-E 2 (OpenAI, 2022) took a distinctive two-stage approach to text-to-image generation, motivated by the observation that CLIP image embeddings capture semantic visual content in a rich, structured way.
The pipeline works as follows. First, a prior network generates a CLIP image embedding from the text prompt. This prior is itself a diffusion model (or autoregressive model) that learns the mapping from text embeddings to image embeddings in CLIP's joint space. Given a text prompt describing a scene, the prior generates a CLIP image embedding that represents the visual concept, without yet specifying pixel-level details.
Second, a decoder (a pixel-space diffusion model) generates an image from that CLIP image embedding. Because CLIP image embeddings encode semantic visual content rather than specific pixels, the decoder can focus on rendering visual content consistently with the embedding rather than interpreting language from scratch.
The prior bridges the gap between language and vision space. CLIP's joint embedding space aligns text and images semantically, but a single text prompt corresponds to many possible images. The prior models this one-to-many mapping, sampling different visual interpretations that are all consistent with the text. The result showed strong semantic faithfulness: images clearly corresponded to the described content, even for novel combinations.
Imagen and the Value of Text Understanding
Imagen (Google, 2022) demonstrated a surprising lesson: text understanding quality matters as much as image modeling quality for text-to-image generation.
Imagen uses a cascade of diffusion models. A base model generates images conditioned on T5-XXL text embeddings. Two super-resolution diffusion models then upscale the image to and then , each conditioned on both the original text and the lower-resolution image. The final output is a high-resolution image generated through this pipeline.
The key finding was that doubling the size of the text encoder (using T5-XXL at 4.6B parameters) improved image quality and text adherence more than doubling the size of the image diffusion model. This implies that the bottleneck in text-to-image generation is often the quality of text understanding, not image rendering. A model that does not understand "the woman standing to the left of the red building" will produce a plausible-looking image that fails to respect the spatial and attribute relationships in the text.
Imagen also introduced dynamic thresholding for classifier-free guidance. At high guidance scales, the predicted clean image values can fall outside the valid range , producing oversaturated images. Dynamic thresholding scales the entire prediction by the maximum absolute value if that exceeds a threshold, preserving relative magnitudes while keeping the image in range. This allows using higher guidance scales without oversaturation.
Stable Diffusion and Latent Diffusion
Stable Diffusion (CompVis/Stability AI, 2022) combined latent diffusion with CLIP text conditioning to create the first major open-source text-to-image model. Its architecture became the standard for the open-source ecosystem due to its balance of computational efficiency and image quality.
The key components are:
- A VAE encoder/decoder that compresses images to latent codes (spatial compression factor 8, with 4 latent channels)
- A CLIP ViT-L/14 text encoder (frozen) that produces 77 token embeddings of dimension 768
- A latent U-Net that performs diffusion in the compressed latent space, with cross-attention layers throughout for text conditioning
- A linear noise schedule with training steps
The open release of model weights enabled a large ecosystem: fine-tuning on specific styles and subjects (DreamBooth, Textual Inversion, LoRA), specialized control adapters (ControlNet, IP-Adapter), GUI tools (AUTOMATIC1111, ComfyUI), and commercial applications. SDXL (2023) scaled this architecture with a larger U-Net and dual CLIP encoders, achieving substantially better image quality and prompt adherence.
Stable Diffusion 3 (2024) moved away from the U-Net architecture toward a Diffusion Transformer (DiT) backbone, using multimodal attention layers that jointly process both image patches and text tokens in a single transformer block, rather than separate processing with cross-attention injection.
Sampling and Inference
Training a diffusion model gives us the denoising network, but generating images requires running the reverse process. This is the sampling procedure, and the design of samplers has been an active area of research because the original DDPM sampler was far too slow for practical use.
DDPM Sampling
The original Denoising Diffusion Probabilistic Models (DDPM) sampler runs the full reverse steps:
where:
- : the inverse scaling factor for this step
- : the denoising correction based on the predicted noise
- : additional stochastic noise injected at each step, where
With , generating one image requires 1000 neural network forward passes through the full U-Net. At 8-30 milliseconds per forward pass on a modern GPU, this translates to 8-30 seconds per image, far too slow for interactive use.
DDIM Sampling
Denoising Diffusion Implicit Models (DDIM), introduced by Song et al. (2021), made sampling dramatically faster by reformulating the reverse process as a deterministic ordinary differential equation (ODE) rather than a stochastic process.
The key insight is that many different reverse processes can share the same marginals as DDPM while taking larger, more efficient steps. DDIM defines a non-Markovian reverse process that allows skipping many timesteps. Instead of going through all 1000 timesteps in order, DDIM can take evenly spaced steps visiting only 20 timesteps out of 1000, while still producing high-quality images.
The deterministic nature of DDIM has an additional useful property: it acts as an encoder as well as a decoder. You can run the DDIM forward process on a real image to find the latent noise representation corresponding to that image (called DDIM inversion), then modify the generation process at that latent to edit the image. This is not possible with stochastic DDPM because the stochasticity makes inversion non-invertible.
DPM-Solver and Modern Samplers
Subsequent work formalized diffusion sampling as solving a differential equation and applied numerical ODE analysis to improve it. DPM-Solver and its successor DPM-Solver++ (Lu et al., 2022) treat the reverse process as an ODE with a known analytical structure and apply higher-order solvers. These can produce high-quality images in as few as 10-20 steps by taking larger, more accurate integration steps.
The key insight is that the diffusion reverse process has a semi-linear structure: the linear component (corresponding to Gaussian noise removal) can be solved analytically, and only the nonlinear component (the U-Net's neural denoising) needs to be approximated numerically. By exploiting this structure, DPM-Solver takes each integration step far more accurately than DDIM, which uses a simpler first-order approximation.
Other popular samplers include PNDM (Pseudo Numerical Methods for Diffusion Models), Euler and Euler a (ancestral Euler methods with optional stochasticity), and Heun (a second-order Runge-Kutta variant). The choice of sampler affects speed and the character of the output: deterministic samplers give reproducible results from a fixed seed, while stochastic samplers add variation and can sometimes escape local optima in the generation trajectory.
Image Editing
Beyond generating new images from scratch, diffusion models can modify existing images in semantically meaningful ways. This is possible because the forward and reverse processes can be applied to noise and to any image.
SDEdit
SDEdit (Meng et al., 2021) is the simplest editing approach. You take a reference image, add noise to a specific level , then run the reverse process with conditioning. This approach does not require any training changes; it uses a pre-trained diffusion model directly.
The noise level controls the tradeoff between fidelity to the original and freedom to change. High means more noise is added, so the model can make larger structural changes but may deviate substantially from the original content. Low preserves the original closely but allows only subtle texture or color edits. This makes an intuitive dial: setting it at keeps the rough composition and changes details, while allows more substantial changes to shape and content.
SDEdit works well for sketch-based generation (start with a rough hand-drawn sketch, add noise, condition on text), style transfer (starting from a content image), and small targeted edits. Its simplicity is both its strength (no additional training needed) and its limitation (it cannot make precise, localized edits).
Inpainting
Inpainting fills a masked region of an image while maintaining consistency with the surrounding context. This is useful for removing unwanted objects, filling in damaged areas, or adding new content to existing images.
During the reverse process, at each timestep, the unmasked regions of the noisy image are replaced with the known noisy image at that level:
where:
- : the binary mask (1 for pixels to keep, 0 for pixels to generate)
- : the original image noised to level using the forward process
- : the image generated so far by the reverse process
- : element-wise multiplication
At each reverse step, the generated regions are denoised by the network while the known regions are reset to their noisy versions (forward-noised at exactly timestep ). This ensures that in the final output, the known regions are exactly the original pixels and the generated region is produced by the diffusion model under the constraint of fitting smoothly with the surroundings.
A limitation of this naive approach is that the seam between generated and preserved regions can be sharp, because the generated region is optimized without full awareness of the boundary constraint. Improved inpainting models are fine-tuned with masked training examples, teaching the model to explicitly condition on both the unmasked image content and the generation task at the boundary.
InstructPix2Pix
Rather than conditioning on a masked region, InstructPix2Pix (Brooks et al., 2022) conditions on both an original image and a text instruction describing the desired edit, such as "make the sky sunset" or "add glasses to the person."
Training this model required a creative data generation pipeline, since pairs of (original image, edit instruction, edited image) are not naturally available at scale. The solution was to synthesize training data: GPT-3 generated pairs of (image caption, edit instruction, edited image caption), and a text-to-image model (DALL-E) generated image pairs from the caption pairs. The diffusion model was then trained to perform the transformation given both the original image and the instruction.
At inference, InstructPix2Pix uses a modified CFG that combines guidance from both the text instruction and the original image, allowing control over how much the edit changes the image versus how literally the instruction is followed.
DreamBooth and LoRA Fine-Tuning
A related class of techniques personalizes image generation to specific subjects. DreamBooth (Ruiz et al., 2022) fine-tunes the entire diffusion model on 3-20 images of a specific person, object, or style. The appearance is bound to a rare token, like "sks dog," that is unlikely to appear in the original training data. After fine-tuning, prompts containing that token generate the specific subject in novel contexts: "a photo of sks dog in the snow" produces the specific dog in a snowy scene.
The challenge with full fine-tuning is the risk of language drift: the model may overfit to the specific images and forget other concepts. DreamBooth addresses this with a prior preservation loss that continues training on text-generated images of the general class (regular dogs, in this example) alongside the specific subject images.
Low-Rank Adaptation (LoRA), which we covered in Part XXXV on parameter-efficient fine-tuning, addresses both the cost and drift concerns by fine-tuning only small low-rank matrices inserted into the attention layers. Instead of updating all weight matrices, LoRA decomposes the update as where and with small rank (typically 4-16). LoRA files for Stable Diffusion are typically 50-200 MB rather than the several GB of full model weights, and they can be combined by summing their contributions with different weights, allowing compositional style mixing.
Textual Inversion
Textual Inversion (Gal et al., 2022) takes an even lighter-weight approach: rather than changing the model weights at all, it learns a new token embedding that represents the target concept. The denoising network is entirely frozen; only a new token embedding is optimized to minimize the denoising loss on the reference images:
where is the text conditioning derived from a string containing the new learned token . The learned token lives in the same embedding space as the model's existing vocabulary, so it can be combined with other words naturally.
Textual Inversion requires even fewer parameters than LoRA (a single token embedding vector rather than low-rank weight matrices) but captures concepts less flexibly because it is limited to the expressiveness of the fixed embedding space. Complex concepts that require weight adaptations to represent well are better suited to LoRA or DreamBooth.
Generation Control and Conditioning
Beyond text, diffusion models can be conditioned on many other signals to give precise control over the output. This is important for professional workflows where the artist needs more control than text alone provides.
ControlNet
ControlNet (Zhang and Agrawala, 2023) adds spatial conditioning signals such as edge maps, pose estimates (OpenPose skeleton), depth maps, surface normal maps, segmentation masks, or text line positions. The architecture addresses a key challenge: how to add conditioning without retraining the entire model.
ControlNet creates a trainable copy of the U-Net encoder. This copy processes the conditioning input (an edge map, for example) and adds its activations to the corresponding layers of the original U-Net's skip connections. The original U-Net is frozen; only the encoder copy and its connection weights are trained.
This design preserves the original U-Net's learned image priors while adding spatial control. Because the conditioning signal is spatial (the same height and width as the image), the network can precisely control image composition.
OpenPose conditioning generates humans in specific poses by feeding a skeleton image as the control signal. Canny edge map conditioning generates images that respect precise boundaries. Depth conditioning preserves 3D structure from a reference image while changing textures or style. This is particularly valuable for architectural rendering (conditioning on floor plans or line drawings), character animation (conditioning on pose sequences from video), and product photography (conditioning on the shape of a product).
Multiple ControlNet models can be stacked, applying combined conditioning. A generation might use edge conditioning for structural fidelity and pose conditioning for character positioning simultaneously.
IP-Adapter
IP-Adapter (Ye et al., 2023) conditions the diffusion model on a reference image's visual style or content, rather than or in addition to a text prompt. You provide an image whose style, subject, or composition you want to replicate in new generations.
IP-Adapter encodes the reference image with a vision encoder (typically CLIP's visual encoder) and injects those embeddings through a separate cross-attention pathway running in parallel to the existing text cross-attention. The image features serve as additional keys and values alongside the text features, with learnable weights determining the balance between text and image conditioning.
This allows tasks that are difficult to specify through text alone: style transfer (generate this scene in the visual style of this reference painting), face consistency (generate different scenarios featuring a specific person's face), and product placement (generate marketing images featuring a specific product). Because IP-Adapter is an adapter added on top of the frozen base model, it is compatible with any ControlNet adapters and LoRA fine-tuned models for the same base.
Evaluation of Generated Images
Evaluating image generation quality is difficult. Human perception is the gold standard, but human evaluation is expensive and slow, with limited reproducibility. Several automated metrics have been developed, each capturing different aspects of quality.
Frechet Inception Distance (FID)
FID is the most widely used automated metric for measuring image generation quality. It measures the distance between the distribution of generated images and the distribution of real images, not in pixel space, but in the feature space of a pre-trained InceptionV3 network.
The idea is that features from a network trained on ImageNet capture semantically meaningful aspects of image content: textures and shapes alongside structures and high-level objects. If the distribution of these features in the generated images matches the distribution in the real images, the generated images capture the same semantic variety and realism as the real data.
Given real images and generated images , their Inception features are extracted and fit to multivariate Gaussians with means , and covariances , . FID computes the Frechet distance between these Gaussians:
where:
- : squared distance between mean feature vectors, measuring distributional shift in average content
- : trace of a matrix (sum of diagonal elements)
- : matrix square root of the product of covariances, measuring structural similarity
Lower FID is better: a perfect model matching the real data distribution exactly would achieve FID of 0. State-of-the-art models achieve FID scores below 3 on standard benchmarks like MS-COCO or ImageNet.
FID has known limitations. It is sensitive to the number of samples used: you need at least 10,000 generated images for reliable estimation, and the value changes substantially with fewer samples. It depends on the specific feature extractor (InceptionV3 may not capture the aspects of quality most relevant to human preference). It measures the full distribution, so a model that generates many diverse but lower-quality images can outscore one that generates fewer but higher-quality images on FID. Despite these limitations, it remains the standard comparison metric in published research.
CLIP Score
For text-to-image evaluation, CLIP Score measures the semantic alignment between a generated image and its text prompt, directly using the CLIP model's joint embedding space.
CLIP Score computes the cosine similarity between the CLIP image embedding of the generated image and the CLIP text embedding of the prompt:
where:
- : the normalized CLIP image embedding of the generated image
- : the normalized CLIP text embedding of the prompt
- : a scaling factor (typically 100) that maps cosine similarities into a human-interpretable score range
- : clips negative similarities to zero, since negative cosine similarity indicates opposite directions in the embedding space
A score of 25-30 is typical for high-quality text-image aligned models. CLIP Score captures semantic fidelity well: does the image visually depict the concepts mentioned in the text? But it does not capture visual quality or photorealism, and it has blind spots wherever CLIP's training data was sparse. It is also susceptible to adversarial inputs: images that score highly on CLIP but look strange to human observers. It is most useful for comparing text-image alignment across models or studying how different prompting strategies affect faithfulness, not as a standalone quality measure.
Human Evaluation
Automated metrics often disagree with human judgment about what makes an image "good." Human evaluation remains the gold standard for capturing subjective quality dimensions that metrics miss.
Human evaluation typically asks annotators to:
- Compare two images generated from the same prompt and indicate which is better (pairwise preference comparison)
- Rate an image on overall quality and photorealism, text-image alignment, diversity, aesthetic appeal
- Evaluate specific failure modes: distorted hands, unnatural textures, implausible lighting, artifacts in background regions, or incorrect object counts
Pairwise comparisons are often more reliable than absolute ratings because they avoid systematic biases in how different annotators interpret rating scales. The Bradley-Terry model can convert pairwise preferences into global scores. Modern leaderboards like HPSv2 and ELO-based systems collect large numbers of human pairwise comparisons and use them to rank models.
Human evaluation is expensive to scale to thousands of prompts and model comparisons, but it remains the most reliable signal for the aspects of quality that matter most in practice: does the image look good to a human? Would a person be satisfied receiving this image in response to their prompt?
Precision and Recall
Precision and recall for generative models, as defined by Kynkaanniemi et al. (2019), measure two distinct aspects of generation quality:
- Precision: Do generated images look realistic? Are they within the manifold of real images? High precision means the model rarely generates implausible images that fall outside what real images look like.
- Recall: Does the model cover the full diversity of the real distribution? High recall means the model can generate images across the full range of visual variety present in the training data.
A model that generates only highly photorealistic images of a narrow slice of the distribution would have high precision but low recall: it generates quality images but misses most of what real images look like. A model that generates highly varied images that are sometimes blurry or contain artifacts would have high recall but low precision.
These metrics are computed using nearest-neighbor distances in the InceptionV3 feature space. For precision, we ask: for each generated image, does a real image exist nearby? For recall: for each real image, does a generated image exist nearby?
Precision and recall provide more detail than FID alone. FID is a single scalar that conflates fidelity and diversity; precision and recall decompose these into separate measurements, allowing diagnosis of whether a model's failure mode is generating unrealistic images (low precision) or missing diversity (low recall). Increasing guidance scale, for example, typically increases precision (higher quality) at the cost of recall (less diversity).
Implementation: Building a Mini Diffusion Pipeline
Let's implement a simplified diffusion training and sampling loop to make these concepts concrete. We will not train a full-scale text-to-image model (that requires significant compute and large datasets), but we will build a complete DDPM pipeline that learns to generate a structured 2D distribution. This allows us to visualize the full forward-reverse process and verify that the math works as described.
Setup and Data
We will work with a toy 2D dataset: a two-moons arrangement. This lets us visualize the full forward-reverse process clearly. Despite its simplicity, this dataset is non-trivial for generative models because it has a bimodal, non-Gaussian shape with a curved manifold.
import numpy as np
import torch
def make_moons_dataset(n_samples=2000, noise=0.05, seed=42):
"""Create a simple two-moons 2D dataset."""
rng = np.random.RandomState(seed)
n_samples_each = n_samples // 2
theta1 = np.linspace(0, np.pi, n_samples_each)
theta2 = np.linspace(np.pi, 2 * np.pi, n_samples_each)
x1 = np.column_stack([np.cos(theta1), np.sin(theta1)])
x2 = np.column_stack([np.cos(theta2) + 1, np.sin(theta2)])
x = np.vstack([x1, x2]) + rng.randn(n_samples, 2) * noise
return x.astype(np.float32)
data = make_moons_dataset(n_samples=3000)
X = torch.tensor(data)
n_samples, d = X.shapeDataset shape: torch.Size([3000, 2]) Data range: [-1.136, 2.134] Data mean: [4.9965525e-01 3.2643635e-05]
Defining the Noise Schedule
The DiffusionSchedule class precomputes all the quantities derived from the noise schedule that we need for both training (forward process) and sampling (reverse process).
class DiffusionSchedule:
"""Linear noise schedule for DDPM."""
def __init__(self, T=200, beta_start=1e-4, beta_end=0.02, device="cpu"):
self.T = T
self.device = device
# Linear noise schedule
self.betas = torch.linspace(beta_start, beta_end, T, device=device)
self.alphas = 1.0 - self.betas
self.alphas_cumprod = torch.cumprod(self.alphas, dim=0)
# Precompute useful quantities
self.sqrt_alphas_cumprod = torch.sqrt(self.alphas_cumprod)
self.sqrt_one_minus_alphas_cumprod = torch.sqrt(
1.0 - self.alphas_cumprod
)
def q_sample(self, x0, t, noise=None):
"""Add noise to x0 at timestep t. Implements the forward process."""
if noise is None:
noise = torch.randn_like(x0)
sqrt_alpha = self.sqrt_alphas_cumprod[t].view(-1, 1)
sqrt_one_minus_alpha = self.sqrt_one_minus_alphas_cumprod[t].view(-1, 1)
return sqrt_alpha * x0 + sqrt_one_minus_alpha * noise, noise
schedule = DiffusionSchedule(T=200)Noise schedule: T=200 Alpha(t=0) = 0.9999 (almost no noise) Alpha(t=100) = 0.5964 (partial noise) Alpha(t=199) = 0.132183 (almost pure noise)
At , so the noisy image is nearly identical to the clean data. At , so the signal has been destroyed and we have nearly pure Gaussian noise. The middle timesteps correspond to partial noise: the data has some structure but is significantly corrupted.
Denoising Network
For our 2D toy problem, a small MLP handles noise prediction. In full image diffusion, this would be a U-Net with millions of parameters. The architecture here captures the essential design: a network that takes both the noisy data and the timestep, and outputs a noise prediction of the same shape as the input.
import torch.nn as nn
class SinusoidalPositionEmbedding(nn.Module):
"""Embed timestep t as a sinusoidal vector."""
def __init__(self, dim):
super().__init__()
self.dim = dim
def forward(self, t):
half_dim = self.dim // 2
emb = torch.log(torch.tensor(10000.0)) / (half_dim - 1)
emb = torch.exp(torch.arange(half_dim, device=t.device) * -emb)
emb = t[:, None].float() * emb[None, :]
return torch.cat([torch.sin(emb), torch.cos(emb)], dim=-1)
class DenoisingMLP(nn.Module):
"""MLP that predicts noise from noisy data and timestep."""
def __init__(self, data_dim=2, hidden_dim=256, time_dim=64):
super().__init__()
self.time_embed = SinusoidalPositionEmbedding(time_dim)
self.net = nn.Sequential(
nn.Linear(data_dim + time_dim, hidden_dim),
nn.GELU(),
nn.Linear(hidden_dim, hidden_dim),
nn.GELU(),
nn.Linear(hidden_dim, hidden_dim),
nn.GELU(),
nn.Linear(hidden_dim, data_dim),
)
def forward(self, x, t):
t_emb = self.time_embed(t)
return self.net(torch.cat([x, t_emb], dim=-1))
model = DenoisingMLP(data_dim=2, hidden_dim=256, time_dim=64)The sinusoidal timestep embedding encodes the scalar as a vector of alternating sines and cosines at different frequencies. This is the same approach used in transformer positional embeddings, and it ensures the model can distinguish all timesteps and interpolate smoothly between them. Without this embedding, the network would not know whether it is being asked to denoise heavily corrupted data (where only global structure matters) or lightly corrupted data (where fine details matter).
Model parameters: 149,250 Input shape: torch.Size([8, 2]), output shape: torch.Size([8, 2])
Training Loop
import torch.nn.functional as F
from torch.optim import Adam
def train_diffusion(model, schedule, X, epochs=3000, batch_size=256, lr=1e-3):
"""Train the denoising network using DDPM objective."""
optimizer = Adam(model.parameters(), lr=lr)
losses = []
for epoch in range(epochs):
# Sample random batch
idx = torch.randint(0, len(X), (batch_size,))
x0 = X[idx]
# Sample random timesteps
t = torch.randint(0, schedule.T, (batch_size,))
# Add noise according to schedule
xt, noise = schedule.q_sample(x0, t)
# Predict noise
noise_pred = model(xt, t)
# Loss: MSE between predicted and actual noise
loss = F.mse_loss(noise_pred, noise)
optimizer.zero_grad()
loss.backward()
optimizer.step()
if epoch % 200 == 0:
losses.append(loss.item())
return losses
losses = train_diffusion(model, schedule, X, epochs=3000, batch_size=256)Training epochs: 3000 Final loss: 0.5023 Initial loss: 1.1503 Loss reduction: 56.3%
The loss measures the mean squared error between the predicted noise and the actual noise added. As training progresses, the model learns to predict noise more accurately across all timesteps, and the loss decreases. A perfect model would achieve loss of 0, but in practice the irreducible loss reflects the stochasticity of the forward process.
DDPM Sampling
With a trained model, we run the reverse process to generate new samples. The sampling loop iterates from down to , denoising at each step.
@torch.no_grad()
def ddpm_sample(model, schedule, n_samples=1000, d=2):
"""Sample from the learned distribution using DDPM reverse process."""
model.eval()
# Start from pure noise
x = torch.randn(n_samples, d)
trajectory = [x.clone()]
for t in reversed(range(schedule.T)):
t_batch = torch.full((n_samples,), t, dtype=torch.long)
# Predict noise
eps_pred = model(x, t_batch)
# Compute denoised estimate
beta_t = schedule.betas[t]
alpha_t = schedule.alphas[t]
alpha_bar_t = schedule.alphas_cumprod[t]
# DDPM reverse step mean
coef = beta_t / torch.sqrt(1 - alpha_bar_t)
mu = (x - coef * eps_pred) / torch.sqrt(alpha_t)
if t > 0:
# Add stochastic noise for all but the last step
noise = torch.randn_like(x)
x = mu + torch.sqrt(beta_t) * noise
else:
x = mu
if t % 50 == 0:
trajectory.append(x.clone())
return x, trajectory
generated, trajectory = ddpm_sample(model, schedule, n_samples=2000)Generated 2000 samples in 2D Generated range: [-1.283, 2.232] Generated mean: [ 0.5127893 -0.04144546]
Notice that we add stochastic noise at every reverse step except the last (). This stochasticity is essential to DDPM sampling: without it, the reverse process would be deterministic from the initial noise, producing less diverse samples and sometimes getting stuck in local optima. At , we take the deterministic mean to obtain the final clean sample.
Visualizing the Noise Schedule

Visualizing Generated vs. Real Data


The generated distribution captures both crescent arcs, confirming that the model has learned the broad structure through denoising alone and avoided collapsing to one mode. The bridge and off-manifold points also show that this small denoising network is not a perfect density model. A GAN on this dataset would often collapse to one mode; diffusion models resist this failure mode more effectively.
Visualizing Training Loss

Visualizing Classifier-Free Guidance Scale
The guidance scale is one of the most impactful hyperparameters in text-to-image generation. The CFG formula decomposes the guided noise prediction into a linear combination of conditional and unconditional predictions. We can visualize exactly how the weights on each component change with different guidance scales.

At , the conditional noise prediction receives a weight of 7.5 and the unconditional prediction receives a weight of . The negative unconditional weight means the model actively moves away from unconditional predictions, amplifying whatever features the text condition adds. This is why high guidance scales can cause oversaturation: the model is maximally exploiting the direction of the text signal, sometimes beyond what produces photorealistic results.
Key Parameters
Diffusion model design involves several critical hyperparameters that jointly determine output quality and generation speed as well as controllability.
The key parameters are:
- T (number of timesteps): Controls granularity of the diffusion process. More steps allow finer-grained denoising but increase sampling time. Typical training values: 100-1000 steps. Typical inference values: 20-50 steps with advanced samplers like DPM-Solver++.
- Beta schedule (linear vs. cosine): The linear schedule is simple but allocates timesteps unevenly. The cosine schedule typically produces better results by spending more steps at informative intermediate noise levels.
- Guidance scale (w): Higher values increase text fidelity but reduce diversity and can cause oversaturation. Typical values: 7-12 for creative generation where strong prompt adherence is desired, 1-3 for realistic variation with more diversity.
- Latent compression factor: For latent diffusion, the ratio between image size and latent size. An 8x spatial compression (standard in Stable Diffusion) reduces compute substantially while maintaining quality. Too much compression loses fine detail; too little is computationally expensive.
- U-Net attention resolution: Self-attention is added at lower resolutions where the spatial dimensions are small. Adding attention at higher resolutions captures finer long-range dependencies but increases memory cost quadratically with the spatial size.
- VAE regularization strength: The KL term in VAE training controls how tightly the latent distribution matches a Gaussian. Too much regularization produces a degenerate latent space; too little allows the latent to become so non-Gaussian that the diffusion model cannot traverse it effectively.
Limitations and Practical Challenges
Diffusion models are impressive but face several real limitations that affect both research progress and practical deployment.
Slow sampling remains the most significant practical challenge for real-time applications. Even with advanced samplers like DPM-Solver++, generating a single high-quality image requires 20-50 neural network forward passes through the full U-Net. Each forward pass is a complete evaluation of a network with hundreds of millions to billions of parameters. On a high-end consumer GPU, this translates to 1-5 seconds per image for a 512x512 output. For comparison, GANs generate images in a single forward pass, typically under 100 milliseconds. Consistency Models (Song et al., 2023) attempt to address this by training a model that can generate directly from noise in one or two steps by distilling a diffusion model's multi-step trajectory. Flow matching approaches (Lipman et al., 2022) reframe diffusion as learning straight-line trajectories in distribution space, which allows larger, more efficient integration steps at inference. Both show promise but still trail full diffusion sampling in the highest quality regime.
Compositional failures are pervasive and poorly understood. Generating an image of "a red cube to the left of a blue sphere" reliably is surprisingly difficult for current models. The model must understand the individual objects, their spatial relationships, and their attribute bindings, and these frequently break down for unusual or complex combinations. "A black cat and a white dog" may produce a black dog and a white cat, because the model does not reliably associate attributes with their intended objects. Models trained with dense spatial supervision (from datasets with bounding boxes or segmentation masks) show improvement, but compositional reasoning remains below human level. The root cause is likely that the model learns statistical associations from image captions, which are inherently unordered and do not provide explicit spatial or relational information.
Text rendering in generated images has historically been poor. Diffusion models process text as visual patterns, not as sequences of characters with specific forms. Images containing text often show plausible-looking letterforms that spell nonsensical words. DALL-E 3 and similar models improved substantially on this through better data curation and training at higher resolutions. But text rendering remains below the reliability expected of professional design tools, especially for longer strings or unusual fonts.
Copyright and provenance present legal and ethical complexity that extends beyond the technical. These models are trained on internet-scale image datasets scraped without explicit consent, and they can reproduce stylistic elements from specific artists. When someone prompts "a painting in the style of [living artist]," the generated image can be commercially valuable and indistinguishable from a human production, raising serious questions about attribution and consent as well as economic impact on the artists whose work was used for training. The legal framework for AI-generated art and training data is unsettled in most jurisdictions, with ongoing litigation in the US and regulatory activity in the EU.
Memorization is a related but distinct concern. Research has demonstrated that diffusion models can memorize specific training images and reproduce them near-exactly when prompted appropriately, particularly for images that appeared multiple times in training data. This is problematic both for privacy (if training data contained personal images like faces or medical data) and copyright. Unlike language model memorization of text sequences, image memorization reproduces recognizable visual content that could have commercial or legal implications.
Evaluation difficulty confounds progress measurement. FID correlates poorly with human preference rankings in several studies: models with better FID scores are sometimes rated worse by human evaluators. CLIP Score does not capture spatial correctness, object count, or fine-grained attribute binding. Developing better automated evaluation metrics that strongly correlate with human judgment is an active research problem, and the lack of reliable metrics makes it difficult to claim that improvements on standard benchmarks translate to better user-facing quality.
Despite these limitations, diffusion models have enabled capabilities with broad practical impact. Professional-grade image generation is now accessible to non-artists. Personalization through fine-tuning on small image sets enables customized visual content. Creative augmentation for film and game production has changed pipeline workflows. Synthesis of training data for other machine learning systems addresses data scarcity. Medical imaging and scientific visualization benefit from diffusion-based super-resolution and synthesis. The combination of accessible open-source models (Stable Diffusion) and capable proprietary APIs (DALL-E 3, Midjourney, Firefly) has made generative image models a widely deployed technology in just three years.
Summary
This chapter covered the core ideas and practical methods of modern image generation:
- Diffusion models define a forward process that gradually corrupts images with Gaussian noise, then train a neural network to reverse this process step by step, converting pure noise into structured samples
- The closed-form forward process allows computing noisy images at any timestep directly: , making training parallelizable across timesteps
- The simplified training objective trains the network to predict added noise: , derived from a variational lower bound but simplified for stability
- Latent diffusion moves the process to a compressed VAE latent space, reducing compute by roughly an order of magnitude without sacrificing image quality
- Text-to-image generation uses cross-attention to inject text embeddings into the U-Net at every spatial scale, with classifier-free guidance amplifying the text-conditioned signal at inference to improve prompt adherence
- Sampling efficiency improved from 1000-step DDPM to 20-50-step DDIM and DPM-Solver++ by exploiting the ODE structure of the reverse process
- Image editing techniques including SDEdit, inpainting, InstructPix2Pix, DreamBooth, LoRA, Textual Inversion extend generation to controlled modification of existing images and subject personalization
- ControlNet and IP-Adapter provide spatial and visual conditioning beyond text, enabling precise compositional control
- Evaluation uses FID for distributional quality, CLIP Score for text alignment, and human evaluation for subjective quality; each metric captures different aspects and has known blind spots
- Key limitations include slow sampling, compositional failures, poor text rendering, copyright challenges, memorization, difficult evaluation
Diffusion models represent the current state of the art in image generation. Their core ideas, noise corruption and learned reversal, have already extended beyond static images into video generation (Sora, Runway Gen-3, Stable Video Diffusion), 3D shape synthesis, protein structure generation, audio synthesis, and molecular design. The mathematical framework of learning to denoise at multiple noise scales has proven to be a remarkably flexible and general principle for generative modeling.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about image generation and diffusion models.
Image Generation: Diffusion Models & Text-to-Image
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!