Image Generation: Diffusion Models, Text-to-Image, Editing

Michael BrenndoerferJanuary 5, 202655 min read

Part of Language AI Handbook

Covers diffusion models from the DDPM objective to latent diffusion, classifier-free guidance, image editing techniques, ControlNet, and FID evaluation metrics.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Image Generation: Diffusion Models and Text-to-Image Synthesis

For most of computing history, creating images required human artists. Generating a novel photorealistic scene, a painting in a specific style, or an illustration from a text description was something only people could do. That changed dramatically in the early 2020s, when diffusion models transformed image synthesis from a niche research curiosity into a technology capable of producing images indistinguishable from photographs and paintings. Today, you can type a sentence and receive a detailed, coherent image in seconds. Understanding how this works, at the level of math and mechanism, is the goal of this chapter.

The story of image generation is a story about probability. Generating an image means sampling from a distribution over pixels. The challenge is that this distribution is astronomically high-dimensional (a 512×512512 \times 512 RGB image has nearly 800,000 dimensions), highly structured, and deeply non-Gaussian. Getting samples that look like photographs, not random noise, requires a model that has learned the complex correlations between pixels at every scale, from fine texture to global composition.

This chapter covers the full arc of modern image generation. We start with the historical context, then examine the mathematical foundations of diffusion: how adding and removing noise creates a tractable generative process. We then examine how language is connected to image synthesis, enabling text-to-image generation. We cover image editing techniques that build on generation, and we discuss how to evaluate the quality of generated images. By the end, you will understand how these systems work, why diffusion became the dominant paradigm, and what its real limitations are.

From GANs to Diffusion: A Brief History

Before diffusion models, the dominant paradigm for image generation was the Generative Adversarial Network (GAN), introduced by Goodfellow et al. in 2014. GANs set up a game between two neural networks: a generator that produces fake images and a discriminator that tries to distinguish real from fake. The generator learns by fooling the discriminator; the discriminator improves by catching the generator. In equilibrium, the generator produces realistic images.

GANs were powerful but notoriously difficult to train. They suffered from mode collapse, a failure mode where the generator produces only a narrow variety of images despite being trained on diverse data. A generator that outputs a single convincing image of a golden retriever technically fools the discriminator, but it has not learned the full distribution. Training instability was equally problematic: the generator and discriminator could fall into unproductive cycles, with the generator oscillating between certain artifact patterns rather than converging. Researchers spent years developing techniques like progressive growing (training first at low resolution and gradually increasing), spectral normalization, gradient penalties, and careful architectural choices just to stabilize training.

Variational Autoencoders (VAEs) offered a different approach. A VAE encodes images into a compact latent distribution, parameterized by a mean and variance, then samples from that distribution to decode new images. VAEs were stable to train and provided an explicit latent space suitable for interpolation and manipulation. However, they produced notoriously blurry outputs. The reason is their reconstruction loss: optimizing pixel-level mean squared error encourages the decoder to average over plausible reconstructions, producing a blurry mean rather than any specific sharp image. The latent bottleneck adds to this averaging pressure by compressing multiple plausible outputs into a single latent point.

Autoregressive models like PixelCNN and VQ-VAE-based approaches (DALL-E v1, ImageGPT) took a third path: model the joint pixel distribution as a product of conditionals, generating one pixel (or patch, or token) at a time. These models can generate very high-quality images but are extremely slow at inference time because generation is sequential. VQ-VAE with a transformer prior (used in the original DALL-E) combined discrete tokenization of images with autoregressive modeling, allowing text-conditioned generation, but still required hundreds to thousands of forward passes.

Diffusion models emerged as an alternative that combined the stability of VAE training with quality approaching or exceeding GANs. Instead of a generator-discriminator game or an encoder-decoder compression, diffusion models learn to reverse a gradual noise process. This framing turns image generation into a sequence of denoising steps, each more tractable than generating an image from scratch. The main point is elegantly simple: if you can reliably remove a small amount of noise, you can chain those removals together to generate images from nothing.

Diffusion Models: The Core Idea

The fundamental insight of diffusion models is this: if you have a process that gradually destroys structure, you can learn its reverse to build structure back.

Consider what happens when you gradually add Gaussian noise to an image. At first, the image is recognizable but slightly grainy. After more noise is added, the details blur and only global structure remains. Continue adding noise, and eventually you have something indistinguishable from pure random static, regardless of what the original image was. This forward process destroys information in a principled, mathematically tractable way.

Now flip the perspective. If you could learn to undo each small noise step, you could start with pure random noise and gradually recover a coherent image. You do not need to generate a complete image in one shot; you only need to denoise slightly at each step. This is dramatically more tractable because each denoising step is a small, local change to a partially formed image, rather than a jump from noise to image.

The Forward Process: Adding Noise

The forward process (also called the diffusion process) takes a real image x0\mathbf{x}_0 and gradually adds Gaussian noise over TT timesteps:

q(xt∣xt−1)=N(xt; 1−βt xt−1, βtI)q(\mathbf{x}_t \mid \mathbf{x}_{t-1}) = \mathcal{N}\left(\mathbf{x}_t;\, \sqrt{1 - \beta_t}\, \mathbf{x}_{t-1},\, \beta_t \mathbf{I}\right)

where:

  • xt\mathbf{x}_t: the noisy image at timestep tt
  • xt−1\mathbf{x}_{t-1}: the image at the previous (less noisy) timestep
  • βt\beta_t: the noise schedule at timestep tt, controlling how much noise is added
  • N(μ,σ2)\mathcal{N}(\mu, \sigma^2): a Gaussian distribution with mean μ\mu and variance σ2I\sigma^2 \mathbf{I}

Each step scales down the signal by 1−βt\sqrt{1 - \beta_t} (to preserve approximate unit variance) and adds Gaussian noise with variance βt\beta_t. The scaling factor ensures that the overall variance of xt\mathbf{x}_t remains roughly constant over time: as signal energy decreases, noise energy increases by the same amount. With a carefully designed noise schedule, after TT steps (typically T=1000T = 1000), xT\mathbf{x}_T is approximately pure Gaussian noise, regardless of what x0\mathbf{x}_0 was.

A mathematical property makes diffusion computationally tractable: you can compute the noisy image at any arbitrary timestep tt directly from x0\mathbf{x}_0, without simulating each intermediate step. This is possible because Gaussian distributions have the property that the composition of Gaussians is itself Gaussian. Define αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^{t} \alpha_s (the cumulative product of all α\alpha values up to step tt). Then:

q(xt∣x0)=N(xt; αˉt x0, (1−αˉt)I)q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}\left(\mathbf{x}_t;\, \sqrt{\bar{\alpha}_t}\, \mathbf{x}_0,\, (1 - \bar{\alpha}_t) \mathbf{I}\right)

which means we can sample xt\mathbf{x}_t directly as:

xt=αˉt x0+1−αˉt ϵ,ϵ∼N(0,I)\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\, \mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\, \boldsymbol{\epsilon}, \quad \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})

where ϵ\boldsymbol{\epsilon} is pure Gaussian noise. This "reparameterization" has major practical implications. During training, you do not need to simulate the full forward chain from t=0t = 0 to t=Tt = T for every training example. Instead, for each training image you pick a random timestep tt, sample noise ϵ\boldsymbol{\epsilon}, compute the corresponding xt\mathbf{x}_t using the formula above, and train the model on that single noisy image. This makes training embarrassingly parallel across timesteps.

The Noise Schedule

The noise schedule {βt}t=1T\{\beta_t\}_{t=1}^{T} controls how quickly structure is destroyed as noise is added. Its design has a significant impact on the quality of the learned model.

A natural first choice is a linear schedule: βt\beta_t increases linearly from a small value (like 10−410^{-4}) to a larger value (like 0.020.02). This is easy to understand, but it has a practical flaw: it destroys too much structure too quickly at low noise levels, where most of the perceptual information lives, and too little at high noise levels. The model ends up spending many timesteps learning to denoise nearly-Gaussian noise, which carries little useful gradient signal for learning image structure.

The cosine schedule, proposed by Nichol and Dhariwal, addresses this by defining αˉt\bar{\alpha}_t directly rather than working through βt\beta_t:

αˉt=f(t)f(0),f(t)=cos⁡(t/T+s1+s⋅π2)2\bar{\alpha}_t = \frac{f(t)}{f(0)}, \quad f(t) = \cos\left(\frac{t/T + s}{1 + s} \cdot \frac{\pi}{2}\right)^2

where:

  • TT: the total number of timesteps
  • ss: a small offset (typically 0.008) that prevents βt\beta_t from being too large near t=0t = 0
  • The cosine function ensures αˉt\bar{\alpha}_t decreases smoothly from 1 to near 0 following a cosine curve

The cosine schedule keeps αˉt\bar{\alpha}_t near 1 for a longer initial period before dropping more smoothly to near zero at the end. This allocates more timesteps to informative intermediate noise levels, where the model can learn meaningful structure. Empirically, cosine schedules produce better image quality than linear ones, especially for higher-resolution images.

More recent work has explored learned noise schedules (optimizing {βt}\{\beta_t\} as parameters) and continuous-time diffusion (defining the noise process as a stochastic differential equation rather than discrete steps), but linear and cosine schedules remain widely used because of their simplicity and reliable performance.

The Reverse Process: Learning to Denoise

The reverse process aims to invert the forward process: starting from pure noise xT∼N(0,I)\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), progressively denoise to recover a clean image x0\mathbf{x}_0. The true reverse conditional distribution q(xt−1∣xt)q(\mathbf{x}_{t-1} \mid \mathbf{x}_t) exists but is intractable to compute directly, because it requires integrating over all possible clean images x0\mathbf{x}_0 consistent with the noisy image xt\mathbf{x}_t. However, if we condition on the clean image, the reverse conditional becomes tractable:

q(xt−1∣xt,x0)=N(xt−1; μ~t(xt,x0), β~tI)q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) = \mathcal{N}\left(\mathbf{x}_{t-1};\, \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0),\, \tilde{\beta}_t \mathbf{I}\right)

where β~t=(1−αˉt−1)1−αˉtβt\tilde{\beta}_t = \frac{(1 - \bar{\alpha}_{t-1})}{1 - \bar{\alpha}_t} \beta_t is the posterior variance. This formula gives the optimal denoising direction when the clean image x0\mathbf{x}_0 is known. In practice, we do not know x0\mathbf{x}_0 during generation, so we learn a neural network pθp_\theta to approximate the reverse conditional:

pθ(xt−1∣xt)=N(xt−1; μθ(xt,t), σt2I)p_\theta(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}\left(\mathbf{x}_{t-1};\, \boldsymbol{\mu}_\theta(\mathbf{x}_t, t),\, \sigma_t^2 \mathbf{I}\right)

where μθ\boldsymbol{\mu}_\theta is a neural network parameterized by θ\theta that predicts the mean of the reverse distribution.

Rather than directly predicting the mean, most implementations train the network to predict the noise ϵ\boldsymbol{\epsilon} that was added. This is the ϵ\boldsymbol{\epsilon}-prediction formulation. Given the predicted noise ϵθ(xt,t)\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t), the reverse step mean becomes:

μθ(xt,t)=1αt(xt−βt1−αˉtϵθ(xt,t))\boldsymbol{\mu}_\theta(\mathbf{x}_t, t) = \frac{1}{\sqrt{\alpha_t}}\left(\mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\right)

where:

  • 1αt\frac{1}{\sqrt{\alpha_t}}: un-scales the signal by reversing the scaling applied in the forward step
  • βt1−αˉtϵθ(xt,t)\frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t): subtracts the predicted noise contribution, pushing the estimate toward the clean image

The ϵ\boldsymbol{\epsilon}-prediction formulation is used in practice because it produces stable training dynamics: the loss is well-scaled across all timesteps, and the noise prediction task is a well-defined, bounded regression problem. An alternative is x0\mathbf{x}_0-prediction (predicting the clean image directly), which is sometimes more stable at low noise levels, and vv-prediction (predicting a linear combination of noise and signal), which was introduced for improved training at high resolutions. Modern models often blend these formulations.

The Training Objective

The full diffusion training objective is derived from a variational lower bound on the log-likelihood of the data. Maximizing log⁡pθ(x0)\log p_\theta(\mathbf{x}_0) directly is intractable, but we can instead maximize a lower bound that decomposes into per-timestep KL divergences between the learned reverse process and the true posterior. Ho et al. (2020) showed that a simplified version of this objective works just as well in practice:

Lsimple=Et,x0,ϵ[∥ϵ−ϵθ ⁣(αˉt x0+1−αˉt ϵ, t)∥2]\mathcal{L}_{\text{simple}} = \mathbb{E}_{t, \mathbf{x}_0, \boldsymbol{\epsilon}}\left[\left\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta\!\left(\sqrt{\bar{\alpha}_t}\, \mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\, \boldsymbol{\epsilon},\, t\right)\right\|^2\right]

where:

  • t∼Uniform(1,T)t \sim \text{Uniform}(1, T): a uniformly randomly sampled timestep
  • x0\mathbf{x}_0: a real training image sampled from the training dataset
  • ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}): noise sampled from a standard Gaussian
  • αˉt x0+1−αˉt ϵ\sqrt{\bar{\alpha}_t}\, \mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\, \boldsymbol{\epsilon}: the noisy image at timestep tt using the reparameterization trick
  • ϵθ\boldsymbol{\epsilon}_\theta: the neural network predicting the noise from the noisy image and timestep

In words: pick a random training image, pick a random timestep, add the corresponding amount of noise, and train the network to predict what noise was added. This is a denoising task framed differently at each timestep. At high tt, the image is mostly noise and the model must infer global structure. At low tt, the image is mostly clean and the model must predict the residual fine-grained noise.

The simplification from the full variational lower bound to Lsimple\mathcal{L}_{\text{simple}} drops per-timestep weighting that appears in the variational objective, trading theoretical optimality for empirical training stability. In practice, models trained with Lsimple\mathcal{L}_{\text{simple}} generate better images than those trained with the full variational bound, likely because the uniform weighting prevents some timesteps from dominating the gradient signal.

Denoising Score Matching

Diffusion models are closely related to score-based generative models. The "score" of a distribution is the gradient of the log probability with respect to the data: ∇xlog⁡p(x)\nabla_{\mathbf{x}} \log p(\mathbf{x}). If you know the score function everywhere, you can generate samples by starting from any point and following the gradient toward regions of higher probability, a procedure called Langevin dynamics. Song and Ermon (2019) showed how to estimate score functions at multiple noise levels through denoising score matching, and Song et al. (2021) unified this with diffusion models by showing that both processes are solutions to stochastic differential equations. Training a diffusion model to predict noise is mathematically equivalent to learning the score of a noisy data distribution at each noise level. This perspective has been productive for theoretical analysis and for developing improved samplers.

The U-Net Architecture for Noise Prediction

The neural network ϵθ\boldsymbol{\epsilon}_\theta needs to accept a noisy image and a timestep, and output a same-sized noise prediction. The dominant architecture is a U-Net, originally developed by Ronneberger et al. for biomedical image segmentation, adapted for the diffusion setting.

The U-Net's key structural property is its encoder-decoder architecture with skip connections. The encoder path progressively downsamples the image using strided convolutions or pooling, building a hierarchy of features from fine-grained local patterns to coarse global structure. The decoder path progressively upsamples back to the original resolution using transposed convolutions or bilinear upsampling. Skip connections copy feature maps directly from each encoder stage to the corresponding decoder stage, allowing fine-grained spatial details to flow directly to the decoder. Without skip connections, the encoder bottleneck would force all spatial information through the narrow representation, causing loss of fine detail.

For diffusion, the U-Net is adapted with several modifications:

  • Timestep embeddings: the scalar timestep tt is embedded via sinusoidal positional encodings (similar to those we discussed in Part XIV on positional encoding) and projected through small feed-forward networks. These embeddings are added to the feature maps at each resolution level, allowing the network to know at which noise level it is operating.
  • Residual blocks: each U-Net block wraps its core convolutions in residual connections, improving gradient flow during training and allowing deeper networks.
  • Self-attention layers: attention layers are added at lower resolutions where the spatial dimensions are small, making attention computationally feasible. These allow the network to model long-range dependencies between distant spatial locations. A global composition constraint, for example a sky being blue and grass being green, requires correlating distant pixels.
  • Group normalization: normalization is applied within feature groups rather than across batch dimensions, which is more stable for the small batch sizes used in diffusion training.

The result is a flexible model that can process images at their full resolution while conditioning on the noise level. Because the same network handles all TT timesteps, it learns a range of denoising behaviors: at high noise (early in generation), it focuses on global structure and low-frequency signals; at low noise (late in generation), it refines fine details and textures. The model implicitly learns a curriculum from coarse to fine.

Latent Diffusion Models

A key limitation of pixel-space diffusion is computational cost. Running a U-Net over full-resolution images for hundreds of timesteps is expensive. A 512×512512 \times 512 image has 786,432 pixels; processing this at each of 1000 timesteps during training requires enormous GPU memory and compute. Pixel-space diffusion models like the one used in DALL-E 2's decoder were trained on thousands of GPU-days.

Latent Diffusion Models (LDMs), introduced by Rombach et al. (2022) and forming the basis of Stable Diffusion, solve this by moving the diffusion process into a compressed latent space.

The approach uses a separate variational autoencoder (VAE) that:

  1. Encodes images to a compact latent representation: z=E(x)\mathbf{z} = \mathcal{E}(\mathbf{x}), where z\mathbf{z} has spatial dimensions roughly 8×8\times smaller than x\mathbf{x}
  2. Decodes latent representations back to images: x^=D(z)\hat{\mathbf{x}} = \mathcal{D}(\mathbf{z})

The VAE encoder-decoder is trained separately with a combination of losses. A reconstruction loss (typically L1 or L2 in pixel space) ensures the decoded image resembles the original. A perceptual loss (computing similarity in the feature space of a pre-trained VGG network rather than raw pixel space) ensures the reconstruction captures high-level semantic content, not just low-level pixel patterns. A GAN loss discriminates between real and decoded images, encouraging sharp textures and preventing blurry averaging. Additionally, a KL regularization term keeps the latent distribution close to a standard Gaussian, enabling sampling.

Once trained, the encoder-decoder is frozen. The diffusion model then operates exclusively on latent codes z\mathbf{z} rather than on pixels. Because latent codes are much smaller (typically 64×6464 \times 64 with 4 channels for a 512×512512 \times 512 image), the U-Net processes a 642=409664^2 = 4096 element feature map rather than a 5122=262144512^2 = 262144 pixel image. This is a 64x reduction in spatial resolution, reducing the computational cost of each forward pass by roughly two orders of magnitude.

The practical consequence is significant. Models like Stable Diffusion can be trained and run on consumer GPUs with 8-24 GB of VRAM, whereas pixel-space models of comparable quality required data center hardware. This democratization of image generation through latent compression was arguably as important as the diffusion algorithm itself.

The latent space also has a useful structural property: it is smoother and more semantically organized than pixel space. Nearby points in latent space correspond to semantically similar images, which means the diffusion model can learn a smoother, more coherent distribution compared to the raw pixel distribution that contains abrupt transitions between adjacent pixel values.

Text-to-Image Generation

Generating images from text descriptions requires connecting language to the visual generation process. This is done through conditioning: the noise-prediction network ϵθ\boldsymbol{\epsilon}_\theta receives the noisy image, the timestep, and a text representation that guides the generation.

The general principle is that the model learns a conditional distribution pθ(x∣c)p_\theta(\mathbf{x} \mid c) where cc is the text condition, rather than the unconditional pθ(x)p_\theta(\mathbf{x}). Every U-Net forward pass receives both the noisy image and the text embedding, allowing the noise prediction to be guided by the semantic content of the prompt.

Classifier-Free Guidance

The key technique for controlling text-conditioned generation is classifier-free guidance (CFG), introduced by Ho and Salimans. The motivating problem is that while a conditional diffusion model generates images corresponding to the text, the correspondence can be weak: the generated image might include elements from the prompt but not strongly emphasize the distinctive aspects that make the text description specific.

The idea is to train the model in two modes simultaneously: conditioned on text, and unconditioned (with the text randomly dropped out, typically 10-20% of the time by replacing the text conditioning with a null embedding). At inference, you run both modes and interpolate their noise predictions:

ϵ~θ(xt,c)=ϵθ(xt,∅)+w⋅(ϵθ(xt,c)−ϵθ(xt,∅))\tilde{\boldsymbol{\epsilon}}_\theta(\mathbf{x}_t, c) = \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing) + w \cdot \left(\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, c) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing)\right)

where:

  • cc: the text conditioning (encoded prompt)
  • ∅\varnothing: the null conditioning (empty or dropped text)
  • ww: the guidance scale, controlling how strongly the generation follows the text
  • ϵθ(xt,c)\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, c): noise prediction conditioned on text
  • ϵθ(xt,∅)\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing): unconditional noise prediction

When w=1w = 1, this reduces to the standard conditioned prediction. When w>1w > 1 (typical values are 7-12), the conditional direction is amplified beyond the raw conditional prediction: the model is pushed more strongly toward the text-conditioned manifold. When w=0w = 0, the result is fully unconditional.

Why does this work? The term ϵθ(xt,c)−ϵθ(xt,∅)\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, c) - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, \varnothing) captures how much the denoising direction changes when conditioned on text versus not. This difference is the direction in noise prediction space that corresponds to "move toward samples that fit the text condition." Scaling this up with large ww amplifies the text-driven signal, pushing the generation more aggressively toward images that match the prompt.

The tradeoff is that higher guidance scales reduce diversity. With w=12w = 12, the model produces images that very closely match the prompt but all tend to look similar; with w=3w = 3, there is more variation at the cost of sometimes weaker prompt adherence. The guidance scale is one of the most important inference-time hyperparameters.

Classifier Guidance (the original approach)

Classifier-free guidance was preceded by classifier guidance, introduced by Dhariwal and Nichol (2021). Their approach used a separately trained image classifier to guide the diffusion process. At each denoising step, the classifier computes ∇xtlog⁡p(c∣xt)\nabla_{\mathbf{x}_t} \log p(c \mid \mathbf{x}_t), the gradient of the log probability that the noisy image belongs to class cc. This gradient is added to the noise prediction, pushing generation toward the target class. Classifier guidance demonstrated that guided diffusion could surpass GANs on image quality metrics. Classifier-free guidance achieves similar or better effects without requiring a separate classifier, making it applicable to any conditioning modality: text, images, poses, style references. It is not limited to predefined class labels.

Text Encoding

The text conditioning cc is typically derived from a pre-trained language model that maps the prompt to a sequence of dense embeddings.

CLIP text encoder: Used in DALL-E 2 and Stable Diffusion 1.x, the CLIP text encoder was trained to align with visual representations through contrastive learning on image-text pairs, as we discussed in the CLIP chapter (Part L). This makes CLIP text embeddings naturally suited for conditioning image generation: they encode language in a space that is semantically linked to visual concepts. The embedding captures semantic meaning in a vision-aware way, which helps the diffusion model bridge language and image.

Large language model encoders: Imagen (Google, 2022) used T5-XXL as its text encoder, a 4.6B parameter language model not specifically trained for visual alignment. Despite lacking CLIP's vision-language grounding, T5 embeddings proved highly effective because T5 handles linguistic structure and negation alongside composition and rare concepts better than CLIP's text encoder. The T5 embeddings carry richer linguistic context, which improves generation of complex, compositional prompts.

Multiple encoder combination: Stable Diffusion XL (SDXL, 2023) and Stable Diffusion 3 (SD3, 2024) use multiple text encoders simultaneously, combining OpenCLIP's ViT-bigG text encoder with the original CLIP ViT-L encoder (in SDXL) or with T5-XXL (in SD3). The outputs are concatenated or pooled together before being fed to the diffusion backbone. This combination captures both vision-aligned CLIP semantics and linguistically rich T5 representations. This provides broader coverage of both visual and textual concepts.

The text is tokenized and passed through the text encoder to produce a sequence of embeddings, one per token. These embeddings are injected into the U-Net through cross-attention layers: the image features at each spatial location act as queries, and the text embeddings serve as keys and values. Each spatial location in the image can thus attend to the text tokens most relevant to what should appear at that location, creating fine-grained text-image alignment. A text token for "red" might be strongly attended to by the image region where the red object should appear.

DALL-E 2 and the CLIP Prior

DALL-E 2 (OpenAI, 2022) took a distinctive two-stage approach to text-to-image generation, motivated by the observation that CLIP image embeddings capture semantic visual content in a rich, structured way.

The pipeline works as follows. First, a prior network generates a CLIP image embedding from the text prompt. This prior is itself a diffusion model (or autoregressive model) that learns the mapping from text embeddings to image embeddings in CLIP's joint space. Given a text prompt describing a scene, the prior generates a CLIP image embedding that represents the visual concept, without yet specifying pixel-level details.

Second, a decoder (a pixel-space diffusion model) generates an image from that CLIP image embedding. Because CLIP image embeddings encode semantic visual content rather than specific pixels, the decoder can focus on rendering visual content consistently with the embedding rather than interpreting language from scratch.

The prior bridges the gap between language and vision space. CLIP's joint embedding space aligns text and images semantically, but a single text prompt corresponds to many possible images. The prior models this one-to-many mapping, sampling different visual interpretations that are all consistent with the text. The result showed strong semantic faithfulness: images clearly corresponded to the described content, even for novel combinations.

Imagen and the Value of Text Understanding

Imagen (Google, 2022) demonstrated a surprising lesson: text understanding quality matters as much as image modeling quality for text-to-image generation.

Imagen uses a cascade of diffusion models. A base model generates 64×6464 \times 64 images conditioned on T5-XXL text embeddings. Two super-resolution diffusion models then upscale the image to 256×256256 \times 256 and then 1024×10241024 \times 1024, each conditioned on both the original text and the lower-resolution image. The final output is a high-resolution image generated through this pipeline.

The key finding was that doubling the size of the text encoder (using T5-XXL at 4.6B parameters) improved image quality and text adherence more than doubling the size of the image diffusion model. This implies that the bottleneck in text-to-image generation is often the quality of text understanding, not image rendering. A model that does not understand "the woman standing to the left of the red building" will produce a plausible-looking image that fails to respect the spatial and attribute relationships in the text.

Imagen also introduced dynamic thresholding for classifier-free guidance. At high guidance scales, the predicted clean image values can fall outside the valid range [−1,1][-1, 1], producing oversaturated images. Dynamic thresholding scales the entire prediction by the maximum absolute value if that exceeds a threshold, preserving relative magnitudes while keeping the image in range. This allows using higher guidance scales without oversaturation.

Stable Diffusion and Latent Diffusion

Stable Diffusion (CompVis/Stability AI, 2022) combined latent diffusion with CLIP text conditioning to create the first major open-source text-to-image model. Its architecture became the standard for the open-source ecosystem due to its balance of computational efficiency and image quality.

The key components are:

  • A VAE encoder/decoder that compresses 512×512512 \times 512 images to 64×64×464 \times 64 \times 4 latent codes (spatial compression factor 8, with 4 latent channels)
  • A CLIP ViT-L/14 text encoder (frozen) that produces 77 token embeddings of dimension 768
  • A latent U-Net that performs diffusion in the compressed 64×6464 \times 64 latent space, with cross-attention layers throughout for text conditioning
  • A linear noise schedule with T=1000T = 1000 training steps

The open release of model weights enabled a large ecosystem: fine-tuning on specific styles and subjects (DreamBooth, Textual Inversion, LoRA), specialized control adapters (ControlNet, IP-Adapter), GUI tools (AUTOMATIC1111, ComfyUI), and commercial applications. SDXL (2023) scaled this architecture with a larger U-Net and dual CLIP encoders, achieving substantially better image quality and prompt adherence.

Stable Diffusion 3 (2024) moved away from the U-Net architecture toward a Diffusion Transformer (DiT) backbone, using multimodal attention layers that jointly process both image patches and text tokens in a single transformer block, rather than separate processing with cross-attention injection.

Sampling and Inference

Training a diffusion model gives us the denoising network, but generating images requires running the reverse process. This is the sampling procedure, and the design of samplers has been an active area of research because the original DDPM sampler was far too slow for practical use.

DDPM Sampling

The original Denoising Diffusion Probabilistic Models (DDPM) sampler runs the full TT reverse steps:

xt−1=1αt(xt−1−αt1−αˉtϵθ(xt,t))+σtz\mathbf{x}_{t-1} = \frac{1}{\sqrt{\alpha_t}}\left(\mathbf{x}_t - \frac{1 - \alpha_t}{\sqrt{1 - \bar{\alpha}_t}} \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\right) + \sigma_t \mathbf{z}

where:

  • 1αt\frac{1}{\sqrt{\alpha_t}}: the inverse scaling factor for this step
  • 1−αt1−αˉtϵθ\frac{1 - \alpha_t}{\sqrt{1 - \bar{\alpha}_t}} \boldsymbol{\epsilon}_\theta: the denoising correction based on the predicted noise
  • σtz\sigma_t \mathbf{z}: additional stochastic noise injected at each step, where z∼N(0,I)\mathbf{z} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})

With T=1000T = 1000, generating one image requires 1000 neural network forward passes through the full U-Net. At 8-30 milliseconds per forward pass on a modern GPU, this translates to 8-30 seconds per image, far too slow for interactive use.

DDIM Sampling

Denoising Diffusion Implicit Models (DDIM), introduced by Song et al. (2021), made sampling dramatically faster by reformulating the reverse process as a deterministic ordinary differential equation (ODE) rather than a stochastic process.

The key insight is that many different reverse processes can share the same marginals q(xt∣x0)q(\mathbf{x}_t \mid \mathbf{x}_0) as DDPM while taking larger, more efficient steps. DDIM defines a non-Markovian reverse process that allows skipping many timesteps. Instead of going through all 1000 timesteps in order, DDIM can take evenly spaced steps visiting only 20 timesteps out of 1000, while still producing high-quality images.

The deterministic nature of DDIM has an additional useful property: it acts as an encoder as well as a decoder. You can run the DDIM forward process on a real image to find the latent noise representation corresponding to that image (called DDIM inversion), then modify the generation process at that latent to edit the image. This is not possible with stochastic DDPM because the stochasticity makes inversion non-invertible.

DPM-Solver and Modern Samplers

Subsequent work formalized diffusion sampling as solving a differential equation and applied numerical ODE analysis to improve it. DPM-Solver and its successor DPM-Solver++ (Lu et al., 2022) treat the reverse process as an ODE with a known analytical structure and apply higher-order solvers. These can produce high-quality images in as few as 10-20 steps by taking larger, more accurate integration steps.

The key insight is that the diffusion reverse process has a semi-linear structure: the linear component (corresponding to Gaussian noise removal) can be solved analytically, and only the nonlinear component (the U-Net's neural denoising) needs to be approximated numerically. By exploiting this structure, DPM-Solver takes each integration step far more accurately than DDIM, which uses a simpler first-order approximation.

Other popular samplers include PNDM (Pseudo Numerical Methods for Diffusion Models), Euler and Euler a (ancestral Euler methods with optional stochasticity), and Heun (a second-order Runge-Kutta variant). The choice of sampler affects speed and the character of the output: deterministic samplers give reproducible results from a fixed seed, while stochastic samplers add variation and can sometimes escape local optima in the generation trajectory.

Image Editing

Beyond generating new images from scratch, diffusion models can modify existing images in semantically meaningful ways. This is possible because the forward and reverse processes can be applied to noise and to any image.

SDEdit

SDEdit (Meng et al., 2021) is the simplest editing approach. You take a reference image, add noise to a specific level t∗t^*, then run the reverse process with conditioning. This approach does not require any training changes; it uses a pre-trained diffusion model directly.

The noise level t∗t^* controls the tradeoff between fidelity to the original and freedom to change. High t∗t^* means more noise is added, so the model can make larger structural changes but may deviate substantially from the original content. Low t∗t^* preserves the original closely but allows only subtle texture or color edits. This makes t∗t^* an intuitive dial: setting it at t∗=0.5Tt^* = 0.5T keeps the rough composition and changes details, while t∗=0.8Tt^* = 0.8T allows more substantial changes to shape and content.

SDEdit works well for sketch-based generation (start with a rough hand-drawn sketch, add noise, condition on text), style transfer (starting from a content image), and small targeted edits. Its simplicity is both its strength (no additional training needed) and its limitation (it cannot make precise, localized edits).

Inpainting

Inpainting fills a masked region of an image while maintaining consistency with the surrounding context. This is useful for removing unwanted objects, filling in damaged areas, or adding new content to existing images.

During the reverse process, at each timestep, the unmasked regions of the noisy image are replaced with the known noisy image at that level:

xtinpaint=m⊙xtknown+(1−m)⊙xtgenerated\mathbf{x}_t^{\text{inpaint}} = m \odot \mathbf{x}_t^{\text{known}} + (1 - m) \odot \mathbf{x}_t^{\text{generated}}

where:

  • mm: the binary mask (1 for pixels to keep, 0 for pixels to generate)
  • xtknown\mathbf{x}_t^{\text{known}}: the original image noised to level tt using the forward process
  • xtgenerated\mathbf{x}_t^{\text{generated}}: the image generated so far by the reverse process
  • ⊙\odot: element-wise multiplication

At each reverse step, the generated regions are denoised by the network while the known regions are reset to their noisy versions (forward-noised at exactly timestep tt). This ensures that in the final output, the known regions are exactly the original pixels and the generated region is produced by the diffusion model under the constraint of fitting smoothly with the surroundings.

A limitation of this naive approach is that the seam between generated and preserved regions can be sharp, because the generated region is optimized without full awareness of the boundary constraint. Improved inpainting models are fine-tuned with masked training examples, teaching the model to explicitly condition on both the unmasked image content and the generation task at the boundary.

InstructPix2Pix

Rather than conditioning on a masked region, InstructPix2Pix (Brooks et al., 2022) conditions on both an original image and a text instruction describing the desired edit, such as "make the sky sunset" or "add glasses to the person."

Training this model required a creative data generation pipeline, since pairs of (original image, edit instruction, edited image) are not naturally available at scale. The solution was to synthesize training data: GPT-3 generated pairs of (image caption, edit instruction, edited image caption), and a text-to-image model (DALL-E) generated image pairs from the caption pairs. The diffusion model was then trained to perform the transformation given both the original image and the instruction.

At inference, InstructPix2Pix uses a modified CFG that combines guidance from both the text instruction and the original image, allowing control over how much the edit changes the image versus how literally the instruction is followed.

DreamBooth and LoRA Fine-Tuning

A related class of techniques personalizes image generation to specific subjects. DreamBooth (Ruiz et al., 2022) fine-tunes the entire diffusion model on 3-20 images of a specific person, object, or style. The appearance is bound to a rare token, like "sks dog," that is unlikely to appear in the original training data. After fine-tuning, prompts containing that token generate the specific subject in novel contexts: "a photo of sks dog in the snow" produces the specific dog in a snowy scene.

The challenge with full fine-tuning is the risk of language drift: the model may overfit to the specific images and forget other concepts. DreamBooth addresses this with a prior preservation loss that continues training on text-generated images of the general class (regular dogs, in this example) alongside the specific subject images.

Low-Rank Adaptation (LoRA), which we covered in Part XXXV on parameter-efficient fine-tuning, addresses both the cost and drift concerns by fine-tuning only small low-rank matrices inserted into the attention layers. Instead of updating all weight matrices, LoRA decomposes the update as ΔW=AB\Delta W = AB where A∈Rd×rA \in \mathbb{R}^{d \times r} and B∈Rr×dB \in \mathbb{R}^{r \times d} with small rank rr (typically 4-16). LoRA files for Stable Diffusion are typically 50-200 MB rather than the several GB of full model weights, and they can be combined by summing their contributions with different weights, allowing compositional style mixing.

Textual Inversion

Textual Inversion (Gal et al., 2022) takes an even lighter-weight approach: rather than changing the model weights at all, it learns a new token embedding that represents the target concept. The denoising network is entirely frozen; only a new token embedding v∗v^* is optimized to minimize the denoising loss on the reference images:

L=Et,x0,ϵ[∥ϵ−ϵθ ⁣(xt, c(S∗))∥2]\mathcal{L} = \mathbb{E}_{t, \mathbf{x}_0, \boldsymbol{\epsilon}}\left[\left\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta\!\left(\mathbf{x}_t,\, c(S^*)\right)\right\|^2\right]

where c(S∗)c(S^*) is the text conditioning derived from a string S∗S^* containing the new learned token v∗v^*. The learned token lives in the same embedding space as the model's existing vocabulary, so it can be combined with other words naturally.

Textual Inversion requires even fewer parameters than LoRA (a single token embedding vector rather than low-rank weight matrices) but captures concepts less flexibly because it is limited to the expressiveness of the fixed embedding space. Complex concepts that require weight adaptations to represent well are better suited to LoRA or DreamBooth.

Generation Control and Conditioning

Beyond text, diffusion models can be conditioned on many other signals to give precise control over the output. This is important for professional workflows where the artist needs more control than text alone provides.

ControlNet

ControlNet (Zhang and Agrawala, 2023) adds spatial conditioning signals such as edge maps, pose estimates (OpenPose skeleton), depth maps, surface normal maps, segmentation masks, or text line positions. The architecture addresses a key challenge: how to add conditioning without retraining the entire model.

ControlNet creates a trainable copy of the U-Net encoder. This copy processes the conditioning input (an edge map, for example) and adds its activations to the corresponding layers of the original U-Net's skip connections. The original U-Net is frozen; only the encoder copy and its connection weights are trained.

This design preserves the original U-Net's learned image priors while adding spatial control. Because the conditioning signal is spatial (the same height and width as the image), the network can precisely control image composition.

OpenPose conditioning generates humans in specific poses by feeding a skeleton image as the control signal. Canny edge map conditioning generates images that respect precise boundaries. Depth conditioning preserves 3D structure from a reference image while changing textures or style. This is particularly valuable for architectural rendering (conditioning on floor plans or line drawings), character animation (conditioning on pose sequences from video), and product photography (conditioning on the shape of a product).

Multiple ControlNet models can be stacked, applying combined conditioning. A generation might use edge conditioning for structural fidelity and pose conditioning for character positioning simultaneously.

IP-Adapter

IP-Adapter (Ye et al., 2023) conditions the diffusion model on a reference image's visual style or content, rather than or in addition to a text prompt. You provide an image whose style, subject, or composition you want to replicate in new generations.

IP-Adapter encodes the reference image with a vision encoder (typically CLIP's visual encoder) and injects those embeddings through a separate cross-attention pathway running in parallel to the existing text cross-attention. The image features serve as additional keys and values alongside the text features, with learnable weights determining the balance between text and image conditioning.

This allows tasks that are difficult to specify through text alone: style transfer (generate this scene in the visual style of this reference painting), face consistency (generate different scenarios featuring a specific person's face), and product placement (generate marketing images featuring a specific product). Because IP-Adapter is an adapter added on top of the frozen base model, it is compatible with any ControlNet adapters and LoRA fine-tuned models for the same base.

Evaluation of Generated Images

Evaluating image generation quality is difficult. Human perception is the gold standard, but human evaluation is expensive and slow, with limited reproducibility. Several automated metrics have been developed, each capturing different aspects of quality.

Frechet Inception Distance (FID)

FID is the most widely used automated metric for measuring image generation quality. It measures the distance between the distribution of generated images and the distribution of real images, not in pixel space, but in the feature space of a pre-trained InceptionV3 network.

The idea is that features from a network trained on ImageNet capture semantically meaningful aspects of image content: textures and shapes alongside structures and high-level objects. If the distribution of these features in the generated images matches the distribution in the real images, the generated images capture the same semantic variety and realism as the real data.

Given real images Xr\mathbf{X}_r and generated images Xg\mathbf{X}_g, their Inception features are extracted and fit to multivariate Gaussians with means μr\boldsymbol{\mu}_r, μg\boldsymbol{\mu}_g and covariances Σr\boldsymbol{\Sigma}_r, Σg\boldsymbol{\Sigma}_g. FID computes the Frechet distance between these Gaussians:

FID=∥μr−μg∥2+Tr(Σr+Σg−2(ΣrΣg)1/2)\text{FID} = \|\boldsymbol{\mu}_r - \boldsymbol{\mu}_g\|^2 + \text{Tr}\left(\boldsymbol{\Sigma}_r + \boldsymbol{\Sigma}_g - 2(\boldsymbol{\Sigma}_r \boldsymbol{\Sigma}_g)^{1/2}\right)

where:

  • ∥μr−μg∥2\|\boldsymbol{\mu}_r - \boldsymbol{\mu}_g\|^2: squared distance between mean feature vectors, measuring distributional shift in average content
  • Tr(⋅)\text{Tr}(\cdot): trace of a matrix (sum of diagonal elements)
  • (ΣrΣg)1/2(\boldsymbol{\Sigma}_r \boldsymbol{\Sigma}_g)^{1/2}: matrix square root of the product of covariances, measuring structural similarity

Lower FID is better: a perfect model matching the real data distribution exactly would achieve FID of 0. State-of-the-art models achieve FID scores below 3 on standard benchmarks like MS-COCO or ImageNet.

FID has known limitations. It is sensitive to the number of samples used: you need at least 10,000 generated images for reliable estimation, and the value changes substantially with fewer samples. It depends on the specific feature extractor (InceptionV3 may not capture the aspects of quality most relevant to human preference). It measures the full distribution, so a model that generates many diverse but lower-quality images can outscore one that generates fewer but higher-quality images on FID. Despite these limitations, it remains the standard comparison metric in published research.

CLIP Score

For text-to-image evaluation, CLIP Score measures the semantic alignment between a generated image and its text prompt, directly using the CLIP model's joint embedding space.

CLIP Score computes the cosine similarity between the CLIP image embedding of the generated image and the CLIP text embedding of the prompt:

CLIP Score=w⋅max⁡(0,cos⁡(eimage,etext))\text{CLIP Score} = w \cdot \max(0, \cos(\mathbf{e}_{\text{image}}, \mathbf{e}_{\text{text}}))

where:

  • eimage\mathbf{e}_{\text{image}}: the normalized CLIP image embedding of the generated image
  • etext\mathbf{e}_{\text{text}}: the normalized CLIP text embedding of the prompt
  • ww: a scaling factor (typically 100) that maps cosine similarities into a human-interpretable score range
  • max⁡(0,⋅)\max(0, \cdot): clips negative similarities to zero, since negative cosine similarity indicates opposite directions in the embedding space

A score of 25-30 is typical for high-quality text-image aligned models. CLIP Score captures semantic fidelity well: does the image visually depict the concepts mentioned in the text? But it does not capture visual quality or photorealism, and it has blind spots wherever CLIP's training data was sparse. It is also susceptible to adversarial inputs: images that score highly on CLIP but look strange to human observers. It is most useful for comparing text-image alignment across models or studying how different prompting strategies affect faithfulness, not as a standalone quality measure.

Human Evaluation

Automated metrics often disagree with human judgment about what makes an image "good." Human evaluation remains the gold standard for capturing subjective quality dimensions that metrics miss.

Human evaluation typically asks annotators to:

  • Compare two images generated from the same prompt and indicate which is better (pairwise preference comparison)
  • Rate an image on overall quality and photorealism, text-image alignment, diversity, aesthetic appeal
  • Evaluate specific failure modes: distorted hands, unnatural textures, implausible lighting, artifacts in background regions, or incorrect object counts

Pairwise comparisons are often more reliable than absolute ratings because they avoid systematic biases in how different annotators interpret rating scales. The Bradley-Terry model can convert pairwise preferences into global scores. Modern leaderboards like HPSv2 and ELO-based systems collect large numbers of human pairwise comparisons and use them to rank models.

Human evaluation is expensive to scale to thousands of prompts and model comparisons, but it remains the most reliable signal for the aspects of quality that matter most in practice: does the image look good to a human? Would a person be satisfied receiving this image in response to their prompt?

Precision and Recall

Precision and recall for generative models, as defined by Kynkaanniemi et al. (2019), measure two distinct aspects of generation quality:

  • Precision: Do generated images look realistic? Are they within the manifold of real images? High precision means the model rarely generates implausible images that fall outside what real images look like.
  • Recall: Does the model cover the full diversity of the real distribution? High recall means the model can generate images across the full range of visual variety present in the training data.

A model that generates only highly photorealistic images of a narrow slice of the distribution would have high precision but low recall: it generates quality images but misses most of what real images look like. A model that generates highly varied images that are sometimes blurry or contain artifacts would have high recall but low precision.

These metrics are computed using nearest-neighbor distances in the InceptionV3 feature space. For precision, we ask: for each generated image, does a real image exist nearby? For recall: for each real image, does a generated image exist nearby?

Precision and recall provide more detail than FID alone. FID is a single scalar that conflates fidelity and diversity; precision and recall decompose these into separate measurements, allowing diagnosis of whether a model's failure mode is generating unrealistic images (low precision) or missing diversity (low recall). Increasing guidance scale, for example, typically increases precision (higher quality) at the cost of recall (less diversity).

Implementation: Building a Mini Diffusion Pipeline

Let's implement a simplified diffusion training and sampling loop to make these concepts concrete. We will not train a full-scale text-to-image model (that requires significant compute and large datasets), but we will build a complete DDPM pipeline that learns to generate a structured 2D distribution. This allows us to visualize the full forward-reverse process and verify that the math works as described.

Setup and Data

We will work with a toy 2D dataset: a two-moons arrangement. This lets us visualize the full forward-reverse process clearly. Despite its simplicity, this dataset is non-trivial for generative models because it has a bimodal, non-Gaussian shape with a curved manifold.

In[4]:
Code
import numpy as np
import torch


def make_moons_dataset(n_samples=2000, noise=0.05, seed=42):
    """Create a simple two-moons 2D dataset."""
    rng = np.random.RandomState(seed)
    n_samples_each = n_samples // 2

    theta1 = np.linspace(0, np.pi, n_samples_each)
    theta2 = np.linspace(np.pi, 2 * np.pi, n_samples_each)

    x1 = np.column_stack([np.cos(theta1), np.sin(theta1)])
    x2 = np.column_stack([np.cos(theta2) + 1, np.sin(theta2)])

    x = np.vstack([x1, x2]) + rng.randn(n_samples, 2) * noise
    return x.astype(np.float32)


data = make_moons_dataset(n_samples=3000)
X = torch.tensor(data)
n_samples, d = X.shape
Out[5]:
Console
Dataset shape: torch.Size([3000, 2])
Data range: [-1.136, 2.134]
Data mean: [4.9965525e-01 3.2643635e-05]

Defining the Noise Schedule

The DiffusionSchedule class precomputes all the quantities derived from the noise schedule that we need for both training (forward process) and sampling (reverse process).

In[6]:
Code
class DiffusionSchedule:
    """Linear noise schedule for DDPM."""

    def __init__(self, T=200, beta_start=1e-4, beta_end=0.02, device="cpu"):
        self.T = T
        self.device = device

        # Linear noise schedule
        self.betas = torch.linspace(beta_start, beta_end, T, device=device)
        self.alphas = 1.0 - self.betas
        self.alphas_cumprod = torch.cumprod(self.alphas, dim=0)

        # Precompute useful quantities
        self.sqrt_alphas_cumprod = torch.sqrt(self.alphas_cumprod)
        self.sqrt_one_minus_alphas_cumprod = torch.sqrt(
            1.0 - self.alphas_cumprod
        )

    def q_sample(self, x0, t, noise=None):
        """Add noise to x0 at timestep t. Implements the forward process."""
        if noise is None:
            noise = torch.randn_like(x0)

        sqrt_alpha = self.sqrt_alphas_cumprod[t].view(-1, 1)
        sqrt_one_minus_alpha = self.sqrt_one_minus_alphas_cumprod[t].view(-1, 1)

        return sqrt_alpha * x0 + sqrt_one_minus_alpha * noise, noise


schedule = DiffusionSchedule(T=200)
Out[7]:
Console
Noise schedule: T=200
Alpha(t=0) = 0.9999 (almost no noise)
Alpha(t=100) = 0.5964 (partial noise)
Alpha(t=199) = 0.132183 (almost pure noise)

At t=0t = 0, αˉt≈1\bar{\alpha}_t \approx 1 so the noisy image is nearly identical to the clean data. At t=199t = 199, αˉt≈0\bar{\alpha}_t \approx 0 so the signal has been destroyed and we have nearly pure Gaussian noise. The middle timesteps correspond to partial noise: the data has some structure but is significantly corrupted.

Denoising Network

For our 2D toy problem, a small MLP handles noise prediction. In full image diffusion, this would be a U-Net with millions of parameters. The architecture here captures the essential design: a network that takes both the noisy data and the timestep, and outputs a noise prediction of the same shape as the input.

In[8]:
Code
import torch.nn as nn


class SinusoidalPositionEmbedding(nn.Module):
    """Embed timestep t as a sinusoidal vector."""

    def __init__(self, dim):
        super().__init__()
        self.dim = dim

    def forward(self, t):
        half_dim = self.dim // 2
        emb = torch.log(torch.tensor(10000.0)) / (half_dim - 1)
        emb = torch.exp(torch.arange(half_dim, device=t.device) * -emb)
        emb = t[:, None].float() * emb[None, :]
        return torch.cat([torch.sin(emb), torch.cos(emb)], dim=-1)


class DenoisingMLP(nn.Module):
    """MLP that predicts noise from noisy data and timestep."""

    def __init__(self, data_dim=2, hidden_dim=256, time_dim=64):
        super().__init__()
        self.time_embed = SinusoidalPositionEmbedding(time_dim)

        self.net = nn.Sequential(
            nn.Linear(data_dim + time_dim, hidden_dim),
            nn.GELU(),
            nn.Linear(hidden_dim, hidden_dim),
            nn.GELU(),
            nn.Linear(hidden_dim, hidden_dim),
            nn.GELU(),
            nn.Linear(hidden_dim, data_dim),
        )

    def forward(self, x, t):
        t_emb = self.time_embed(t)
        return self.net(torch.cat([x, t_emb], dim=-1))


model = DenoisingMLP(data_dim=2, hidden_dim=256, time_dim=64)

The sinusoidal timestep embedding encodes the scalar tt as a vector of alternating sines and cosines at different frequencies. This is the same approach used in transformer positional embeddings, and it ensures the model can distinguish all timesteps and interpolate smoothly between them. Without this embedding, the network would not know whether it is being asked to denoise heavily corrupted data (where only global structure matters) or lightly corrupted data (where fine details matter).

Out[9]:
Console
Model parameters: 149,250
Input shape: torch.Size([8, 2]), output shape: torch.Size([8, 2])

Training Loop

In[10]:
Code
import torch.nn.functional as F
from torch.optim import Adam


def train_diffusion(model, schedule, X, epochs=3000, batch_size=256, lr=1e-3):
    """Train the denoising network using DDPM objective."""
    optimizer = Adam(model.parameters(), lr=lr)
    losses = []

    for epoch in range(epochs):
        # Sample random batch
        idx = torch.randint(0, len(X), (batch_size,))
        x0 = X[idx]

        # Sample random timesteps
        t = torch.randint(0, schedule.T, (batch_size,))

        # Add noise according to schedule
        xt, noise = schedule.q_sample(x0, t)

        # Predict noise
        noise_pred = model(xt, t)

        # Loss: MSE between predicted and actual noise
        loss = F.mse_loss(noise_pred, noise)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        if epoch % 200 == 0:
            losses.append(loss.item())

    return losses


losses = train_diffusion(model, schedule, X, epochs=3000, batch_size=256)
Out[11]:
Console
Training epochs: 3000
Final loss: 0.5023
Initial loss: 1.1503
Loss reduction: 56.3%

The loss measures the mean squared error between the predicted noise and the actual noise added. As training progresses, the model learns to predict noise more accurately across all timesteps, and the loss decreases. A perfect model would achieve loss of 0, but in practice the irreducible loss reflects the stochasticity of the forward process.

DDPM Sampling

With a trained model, we run the reverse process to generate new samples. The sampling loop iterates from t=T−1t = T-1 down to t=0t = 0, denoising at each step.

In[12]:
Code
@torch.no_grad()
def ddpm_sample(model, schedule, n_samples=1000, d=2):
    """Sample from the learned distribution using DDPM reverse process."""
    model.eval()

    # Start from pure noise
    x = torch.randn(n_samples, d)
    trajectory = [x.clone()]

    for t in reversed(range(schedule.T)):
        t_batch = torch.full((n_samples,), t, dtype=torch.long)

        # Predict noise
        eps_pred = model(x, t_batch)

        # Compute denoised estimate
        beta_t = schedule.betas[t]
        alpha_t = schedule.alphas[t]
        alpha_bar_t = schedule.alphas_cumprod[t]

        # DDPM reverse step mean
        coef = beta_t / torch.sqrt(1 - alpha_bar_t)
        mu = (x - coef * eps_pred) / torch.sqrt(alpha_t)

        if t > 0:
            # Add stochastic noise for all but the last step
            noise = torch.randn_like(x)
            x = mu + torch.sqrt(beta_t) * noise
        else:
            x = mu

        if t % 50 == 0:
            trajectory.append(x.clone())

    return x, trajectory


generated, trajectory = ddpm_sample(model, schedule, n_samples=2000)
Out[13]:
Console
Generated 2000 samples in 2D
Generated range: [-1.283, 2.232]
Generated mean: [ 0.5127893  -0.04144546]

Notice that we add stochastic noise at every reverse step except the last (t=0t = 0). This stochasticity is essential to DDPM sampling: without it, the reverse process would be deterministic from the initial noise, producing less diverse samples and sometimes getting stuck in local optima. At t=0t = 0, we take the deterministic mean to obtain the final clean sample.

Visualizing the Noise Schedule

Out[14]:
Visualization
Line plot showing signal retention decreasing and noise fraction increasing over 200 timesteps, with equal contribution at timestep 116.
Linear noise schedule for T=200 timesteps showing cumulative signal retention ($\bar{\alpha}_t$, blue) decreasing from 1.0 to about 0.13 and noise fraction ($1 - \bar{\alpha}_t$, teal) rising to about 0.87. The crossover at t=116 marks where signal and noise contribute equally. Early timesteps retain most signal; later timesteps are noise-dominated without yet reaching pure Gaussian noise.

Visualizing Generated vs. Real Data

Out[15]:
Visualization
Scatter plot of real two-moons training data in blue showing two crescent arcs
Training data for the two-moons dataset, showing the bimodal crescent structure the diffusion model must learn. The two arcs are close but distinct, requiring the model to learn a non-convex, multimodal distribution that a Gaussian model could not capture.
Scatter plot of diffusion-generated samples in orange showing recovered two crescent arcs
Samples generated by the trained diffusion model via the 200-step DDPM reverse process. The model recovers both crescent arcs from pure Gaussian noise, while a small number of bridge and off-manifold samples reveal residual approximation error.

The generated distribution captures both crescent arcs, confirming that the model has learned the broad structure through denoising alone and avoided collapsing to one mode. The bridge and off-manifold points also show that this small denoising network is not a perfect density model. A GAN on this dataset would often collapse to one mode; diffusion models resist this failure mode more effectively.

Visualizing Training Loss

Out[16]:
Visualization
Line plot showing decreasing MSE training loss over training epochs
Training loss curve for the diffusion denoising MLP over 3000 epochs, sampled every 200 steps. The loss measures mean squared error between predicted and actual noise (the DDPM objective). Despite stochastic batch-to-batch fluctuations, the overall decline shows that the model is learning to predict noise across timestep levels.

Visualizing Classifier-Free Guidance Scale

The guidance scale ww is one of the most impactful hyperparameters in text-to-image generation. The CFG formula decomposes the guided noise prediction into a linear combination of conditional and unconditional predictions. We can visualize exactly how the weights on each component change with different guidance scales.

Out[17]:
Visualization
Line plot showing guidance scale effect on conditional vs unconditional noise mixing
Effect of classifier-free guidance scale on the balance between conditional and unconditional noise predictions. Higher guidance scales (w) amplify the conditional signal (text-guided direction) relative to the unconditional baseline, increasing semantic fidelity at the cost of diversity. At w=1 the conditional prediction is used as-is; at w=12 the conditional direction dominates strongly.

At w=7.5w = 7.5, the conditional noise prediction receives a weight of 7.5 and the unconditional prediction receives a weight of −6.5-6.5. The negative unconditional weight means the model actively moves away from unconditional predictions, amplifying whatever features the text condition adds. This is why high guidance scales can cause oversaturation: the model is maximally exploiting the direction of the text signal, sometimes beyond what produces photorealistic results.

Key Parameters

Diffusion model design involves several critical hyperparameters that jointly determine output quality and generation speed as well as controllability.

The key parameters are:

  • T (number of timesteps): Controls granularity of the diffusion process. More steps allow finer-grained denoising but increase sampling time. Typical training values: 100-1000 steps. Typical inference values: 20-50 steps with advanced samplers like DPM-Solver++.
  • Beta schedule (linear vs. cosine): The linear schedule is simple but allocates timesteps unevenly. The cosine schedule typically produces better results by spending more steps at informative intermediate noise levels.
  • Guidance scale (w): Higher values increase text fidelity but reduce diversity and can cause oversaturation. Typical values: 7-12 for creative generation where strong prompt adherence is desired, 1-3 for realistic variation with more diversity.
  • Latent compression factor: For latent diffusion, the ratio between image size and latent size. An 8x spatial compression (standard in Stable Diffusion) reduces compute substantially while maintaining quality. Too much compression loses fine detail; too little is computationally expensive.
  • U-Net attention resolution: Self-attention is added at lower resolutions where the spatial dimensions are small. Adding attention at higher resolutions captures finer long-range dependencies but increases memory cost quadratically with the spatial size.
  • VAE regularization strength: The KL term in VAE training controls how tightly the latent distribution matches a Gaussian. Too much regularization produces a degenerate latent space; too little allows the latent to become so non-Gaussian that the diffusion model cannot traverse it effectively.

Limitations and Practical Challenges

Diffusion models are impressive but face several real limitations that affect both research progress and practical deployment.

Slow sampling remains the most significant practical challenge for real-time applications. Even with advanced samplers like DPM-Solver++, generating a single high-quality image requires 20-50 neural network forward passes through the full U-Net. Each forward pass is a complete evaluation of a network with hundreds of millions to billions of parameters. On a high-end consumer GPU, this translates to 1-5 seconds per image for a 512x512 output. For comparison, GANs generate images in a single forward pass, typically under 100 milliseconds. Consistency Models (Song et al., 2023) attempt to address this by training a model that can generate directly from noise in one or two steps by distilling a diffusion model's multi-step trajectory. Flow matching approaches (Lipman et al., 2022) reframe diffusion as learning straight-line trajectories in distribution space, which allows larger, more efficient integration steps at inference. Both show promise but still trail full diffusion sampling in the highest quality regime.

Compositional failures are pervasive and poorly understood. Generating an image of "a red cube to the left of a blue sphere" reliably is surprisingly difficult for current models. The model must understand the individual objects, their spatial relationships, and their attribute bindings, and these frequently break down for unusual or complex combinations. "A black cat and a white dog" may produce a black dog and a white cat, because the model does not reliably associate attributes with their intended objects. Models trained with dense spatial supervision (from datasets with bounding boxes or segmentation masks) show improvement, but compositional reasoning remains below human level. The root cause is likely that the model learns statistical associations from image captions, which are inherently unordered and do not provide explicit spatial or relational information.

Text rendering in generated images has historically been poor. Diffusion models process text as visual patterns, not as sequences of characters with specific forms. Images containing text often show plausible-looking letterforms that spell nonsensical words. DALL-E 3 and similar models improved substantially on this through better data curation and training at higher resolutions. But text rendering remains below the reliability expected of professional design tools, especially for longer strings or unusual fonts.

Copyright and provenance present legal and ethical complexity that extends beyond the technical. These models are trained on internet-scale image datasets scraped without explicit consent, and they can reproduce stylistic elements from specific artists. When someone prompts "a painting in the style of [living artist]," the generated image can be commercially valuable and indistinguishable from a human production, raising serious questions about attribution and consent as well as economic impact on the artists whose work was used for training. The legal framework for AI-generated art and training data is unsettled in most jurisdictions, with ongoing litigation in the US and regulatory activity in the EU.

Memorization is a related but distinct concern. Research has demonstrated that diffusion models can memorize specific training images and reproduce them near-exactly when prompted appropriately, particularly for images that appeared multiple times in training data. This is problematic both for privacy (if training data contained personal images like faces or medical data) and copyright. Unlike language model memorization of text sequences, image memorization reproduces recognizable visual content that could have commercial or legal implications.

Evaluation difficulty confounds progress measurement. FID correlates poorly with human preference rankings in several studies: models with better FID scores are sometimes rated worse by human evaluators. CLIP Score does not capture spatial correctness, object count, or fine-grained attribute binding. Developing better automated evaluation metrics that strongly correlate with human judgment is an active research problem, and the lack of reliable metrics makes it difficult to claim that improvements on standard benchmarks translate to better user-facing quality.

Despite these limitations, diffusion models have enabled capabilities with broad practical impact. Professional-grade image generation is now accessible to non-artists. Personalization through fine-tuning on small image sets enables customized visual content. Creative augmentation for film and game production has changed pipeline workflows. Synthesis of training data for other machine learning systems addresses data scarcity. Medical imaging and scientific visualization benefit from diffusion-based super-resolution and synthesis. The combination of accessible open-source models (Stable Diffusion) and capable proprietary APIs (DALL-E 3, Midjourney, Firefly) has made generative image models a widely deployed technology in just three years.

Summary

This chapter covered the core ideas and practical methods of modern image generation:

  • Diffusion models define a forward process that gradually corrupts images with Gaussian noise, then train a neural network to reverse this process step by step, converting pure noise into structured samples
  • The closed-form forward process allows computing noisy images at any timestep tt directly: xt=αˉt x0+1−αˉt ϵ\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\, \mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\, \boldsymbol{\epsilon}, making training parallelizable across timesteps
  • The simplified training objective trains the network to predict added noise: L=E[∥ϵ−ϵθ(xt,t)∥2]\mathcal{L} = \mathbb{E}\left[\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\|^2\right], derived from a variational lower bound but simplified for stability
  • Latent diffusion moves the process to a compressed VAE latent space, reducing compute by roughly an order of magnitude without sacrificing image quality
  • Text-to-image generation uses cross-attention to inject text embeddings into the U-Net at every spatial scale, with classifier-free guidance amplifying the text-conditioned signal at inference to improve prompt adherence
  • Sampling efficiency improved from 1000-step DDPM to 20-50-step DDIM and DPM-Solver++ by exploiting the ODE structure of the reverse process
  • Image editing techniques including SDEdit, inpainting, InstructPix2Pix, DreamBooth, LoRA, Textual Inversion extend generation to controlled modification of existing images and subject personalization
  • ControlNet and IP-Adapter provide spatial and visual conditioning beyond text, enabling precise compositional control
  • Evaluation uses FID for distributional quality, CLIP Score for text alignment, and human evaluation for subjective quality; each metric captures different aspects and has known blind spots
  • Key limitations include slow sampling, compositional failures, poor text rendering, copyright challenges, memorization, difficult evaluation

Diffusion models represent the current state of the art in image generation. Their core ideas, noise corruption and learned reversal, have already extended beyond static images into video generation (Sora, Runway Gen-3, Stable Video Diffusion), 3D shape synthesis, protein structure generation, audio synthesis, and molecular design. The mathematical framework of learning to denoise at multiple noise scales has proven to be a remarkably flexible and general principle for generative modeling.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about image generation and diffusion models.

Image Generation: Diffusion Models & Text-to-Image

Question 1 of 80 of 8 completed
What does the simplified DDPM training objective train the network to predict?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026imagegeneration, author = {Michael Brenndoerfer}, title = {Image Generation: Diffusion Models, Text-to-Image, Editing}, year = {2026}, url = {https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Image Generation: Diffusion Models, Text-to-Image, Editing. Retrieved from https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image
MLAAcademic
Michael Brenndoerfer. "Image Generation: Diffusion Models, Text-to-Image, Editing." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image>.
CHICAGOAcademic
Michael Brenndoerfer. "Image Generation: Diffusion Models, Text-to-Image, Editing." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Image Generation: Diffusion Models, Text-to-Image, Editing'. Available at: https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Image Generation: Diffusion Models, Text-to-Image, Editing. https://mbrenndoerfer.com/writing/image-generation-diffusion-models-text-to-image

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.