Skip to content
GenAI Learn/Diffusion Models & Multimodal Generation
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

How Diffusion Models Generate Images

10 min read

You'll learn to

  • -Explain the forward diffusion process as a fixed, no-training-involved way to noise an image to any degree
  • -Explain the reverse process as what a diffusion model actually learns: predicting the noise to remove
  • -Implement classifier-free guidance and explain why it makes text-to-image models follow prompts more literally

Every module before this one assumed the goal was generating text. Stable Diffusion, DALL-E, and Midjourney generate images with a completely different mechanism - one that has nothing to do with predicting the next token. Instead, a diffusion model learns to reverse a process of progressively destroying an image with noise. Understanding that one idea unlocks the entire family.

The Forward Process: Destroying an Image on a Schedule

Training a diffusion model needs images at every stage of being progressively noised. The forward process is a fixed formula - no learning involved - that takes a clean image x0 and a timestep t, and produces a noised version x_t directly, in one step, with no need to simulate t small noise-additions in sequence.

DDPM Forward Process
xₜ = √ᾱₜ · x₀ + √(1 - ᾱₜ) · noise
alpha_bar_t is a scalar between 0 and 1 from a precomputed noise schedule: 1 means no noise added yet (x_t = x0), 0 means the signal is completely destroyed (x_t = pure noise).
Forward diffusion: noising a data point to a given timestep directly

The Reverse Process: A Model Learns to Undo It

Generation runs the forward process backward: starting from pure noise, and repeatedly removing a little of it until a coherent image emerges. A trained model's actual job is narrow - look at a noisy x_t and predict what noise was added - and that predicted noise plugs directly into a formula that nudges x_t toward x_{t-1}, slightly less noisy. Chain enough of these steps (real models take anywhere from a handful to thousands, depending on the sampler) and static becomes an image.

DDPM Reverse Step
xₜ₋₁ = (xₜ - (βₜ/√(1-ᾱₜ))·ε̂) / √αₜ + √βₜ · z
ε̂ is the model's predicted noise; z is a fresh random sample (typically zeroed on the final step so the last image is deterministic).
One reverse denoising step, given a predicted noise

Classifier-Free Guidance: Making the Model Actually Follow the Prompt

A text-to-image model can predict noise two ways for the same noisy image: conditioned on the prompt, or completely unconditioned (ignoring it). Classifier-Free Guidance runs both and extrapolates past the conditional prediction, in the direction away from the unconditional one - exaggerating exactly what the prompt contributed. This is the "CFG scale" slider in every Stable Diffusion-style UI.

Classifier-Free Guidance
guided = uncond + scale · (cond - uncond)
scale=1 recovers plain conditional generation; scale=0 recovers plain unconditional generation; typical UIs use 7-15, well past 1.
Classifier-free guidance: exaggerating the prompt's effect

AI Lab's Levels 45-47 implement all three formulas in numpy, over full arrays instead of single points - the exact same math, vectorized.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement all three from scratch: Level 68 (Forward Diffusion), Level 69 (Reverse Denoising Step), and Level 70 (Classifier-Free Guidance) in AI Lab.