Diffusion & Multimodal Models

Denoizing diffusion, latent diffusion, and the models that put images, audio and text in one system.

Module 06.7 · Advanced
Free

8h · 8 lessons · 1 challenge

Before this: Convolutional Networks & Vision, Attention & the Transformer

About this module

The other half of generative AI. Diffusion is derived as iterative denoizing, which makes the sampler choices and guidance scales interpretable rather than magical.

The multimodal half covers vision-language models and the architectures that fuse modalities, which is increasingly what "frontier model" means.

After this module you can

8 lessons

Lesson 1
Denoizing diffusion, forward and reverse
Kind
Concept
Length
60 min
Lesson 2
Latent diffusion and the compute saving
Kind
Concept
Length
45 min
Lesson 3
Conditioning, guidance and control
Kind
Concept
Length
50 min
Lesson 4
Samplers, steps and quality tradeoffs
Kind
Concept
Length
40 min
Lesson 5
Vision-language models and modality fusion
Kind
Concept
Length
55 min
Lesson 6
Fine-tune a diffusion model on a small set
Kind
Lab
Length
100 min
Lesson 7
Denoize Derby: reach the image in fewest steps
Kind
Interactive
Length
35 min
Lesson 8
Checkpoint: generative model families
Kind
Checkpoint
Length
25 min

The challenge: Denoize Derby

A target image and a step budget. Choose sampler, schedule and guidance to get closest before the budget runs out.

Worth 500 XP.

Related modules

All modules