Denoizing diffusion, latent diffusion, and the models that put images, audio and text in one system.
8h · 8 lessons · 1 challenge
Before this: Convolutional Networks & Vision, Attention & the Transformer
The other half of generative AI. Diffusion is derived as iterative denoizing, which makes the sampler choices and guidance scales interpretable rather than magical.
The multimodal half covers vision-language models and the architectures that fuse modalities, which is increasingly what "frontier model" means.
A target image and a step budget. Choose sampler, schedule and guidance to get closest before the budget runs out.
Worth 500 XP.