Optimizers & Training Dynamics

Learning rate schedules, warmup, batch size effects, gradient clipping and mixed precision.

Module 04.3 · Intermediate
Free

8h · 8 lessons · 1 challenge

Before this: Making Deep Networks Train

About this module

The learning rate is the hyperparameter that matters most and the one people tune least methodically. This module covers schedules and warmup, the interaction between batch size and learning rate, clipping, and the precision choices that let large runs fit.

Every claim is checked empirically — you run the sweep and see the shape of the answer.

After this module you can

8 lessons

Lesson 1
SGD, momentum, Adam and AdamW compared
Kind
Concept
Length
50 min
Lesson 2
Learning rate range tests and schedules
Kind
Concept
Length
50 min
Lesson 3
Warmup, cosine decay and why they help
Kind
Concept
Length
40 min
Lesson 4
Batch size, gradient noise and the linear scaling rule
Kind
Concept
Length
50 min
Lesson 5
Clipping, accumulation and mixed precision
Kind
Concept
Length
45 min
Lesson 6
LR Roulette: find the sweet spot in five runs
Kind
Interactive
Length
40 min
Lesson 7
Run the sweep and read the surface
Kind
Lab
Length
80 min
Lesson 8
Checkpoint: tune the run
Kind
Checkpoint
Length
25 min

The challenge: LR Roulette

A fixed compute budget and a model that has not converged. Five runs to find a learning rate and schedule that gets there.

Worth 450 XP.

Related modules

All modules
Structuring ML Projects
8h · 8 lessons
Sequence Models
10h · 7 lessons
Tokenization & Embeddings
8h · 7 lessons