Making Deep Networks Train

Initialization, normalization, dropout, weight decay and the vanishing gradient -- everything between a network that exists and one that converges.

Module 04.2 · Intermediate
Free

10h · 8 lessons · 1 challenge

Before this: Neural Networks from Scratch

About this module

A deep network that will not train is the standard experience, and the fixes are not folklore — they follow from what happens to gradient and activation statistics as depth increases.

This module derives why Xavier and He initialization have the constants they do, what batch and layer normalization each stabilize, and how to tell from a training run which one you need.

After this module you can

8 lessons

Lesson 1
Vanishing and exploding gradients, quantified
Kind
Concept
Length
50 min
Lesson 2
Xavier and He initialization, derived
Kind
Concept
Length
50 min
Lesson 3
Batch, layer and RMS normalization
Kind
Concept
Length
55 min
Lesson 4
Dropout, weight decay and modern regularization
Kind
Concept
Length
45 min
Lesson 5
Residual connections and why depth became possible
Kind
Concept
Length
45 min
Lesson 6
Training Triage: the run is not converging
Kind
Interactive
Length
45 min
Lesson 7
Fix five broken training runs
Kind
Lab
Length
100 min
Lesson 8
Checkpoint: stabilize the network
Kind
Checkpoint
Length
25 min

The challenge: Training Triage

Loss curves, gradient norms and activation histograms from a failing run. One fix per case, twelve cases, no re-runs.

Worth 550 XP.

Related modules

All modules
Structuring ML Projects
8h · 8 lessons
Sequence Models
10h · 7 lessons