Numerical Computing & Stability

Floating point, overflow, log-sum-exp and the reasons a correct formula still returns NaN.

Module 01.5 · Intermediate
Free

8h · 6 lessons · 1 challenge

Before this: Calculus & Optimization

About this module

Your softmax is mathematically correct and returns NaN. Your loss was fine in float32 and diverges in float16. This module covers the gap between the equation and the machine, which is where a surprising share of real training failures live.

Short, and disproportionately useful: the log-sum-exp trick alone has saved more training runs than any architecture choice.

After this module you can

6 lessons

Lesson 1
Floating point, precision and representable numbers
Kind
Concept
Length
45 min
Lesson 2
Overflow, underflow and catastrophic cancellation
Kind
Concept
Length
45 min
Lesson 3
The log-sum-exp trick, derived
Kind
Concept
Length
35 min
Lesson 4
NaN Hunt: find the operation that poisoned the run
Kind
Interactive
Length
35 min
Lesson 5
Mixed precision: float16, bfloat16 and loss scaling
Kind
Concept
Length
50 min
Lesson 6
Checkpoint: stabilize the computation
Kind
Checkpoint
Length
20 min

The challenge: NaN Hunt

A training run has gone to NaN at step 4,120. Binary-search the computation graph to the offending op. Fewer probes, more points.

Worth 350 XP.

Related modules

All modules
Python for ML Engineers
10h · 7 lessons
NumPy & Vectorization
8h · 7 lessons
SQL & Data Modeling
8h · 7 lessons