Floating point, overflow, log-sum-exp and the reasons a correct formula still returns NaN.
8h · 6 lessons · 1 challenge
Before this: Calculus & Optimization
Your softmax is mathematically correct and returns NaN. Your loss was fine in float32 and diverges in float16. This module covers the gap between the equation and the machine, which is where a surprising share of real training failures live.
Short, and disproportionately useful: the log-sum-exp trick alone has saved more training runs than any architecture choice.
A training run has gone to NaN at step 4,120. Binary-search the computation graph to the offending op. Fewer probes, more points.
Worth 350 XP.