Training Transformers at Scale

Data pipelines, curricula, stability tricks and the failure modes of long training runs.

Module 05.3 · Advanced
Free

10h · 8 lessons · 1 challenge

Before this: Attention & the Transformer

About this module

Training a transformer for a week is a different discipline from training one for an hour. Loss spikes, dead runs, corrupted checkpoints and data ordering effects all become real problems with real remedies.

This module is the practical craft: data mixing and deduplication, curriculum, checkpointing, and reading the telemetry of a run you cannot afford to restart.

After this module you can

8 lessons

Lesson 1
Corpora: sourcing, mixing, deduplication, filtering
Kind
Concept
Length
55 min
Lesson 2
Sequence packing, masking and throughput
Kind
Concept
Length
45 min
Lesson 3
Loss spikes: causes and recoveries
Kind
Concept
Length
50 min
Lesson 4
Checkpointing, resumption and determinism
Kind
Concept
Length
45 min
Lesson 5
Run telemetry and what to log
Kind
Concept
Length
40 min
Lesson 6
Spike Watch: save the run
Kind
Interactive
Length
40 min
Lesson 7
Train a small language model end to end
Kind
Lab
Length
140 min
Lesson 8
Checkpoint: diagnose the long run
Kind
Checkpoint
Length
25 min

The challenge: Spike Watch

A month-long run, compressed. Loss spikes arrive; roll back, skip the batch, lower the rate, or hold your nerve. Reach the end.

Worth 600 XP.

Related modules

All modules