Training Infrastructure & Distributed Training

Data, tensor and pipeline parallelism, ZeRO and FSDP, and making a run survive a failed node.

Module 07.2 · Advanced
Free

10h · 8 lessons · 1 challenge

Before this: Optimizers & Training Dynamics

About this module

When a model stops fitting on one device you need to know which axis to split and what that costs in communication. This module covers the parallelism strategies, the sharding schemes, and the collective operations underneath them.

Reliability gets real attention: long multi-node runs fail, and the difference between losing an hour and losing a week is entirely preparation.

After this module you can

8 lessons

Lesson 1
Memory arithmetic: weights, gradients, optimizer, activations
Kind
Concept
Length
55 min
Lesson 2
Data parallelism and gradient synchronization
Kind
Concept
Length
50 min
Lesson 3
Tensor and pipeline parallelism
Kind
Concept
Length
60 min
Lesson 4
ZeRO, FSDP and sharded optimizer state
Kind
Concept
Length
55 min
Lesson 5
Collectives, interconnect and where time goes
Kind
Concept
Length
45 min
Lesson 6
Fault tolerance and elastic training
Kind
Concept
Length
45 min
Lesson 7
Shard Shuffle: fit the model across the cluster
Kind
Interactive
Length
45 min
Lesson 8
Checkpoint: plan the distributed run
Kind
Checkpoint
Length
30 min

The challenge: Shard Shuffle

A model, a cluster and an interconnect. Choose the parallelism split; the game reports whether it fit and what throughput you left on the table.

Worth 650 XP.

Related modules

All modules