Objectives, scaling laws, compute-optimal training and how frontier models are actually built.
10h · 7 lessons · 1 challenge
Before this: Training Transformers at Scale
Scaling laws turn "should we train longer or bigger" from an argument into an arithmetic problem. This module works that arithmetic: Kaplan, Chinchilla, and what the compute-optimal frontier implies for a real budget.
Also covers mixture-of-experts, long-context methods and where the published scaling picture is contested — because a candidate who knows the disagreements sounds different from one who has read a summary.
Ten million dollars of compute. Choose parameters, tokens and context. The game trains your configuration against the compute-optimal one and shows you the gap.
Worth 650 XP.