LLM Pretraining & Scaling Laws

Objectives, scaling laws, compute-optimal training and how frontier models are actually built.

Module 06.1 · Advanced
Free

10h · 7 lessons · 1 challenge

Before this: Training Transformers at Scale

About this module

Scaling laws turn "should we train longer or bigger" from an argument into an arithmetic problem. This module works that arithmetic: Kaplan, Chinchilla, and what the compute-optimal frontier implies for a real budget.

Also covers mixture-of-experts, long-context methods and where the published scaling picture is contested — because a candidate who knows the disagreements sounds different from one who has read a summary.

After this module you can

7 lessons

Lesson 1
Pretraining objectives across model families
Kind
Concept
Length
50 min
Lesson 2
Scaling laws: Kaplan, Chinchilla and after
Kind
Concept
Length
60 min
Lesson 3
Compute arithmetic: FLOPs, tokens, parameters, dollars
Kind
Concept
Length
55 min
Lesson 4
Mixture-of-experts and conditional computation
Kind
Concept
Length
50 min
Lesson 5
Long context: attention variants and position extension
Kind
Concept
Length
50 min
Lesson 6
Budget Frontier: allocate the training run
Kind
Interactive
Length
40 min
Lesson 7
Checkpoint: scaling arithmetic
Kind
Checkpoint
Length
30 min

The challenge: Budget Frontier

Ten million dollars of compute. Choose parameters, tokens and context. The game trains your configuration against the compute-optimal one and shows you the gap.

Worth 650 XP.

Related modules

All modules
Instruction Tuning & Alignment
10h · 9 lessons
Retrieval-Augmented Generation
10h · 8 lessons
Agents & Tool Use
10h · 8 lessons