Instruction Tuning & Alignment

SFT, reward models, RLHF, DPO and the constitutional methods — how a base model becomes an assistant.

Module 06.2 · Advanced
Free

10h · 9 lessons · 1 challenge

Before this: Reinforcement Learning Foundations, LLM Pretraining & Scaling Laws

About this module

A base model completes text; an assistant follows instructions and declines some of them. The gap is closed by supervised fine-tuning, a reward model, and a policy optimization step — and this module walks the whole pipeline with the RL foundations already in place.

DPO and its relatives get equal weight, since they are what most teams now reach for, and the tradeoff against PPO is a live interview question.

After this module you can

9 lessons

Lesson 1
Supervised fine-tuning and instruction data
Kind
Concept
Length
50 min
Lesson 2
Preference data collection and its biases
Kind
Concept
Length
45 min
Lesson 3
Reward models and reward hacking
Kind
Concept
Length
55 min
Lesson 4
RLHF with PPO, step by step
Kind
Concept
Length
60 min
Lesson 5
DPO and direct preference methods
Kind
Concept
Length
55 min
Lesson 6
Constitutional and AI-feedback approaches
Kind
Concept
Length
45 min
Lesson 7
Reward Hack: break the reward model
Kind
Interactive
Length
45 min
Lesson 8
Align a small model with DPO
Kind
Lab
Length
120 min
Lesson 9
Checkpoint: the alignment pipeline
Kind
Checkpoint
Length
30 min

The challenge: Reward Hack

You are the policy. Maximize a reward model without doing the task it was meant to reward. Then patch the reward model against your own exploit.

Worth 650 XP.

Related modules

All modules
Retrieval-Augmented Generation
10h · 8 lessons
Agents & Tool Use
10h · 8 lessons
Evaluating LLM Systems
8h · 8 lessons