Reinforcement Learning Foundations

MDPs, value and policy methods, Q-learning and policy gradients — the machinery RLHF is built from.

Module 03.7 · Advanced
Free

10h · 8 lessons · 1 challenge

Before this: Calculus & Optimization, Probability Foundations

About this module

Reinforcement learning matters twice over: as its own field, and as the foundation of the alignment methods in the LLM track. PPO on a language model makes very little sense until policy gradients make sense on a gridworld.

The module stays hands-on — you implement tabular Q-learning, then DQN, then REINFORCE with a baseline — and ends pointed directly at RLHF.

After this module you can

8 lessons

Lesson 1
Markov decision processes, returns and discounting
Kind
Concept
Length
50 min
Lesson 2
Value iteration and policy iteration
Kind
Concept
Length
50 min
Lesson 3
Q-learning, exploration and the epsilon schedule
Kind
Concept
Length
55 min
Lesson 4
Deep Q-networks and what breaks
Kind
Lab
Length
90 min
Lesson 5
Policy gradients and REINFORCE
Kind
Concept
Length
60 min
Lesson 6
Actor-critic and PPO, on the way to RLHF
Kind
Concept
Length
55 min
Lesson 7
Bandit Arena: explore against exploit
Kind
Interactive
Length
35 min
Lesson 8
Checkpoint: formalize the environment
Kind
Checkpoint
Length
25 min

The challenge: Bandit Arena

Your exploration policy against epsilon-greedy, UCB and Thompson sampling over a thousand pulls. Regret is the score.

Worth 500 XP.

Recommender Systems

End of the track

Related modules

All modules
Neural Networks from Scratch
12h · 8 lessons
Making Deep Networks Train
10h · 8 lessons
Structuring ML Projects
8h · 8 lessons