Evaluating LLM Systems

Benchmarks, model-graded evaluation, human review and building an eval set that actually predicts production.

Module 06.6 · Advanced
Free

8h · 8 lessons · 1 challenge

Before this: Agents & Tool Use

About this module

"It seems better" is not shippable. Evaluation is the discipline that separates teams who improve their systems from teams who change them, and it is the skill hiring managers most consistently say is missing.

Covers benchmark contamination, LLM-as-judge and its biases, inter-rater agreement, and the statistics — from the inference module — needed to say whether a change is real.

After this module you can

8 lessons

Lesson 1
Benchmarks, contamination and what they miss
Kind
Concept
Length
50 min
Lesson 2
Building a task-specific eval set
Kind
Concept
Length
55 min
Lesson 3
LLM-as-judge: setup, biases, calibration
Kind
Concept
Length
55 min
Lesson 4
Human evaluation and inter-rater agreement
Kind
Concept
Length
45 min
Lesson 5
Online evaluation and guarded rollout
Kind
Concept
Length
45 min
Lesson 6
Judge the Judge: catch the biased grader
Kind
Interactive
Length
40 min
Lesson 7
Build an eval harness for a real system
Kind
Lab
Length
100 min
Lesson 8
Checkpoint: is this model better?
Kind
Checkpoint
Length
25 min

The challenge: Judge the Judge

A model-graded eval prefers verbose, confident, wrong answers. Find the bias, fix the rubric, and prove the fix.

Worth 550 XP.

Related modules

All modules