Inference Optimization & Serving

Quantization, batching, KV cache management, speculative decoding and the latency budget behind them.

Module 07.4 · Advanced
Free

10h · 8 lessons · 1 challenge

Before this: GPUs, Kernels & Performance

About this module

Inference is where the bill arrives. This module covers the levers — quantization, continuous batching, paged KV cache, speculative decoding, distillation — and what each costs in quality.

Framed throughout as a budget problem: a latency target, a throughput target and a cost ceiling, with the levers traded against each other explicitly.

After this module you can

8 lessons

Lesson 1
The serving stack and where latency accumulates
Kind
Concept
Length
50 min
Lesson 2
Quantization: int8, int4, and quality loss
Kind
Concept
Length
55 min
Lesson 3
Continuous batching and paged KV cache
Kind
Concept
Length
55 min
Lesson 4
Speculative decoding and draft models
Kind
Concept
Length
50 min
Lesson 5
Distillation and smaller models
Kind
Concept
Length
45 min
Lesson 6
Serve a model to a latency and cost target
Kind
Lab
Length
120 min
Lesson 7
Latency Ladder: hit p99 under budget
Kind
Interactive
Length
45 min
Lesson 8
Checkpoint: optimize the deployment
Kind
Checkpoint
Length
30 min

The challenge: Latency Ladder

A p99 target, a quality floor and a cost ceiling. Pull the levers in any order; the game charges you for every one.

Worth 650 XP.

Related modules

All modules
Cost, Capacity & Tradeoffs
6h · 6 lessons
The ML Breadth Interview
8h · 6 lessons