Quantization, batching, KV cache management, speculative decoding and the latency budget behind them.
10h · 8 lessons · 1 challenge
Before this: GPUs, Kernels & Performance
Inference is where the bill arrives. This module covers the levers — quantization, continuous batching, paged KV cache, speculative decoding, distillation — and what each costs in quality.
Framed throughout as a budget problem: a latency target, a throughput target and a cost ceiling, with the levers traded against each other explicitly.
A p99 target, a quality floor and a cost ceiling. Pull the levers in any order; the game charges you for every one.
Worth 650 XP.