Advanced · every module free · 5 challenges

Self-attention, multi-head attention and a complete transformer implemented from the paper, with every shape written out.

Data pipelines, curricula, stability tricks and the failure modes of long training runs.

Full fine-tuning, LoRA, QLoRA and adapters — adapting a pretrained model on a budget you actually have.