Attention & the Transformer

Self-attention, multi-head attention and a complete transformer implemented from the paper, with every shape written out.

Module 05.2 · Advanced
Free

14h · 10 lessons · 1 challenge

Before this: Tokenization & Embeddings

About this module

The center of the curriculum. You implement a transformer end to end — scaled dot-product attention, multiple heads, positional encoding, the residual and normalization structure, the feedforward block — and train it on a real task.

Every tensor shape is written on the board at every step. Being able to reproduce that diagram from memory is close to a guarantee of passing a transformer depth interview.

After this module you can

10 lessons

Lesson 1
Queries, keys, values and the attention operation
Kind
Concept
Length
60 min
Lesson 2
Why scaled, and what the softmax temperature does
Kind
Concept
Length
40 min
Lesson 3
Multi-head attention and what heads specialize in
Kind
Concept
Length
55 min
Lesson 4
Positional encoding: sinusoidal, learned, rotary, ALiBi
Kind
Concept
Length
55 min
Lesson 5
The block: residuals, normalization, feedforward
Kind
Concept
Length
50 min
Lesson 6
Implement a transformer from the paper
Kind
Lab
Length
150 min
Lesson 7
Encoder-only, decoder-only, encoder-decoder
Kind
Concept
Length
50 min
Lesson 8
KV caching and how generation actually runs
Kind
Concept
Length
50 min
Lesson 9
Shape Sprint: annotate the block against the clock
Kind
Interactive
Length
40 min
Lesson 10
Checkpoint: the transformer, in full
Kind
Checkpoint
Length
35 min

The challenge: Shape Sprint

An unlabeled transformer block. Annotate every tensor shape before the timer runs out. The interview version of this question, made into a drill.

Worth 700 XP.

Related modules

All modules