GPUs, Kernels & Performance

The memory hierarchy, arithmetic intensity, fused kernels and writing your own in Triton.

Module 07.3 · Advanced
Free

10h · 8 lessons · 1 challenge

Before this: Training Infrastructure & Distributed Training

About this module

Almost every model is memory-bandwidth bound, and understanding why is what separates engineers who can make a model faster from engineers who file a ticket.

You profile a real workload, find the bound, and write a fused kernel in Triton. Flash Attention is studied as the worked example of exactly this reasoning.

After this module you can

8 lessons

Lesson 1
GPU architecture and the memory hierarchy
Kind
Concept
Length
55 min
Lesson 2
Arithmetic intensity and the roofline model
Kind
Concept
Length
50 min
Lesson 3
Kernel launch, occupancy and memory coalescing
Kind
Concept
Length
50 min
Lesson 4
Fusion, and Flash Attention as the worked example
Kind
Concept
Length
60 min
Lesson 5
Profile a training step and find the bound
Kind
Lab
Length
90 min
Lesson 6
Write a fused kernel in Triton
Kind
Lab
Length
120 min
Lesson 7
Roofline Rally: find the bound, fix it
Kind
Interactive
Length
45 min
Lesson 8
Checkpoint: performance reasoning
Kind
Checkpoint
Length
30 min

The challenge: Roofline Rally

Profiles from eight workloads. Name the bound — bandwidth, compute, launch overhead, synchronization — and pick the fix that buys the most.

Worth 650 XP.

Related modules

All modules