Convolutional Networks & Vision

Convolutions, the architectures that mattered, detection and segmentation, and vision transformers.

Module 04.5 · Intermediate
Free

12h · 9 lessons · 1 challenge

Before this: Making Deep Networks Train

About this module

Convolution as an operation, then the architectural line from LeNet through ResNet to EfficientNet, then the tasks — classification, detection, segmentation — and finally vision transformers and the CLIP-style contrastive models that connect vision to the language track.

You build a detector, which is where the abstractions become concrete.

After this module you can

9 lessons

Lesson 1
Convolution, stride, padding and receptive fields
Kind
Concept
Length
55 min
Lesson 2
Pooling, and the architectures that dropped it
Kind
Concept
Length
40 min
Lesson 3
LeNet to ResNet: what each step solved
Kind
Concept
Length
55 min
Lesson 4
Augmentation and vision-specific regularization
Kind
Concept
Length
45 min
Lesson 5
Detection and segmentation heads
Kind
Concept
Length
55 min
Lesson 6
Fine-tune a detector on your own images
Kind
Lab
Length
120 min
Lesson 7
Vision transformers and contrastive pretraining
Kind
Concept
Length
55 min
Lesson 8
Kernel Craft: build the filter that finds it
Kind
Interactive
Length
35 min
Lesson 9
Checkpoint: shapes, receptive fields, choices
Kind
Checkpoint
Length
30 min

The challenge: Kernel Craft

Hand-design convolution kernels to detect edges, corners and textures, then watch a trained network arrive at yours.

Worth 500 XP.

Related modules

All modules
Sequence Models
10h · 7 lessons
Tokenization & Embeddings
8h · 7 lessons
Attention & the Transformer
14h · 10 lessons
Training Transformers at Scale
10h · 8 lessons