Tokenization & Embeddings

Byte-pair encoding, vocabulary design, and what an embedding space actually contains.

Module 05.1 · Advanced
Free

8h · 7 lessons · 1 challenge

Before this: Sequence Models

About this module

Tokenization is where a surprising number of model behaviors originate: why arithmetic is hard, why some languages cost three times more, why a trailing space changes an answer. You implement BPE, so none of that stays mysterious.

The embeddings half goes from word2vec to contextual representations, with the geometry made visible rather than asserted.

After this module you can

7 lessons

Lesson 1
Characters, words, subwords and bytes
Kind
Concept
Length
45 min
Lesson 2
Implement byte-pair encoding
Kind
Lab
Length
80 min
Lesson 3
Vocabulary size, coverage and cost per language
Kind
Concept
Length
45 min
Lesson 4
Token Trap: why the model cannot count
Kind
Interactive
Length
30 min
Lesson 5
word2vec, GloVe and the geometry of meaning
Kind
Concept
Length
50 min
Lesson 6
Contextual embeddings and what changed
Kind
Concept
Length
45 min
Lesson 7
Checkpoint: tokenization and its consequences
Kind
Checkpoint
Length
25 min

The challenge: Token Trap

Prompts that fail for tokenization reasons alone. Predict the failure, then fix the prompt without changing its meaning.

Worth 400 XP.

Related modules

All modules