Feature & Data Infrastructure

Feature stores, streaming versus batch, point-in-time correctness and lineage.

Module 07.5 · Advanced
Free

8h · 7 lessons · 1 challenge

Before this: Data Pipelines & Feature Engineering

About this module

The infrastructure that makes the same feature available offline for training and online for serving, computed identically, at the right point in time. Get it wrong and every model on top of it is quietly wrong.

Covers the build-versus-buy question honestly — most teams do not need a feature store, and knowing when you do is the actual skill.

After this module you can

7 lessons

Lesson 1
Offline and online stores, and the consistency problem
Kind
Concept
Length
50 min
Lesson 2
Point-in-time joins done properly
Kind
Concept
Length
55 min
Lesson 3
Streaming versus batch, feature by feature
Kind
Concept
Length
45 min
Lesson 4
Lineage, versioning and reproducibility
Kind
Concept
Length
45 min
Lesson 5
Build a minimal feature store
Kind
Lab
Length
100 min
Lesson 6
Time Travel: rebuild the training set as of last March
Kind
Interactive
Length
40 min
Lesson 7
Checkpoint: feature architecture
Kind
Checkpoint
Length
25 min

The challenge: Time Travel

Reconstruct a training set exactly as it was on a past date, with no future information leaking in. The grader checks every row.

Worth 550 XP.

Related modules

All modules
Cost, Capacity & Tradeoffs
6h · 6 lessons
The ML Breadth Interview
8h · 6 lessons