Data Pipelines & Feature Engineering

Encoding, scaling, splits that do not leak, and pipelines that produce the same features in training and serving.

Module 02.5 · Intermediate
Free

10h · 7 lessons · 1 challenge

Before this: Dataframes & Exploratory Analysis, SQL & Data Modeling

About this module

Training/serving skew is the defect that survives every code review and kills the model in production. It is nearly always a feature computed one way offline and another way online.

This module builds features the way that avoids it: one transformation defined once, split correctly, validated, versioned. It is also where the leakage material from the dataframes module becomes a discipline rather than a warning.

After this module you can

7 lessons

Lesson 1
Splits: random, grouped, temporal, and what each protects
Kind
Concept
Length
50 min
Lesson 2
Categorical encoding and target leakage
Kind
Concept
Length
50 min
Lesson 3
Scaling, binning and monotonic transforms
Kind
Concept
Length
40 min
Lesson 4
Skew Trap: same feature, two answers
Kind
Interactive
Length
40 min
Lesson 5
Build a pipeline that serves what it trained on
Kind
Lab
Length
100 min
Lesson 6
Data validation and contracts in CI
Kind
Concept
Length
45 min
Lesson 7
Checkpoint: audit this pipeline
Kind
Checkpoint
Length
25 min

The challenge: Skew Trap

Offline metrics are excellent, online is a disaster. Diff the two feature paths and find the divergence in each of eight scenarios.

Worth 500 XP.

SQL & Data Modeling

End of the track

Related modules

All modules