Atlas · GenAI 2026
Data Preprocessing for ML
Building repeatable transformations that prepare raw features for machine-learning estimators, including encoding, scaling and missing-value treatment, while fitting learned transformations only on training data.
Also searchable as: ML Data Preprocessing, Machine Learning Data Preprocessing
conceptFeature EngineeringAI consensus: 0/3
Prerequisites
Understanding feature distributions and missingness informs preprocessing choices.
Recommended reference
scikit-learn: Preprocessing data — https://scikit-learn.org/stable/modules/preprocessing.html; scikit-learn: Common pitfalls and recommended practices — https://scikit-learn.org/stable/common_pitfalls.html#data-leakage
Reviewed sources
Primary and first-party material reviewed for this editorial summary. These citations are separate from the AI consensus score above.
- scikit-learn: Preprocessing data
Feature transformations, scaling and categorical encoding.
- scikit-learn: Common pitfalls and recommended practices
Training-only fitting of transformations and consistent use at evaluation and prediction time.
- scikit-learn: Pipeline
Composing fitted transformations and an estimator for consistent evaluation and prediction.
Notes from AI deep research
Related skills
- → is part of: Feature Engineering