Feature Engineering
Feature engineering designs useful model inputs from available observations while respecting when and how those observations become available. It combines domain reasoning, transformations and validation to create representations that improve the intended prediction task without introducing target leakage or inconsistent behavior between training and serving.
Also searchable as: inżynieria cech
What it is
A feature is a variable or representation supplied to an estimator. Engineering may derive ratios, historical aggregates, categorical encodings, text representations or indicators of missing information. Its scope includes choosing the observation unit, time window, source joins and transformation parameters. Feature extraction creates representations; feature selection chooses a subset; preprocessing prepares data for an estimator. Feature engineering coordinates these decisions around a task. A feature's statistical association is insufficient if it uses information unavailable at prediction time or encodes a process that changes after deployment.
What the work involves
The practitioner identifies candidate signals with domain experts, records their definitions and availability, and builds transformations that can be repeated on new data. They enforce point-in-time joins for historical features, fit learned transformations within training folds and compare candidates with a simple baseline. They inspect missingness, stability, acquisition cost and subgroup behavior. Useful work yields a versioned feature specification and executable pipeline, with evidence for keeping or removing each important signal rather than a large collection of unexplained columns.
Illustrative example
For forecasting equipment failure, an engineer derives recent temperature variation and time since the last completed maintenance visit. Each training row is anchored to a prediction timestamp. Maintenance records entered after that timestamp are excluded, even if they describe an earlier event. The team compares the added features on a chronological holdout and verifies that the serving system can compute identical windows from live readings.
Limits and common mistakes
More features can increase leakage, redundancy and maintenance burden. Target encoding, aggregate windows and source joins are common routes for hidden access to future outcomes. A feature useful in one operating regime may fail after a policy or sensor change. Validate availability, transformation fit boundaries and train-serving parity separately from model scores. Predictive usefulness also does not establish that manipulating a feature will change the outcome.
Prerequisites
- mediumETL Pipeline Design
Feature stores are served by pipelines that compute and refresh features — pipeline design enables feature engineering at scale
Related skills
- → is part of: Machine Learning
- → is part of: Predictive Analytics
- → is part of: Data Science
- ← is an instance of: Feast
- ← is part of: Data Preprocessing for ML
- ← is part of: Feature Scaling
- ← is subcategory of: Feature Extraction
- ← is subcategory of: Feature Selection
Sources and further reading
- Rules of Machine Learning
Supports feature design, training-serving consistency and production measurement.
- Scikit-learn common pitfalls
Documents leakage and fitting transformations only within training data.
Last updated: 2026-10-10