Atlas · skill

Ensemble Learning

Ensemble learning combines multiple models to produce one prediction. Different constructions reduce instability, correct errors sequentially or learn how to mix complementary predictors. The skill is creating genuinely useful diversity and evaluating the combined system, including the added computation and the danger of leakage when a second model learns from first-stage predictions.

conceptSupervised Learning

What it is

Bagging trains learners on resampled data and aggregates them, while boosting builds an additive sequence that improves an objective. Voting and averaging combine predictions directly; stacking trains another model on predictions from base learners. These approaches have different assumptions and sources of benefit. Diversity matters because models making the same errors offer little new information to combine. An ensemble may improve predictive quality but also increase complexity, latency and storage. Competence includes understanding how training data reach every layer, especially constructing out-of-fold predictions for stacking so the meta-model does not learn from overly optimistic in-sample outputs.

What the work involves

Begin with independently evaluated base models and inspect whether their errors differ on relevant cases. Choose an aggregation rule suitable for class scores, probabilities or continuous outcomes and align output scales. For stacking, produce leakage-safe training predictions and preserve the full preprocessing path for each learner. Compare the ensemble with its strongest component and measure operational cost. The result should explain which complementary behavior justifies the combination, rather than attributing a small score increase to the number of models alone.

Illustrative example

Suppose, illustratively, a demand predictor combines a seasonal baseline with a tree model using weather and calendar features. The analyst observes that their errors differ across ordinary days and special events. A weighted average is fitted on development data and evaluated on later periods. If stacking is considered, the meta-model receives historical out-of-fold predictions. The team checks whether the improvement survives the final test and whether maintaining both pipelines is worth the added operational effort.

Limits and common mistakes

Highly correlated models may add cost without useful improvement. Incompatible probability calibration or preprocessing can undermine aggregation. A stacking model trained on in-sample predictions leaks fitting information and can appear deceptively accurate. Ensemble interpretation is more difficult than interpreting each component separately. Gains can disappear after distribution shift, and averaging can dilute a model that performs well on a critical subgroup. Compare the complete system against a clear baseline with the same information and resource constraints.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10