Atlas · skill

Hyperparameter Optimization

Hyperparameter optimization searches configurations that control model structure or training. It uses validation performance and a resource budget to choose among candidates. The competence is defining a meaningful search space and reliable objective while preventing leakage, overfitting to validation results and expensive searches that offer little improvement over sensible baselines.

conceptModel Selection & Tuning

What it is

Hyperparameters are settings not estimated by the ordinary fitting procedure, such as tree depth, regularization strength or learning rate. Grid and random search explore predefined spaces; adaptive methods use previous results to propose later configurations. Multi-fidelity approaches allocate less computation to weak candidates, often using partial training as evidence. The search objective is an estimate from a validation procedure, so its noise and bias affect selection. Optimization can cover an entire preprocessing and model pipeline. Searching more candidates also creates more opportunities to select a configuration that fits quirks of validation data rather than the underlying task.

What the work involves

Choose a score aligned with the decision, specify leakage-safe folds and include preprocessing within each trial. Define ranges on appropriate scales and represent conditional settings so invalid combinations are excluded. Allocate a budget, use pruning only when intermediate scores predict final usefulness and record every trial's configuration and outcome. Compare the best configuration with a default or expert baseline, then evaluate the final choice independently. The deliverable is a reproducible search study with an honest assessment of improvement and its computation cost.

Illustrative example

Suppose, illustratively, a gradient-boosted classifier has learning rate, depth and tree count to tune. An engineer uses grouped validation because multiple records belong to the same account. Trial results reveal that deeper trees improve training scores but not validation. The search narrows toward simpler configurations and uses early stopping on development data. A final untouched test measures the selected pipeline; the best trial's validation score is retained as selection evidence rather than reported as unbiased final performance.

Limits and common mistakes

An adaptive search can overfit its validation set, and unstable scores can lead it toward lucky trials. Pruning may discard slow-starting configurations. An overly broad space wastes resources, while a narrow space can exclude useful settings. Reproducibility requires seeds, data splits and software configuration, not only the winning parameters. Nested validation or a separate final test may be needed to assess the full selection process. Greater search effort is not evidence that the selected model is more trustworthy.

Prerequisites

Related skills

Sources and further reading

Last updated: 2026-10-10