Cross-Validation
Cross-validation estimates predictive performance by repeatedly fitting on one part of the data and evaluating on another. The skill is selecting splits that reflect future use and keeping every learned transformation inside each training partition. It supports model comparison, but requires a separate view of the uncertainty and bias introduced by selection.
Also searchable as: Cross Validation, Cross validation techniques
What it is
A cross-validation procedure defines training and validation indices for several fits. Ordinary folds suit some independent-sample settings; grouped splits prevent shared entities crossing boundaries; temporal splits preserve information order. Stratification maintains class proportions where appropriate but does not solve entity or temporal leakage. Each fold assesses a newly fitted pipeline, including preprocessing, feature selection and resampling. Aggregated scores describe performance under the specified split design. Reusing them to choose models creates selection effects, so the best score is not automatically an unbiased assessment of the selected procedure. Nested validation or a reserved final test can assess that broader process.
What the work involves
Identify the unit of independence and the deployment boundary before selecting a splitter. Fit imputers, scalers, vocabularies and feature selectors within each training fold, and keep related records or future information out of validation. Choose a scoring rule, inspect variation and errors across folds and compare candidates on the same splits. Record indices or reproducible split logic. The deliverable explains what scenario the procedure simulates and how final performance will be assessed after model selection, rather than reporting an average without describing how data were separated.
Illustrative example
Suppose, illustratively, a model predicts outcomes from multiple visits per customer. Randomly splitting rows would let earlier or similar visits from the same customer appear on both sides. The analyst uses grouped folds when evaluating new-customer use, fitting preprocessing afresh in each fold. If the application instead predicts later visits for existing customers, a time-aware design answers that different question. The validation plan is chosen from the intended deployment, not from whichever split yields a higher score.
Limits and common mistakes
Fold scores are dependent because training sets overlap, so their spread is not automatically a confidence interval for final performance. Cross-validation cannot repair biased labels or an unrepresentative sample. Repeatedly tuning against the same folds can overfit selection. Temporal gaps or embargoes may be needed when outcomes overlap in time. Grouped and temporal designs answer different generalization questions. Document the boundary and evaluate the complete pipeline, including feature selection and resampling, to avoid optimistic leakage.
Prerequisites
- mediumModel Evaluation
Correct fold design and interpretation depend on evaluation metrics, leakage control and validation objectives.
Related skills
- → is subcategory of: Model Evaluation
Sources and further reading
- scikit-learn: Cross Validation
Fold strategies, groups, time dependence, pipelines and model-selection assessment.
Last updated: 2026-10-10