Hypothesis Testing
Hypothesis testing evaluates how compatible observed data are with a specified null model. The skill is choosing a defensible test, understanding error rates and reporting the evidence in relation to a practical question. It requires explicit hypotheses, design-aware assumptions and restraint when interpreting a p-value or a nonsignificant result.
Also searchable as: hypothesis-testing, Statistical Hypothesis Testing
What it is
A statistical test defines a null hypothesis, an alternative, a statistic and its distribution under the null assumptions. A p-value measures the probability of a result at least as extreme as the observed one under that null model; it is not the probability that the null is true. A decision threshold controls a specified error rate under the test's conditions. Power concerns sensitivity to particular alternatives. Tests can address means, distributions, proportions or model parameters, with parametric and resampling approaches. The method organizes evidence against a model, but its relevance depends on whether the tested hypothesis corresponds to the scientific or operational question.
What the work involves
State the hypothesis and the smallest effect that would matter before selecting a method. Check randomization or sampling assumptions, dependence, sample size and the suitability of the statistic. Account for multiple comparisons and any sequential examination of results. Report estimates and uncertainty alongside the test outcome, distinguishing evidence of a difference from its operational importance. The analysis should also explain whether it was designed to detect a difference, demonstrate equivalence or establish noninferiority, because these require different hypotheses and decision rules.
Illustrative example
For an illustrative packaging comparison, a team asks whether a new process changes average defect rate. The analyst defines the direction and relevant effect size, checks whether batches rather than individual items are the independent units and selects a corresponding test. A small p-value prompts examination of the effect estimate and interval. If the result is nonsignificant, the team asks whether the experiment had adequate sensitivity rather than concluding that the two processes are identical.
Limits and common mistakes
Rejecting a null does not prove a causal mechanism, and failing to reject it does not establish equivalence. Large samples can detect effects too small to matter. Repeated testing, subgroup searches and outcome switching can make nominal error rates misleading. Test assumptions include more than a distributional check; biased sampling or dependent observations can dominate the conclusion. Separate statistical evidence from the decision cost, and use appropriately designed equivalence or noninferiority procedures when similarity is the actual question.
Prerequisites
- mediumProbability Theory
Sampling distributions, error probabilities and p-values depend on probability concepts.
Related skills
- → is subcategory of: Statistical Inference
Sources and further reading
- NIST: Product and Process Comparisons
Hypothesis tests, comparisons and interpretation of uncertainty.
- SciPy: Statistical functions
Parametric, nonparametric, resampling and multiple-testing tools.
Last updated: 2026-10-10