A/B Testing
A/B testing compares alternatives through randomized assignment and a predefined outcome. The skill covers experiment design, reliable measurement and interpretation of uncertainty. A good test answers a specific decision question while accounting for assignment units, sample requirements, guardrail outcomes and the consequences of repeated or selective analysis.
What it is
In an A/B test, eligible units are assigned to a control or treatment, and their outcomes are compared under a specified analysis. Randomization aims to prevent systematic differences in observed and unobserved characteristics from driving the comparison. The unit can be a user, account or cluster, and should match how the intervention reaches people. Design also defines the population, exposure, outcome window and estimand. A simple difference in averages is only one possible analysis; blocking, covariate adjustment or cluster-aware inference may be appropriate. The method provides evidence about the tested change under the experiment's conditions, not an unrestricted claim about product quality.
What the work involves
Translate the decision into a primary metric and a minimum effect worth acting on. Choose an assignment unit that avoids spillover and specify exclusions, stopping rules and analysis before collecting outcomes. Check that assignment and exposure are logged correctly, compare allocated counts and inspect missingness or attrition. Estimate the effect with an interval, examine guardrails and explain practical significance alongside statistical uncertainty. The output should connect a deployment decision to an auditable design, including what would count as insufficient or conflicting evidence.
Illustrative example
For an illustrative checkout experiment, a retailer compares the existing form with a shorter version. Assignment occurs at account level so the same customer does not encounter both variants. Completed orders are the primary outcome, while payment failures and support contacts are guardrails. The team chooses a run period that covers ordinary weekly variation and reviews instrumentation before interpreting the difference. A rise in completion accompanied by more payment failures would require a decision beyond merely selecting the variant with the higher primary metric.
Limits and common mistakes
Repeatedly checking ordinary fixed-horizon significance tests and stopping when a result looks favorable can distort error rates. Multiple outcomes and segments create additional opportunities for selective conclusions. Network effects, novelty, inconsistent exposure and missing outcomes can weaken the randomized comparison. A statistically detectable effect may be too small to justify a change, while an inconclusive test does not establish equivalence. Generalization should consider seasonal conditions, the tested population and whether implementation after the experiment matches the treatment that was actually evaluated.
Prerequisites
A/B testing requires understanding of p-values, confidence intervals, power analysis, and multiple comparison correction — all core statistical concepts
Related skills
- → is an instance of: Experimental Design
Sources and further reading
- NIST: Choosing an experimental design
Randomization, blocking and choosing a design for the experimental question.
- NIST: Product and Process Comparisons
Statistical comparison of groups and interpretation of experimental evidence.
Last updated: 2026-10-10