Atlas · skill

Hypothesis

Hypothesis is a Python library for property-based testing that generates examples from specified input strategies and searches for failures of stated properties. The competency is defining useful invariants and realistic input spaces, then interpreting a minimized counterexample as evidence of a faulty assumption or implementation.

toolTesting & Quality

What it is

A property-based test describes behavior that should hold for a class of inputs rather than listing only individual examples. Hypothesis draws inputs from composable strategies, runs the test and attempts to simplify a failing example. Shrinking makes a counterexample easier to investigate, while retained examples can help reproduce regressions. The framework does not invent the oracle: the author decides what property should hold. It complements example-based tests and differs from formal proof, because generated exploration is finite and bounded by the supplied strategies and execution settings.

What the work involves

The practitioner selects invariants meaningful to the component, such as round-trip equivalence, conservation of totals or agreement with a simpler implementation. They construct strategies that represent valid structures and deliberately include edge conditions. They avoid filtering away difficult inputs and distinguish unsupported values from implementation defects. On failure, they inspect the reduced example, fix the root cause and retain a focused regression test where helpful. Useful work expands coverage of assumptions while keeping failures understandable and test execution practical.

Illustrative example

An engineer tests a serializer for nested feature records. A strategy generates nullable values, Unicode keys and lists of varying length. The property requires decoding an encoded valid record to preserve its contents. Hypothesis finds a small record where an empty list becomes null. The engineer repairs the encoding rule and adds a specific example documenting why the two values must remain distinct.

Limits and common mistakes

Weak properties can pass broken code, and unrealistic strategies can spend effort on irrelevant inputs. A round-trip test can miss paired errors when both encoder and decoder make the same mistake. External services or uncontrolled randomness can make results unstable. Check oracle independence, input validity and shrinking behavior. Hypothesis explores the behavior expressed by a test; it cannot establish an unspecified requirement or prove correctness for all possible execution environments.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10