Software Testing
Software testing establishes executable evidence about specified behavior, including failure handling and integration boundaries. In AI systems, it covers ordinary code, data transformations and service contracts while distinguishing these deterministic checks from evaluations of a model's probabilistic usefulness or factual quality.
What it is
Tests compare observed behavior with an oracle: an expected result, an invariant or a defined interaction. Unit tests isolate a small component; integration tests exercise collaboration between components; property-based tests explore classes of input through stated rules. Fixtures and controlled substitutes make scenarios repeatable, but they can also hide real integration problems. AI applications include deterministic infrastructure around nondeterministic model calls, so a test strategy needs separate treatment of parsing, permissions and state transitions versus task-level output quality. Coverage measures exercised code, not the adequacy of the oracle.
What the work involves
The practitioner derives tests from requirements and failure modes, chooses the smallest useful level and builds realistic fixtures. They verify edge cases, cleanup, authorization and error propagation rather than only the happy path. They keep external dependencies controlled for fast checks, then add integration evidence for important interfaces. Continuous integration runs an appropriate suite before changes advance. The deliverable is a maintainable set of checks that catches meaningful regressions and explains failures without simply reproducing the implementation's internal sequence.
Illustrative example
For a retrieval assistant, unit tests verify chunk offsets and citation formatting, integration tests check that unauthorized documents cannot enter results, and model evaluations assess whether answers use retrieved evidence. A fixture includes two documents with identical titles but different permissions. The team deliberately changes a filtering rule and confirms that the integration test fails even though the answer still reads fluently.
Limits and common mistakes
Mocks can validate an imagined dependency contract, and generated expected outputs can inherit the same error as the code. Flaky tests obscure real regressions; excessive implementation coupling makes harmless refactoring expensive. Passing deterministic checks does not establish model accuracy or security in every situation. Inspect oracle independence, representative boundary cases and integration realism. Use model evaluations alongside software tests where acceptance depends on generated behavior.
Prerequisites
Related skills
- ← is an instance of: Hypothesis
Sources and further reading
- Pytest good integration practices
Supports test organization, integration and reproducible test execution.
- Hypothesis documentation
Documents generated test inputs and property-based checking.
Last updated: 2026-10-10