Atlas · skill

AI Requirements Engineering

AI requirements engineering specifies what an AI-enabled system must do, under which conditions and with what evidence of acceptance. The competency is translating user needs and uncertain model behavior into testable contracts, including data, permissions, quality thresholds, fallback and operational constraints that cannot be inferred from a prompt alone.

conceptRequirements & Specs

What it is

A requirement states a needed system property or behavior and connects it to a verification or validation method. AI requirements combine deterministic rules, such as permitted actions and response schemas, with statistical expectations assessed over defined cases. They must identify the population and context in which quality is measured. This differs from product management's selection of value and scope, and from prompt writing's configuration of one model interaction. The specification covers the complete system, including retrieval, tools, interface and human workflow, rather than assuming a model's general capability becomes an application guarantee.

What the work involves

The practitioner elicits concrete tasks and unacceptable outcomes, records assumptions and resolves ambiguous terms through examples. They define inputs, outputs, data availability, permission boundaries and error behavior. For variable outputs, they choose a representative evaluation set, review criteria and release conditions. They trace requirements to implementation and checks, assigning owners for unresolved questions. Useful work yields an actionable specification and acceptance plan that engineering, design and reviewers can apply consistently, including what the system should do when it cannot satisfy a request.

Illustrative example

A team specifies a policy-answering assistant. Requirements state which document collection it may use, how citations identify supporting passages and when it must report that evidence is insufficient. Access filtering has deterministic integration tests; answer usefulness has a documented review rubric and representative questions. A scenario with conflicting policy versions checks that the assistant exposes the conflict instead of inventing one definitive answer.

Limits and common mistakes

Terms such as accurate, helpful or safe remain ambiguous without context and checks. A benchmark threshold can hide failure on important subgroups or rare consequential cases. Overly narrow acceptance can reward formatting while missing the real user need. Inspect traceability, evaluation coverage and fallback requirements. Requirements cannot make a probabilistic model deterministic; they define acceptable system behavior and the evidence needed to decide whether a particular implementation meets it.

Prerequisites

  • Specs operationalize product thinking — you must understand the value proposition before you can specify acceptance criteria

  • Acceptance criteria for GenAI often map to evaluation metrics — understanding automated eval helps write testable specs

Related skills

Sources and further reading

Last updated: 2026-10-10