AI Guardrails
AI guardrails are controls that check or constrain model inputs, outputs and proposed actions against application rules. They can combine deterministic validation, classifiers and review steps, with a defined response to violations, so acceptable behavior is enforced at specific points in an AI workflow.
What it is
A guardrail evaluates a boundary condition such as an output schema, prohibited content category, tool permission or required evidence. Some checks operate before generation, others examine a completed response, and action checks can run before a tool executes. Deterministic rules are useful for explicit constraints; model-based checks handle less easily specified judgments but introduce uncertainty. A guardrail is separate from the model's general instructions and from the underlying authorization system. It can block, redact, request revision or escalate a result. The choice of intervention matters because detecting a violation after an external action cannot undo the action.
What the work involves
The practitioner converts requirements into named checks, identifies where each check runs and defines fail-open or fail-closed behavior deliberately. Test cases should include acceptable requests that resemble violations, not only obvious attacks. Action guardrails need to inspect actual tool arguments and user authority. A useful result is a policy map with tested interventions and observable decisions. Monitoring records false positives, missed violations and latency so a control can be adjusted without silently weakening the protected boundary.
Illustrative example
An assistant drafts customer-facing responses and can request a refund through a tool. One check validates that the draft contains no unsupported delivery promise; a separate check verifies the refund amount against the authenticated customer's order and approval policy. A blocked tool call is returned as a structured policy failure. Tests include a legitimate high-value order and a malicious request embedded in a quoted email, ensuring that content classification does not replace authorization.
Limits and common mistakes
Model-based guardrails can inherit the errors or susceptibility of the system they check. A broad refusal rule can prevent legitimate work, while a narrow phrase filter can be bypassed through paraphrase. Schema validity says little about truth or permission. Quality requires checks tied to concrete requirements, enforcement before consequential effects and a measured response to uncertainty. Guardrails reduce particular failure modes; they do not certify the entire application or replace careful tool design and access control.
Prerequisites
- mediumStructured LLM Outputs
Guardrails validate that outputs conform to expected schemas and policies — structured output concepts are the foundation
- mediumPrompt Injection Defense
Guardrails are a defensive layer that includes prompt injection mitigation — understanding attacks informs defense design
Related skills
- → is part of: AI Safety
- ← is an instance of: NeMo Guardrails
Sources and further reading
- LangChain guardrails
Describes deterministic and model-based controls around input, output and agent execution.
Last updated: 2026-10-10