Atlas · skill

OWASP Top 10 for LLM Applications

The OWASP Top 10 for LLM Applications is a community-maintained taxonomy of major security risks in language-model applications. The skill uses that taxonomy to structure design reviews and testing, while translating broad risk categories into specific attack paths, controls and evidence for the system being built.

conceptSecurity Frameworks

What it is

The list addresses risks such as prompt injection, information disclosure, supply chain weaknesses and excessive agency across an LLM application's lifecycle. It focuses attention on failure patterns that conventional web security reviews may not fully capture. Each category describes a family of problems, not a single vulnerability with one mandatory fix. Categories can overlap in an attack: a poisoned document may inject instructions that exploit broad tool permissions and disclose data. Versions of the list change as the field develops, so a review must identify the version it uses and consider additional threats specific to its architecture.

What the work involves

A practitioner maps the system's components and capabilities to relevant OWASP categories, then writes concrete abuse cases and mitigation checks. A useful review records where each risk can arise, which boundary prevents impact and how that boundary is tested. Teams should link findings to owners and deployment decisions rather than merely marking categories as considered. The taxonomy complements conventional controls for authentication, authorization and dependency security. It also helps communicate findings across model, application and security teams using a shared vocabulary without assuming that every category applies equally.

Illustrative example

A team reviews an assistant with database access and email tools. Prompt-injection analysis identifies hostile text entering through search results; excessive-agency analysis examines unnecessary write permissions; disclosure analysis examines outbound messages. The review leads to read-only queries, restricted destinations and regression tests for combined attacks. One attack spans several categories, so the team documents the complete path rather than counting it as three unrelated checklist findings.

Limits and common mistakes

Covering the Top 10 does not prove security or exhaust all possible threats. A checklist can overlook deployment-specific trust boundaries, and category names cannot replace reproducible evidence. The list also addresses application security rather than every ethical or legal concern about AI. Good use prioritizes risks according to actual capabilities and impact, while recording the version, assumptions and areas the review did not examine.

Prerequisites

Sources and further reading

Last updated: 2026-10-10