AI Safety, Security, Governance & Ethics
24 skills · ontology graph below shows relations within this section.
What this domain covers
This edition groups 24 capabilities in AI Safety, Security, Governance & Ethics across 13 named categories. The inventory contains 18 concepts and 6 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.
Current category labels: AI Application Security · AI Security · Content Safety · Ethics · Explainability & Fairness · Governance & Standards · Guardrails · LLM Security · and 5 more
Frequent learning foundations
- Prompt Injection Defense supports 6 mapped skills
- EU AI Act Compliance supports 3 mapped skills
- LLM Observability supports 2 mapped skills
- Adversarial AI Testing supports 1 mapped skill
- AI Agent Design supports 1 mapped skill
Skills in this section
Secure RAG is the practice of making retrieval-augmented generation respect information boundaries throughout ingestion, retrieval and answer delivery. It combines document authorization, trusted identity handling and adversarial testing so that an assistant can use relevant evidence without exposing material its user is not allowed to read.
AI data security protects the confidentiality and integrity of information as it moves through datasets, model services, retrieval systems and agents. The skill focuses on preventing unauthorized access, disclosure or modification, including exposures introduced by prompts, generated outputs, tool calls and operational logs.
AI rate limiting controls how quickly users, tenants or processes consume model and agent resources. It protects availability and budgets by enforcing request, token, concurrency or work limits, especially where a small input can trigger expensive generation, retrieval or repeated tool execution.
AI supply chain security addresses threats introduced through third-party models, datasets, libraries, containers and deployment components. It asks whether an AI artifact has trustworthy provenance, can be verified before use and remains protected from tampering as it moves from development into production.
AI toxicity analysis evaluates whether model inputs or outputs contain language that is abusive, hateful, threatening or otherwise harmful under a defined content policy. It combines automated detection with contextual review to measure harmful generation and choose moderation actions without treating every sensitive discussion as abuse.
AI ethics examines how AI systems affect people, institutions and the distribution of benefits and harms. It guides choices about purpose, data, oversight and deployment by making values and tradeoffs explicit, including questions that legal compliance or predictive performance alone cannot answer.
AI fairness is the practice of identifying and reducing unjust differences in an AI system's treatment or effects across people and groups. It connects statistical assessment to the deployment context, because equal aggregate accuracy does not establish that errors, opportunities or service quality are distributed fairly.
Explainable AI produces information that helps people understand an AI system's behavior for a particular purpose. It includes interpretable models and post-hoc explanations, with attention to what an explanation actually supports, who needs it and whether it faithfully reflects the system rather than merely sounding plausible.
Mechanistic interpretability investigates how a neural network's internal components produce particular behaviors. It studies activations, learned features and computational circuits, using interventions to test hypotheses about mechanisms rather than relying only on input-output correlations or a model's verbal explanation of itself.
AI auditability is the ability to reconstruct and examine how an AI system was built, evaluated and used. It connects decisions to evidence through documentation, versioning and appropriately governed records, so reviewers can assess a specific system rather than relying on general claims about its model family.
ISO/IEC 42001 is a standard for establishing, maintaining and improving an organizational AI management system. Competence in it means translating AI governance responsibilities into repeatable processes and evidence, including how the organization evaluates risks, manages lifecycle changes and checks whether its controls work.
NeMo Guardrails is NVIDIA's open-source toolkit for adding programmable behavioral controls to LLM applications. The skill involves configuring and evaluating checks around conversation flows, inputs and outputs, so application policies become testable behavior with explicit handling for rejected, redirected or modified responses.
Prompt injection defense protects an LLM application when untrusted content tries to redirect its behavior. It focuses on preserving the boundary between instructions and data, while limiting the information and actions an attacker could obtain even if the model follows a malicious instruction.
Presidio is an open-source toolkit, originally developed at Microsoft, for detecting and transforming personally identifiable information. Its analyzer locates candidate entities in text, and its anonymizer applies selected operations to those spans, enabling configurable privacy preprocessing while leaving detection quality and residual disclosure risk for the application to evaluate.
AI watermarking embeds a detectable signal in generated content so an authorized detector can assess whether a particular generation process likely produced it. The signal may be statistical or encoded in media; its usefulness depends on detection accuracy, content length and robustness to editing or transformation.
AI red teaming is an organized attempt to discover harmful or exploitable behavior before it affects real users. It uses adversarial scenarios, realistic attacker goals and careful evidence collection to challenge the system's assumptions, then turns findings into fixes, risk decisions and repeatable regression tests.
Adversarial AI testing measures how an AI system behaves under deliberately challenging inputs or attack conditions. It converts specified threats into repeatable experiments, such as evasion, poisoning or prompt injection, so teams can compare defenses and detect regressions against a known attacker model.
EU AI Act compliance is the practice of identifying which obligations apply to an AI activity and producing evidence that the responsible organization meets them. It connects intended use, operator roles, risk classification and lifecycle controls to the applicable legal text rather than treating a technical checklist as compliance.
The NIST AI Risk Management Framework is a voluntary framework for managing risks associated with AI throughout its lifecycle. Competence in it means using its Govern, Map, Measure and Manage functions to connect organizational responsibilities, deployment context, evidence and risk treatment in a repeatable process.
Agent threat modeling with MAESTRO analyzes how an AI agent's models, data, frameworks and integrations can be exploited together. The Cloud Security Alliance framework provides a layered way to identify attack paths, especially where memory, delegated actions and multiple agents create risks beyond an isolated model request.
The OWASP Top 10 for LLM Applications is a community-maintained taxonomy of major security risks in language-model applications. The skill uses that taxonomy to structure design reviews and testing, while translating broad risk categories into specific attack paths, controls and evidence for the system being built.
SAIF, Google's Secure AI Framework, helps organizations connect AI-specific security risks to lifecycle components and controls. Competence in it means adapting a security framework to models, data, infrastructure and applications, then checking that the chosen controls address the organization's actual attack surfaces.
LIME explains an individual model prediction by fitting a simpler surrogate around that input. It perturbs interpretable parts of the example, observes the original model's outputs and learns which local changes matter, producing an explanation whose meaning depends on the chosen neighborhood and representation.
SHAP explains model predictions through additive feature attributions based on Shapley-value ideas. It allocates the difference between a prediction and a reference value across input features, providing a common explanation form whose interpretation depends on the background data and how missing features are modeled.