Atlas · skill

Presidio

Presidio is an open-source toolkit, originally developed at Microsoft, for detecting and transforming personally identifiable information. Its analyzer locates candidate entities in text, and its anonymizer applies selected operations to those spans, enabling configurable privacy preprocessing while leaving detection quality and residual disclosure risk for the application to evaluate.

toolPII & Privacy Tooling

What it is

Presidio separates detection from transformation. Recognizers may combine patterns, checksums, context and named-entity models to identify spans such as phone numbers or personal names. The analyzer returns entity types, offsets and confidence information; anonymization operators then redact, replace, hash or otherwise transform the detected spans. This separation lets an application choose different treatment for different identifiers and extend recognition for its domain. The toolkit's use of the term anonymizer does not establish legal anonymization: some transformations are reversible, and indirect attributes can still identify a person even when explicit identifiers are removed. The project is transitioning to independent community governance under Data Privacy Stack, whose documentation and release channels should be used when checking current deployment instructions.

What the work involves

The practitioner selects recognizers and supported language resources, adds domain-specific patterns and evaluates them on representative labeled text. They define per-entity thresholds and transformation policies, then verify that offsets remain correct through text processing. Useful artifacts include detection precision and recall by entity type, a transformation specification and tests for overlapping spans. Reversible mappings or encryption keys need separate protection. The implementation should also examine what happens to raw text in logs, error responses and temporary files rather than checking only the returned sanitized string.

Illustrative example

A team prepares support conversations for an evaluation dataset. Presidio detects common email addresses and phone numbers, while a custom recognizer finds the organization's customer identifiers. The pipeline replaces these with consistent placeholders so dialogue references remain understandable. Reviewers inspect sampled failures and challenge the pipeline with unusual formatting. Raw conversations stay restricted, and only the reviewed transformed dataset enters the broader evaluation workflow.

Limits and common mistakes

Recognizers can miss rare names, images, misspellings or contextual identifiers and can remove ordinary words by mistake. Hashing predictable identifiers may still permit linkage, while replacement does not remove identifying combinations of attributes. Presidio is a component in a privacy workflow, not evidence of compliance by itself. Evaluate residual information and downstream usefulness together, and document entity types or languages the configured detector does not reliably cover.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10