Presidio: Data Protection and De-identification SDK
Presidio (Origin from Latin praesidium ‘protection, garrison’) helps to ensure sensitive data is properly managed and governed. It provides fast identification and anonymization modules for private entities in text and images such as credit card numbers, names, locations, social security numbers, bitcoin wallets, US phone numbers, financial data and more.
- Allow organizations to preserve privacy in a simpler way by democratizing de-identification technologies and introducing transparency in decisions.
- Embrace extensibility and customizability to a specific business need.
- Facilitate both fully automated and semi-automated PII de-identification flows on multiple platforms.
How it works
- Predefined or custom PII recognizers leveraging Named Entity Recognition, regular expressions, rule based logic and checksum with relevant context in multiple languages.
- Options for connecting to external PII detection models.
- Multiple usage options, from Python or PySpark workloads through Docker to Kubernetes.
- Customizability in PII identification and anonymization.
- Module for redacting PII text in images.
Presidio can help identify sensitive/PII data in un/structured text. However, because it is using automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information. Consequently, additional systems and protections should be employed.
- Presidio analyzer: PII identification in text
- Presidio anonymizer: De-identify detected PII entities using different operators
- Presidio image redactor: Redact PII entities from images using OCR and PII identification
- Samples for running Presidio via code
- Running Presidio as an HTTP service
- Setting up a development environment
- Perform PII identification using presidio-analyzer
- Perform PII de-identification using presidio-anonymizer
- Perform PII identification and redaction in images using presidio-image-redactor
- Example deployments