Streaming-first
Process continuous text streams with boundary buffering, context windows, and safe flush points. No need to buffer entire inputs in memory.
Detect and redact sensitive data with 129 built-in detectors, AI-powered semantic confirmation, and full streaming support.
Process continuous text streams with boundary buffering, context windows, and safe flush points. No need to buffer entire inputs in memory.
Email, phone, payment cards, national IDs, passports, person names, cloud keys, JWT tokens, private keys, healthcare identifiers, HR data, legal references, crypto addresses, tracking numbers, and more across 13 domains.
The core library has no runtime dependencies. Person name detection uses a lightweight bloom filter (person_name_lite). Optional NER via compromise.js (person_name) for higher recall.
Enable restoration mode to get numbered placeholders and a restoration map. Reverse redaction back to original text with a single call.
Redact to type labels, mask with configurable preservation, remove entirely, format-preserve structure, or token-replace with deterministic fakes.
Start with pii, gdpr, hipaa, ccpa, pci-dss, healthcare, finance, education, soc2, or security. Combine and override with explicit rules.
Re-redacting already-redacted text is a no-op. Placeholder filtering ensures safe pipeline chaining without double-processing.
Grapheme cluster boundary enforcement and UTF-16 offset reporting. Handles emoji, combining marks, and multi-codepoint characters correctly.
Opt-in Jev (TypeSafe System One) integration confirms detected PII candidates before redacting. Eliminates false positives without sacrificing recall. 100% precision in independent evaluation.