Insurance Runs on Documents Full of Other People's Lives
No industry accumulates a stranger mix of personal data than insurance. A single auto claim file can contain the policyholder's SSN and driver's license, the other driver's name and plate, a witness's phone number, a police report with everyone's addresses, medical records describing injuries, and an adjuster's free-text narrative tying it all together. Multiply that by millions of claims, decades of retention obligations, and a growing web of TPAs, reinsurers, and analytics vendors — and "where is our PII?" becomes one of the hardest questions a carrier can be asked.
The uncomfortable truth is that most of this data lives in unstructured text and scanned documents that traditional data-governance tools never see. Claims systems index the structured fields, but the risk sits in the FNOL description, the adjuster diary, the ISO claim search response pasted into a note, the demand letter from opposing counsel. When a regulator, an auditor, or a breach-response team needs an inventory, schema-level tools come back empty-handed.
Our PII Detection API reads that unstructured layer. Send any text — a claim note, an OCR'd police report, an underwriting memo — and receive every detected entity with its type, exact character offsets, and confidence score, plus an optional masked version via mask_mode. It is the discovery engine that turns insurance privacy programs from policy binders into running software.