piidetectionapi.com
Home
Solutions - Fundamentals
What Is PII Detection? NER vs Regex vs Rules Accuracy, Precision & Recall PII in Test Data
Solutions - Compliance
GDPR Personal Data HIPAA PHI Detection CCPA / CPRA PCI DSS Card Data
Solutions - AI & LLM Safety
LLM Guardrails Chatbot PII Filtering RAG Pipelines
Solutions - Data Discovery & DLP
Data Loss Prevention Log File Scanning Support Tickets Email Scanning Documents & PDFs Database Discovery ETL & Streaming Pipelines
Industries - Financial
Banking Fintech Insurance
Industries - Healthcare
Healthcare Pharma & Clinical Trials Telehealth
Industries - Public Sector & Legal
Government & FOIA Law Enforcement Law Firms & eDiscovery Education (FERPA)
Industries - Technology
SaaS Platforms Cybersecurity & IR Telecommunications Gaming & Platforms
Industries - Other
HR & Recruiting Retail & E-commerce Call Centers & BPO Real Estate Travel & Hospitality Marketing & AdTech
How-to Guides - Identity & Contact
Detect Names Detect Email Addresses Detect Phone Numbers Detect Physical Addresses Detect Dates of Birth
How-to Guides - IDs & Financial
Detect SSNs Detect Passport Numbers Detect Drivers Licenses Detect Credit Card Numbers Detect Bank Accounts & IBAN
How-to Guides - Technical & Health
Detect IP & Device IDs Detect Medical Records & PHI
Resources
Pricing API Docs Supported Entities Languages About Contact Sign In Try the Live Demo Get Started
Insurance Solutions

PII Detection for Insurance

Locate and classify personal data across claims files, adjuster notes, medical records, and underwriting submissions. Meet NAIC model law expectations and state privacy rules with AI that reads insurance documents the way your adjusters do.

Claims documents Medical records NAIC model laws 150+ entity types GDPR-native

Insurance Runs on Documents Full of Other People's Lives

No industry accumulates a stranger mix of personal data than insurance. A single auto claim file can contain the policyholder's SSN and driver's license, the other driver's name and plate, a witness's phone number, a police report with everyone's addresses, medical records describing injuries, and an adjuster's free-text narrative tying it all together. Multiply that by millions of claims, decades of retention obligations, and a growing web of TPAs, reinsurers, and analytics vendors — and "where is our PII?" becomes one of the hardest questions a carrier can be asked.

The uncomfortable truth is that most of this data lives in unstructured text and scanned documents that traditional data-governance tools never see. Claims systems index the structured fields, but the risk sits in the FNOL description, the adjuster diary, the ISO claim search response pasted into a note, the demand letter from opposing counsel. When a regulator, an auditor, or a breach-response team needs an inventory, schema-level tools come back empty-handed.

Our PII Detection API reads that unstructured layer. Send any text — a claim note, an OCR'd police report, an underwriting memo — and receive every detected entity with its type, exact character offsets, and confidence score, plus an optional masked version via mask_mode. It is the discovery engine that turns insurance privacy programs from policy binders into running software.

Third parties in your files are data subjects too

Claimants, witnesses, medical providers, and other drivers never signed your privacy notice, yet their PII fills your claim files. Detection identifies every person's identifiers — not just the policyholder's — which is exactly what state privacy laws and defensible redaction workflows require. Browse the full list of supported identifiers in the entity catalog.

NAIC Model Laws and the State Patchwork

Insurance privacy is regulated state by state, but nearly every regime starts with the same demand: know what personal information you hold and where

NAIC Insurance Data Security Model Law (#668)

Adopted in the majority of US states, Model Law 668 requires licensees to run a written information security program grounded in a risk assessment of the nonpublic information they hold — and to notify regulators of cybersecurity events quickly, often within 72 hours. Both duties presuppose an accurate data inventory. Automated detection across claims, underwriting, and correspondence systems produces that inventory and keeps it current between assessments.

NAIC Privacy Models (#670 / #672 / #674 work)

The older privacy of consumer financial and health information models — and the ongoing consumer privacy protections modernization work — govern how carriers collect, use, and disclose policyholder information, including opt-out rights for third-party sharing. Enforcing a disclosure policy in practice means scanning what actually leaves your systems, not what the data-sharing agreement says should leave.

HIPAA at the Claims Boundary

Health insurers are HIPAA covered entities outright, and P&C carriers constantly receive medical records for injury claims. Records and adjuster summaries describing diagnoses and treatments deserve PHI-grade handling: detection of MEDICAL_RECORD_NUMBER, DIAGNOSIS, TREATMENT, and HEALTH_INSURANCE_ID alongside standard PII. See our HIPAA 18-identifiers guide.

CCPA/CPRA & State Privacy Laws

GLBA- and state-insurance-code exemptions cover some insurance data, but not all of it — website analytics, marketing lists, employee data, and much claims-adjacent processing remain in scope for CCPA/CPRA and its siblings. Responding to access and deletion requests means finding a named person's identifiers across systems fast. Start with the CCPA/CPRA detection guide.

The 72-Hour Problem

Model Law 668's tight notification windows expose the real weakness in most carriers' programs: scoping speed. When a claims share drive or an email archive is compromised, the clock starts before you know what was in it. Carriers that have already scanned those repositories answer the two questions regulators ask first — how many consumers, which data elements — from an existing index instead of launching a forensic document review under deadline pressure.

The same index pays for itself outside of incidents. Market-conduct exams, reinsurance audits, and TPA oversight reviews all ask for evidence that nonpublic information is identified and controlled. A detection pipeline generates that evidence continuously, with entity-type counts per system that a compliance officer can hand over without engineering help.

Personal Data in the Insurance Lifecycle

From quote to claim to subrogation, each stage accumulates its own identifier mix — and its own regulatory hooks

PERSON_NAME
Insureds, claimants, witnesses
SSN
Applications, 1099s, liens
DRIVERS_LICENSE_NUMBER
Auto policies & claims
DATE_OF_BIRTH
Rating & identity checks
ADDRESS
Risk locations, mailing
MEDICAL_DATA / DIAGNOSIS
Injury & health claims
MEDICAL_RECORD_NUMBER
Provider records in files
HEALTH_INSURANCE_ID
Coordination of benefits
PHONE / EMAIL
All parties & providers
FINANCIAL_ACCOUNT_NUMBER
Premium & payout details
EMPLOYMENT
Workers' comp files
PRESCRIPTION / TREATMENT
Life & disability records
Insurance Data Source Typical PII / PHI Found Regulatory Hooks Recommended Entity Filter
FNOL & claims intake Names, phone numbers, plate and license numbers, incident addresses NAIC 668, state UDAP rules PERSON_NAME, PHONE_NUMBER, DRIVERS_LICENSE_NUMBER, ADDRESS
Adjuster notes & diaries Every party's identifiers, medical snippets, SSNs pasted from documents NAIC 668, HIPAA (health claims), CCPA PERSON_NAME, SSN, DIAGNOSIS, MEDICAL_DATA, PHONE_NUMBER
Medical records in claim files MRNs, diagnoses, treatments, prescriptions, provider and patient identity HIPAA, state health privacy laws MEDICAL_RECORD_NUMBER, DIAGNOSIS, TREATMENT, PRESCRIPTION, HEALTH_INSURANCE_ID, DATE_OF_BIRTH
Underwriting submissions Applicant identity, DOB, SSN, financials, health questionnaires, loss runs NAIC privacy models, FCRA, GLBA PERSON_NAME, SSN, DATE_OF_BIRTH, FINANCIAL_ACCOUNT_NUMBER, MEDICAL_DATA, EMPLOYMENT
Policy correspondence & email Policy numbers tied to names, addresses, payment details NAIC 668, CCPA/CPRA, GDPR (intl.) PERSON_NAME, ADDRESS, EMAIL_ADDRESS, CREDIT_CARD_NUMBER
TPA / reinsurer / vendor feeds Bulk claim and policyholder extracts, bordereaux with full identity fields NAIC 668 third-party oversight, DPAs All entities (default), mask_mode: "hash" for analytics feeds

Reading Claims Files the Way Adjusters Write Them

Adjuster prose is its own dialect. "Clmt" for claimant, "IV" and "CV" for insured and claimant vehicles, "atty" for attorney, medical shorthand copied from records, and names introduced once then referred to by role for the rest of the file. A regex engine sees noise; a context-aware model sees that "spoke w/ clmt Rodriguez re: lumbar MRI, cb 555-0182" contains a person, a medical finding, and a callback number.

Detection quality on this dialect is what separates a useful scan from an unusable one. Our transformer NER models handle abbreviations, truncations, OCR artifacts from scanned police reports, and mixed-language files — with a confidence score per entity so that automated redaction can run strict while discovery sweeps run broad. For scanned documents, pair the API with OCR as described in the document and PDF scanning guide; for older license formats, our driver's license detection guide covers state-by-state patterns.

Insurance PII Detection Use Cases

Where carriers, TPAs, and insurtechs put detection to work across the policy lifecycle

1

Claim File Inventory & Redaction Prep

Scan claim documents and diary notes as they are created or migrated, building a per-file index of which identifiers appear where. When a file must be produced — to opposing counsel, a DOI examiner, or the insured under an access request — redaction teams start from a machine-generated map with offsets instead of reading every page.

Input
Clmt Ana Rodriguez DOB 03/22/1985 struck by IV. Clmt SSN 456-78-9012 for lien. Witness Bob Feld 555-018-2231.
Detected & Masked
Clmt [PERSON_NAME] DOB [DATE_OF_BIRTH] struck by IV. Clmt SSN [SSN] for lien. Witness [PERSON_NAME] [PHONE_NUMBER].
2

Medical Records in Injury Claims

Bodily-injury and workers' comp files pull in provider records dense with PHI. Detection classifies MRNs, diagnoses, treatments, and insurance IDs, letting you enforce PHI-grade access controls on exactly the documents that need them — and strip clinical identifiers from data used for reserving models and severity analytics.

Input
St. Mary's MRN 8834412: pt Daniel Ochoa, dx L4-L5 herniation, rx oxycodone 5mg, member ID BCB-88213377
Detected & Masked
St. Mary's MRN [MEDICAL_RECORD_NUMBER]: pt [PERSON_NAME], dx [DIAGNOSIS], rx [PRESCRIPTION], member ID [HEALTH_INSURANCE_ID]
3

Underwriting Data Minimization

Broker submissions, loss runs, and application packets arrive with more identity data than rating requires. Scan submissions at intake and route a minimized version to pricing and analytics — full identity stays in the underwriting file of record, while models train on data with names, SSNs, and DOBs consistently hashed.

Input
Applicant Grace Lin, SSN 234-56-7890, DOB 09/14/1971, 3 prior claims, requested limit $2M
Detected & Masked
Applicant [PERSON_NAME], SSN [SSN], DOB [DATE_OF_BIRTH], 3 prior claims, requested limit $2M
4

TPA & Reinsurance Bordereaux

Bordereaux and claim extracts shared with reinsurers and TPAs rarely need full policyholder identity. Gate the export pipeline: each row's free-text fields are scanned, out-of-policy identifiers are masked or hashed, and the transfer log records what was suppressed. Third-party oversight stops being a paragraph in a contract and becomes an enforced control.

Input
CLM-20441, Marcus Webb, 88 Pine Rd Dayton OH, reserve 45000, atty demand received
Detected & Masked
CLM-20441, [PERSON_NAME], [ADDRESS], reserve 45000, atty demand received
5

SIU & Fraud Analytics

Special investigation units mine claim narratives for staged-accident and provider-fraud patterns. Detection with mask_mode: "hash" lets analysts link the same phone number or provider across claims without exposing raw identity in the analytics environment — re-identification stays a controlled step inside the SIU case system.

Input
Same chiro clinic 555-441-0092 appears in claims by T. Vaughn and R. Vaughn, both DOB 1990
Detected & Masked
Same chiro clinic [PHONE_a91f] appears in claims by [NAME_2c8b] and [NAME_77e1], both DOB [DOB_hash]
6

Policyholder Chat & Email Channels

Customers email photos of licenses and type card numbers into chat. Scanning inbound messages masks payment data before transcripts reach the CRM, and flags conversations where PHI or SSNs appeared so retention rules can be applied. Our email PII scanning guide covers the mailbox-side pattern.

Input
Please charge my new card 4716 5522 0091 3348 exp 11/27 for the premium, thanks - Joan
Detected & Masked
Please charge my new card [CREDIT_CARD_NUMBER] exp [CREDIT_CARD_EXPIRATION_DATE] for the premium, thanks - [PERSON_NAME]
150+
Entity Types incl. Medical
60+
Languages for Global Books
50k
Characters per Request
<200ms
Typical Processing Time

Integrating Detection into Claims & Policy Systems

One endpoint fits every insertion point: intake webhooks, document pipelines, export gates, and nightly archive sweeps

cURL — Scan an Adjuster Note

# Detect PII and PHI in a claims diary entry before it is stored
curl -X POST https://piidetectionapi.com/api/moderate.php \
  -H "Content-Type: application/json" \
  -d '{
    "api_key": "YOUR_API_KEY",
    "api_type": "pii_detection",
    "text": "Spoke w/ clmt Ana Rodriguez re: lumbar MRI results. SSN 456-78-9012 needed for Medicare lien. CB 555-018-2231. Witness Bob Feld confirms IV ran the light at 88 Pine Rd.",
    "entities": ["PERSON_NAME", "SSN", "PHONE_NUMBER", "ADDRESS", "MEDICAL_DATA", "DIAGNOSIS"],
    "mask_mode": "replace",
    "threshold": 0.6
  }'
# Response
{
  "detected_entities": [
    {"type": "PERSON_NAME", "text": "Ana Rodriguez", "start": 14, "end": 27, "confidence": 0.97},
    {"type": "MEDICAL_DATA", "text": "lumbar MRI results", "start": 32, "end": 50, "confidence": 0.88},
    {"type": "SSN", "text": "456-78-9012", "start": 56, "end": 67, "confidence": 0.99},
    {"type": "PHONE_NUMBER", "text": "555-018-2231", "start": 101, "end": 113, "confidence": 0.98},
    {"type": "PERSON_NAME", "text": "Bob Feld", "start": 123, "end": 131, "confidence": 0.95},
    {"type": "ADDRESS", "text": "88 Pine Rd", "start": 160, "end": 170, "confidence": 0.9}
  ],
  "anonymized_text": "Spoke w/ clmt [PERSON_NAME] re: [MEDICAL_DATA]. SSN [SSN] needed for Medicare lien. CB [PHONE_NUMBER]. Witness [PERSON_NAME] confirms IV ran the light at [ADDRESS].",
  "entities_detected": 6,
  "processing_time_ms": 178,
  "mask_mode_used": "replace",
  "status": 200
}

Python — Index a Claim File for Redaction

import requests

API_URL = "https://piidetectionapi.com/api/moderate.php"

def index_document(doc_id: str, text: str) -> list:
    """Return an offset map of identifiers for a claim document."""
    entities = []
    # chunk long OCR output to stay under 50k chars/request
    for offset in range(0, len(text), 45000):
        chunk = text[offset:offset + 45000]
        resp = requests.post(API_URL, json={
            "api_key": "YOUR_API_KEY",
            "api_type": "pii_detection",
            "text": chunk,
            "entities": [
                "PERSON_NAME", "SSN", "DATE_OF_BIRTH",
                "DRIVERS_LICENSE_NUMBER", "ADDRESS",
                "MEDICAL_RECORD_NUMBER", "DIAGNOSIS",
                "HEALTH_INSURANCE_ID", "PHONE_NUMBER",
            ],
            "threshold": 0.55,   # broad sweep for discovery
        }, timeout=30)
        for e in resp.json()["detected_entities"]:
            entities.append({
                "doc_id": doc_id,
                "type": e["type"],
                "start": e["start"] + offset,
                "end": e["end"] + offset,
                "confidence": e["confidence"],
            })
    return entities

# store the map; redaction and access decisions read from it
offset_map = index_document("CLM-20441-doc7", ocr_text)
print(f"{len(offset_map)} identifiers indexed")

JavaScript — Gate a TPA Bordereau Export

// Node.js: mask identity fields in rows bound for a reinsurer
async function sanitizeRow(row) {
  const res = await fetch(
    "https://piidetectionapi.com/api/moderate.php",
    {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({
        api_key: process.env.PII_API_KEY,
        api_type: "pii_detection",
        text: row.join(" | "),
        entities: ["PERSON_NAME", "SSN", "ADDRESS",
                   "DATE_OF_BIRTH", "PHONE_NUMBER",
                   "MEDICAL_RECORD_NUMBER"],
        mask_mode: "hash",   // stable tokens keep actuarial joins
        threshold: 0.65
      })
    }
  );
  const data = await res.json();

  // audit trail for third-party oversight reviews
  auditLog.write({
    export: "reinsurer-bordereau",
    masked: data.entities_detected,
    at: new Date().toISOString()
  });
  return data.anonymized_text.split(" | ");
}

const safeRows = [];
for (const row of bordereauRows) {
  safeRows.push(await sanitizeRow(row));
}
Tuning for insurance documents

Run discovery sweeps at threshold: 0.55 to catch borderline OCR-damaged identifiers, and automated masking at 0.7+ for precision. Use custom_instruction for carrier-specific rules — e.g. "do not flag adjuster initials or claim numbers in the CLM- format as personal identifiers." Validate on your own files in the live demo, then see pricing for archive-scale volume tiers.

From Underwriting Desk to Data Lake, Safely

The modern carrier's competitive push — straight-through underwriting, AI claim triage, GenAI summarization of claim files — all depends on feeding historical documents into models and pipelines. That is precisely where privacy programs and innovation teams collide: the documents most valuable for training are the ones most saturated with personal data.

Detection resolves the collision. Run every document through the API on its way to the data lake: the pipeline persists the masked text plus the entity metadata, and raw identifiers stay only in the system of record with its existing access controls. Data scientists get realistic, structure-preserving text; the privacy office gets a provable minimization step; and GenAI features can be built on corpora that were de-identified before any model saw them — the same pattern our RAG pipeline protection guide recommends for retrieval systems.

Deployment fits carrier constraints: cloud API for speed, on-premise for books of business that cannot leave your environment, both behind the same contract. Start with a free API key, or talk through architecture with us via the contact page.

Insurance PII Detection FAQ

Questions we hear from carrier IT, claims operations, and compliance teams

Can the API distinguish PHI from ordinary PII inside a claim file?

Yes. Medical entity types — MEDICAL_RECORD_NUMBER, DIAGNOSIS, TREATMENT, PRESCRIPTION, HEALTH_INSURANCE_ID, MEDICAL_DATA — are detected as distinct classes from general identifiers like names and phone numbers. That lets you apply PHI-grade handling (restricted access, HIPAA-aligned retention) to exactly the documents and passages that contain clinical content, rather than treating every claim file as maximally sensitive.

How does detection handle scanned documents like police reports and medical records?

Run OCR first, then send the extracted text to the API. The models tolerate typical OCR noise — broken words, misread characters in numbers, inconsistent spacing — and the confidence score tells you when a detection is uncertain enough to warrant human review. For multi-hundred-page files, chunk the text and merge results using the returned offsets, as shown in the Python example above and detailed in our document scanning guide.

Does this help with NAIC Model Law 668 compliance specifically?

Directly. The model law requires a risk assessment based on the nonpublic information you hold, safeguards proportional to that assessment, and rapid event notification. Detection supplies the factual foundation for all three: a current inventory of identifier types per system, evidence that minimization and masking controls operate continuously, and pre-built indexes that let you scope an incident in hours. Your legal team defines the program; detection makes its claims true.

Witnesses and third parties appear throughout our files. Are their identifiers detected too?

Yes — detection is content-based, not profile-based. Every name, phone number, address, or license number in the text is found regardless of whose it is. This is essential for insurance, where state privacy laws grant rights to claimants and other third parties, and where producing a file without redacting a witness's phone number is a routine but real privacy failure.

Can SIU still investigate effectively if analytics data is masked?

Yes, using mask_mode: "hash". Each identifier maps to a stable token, so the same phone number or provider recurring across suspicious claims still clusters — the linkage signal fraud models need survives. When a pattern justifies escalation, re-identification happens inside the SIU case system, where access is logged and controlled. Analysts explore freely; identity exposure happens only on documented cause.

What does it cost to scan a large claims archive?

Pricing is usage-based per request, and each request carries up to 50,000 characters — roughly 15–20 pages of typical claim text — so archives cost less to sweep than teams expect. Most carriers pilot on one line of business, measure precision and recall on a labeled sample (see our accuracy guide), then schedule the full backfile. Current tiers are on the pricing page.

Related Resources

Guides for the identifiers and regulations at the heart of insurance data

See What's Hiding in Your Claim Files

Paste a real adjuster note or OCR'd claim document into the live demo and watch every identifier surface with type, offsets, and confidence. NAIC-ready evidence, from day one.