piidetectionapi.com
Home
Solutions - Fundamentals
What Is PII Detection? NER vs Regex vs Rules Accuracy, Precision & Recall PII in Test Data
Solutions - Compliance
GDPR Personal Data HIPAA PHI Detection CCPA / CPRA PCI DSS Card Data
Solutions - AI & LLM Safety
LLM Guardrails Chatbot PII Filtering RAG Pipelines
Solutions - Data Discovery & DLP
Data Loss Prevention Log File Scanning Support Tickets Email Scanning Documents & PDFs Database Discovery ETL & Streaming Pipelines
Industries - Financial
Banking Fintech Insurance
Industries - Healthcare
Healthcare Pharma & Clinical Trials Telehealth
Industries - Public Sector & Legal
Government & FOIA Law Enforcement Law Firms & eDiscovery Education (FERPA)
Industries - Technology
SaaS Platforms Cybersecurity & IR Telecommunications Gaming & Platforms
Industries - Other
HR & Recruiting Retail & E-commerce Call Centers & BPO Real Estate Travel & Hospitality Marketing & AdTech
How-to Guides - Identity & Contact
Detect Names Detect Email Addresses Detect Phone Numbers Detect Physical Addresses Detect Dates of Birth
How-to Guides - IDs & Financial
Detect SSNs Detect Passport Numbers Detect Drivers Licenses Detect Credit Card Numbers Detect Bank Accounts & IBAN
How-to Guides - Technical & Health
Detect IP & Device IDs Detect Medical Records & PHI
Resources
Pricing API Docs Supported Entities Languages About Contact Sign In Try the Live Demo Get Started
Legal & eDiscovery Solutions

PII Detection for Law Firms & eDiscovery

Locate every name, SSN, account number and address across matter files, review sets and court filings. Automate FRCP 5.2 redaction prep, protective-order compliance and client-data governance with AI-powered entity detection.

Why Legal Teams Need Automated PII Detection

Law firms are custodians of other people's secrets at industrial scale. A single commercial litigation matter can involve hundreds of thousands of documents — emails, contracts, HR files, medical records, financial statements — each potentially carrying Social Security numbers, dates of birth, account numbers, and home addresses belonging to parties, employees, and complete strangers to the case. eDiscovery multiplies the problem: collection sweeps in everything on a custodian's laptop, relevant or not.

The traditional answer — associates and contract reviewers eyeballing documents page by page — fails on both cost and consistency. Redaction studies in sanctions decisions read like a catalog of human error: an SSN missed on page 412 of an exhibit, a minor's name left in a footnote, metadata that undid a black box drawn over text. Courts have little patience for "we missed it" when Rule 5.2 or a protective order required the redaction, and clawback under FRE 502 does not un-ring the privacy bell for the person whose identifiers were published on PACER.

The PII Detection API gives litigation-support and innovation teams a programmatic first pass. Send document text to one endpoint and receive every detected identifier with its type, exact character offsets, and confidence score — machine-readable coordinates your review platform can convert directly into redaction boxes, privilege-log annotations, or breach-notification datasets. It is transformer-based NER, not a regex list, so it catches "the claimant's daughter, Emily, born in March of 2011" as readily as a formatted SSN.

Where PII Hides in a Law Practice

Review sets and productions are the highest-stakes surface. Opposing counsel receives whatever your redaction workflow lets through, and inadvertently produced third-party PII can trigger breach-notification duties independent of the litigation. Court filings are worse: anything e-filed is presumptively public, and Rule 5.2 violations sit in the public record until someone notices.

Client intake and conflicts data concentrates identity information — names, DOBs, SSNs for conflicts checks, financial details for engagement letters — in CRM and intake systems that rarely get the same security scrutiny as the document management system. Matter files accumulate PII for decades; firms performing information-governance cleanups routinely discover terabytes of legacy documents stuffed with identifiers nobody knew were there.

Internal knowledge systems are the newest risk. Firms building brief banks, precedent databases, and LLM-assisted drafting tools on top of past work product need those corpora scrubbed of client and third-party identifiers before a model or a search index memorializes them.

Each surface has a different tolerance for friction: an e-filing gate must run in seconds, a legacy-archive sweep can run for a weekend, and an intake scan has to be invisible to the client filling in the form. A single detection endpoint that scales from one interactive request to millions of batch calls lets one policy serve all three.

Detection output is litigation-support-ready

Each detected entity returns start and end character offsets. Litigation-support teams map those offsets back to page and coordinate positions in the underlying PDF or TIFF, generating draft redaction boxes for reviewer approval instead of asking reviewers to find identifiers themselves.

The Rules That Make PII Detection Mandatory

Court rules, ethical duties, and privacy statutes converge on the same requirement: know where the identifiers are before the document leaves your hands.

FRCP 5.2 & Court Redaction Rules

Federal Rule of Civil Procedure 5.2 (and its criminal and bankruptcy counterparts, Fed. R. Crim. P. 49.1 and Fed. R. Bankr. P. 9037) requires filings to truncate SSNs and taxpayer IDs to the last four digits, reduce birth dates to the year, identify minors by initials only, and truncate financial account numbers. Many state and local rules go further. Automated detection flags every instance before e-filing, including the ones buried in exhibits.

Protective Orders & Confidentiality

Stipulated protective orders routinely require redaction of personal identifiers from documents designated for filing or shared beyond outside counsel. Violating one is a sanctionable event. Entity-level detection lets you verify — document by document, with an audit trail — that produced and filed materials conform to the order's redaction categories.

Ethical Duties & Client Confidences

ABA Model Rules 1.6(c) and 1.1's technology-competence comment obligate lawyers to take reasonable measures against unauthorized disclosure of client information. When firms adopt cloud tools, AI drafting assistants, or offshore review, knowing exactly which identifiers a document contains is the precondition for deciding what may leave the firm's boundary.

GDPR, CCPA & Breach Duties

Firms holding EU or California resident data are controllers and processors in their own right. Data-subject access requests, cross-border discovery conflicts, and post-breach notification all require an inventory of whose personal data appears where in the matter file — precisely what large-scale entity detection produces. See our GDPR PII detection guide for the discovery-side analysis.

The Cost of a Missed Redaction

Law firms have become prime breach targets precisely because their files aggregate the sensitive data of many organizations and individuals at once. Beyond breach costs, redaction failures carry uniquely legal consequences: sanctions motions, disqualification fights, malpractice exposure, and bar complaints. Several high-profile incidents — improperly flattened PDF redactions revealing sealed material, exhibits filed with intact SSNs — became national news within hours of filing.

The economics also cut in favor of automation. Manual PII review runs at roughly 50 documents per reviewer-hour; an API call processes the same text in milliseconds for a fraction of a cent. Firms that automate the first pass redeploy reviewer time to judgment calls — privilege, responsiveness, designation — where human lawyers actually add value.

Clients have noticed. Corporate legal departments now send outside-counsel guidelines that ask, in writing, how the firm identifies and protects personal data inside matter files, and cyber insurers price law-firm policies on the same answers. A documented, automated detection layer converts those questionnaire items from awkward paragraphs into a one-line answer with logs behind it.

Legal Data Types We Detect

Entity coverage tuned to what actually appears in matter files, review sets, and filings

Parties & Third Parties
PERSON_NAME
Social Security Numbers
SSN, TAX_ID
Dates of Birth
DATE_OF_BIRTH, AGE
Financial Accounts
FINANCIAL_ACCOUNT_NUMBER, IBAN_CODE
Payment Card Data
CREDIT_CARD_NUMBER, CVV_NUMBER
Home Addresses
ADDRESS, CITY, ZIP_CODE
Contact Details
PHONE_NUMBER, EMAIL_ADDRESS
Medical Information
MEDICAL_DATA, DIAGNOSIS
Government IDs
PASSPORT_NUMBER, DRIVERS_LICENSE_NUMBER
Employment Records
EMPLOYMENT
Digital Identifiers
IP_ADDRESS, DEVICE_ID, URL
Sensitive Categories
RELIGION, SEXUAL_ORIENTATION

Mapping FRCP 5.2 to Detection Entities

Rule 5.2(a) names four identifier categories that must be truncated or abbreviated in federal filings unless an exemption applies. The table below maps each category — plus common protective-order additions — to the entity types you would pass in a detection request. Browse all 150+ types on the entities page.

Note that Rule 5.2 permits partial disclosure (last four of the SSN, year of birth). A typical workflow detects the full identifier, then applies the truncation your local rules require during the redaction step, using the returned offsets to edit precisely.

Redaction Requirement Rule Source API Entity Types Permitted Form in Filing
Social Security / taxpayer ID numbers FRCP 5.2(a)(1) SSN, TAX_ID Last four digits only
Birth dates FRCP 5.2(a)(2) DATE_OF_BIRTH Year of birth only
Names of minors FRCP 5.2(a)(3) PERSON_NAME + AGE context Initials only
Financial account numbers FRCP 5.2(a)(4) FINANCIAL_ACCOUNT_NUMBER, CREDIT_CARD_NUMBER, IBAN_CODE Last four digits only
Home addresses (criminal cases) Fed. R. Crim. P. 49.1(a)(5) ADDRESS, CITY, ZIP_CODE City and state only
Medical and treatment information Protective orders; HIPAA in discovery MEDICAL_DATA, DIAGNOSIS, MEDICAL_RECORD_NUMBER Redact or file under seal
Contact information of non-parties Protective orders; ESI protocols PHONE_NUMBER, EMAIL_ADDRESS, ADDRESS Redact per order terms
Driver's license & passport numbers State analogs; protective orders DRIVERS_LICENSE_NUMBER, PASSPORT_NUMBER Redact in full

Legal PII Detection Use Cases

Where entity detection slots into litigation, transactional, and firm-governance workflows

1

Production Redaction Prep

Run every document in the production queue through the API before redaction QC. Detected entities become draft redaction annotations in your review platform, and a zero-hit verification pass on the final production set documents that nothing enumerated in the ESI protocol slipped through.

Before Detection
Payroll register: Daniel Okonkwo, SSN 512-33-8974, Acct 004482913344, 1817 Fairview Ave.
After Masking
Payroll register: [PERSON_NAME], SSN [SSN], Acct [FINANCIAL_ACCOUNT_NUMBER], [ADDRESS].
2

Pre-Filing FRCP 5.2 Check

Add an automated gate to your e-filing workflow: extract text from the brief and every exhibit, scan for Rule 5.2 categories, and block filing while any full SSN, DOB, minor's name, or account number remains. The check takes seconds and removes the single most embarrassing class of filing error.

Before Detection
Exhibit C: Judgment debtor Anna Whitfield, born 06/02/1979, SSN 431-88-2205.
After Masking
Exhibit C: Judgment debtor [PERSON_NAME], born [DATE_OF_BIRTH], SSN [SSN].
3

Client Intake & Conflicts Hygiene

Scan intake notes, web-form submissions, and conflicts memos at capture time to classify what identity data the firm has just taken custody of. Detection output drives field-level access controls and retention schedules, and keeps prospective-client PII out of general-purpose email and chat.

Before Detection
New PNC: Rosa Marchetti, DOB 11/30/1965, cell 555-882-0417, re: dispute with her employer.
After Masking
New PNC: [PERSON_NAME], DOB [DATE_OF_BIRTH], cell [PHONE_NUMBER], re: dispute with her employer.
4

Privilege Review Support

Privilege logs and privilege-review QC often must describe documents without exposing personal details of employees and third parties. Masked document text lets junior reviewers, contract attorneys, and offshore teams see the substance needed for privilege calls while identifiers stay hidden until designation is complete.

Before Detection
Email: GC to HR re termination of Kevin Aldana (DOB 04/17/1990) — legal advice re severance.
After Masking
Email: GC to HR re termination of [PERSON_NAME] (DOB [DATE_OF_BIRTH]) — legal advice re severance.
5

Knowledge Systems & AI Drafting

Before past work product enters a brief bank, search index, or LLM fine-tuning corpus, strip client and third-party identifiers so institutional knowledge can be reused without re-disclosing confidences. The same gate belongs in front of prompts sent to external AI drafting tools — the pattern covered in our LLM guardrails guide.

Before Detection
Precedent: settlement agreement for claimant Marcus Yee, 88 Delancey St, paid via acct 3300127745.
After Masking
Precedent: settlement agreement for claimant [PERSON_NAME], [ADDRESS], paid via acct [FINANCIAL_ACCOUNT_NUMBER].
150+
Entity Types Detected
50k
Characters per Request
4
FRCP 5.2 Categories Covered
60+
Languages for Cross-Border Matters

Integrate Detection Into Your Review Stack

One JSON endpoint that slots behind any review platform, DMS, or filing workflow. Full reference in the API documentation and API overview.

cURL — FRCP 5.2 Pre-Filing Scan

curl -X POST https://piidetectionapi.com/api/moderate.php \
  -H "Content-Type: application/json" \
  -d '{
    "api_key": "YOUR_API_KEY",
    "api_type": "pii_detection",
    "text": "Exhibit C: Judgment debtor Anna Whitfield, born 06/02/1979, SSN 431-88-2205, account 004482913344 at First Meridian Bank.",
    "entities": ["PERSON_NAME", "DATE_OF_BIRTH", "SSN", "FINANCIAL_ACCOUNT_NUMBER"],
    "mask_mode": "replace",
    "threshold": 0.4
  }'

Python — Batch Scan a Production Set

import requests, json, pathlib

API_URL = "https://piidetectionapi.com/api/moderate.php"

PROTECTIVE_ORDER_ENTITIES = [
    "PERSON_NAME", "SSN", "TAX_ID", "DATE_OF_BIRTH",
    "FINANCIAL_ACCOUNT_NUMBER", "CREDIT_CARD_NUMBER",
    "ADDRESS", "PHONE_NUMBER", "EMAIL_ADDRESS",
    "DRIVERS_LICENSE_NUMBER", "PASSPORT_NUMBER", "MEDICAL_DATA",
]

def scan_document(doc_id: str, text: str) -> dict:
    resp = requests.post(
        API_URL,
        json={
            "api_key": "YOUR_API_KEY",
            "api_type": "pii_detection",
            "text": text[:50000],  # chunk longer documents
            "entities": PROTECTIVE_ORDER_ENTITIES,
            "mask_mode": "replace",
            "threshold": 0.4,  # recall-first: over-flag, let reviewers clear
            "custom_instruction": "Do not flag names of attorneys of record or the presiding judge.",
        },
        timeout=30,
    )
    return {"doc_id": doc_id, **resp.json()}

# Emit a redaction worksheet: one row per hit, with offsets for the review tool
for path in pathlib.Path("production_vol_003").glob("*.txt"):
    report = scan_document(path.stem, path.read_text())
    for e in report["detected_entities"]:
        print(json.dumps({
            "doc": report["doc_id"], "type": e["type"],
            "start": e["start"], "end": e["end"], "conf": e["confidence"],
        }))

JavaScript — Intake Form Gate

// Classify identifiers in a web intake submission before it is stored or emailed
async function gateIntakeSubmission(formText) {
  const resp = await fetch("https://piidetectionapi.com/api/moderate.php", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      api_key: process.env.PII_API_KEY,
      api_type: "pii_detection",
      text: formText,
      entities: [
        "PERSON_NAME", "SSN", "DATE_OF_BIRTH",
        "PHONE_NUMBER", "EMAIL_ADDRESS", "ADDRESS"
      ],
      mask_mode: "redact"  // strip identifiers from the copy routed to shared inboxes
    })
  });

  const data = await resp.json();
  return {
    piiTypes: [...new Set(data.detected_entities.map(e => e.type))],
    sanitized: data.anonymized_text,
    count: data.entities_detected
  };
}

gateIntakeSubmission(
  "New PNC: Rosa Marchetti, DOB 11/30/1965, cell 555-882-0417, employment dispute."
).then(console.log);

Best Practices for Legal Deployments

Detect broadly, redact narrowly. Scan with the full protective-order entity list at a low threshold (0.4), then apply your jurisdiction's specific truncation rules at redaction time. Over-detection costs a reviewer a second per hit; under-detection costs a motion.

Use custom instructions to encode matter context. The custom_instruction field accepts plain-language exclusions such as "do not flag names of counsel of record, the court, or the testifying expert." This keeps the hit list focused on the third-party identifiers that actually require action.

Keep an audit trail per document. Store the full detected_entities array with the document record. When redaction adequacy is challenged, you can show what was detected, at what confidence, and what the reviewer did with each hit — a defensibility posture no manual-only process can match.

Verify the final artifact. Re-scan the text layer of the redacted PDF before it goes out. A zero-entity result on the outbound file is the cheapest malpractice insurance a firm can buy, and it catches the classic flattening failures where a black rectangle hides text that is still extractable underneath.

Redact the text layer, not just the image

Courts have repeatedly seen "redacted" filings where copy-paste revealed the hidden text. Detection offsets operate on the extracted text itself, which forces your pipeline to confront the text layer directly — scan, redact at the character level, regenerate, then re-scan to confirm zero residual entities.

Legal PII Detection FAQ

Common questions from litigation support, innovation, and risk teams

Can the API drive redactions in our eDiscovery review platform?

Yes. The API is platform-agnostic: extract text from the document (native, OCR, or the platform's extracted-text field), send it to the endpoint, and use the returned entity types and character offsets to create draft redaction annotations through your platform's API — Relativity, Everlaw, DISCO, Reveal, and similar tools all support programmatic annotation. Reviewers then approve, adjust, or reject each proposed redaction rather than hunting for identifiers manually.

How does detection handle scanned documents and OCR noise?

Litigation documents are full of imperfect OCR — "S0cial Security N0: 431-88-22O5" with zeros for O's and vice versa. Because detection is context-aware rather than pattern-locked, entities survive moderate OCR corruption: the model reads the surrounding words and flags the identifier even when a strict regex would fail. For heavily degraded scans, we recommend a lower threshold plus targeted re-OCR of pages where detection density looks anomalously low. See the document and PDF scanning guide for the full OCR pipeline.

Is sending client documents to a cloud API consistent with our confidentiality duties?

The cloud API runs on GDPR-native certified infrastructure, request content is processed transiently and not used for model training, and data-processing agreements are available. That satisfies most firms' outside-vendor analysis under Model Rule 1.6(c). Firms and matters with stricter obligations — sealed materials, ITAR, banking secrecy — can deploy the same detection engine on-premise or in the firm's private cloud so document text never leaves your environment. Contact us via the contact page to discuss deployment options.

Does the API distinguish minors' names for Rule 5.2(a)(3)?

The API detects all person names plus AGE and DATE_OF_BIRTH entities. Minor status is a legal conclusion drawn from context, so the recommended pattern is: flag every name, and where an associated age or birth date entity indicates a person under 18, route that document to a reviewer for the initials-only treatment. The co-occurrence of name and age entities in the response makes that routing rule a few lines of code.

Can we scan documents in languages other than English for cross-border matters?

Yes — detection covers more than 60 languages, and mixed-language documents (a German email chain with English attachments, say) are handled in a single request. This matters for cross-border discovery, where GDPR and blocking statutes often require identifying and minimizing EU personal data before transfer to a US review. The supported languages page lists current coverage.

What does it cost to scan a large review set?

Pricing is per-request and volume-tiered, so cost scales with document count rather than reviewer hours — typically orders of magnitude below manual first-pass PII review. A hundred-thousand-document scan is a batch job measured in hours and hundreds of dollars, not weeks and six figures. See the pricing page for tiers, or run a sample through the live demo to estimate hit density on your own data.

Related Resources

Go deeper on document scanning and identifier-specific detection

Make Missed Redactions a Solved Problem

Paste a sample filing or production document into the live demo and see every Rule 5.2 identifier flagged with offsets — then wire the same call into your review workflow.