piidetectionapi.com
Home
Solutions - Fundamentals
What Is PII Detection? NER vs Regex vs Rules Accuracy, Precision & Recall PII in Test Data
Solutions - Compliance
GDPR Personal Data HIPAA PHI Detection CCPA / CPRA PCI DSS Card Data
Solutions - AI & LLM Safety
LLM Guardrails Chatbot PII Filtering RAG Pipelines
Solutions - Data Discovery & DLP
Data Loss Prevention Log File Scanning Support Tickets Email Scanning Documents & PDFs Database Discovery ETL & Streaming Pipelines
Industries - Financial
Banking Fintech Insurance
Industries - Healthcare
Healthcare Pharma & Clinical Trials Telehealth
Industries - Public Sector & Legal
Government & FOIA Law Enforcement Law Firms & eDiscovery Education (FERPA)
Industries - Technology
SaaS Platforms Cybersecurity & IR Telecommunications Gaming & Platforms
Industries - Other
HR & Recruiting Retail & E-commerce Call Centers & BPO Real Estate Travel & Hospitality Marketing & AdTech
How-to Guides - Identity & Contact
Detect Names Detect Email Addresses Detect Phone Numbers Detect Physical Addresses Detect Dates of Birth
How-to Guides - IDs & Financial
Detect SSNs Detect Passport Numbers Detect Drivers Licenses Detect Credit Card Numbers Detect Bank Accounts & IBAN
How-to Guides - Technical & Health
Detect IP & Device IDs Detect Medical Records & PHI
Resources
Pricing API Docs Supported Entities Languages About Contact Sign In Try the Live Demo Get Started
Telehealth Solutions

PII Detection for Telehealth

Automatically detect and classify PHI and PII in video-visit transcripts, chat-based care conversations, remote monitoring feeds, and patient intake forms. Build HIPAA-compliant virtual care on an AI detection engine with 150+ entity types.

Why Telehealth Needs Purpose-Built PII Detection

Telehealth turned the clinical encounter into a stream of unstructured data. A single video visit produces an audio recording, an automatic speech-recognition transcript, chat sidebar messages, screen-share captures, an after-visit summary, and platform metadata such as IP addresses and device identifiers. Every one of those artifacts can contain protected health information, and none of it arrives in the neat, structured fields that traditional data-protection tooling expects.

The scale is unforgiving. A mid-sized virtual care group running a few hundred visits a day generates millions of words of transcript text each month. Asynchronous chat-based care programs add a continuous flow of patient messages describing symptoms, medications, and life circumstances. Remote patient monitoring pushes device readings tagged with names, member IDs, and timestamps into analytics pipelines around the clock. Manually reviewing this volume for PHI before it reaches vendors, data warehouses, QA teams, or AI models is impossible.

The PII Detection API solves the discovery problem at the API layer. Send any telehealth text — a diarized visit transcript, a triage chat, an intake form, a device alert — to a single endpoint and receive back every detected entity with its type, exact character offsets, and a confidence score, plus an optionally masked version of the input. Detection is powered by transformer-based NER that understands clinical conversation, not brittle regex lists, so "the patient's sister, Angela, also takes metformin" is caught even though no pattern would match it.

Where PHI Hides in Virtual Care

Video-visit transcripts are the densest PHI surface in telehealth. Patients state their full name and date of birth for identity verification in the first thirty seconds, then discuss diagnoses, prescriptions, family history, and home addresses in free-flowing speech. ASR errors make things harder: "John Kowalski" may be transcribed as "John Koval ski", defeating exact-match tools while remaining obvious to a context-aware model.

Chat-based care mixes identifiers into casual language — "this is Maria, DOB 3/14/1988, my insurance ID is BCB-448291022, and the rash from the amoxicillin is back." Messages also carry phone numbers for callbacks, pharmacy addresses, and photos' file names that embed patient names.

Remote monitoring data looks structured but leaks PII through free-text fields: device alert notes, nurse annotations, and escalation emails routinely contain names, medical record numbers, and readings tied to timestamps that function as identifiers. Patient intake forms concentrate nearly all 18 HIPAA identifiers into one document — name, address, phone, email, SSN, insurance ID, emergency contacts, and medication lists.

Session metadata counts too

HIPAA's identifier list explicitly includes IP addresses and device identifiers. Telehealth platform logs that tie an IP address or device ID to a visit are PHI. The API detects IP_ADDRESS, DEVICE_ID, MAC_ADDRESS, and URL entities in log text so your observability stack stays out of scope.

HIPAA Compliance for Virtual Care

Virtual care lost its pandemic-era enforcement discretion. These are the obligations telehealth data flows must satisfy today.

HIPAA Privacy Rule

The Privacy Rule governs every use and disclosure of PHI from a telehealth encounter — transcripts, chat logs, and recordings included. De-identification under the Safe Harbor method requires removing all 18 identifier categories. Automated detection with character-level offsets gives you a defensible, auditable process for locating those identifiers before any secondary use of visit data.

HIPAA Security Rule & BAAs

Every vendor that touches identifiable visit data — transcription services, analytics platforms, cloud LLM providers — needs a business associate agreement. Detecting and masking PHI before data leaves your boundary shrinks the BAA surface dramatically: a QA vendor reviewing masked transcripts never becomes a business associate for that data flow.

42 CFR Part 2

Tele-behavioral health and virtual substance-use-disorder programs face stricter-than-HIPAA confidentiality rules under 42 CFR Part 2. Records that identify a patient as receiving SUD treatment require specific consent to disclose. Detecting names, contact details, and treatment references in therapy transcripts is a precondition for any analytics or supervision workflow in these programs.

State Laws & FTC Enforcement

State telehealth privacy statutes, consumer health data laws such as Washington's My Health My Data Act, and the FTC's Health Breach Notification Rule reach telehealth apps even where HIPAA does not. Direct-to-consumer platforms have faced FTC actions for sharing user health data with ad networks — data flows that PII detection at the egress point would have flagged.

The Cost of Getting It Wrong

Healthcare has held the top spot for breach costs for more than a decade, averaging over $9 million per incident — and telehealth breaches are disproportionately visible because they involve recordings and transcripts of intimate conversations. OCR settlements for impermissible disclosures regularly run into seven figures, and 42 CFR Part 2 violations now carry HIPAA-equivalent penalties.

The quieter cost is stalled innovation. Teams that cannot prove their transcripts are de-identified cannot use them: no model training, no quality analytics, no outsourced coding review. Reliable PHI detection is what converts an archive of risky recordings into a usable clinical data asset.

There is also a contractual dimension. Health systems buying telehealth platforms now ask pointed diligence questions about where transcripts go, which subprocessors see them, and how de-identification is verified. Being able to answer with a documented, entity-level detection step — complete with confidence scores and audit logs — shortens security reviews and wins enterprise deals that a vague "we encrypt everything" answer loses.

Telehealth Data Types We Detect

Purpose-built entity coverage for the identifiers that appear in virtual care conversations and records

Patient & Family Names
PERSON_NAME
Medical Record Numbers
MEDICAL_RECORD_NUMBER
Insurance Member IDs
HEALTH_INSURANCE_ID
Diagnoses & Conditions
DIAGNOSIS, MEDICAL_DATA
Prescriptions
PRESCRIPTION, TREATMENT
Dates of Birth & Ages
DATE_OF_BIRTH, AGE
Contact Details
PHONE_NUMBER, EMAIL_ADDRESS
Home Addresses
ADDRESS, ZIP_CODE
Session Identifiers
IP_ADDRESS, DEVICE_ID
Government IDs
SSN, NATIONAL_ID
Biometric References
BIOMETRIC_DATA, BLOOD_TYPE
Employment & Demographics
EMPLOYMENT, ETHNIC_GROUP

Mapping HIPAA Identifiers to API Entities

HIPAA Safe Harbor de-identification enumerates 18 identifier categories. The table below maps the ones most common in telehealth artifacts to the entity types you pass in the entities array of a detection request. For the full 18-identifier walkthrough, see our HIPAA PHI detection guide, and browse every supported type on the entities page.

Because detection is context-aware, the same model distinguishes a clinician's name from a drug name, and a callback number from a blood-pressure reading — a distinction pure pattern-matching cannot make in conversational transcripts.

HIPAA Identifier API Entity Types Typical Telehealth Occurrence
Names PERSON_NAME Identity verification at visit start; family members mentioned in history
Geographic subdivisions smaller than state ADDRESS, CITY, ZIP_CODE Pharmacy selection, home-health scheduling, intake forms
Dates related to an individual DATE_OF_BIRTH, AGE DOB confirmation on every visit; ages over 89 flagged for aggregation
Telephone / fax numbers PHONE_NUMBER Callback numbers in chat, emergency contacts on intake
Email addresses EMAIL_ADDRESS Portal registration, visit-summary delivery preferences
Social Security numbers SSN Intake and billing forms, insurance eligibility checks
Medical record numbers MEDICAL_RECORD_NUMBER EHR cross-references in nurse notes and escalations
Health plan beneficiary numbers HEALTH_INSURANCE_ID Member IDs quoted in chat for billing questions
Device identifiers & serial numbers DEVICE_ID, SERIAL_NUMBER, IMEI RPM device pairing records, glucometer and BP-cuff serials
IP addresses & URLs IP_ADDRESS, URL Video platform session logs, join-link audit trails

Telehealth PII Detection Use Cases

How virtual care teams put entity-level detection to work across the visit lifecycle

1

Visit Transcript De-Identification

Run every ASR transcript through the API before it lands in your analytics warehouse or QA review queue. Detected entities are replaced with typed placeholders while the clinical narrative — symptoms, dosages, care plans — stays intact for quality scoring and outcome research.

Before Detection
Dr. Patel: Can you confirm your date of birth? Patient: March 14, 1988, and I'm still at 2214 Birchwood Lane.
After Masking
Dr. [PERSON_NAME]: Can you confirm your date of birth? Patient: [DATE_OF_BIRTH], and I'm still at [ADDRESS].
2

Real-Time Chat-Based Care Filtering

Screen asynchronous care messages in-flight before they are routed to third-party triage tools or LLM-assisted drafting. Sub-second responses let you mask identifiers without adding perceptible latency to the conversation, the same pattern described in our chatbot PII filtering guide.

Before Detection
Hi, it's Maria Delgado, member ID BCB-448291022 — the amoxicillin rash is back, call me at 555-201-8834.
After Masking
Hi, it's [PERSON_NAME], member ID [HEALTH_INSURANCE_ID] — the amoxicillin rash is back, call me at [PHONE_NUMBER].
3

Remote Monitoring Pipelines

Scan RPM alert notes, nurse annotations, and escalation emails as they flow into dashboards and data lakes. Structured vitals pass through untouched; embedded names, MRNs, and device serials are caught and masked so population-health analytics stay de-identified.

Before Detection
ALERT: BP 182/104 for Robert Chen, MRN 00483920, device SN GLU-99213 — escalate to on-call.
After Masking
ALERT: BP 182/104 for [PERSON_NAME], MRN [MEDICAL_RECORD_NUMBER], device SN [SERIAL_NUMBER] — escalate to on-call.
4

Patient Intake Form Scanning

Intake forms concentrate the densest PII on your platform. Detect and classify every identifier at ingestion to drive field-level encryption, retention policies, and data maps for HIPAA and state privacy audits — before the form text is ever copied into downstream systems.

Before Detection
Name: Alicia Fontaine, SSN 527-44-9081, Insurance: Aetna W220-4491, Emergency contact: 555-903-2214
After Masking
Name: [PERSON_NAME], SSN [SSN], Insurance: Aetna [HEALTH_INSURANCE_ID], Emergency contact: [PHONE_NUMBER]
5

Training AI Scribes on Clean Data

Ambient documentation and AI scribe models need thousands of real visit transcripts for fine-tuning and evaluation. Detecting and masking PHI first lets your ML team build on genuine clinical dialogue without ever holding identifiable data in the training environment.

Before Detection
Follow up with Mrs. Okafor at [email protected] regarding the lisinopril titration next Tuesday.
After Masking
Follow up with Mrs. [PERSON_NAME] at [EMAIL_ADDRESS] regarding the lisinopril titration next Tuesday.
150+
Entity Types Detected
60+
Languages Supported
<300ms
Typical Chat Message Latency
18/18
HIPAA Identifier Categories Covered

Integrate PHI Detection Into Your Telehealth Stack

One JSON endpoint, three lines of integration. Full request and response reference in the API documentation and API overview.

cURL — Scan a Visit Transcript Snippet

curl -X POST https://piidetectionapi.com/api/moderate.php \
  -H "Content-Type: application/json" \
  -d '{
    "api_key": "YOUR_API_KEY",
    "api_type": "pii_detection",
    "text": "Dr. Patel: Confirming this is Maria Delgado, DOB 03/14/1988, MRN 00483920? Patient: Yes, and my new number is 555-201-8834.",
    "entities": ["PERSON_NAME", "DATE_OF_BIRTH", "MEDICAL_RECORD_NUMBER", "PHONE_NUMBER"],
    "mask_mode": "replace",
    "threshold": 0.5
  }'

Python — Batch De-Identify ASR Transcripts

import requests

API_URL = "https://piidetectionapi.com/api/moderate.php"

TELEHEALTH_ENTITIES = [
    "PERSON_NAME", "DATE_OF_BIRTH", "AGE", "PHONE_NUMBER",
    "EMAIL_ADDRESS", "ADDRESS", "ZIP_CODE", "SSN",
    "MEDICAL_RECORD_NUMBER", "HEALTH_INSURANCE_ID",
    "DEVICE_ID", "SERIAL_NUMBER", "IP_ADDRESS",
]

def deidentify_transcript(transcript_text: str) -> dict:
    # Each request accepts up to 50,000 characters; chunk longer visits
    resp = requests.post(
        API_URL,
        json={
            "api_key": "YOUR_API_KEY",
            "api_type": "pii_detection",
            "text": transcript_text,
            "entities": TELEHEALTH_ENTITIES,
            "mask_mode": "replace",
            "threshold": 0.4,  # recall-first for Safe Harbor workflows
            "custom_instruction": "Do not mask drug names, dosages, or vital-sign readings.",
        },
        timeout=30,
    )
    return resp.json()

result = deidentify_transcript(open("visit_20260825.txt").read())
for e in result["detected_entities"]:
    print(e["type"], e["text"], e["start"], e["end"], e["confidence"])
print(result["anonymized_text"])

JavaScript — Real-Time Chat Message Filter

// Filter each care-chat message before routing to downstream tools
async function filterCareMessage(message) {
  const resp = await fetch("https://piidetectionapi.com/api/moderate.php", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      api_key: process.env.PII_API_KEY,
      api_type: "pii_detection",
      text: message,
      entities: [
        "PERSON_NAME", "PHONE_NUMBER", "EMAIL_ADDRESS",
        "HEALTH_INSURANCE_ID", "DATE_OF_BIRTH", "ADDRESS"
      ],
      mask_mode: "hash"  // consistent hashes keep threads linkable per patient
    })
  });

  const data = await resp.json();
  console.log(`Masked ${data.entities_detected} entities in ${data.processing_time_ms}ms`);
  return data.anonymized_text;
}

filterCareMessage(
  "Hi, it's Maria Delgado, member ID BCB-448291022 — call me at 555-201-8834."
).then(console.log);

Best Practices for Telehealth Deployments

Tune the threshold to the destination. For Safe Harbor de-identification feeding research or model training, lower threshold to 0.4 and accept a few over-masks — recall matters more than precision. For live chat where over-masking hurts usability, keep the default 0.5 and route low-confidence hits to human review instead.

Preserve clinical utility with custom instructions. The custom_instruction field lets you say, in plain language, "do not mask medication names, dosages, or vital signs." You keep the analytical signal that makes transcripts valuable while removing everything that identifies the patient.

Choose mask modes deliberately. Use replace for human-readable QA transcripts, hash when the same patient must remain linkable across a longitudinal chat thread without being identifiable, and redact for exports leaving your trust boundary entirely.

Design for auditability from day one. Persist the detected_entities array — types, offsets, and confidence scores — alongside each processed artifact. When an auditor or a customer's security team asks how a particular transcript was de-identified, you can show exactly which identifiers were found, where they sat in the text, and how confident the model was, rather than pointing at a policy document. Start with the getting-started guide to wire this into a proof of concept in an afternoon.

Handle recordings through their transcripts. For audio and video assets, the practical pattern is to detect on the ASR transcript, use the returned character offsets to locate identifier utterances in the time-aligned transcript, and then bleep or cut the corresponding media segments. The same detection call therefore drives both text and media redaction without a second tool.

Detection precedes every disclosure

Place the API call at the last point inside your compliance boundary — before the transcription vendor webhook fires, before the warehouse load, before the LLM prompt is assembled. Detection after disclosure is documentation of a breach, not prevention of one. On-premise deployment is available for programs that cannot send raw PHI to any external endpoint.

Telehealth PII Detection FAQ

Common questions from virtual care engineering and compliance teams

How accurate is detection on conversational, ASR-generated transcripts?

The models are transformer-based and evaluated on conversational text, so they handle disfluencies, diarization tags, and the informal phrasing typical of speech-to-text output. Crucially, context-aware NER tolerates ASR spelling errors: a misrecognized surname still sits in a name-shaped context and is flagged as PERSON_NAME even when no dictionary or regex would match it. For de-identification pipelines we recommend a lower threshold (around 0.4) plus spot-check sampling to validate recall on your own audio mix.

Does masking API output satisfy HIPAA Safe Harbor de-identification?

Safe Harbor requires removal of all 18 identifier categories and no actual knowledge that the remainder could identify the individual. The API detects entity types covering all 18 categories, and the offsets and confidence scores give you an auditable record of what was found and removed. Most organizations pair automated detection with a documented sampling review to establish the "no actual knowledge" prong, or use detection output as evidence in an Expert Determination process. Your privacy officer or expert makes the final call; the API supplies the coverage and the paper trail.

Is the API fast enough to filter live chat-based care conversations?

Yes. Typical chat messages return in well under half a second, which is imperceptible inside an asynchronous care workflow and acceptable even for synchronous chat. Teams commonly call the API from the message-routing service, mask identifiers, and forward the cleaned text to triage queues, analytics, or LLM drafting tools. For long visit transcripts, chunk the text under the 50,000-character request limit and process chunks in parallel.

Can we keep clinical content — drugs, symptoms, vitals — while removing identifiers?

That is the default posture we recommend. Scope the entities array to identifier types only (names, dates, contact details, MRNs, insurance IDs, device IDs) and leave MEDICAL_TERM, PRESCRIPTION, and DIAGNOSIS out of the list — or include them but add a custom_instruction such as "do not mask medication names or dosages." The result is a transcript that is useless to an attacker but fully useful to a quality analyst or a model.

What about visits conducted in languages other than English?

The API detects PII in more than 60 languages, including Spanish, Mandarin, Vietnamese, Tagalog, and Arabic — the languages most requested in US telehealth interpretation. Mixed-language transcripts (an interpreter-mediated visit, for example) are handled in a single request without language hints. See the supported languages page for the full list.

Will you sign a BAA, and is on-premise deployment available?

Yes on both counts. The cloud API runs on GDPR-native certified infrastructure and business associate agreements are available on paid plans for covered entities and business associates. Telehealth platforms with stricter data-residency or Part 2 obligations can deploy the detection engine on-premise or in a private cloud, so raw PHI never leaves your environment. Contact us via the contact page to scope a deployment.

Related Resources

Go deeper on PHI detection for healthcare and virtual care

Ready to De-Identify Your Virtual Care Data?

Paste a sample transcript into the live demo and watch every identifier light up in seconds — then scale the same call to millions of visits.