Independent editorial research

Screening Data Quality and False Positives

An evidence-led explanation of how source structure, deduplication, identifiers, customer data, updates, and matching settings shape screening candidates and false-positive workload.

By AML Tech Reviews Editorial TeamPublished Updated

“Mohamed Ali” can produce many screening candidates. A passport number can narrow them quickly—if both the customer record and the reference record contain it, the formats are compatible, and the investigator can see the comparison. That short example shows why false-positive workload is a data problem as much as a matching problem.

A false positive is a candidate that the reviewing institution determines is not the relevant listed or risk-associated subject under its procedure. The term is often used loosely for any alert closed without escalation. Buyers should define it precisely because duplicate alerts, stale prior decisions, out-of-scope records, and low-quality candidates have different causes.

Source records need structure

Reference data can contain names, aliases, scripts, entity type, date and place of birth, nationality, addresses, registration numbers, passport or tax identifiers, roles, relationships, programs, vessels, and source history. A flat name string discards much of that context.

OFAC’s FAQ on determining a valid match tells users to compare entity type, the full name, address, nationality, passport or tax identifier, place and date of birth, former names, and aliases. The guidance does not say that every field will be present. It shows that a match decision should use the available identifying information rather than a name score alone.

Structure lets a system treat fields according to meaning. A date is not just another string. An identifier may need issuer and country. A corporate suffix may carry less distinguishing weight than a registration number. A place can have several spellings or levels of geography. If normalization erases the original value, an investigator may lose the ability to explain a candidate.

Deduplication changes workload

One person may appear in several official lists, source files, aliases, or data-provider records. Deduplication can consolidate that material into a coherent subject profile and reduce repeated review. It can also make a serious error: two distinct people can be merged because they share a name and limited attributes.

The safer design preserves source records and shows why they were linked. A consolidated profile should retain source identifiers, list and program information, effective dates, aliases, conflicting attributes, and change history. The system should be able to split a mistaken merge without losing decisions.

Deduplication also applies to customer data and alerts. Duplicate customer profiles can create repeat candidates. Rescreening every unchanged customer against every unchanged record can reopen resolved work unless prior decisions are reused under controlled rules. Reuse must respond to material changes in customer data, reference data, policy, or matching configuration.

Identifiers improve discrimination

Names are common screening inputs but often weak discriminators. Some sanctions controls also evaluate non-name indicators such as geographic terms, bank identifiers, securities identifiers, vessel or aircraft details, account data, and transaction text. Dates of birth, passport numbers, national identifiers, company registration numbers, addresses, nationality, and entity type can help confirm or contradict a candidate. The Wolfsberg sanctions guidance describes screening operational data against lists of names and other indicators. More data is not automatically better. An incorrect date can suppress a true candidate if the system treats it as a hard mismatch.

Buyers should ask how each field contributes, what happens when it is absent or conflicting, and whether the behavior changes by list, entity type, or jurisdiction. They should also inspect data collection upstream. A screening engine cannot compare a date of birth that onboarding never captured.

Updates are part of data quality

Reference data changes when authorities add, amend, or remove records. Customer data changes when names, ownership, addresses, documents, or risk information are updated. Data quality therefore includes time: when the source changed, when the provider received it, when the system processed it, and when affected customers or payments were evaluated.

The Wolfsberg sanctions guidance discusses reference data, timing, screening technology, alert investigation, testing, and quality assurance as parts of the control. Procurement should test late, failed, duplicate, corrected, and removed updates—not only a clean initial load.

Stable source identifiers help determine what changed. Without them, a corrected record may look like a new subject and reopen work. Update reconciliation should prove which files or messages arrived, which records were accepted or rejected, which population was screened, and which exceptions remain.

Matching settings make trade-offs

Matching methods can compare spelling, phonetics, token order, aliases, transliterations, abbreviations, dates, identifiers, and other context. OFAC explains that its public Sanctions List Search score represents name similarity and that lower scores indicate potential matches. Its separate calculation FAQ describes specific string and phonetic algorithms used by that public tool.

Those pages explain OFAC’s search service, not a universal screening design. A buyer should not copy one threshold into every customer population or payment field. Lowering a name threshold may find more variations and create more candidates. Raising it may reduce workload and miss relevant variations. Secondary identifiers can improve discrimination, but only if field quality and matching behavior are understood.

Measure the trade-off on labelled, representative data. Report candidate counts, relevant-match capture, false positives under the agreed definition, results by segment, and unresolved cases. A reduction percentage is meaningless without the baseline, population, configuration, time period, and treatment of missed relevant matches.

Procurement implications

Evaluate the reference-data service, customer-data pipeline, matching engine, and investigation workflow as separate layers. Request source inventories, sample records, provenance, deduplication logic, update logs, normalization rules, field treatment, configuration history, and alert explanations. Test common names, sparse records, entities, non-Latin scripts, aliases, close dates, conflicting identifiers, source corrections, and full-population rescreening.

Cost models should include data licenses, screening events, rescreening, investigator workload, quality assurance, data remediation, and change management. A lower subscription price can be outweighed by repeated review. A low alert count can hide weak coverage. Use paired detection and workload measures.

Evidence limitations

No public dataset fully represents an institution’s customers, jurisdictions, scripts, and policy. Labelled cases can be uncertain, and past decisions may contain error. Official records can be incomplete. Customer identifiers can be missing, stale, or wrong. Test results therefore describe a bounded sample and configuration, not permanent screening accuracy.

The defensible objective is not zero false positives. It is a risk-based control that can find relevant candidates, use available context, explain decisions, manage updates, and measure both missed risk and unnecessary work over time.