Your screening system only finds what your data lets it see. With 59% of compliance professionals spending most of their time on data quality issues, the problem isn't the screening engine, it's what you're feeding it.
This checklist helps you assess and fix the data quality issues that create false positives, hide true matches, and turn AI-enhanced screening into a liability. Each item includes the regulatory basis where applicable and defines what "done" looks like.
What This Checklist Covers
This checklist focuses on the data preparation needed before customer records reach your sanctions and Name Screening engine. It aligns with the regulatory expectation that data quality is part of the screening control itself. The European Banking Authority's guidelines on restrictive measures state this explicitly: incomplete or mistaken customer data produces inaccurate outcomes even when the screening system is technically sound.
Prerequisites
Before starting this checklist, ensure you have:
- Access to all source systems feeding customer data into screening
- Authority to profile data across those systems
- A sample period covering recent acquisitions, migrations, or system changes
- Your current false positive rate and analyst review time per alert
Data Quality Checklist
1. Profile Name Field Structure
Requirement: Identify how often party names appear in non-name fields or multiple parties exist in a single name string.
Regulatory basis: SR 11-7 requires rigorous assessment of data quality and relevance in model development, including validation of data inputs.
Action: Run profiling tools against name fields to quantify:
- Joint account names entered as single strings (e.g., "Kim and Jim Tynan")
- Names embedded in address or reference fields
- Titles without given names
- Name variations across systems (Robert vs. Bob)
What good looks like: You have a quantified count of how many records contain multiple parties in one field, and you know which source systems contribute most to the problem.
2. Separate Hidden Parties Into Discrete Records
Requirement: Extract each party from joint strings and incorrect fields so the screening engine evaluates them individually.
Regulatory basis: EBA guidelines require payment service providers to assess whether data is sufficiently detailed to determine if a party is subject to restrictive measures.
Action: Apply cleansing logic that:
- Splits joint account holders into separate screening records
- Extracts names from address lines (e.g., "care of Dave Macy")
- Removes non-party content from name fields (e.g., "Bank refund")
What good looks like: Each natural person and legal entity in your customer base generates its own screening record, regardless of where that party appeared in the source data.
3. Validate and Standardize Address Data
Requirement: Resolve inconsistent street names, postal codes, and country formats that prevent accurate geographic risk rating.
Regulatory basis: Address determines geographic risk and helps distinguish customers from sanctioned parties with similar names.
Action:
- Apply address validation tools to standardize street names and postal codes
- Geocode addresses to assign consistent country codes
- Parse free-text address fields into structured components (street, city, state, postal code, country)
What good looks like: The same address entered as "United States" in one system and "USA" in another resolves to a single standardized country code. Postal codes and cities match official formats.
4. Correct Date of Birth Errors
Requirement: Identify and fix date of birth values that are missing, contradictory, or formatted inconsistently.
Regulatory basis: Date of birth is a primary attribute for distinguishing customers from sanctioned individuals who share a name.
Action:
- Flag records where date of birth is missing
- Identify contradictions between title and gender fields
- Standardize date formats across source systems
What good looks like: Every individual customer record includes a date of birth in a consistent format, and demographic fields align logically.
5. Deduplicate and Link Related Records
Requirement: Identify when the same customer appears multiple times across systems and link related parties.
Regulatory basis: Accurate risk identification requires a complete view of customer relationships, which duplicate records prevent.
Action: Apply matching and linking logic to:
- Resolve duplicate customer records (e.g., "Martha R. Parks" and "Martha Parks")
- Identify relationships between parties (e.g., "care of John Parks")
- Create a golden record that represents the single truth about each customer
What good looks like: Each customer exists once in your screening dataset, with all source system data consolidated and relationships between parties documented.
6. Document Data Lineage
Requirement: Maintain an audit trail showing how each golden record was assembled from source systems.
Regulatory basis: SR 11-7 requires validation of model inputs. Regulators expect to see how screened data was sourced, reformatted, and handled.
Action:
- Map each field in your golden record back to its source system
- Document transformation logic applied during cleansing
- Retain evidence of profiling findings and remediation decisions
What good looks like: For any customer record that generates a screening alert, you can show an auditor or examiner exactly which source systems contributed which fields and what cleansing rules were applied.
7. Validate Data Quality Against Screening Outcomes
Requirement: Measure whether data remediation reduces false positives and improves match accuracy.
Regulatory basis: Wolfsberg Group guidance states that data accuracy and completeness are central to effective and efficient screening.
Action:
- Establish baseline metrics before remediation: false positive rate, average alerts per analyst per day, match review time
- Measure the same metrics after data cleansing
- Identify remaining data quality issues in high-volume false positive categories
What good looks like: Your false positive rate decreases measurably, and you can demonstrate the reduction with before-and-after metrics tied to specific data quality improvements.
Common Mistakes
Treating data quality as an IT project. General-purpose data quality tools optimize for operational use cases like accounts receivable. Screening depends on a narrower set of attributes, party names, parsed addresses, dates of birth, that require compliance-specific remediation logic.
Cleansing data after it reaches the screening system. By that point, the hidden party was never screened. Remediation must happen before the record reaches the engine.
Assuming AI will compensate for poor data. AI models inherit the same limitations as rules-based engines. If a name is in an address field, it's invisible to both. Implementing AI without fixing underlying data quality scales the problem.
Skipping lineage documentation. Creating a golden record isn't enough. Regulators expect to see how you built it.
Next Steps
Start with profiling. You can't scope remediation work or build a business case without quantifying how often each issue occurs and how deep it runs. If your organization has been through an acquisition, integration, or system migration in the past two to three years, data quality problems are almost certain.
Once you've profiled the data, prioritize cases where a party exists in your records but not in a place the screening system can read. Those are true match risks, not just efficiency problems.
Finally, if a regulator or auditor has flagged match volume or data quality recently, treat that as evidence that your screening control is incomplete until the underlying data is fixed.



