Duplicates by matching rule · 5,100 records
Chart data
| Rule | Duplicates |
|---|---|
| Tax ID after normalisation | 200 |
| Name + postcode (no tax ID) | 75 |
| Not found (typos, no tax ID) | 25 |
Duplicates by matching rule · 5,100 records
Chart data
| Rule | Duplicates |
|---|---|
| Tax ID after normalisation | 200 |
| Name + postcode (no tax ID) | 75 |
| Not found (typos, no tax ID) | 25 |
Customer database clean-up
- Problem
- The same customer recorded in the CRM several times, with the tax ID and name written differently.
- Data
- 5,100 records including 300 hidden duplicates: tax IDs with dashes, spaces or a PL prefix, names in capitals, different legal form spellings.
- Work
- Normalised tax IDs and names, two matching rules (tax ID; name plus postcode for records without one) and a quality report with a review list.
- Result
- 275 of 300 duplicates found (91.7%), with no false matches. The other 25 are names with a typo and no tax ID — they need fuzzy matching or manual review.
- Tools
- Power Query, R