Merging contact lists gives me duplicates
Three lists went in and the combined file has the same people more than once, often with slightly different spellings, capitalisation, or company names.
Why it happens
- Names are not identifiers.
- Even a stable field fails on presentation.
- Each source also has its own idea of the columns.
What to do now
- Pick one field that identifies a person and merge only on that.
- Normalise before comparing: trim, lowercase, and strip any display wrapper.
- Decide which source wins per field, not per row.
- Keep a column recording where each row came from.
Next time
Mark the email field unique and merging stops being an operation you perform: rows join on it, so three partner lists produce one contact per person rather than three.
Start from the contact list template.
Related problems
“N/A” and dashes in a column that should be numbersA numeric column contains “N/A”, “-”, “n.a.”, or “(blank)” alongside real figures.My CSV columns are shifted or split in the wrong placesMost rows look right, but some have values in the wrong columns: a postcode in the country column, a phone number where the email should be.My supplier's prices are a thousand times wrongPrices from a European supplier are out by a factor of about a thousand, or occasionally by a factor of a hundred.The plus sign disappeared from my phone numbersInternational phone numbers have lost their leading plus.