Duplicate & Fuzzy Match Scout
What This Prompt Does:
Clean your mailing lists or databases of duplicate entries. This prompt identifies not just exact matches, but ‘fuzzy’ duplicates caused by typos or slightly different naming conventions.
Key Tips for Best Results:
- Define the sensitivity level (strict vs. loose matching).
- Mention if you want to keep the ‘oldest’ or ‘newest’ record.
- Useful for combining customer lists from different sources.
- Description
Description
About This Prompt: Duplicate & Fuzzy Match Scout
Verified precision.
How To Use The Prompt:
- Copy the entire prompt from the window and paste it directly into ChatGPT, Claude, or Gemini.
- Locate and replace the main placeholders with your specific details:
[COLUMN LIST],[MATCH SENSITIVITY], and[RAW DATA]. - Input Example: Use the following format for your input:
Example Input:
Columns: Name, Email, Address, Sensitivity: High (typo-resistant), Data: [Paste Here].
Additional Information:
Duplicates waste storage and confuse analytics. This scout finds the subtle ones that standard ‘Distinct’ filters miss.
This prompt provides the foundation to:
- Fuzzy Logic: Finds ‘John Doe’ vs ‘Jon Doe’ automatically.
- Data Cleaning: Suggests which records to merge for a ‘Golden Record’.
- Cost Savings: Reduces database bloat and prevents duplicate communication.
Safety Note: As a professional colleague, always review AI-generated assessments for potential bias or inaccuracies before making final business or management decisions.
You are a Database Administrator specializing in data deduplication. Analyze the [RAW DATA] focusing on the following [COLUMN LIST].
Your goal is to find exact duplicates and ‘fuzzy’ matches (typos, varied spellings) using a [MATCH SENSITIVITY] approach. Output a report listing the duplicate clusters found and provide a clear recommendation on which record should be preserved as the primary entry based on data completeness.
