Duplicate & Fuzzy Match Scout

What This Prompt Does:

Clean your mailing lists or databases of duplicate entries. This prompt identifies not just exact matches, but ‘fuzzy’ duplicates caused by typos or slightly different naming conventions.

Key Tips for Best Results:

  • Define the sensitivity level (strict vs. loose matching).
  • Mention if you want to keep the ‘oldest’ or ‘newest’ record.
  • Useful for combining customer lists from different sources.

Description

About This Prompt: Duplicate & Fuzzy Match Scout

Optimized for:
ChatGPT
Claude
Gemini
Grok
Verified precision.

How To Use The Prompt:

  1. Copy the entire prompt from the window and paste it directly into ChatGPT, Claude, or Gemini.
  2. Locate and replace the main placeholders with your specific details: [COLUMN LIST], [MATCH SENSITIVITY], and [RAW DATA].
  3. Input Example: Use the following format for your input:

Example Input: Columns: Name, Email, Address, Sensitivity: High (typo-resistant), Data: [Paste Here].

Additional Information:

Duplicates waste storage and confuse analytics. This scout finds the subtle ones that standard ‘Distinct’ filters miss.

This prompt provides the foundation to:

  • Fuzzy Logic: Finds ‘John Doe’ vs ‘Jon Doe’ automatically.
  • Data Cleaning: Suggests which records to merge for a ‘Golden Record’.
  • Cost Savings: Reduces database bloat and prevents duplicate communication.

Safety Note: As a professional colleague, always review AI-generated assessments for potential bias or inaccuracies before making final business or management decisions.

PROMPT WINDOW

You are a Database Administrator specializing in data deduplication. Analyze the [RAW DATA] focusing on the following [COLUMN LIST].

Your goal is to find exact duplicates and ‘fuzzy’ matches (typos, varied spellings) using a [MATCH SENSITIVITY] approach. Output a report listing the duplicate clusters found and provide a clear recommendation on which record should be preserved as the primary entry based on data completeness.