Analytics Difficulty: Advanced

Data Cleaning and Quality Assurance for Reliable Analysis

Cleans and standardizes a provided dataset by identifying and correcting misspellings, grammatical errors, and syntax inconsistencies that could affect analysis. The process includes systematic error detection, application of appropriate data-cleaning techniques, and clear documentation of all issues found and corrections made to ensure transparency, reproducibility, and high data integrity.

The prompt

246 words3 blanks to fill in

Act as a data analysis expert responsible for data integrity and quality assurance.
Task: Clean and standardize the following dataset to prepare it for accurate analysis.
– Dataset (raw data or file contents): [DATASET], placeholder
– Data format (e.g., CSV, Excel, SQL table, JSON, text): [DATA_FORMAT], placeholder
– Intended use or analysis goal (optional): [ANALYSIS_GOAL], placeholder
Data cleaning objectives:
– Identify and correct misspellings and typographical errors.
– Fix grammatical or textual inconsistencies in categorical or text fields.
– Resolve syntax issues (inconsistent delimiters, casing, spacing, encoding problems).
– Standardize formats (dates, numbers, units, categories, naming conventions).
– Detect anomalies or inconsistencies that may impact downstream analysis.
Process requirements:
1) Initial data review and profiling (patterns, frequency checks, nulls, duplicates).
2) Detailed list of issues found, categorized by type (spelling, grammar, syntax, formatting).
3) Specific corrections applied, including standardization rules used.
4) Validation steps taken to confirm data accuracy after cleaning.
5) A concise data cleaning log documenting:
– Original issue
– Location/field
– Correction made
– Rationale
Output format:
– Cleaned dataset (presented in the same format as input where possible).
– A clearly labeled data cleaning report documenting all actions taken.
– Notes on any assumptions made or unresolved data quality concerns.
Guidelines:
– Preserve original meaning and intent of the data.
– Do not invent or infer missing values unless explicitly instructed.
– Prioritize reproducibility and clarity in documentation.
If any required input is missing, request clarification or proceed using best-practice assumptions and label them clearly.

Highlighted text is a blank. Swap it for your own detail before you send the prompt.

What it produces

A cleaned CSV file with standardized column names, corrected spelling errors, and consistent date formats.

A data cleaning report listing detected anomalies, corrections applied, and validation checks performed.

Documentation explaining standardization rules (e.g., country names, category labels) used across the dataset.

How to use this prompt

Four steps, about a minute. Nothing to install and no account needed.

  1. Copy the prompt

    Use the Copy button above. The whole prompt goes to your clipboard exactly as written. Nothing is trimmed.

  2. Fill in the blanks

    Replace [DATASET] [DATA_FORMAT] [ANALYSIS_GOAL] with your own details. Everything in square brackets is a blank for you to fill in.

  3. Paste it into your assistant

    Checked against Perplexity, ChatGPT, Claude. It is written as plain instructions, so paste it into whichever of them you already use.

  4. Read the result, then push back

    Compare what you get to the example below. If it is close but not right, say what to change: shorter, warmer, more specific, then ask again in the same conversation.

More prompts like this one

Close neighbours in Analytics, by shared subject and tagging.

Share Data Cleaning and Quality Assurance for Reliable Analysis:

Reviews and questions

What other people got out of this prompt, and what they asked about it.

Discussion is closed for this prompt.

No reviews yet

Be the first to share how this prompt worked for you.