Skip to content

Data quality check

Is your survey data high quality?

Check your data like how social scientists check data quality before analysis. Applicable for both human panel data and simulated data.

Why check internal quality of survey data

And how it applies to both human and simulated respondents.

Pew Research Center found that 4–7% of interviews in online opt-in panels came from ‘bogus’ respondents — and that traditional quality checks missed most of them: 84% passed the survey’s trap question and 87% passed a speeding check. In 2026, Pew’s methodologists warned that AI now makes it cheap to fake a respondent at scale in panels that claim to be fully human.

Artificial Societies uses large networks of AI to model important groups of people. We’re social and behavioural scientists, so we hold our simulations to the same standard as how we would check a human dataset before conducting social science research. A human panel can be contaminated with low-attention respondents or bot-farm participants. A carefully built simulation can be more internally coherent. The only way to tell is to look at the raw, respondent-level data.

  • What the check looks for

    • Coherent respondents that don’t hallucinate and contradict themselves
    • Realistic diversity of opinions by gender, generation and geography
    • Opinions that correlate into themes, the way real attitudes do
    • Open text with depth and nuance, written by many different people
  • What it doesn’t do

    • It doesn’t check external accuracy of the findings
    • It is not domain-specific, and focuses on universal internal quality metrics
    • It doesn’t tell you whether the data is current, or whether the questions were well designed
    • It doesn’t tell you whether the dataset is human or synthetic, just its quality

What your data gets compared against

Metrics social scientists use to check for data quality.

High-quality data has a recognisable shape. Answers spread across the options, differ from question to question, correlate into themes, and vary between different people. Good data is messy with patterns. Low-quality data — rushed, straightlined, automated or generated — can sometimes be too inconsistent, too uniform, or conversely too templated.

  • High-quality data

    Varied, spread, untidy

  • Mixed signals

    Two questions spike

  • Low quality / unlikely human

    One shape, repeated

Share choosing each option — rows are the questions (awareness, purchase intent, motivation, repurchase intent, perceived value, barriers, comfort level), columns the seven answer options.lowerhigherIllustrative
  • Too uniform

    Individuals barely differ from each other, and no one has a consistent belief across their answers.

  • The high-quality band

    Messy but with pattern: clear themes, clear outliers.

  • Too templated

    Answers collapse into a few repeated shapes, losing the noise and nuance that make us human.

The four measures of quality

Based on the checks social scientists run before trusting a dataset.

  1. Respondent coherence

    How often do respondents contradict themselves?

    Needs: Questions that are related

  2. Subgroup diversity

    Do opinions differ by gender, generation, and geography realistically?

    Needs: Demographic columns

  3. Opinion correlation

    Do opinions form thematic structures organically?

    Needs: Several rating-scale questions

  4. Open text richness

    Does the free text have depth and nuance?

    Needs: At least one free-text column

What comes back

A short report you can circulate.

  • A score and a band

    High quality, mixed signals, low quality / unlikely human, or not assessable.

  • Coverage

    How much of the check could run on your file. A file with no free text can’t be scored on open text richness, so we report it as skipped to help contextualise the score.

  • Findings

    One per measure explaining how to interpret the metric, or reason why the check was skipped.

  • Caveats

    What we couldn’t measure, and a note that this checks internal quality, not external quality or source of data.

Quality report

Illustrative

61

Mixed signals

out of 100

Coverage3 of 4 checks ran
  • Respondent coherence

    Flagged

    11 respondents (2.1%) contradict themselves across related questions.

  • Subgroup diversity

    Pass

    Age and region shift opinions realistically compared to high-quality samples.

  • Opinion correlation

    Pass

    Answers are correlated into two thematic clusters in realistic ways.

  • Open text richness

    Skipped

    Skipped — the file has no free-text column to read.

Caveat: This report checks the dataset’s internal quality, not its accuracy or where the data came from.

What to send

The more of your file arrives intact, the more we can measure.

  • Raw responses, one row per person

    One column per question. We are not able to measure these metrics based on crosstab or topline summaries.

  • Include the extra columns

    Free text unlocks open text richness checks, and demographics unlock subgroup diversity checks.

  • Strip direct identifiers

    Remember to remove names, emails, phone numbers, addresses and panellist IDs first.

Send a dataset for a quality check

Get results within two working days.

CSV, TSV, Excel, JSON, .sav or .dta · up to 25MB

By submitting you agree to our privacy policy, terms and DPA, and confirm you’re allowed to share this file.

Artificial Societies is SOC 2 certified and has strict GDPR compliance. We use your file to run this check and send you the report. Your file will never be shared outside Artificial Societies, and can be deleted on request. See our privacy policy and DPA.

Questions

  • Does a low score mean my data is fake?

    Not necessarily. The score is developed based on how social scientists check data quality. Real human panel data can also have a low score when it is contaminated with low-attention respondents, or even bot-farm ‘bogus’ participants. Conversely, high quality simulations can have a high score.

  • Does it work on my topic?

    Yes. The check focuses on internal quality, rather than whether the absolute numbers are accurate in reflecting current opinions. As such, it focuses on universal metrics rather than domain-specific benchmarks. We recommend domain-specific benchmarks to be conducted separately.

  • Can well-built simulations score highly?

    Yes. The check measures the internal quality of the dataset, such as how much do respondents contradict themselves, how much do relevant opinions correlate, and how deep and nuanced are the open text responses. A carefully built simulation can pass these tests, ours included, and a carelessly run human panel can fail them.

  • What if my file has no free text or demographics?

    We will report that those checks are skipped — “we couldn’t measure this because the file has no free-text column”. Our coverage reporting helps qualify the score in the context of the specific dataset.

  • Is it useful for human panel data too?

    Yes. According to Pew Research, traditional survey quality checks, such as a trap question or a speeding check, can fail to catch 80%+ of ‘bogus’ respondents. These checks are based on how social scientists check data quality before conducting any further analyses, which provide a more rigorous and holistic quality benchmark.

  • What does it cost, and what does Artificial Societies get out of it?

    It’s free. We built it to hold our own simulations and the human data we buy to the same standard, and we’d rather the whole field got better at this than keep it to ourselves. If the report is useful and you want to talk about the quality of the data behind your decisions, we’d love to — but there’s no obligation either way.

These are the same quality checks we hold our own simulations to. How we build and evaluate them is written up in our method and evaluation.