Data quality check
Is your survey data high quality?
Check your data like how social scientists check data quality before analysis. Applicable for both human panel data and simulated data.
Why check internal quality of survey data
And how it applies to both human and simulated respondents.
Pew Research Center found that 4–7% of interviews in online opt-in panels came from ‘bogus’ respondents — and that traditional quality checks missed most of them: 84% passed the survey’s trap question and 87% passed a speeding check. In 2026, Pew’s methodologists warned that AI now makes it cheap to fake a respondent at scale in panels that claim to be fully human.
Artificial Societies uses large networks of AI to model important groups of people. We’re social and behavioural scientists, so we hold our simulations to the same standard as how we would check a human dataset before conducting social science research. A human panel can be contaminated with low-attention respondents or bot-farm participants. A carefully built simulation can be more internally coherent. The only way to tell is to look at the raw, respondent-level data.
What the check looks for
- Coherent respondents that don’t hallucinate and contradict themselves
- Realistic diversity of opinions by gender, generation and geography
- Opinions that correlate into themes, the way real attitudes do
- Open text with depth and nuance, written by many different people
What it doesn’t do
- It doesn’t check external accuracy of the findings
- It is not domain-specific, and focuses on universal internal quality metrics
- It doesn’t tell you whether the data is current, or whether the questions were well designed
- It doesn’t tell you whether the dataset is human or synthetic, just its quality
What your data gets compared against
Metrics social scientists use to check for data quality.
High-quality data has a recognisable shape. Answers spread across the options, differ from question to question, correlate into themes, and vary between different people. Good data is messy with patterns. Low-quality data — rushed, straightlined, automated or generated — can sometimes be too inconsistent, too uniform, or conversely too templated.
High-quality data
Varied, spread, untidy
Mixed signals
Two questions spike
Low quality / unlikely human
One shape, repeated
Too uniform
Individuals barely differ from each other, and no one has a consistent belief across their answers.
The high-quality band
Messy but with pattern: clear themes, clear outliers.
Too templated
Answers collapse into a few repeated shapes, losing the noise and nuance that make us human.
The four measures of quality
Based on the checks social scientists run before trusting a dataset.
Respondent coherence
How often do respondents contradict themselves?
Needs: Questions that are related
Subgroup diversity
Do opinions differ by gender, generation, and geography realistically?
Needs: Demographic columns
Opinion correlation
Do opinions form thematic structures organically?
Needs: Several rating-scale questions
Open text richness
Does the free text have depth and nuance?
Needs: At least one free-text column
What comes back
A short report you can circulate.
A score and a band
High quality, mixed signals, low quality / unlikely human, or not assessable.
Coverage
How much of the check could run on your file. A file with no free text can’t be scored on open text richness, so we report it as skipped to help contextualise the score.
Findings
One per measure explaining how to interpret the metric, or reason why the check was skipped.
Caveats
What we couldn’t measure, and a note that this checks internal quality, not external quality or source of data.
Quality report
Illustrative
61
Mixed signals
out of 100
Respondent coherence
Flagged
11 respondents (2.1%) contradict themselves across related questions.
Subgroup diversity
Pass
Age and region shift opinions realistically compared to high-quality samples.
Opinion correlation
Pass
Answers are correlated into two thematic clusters in realistic ways.
Open text richness
Skipped
Skipped — the file has no free-text column to read.
Caveat: This report checks the dataset’s internal quality, not its accuracy or where the data came from.
What to send
The more of your file arrives intact, the more we can measure.
Raw responses, one row per person
One column per question. We are not able to measure these metrics based on crosstab or topline summaries.
Include the extra columns
Free text unlocks open text richness checks, and demographics unlock subgroup diversity checks.
Strip direct identifiers
Remember to remove names, emails, phone numbers, addresses and panellist IDs first.
Send a dataset for a quality check
Get results within two working days.
Questions
Does a low score mean my data is fake?
Not necessarily. The score is developed based on how social scientists check data quality. Real human panel data can also have a low score when it is contaminated with low-attention respondents, or even bot-farm ‘bogus’ participants. Conversely, high quality simulations can have a high score.
Does it work on my topic?
Yes. The check focuses on internal quality, rather than whether the absolute numbers are accurate in reflecting current opinions. As such, it focuses on universal metrics rather than domain-specific benchmarks. We recommend domain-specific benchmarks to be conducted separately.
Can well-built simulations score highly?
Yes. The check measures the internal quality of the dataset, such as how much do respondents contradict themselves, how much do relevant opinions correlate, and how deep and nuanced are the open text responses. A carefully built simulation can pass these tests, ours included, and a carelessly run human panel can fail them.
What if my file has no free text or demographics?
We will report that those checks are skipped — “we couldn’t measure this because the file has no free-text column”. Our coverage reporting helps qualify the score in the context of the specific dataset.
Is it useful for human panel data too?
Yes. According to Pew Research, traditional survey quality checks, such as a trap question or a speeding check, can fail to catch 80%+ of ‘bogus’ respondents. These checks are based on how social scientists check data quality before conducting any further analyses, which provide a more rigorous and holistic quality benchmark.
What does it cost, and what does Artificial Societies get out of it?
It’s free. We built it to hold our own simulations and the human data we buy to the same standard, and we’d rather the whole field got better at this than keep it to ourselves. If the report is useful and you want to talk about the quality of the data behind your decisions, we’d love to — but there’s no obligation either way.
These are the same quality checks we hold our own simulations to. How we build and evaluate them is written up in our method and evaluation.