Knowledge base
How to evaluate a synthetic research vendor: eight validity tests
By Artificial Societies
Published
To evaluate a synthetic research vendor, ask for evidence on eight validity tests that answer three questions: whether the stimulus caused the response shift, whether the questions measure the intended trait, and whether the results hold for real populations. Artificial Societies, which builds audiences as networks of AI personas, reports 86% distribution accuracy across 1,000 surveys against a 91% human self-replication ceiling (Survey Evaluation Report, January 2026), which is evidence for one of the eight. Artificial Societies’ September 2026 validity framework sets out all eight and counts two as today’s industry standard.
Why do headline accuracy numbers mislead?
Headline accuracy numbers mislead because they leave four things open: what was measured, against which people, on which questions, and who ran the test. Edoardo Chidichimo and Felix Wallis start from that gap in Artificial Societies’ validity framework, Distribution Matching Is a Bad Way to Evaluate Simulations, published on 9 September 2026. In the framework’s reading, the usual answer to “accurate at what?” is distributional fit: how far simulated and human answer distributions sit apart, question by question, measured with total variation distance (TVD), Jensen–Shannon divergence (JSD) or Spearman’s rank correlation.
Distributional fit is easy to satisfy without modelling anyone. The framework’s first figure sets a human audience beside a synthetic one with identical answer distributions for “I intend to buy this” and “It is worth the price”, so every per-question metric scores zero distance. Their joint distributions, which record who gave which combination of answers, differ by a TVD of 0.4: 40% of one audience would have to change its combination of answers to match the other. Among the humans one attitude drives both answers, and among the synthetic respondents the two are unrelated.
Bisbee and colleagues found the same gap in real data (Political Analysis, May 2024 (opens in a new tab)). Their ChatGPT personas matched the averages of the American National Election Studies (ANES), with less variation than real respondents. Yet 48% of the regression coefficients, which measure how a trait such as age or party relates to an answer, differed significantly from the human estimates. A vendor whose evidence stops at matched averages has shown you the part of the result that Bisbee’s personas got right.
What are the three kinds of validity in synthetic research?
Synthetic research borrows three kinds of validity from social science, and each asks a separate question. Internal validity asks whether the stimulus caused the response shift, rather than noise, a quirk of the model or sycophancy, a language model’s habit of siding with whatever the prompt appears to want. Construct validity asks whether the questions measure the trait they are meant to measure, such as economic ideology or purchase intent. External validity asks whether the results hold for real populations, settings and times.
Each failure has a recognisable form in the framework. Without internal validity, a message can seem to persuade personas that were only being agreeable. Without construct validity, a synthetic audience rates a new product high on quality, value, trust and recommendation at once, a halo that real consumer data does not show. Without external validity, a coherent population can still be the wrong one, which is why the framework asks for effect sizes from human experiments.
What are the eight validity tests, and what should you ask for?
Artificial Societies’ September 2026 framework groups the eight tests below and marks repeated questions and distributional fit as the current industry standard. The last column is the evidence to request from any vendor, including us.
| Test | Validity | What it catches | What to ask the vendor for |
|---|---|---|---|
| 1. Repeated questions (Cohen’s kappa, agreement corrected for chance) | Internal | A persona whose answer changes when a question is reworded, or never changes at all | Agreement between equivalent phrasings, scored against human test–retest bands |
| 2. Linked questions (logical violation rate) | Internal | Answers that contradict each other across related questions | The rate at which personas break valid chains of answers |
| 3. Perturbed questions (order, scale, framing) | Internal | Results that move under cosmetic changes, or stay still under substantive ones | Results with options reordered, the scale reversed and question order changed, plus a claim scaled from a 5% to a 25% saving |
| 4. Item covariance (Cronbach’s alpha, McDonald’s omega: how tightly items on one trait agree) | Construct | Items meant to measure one trait that do not agree, or agree too tightly | Alpha or omega on established multi-item scales, against the human band of 0.7–0.9; a score near 1.0 is a warning |
| 5. Trait separation (discriminant validity: distinct traits stay distinct) | Construct | A halo: one product rated high on quality, value, trust and recommendation at once | Correlations between distinct measures, such as purchase intent and unaided brand recall, compared with two measures of intent |
| 6. Distributional fit (TVD, JSD, Spearman) | External | Wrong headline answers to individual questions | Which surveys, whether they were held out, and fit for subgroups as well as totals |
| 7. Effect-size match (held-out experiments) | External | A simulation that gets the direction of an effect right and its size wrong | Replications of human experiments, with the human effect and its confidence interval beside the synthetic one |
| 8. Structure match (factor loadings: how strongly each question ties to an underlying trait) | External | A hollow simulation: plausible answers with no underlying structure | A factor analysis, which finds the few traits behind many answers, compared with human data, and whether persona attributes predict those traits as they do in people |
Source: Chidichimo and Wallis, Distribution Matching Is a Bad Way to Evaluate Simulations, Artificial Societies, September 2026. In the framework’s figure, tests 6 to 8 are the ones that require human study data.
For message testing, the perturbation test matters most. Language model personas agree with confident framing more readily than people do, so a message can appear to persuade a synthetic audience that has only been agreeable. The framework (Chidichimo and Wallis, September 2026) therefore asks for checks in both directions: cosmetic changes should leave results where they are, while changing the messenger or removing the evidence behind a claim should move them.
What counts as held-out data?
Held-out data is human data a model could not have seen during training, so agreement with it tests the model rather than its memory. The framework sets a strict bar: major public survey programmes sit in language model training data, so only unpublished or newly fielded studies qualify. A match against ANES, the General Social Survey (GSS) or Pew may show that the model has read the published results. It says less about how the model will handle your question.
The American Association for Public Opinion Research (AAPOR) takes a compatible position in its May 2026 task force report, Responsible AI Integration in Survey Research (opens in a new tab). It suggests checking synthetic responses for leakage or memorisation, such as unusually specific strings reproduced from training sources, and recommends keeping at least some human respondents in a study so that simulated and human answers can be compared.
Timing gives the cleanest held-out design, because a prediction logged before an event cannot draw on an outcome that did not yet exist. Artificial Societies used it in its September 2026 earnings call study, preregistering the questions its simulated sell-side analysts would ask at the NVIDIA, Marvell and Credo calls on the Open Science Framework (OSF), a time-stamped research registry, before each call. Ask any vendor for the fielding date of every benchmark and the training cut-off of the model that produced the answers.
How does Artificial Societies’ evidence map onto the eight tests?
Artificial Societies’ evidence sits in three of the sources this article draws on. The January 2026 report predates the framework, so the mapping below is ours.
| Test | Artificial Societies’ evidence | Source | Limit |
|---|---|---|---|
| 6. Distributional fit | 86% distribution accuracy across 1,000 surveys, against 67% for biography-prompted models and a 91% human self-replication ceiling | Survey Evaluation Report, January 2026 | The evaluation page does not say which of the 1,000 surveys were held out |
| 4. Item covariance | 89% internal coherence on Cronbach’s alpha | Survey Evaluation Report, January 2026 | The evaluation page gives a human range of 60–95%, wider than the framework’s 0.7–0.9 band |
| 2. Linked questions | Under 2% self-contradiction, against about 9% for human panels | Survey Evaluation Report, January 2026 | A rate far below the human one also needs checking for rigidity |
| 7. Effect-size match | Misinformation cut the share of Britons who would definitely accept a COVID-19 vaccine by 6.6 points, and the synthetic share by 8.2, inside the authors’ 3.8–9.0 interval | Misinformation study, 23 September 2026 | The experiment was published, so it is not held out; a synthetic control group drifted about four points where the human one did not |
| Held-out design (not one of the eight) | Questions preregistered before each call; detailed personas came closer than basic profiles in substance and phrasing, and an evaluator matched Credo analysts more often than chance | Earnings call study, 21 September 2026 | No gain on Credo’s broad topics; analysts’ past questions remained a strong baseline for wording and identity |
Source: Artificial Societies, method and evaluation page, misinformation study and earnings call study, read 28 September 2026.
Artificial Societies hashed and time-stamped the misinformation study’s design before the first survey call, which rules out tuning the design after seeing results. It cannot rule out a model having read the published papers, which is why the table does not count the replication as held out. Artificial Societies’ framework states the commitment as “We hold ourselves to all eight.” As of 28 September 2026, four of the eight have no published Artificial Societies result: repeated questions, perturbed questions, trait separation and structure match. Ask us for that evidence as you would any other vendor.
What questions should you put to any synthetic research vendor?
ESOMAR already asks suppliers how they measure and assess validity in its 20 Questions to Help Buyers of AI-Based Services for Market Research and Insights (opens in a new tab), issued in March 2024. For tools that use synthetic data, it also asks what validation compares synthetic outputs with primary research or real-world results. The questions below apply that request to the eight tests:
- 1.
What exactly does your accuracy figure measure, against which human data, on which questions, and who ran the evaluation?
- 2.
Which of your benchmark studies were unpublished, or fielded after the model’s training cut-off?
- 3.
Can you show joint distributions or subgroup cross-tabulations as well as per-question totals?
- 4.
What happens to results when answer options are reordered, the scale is reversed or the messenger changes?
- 5.
Have you reproduced the effect size of a human experiment, and where does your estimate fall against the human confidence interval?
- 6.
Have you logged any prediction before the outcome was known, and where is the time-stamped record?
- 7.
How many human respondents does a study include, and how are synthetic responses labelled in the deliverable?
- 8.
When the underlying model changes, which validations do you repeat?
Questions 7 and 8 follow AAPOR’s May 2026 advice on disclosure and on repeating validation when a model changes. Our guide to AI market research tools sets out what named vendors publish about their methods and evidence, and our explainer on silicon sampling covers the academic studies behind the method.
What do validity tests leave unanswered?
Validity tests leave two problems open. Human data, the standard they grade against, is itself unstable: answers shift with wording, timing and the audience a respondent imagines, and stated intentions predict behaviour weakly, the intention–behaviour gap every survey inherits. Subgroups are the other gap. A preprint by Florian Kutzner and colleagues, including Artificial Societies researchers, posted to arXiv on 23 September 2026, argues that ad hoc comparisons with human surveys test the wrong thing when decision makers need to anticipate consequential behaviour. It asks for validity claims by subgroup, since aggregate accuracy can hide the groups a decision affects most. A simulation that passes all eight tests is still graded against self-reports, which makes it, in the framework’s phrase, an imitation of an imitation.
Frequently asked questions
Is a higher accuracy percentage always better?
A higher percentage is better only up to the human ceiling. Artificial Societies’ January 2026 Survey Evaluation Report uses a 91% human self-replication ceiling, the rate at which people repeat their own answer to the same question. Artificial Societies’ September 2026 framework says a simulation should land near that figure, not above it, because real respondents do not repeat every answer. Ask where a vendor’s figure sits against the human benchmark, and how the vendor measured that benchmark.
What is the difference between a marginal and a joint distribution?
A marginal distribution counts how people answered one question, such as the share who agree. A joint distribution counts combinations across questions: who agreed with one statement and disagreed with another. Beliefs connect through those combinations, so a vendor that reports only marginals, however closely they match, has not tested whether its personas hold beliefs that fit together.
Which validity tests can a vendor run without new human study data?
Five of the eight, according to Artificial Societies’ September 2026 framework. Repeated, linked and perturbed questions compare the simulation with itself, and item covariance and trait separation examine how its answers relate to each other, although several are scored against human benchmark bands. Distributional fit, effect-size match and structure match require human study data to compare against, and the framework’s figure marks them that way.
How can you check a vendor’s claim of preregistration?
Ask for the registry link and compare its timestamp with the event. Artificial Societies registered its earnings call predictions on the Open Science Framework, a free platform run by the non-profit Center for Open Science, keeping the prompts, generated answers and evaluation rules, with digital fingerprints for any withheld proprietary inputs. Check that the evaluation rules were registered too; without them, a flattering analysis can be chosen afterwards.
Should a synthetic study be called a survey or a poll?
AAPOR’s May 2026 task force report says it should not. Poll, polling, survey and surveying imply that the data come from human respondents. The report therefore asks that data created by AI be identified as such, and that the number of human respondents be disclosed. ESOMAR’s 2024 buyer questions ask the same of suppliers: how they distinguish data derived from natural persons from data derived synthetically.
How often should a vendor repeat its validation?
A vendor should repeat its validation whenever the model, the prompts or the audience changes. AAPOR’s task force advises against assuming that performance on one task carries over to another, or that past performance holds, because providers update models and prompt templates in ways that can alter behaviour. Bisbee and colleagues found the same ChatGPT prompt gave significantly different results over three months in 2023. Ask for the date and model version behind every figure.
Sources
- Chidichimo, E. and Wallis, F., Artificial Societies, Distribution Matching Is a Bad Way to Evaluate Simulations, 9 September 2026. Read 28 September 2026.
- Artificial Societies, Survey Evaluation Report, January 2026, as presented on the method and evaluation page. Read 28 September 2026.
- Artificial Societies, Simulating Misinformation, Fact-Checking, and Prebunking, 23 September 2026. Read 28 September 2026.
- Wei, L. and Chidichimo, E., Artificial Societies, Testing Simulations Against Reality: Learnings from Earnings Calls, 21 September 2026. Read 28 September 2026.
- Kutzner, F., Kacperski, C., de Molière, L., Chidichimo, E., Jung, M. J., Wallis, F. P. S. and He, J. K., Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research, arXiv preprint, 23 September 2026. https://doi.org/10.48550/arXiv.2609.27690 (opens in a new tab) Read 28 September 2026.
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B. and Larson, J. M., Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4), 401–416, May 2024. https://doi.org/10.1017/pan.2024.5 (opens in a new tab) Read 28 September 2026.
- AAPOR Task Force on Responsible AI Integration in Survey Research (co-chairs Rothschild, D. M. and Marlar, J.), Responsible AI Integration in Survey Research, American Association for Public Opinion Research, May 2026. https://aapor.org/wp-content/uploads/2026/05/Responsible-AI-Integration-In-Survey-Research.pdf (opens in a new tab) Read 28 September 2026.
- ESOMAR, 20 Questions to Help Buyers of AI-Based Services for Market Research and Insights, March 2024. https://esomar.org/20-questions-to-help-buyers-of-ai-based-services (opens in a new tab) Read 28 September 2026.