Suppose a competitor tells you their synthetic audience is 98% accurate. Accurate at what? Against which humans, on which questions, measured how? Currently, there is no published methodology for evaluating synthetic research architectures. For a technology informing the world’s most consequential decisions, this isn’t good enough. At Artificial Societies, we simulate audiences of high-value decision makers so enterprise customers can test strategy and messaging before committing to it in the real world. The stakes of their consequential decisions are necessarily the stakes of our evaluation.
Large language models (LLMs) write conspicuously plausible text. A persona enacted by an LLM can sound exactly like the person it is meant to represent while holding no stable beliefs at all, and a synthetic population can reproduce a survey’s headline distribution while the relationships between its respondents’ answers remain incoherent. Since consequential decisions have little margin for error, we’ve become obsessed with evaluating these synthetic samples. How do we ensure consistent outputs when working with stochastic models? Or that persona responses reflect human belief systems? And when we test a message, does the observed effect reflect a real mechanism of human behaviour that will generalise to new settings and unseen populations?
Given their exceptional language understanding and fluency, LLMs can capture patterns of cultural knowledge, linguistic idiosyncrasies, common-sense reasoning, and even human behaviour (Liu et al., 2022; Park et al., 2022, 2023; Horton, 2023; Ashokkumar et al., 2026). With their similarly impressive capabilities in classifying and generating text, industry and academia have used these models to generate both survey (Argyle et al., 2023) and behavioural responses (Aher et al., 2022). Prompt engineering has enabled LLMs to adopt personas (with specified characteristics such as age, gender, ethnicity, and political orientation) and generate representative response distributions that capture the complex relationship between human subpopulations’ beliefs, attitudes, and cultural influences (Argyle et al., 2023). Our work shows that this performance extends to large populations of LLM agents that spontaneously reproduce homophily, a core feature of human community formation (He et al., 2024). In any case, synthetic surveys have become an increasingly popular way to learn about society and interrogate the hardest-to-reach populations that previously would have been impossible to research.1 (footnote) The field, as we see it, has now reached an evaluation bottleneck.
Beyond Superficial Benchmarks
Current industry standards for evaluating synthetic research are underwhelming. Naive metrics that measure how well simulated response distributions match held-out distributions of real human responses (i.e., Total Variation Distance, Jensen-Shannon divergence, Spearman’s ρ on rankings of a survey question’s options) are reported as proof of validity. In fact, these metrics are easily satisfied when a simulation reproduces the (correct) marginal distribution of every question in a survey, even when the relationship between them is instead incorrect. In other words, a simulation can reproduce how many people gave each answer without reproducing who gave which combination of them. Marginal fit is necessary but, alone, it provides no information about whether the population’s underlying response structure has been captured (Figure 1).
Exact same answers ≠ the same beliefs behind them
Two bar charts show humans and a synthetic audience giving identical distributions of answers to “I intend to buy this” and “It is worth the price”. Two heatmaps of the joint distribution show the human answers concentrated along the diagonal (one attitude drives both answers, r = 0.8) and the synthetic answers spread evenly across the grid (the two answers are unrelated, r = 0).
Q1 “I intend to buy this”
Q2 “It is worth the price”
Real humans
One attitude drives both answers; r = 0.8
Synthetic audience
The two answers are unrelated; r = 0
A more demanding test, and one we report ourselves, is producing a simulation that is as accurate as human self-replication. When humans answer the same question twice, they only reproduce the same answer roughly 91% of the time (Park et al., 2026). Our own architecture reproduces human answers 86% of the time across 1,000 diverse surveys, achieving up to 95% of the human ceiling. That said, we shouldn't aim to reproduce human answers with perfect consistency. People are reliably inconsistent, with their response distributions shifting when questions are reworded. Our goal is therefore to produce simulation outputs with human-calibrated consistency (Figure 2).
Human-calibrated (not maximal) consistency is the target
A curve of behavioural realism against test–retest reproduction rate rises from 50% to a peak at the human ceiling of 91% and falls steeply beyond it. A shaded human-calibrated range spans about 80% to 95%; our architecture is marked at 86%. Below the band the simulation is too noisy; above it, too rigid.
Three Kinds of Validity
To move beyond superficial survey matching and assess how well a system models human behaviour, we engage social science and psychology. These disciplines proffer empirically validated theories about how groups of people behave, form worldviews, and are influenced by those around them, while equipping us with a repertoire of psychometric tools to measure various underlying (or “latent”) traits and assess their observable implications. We ground our evaluation architecture in three foundational tenets of social science (Campbell, 1957; Cronbach & Meehl, 1955): internal validity (did the stimulus cause the response shift?), construct validity (do the questions measure the trait we think they measure?), and external validity (do the results hold for real populations?). Figure 3 lays out the eight tests that fall under them.
Internal Validity
In classical behavioural science, internal validity asks whether an observed effect is genuinely caused by the stimulus rather than system noise or artefacts. In other words, in a simulation, a shift in persona opinion must be strictly attributable to the tested stimulus, not to model hallucination or stochastic noise. Take sycophancy: humans are prone to agreeing with whatever a question might imply, but LLM personas do so far more often. In simulation, a confidently framed stimulus can shift responses because the persona is being agreeable (an artefact of its training) rather than because it was persuaded.
We test for internal validity at an individual level by checking that our personas’ responses are anchored to a stable belief system. To mimic humans, these responses should be generated by a small set of unobserved (“latent”) constructs, such as political ideology, risk tolerance, and institutional trust. Personas should hold these traits and conduct themselves accordingly. To test for this, we ask semantically equivalent versions of the same question, and measure agreement with Cohen’s kappa. We also define sets of conceptually linked questions with valid response chains and measure the rate at which a persona violates them. Both metrics are scored against established human test–retest bands, so a persona whose answers do not shift under reframing fails the same test as one whose answers move at random.
The same logic extends to the survey instrument itself. A simulation architecture is coherent and robust if it produces the same results when we change meaningless features of a survey. For example, reordering response options should not shift a distribution; reversing a Likert scale should only reverse the answer; and changing question order should not affect a marginal. Conversely, increasing the magnitude of a claim (a stimulus promising a 5% vs 25% saving) should shift responses further and monotonically across that range. Likewise, changing the messenger should change the reception, and the removal of substantiating evidence from a narrative should weaken scores. Thus, evaluation must strike a balance between a volatile simulation that reports significant differences in response to menial changes, and an inert one that stays the same even when we perturb the elements that matter.
Construct Validity
As the science of measuring latent traits and attitudes, psychometrics powers construct validity. In human data, latent constructs leave a signature in the covariance among survey questions. The logic here is that related survey questions should load on the same latent traits, generating correlations in humans’ responses to them (Cronbach & Meehl, 1955). To ensure that our personas exhibit the same correlations, we administer sets of items known from human studies to measure a single construct (e.g., items on redistribution, regulation, and public spending that together measure economic ideology), and then assess their covariance with Cronbach’s alpha and McDonald’s omega. Human instruments typically land between 0.7 and 0.9, and we hold our simulations to that same band.
Psychometrics also helps us distinguish convergent from discriminant validity (Campbell & Fiske, 1959). Convergent validity requires that measures of the same underlying trait correlate strongly (e.g., stated purchase intent should track willingness to pay for the same product), whilst discriminant validity requires that measures of distinct traits correlate more weakly. For example, purchase intent and unaided brand recall should be related, but far less tightly than two measures of intent.
Synthetic populations frequently satisfy convergent validity, but often fail to discriminate: a bad LLM persona shown a new product will rate it high on quality, value, trustworthiness, and likelihood of recommending it, even though these judgements come apart in real consumer data (Lukauskas & Šarkauskaitė, 2026). Together, the internal and construct validity tests establish what we call multi-level coherence: personas must maintain logically stable belief systems across paraphrased questions (persona level) and exhibit authentic latent psychometric covariance across the population (audience level).
External Validity
External validity asks whether findings generalise across populations, settings, and timeframes (Campbell, 1957). A simulation is externally valid if its outputs accurately predict human responses beyond a single isolated prompt. At its most basic, this means matching real-world response distributions across granular population sub-samples (Argyle et al., 2023; Santurkar et al., 2023). Beyond that, it means demonstrating that when we test a stimulus (e.g., a corporate narrative), the effect size produced by the synthetic audience matches the effect size produced in real-world human experiments (Ashokkumar et al., 2026; Aher et al., 2022). Replicating treatment effects provides evidence that the simulation has a causal understanding of the target population, increasing the chance that it will reliably generalise to unseen settings.
A third external test is structure match. Here, exploratory factor analysis can help decompose the shared variance among items into unlabelled factors, allowing us to model each response as a linear combination of those factors plus a unique error term (Spearman, 1904; Fabrigar et al., 1999). Two outputs should follow. First, the recovered factors should correspond to constructs we can name (Messick, 1995): a factor loading on nationalisation, tax-funded services, and business regulation is recognisably “support for state intervention”. Second, persona attributes — age, gender, occupation, etc. — should predict scores on those factors as they do in human data (Bisbee et al., 2024). In this way, convincing marginal distributions with no recoverable factor structure indicate a hollow simulation that’s unlikely to generalise. In other words, a simulation architecture with no meaningful response structure will not generalise to truly held-out samples or unseen questions, no matter its distributional accuracy against ANES, GSS, and Pew (ibid.).
Does the simulation model human behaviour?
Internal validity
Did the stimulus cause the response shift?
Construct validity
Do the questions measure the trait?
External validity
Do the results hold for real populations?
Internal validity
Did the stimulus cause the response shift?
Repeated questions (current industry standard)
Cohen’s kappa
Linked questions
Logical violation rate
Perturbed questions
Order, scale, framing
Construct validity
Do the questions measure the trait?
Item covariance
Alpha and omega
Trait separation
Discriminant validity
External validity
Do the results hold for real populations?
Distributional fit (current industry standard)
TVD, JSD, Spearman
Effect size match
Held-out experiments
Structure match
Human factor loadings
Internal and construct validity together: multi-level coherence.
Two of eight boxes are the current industry standard
Dashed boxes require human study data
An Imitation of an Imitation
A harder problem brews beneath all this. Though we argue that matching ground truth (viz. human-answered surveys) is important, whether by external criteria (matching distributions and effect sizes) or internal ones (coherence and psychometric structure), human survey responses are themselves messy and inherently unstable. People answer differently depending on how you ask, when you ask, and who they think is listening. And, of course, what people say is often a weak predictor of what they will do, the age-old intention–behaviour gap that is inevitably inherited in synthetic research. Evaluating a synthetic population purely against human self-reports therefore means building an imitation of an imitation, and grading it against that very imitation.
To be clear, this is not an argument against synthetic research. Every survey ever fielded has run on the same proxy, and market research has spent decades managing the gap between stated and revealed preference. However, the critical benefit for social simulations is their capacity for interrogation in ways that are impossible for human respondents. We can gauge whether the persona’s answers come from a stable belief system, whether that structure matches the structure in human data, and whether it moves under a stimulus the way real people move.
Synthetic audiences will never be a true substitute for the people they seek to represent. But we can test them more thoroughly than any panel of respondents. Do personas keep their beliefs when a question is reworded, do those beliefs hang together across the audience as human beliefs do, does the audience move for substantive changes and stay still for cosmetic ones, and does it reproduce effects from experiments the model has never seen?
Of the eight tests in Figure 3, the two marked with an asterisk are the industry standard today. We hold ourselves to all eight. Internally, we make a commitment to consistently evaluating our methodology so that alongside contributing to the frontier of synthetic research, others can check our numbers and compare where we fall short. The decisions our customers bring to us deserve nothing less than results they can act on with confidence.
Notes
- 1.
Synthetic research has also received well-warranted criticism (Bisbee et al., 2024; Sarstedt et al., 2024; Harding et al., 2024).
Glossary
- Marginal distribution
- How answers to a single question are spread across the response options (e.g., 20% strongly agree, 35% agree…). Matching marginals means getting each question’s headline result right, considered on its own.
- Joint distribution
- How answers to two or more questions are spread together: not just how many people agree with each statement, but how many agree with both, with one but not the other, and so on. This is where the relationships between beliefs live, and where marginal matching goes blind.
- Total Variation Distance (TVD)
- A measure of how far apart two distributions are, from 0 (identical) to 1 (no overlap). A TVD of 0.1 means you would need to move 10% of respondents between answer categories to turn one distribution into the other.
- Jensen–Shannon divergence (JSD)
- Another measure of the distance between two distributions, also bounded between 0 and 1. It weights disagreements differently from TVD but answers the same question: how similar are these two sets of answers?
- Spearman’s ρ (rho)
- A correlation measure for rankings. If you order a question’s answer options from most to least popular for both humans and the simulation, Spearman’s ρ tells you how well the two orderings agree.
- Test–retest reliability
- The rate at which a respondent gives the same answer when asked the same question twice. Humans reproduce their own answers roughly 91% of the time; a simulation should land near that figure, not above it.
- Cohen’s kappa (κ)
- A measure of agreement between two sets of answers that corrects for agreement you would expect by chance. Used here to check whether a persona gives the same answer to two differently worded versions of the same question.
- Latent construct
- An underlying trait that cannot be observed directly but shapes many observable answers, such as political ideology, risk tolerance, or trust in institutions. Survey questions are attempts to measure these indirectly.
- Covariance
- The tendency of two measures to move together. If people who rate a product as overpriced also tend to report weak renewal intent, those two items covary. Human beliefs leave a characteristic covariance signature that a simulation should reproduce.
- Cronbach’s alpha (α) and McDonald’s omega (ω)
- Measures of internal consistency: how strongly a set of questions designed to measure the same construct actually hang together. Human instruments typically score between 0.7 and 0.9. Scores near 1.0 are a warning sign, indicating respondents are repeating themselves rather than reasoning.
- Exploratory factor analysis (EFA)
- A statistical method that looks at the pattern of covariance among many questions and identifies a smaller number of underlying factors that explain it. Used to check whether the simulation’s hidden structure resembles the structure found in human data.
- Factor loading
- How strongly a given question relates to a given underlying factor. If a persona’s age predicts how it loads on a factor capturing support for state intervention, the simulation is reproducing a relationship found in real populations.
- Internal validity
- Whether an observed effect was genuinely caused by the thing being tested, rather than by noise, artefacts, or quirks of the system.
- External validity
- Whether findings generalise beyond the specific test to real populations, settings, and time periods.
- Construct validity
- Whether the questions asked actually measure the trait they are meant to measure.
- Convergent validity
- Different measures of the same trait should agree with one another (stated purchase intent should track willingness to pay).
- Discriminant validity
- Measures of different traits should be distinguishable from one another (purchase intent and brand recall should be related, but not identical). Synthetic audiences often pass convergent tests and fail discriminant ones, rating everything about a product uniformly high.
- Effect size
- The magnitude of a change caused by a stimulus, not just whether a change occurred. A simulation with external validity should reproduce not only the direction but the size of effects observed in human experiments.
- Held-out data
- Human data the model has never been exposed to, so that agreement with it counts as a real test. Because major public survey programmes appear in LLM training data, only unpublished or newly fielded studies are truly held out.
- Sycophancy
- The tendency of language models to agree with whatever a prompt implies. Humans do this too, but LLM personas do it far more, so a confidently framed message can shift responses without any real persuasion taking place.
- Stochastic
- Involving randomness. LLMs are stochastic: the same prompt can produce different outputs on different runs, which is why consistency has to be measured rather than assumed.
- Perturbation
- A deliberate change to a survey to test robustness. Cosmetic perturbations (reordering options, reversing a scale) should leave results unchanged; substantive perturbations (changing a claim from 5% to 25%) should move them predictably.
- Intention–behaviour gap
- The well-documented difference between what people say they will do and what they actually do. Every survey inherits it; synthetic research inherits it, too.
References
Aher, G., Arriaga, R. I., & Kalai, A. T. (2022). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies (Version 5). arXiv.
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3), 337–351.
Ashokkumar, A., Hewitt, L., Ghezae, I., & Willer, R. (2026). Large language models can predict the results of social science experiments. Nature, 656(8126), 115–122.
Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis, 32(4), 401–416.
Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297–312.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Dillion, D., Tandon, N., Gu, Y., & Gray, K. (2023). Can AI language models replace human participants? Trends in Cognitive Sciences, 27(7), 597–600.
Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J. (1999). Evaluating the use of exploratory factor analysis in psychological research. Psychological Methods, 4(3), 272–299.
Harding, J., D’Alessandro, W., Laskowski, N. G., & Long, R. (2024). AI language models cannot replace human research participants. AI & SOCIETY, 39(5), 2603–2605.
He, J. K., Wallis, F. P. S., Gvirtz, A., & Rathje, S. (2024). Artificial intelligence chatbots mimic human collective behaviour. British Journal of Psychology, 117(2), 761–776.
Horton, J. J. (2023). Large language models as simulated economic agents: What can we learn from Homo silicus? (Version 1). arXiv.
Liu, J., Liu, A., Lu, X., Welleck, S., West, P., Bras, R. L., Choi, Y., & Hajishirzi, H. (2022). Generated Knowledge Prompting for Commonsense Reasoning (arXiv:2110.08387). arXiv.
Lukauskas, M., & Šarkauskaitė, V. (2026). Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents (Version 1). arXiv.
Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749.
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior (Version 2). arXiv.
Park, J. S., Popowski, L., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2022). Social simulacra: Creating populated prototypes for social computing systems (Version 1). arXiv.
Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., & Bernstein, M. S. (2026). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (arXiv:2411.10109). arXiv.
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? (Version 1). arXiv.
Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41(6), 1254–1270.
Spearman, C. (1904). ‘General Intelligence,’ Objectively Determined and Measured. The American Journal of Psychology, 15(2), 201.