Skip to content

Blog · · 4 min read

Artificial Societies Is More Accurate Than State-of-the-Art LLMs

By Artificial Societies Team

Last week, we published the Artificial Societies Benchmark, a validation framework for synthetic research. It asks whether simulations hold up as research evidence beyond matching the overall percentages choosing each survey answer. The benchmark’s eleven tests span three kinds of validity. External validity asks whether populations, subgroups, and effects match what real respondents did (i.e., top-line accuracy). Internal validity asks whether each respondent’s answers are stable and coherent. Construct validity asks whether the relationships between answers match those found in people. The framework is set out in two papers, Artificial Societies Benchmark: A Validation Framework for Synthetic Research and Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research, and in shorter form on our validation blog. To our knowledge, it is the most comprehensive validation framework published for synthetic respondents.

The public benchmark tests nine off-the-shelf large language models. Here, we evaluated Artificial Societies on the benchmark’s measures, with the same respondents, the same items, and the same scoring code, and set the results beside the nine LLMs.

Artificial Societies has the lowest error on every measure

Artificial Societies beside the nine benchmarked LLMs. Lower error is better.

Four charts of the error of Artificial Societies and the nine benchmarked LLMs on the benchmark’s measures; lower error is better. Population answer distributions: Artificial Societies 0.08, then Qwen3.8 0.31, Mistral Small 2603 0.34, Claude Opus 5 0.35, Grok 4.3 0.36, Gemini 3.8 Flash 0.36, Gemma 4 0.37, DeepSeek Flash 0.39, Llama 4 Scout 0.41, GPT-5.6 Sol 0.44. Subgroup answer distributions: Artificial Societies 0.06, then Qwen3.8 0.22, Mistral Small 2603 0.24, Gemini 3.8 Flash 0.27, Gemma 4 0.29, GPT-5.6 Sol 0.31, Llama 4 Scout 0.32, Claude Opus 5 0.33, DeepSeek Flash 0.36, Grok 4.3 0.39. Question-wording effect: Artificial Societies 1.4 pp, then Llama 4 Scout 6.2 pp, Gemma 4 11.1 pp, Qwen3.8 18.0 pp, DeepSeek Flash 30.9 pp, Grok 4.3 36.7 pp, Claude Opus 5 46.1 pp, Mistral Small 2603 46.5 pp, GPT-5.6 Sol 49.6 pp, Gemini 3.8 Flash 49.8 pp. Spread of opinion: Artificial Societies 0.006, then Qwen3.8 0.091, Mistral Small 2603 0.092, Gemma 4 0.106, Gemini 3.8 Flash 0.117, GPT-5.6 Sol 0.153, Llama 4 Scout 0.156, Claude Opus 5 0.164, DeepSeek Flash 0.175, Grok 4.3 0.198.

Fig. 1 | Artificial Societies has the lowest error on every measure. Each panel measures how far simulated answers fall from real survey responses. Shorter bars and dots further left mean less error. Artificial Societies is in ember and the nine benchmarked LLMs are in grey, with the black dot marking the best LLM.1 (footnote)

What We Measured

Every model answered a set of survey questions by enacting each of the real respondents in seven survey sources, having been given that respondent’s recorded profile. Specifically, we used the 2024 General Social Survey batteries on confidence in institutions and interpersonal trust, the SAPA personality traits and facets, the Twin retest study, and the Jordan and Albertson survey experiments as sources. This produced 2,796 synthetic respondents and 68,280 answers, scored against what those people actually said.

Results

Artificial Societies is more accurate than every benchmarked LLM on every measure shown.

Population answer distributions. When measuring the top-line difference between a model’s predictions and real respondents, Artificial Societies is more accurate than state-of-the-art LLMs across every survey tested.

SourceBest LLMArtificial Societies
GSS 2024, confidence in institutions0.207 (Qwen3.8)0.042
GSS 2024, interpersonal trust0.155 (Gemini 3.8 Flash)0.058
SAPA personality traits0.376 (Qwen3.8)0.090
SAPA personality facets0.336 (Qwen3.8)0.117
Twin retest, categorical items0.268 (Qwen3.8)0.147
Jordan survey0.145 (Grok 4.3)0.096
Albertson survey0.199 (Grok 4.3)0.025

Across the seven sources, Artificial Societies’ simulated answer distributions needed an average adjustment of 8 percentage points to match human responses, compared with 31 percentage points for the next best model, Qwen3.8. This measure, called total variation distance (TVD), captures how closely the proportions choosing each answer match, where lower scores mean closer agreement. The rest of the LLMs fall between 0.34 and 0.44. The best LLM changes from source to source, while our margins range from 34% lower error on the Jordan experiment to 87% lower on the Albertson experiment.

Groups within the population. Clients often want to know how particular subgroups, such as women under thirty, rural voters, or graduates, will answer their questions. On the General Social Survey, where the benchmark scores answer distributions within demographic groups, Artificial Societies’ TVD is 0.063 against 0.222 for the best LLM.

The range of opinion. Real people disagree with one another. In contrast, LLMs cluster around an average answer, so a simulated population looks more unanimous than the real one. Here, error is how far the spread of simulated answers to each question differs from the spread of real answers. Artificial Societies reproduces how spread out real opinion is with less than a tenth of the best LLM’s error.

Reaction to the question asked. To conduct message testing, clients need to be confident that changes in question wording have the same effects in simulation as they do in real humans. The Albertson survey asks the same question with two separate wordings, and Artificial Societies reproduces the resulting shift in answers to within 1.4 percentage points, against 6.2 for the best LLM. In the Twin retest study, where questions were asked in two forms, Artificial Societies’ error rate was less than half of the best LLM.

Relationships between answers. Real answers hang together. Related answers should stay linked as strongly as they are in real people, and personality traits that are unrelated in people should stay unrelated in a simulation. On both tests, the link between answers and independently measured criteria and keeping different traits distinct, Artificial Societies is first.

What It Means

Matching a distribution is not, on its own, evidence that a simulation can stand in for research for real people. A general-purpose LLM asked to answer as a respondent produces plausible individual answers whose distributions drift a long way from the real population’s, whose range is too narrow, and whose reaction to a change of wording is far larger than real people’s. Artificial Societies is built to estimate what populations and their groups say and how they respond, and on the benchmark’s public sources it has the lowest error on every measure reported here. The framework, the benchmark, and the papers are public, so the same tests can be applied to any simulation that claims to represent people.

Notes

  1. 1.

    Population answer distributions is the gap between simulated and real answer distributions, averaged over seven survey sources (total variation distance). Subgroup answer distributions is the same gap within demographic groups (GSS 2024). Spread of opinion is the gap in how widely answers vary (standard deviation, GSS 2024). Question-wording effect is the error in how far answers shift when a question is reworded, in percentage points (Albertson experiment).