Knowledge base
What is silicon sampling?
By Artificial Societies
Published
Silicon sampling is the use of a large language model (LLM) as a stand-in survey sample: the model reads a respondent’s demographic backstory and answers a questionnaire as that person would. Artificial Societies, which builds audiences as networks of AI personas grounded in real individuals, reports 86% distribution accuracy across 1,000 surveys against a 91% human self-replication ceiling, the rate at which people repeat their own answers (Survey Evaluation Report, January 2026). Artificial Societies’ September 2026 validity framework counts that kind of match as one of eight tests.
Where does the term silicon sampling come from?
Silicon sampling comes from a paper by Lisa Argyle and five co-authors, Out of One, Many: Using Language Models to Simulate Human Samples (opens in a new tab), published online in Political Analysis on 21 February 2023. The authors conditioned GPT-3 on thousands of sociodemographic backstories taken from real participants in large US surveys, including the American National Election Studies (ANES). They called the resulting synthetic respondents silicon samples, and reported that the model’s answers went well beyond surface similarity to the human ones.
The same paper coined algorithmic fidelity, the property a silicon sample depends on. Argyle and colleagues argued that the biases inside GPT-3 are fine-grained and correlated with demographics, so conditioning the model on the right backstory reproduces that subgroup’s response distribution. Algorithmic fidelity is therefore a claim about the model for a given group, and only a comparison with human data can establish it. Where a model lacks fidelity for a group, its silicon sample still arrives complete, with every cell filled, and describes nobody in particular.
John Horton made a parallel case for economics in a January 2023 arXiv paper, Large language models as simulated economic agents: What can we learn from Homo silicus? (opens in a new tab) He described LLMs as implicit computational models of humans that researchers can give endowments, information and preferences, as economists do with Homo economicus, the rational agent of economic theory. Using GPT-3, he re-ran three classic experiments published between 1986 and 2002 and found results qualitatively similar to the originals.
How does a silicon sample work?
A silicon sample replaces the respondent and keeps the questionnaire. The researcher takes each respondent’s profile from an existing survey, writes it as a backstory, gives the model that backstory and a survey item, and records the answer. Bisbee and colleagues ran it at scale (Political Analysis, May 2024), with ChatGPT personas matched to 7,530 respondents from the 2016–2020 ANES and 30 synthetic responses per person.
The usual benchmark then measures how far the silicon distribution for each question sits from the human distribution for the same question. Artificial Societies’ September 2026 validity framework counts that test, distributional fit, as one of two that make up today’s industry standard. The framework argues that it is easy to satisfy without modelling the population at all.
What have studies of silicon samples found?
Three studies from 2023 and 2024 mark out what a silicon sample can support.
| Study | What it tested | What it found | What it means for your study |
|---|---|---|---|
| Argyle et al., Political Analysis, 2023 | GPT-3 conditioned on backstories of real US survey respondents | Silicon samples reproduced subgroup response patterns beyond surface similarity | The method can work for groups where the model has fidelity |
| Santurkar et al., arXiv, March 2023 | Model opinions against 60 US demographic groups, drawn from public opinion polls (OpinionsQA) | Substantial misalignment that persisted after steering a model towards a group; over-65s and widowed people among the groups poorly reflected | Naming a demographic in the prompt does not guarantee fidelity for it |
| Bisbee et al., Political Analysis, May 2024 | ChatGPT personas matched to 7,530 ANES respondents, rating their warmth towards 11 groups on feeling thermometers | Averages close to ANES; less variation than real surveys; 48% of regression coefficients significantly different, with the sign reversed in 32% of those; results changed with prompt wording and over three months | Matching averages does not license inference about relationships between answers |
Sources: Argyle et al., 2023 (opens in a new tab); Santurkar et al., 2023 (opens in a new tab); Bisbee et al., 2024 (opens in a new tab). All read 28 September 2026.
Bisbee’s result matters most to anyone commissioning synthetic research. The averages matched ANES, so a benchmark on averages would have passed the sample. The relationships did not. Regression coefficients measure how a trait such as age or party relates to an answer. In Bisbee’s data, 48% of them differed significantly from ANES, and about 15% of all coefficients (32% of that 48%) pointed the wrong way. A team asking which kinds of voter feel warmest towards a group would have reached the opposite conclusion in those cases.
Santurkar and colleagues compared the opinions of language models with those of 60 US demographic groups, on topics ranging from abortion to automation. The gap was about as large as the divide between Democrats and Republicans on climate change, and explicitly steering a model towards a group did not close it.
What did the 2026 argument about AI polls say?
Two publications in spring 2026 set out the case against treating silicon samples as polls. The first was Eli McKown-Dawson’s Silver Bulletin article “AI polls” are fake polls (opens in a new tab), published on 11 April 2026. It draws its line in one sentence: “Silicon sampling, on the other hand, produces no new data.” McKown-Dawson still concedes evidence that some techniques replicate topline results quickly and cheaply, and the subtitle proposes treating them as models.
The American Association for Public Opinion Research (AAPOR) followed in May 2026 with a task force report, Responsible AI Integration in Survey Research (opens in a new tab). The task force ranks synthetic responses as the riskiest of the core AI tasks it considers, and separates three uses with different levels of risk: testing a questionnaire, filling in missing answers and replacing human respondents. The report prefers “synthetic responses” to “synthetic samples”, because the method estimates what respondents would say and is not a sampling design. It also asks researchers not to call AI-generated data a poll or a survey, since both words imply human respondents.
The task force’s methodological point is the one silicon sampling benchmarks struggle with. LLM responses can sometimes approximate the marginal distribution of a single item, meaning the spread of answers to one question. They frequently fail, the report adds, to capture the nuance and heterogeneity of real respondents. It expects evaluation to move from matching the average answer to one question towards reproducing correlations, joint distributions (how answers to several questions combine) and subgroup interactions.
How does Artificial Societies’ approach differ from silicon sampling?
Artificial Societies starts from a different unit than a silicon sample, which describes a person by category. Our method and evaluation page describes the standard way to build a synthetic persona: prompt an LLM with an invented biography, such as “you are a 34-year-old teacher from Ohio”, which the page says produces generic, stereotyped answers. Argyle’s backstories came from real respondents, so the two methods differ, but both hand the model a demographic summary and ask it to fill in the rest.
We ground each persona in a real individual instead. The inputs are observations of how people communicate and behave: published research, online commentary, reviews, anonymised social data, and public records of voting and investment. Psychometric methods reconstruct how each individual reasons, and every persona is checked against audience benchmarks and prior behaviour before it joins a society. Artificial Societies then connects personas through relationships and patterns of influence, at 12 to 3,500 personas per society. A silicon sample is a set of independent respondents; a society lets you study how opinion moves between people, the question generative agents research explores.
| Measure | Artificial Societies | Biography-prompted LLMs | Human comparator |
|---|---|---|---|
| Distribution accuracy (overlap with human opinion distributions) | 86% | 67% | 91% self-replication ceiling |
| Hallucination rate (personas contradicting themselves; lower is better) | Under 2% | Up to 35% | About 9% in human panels |
| Internal coherence (Cronbach’s alpha across questions on one attitude) | 89% | Not reported | 60–95% human range; 70% threshold |
Source: Artificial Societies, Survey Evaluation Report, January 2026, across 1,000 real-world surveys, as presented on our method and evaluation page.
Artificial Societies’ 86% belongs to the family of marginal-matching benchmarks that silicon sampling studies report, so on its own it cannot show that the relationships between answers are right. Artificial Societies’ evaluation page, as of 28 September 2026, reports no result for five of the framework’s eight tests: repeated questions, perturbation, trait separation, structure match and effect-size match.
Why is matching survey marginals not validity?
Matching survey marginals falls short of validity because a marginal describes one question at a time. A marginal distribution is the spread of answers to a single question across its options, for example 20% strongly agree and 35% agree. A simulation can reproduce how many people gave each answer without reproducing who gave which combination of answers. Artificial Societies’ September 2026 framework, Distribution Matching Is a Bad Way to Evaluate Simulations by Edoardo Chidichimo and Felix Wallis, gives an example. In it, a human audience and a synthetic one give identical answer distributions to “I intend to buy this” and “It is worth the price”, yet their joint distributions differ by a total variation distance of 0.4. Put plainly, 40% of one audience would have to change its combination of answers to match the other. In the human audience one attitude drives both answers; in the synthetic one the two are unrelated.
Even a matched joint distribution counts only if the human data was held out, meaning the model has never seen it, so that agreement tests the model rather than its memory. The framework notes that major public survey programmes appear in LLM training data, so only unpublished or newly fielded studies count as held out. ANES, the benchmark silicon sampling grew up on, is the kind of published programme a model may already have read. The rule applies to Artificial Societies too: its evaluation page does not say which of the 1,000 surveys were held out of model training data, as of 28 September 2026. The published experiments in Artificial Societies’ September 2026 misinformation study do not meet it either, although one matched a 6.6-point human drop in firm COVID-19 vaccine intent with an 8.2-point synthetic one.
Consistency needs its own benchmark, because people do not repeat themselves perfectly. Artificial Societies’ January 2026 Survey Evaluation Report uses a 91% human self-replication ceiling: asked the same question twice, people give the same answer roughly 91% of the time. A simulation that repeats itself more reliably than that behaves less like a person. The framework therefore fails a persona that never varies, just as it fails one that varies at random. Our guide to evaluating a synthetic research vendor turns the framework’s eight tests into questions for any supplier.
When is a silicon sample the right tool?
A silicon sample is the right tool when you need a fast estimate and will label it as a model. AAPOR’s task force says synthetic responses carry particularly serious validity and disclosure risks once used beyond clearly labelled pretesting, pilot work and exploratory diagnostics. A pretest of that kind runs a questionnaire through a model to catch broken skip patterns and ambiguous wording before any human sees it.
For a consequential decision, ask which validity tests a simulation has passed for your audience, on data it has not seen. A synthetic audience built from observations of real people and tested against held-out experiments answers a narrower question than a poll: how a specific group is likely to react to something new, and why.
The unsolved part concerns the groups a decision can turn on. Santurkar’s study found over-65s and widowed people among the groups language models reflected poorly. AAPOR’s task force judges it likely that the populations most expensive to reach by survey are the ones synthetic responses represent least accurately, and those are the audiences for which a silicon sample is most tempting to use.
Frequently asked questions
Is silicon sampling the same as a synthetic audience?
Silicon sampling is one way to build a synthetic audience: it prompts a model with a demographic description and collects independent answers. AAPOR’s May 2026 report describes work that goes further. Park and colleagues (2024) built synthetic respondents from interview transcripts, and those outperformed direct demographic prompts on survey items and personality inventories. Our explainer on what a synthetic audience is sets out the main types.
Do silicon samples lean in a political direction?
The studies cited here point left. Santurkar and colleagues confirmed left-leaning tendencies in some language models tuned on human feedback. AAPOR’s May 2026 report cites a German study in which ChatGPT 3.5 estimates of 2017 vote choice leaned towards the Green and Left parties, and calls liberal bias one of the prevailing findings across models and countries. Ask a vendor for results by political subgroup as well as totals.
Can a silicon sample replace a poll?
No, on AAPOR’s May 2026 guidance. The report notes that synthetic responses lack the properties of random samples, so their errors are systematic and can pass standard quality checks unnoticed. A reported sample size or confidence interval means little without the number of human respondents, which the task force asks researchers to disclose, along with a clear label on any data created by AI.
Does a newer language model fix the problems Bisbee found?
None of the sources behind this article re-tests Bisbee’s coefficient problem on a newer model, so the question is open. What the sources do show is drift: the same prompt gave significantly different results within three months, and AAPOR’s task force notes that providers retrain and realign models without notice. A validation therefore holds for the model version tested, and a vendor should repeat it when the model changes.
How is algorithmic fidelity measured?
Algorithmic fidelity is measured by comparing silicon samples with human samples on the same questions. Argyle and colleagues proposed criteria for it and compared response patterns across subgroups. AAPOR adds practical checks: unrealistically high refusal or don’t-know rates, answer options that do not exist, contradictory combinations across items, and signs that the model is reproducing memorised training text. Its preferred design keeps at least some human respondents in the study for comparison.
Is Artificial Societies the same as the academic term artificial societies?
Artificial Societies is a company, founded in October 2024 and headquartered in London; its product, Radiant, simulates audiences as networks of AI personas. The academic term comes from Joshua Epstein and Robert Axtell’s Growing Artificial Societies (opens in a new tab) (1996), which used the Sugarscape model to show social structures emerging from interacting agents. The company and the research tradition are distinct, although both study how individual behaviour produces group outcomes.
Sources
- Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C. and Wingate, D., Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis 31(3), 337–351, published online 21 February 2023. https://doi.org/10.1017/pan.2023.2 (opens in a new tab) Read 28 September 2026.
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B. and Larson, J. M., Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4), 401–416, May 2024. https://doi.org/10.1017/pan.2024.5 (opens in a new tab) Read 28 September 2026.
- Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P. and Hashimoto, T., Whose Opinions Do Language Models Reflect?, arXiv, March 2023. https://arxiv.org/abs/2303.17548 (opens in a new tab) Read 28 September 2026.
- Horton, J. J., Large language models as simulated economic agents: What can we learn from Homo silicus?, arXiv, January 2023. https://arxiv.org/abs/2301.07543 (opens in a new tab) Read 28 September 2026.
- McKown-Dawson, E., “AI polls” are fake polls, Silver Bulletin, 11 April 2026. https://www.natesilver.net/p/ai-polls-are-fake-polls (opens in a new tab) Read 28 September 2026.
- AAPOR Task Force on Responsible AI Integration in Survey Research (co-chairs Rothschild, D. M. and Marlar, J.), Responsible AI Integration in Survey Research, American Association for Public Opinion Research, May 2026. https://aapor.org/wp-content/uploads/2026/05/Responsible-AI-Integration-In-Survey-Research.pdf (opens in a new tab) Read 28 September 2026.
- Artificial Societies, Survey Evaluation Report, January 2026, as presented on the method and evaluation page. Read 28 September 2026.
- Chidichimo, E. and Wallis, F., Artificial Societies, Distribution Matching Is a Bad Way to Evaluate Simulations, 9 September 2026. Read 28 September 2026.
- Artificial Societies, Simulating Misinformation, Fact-Checking, and Prebunking, 23 September 2026. Read 28 September 2026.
- Epstein, J. M. and Axtell, R. L., Growing Artificial Societies: Social Science From the Bottom Up, Brookings Institution Press and MIT Press, 1996. https://www.brookings.edu/books/growing-artificial-societies/ (opens in a new tab) Read 28 September 2026.