Knowledge base
Can AI replace market research? Evidence explained
By Artificial Societies
Published
No, AI cannot replace market research or the contact with real people that research requires. Artificial Societies uses networks of AI personas to simulate high-value audiences. We test decisions involving confidential material or audiences you cannot recruit, while human research remains appropriate for general population surveys and in-person ethnographic research. The useful distinction is between a study you can responsibly field and a decision you need to rehearse before anyone encounters it.
How accurate is AI simulation for market research?
Our 86% distribution accuracy is 95% of the 91% human self-replication ceiling, measured against 1,000 real-world surveys in the Artificial Societies Survey Evaluation Report, January 2026 (opens in a new tab). Distribution accuracy measures the overlap between simulated and human opinion distributions. The ceiling reflects people’s agreement with their own answers when asked again. It gives the result a human reference point, while the evaluation’s biography-prompted large language models (LLMs) provide a baseline for personas created from invented biographies.
| Approach | Distribution accuracy: overlap with human opinion | Self-contradiction rate |
|---|---|---|
| Artificial Societies | 86% | Under 2% |
| Biography-prompted LLMs | 67% | Up to 35% |
| Human comparators | Self-replication ceiling 91% | Human panels about 9% |
Source: Artificial Societies, Survey Evaluation Report, January 2026, presented on our method and evaluation page. Self-contradiction means inconsistent answers; the measure does not test factual errors about the world.
Distribution overlap gives you evidence about the spread of opinions. It leaves a separate question about whether people holding one belief also hold another. Our September 2026 validity framework explains why matching answer totals cannot establish those relationships. For a researcher examining segments, the combinations of answers matter alongside the totals.
What does Nielsen Norman Group say about synthetic user research?
On research with reachable users, we agree with Nielsen Norman Group’s critique. Edoardo Chidichimo and Felix Wallis call criticism of synthetic research “well-warranted” in Distribution Matching Is a Bad Way to Evaluate Simulations, September 2026. That footnote cites Bisbee and colleagues, Sarstedt and colleagues, and Harding and colleagues. We regard those objections as questions our method must answer.
Moran and Rosala’s Nielsen Norman Group article (opens in a new tab), published in June 2024 and updated in September 2025, states: “Synthetic users cannot replace the depth and empathy gained from studying and speaking with real people.” We agree. A simulation of a reachable user does not provide the contact with that person that the authors defend.
The article also gives synthetic users a role: “Use synthetic users to help you prepare for research studies with real users.” It asks researchers to treat synthetic findings as hypotheses for testing. Preparation is a defensible use, and it fits our recommendation to explore possibilities before commissioning targeted human research. Our own published boundary is explicit: “Synthetic audiences will never be a true substitute for the people they seek to represent.” That is the position in our September 2026 post, including when simulation lets us run a study recruitment cannot support.
How does AI audience simulation work?
Artificial Societies starts with observations of people. Our method page describes published research, online commentary and reviews, anonymised social data, and public records of voting and investment. Persona construction uses psychometric methods to reconstruct belief systems: how an individual reasons, rather than a biography alone. We check each persona against audience benchmarks and prior behaviour before connecting it to a society through observed relationships and patterns of influence.
The Simulation Engine exposes those personas to scenarios and stimuli. You can examine the resulting opinions at population, segment or persona level. This is the mechanism behind testing a strategic narrative: the audience has a basis in observed people, and the analysis can examine differences within that audience. Diverse observations are a prerequisite. Where neither public nor client-held observations exist, the method has nothing to ground personas in; accuracy on an observed population does not transfer to an unobserved one.
Our peer-reviewed paper on group behaviour (He, Wallis, Gvirtz and Rathje, British Journal of Psychology (opens in a new tab), December 2024) supports the network modelling, not the survey accuracy figures; the science page explains what it found.
What do the MRS and Esomar say about synthetic research?
The Market Research Society (MRS) treats representation and sampling quality as central concerns in its Delphi Group report, Using synthetic respondents for market research (opens in a new tab), published in 2024. Its technical objection is direct: “LLMs are good at averaging, so deliver few surprises and level results into generic information.” The report also raises concern about obscuring the distinction between primary research and synthetic data. Those concerns make the origin of a persona and the tests applied to its answers relevant to a buyer.
The ICC/Esomar International Code on Market, Opinion and Social Research and Data Analytics (opens in a new tab), effective September 2025, sets disclosure duties. Article 7(e) requires clients to be informed about AI use, including synthetic data and personas, and about the extent of human oversight. Article 9(b) extends disclosure to significant use in sampling, deployment, analysis or interpretation.
The MRS document is a report; the ICC/Esomar document is a code. Neither calls for a ban on synthetic research. We publish our method, data sources and evaluation so a researcher can inspect how the work is produced. For your own study, specify which findings come from simulation and describe the human oversight involved. A presentation should preserve that distinction when its findings leave the research team.
What are the limits of synthetic market research?
Synthetic market research inherits the gap between stated intentions and behaviour. Our September 2026 validity framework acknowledges that answers to a survey can be weak predictors of action. Matching those answers is evidence about reported opinion. A decision that requires observed behaviour therefore still needs research capable of observing it, even when simulation has helped identify the question to investigate.
Bisbee and colleagues’ Synthetic Replacements for Human Survey Data? (opens in a new tab) appeared in Political Analysis in 2024; Gao and colleagues published Take caution in using LLMs as human surrogates (opens in a new tab) in the Proceedings of the National Academy of Sciences in 2025. Lukauskas and Šarkauskaitė’s 2026 preprint (opens in a new tab), which is not peer-reviewed, states: “LLM samples are not a drop-in replacement for human survey data”. A fluent answer alone cannot settle those concerns.
Our response is to examine how answers behave under testing. The September 2026 framework sets out eight tests of whether a stimulus caused a response shift, whether questions measure the intended attitude, and whether results hold for real populations. We state: “We hold ourselves to all eight.” Reordering answers should not move opinion; changing the evidence behind a claim should. The distinction matters when a simulated audience appears persuaded: you need to establish whether it responded to the argument or to the way the question invited agreement.
When should you use audience simulation instead of human participants?
Audience simulation addresses barriers that occur in our case studies: confidential content, the effect of asking, and access to the audience. For an earnings-call engagement, restrictions around selective disclosure prevented testing draft material with real investors. In a transport engagement, exposing the audience to draft messages would have changed the debate under study. These are reasons to simulate before acting, with the results treated as forecasts of reactions.
Teneo’s engagement concerned audiences that were impossible to reach. Teneo’s own account of the engagement is in our Teneo case study.
Simulation also permits interrogation of belief systems and their response to a stimulus, as our September 2026 validity framework describes. In the earnings-call case study, simulated analysts’ questions matched the live call, and simulated press coverage matched the next morning’s framing. That account illustrates a decision rehearsal followed by a real event. It is a single engagement, separate from the survey evaluation, and should be read at that scope.
How should you combine AI simulation with traditional market research?
We would use Artificial Societies to explore strategic options, then commission targeted human research to test the findings that determine the decision. You can also reverse the sequence: establish a baseline through traditional research, then use simulation to test confidential content against it. The choice follows the material you can disclose and the audience you can reach.
For a straightforward general population survey about a publicly available product, we would recommend a traditional panel. Biography-prompted LLMs may suffice for low-stakes directional feedback. Networks of enriched personas fit decisions where you need to examine stakeholder reasoning and relationships.
For validation, our September 2026 framework calls for unpublished or newly fielded studies, since public survey programmes can appear in model training data. Agree on the human evidence that will test your finding before treating simulation as support for a strategic recommendation. With a confidential announcement, that means distinguishing what you can test with people before release from the reactions you will compare with the simulation after it becomes public.
Frequently asked questions
How should you assess claims about synthetic user research on Reddit?
Ask for the method and the human comparison behind the claim. Our evaluation defines its metrics, while our validity framework explains why plausible speech and matching answer totals leave questions about belief structure. Apply those questions to us too. A discussion can identify a concern worth testing; inspect the underlying study before treating its conclusion as evidence for your audience.
What is the difference between synthetic data and a synthetic persona?
The 2025 ICC/Esomar Code defines synthetic data as information generated to replicate characteristics of real-world data. A synthetic persona is a digital representation intended to mimic people’s behaviours, preferences and characteristics. The distinction identifies what you are evaluating: generated information or a representation of a person. Our method uses personas grounded in observations and connected through relationships.
Can Artificial Societies use our existing customer research?
Our method page allows proprietary first-party data to be added to the persona set. We segregate it by engagement, hold it in the EU and do not use it to train our models. When discussing a study, identify which observations your research provides about the audience and which questions you still need the simulation to examine.
Can the free data-quality check establish that a simulation is accurate?
Our free data-quality check examines dataset quality. It does not establish external accuracy, whether the data is current, or whether the questions were well designed. It also does not determine whether a dataset is human or synthetic. Use the check for its stated purpose, then assess the findings against evidence from the population you want to understand.
How can researchers test whether an AI persona is just agreeing with them?
Our September 2026 validity framework discusses sycophancy: personas agreeing with what a question implies. It describes tests using equivalent question wordings and changes to response order, alongside changes to the substance of a claim. The research task is to separate agreement prompted by presentation from a response to evidence, with persona agreement assessed against human test–retest bands.
Sources
- Artificial Societies, method and evaluation. Source record dated 17 September 2026.
- Artificial Societies, Survey Evaluation Report (opens in a new tab), January 2026. Source record dated 17 September 2026.
- Edoardo Chidichimo and Felix Wallis, Distribution Matching Is a Bad Way to Evaluate Simulations, Artificial Societies, September 2026. Source record dated 17 September 2026.
- Moran, K., and Rosala, M., Synthetic Users: If, When, and How to Use AI-Generated “Research” (opens in a new tab), Nielsen Norman Group, 21 June 2024; updated 24 September 2025. Accessed 17 September 2026.
- MRS Delphi Group, Using synthetic respondents for market research (opens in a new tab), 2024. Accessed 17 September 2026.
- International Chamber of Commerce and Esomar, ICC/Esomar International Code on Market, Opinion and Social Research and Data Analytics (opens in a new tab), 2025 revision. Accessed 17 September 2026.
- Bisbee et al., Synthetic Replacements for Human Survey Data? The Perils of Large Language Models (opens in a new tab), Political Analysis, 2024. Accessed 17 September 2026.
- Gao et al., Take caution in using LLMs as human surrogates (opens in a new tab), Proceedings of the National Academy of Sciences, 2025. Accessed 17 September 2026.
- Lukauskas, M., and Šarkauskaitė, V., Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents (opens in a new tab), preprint, 2026. Accessed 17 September 2026.
- He, Wallis, Gvirtz and Rathje, Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), British Journal of Psychology, published online December 2024. Source record dated 17 September 2026.
- Artificial Societies, case studies, including Teneo, the earnings-call engagement and the transport engagement. Source record dated 17 September 2026.
- Artificial Societies, data-quality check. Source record dated 17 September 2026.