Knowledge base
What are synthetic personas?
By Artificial Societies
Published
Synthetic personas are simulated people: AI models, built on large language models, that answer questions and react to material as a specified person or type of person would. Artificial Societies, which builds networks of AI personas grounded in real individuals, measures 86% distribution accuracy for its personas across 1,000 surveys, against 67% for personas prompted with an invented biography and a 91% human ceiling (Survey Evaluation Report, January 2026). The 19-point gap between the two reflects how the personas were built, not what they are called.
The ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics treats a synthetic persona as a generated representation of a person that imitates how real people or groups behave and what they prefer (2025 edition). That definition covers a one-line chatbot prompt and an agent built from an interview with the person it stands for, and the two behave very differently under questioning.
How are synthetic personas built?
A synthetic persona can be built from an invented biography, from observations of a real person, or from that person’s own interview and survey answers. The method decides what you can later check the persona against.
Invented-biography prompting needs no data about any real person. Our method and evaluation page describes it as the standard approach: a language model receives a biography such as “you are a 34-year-old teacher from Ohio” and answers in character, which produces generic, stereotyped answers. The model supplies everything the biography leaves out from its training data, which is the likely source of that stereotyping.
Grounded personas start from evidence about a real person. At Artificial Societies each persona begins with a real individual, observed through published research, online commentary, reviews, anonymised social data and public records of voting and investment, from sources including LinkedIn, X, Reddit, YouTube and Amazon. Psychometric methods reconstruct how that individual reasons, and each persona is validated against audience benchmarks and prior behaviour. The persona is judged as a member of an audience, not scored against its original person, which is what separates it from a digital twin. Our explainer on AI persona simulation sets out the full pipeline.
Academic work adds agents built from a person’s own self-reports. Park and colleagues built agents for 1,052 Americans by applying language models to qualitative interviews about their lives, and in the paper’s June 2026 version from survey answers as well (arXiv 2411.10109 (opens in a new tab)). Each agent stands for one participant and is scored against that person’s answers to questions it never saw, which makes it a digital twin.
Output quality tracks the method. In Artificial Societies’ January 2026 Survey Evaluation Report, biography-prompted LLM personas contradicted themselves in up to 35% of cases, against under 2% for Artificial Societies’ grounded personas and about 9% for human panels. On distinct-word richness, a measure of how varied the vocabulary of open answers is, the same report scores biography-prompted personas at 61%, grounded personas at 93% and a benchmark of 120,000 human social-media posts at 91%. The contradiction rate matters once you read a segment’s reasoning instead of its totals.
How do synthetic personas differ from synthetic respondents, digital twins and AI personas?
The four terms overlap because each names a different aspect of one object. The persona is the simulated person, and a respondent is the role it plays when it answers a questionnaire. A twin is a persona built from one individual’s own data and judged against that individual. “AI persona” is the everyday synonym, and the one our own pages use.
| Term | What it names | Unit | Tied to one real individual? | Evidence you can ask for |
|---|---|---|---|---|
| Synthetic persona | A simulated person, however it was built | One person | Only if grounded or interview-based | Depends on the construction method |
| Synthetic respondent | A synthetic persona in the role of survey or interview participant | One set of answers | Only if grounded or interview-based | Match to human answer distributions |
| Digital twin | A persona built from one person’s own self-reports | One specific individual | Always, and judged against that individual | Match to that person’s answers on questions not used to build it |
| AI persona | Everyday synonym for synthetic persona, used on our pages | One person | Only if grounded or interview-based | As for a synthetic persona |
| Synthetic audience | A group of personas; in Artificial Societies’ construction, connected by relationships and influence | A group (12 to 3,500 personas at Artificial Societies) | Its members can be; results are judged at group level | Group distributions, segment results and effect sizes |
Sources: ICC/ESOMAR International Code, 2025 edition; AAPOR task force report, 2026; Park and colleagues, June 2026; Artificial Societies, method and evaluation, read 28 September 2026.
A 2026 task force report accepted by the American Association for Public Opinion Research (AAPOR) prefers “synthetic responses” to “synthetic samples”, because generating answers estimates what a set of people would say; it does not sample them. Our page on synthetic respondents compares these terms against industry definitions.
How accurate are synthetic personas?
The accuracy figures for synthetic personas in this section answer two different questions. Individual accuracy asks whether a persona gives the answers its real counterpart gave. Distribution accuracy asks whether a population of personas splits across the answer options the way a population of people does. Park and colleagues and Artificial Societies read their figures against human test–retest consistency, the rate at which people repeat their own answer when asked again; Peng and colleagues report a raw correlation.
| Study | Personas | What was measured | Result |
|---|---|---|---|
| Park and colleagues, version 1, November 2024 | Agents for 1,052 Americans, built from qualitative interviews | General Social Survey answers, relative to participants’ own consistency two weeks later | 85% |
| Park and colleagues, version 3, June 2026 | Agents for 1,052 Americans, from interviews, surveys or both | Survey items not used to build the agents, relative to the same two-week consistency | 86% combined; 83% interview only; 82% survey only; 74% demographics only |
| Peng, Toubia and colleagues, April 2026 | Twins of over 2,000 people, from answers to over 500 questions | Correlation with each person’s responses across 164 outcomes in 19 pre-registered studies | Average r = 0.20 |
| Artificial Societies, January 2026 | Networks of personas grounded in real individuals | Overlap with human opinion distributions across 1,000 surveys | 86%, against 67% for biography-prompted LLMs and a 91% human ceiling |
Sources: arXiv 2411.10109 (opens in a new tab); arXiv 2509.19088 (opens in a new tab); Artificial Societies, Survey Evaluation Report, January 2026, as published on the method and evaluation page. All read 28 September 2026.
The two 86% figures in that table measure different things. Park and colleagues’ figure is a ratio: the agents’ match rate divided by the participants’ own two-week consistency, so the share of answers an agent actually matched is lower than 86%. Artificial Societies’ figure is the overlap between simulated and human opinion distributions, five points below the 91% at which people reproduce their own answers. It establishes that a population of personas splits the right way, but not that any single persona answered as a particular person would; our accuracy page covers Artificial Societies’ other published measures.
Peng, Toubia and colleagues report a weaker result for digital twins. Twins built from each person’s answers to over 500 questions correlated with that person’s responses at an average r of 0.20, where 1 is perfect agreement and 0 is none (Digital Twins as Funhouse Mirrors (opens in a new tab), version of 19 April 2026). They were only modestly more accurate than a base language model given no individual data. At that correlation, a twin’s answer tells you little about which way its individual will go.
Where do synthetic personas fail?
Peng, Toubia and colleagues’ April 2026 paper, Digital Twins as Funhouse Mirrors, named five distortions in digital twins. They are the failure modes to test any synthetic persona for:
Insufficient individuation: twin answers were less spread out than human answers in 154 of 164 outcomes (93.9%), so the simulated population is more uniform than the real one.
Stereotyping: twins relied heavily on demographic characteristics and made little use of the individual information they were given.
Representation bias: twins were more accurate for participants with higher education and higher income.
Ideological bias: twins expressed more pro-human views than the people they modelled, and also showed pro-technology tendencies.
Hyper-rationality: twins tended to display perfect knowledge on questions with a correct answer and were more likely than people to choose the textbook-rational answer.
Artificial Societies’ September 2026 validity framework by Edoardo Chidichimo and Felix Wallis, Distribution Matching Is a Bad Way to Evaluate Simulations, describes three further failures that a fluent answer hides. Sycophancy is a language-model persona agreeing with whatever a question implies, far more often than people do, so confident wording can move a persona that was never persuaded. Weak discrimination is a persona rating a new product high on quality, value, trustworthiness and likelihood to recommend at once, although in real consumer data those judgements come apart. The third is exposure: major public survey programmes appear in language-model training data, so matching them tests memory rather than prediction.
Two limits sit outside the model. What people say is often a weak predictor of what they later do, and synthetic personas inherit that intention–behaviour gap from the surveys they are scored against. An audience with no public trace and no first-party research gives a persona nothing to be grounded in, and a taste trial or other sensory test needs real people handling the real product.
How does a synthetic persona differ from a synthetic audience?
A synthetic persona is one simulated person; a synthetic audience is a group of them, and in a networked build its members influence each other. Artificial Societies clusters validated personas by shared characteristics and connects them through real relationships and patterns of influence, in societies of 12 to 3,500, according to its method page (read 28 September 2026). Artificial Societies’ crisis simulation study for a consumer goods company built 1,174 personas across five societies and mapped each audience’s social influence network before testing 24 messages.
Questioning personas one at a time cannot show whether a message that persuades individuals still holds once they hear from the people they trust. Testing that needs a group with social structure. A paper by Artificial Societies researchers, He, Wallis, Gvirtz and Rathje in the British Journal of Psychology (opens in a new tab) (December 2024), found that 33,299 chatbots formed communities around a shared language, as people do. That shows a society of language-model agents can form human-like structure; it does not by itself show how persuasion changes inside it. What is a synthetic audience? and our article on networked audience simulation take the group question further.
When should you use synthetic personas?
Synthetic personas fit a question about opinion in an observable audience that is slow, impossible or unwise to ask directly. Artificial Societies’ hyperscaler earnings-call study is one such case: rules on selective disclosure meant unreleased numbers and draft scripts could never be shown to a real investor. For that study, Artificial Societies built more than 2,400 personas, each matched one-to-one to a real individual using public data and social activity, including 14 for the sell-side analysts who had asked questions on the company’s last eight calls. According to the case study, the live call’s questions ran the way the simulated analysts had predicted, which is one engagement reported qualitatively, not an accuracy score.
AAPOR’s 2026 task force report names a lower-stakes use: a survey pretest. Before fielding, a team prompts a language model, optionally under persona constraints, to complete the questionnaire repeatedly and find broken skip patterns, ambiguous wording and excessive completion time. That use needs plausible personas, not accurate ones.
Synthetic personas are the wrong tool when the decision needs proof of behaviour, when the test is physical, and when the result will be published as public opinion. The same AAPOR report says AI-generated cases are not research participants and must be identified as AI-created in any purported study of public opinion.
Frequently asked questions
Is a synthetic persona the same as a marketing or UX persona?
A marketing or user-experience persona is a written profile of an archetypal customer, used to keep a team’s decisions anchored to the people it serves. It describes a customer and cannot answer a question. A synthetic persona is a working model that responds to new material, so it can be surveyed, interviewed and shown a stimulus, and its answers can be checked against human data.
Do you have to disclose that research used synthetic personas?
Under the 2025 edition of the ICC/ESOMAR International Code, clients must be told when AI or other emerging technologies are used to compile datasets, analyse, report or interpret findings, including synthetic data and synthetic personas, and the extent of human oversight must be stated. Where a synthetic persona collects data from real people, for example as an AI interviewer, the Code requires that those people are told at the start.
Are synthetic personas biased?
Synthetic personas can be more accurate for some groups than for others. AAPOR’s 2026 task force report warns that synthetic responses miss nuance most often for marginalised groups, culturally specific concepts and emotionally charged topics. Park and colleagues report that agents grounded in interviews reduced accuracy gaps across racial and ideological groups, compared with agents given only demographic descriptions. Our article on AI persona bias covers how to test a persona set for representational fairness.
Can you build synthetic personas from your own customer research?
You can layer proprietary research into a persona set alongside public observations. Artificial Societies’ method page states that first-party data is segregated for each engagement, held in the EU and never used to train its models. Your own interviews or survey data give personas evidence that no public source holds, which matters most for customers who leave little trace online.
Can synthetic personas replace human research participants?
AAPOR’s 2026 task force report does not treat synthetic personas as a replacement for people. It states that the field has not reached consensus on when, if ever, AI-generated responses can stand in for human ones without changing what is being measured. Until that changes, a synthetic persona’s answer is evidence about a model of a person, and a decision that needs the person still needs a human study.
Is Artificial Societies the same as the academic term “artificial societies”?
The academic term refers to agent-based social simulation, associated with Epstein and Axtell’s Growing Artificial Societies (opens in a new tab) (1996), in which simple rule-following agents produce social patterns. Artificial Societies is a company, founded in October 2024 and headquartered in London, that builds networks of AI personas grounded in real individuals, through its product Radiant, so organisations can test decisions before making them.
Sources
- Artificial Societies, Survey Evaluation Report, January 2026, as published on the method and evaluation page. Read 28 September 2026.
- Edoardo Chidichimo and Felix Wallis, Artificial Societies, Distribution Matching Is a Bad Way to Evaluate Simulations, 9 September 2026. Read 28 September 2026.
- Artificial Societies, hyperscaler earnings-call case study and crisis simulation case study. Read 28 September 2026.
- ICC/ESOMAR, International Code on Market, Opinion and Social Research and Data Analytics (opens in a new tab), 2025 edition. Read 28 September 2026.
- Rothschild, Marlar and colleagues, AAPOR Task Force on Responsible AI Integration in Survey Research, Responsible AI Integration in Survey Research (opens in a new tab), American Association for Public Opinion Research, 2026. Read 28 September 2026.
- Park, Zou and colleagues, Generative Agent Simulations of 1,000 People (opens in a new tab), arXiv 2411.10109, version 1, 15 November 2024; current version retitled LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (opens in a new tab), version 3, 28 June 2026. Read 28 September 2026.
- Peng, Gui, Brucks and colleagues, including Toubia, Digital Twins as Funhouse Mirrors: Five Key Distortions (opens in a new tab), arXiv 2509.19088, version 5, 19 April 2026 (first version 23 September 2025). Read 28 September 2026.
- He, Wallis, Gvirtz and Rathje, Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), British Journal of Psychology, published online December 2024. Read 28 September 2026.
- Epstein and Axtell, Growing Artificial Societies: Social Science from the Bottom Up (opens in a new tab), Brookings Institution Press and MIT Press, 1996. Read 28 September 2026.