Skip to content

Synthetic personas are simulated people: AI models, built on large language models, that answer questions and react to material as a specified person or type of person would. Artificial Societies, which builds networks of AI personas grounded in real individuals, measures 86% distribution accuracy for its personas across 1,000 surveys, against 67% for personas prompted with an invented biography and a 91% human ceiling (Survey Evaluation Report, January 2026). The 19-point gap between the two reflects how the personas were built, not what they are called.

The ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics treats a synthetic persona as a generated representation of a person that imitates how real people or groups behave and what they prefer (2025 edition). That definition covers a one-line chatbot prompt and an agent built from an interview with the person it stands for, and the two behave very differently under questioning.

How are synthetic personas built?

A synthetic persona can be built from an invented biography, from observations of a real person, or from that person’s own interview and survey answers. The method decides what you can later check the persona against.

Invented-biography prompting needs no data about any real person. Our method and evaluation page describes it as the standard approach: a language model receives a biography such as “you are a 34-year-old teacher from Ohio” and answers in character, which produces generic, stereotyped answers. The model supplies everything the biography leaves out from its training data, which is the likely source of that stereotyping.

Grounded personas start from evidence about a real person. At Artificial Societies each persona begins with a real individual, observed through published research, online commentary, reviews, anonymised social data and public records of voting and investment, from sources including LinkedIn, X, Reddit, YouTube and Amazon. Psychometric methods reconstruct how that individual reasons, and each persona is validated against audience benchmarks and prior behaviour. The persona is judged as a member of an audience, not scored against its original person, which is what separates it from a digital twin. Our explainer on AI persona simulation sets out the full pipeline.

Academic work adds agents built from a person’s own self-reports. Park and colleagues built agents for 1,052 Americans by applying language models to qualitative interviews about their lives, and in the paper’s June 2026 version from survey answers as well (arXiv 2411.10109 (opens in a new tab)). Each agent stands for one participant and is scored against that person’s answers to questions it never saw, which makes it a digital twin.

Output quality tracks the method. In Artificial Societies’ January 2026 Survey Evaluation Report, biography-prompted LLM personas contradicted themselves in up to 35% of cases, against under 2% for Artificial Societies’ grounded personas and about 9% for human panels. On distinct-word richness, a measure of how varied the vocabulary of open answers is, the same report scores biography-prompted personas at 61%, grounded personas at 93% and a benchmark of 120,000 human social-media posts at 91%. The contradiction rate matters once you read a segment’s reasoning instead of its totals.

How do synthetic personas differ from synthetic respondents, digital twins and AI personas?

The four terms overlap because each names a different aspect of one object. The persona is the simulated person, and a respondent is the role it plays when it answers a questionnaire. A twin is a persona built from one individual’s own data and judged against that individual. “AI persona” is the everyday synonym, and the one our own pages use.

TermWhat it namesUnitTied to one real individual?Evidence you can ask for
Synthetic personaA simulated person, however it was builtOne personOnly if grounded or interview-basedDepends on the construction method
Synthetic respondentA synthetic persona in the role of survey or interview participantOne set of answersOnly if grounded or interview-basedMatch to human answer distributions
Digital twinA persona built from one person’s own self-reportsOne specific individualAlways, and judged against that individualMatch to that person’s answers on questions not used to build it
AI personaEveryday synonym for synthetic persona, used on our pagesOne personOnly if grounded or interview-basedAs for a synthetic persona
Synthetic audienceA group of personas; in Artificial Societies’ construction, connected by relationships and influenceA group (12 to 3,500 personas at Artificial Societies)Its members can be; results are judged at group levelGroup distributions, segment results and effect sizes

Sources: ICC/ESOMAR International Code, 2025 edition; AAPOR task force report, 2026; Park and colleagues, June 2026; Artificial Societies, method and evaluation, read 28 September 2026.

A 2026 task force report accepted by the American Association for Public Opinion Research (AAPOR) prefers “synthetic responses” to “synthetic samples”, because generating answers estimates what a set of people would say; it does not sample them. Our page on synthetic respondents compares these terms against industry definitions.

How accurate are synthetic personas?

The accuracy figures for synthetic personas in this section answer two different questions. Individual accuracy asks whether a persona gives the answers its real counterpart gave. Distribution accuracy asks whether a population of personas splits across the answer options the way a population of people does. Park and colleagues and Artificial Societies read their figures against human test–retest consistency, the rate at which people repeat their own answer when asked again; Peng and colleagues report a raw correlation.

StudyPersonasWhat was measuredResult
Park and colleagues, version 1, November 2024Agents for 1,052 Americans, built from qualitative interviewsGeneral Social Survey answers, relative to participants’ own consistency two weeks later85%
Park and colleagues, version 3, June 2026Agents for 1,052 Americans, from interviews, surveys or bothSurvey items not used to build the agents, relative to the same two-week consistency86% combined; 83% interview only; 82% survey only; 74% demographics only
Peng, Toubia and colleagues, April 2026Twins of over 2,000 people, from answers to over 500 questionsCorrelation with each person’s responses across 164 outcomes in 19 pre-registered studiesAverage r = 0.20
Artificial Societies, January 2026Networks of personas grounded in real individualsOverlap with human opinion distributions across 1,000 surveys86%, against 67% for biography-prompted LLMs and a 91% human ceiling

Sources: arXiv 2411.10109 (opens in a new tab); arXiv 2509.19088 (opens in a new tab); Artificial Societies, Survey Evaluation Report, January 2026, as published on the method and evaluation page. All read 28 September 2026.

The two 86% figures in that table measure different things. Park and colleagues’ figure is a ratio: the agents’ match rate divided by the participants’ own two-week consistency, so the share of answers an agent actually matched is lower than 86%. Artificial Societies’ figure is the overlap between simulated and human opinion distributions, five points below the 91% at which people reproduce their own answers. It establishes that a population of personas splits the right way, but not that any single persona answered as a particular person would; our accuracy page covers Artificial Societies’ other published measures.

Peng, Toubia and colleagues report a weaker result for digital twins. Twins built from each person’s answers to over 500 questions correlated with that person’s responses at an average r of 0.20, where 1 is perfect agreement and 0 is none (Digital Twins as Funhouse Mirrors (opens in a new tab), version of 19 April 2026). They were only modestly more accurate than a base language model given no individual data. At that correlation, a twin’s answer tells you little about which way its individual will go.

Where do synthetic personas fail?

Peng, Toubia and colleagues’ April 2026 paper, Digital Twins as Funhouse Mirrors, named five distortions in digital twins. They are the failure modes to test any synthetic persona for:

  • Insufficient individuation: twin answers were less spread out than human answers in 154 of 164 outcomes (93.9%), so the simulated population is more uniform than the real one.

  • Stereotyping: twins relied heavily on demographic characteristics and made little use of the individual information they were given.

  • Representation bias: twins were more accurate for participants with higher education and higher income.

  • Ideological bias: twins expressed more pro-human views than the people they modelled, and also showed pro-technology tendencies.

  • Hyper-rationality: twins tended to display perfect knowledge on questions with a correct answer and were more likely than people to choose the textbook-rational answer.

Artificial Societies’ September 2026 validity framework by Edoardo Chidichimo and Felix Wallis, Distribution Matching Is a Bad Way to Evaluate Simulations, describes three further failures that a fluent answer hides. Sycophancy is a language-model persona agreeing with whatever a question implies, far more often than people do, so confident wording can move a persona that was never persuaded. Weak discrimination is a persona rating a new product high on quality, value, trustworthiness and likelihood to recommend at once, although in real consumer data those judgements come apart. The third is exposure: major public survey programmes appear in language-model training data, so matching them tests memory rather than prediction.

Two limits sit outside the model. What people say is often a weak predictor of what they later do, and synthetic personas inherit that intention–behaviour gap from the surveys they are scored against. An audience with no public trace and no first-party research gives a persona nothing to be grounded in, and a taste trial or other sensory test needs real people handling the real product.

How does a synthetic persona differ from a synthetic audience?

A synthetic persona is one simulated person; a synthetic audience is a group of them, and in a networked build its members influence each other. Artificial Societies clusters validated personas by shared characteristics and connects them through real relationships and patterns of influence, in societies of 12 to 3,500, according to its method page (read 28 September 2026). Artificial Societies’ crisis simulation study for a consumer goods company built 1,174 personas across five societies and mapped each audience’s social influence network before testing 24 messages.

Questioning personas one at a time cannot show whether a message that persuades individuals still holds once they hear from the people they trust. Testing that needs a group with social structure. A paper by Artificial Societies researchers, He, Wallis, Gvirtz and Rathje in the British Journal of Psychology (opens in a new tab) (December 2024), found that 33,299 chatbots formed communities around a shared language, as people do. That shows a society of language-model agents can form human-like structure; it does not by itself show how persuasion changes inside it. What is a synthetic audience? and our article on networked audience simulation take the group question further.

When should you use synthetic personas?

Synthetic personas fit a question about opinion in an observable audience that is slow, impossible or unwise to ask directly. Artificial Societies’ hyperscaler earnings-call study is one such case: rules on selective disclosure meant unreleased numbers and draft scripts could never be shown to a real investor. For that study, Artificial Societies built more than 2,400 personas, each matched one-to-one to a real individual using public data and social activity, including 14 for the sell-side analysts who had asked questions on the company’s last eight calls. According to the case study, the live call’s questions ran the way the simulated analysts had predicted, which is one engagement reported qualitatively, not an accuracy score.

AAPOR’s 2026 task force report names a lower-stakes use: a survey pretest. Before fielding, a team prompts a language model, optionally under persona constraints, to complete the questionnaire repeatedly and find broken skip patterns, ambiguous wording and excessive completion time. That use needs plausible personas, not accurate ones.

Synthetic personas are the wrong tool when the decision needs proof of behaviour, when the test is physical, and when the result will be published as public opinion. The same AAPOR report says AI-generated cases are not research participants and must be identified as AI-created in any purported study of public opinion.

Frequently asked questions

Is a synthetic persona the same as a marketing or UX persona?

A marketing or user-experience persona is a written profile of an archetypal customer, used to keep a team’s decisions anchored to the people it serves. It describes a customer and cannot answer a question. A synthetic persona is a working model that responds to new material, so it can be surveyed, interviewed and shown a stimulus, and its answers can be checked against human data.

Do you have to disclose that research used synthetic personas?

Under the 2025 edition of the ICC/ESOMAR International Code, clients must be told when AI or other emerging technologies are used to compile datasets, analyse, report or interpret findings, including synthetic data and synthetic personas, and the extent of human oversight must be stated. Where a synthetic persona collects data from real people, for example as an AI interviewer, the Code requires that those people are told at the start.

Are synthetic personas biased?

Synthetic personas can be more accurate for some groups than for others. AAPOR’s 2026 task force report warns that synthetic responses miss nuance most often for marginalised groups, culturally specific concepts and emotionally charged topics. Park and colleagues report that agents grounded in interviews reduced accuracy gaps across racial and ideological groups, compared with agents given only demographic descriptions. Our article on AI persona bias covers how to test a persona set for representational fairness.

Can you build synthetic personas from your own customer research?

You can layer proprietary research into a persona set alongside public observations. Artificial Societies’ method page states that first-party data is segregated for each engagement, held in the EU and never used to train its models. Your own interviews or survey data give personas evidence that no public source holds, which matters most for customers who leave little trace online.

Can synthetic personas replace human research participants?

AAPOR’s 2026 task force report does not treat synthetic personas as a replacement for people. It states that the field has not reached consensus on when, if ever, AI-generated responses can stand in for human ones without changing what is being measured. Until that changes, a synthetic persona’s answer is evidence about a model of a person, and a decision that needs the person still needs a human study.

Is Artificial Societies the same as the academic term “artificial societies”?

The academic term refers to agent-based social simulation, associated with Epstein and Axtell’s Growing Artificial Societies (opens in a new tab) (1996), in which simple rule-following agents produce social patterns. Artificial Societies is a company, founded in October 2024 and headquartered in London, that builds networks of AI personas grounded in real individuals, through its product Radiant, so organisations can test decisions before making them.

Sources