Skip to content

Knowledge base

Are AI personas biased? Representational fairness in synthetic research

By Artificial Societies
Published

AI personas are biased by default: a language model asked to answer as a person leans towards some groups’ views, narrows the range of opinion inside each group and fills invented biographies with stereotypes. Four researchers at Artificial Societies, which builds networks of AI personas grounded in real individuals, co-authored a September 2026 arXiv paper (Kutzner et al.) that treats this as a validity failure and requires accuracy to be reported for subgroups chosen in advance. It makes fairness checkable as three quantities: prediction error across groups, documentation of the process and variety preserved inside each group.

What does bias mean for an AI persona?

Bias in an AI persona is a systematic difference between its answers and those of the people it stands for, larger for some groups than for others. A synthetic persona inherits it from its language model, and published research documents at least three forms.

Form of biasWhat goes wrongPublished evidence
Whose opinionsThe model’s default answers match some groups better than othersSanturkar et al., ICML 2023
FlatteningA group’s range of views collapses towards one typical positionWang, Morgenstern and Dickerson, 2025; Bisbee et al., 2024
MisportrayalAnswers sound like outsiders’ view of a group, not its members’ ownWang, Morgenstern and Dickerson, 2025

Santurkar and colleagues found language model opinions as far from those of US demographic groups as Democrats are from Republicans on climate change. People aged 65 and over and widowed people were among the groups whose opinions the models reflected poorly. Kutzner and colleagues, citing this and later work, add that a model prompted with no individual information leans towards the views of younger, better-educated and more liberal respondents.

Wang, Morgenstern and Dickerson tested four large language models in human studies with 3,200 participants across 16 demographic identities (Nature Machine Intelligence, February 2025). They trace misportrayal and flattening to training. Text rarely records who wrote it, so what a model learns about a group can come from people outside it; and training rewards the most likely output, which erases variety within a subgroup. Bisbee and colleagues found the same shape in ChatGPT’s ratings of 11 sociopolitical groups (Political Analysis, 2024). Averages came close to the American National Election Studies, but varied less than the real surveys and shifted with small wording changes and over three months.

Why does an accurate average hide biased subgroups?

An accuracy figure computed over a full sample is a weighted average, dominated by its largest groups. The September 2026 paper by Kutzner and colleagues shows the consequence: a synthetic sample can reproduce a population mean while being wrong for every subgroup inside it, as long as the errors offset.

Kutzner and colleagues name three mechanisms behind subgroup failure. A model can learn population-level statistics from published polling in its training data without learning the group-level structure beneath them. Its performance tracks training-data coverage, which is thin for older people, people on low incomes, speakers of less-represented languages and people with limited digital participation. And a model trained to be helpful struggles to act out incapacity. Non-uptake often comes from not knowing the option exists, not understanding it or not managing the process, so a persona set that cannot represent those states over-predicts uptake in the groups where they are most common.

One of the paper’s four testable predictions follows: aggregate fidelity will overstate worst-subgroup fidelity, and the gap will grow as training-data coverage falls. The group your decision affects most may be the one a headline figure describes worst.

How do you measure bias in synthetic respondents?

You measure bias by repeating each accuracy test inside subgroups fixed before anyone looks at the synthetic data. Kutzner and colleagues (arXiv, September 2026) ask for groups defined by the decision, not by the data: for an energy tariff, income band, tenure, dwelling type and household composition. Where samples allow, they ask for cells crossing the two or three most decision-relevant attributes, and for the worst-group result beside the average. Baselines are declared too: demographic base rates and a model given no persona at all, not chance.

A behavioural criterion (C) asks whether personas predict what the represented people do. Four diagnostic levels then locate failures. Location (L0) asks whether averages match per group, and dispersion (L1) whether the spread of answers is human-like. Response process (L2) asks whether personas react to wording, order and format as people do, and structure (L3) whether the relationships between answers match human data. A cross-cutting experimental check (E) asks whether a manipulation moves the synthetic sample as it moves people.

Two of the three forms map onto particular levels. Flattening is an L1 failure, and the paper calls dispersion the level where models fail most consistently; the test is the ratio of synthetic to human variance within each subgroup. Associations can also be reproduced for the wrong reasons, from stereotypes in outsiders’ descriptions of a group. At L3 the paper therefore treats a structure cleaner than the human data as a failure, not a success. The human benchmark bounds every claim: the authors note that a subgroup with 30 human respondents cannot support a strong claim either way.

What does the representational fairness paper propose?

“Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research” was posted to arXiv on 23 September 2026. Its authors are Florian Kutzner, Celina Kacperski and Laura de Molière of decision-context, and Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis and James K. He of Artificial Societies. Artificial Societies funded the decision-context authors’ working time, and the paper states that final authority over content rested with them.

Its two requirements: every validity claim states its level of correspondence with human data, and every claim is reported for subgroups. It turns three dimensions from energy-justice research into quantities you can read from a results table:

Justice dimensionQuestion it asksMeasured as
DistributionalWho bears the cost of prediction error?Prediction error across pre-specified subgroups
ProceduralCan others reproduce or contest the result?Documentation of model and version, prompt text, sampling temperature, draws per persona, persona-set construction and the human benchmark
RecognitionWhose situation is acknowledged rather than overlooked?Variance preserved within each subgroup

The authors add that these indicators do not establish that a deployment is fair. They also define within-persona counterfactual experiments, which copy one persona into several conditions and read the contrast for that profile; human research cannot do this, because a person carries one condition into the next. Because a prompt change can alter the model’s picture of the respondent, and identical copies vary at random, the paper limits such claims to rough effect sizes.

The worked case is an electricity supplier testing a time-of-use tariff meant to move electric vehicle charging into hours of high renewable generation, with charging start times from meter data as the behaviour. If a synthetic sample reproduced the aggregate shift but overstated it for low-income and renting households, the supplier could adopt a tariff whose benefits go to households with the flexibility to respond. The paper recommends subgroups that track constraints, such as access to a private charging point, over plain demographics. The case is an illustration; the paper reports no new data.

Does grounding personas in real individuals address each kind of bias?

Artificial Societies grounds each persona in a real individual rather than an invented biography such as “you are a 34-year-old teacher from Ohio”. Our method and evaluation page calls that biography prompt the standard approach, one that produces generic, stereotyped answers. The page lists the observations behind each persona: published research, online commentary, reviews, anonymised social data and public records of voting and investment.

That design targets misportrayal directly, since it replaces the invented biography. It is also meant to draw variety within a group from observed individuals rather than from the model’s most likely answer, though no published within-group variance figure tests that yet. Artificial Societies’ Survey Evaluation Report (January 2026) gives the aggregate evidence: 86% distribution accuracy across 1,000 surveys, against 67% for biography-prompted models and a 91% human self-replication ceiling. The same report puts Artificial Societies’ internal coherence at 89% (Cronbach’s alpha, a measure of whether related questions hang together), inside the 60–95% human range. That meets the rule in Kutzner and colleagues’ paper that structure be no cleaner than human data, in aggregate only.

The paper’s equivalence logic cuts against Artificial Societies on self-contradiction. Artificial Societies’ hallucination rate of under 2% sits below the roughly 9% measured for human panels (January 2026 report), and the paper’s response-process test asks for human-level inconsistency, not less. For whose-opinions bias, grounding replaces the model’s default voice with observed people, but only the people the sources reach. Three of the five source types (online commentary, reviews and anonymised social data) come from online platforms. Kutzner and colleagues name people with limited digital participation among the groups training data miss, and more online observation does not reach them. In our reading of the paper’s levels, this is what Artificial Societies’ evaluation page reported on 28 September 2026:

Paper’s levelClosest measure Artificial Societies publishesNot reported on the evaluation page
C: behaviourNo figure; the method page states personas are validated against prior behaviourA behavioural prediction result per subgroup
L0–L1: location and dispersion86% distribution accuracy, aggregate across 1,000 surveys (January 2026 report)Per-subgroup accuracy, worst-group result, within-group variance ratio
L2: response processUnder 2% hallucination rate against about 9% for human panels (January 2026 report)Order, scale and framing effects against human bands
L3: structure89% internal coherence, inside the 60–95% human range (January 2026 report)Evidence that questions measure the same thing in every subgroup
E: experimentalNo figure; the September 2026 validity framework commits to matching effect sizes on experiments the model has not seenPublished effect-size results

What should you ask a synthetic research vendor about bias?

Kutzner and colleagues close their September 2026 paper with a 12-item reporting checklist, which translates into questions you can put to any vendor, including us. Ask for written answers with numbers:

  1. 1.

    Which subgroups did you define before validation, and why those for my decision?

  2. 2.

    What is the result for the worst subgroup, next to the average?

  3. 3.

    What is the ratio of synthetic to human variance within each subgroup?

  4. 4.

    How many human respondents sit behind each subgroup benchmark, and from which source and date?

  5. 5.

    Does the claim concern what people say or what they do, and is a survey-only claim labelled as one?

  6. 6.

    How were the personas built, from which data, and on what legal and ethical basis?

  7. 7.

    Which model and version produced the results, and when was it accessed?

  8. 8.

    Who ran the validation, and what is their financial relationship to the vendor?

On questions 2 and 3, Artificial Societies’ evaluation page reported no subgroup or variance-ratio figures as of 28 September 2026. The eight validity tests from Artificial Societies’ September 2026 framework are in how to evaluate synthetic research.

What are the limits of measuring bias this way?

In our view the hardest limit is circular. Subgroup validation needs human benchmark data for each subgroup, and synthetic samples are attractive precisely for groups where human data are scarce or do not exist. For those settings Kutzner and colleagues (September 2026) point to algorithmic humility: a system should recognise when it works at the edge of its evidence and defer to human review.

The authors call a different limit perhaps the most crucial: the field still lacks a full set of practical tests. They also note that a model may reproduce a classic result because the paper describing it sits in its training data. And providers update models without notice, so a validity result belongs to one model at one time. The paper ends by asking what synthetic respondents “can predict, and for whom”, and for a group no human study has measured, a fairness claim is one that nobody can yet check.

Frequently asked questions

Can better prompting remove AI persona bias?

Better prompting reduces some bias but does not remove it. Santurkar and colleagues (ICML 2023) found that misalignment with US demographic groups persisted after models were explicitly steered towards those groups. Wang, Morgenstern and Dickerson (Nature Machine Intelligence, 2025) tested inference-time techniques, such as identity-coded names in place of explicit labels, and report that these reduce, but do not remove, misportrayal and flattening.

Is a biased synthetic sample always useless?

A biased sample can still serve a narrower use. Kutzner and colleagues (September 2026) set requirements by use: piloting a questionnaire needs location, dispersion and response process inside human bands for each relevant subgroup. Standing in for a human sample adds evidence that answers relate the same way in every compared group, experimental correspondence and a human anchor. A subgroup without that evidence cannot inherit conclusions drawn from the whole sample.

Can a small human sample correct synthetic research bias?

Kutzner and colleagues (September 2026) recommend it where human data are scarce: researchers get more from correcting synthetic subgroup estimates against a small human sample than from fine-tuning the model. The validity guarantee then covers the corrected subgroup estimates only. The individual generated responses remain uncorrected, which matters wherever they feed follow-up questions.

Does Artificial Societies report results by segment?

Artificial Societies’ method page says study results can be analysed at market, segment or persona level, and lists reliable segments and crosstabs as a strength of the approach. As of 28 September 2026 the page publishes no accuracy figure for a single segment. If your decision turns on one segment, ask which human benchmark stands behind it, and treat a segment without one as unvalidated for that decision.

Why does the representational fairness paper focus on behaviour?

Kutzner and colleagues focus on behaviour because what people say and what they do differ, and the gap widens when acting is costly. The paper’s figure puts the human coupling at roughly half of people with an intention acting on it, citing Sheeran and Webb (2016). A persona set can match survey answers and still over-predict uptake for low-income households, which is a bias in the outcome a decision maker cares about.

Does grounding a persona in a real individual raise privacy issues?

Kutzner and colleagues (September 2026) note that a persona built to represent a named living person, from material about that person, can raise data-protection and profiling issues under the General Data Protection Regulation and the EU AI Act. They restrict their own definition of a persona to a hypothetical archetype. Artificial Societies’ method page describes anonymised social data, and first-party data segregated per engagement and held in the EU. Ask any vendor for the legal basis of its persona data.

Does the fairness framework favour one way of building personas?

The framework in Kutzner and colleagues’ September 2026 paper is agnostic, by design, to how personas are built, whether from prompts, demographic profiles or interview data. The paper leaves open whether more individual detail, such as interview transcripts, improves behavioural prediction, and presents its levels as the way to test that. It also states that the framework is vendor-agnostic, and it requires every study to disclose who ran the validation.

Sources