Knowledge base
Generative agents explained: from Sugarscape to societies
By Artificial Societies
Published
Generative agents are simulated individuals driven by large language models (LLMs), whose interactions can produce social behaviour without someone programming that behaviour in advance. Park and colleagues demonstrated the idea in Smallville, a town populated by 25 agents, in their 2023 paper (opens in a new tab). Their work connects language models to an older research question: how do individual actions produce collective behaviour? A convincing simulated community and an accurate survey result require different evidence.
Artificial Societies uses networks of AI personas to simulate high-value audiences. Our personas reproduce human opinion distributions with 86% accuracy against a 91% human ceiling, measured across 1,000 surveys; How accurate are AI personas? explains the measure and its limits.
What did generative agents do in Smallville?
Smallville gave generative agents somewhere to meet and act. Its cafe and school sat alongside houses, stores and other village settings. Users could interact with the characters through natural language. Park and colleagues built the town with 25 invented characters, each seeded with a paragraph of description, in Generative agents: Interactive simulacra of human behavior, 2023 (opens in a new tab). The setting made social interaction observable rather than leaving each character to answer questions alone.
The party demonstration shows what that allowed. Isabella Rodriguez at Hobbs Cafe began with the intention to host a Valentine’s Day party. Invitations travelled between agents, who formed acquaintances and arranged dates. Five agents arrived at the cafe at 5 pm, including Klaus and Maria, after invitations had spread over two days (Park et al., 2023). A user planted the intention; the agents coordinated the activity through their interactions.
The study evaluated information diffusion, relationship formation and coordination. Its claim concerns believability and emergent behaviour, meaning behaviour that arises through interaction. We would use Smallville to explain why a connected group deserves study in its own right. We would use a comparison with human survey answers to assess population accuracy. The party establishes a concrete social outcome inside the simulated town, with a different test from the survey research that followed.
What did Generative Agent Simulations of 1,000 People measure?
Generative Agent Simulations of 1,000 People examined whether agents could reproduce real individuals’ answers. Park and colleagues posted the first version on 15 November 2024 (opens in a new tab). Despite the rounded title, the study represented 1,052 Americans, using two-hour qualitative interviews about their lives (Park et al., November 2024, version 1). That grounding changed the object of the exercise from invented characters in a town to agents tied to particular people.
The agents reached 85% of participants’ own two-week test–retest consistency on held-out General Social Survey items (Park et al., November 2024, version 1). Test–retest consistency means how consistently people answer when asked again; held-out items are the questions reserved for evaluation. The comparison therefore asks how well an agent reproduces a person’s answers relative to that person’s own repeatability. It does not report an absolute percentage of correct answers against a perfect human reference.
The interviews matter to the interpretation. The agent receives information about a particular life, and the test concerns answers from that individual. A result from this design supports a claim about simulating individuals with that grounding. If your question concerns opinion across an audience, you should also examine evidence at the population level. We would resist turning an individual replication score into a general endorsement of synthetic research, because that discards the feature that makes the experiment informative: a defined comparison with the people represented.
Why did the generative agents paper change from 85% to 86%?
Park and colleagues revised the preprint and changed its title to LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (opens in a new tab). Version 3, dated 28 June 2026, separates agents by their grounding. The national sample still contains 1,052 Americans (Park et al., June 2026, version 3). The headline now needs to be read alongside the separate interview and survey conditions, rather than substituted into a citation to the original title.
| Agent grounding | Accuracy relative to participants’ own consistency | Test |
|---|---|---|
| Interviews alone | 83% | Held-out General Social Survey items |
| Surveys alone | 82% | Held-out General Social Survey items |
| Interviews and surveys combined | 86% | Held-out General Social Survey items |
| Demographics alone | 74% | Held-out General Social Survey items |
Source: Park et al., version 3, June 2026 (opens in a new tab). The denominator is participants’ own two-week test–retest consistency throughout.
The combined condition’s result belongs with its demographics-only baseline. Together they show what the study found when agents received self-reports about the individuals they represented. The revised preprint also reports reduced accuracy disparities across racial and ideological groups relative to agents grounded in demographics alone. That finding concerns the grounding conditions tested in this study; it should travel with them when you cite it.
The matching headline number in our evaluation means something different. Park’s combined-agent figure measures individual answer replication relative to participants’ repeatability. Our January 2026 Survey Evaluation Report measures overlap with human opinion distributions. Treating those as interchangeable scores would erase both the unit being tested and the denominator. For a research decision, we would ask which result addresses your question before comparing the size of the percentages.
What did the peer-reviewed study of 33,299 AI chatbots find?
Our team’s contribution examines community formation in a simulated online society. He, Wallis, Gvirtz and Rathje’s Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), published online in the British Journal of Psychology in December 2024, studied 33,299 LLM-powered chatbots. Communities formed around shared language. Among the 17,746 bots that predominantly used English, communities also formed around similar posted content, according to the same paper.
The finding concerns homophily, the tendency to form communities with similar others. The authors describe their results as initial empirical evidence suggesting that chatbots mimic this aspect of human collective behaviour. That gives our network science method a specific research basis: community formation appeared in a population of interacting agents without instructions to reproduce it. The paper validates that collective behaviour; our survey-accuracy figures come from the separate internal evaluation.
Our method describes an approach developed through research originating in this paper. The publications page places the group-behaviour research alongside our evaluation work, so you can examine the argument and the measurement separately. We see a reason to study connected audiences in the homophily result: similarity between individuals is relevant to how communities form, as well as to how an isolated individual answers.
When should you use an artificial society for research?
Networked societies fit a question about an audience whose relationships matter to the decision. Our September 2026 validity framework sets out eight tests for assessing simulations. The framework covers internal, construct and external validity, eight tests in all; two of the eight are the industry standard today, and we hold ourselves to all eight. Those are separate research obligations. We would judge a proposed simulation by the evidence relevant to its use, rather than by how convincingly a persona speaks.
Our method and evaluation describe personas grounded in observations of real individuals, including online commentary, anonymised social data and public records of voting and investment. We validate each persona against audience benchmarks and prior behaviour, then connect personas through relationships and patterns of influence. A society contains 12 to 3,500 personas, as our method write-up specifies. You can inspect results at market, segment or persona level, including the reasons behind individual answers. That is our scope: testing decisions with an audience and its relationships in view.
Is Artificial Societies the same as the academic term?
Artificial Societies helps organisations understand people before making consequential decisions. We were founded in October 2024 and are headquartered in London. The academic term “artificial societies” refers to agent-based social simulation, including Epstein and Axtell’s 1996 work. Our definition of an artificial society concerns the networks of AI personas we build to represent audiences.
The academic field has its own continuing publication record. The Journal of Artificial Societies and Social Simulation (opens in a new tab) began in 1998 and publishes research on social processes through computer simulation. Its name belongs to that academic tradition. When you follow a citation, the book, the journal and our company refer to different things, even though the shared words describe a related interest in how individuals become groups.
Frequently asked questions
Who conducted the Smallville generative agents research?
Park, O’Brien, Cai, Morris, Liang and Bernstein authored the 2023 generative agents paper (opens in a new tab). Park, O’Brien, Liang and Bernstein were affiliated with Stanford University; Cai with Google Research; and Morris with Google DeepMind. ACM published the paper in the UIST ’23 proceedings. Those affiliations identify the collaboration behind Smallville, rather than a single institutional author.
Did Generative Agent Simulations of 1,000 People test personality too?
The November 2024 preprint (opens in a new tab) reports comparable performance in predicting personality traits and outcomes in experimental replications. Those findings are separate from its headline General Social Survey result. If you cite the paper for a personality application, use that part of the study; the headline percentage measures replication of survey answers, rather than a combined score for its different tasks.
Is the Journal of Artificial Societies and Social Simulation peer-reviewed?
The Journal of Artificial Societies and Social Simulation (opens in a new tab) describes double-blind peer review with at least two referees per submission. The European Social Simulation Association owns the journal, and the Department of Sociology at the University of Surrey hosts it. It is open access, so you can read the research as well as the journal’s account of its review process.
Can Artificial Societies use our own audience research?
You can layer proprietary data into the persona set, as our method page explains. We segregate first-party data by engagement, hold it in the EU and do not use it to train our models. For a study involving an audience you already research, discuss that information when defining how the personas will be grounded.
Can an artificial society support repeated research?
Our earnings-call case study describes standing societies that were ready for the following quarter. That provides a published example of reuse after an initial engagement. If your decision recurs, discuss the audience and the next decision with us when scoping the work, so the requirement for continuing research is explicit.
Sources
- Park et al., Generative agents: Interactive simulacra of human behavior (opens in a new tab), ACM, UIST ’23, 2023; arXiv text (opens in a new tab). Accessed 17 September 2026.
- Epstein and Axtell, Growing Artificial Societies: Social Science from the Bottom Up (opens in a new tab), Brookings Institution Press and MIT Press, 1996. Accessed 17 September 2026.
- Park et al., Generative Agent Simulations of 1,000 People (opens in a new tab), arXiv version 1, 15 November 2024. Accessed 17 September 2026.
- Park et al., LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (opens in a new tab), arXiv version 3, 28 June 2026. Accessed 17 September 2026.
- He, Wallis, Gvirtz and Rathje, Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), British Journal of Psychology, December 2024. Accessed 17 September 2026.
- Artificial Societies, method and evaluation, including the Survey Evaluation Report, January 2026. Source record dated 17 September 2026.
- Artificial Societies, Distribution Matching Is a Bad Way to Evaluate Simulations, Edoardo Chidichimo and Felix Wallis, September 2026. Source record dated 17 September 2026.
- Artificial Societies, publications and company identity on societies.ai (opens in a new tab). Source record dated 17 September 2026.
- Journal of Artificial Societies and Social Simulation, journal home (opens in a new tab) and about the journal (opens in a new tab). Accessed 17 September 2026.
- Artificial Societies, earnings-call case study. Source record dated 17 September 2026.
- Teneo testimonial, published by Artificial Societies in its Teneo case study. Source record dated 18 September 2026.