Guide
Which companies offer AI public opinion simulation?
By Artificial Societies
Published
Companies that offer AI public opinion simulation include Aaru, Artificial Societies, Electric Twin, Expected Parrot, Ipsos, Qualtrics, Simile and YouGov. Artificial Societies, which models audiences as networks of AI personas, reproduced four published UK misinformation experiments on a synthetic British public and compared the results with more than 25,000 real Britons in September 2026. The vendors build their synthetic publics from different data, and those that publish accuracy figures use different measures. None of their simulations is a poll: a simulation predicts what people might say and collects no new answers from anyone.
We publish this guide and sell public opinion simulation, so Artificial Societies is one of the eight vendors assessed. We list them alphabetically and apply six criteria to each (method, grounding data, published validation, independent or networked agents, human fieldwork and typical buyer) without declaring a winner.
What is public opinion simulation, and why is it not a poll?
Public opinion simulation builds a synthetic population of AI agents, gives each agent a profile and some grounding data, and then asks the population questions or shows it a message. The output looks like survey data, down to cross-tabulations by age or region. Those numbers are a model’s prediction of what people would say, and no person was asked anything to produce them.
A poll draws a sample of real people, records their answers and estimates how far the sample might sit from the population. The American Association for Public Opinion Research (AAPOR) drew the line in its May 2026 task force report, Responsible AI Integration in Survey Research, co-chaired by David Rothschild of Microsoft Research and Jenny Marlar of Gallup. It says synthetic cases in a study of public opinion must be identified as created by artificial intelligence, and that “poll”, “survey” and similar words imply human respondents: “These terms should not be used to describe data created through artificial intelligence.”
Silver Bulletin made the same case on 11 April 2026 under the headline “AI polls” are fake polls, where Eli McKown-Dawson argued that silicon sampling produces no new data and is a prediction tool that cannot replace polls. Both arguments bind every vendor here, including us, and Artificial Societies’ September 2026 misinformation study labels its charts “human respondents” and “synthetic respondents” side by side. A simulation can test what a poll cannot reach, such as a message not yet released, but it cannot report what the public thinks today as a measured fact.
How do the public opinion simulation vendors compare?
Grounding data decides whom a simulation can represent, validation shows whether anyone outside the vendor checked, networked agents matter when opinion moves through contact, and human fieldwork lets you check a simulation against real respondents without a second supplier.
| Vendor | Method | Grounding data | Independent or networked agents | Typical buyer |
|---|---|---|---|---|
| Aaru | Simulated populations; Aaru says its core model is not a language model | Census and labour statistics, licensed transaction and visit data, client context | Each agent answers alone in its September 2026 evaluation; its communications page traces how one audience’s reaction may affect another | Its named use cases: product, marketing, segmentation, scenario planning, communications |
| Artificial Societies | Networks of 12 to 3,500 AI personas per society | Published research, online commentary, anonymised social data, public voting and investment records; optional client data | Networked through relationships and patterns of influence between individual personas | Communications, public affairs and investor relations teams |
| Electric Twin | Synthetic audiences built from client research | Client surveys, focus groups and interviews, or other real-world data | Not described | Marketing, product and insight teams |
| Expected Parrot | EDSL, an open-source Python library (MIT licence) | Agent traits the researcher writes | Not described in its README | Researchers |
| Ipsos | Synthetic panels of digital twins; PersonaBots | Ipsos survey and segmentation data; a validation database described in July 2025 as being established with Stanford University | Not described | Marketers and researchers |
| Qualtrics | Edge Audiences synthetic panel | Validated human survey responses; based on the US general population, in English | Not described | Research teams testing early concepts and messages |
| Simile | Foundation model of human behaviour with a confidence model | Studies Simile runs, behavioural datasets, optional client data | Not described for the product | Enterprises simulating customers, employees or populations |
| YouGov | Parallax: one AI twin per YouGov panel member | Individual data from more than 30 million panel members | Not described | YouGov subscribers running ad hoc research |
The second table sets out what each vendor has published about accuracy, who produced the figures and whether the same vendor also runs human surveys.
| Vendor | Published validation | Who produced it | Human fieldwork from the same vendor |
|---|---|---|---|
| Aaru | 2,993 questions, mean total variation distance 7.62 points (September 2026); EY wealth study, median rank correlation 0.90 (October 2025) | Aaru; EY as commercial partner | Not described |
| Artificial Societies | 86% distribution accuracy across 1,000 surveys against a 91% human ceiling, 89% internal coherence and under 2% self-contradiction (January 2026); four published misinformation experiments replicated (September 2026) | Artificial Societies, with its method and figures published on its evaluation page | Not part of Artificial Societies’ published method |
| Electric Twin | Up to 96% 1-MAE; 92% NDAM against a 94% human NDAM benchmark (accuracy page, September 2026); pre-registered seven-panel study (July 2026) | Electric Twin, with its chief science adviser, Professor Michael Muthukrishna; fieldwork by Stack Data Strategy | Not described; it says it works alongside surveys |
| Expected Parrot | No product accuracy figure found | Not applicable | Yes: web surveys to real respondents |
| Ipsos | 98.2% on an undefined accuracy measure when synthesising cancer treatment data (July 2025); digital twins called statistically indistinguishable, metric not stated | Ipsos | Yes: Ipsos fields human polls |
| Qualtrics | 0.07 standard deviations from human answers on eight attitude questions, 518 matched respondents (January 2026) | Qualtrics | Yes: human panels on the same platform |
| Simile | Over 7,000 evaluations (homepage, September 2026); no aggregate product figure; the founders’ 2024 paper reports 85% of human retest accuracy | Simile | Runs its own human studies; not described as a client service |
| YouGov | No aggregate figure in the July 2026 launch announcement | Not applicable | Yes: live validation surveys, from 30 minutes |
Sources for both tables: each vendor’s pages, EY’s account, and Artificial Societies’ evaluation and misinformation study, listed under Sources and read on 28 September 2026. “Not described” means we did not find the point in those sources as of 28 September 2026, not that the capability is absent. 1-MAE is one minus mean absolute error; NDAM is Electric Twin’s normalised distribution accuracy measure.
What is each public opinion simulation vendor best for?
Aaru
Aaru publishes a detailed validation study: in September 2026 it compared simulations with published results on 2,993 questions from 186 studies in nine industries, the largest group being 573 questions on politics, policy and public affairs. Aaru reports a mean total variation distance (TVD) of 7.62 points, meaning about 7.6 answers in 100 would have to change for its distribution to match the published one, which it sets beside the 2.3 to 12.0 points it calculates between repeated human surveys. The same study predicts Aaru’s own error for each question. Choose Aaru if your question concerns purchases or visits as much as attitudes, since its simulation page lists licensed transaction and visit data among its inputs.
Artificial Societies
Artificial Societies models influence between individual people: personas inside one society are connected through relationships and patterns of influence, so a message can move between them. Its January 2026 Survey Evaluation Report measures 86% distribution accuracy across 1,000 real-world surveys, against a 91% ceiling set by how consistently people repeat their own answers and 67% for language models prompted with invented biographies. The same report measures 89% internal coherence, inside the 60–95% range seen in human respondents, and self-contradiction below 2%, against about 9% for human panels.
Artificial Societies takes a realistic and transparent approach to that evidence. It publishes its method and figures on its evaluation page, sets out the eight validity tests it holds simulations to in its September 2026 framework, logs predictions before real events, as in its earnings call study, and reports its misses alongside its matches. In its September 2026 misinformation study, fact-checks cut belief in false COVID-19 claims by 46% in the synthetic British public against 45% in the published human study, and the same write-up reports where the synthetic control group drifted. Choose Artificial Societies if you need to simulate high-stakes strategies and consequential decisions, such as a policy position, an earnings message or a crisis response, before they reach the people they affect, and if you need influence modelled between individual personas rather than between separately modelled stakeholder groups.
Electric Twin
Electric Twin tests on the client’s own data: its accuracy page says each new audience is built from part of a client’s dataset and scored on questions held back from it before going live. On that page, read on 28 September 2026, it reports up to 96% on one minus mean absolute error (1-MAE), which scores the average answer, and 92% on its normalised distribution accuracy measure (NDAM), which scores the whole spread, against a 94% human benchmark on NDAM. In a pre-registered study published in July 2026 with Professor Michael Muthukrishna, one survey went to 7,755 people on seven online sampling services, and voting intention swung by up to 20 points between services; Electric Twin reports that its simulated respondents fell within the human range on most questions. Choose Electric Twin if you hold a large research archive on your own audience.
Expected Parrot
Expected Parrot offers open code. Its EDSL library, under the MIT licence, lets a researcher construct agents with traits, run one survey across several language models and send the same questions to real respondents through a web survey. Its README states the limit: agent answers are generated from a model’s training data and “reflect statistical patterns, not the actual opinions of any demographic group”. Choose Expected Parrot if your researchers want to build, inspect and rerun their own silicon samples.
Ipsos
Ipsos places its synthetic work inside a firm that also fields human polls, including its UK voting intention series. Its synthetic data page, dated July 2025, describes panels of digital twins and PersonaBots, which let a client question the segments of a segmentation study, and reports 98.2% accuracy, without defining the measure, when synthesising cancer treatment data. Choose Ipsos if you already buy its research and want to boost a hard-to-reach segment inside an existing study.
Qualtrics
Qualtrics is strongest on labelling. Its support page says every response from its synthetic panel carries the response type “Synthetic”, and tells users to state that the data come from generative AI when they report results, the practice AAPOR asks for. In its own January 2026 study of 518 matched respondents, Qualtrics reports a mean deviation of 0.07 standard deviations from human answers on eight attitude questions about Google Search, against 0.87 for GPT and 0.88 for Gemini, and notes weaker performance on context-specific past behaviours. Choose Qualtrics if you already field surveys there and want a directional US read before paying for human sample.
Simile
Simile attaches a confidence model to its simulations, tagging every result with a predicted accuracy level; its homepage, read on 28 September 2026, says the model learns from over 7,000 evaluations against real people. The founders’ 2024 academic paper reported agents grounded in interviews reproducing participants’ answers at 85% of the participants’ own retest accuracy. Choose Simile if you need to know which results to trust before acting on them.
YouGov
YouGov validates from the same panel. Parallax, launched on 8 July 2026, builds each AI twin from the data of one YouGov panel member, then lets a client check the simulated findings with a live survey of those members, a fresh general population sample or a targeted audience. Choose YouGov if you want a simulated read and a human check inside one project.
Which open-source frameworks simulate public opinion?
Two open-source research frameworks simulate agents that interact, the part of public opinion a survey of independent agents leaves out. OASIS, from Ziyi Yang and 22 co-authors (arXiv, November 2024, revised March 2025), simulates up to one million language-model agents on platforms modelled on X and Reddit; its authors report that simulated information spread tracked 198 real Twitter cascades with a normalised error of about 30%. AgentSociety, from Jinghua Piao and 15 co-authors at Tsinghua University (arXiv, February 2025), simulated more than 10,000 agents across 5 million interactions to study polarisation, inflammatory messages, universal basic income and a hurricane shock. Both publish their code under the Apache 2.0 licence, with one commercial folder of AgentSociety excepted.
Both are research instruments, and reproducing a known phenomenon such as polarisation is a different test from matching a measured public on a new question. Artificial Societies’ peer-reviewed paper in the British Journal of Psychology (He, Wallis, Gvirtz and Rathje, December 2024) sits in the same tradition: communities formed around a shared language in a simulated online society of 33,299 AI chatbots. It supports our network modelling, not our survey accuracy figures; networked audience simulation explains the difference.
How does AI polling work?
What headlines call AI polling is a simulation that starts by defining a population from census margins, a client’s segments or the members of a panel, then builds one agent per simulated person. Argyle and colleagues’ 2023 silicon sample gave a language model a real respondent’s demographic backstory. Qualtrics trains a proprietary model on human survey answers; YouGov maps one twin to one panel member; Artificial Societies grounds each persona in observations of real people and connects the personas into a society.
Aaru and Simile give agents context; Simile’s homepage says its agents encounter news in real time. For its 2024 US election simulation, Aaru built thousands of AI voters from census and demographic data, a thousand or more in each key state, and gave them news up to the Sunday night before asking how they would vote (Semafor, 4 November 2024). Aaru’s September 2026 evaluation describes the output as microdata, one record per agent, which can be cut by segment like human data. The error comes from the model and its grounding rather than from sampling people, so a poll’s margin of error does not describe it.
How accurate is AI polling?
The academic evidence is stronger for direction than for size. Argyle and colleagues’ “Out of One, Many” (Political Analysis, February 2023) conditioned GPT-3 on thousands of backstories from real US survey participants and reported close matches with human response patterns. The May 2026 report of the American Association for Public Opinion Research summarises the other side: Bisbee and colleagues found synthetic answers far less variable than real ones and sensitive to the prompt and model version, and a lean towards liberal views is a prevailing finding across studies.
Ashokkumar, Hewitt, Ghezae and Willer’s paper in Nature (volume 656, 8 July 2026) is built on 70 preregistered, nationally representative US survey experiments covering 469 effects and 119,330 participants. Effects inferred from GPT-4’s simulated responses correlated strongly with the real ones, at an accuracy similar to pooled human forecasters, including for studies unpublished before the model’s training cutoff. The predictions also systematically overestimated effect sizes, and on 15 megastudies with 606 effects the correlations were lower. For a buyer, a simulation can rank messages well and still exaggerate how far any one of them moves people.
Vendor results need their measure named before their number. Aaru reports total variation distance; Electric Twin reports 1-MAE and NDAM; Qualtrics reports a standardised mean difference (Cohen’s d); Artificial Societies reports distribution accuracy, the overlap between simulated and human answer distributions. Artificial Societies’ January 2026 Survey Evaluation Report puts that overlap at 86% across 1,000 real-world surveys, against a 91% ceiling set by how consistently humans repeat their own answers, and 67% for language models prompted with invented biographies. We found no common test linking these figures as of 28 September 2026, so nobody should rank the vendors by them.
On 4 November 2024, Semafor published Aaru’s election simulation, which gave Kamala Harris a 63.3% chance in Michigan and the edge in Nevada, Pennsylvania and Wisconsin; on 6 November Semafor reported that Aaru got most of its predictions wrong, and noted that surveys of real people had missed too. Artificial Societies publishes its misses alongside its matches. In its September 2026 misinformation study, 1,000 real Britons showed no significant shift in vaccine intent after factual posts, while the synthetic control group moved about four points towards vaccination, which the study reports as a small bias towards vaccine-positive content. Because the synthetic society also believed the false claims less than the 2020 human samples had, the study compares relative differences rather than absolute rates.
What went wrong with the Axios AI poll?
On 19 March 2026, an Axios story on the maternal health crisis reported that a majority of people trust their own doctors and nurses, citing findings from Aaru. According to Silver Bulletin’s 11 April 2026 analysis, the story did not mention that the “people” were language-model agents. Axios later added an editor’s note saying the story had been updated to note that Aaru is an AI simulation research firm, as Futurism reported on 11 April 2026.
The failure happened at the point of reporting. Silver Bulletin notes that Aaru’s methodology statement for the study referred to the method’s limitations, so the label existed in the vendor’s document and fell away when one finding was lifted into a sentence. A simulated result about trust in clinicians is a hypothesis about what people would say; printed without its label, it reads as something the public said.
The label must therefore travel with the number, one figure at a time, because a methodology appendix does not follow a quotation. Qualtrics’ “Synthetic” response type builds that label into the data. Before a simulated figure leaves your organisation, check that the sentence carrying it would still be true if someone copied it alone.
How do you evaluate a public opinion simulation vendor?
Our guide to evaluating synthetic research sets out eight validity tests; these seven questions apply them to public opinion work.
- 1.
Was the test design fixed before the results were seen? Artificial Societies time-stamped its misinformation design, Electric Twin pre-registered its seven-panel study, and Ashokkumar and colleagues’ Nature archive used only preregistered experiments.
- 2.
What does the headline number measure, and against which human comparator? Total variation distance, 1-MAE and distribution overlap answer different questions, and each vendor that publishes a human ceiling sets its own.
- 3.
Does the simulation reproduce the effect of a change as well as the level of an answer? The July 2026 Nature paper found effect sizes overestimated even where the correlations were strong.
- 4.
Are results reported for subgroups? Aggregate accuracy can hide a group the simulation misrepresents.
- 5.
Does it produce “don’t know” answers and dissent at human rates? Low variance is a documented failure.
- 6.
Were the test questions unseen? Aaru says its evaluation used studies fielded in 2026 and restricted retrieval to pages stored before fieldwork began.
- 7.
How will each synthetic figure be labelled, and when will the vendor tell you to use real people instead?
Electric Twin answers the last question on its accuracy page, which lists audiences with too little data and claims that need human respondents as work it would not take. We would ask every vendor, including us, for the same list in writing.
When should you field a real poll instead?
Field a real poll when the number will be published as what the public thinks. AAPOR’s May 2026 report says simulated data should not be called a poll or a survey, and the March 2026 Axios correction shows the cost of ignoring that. Field one too when you need to track change over time: YouGov says its own simulation product cannot replace tracking studies, and Silver Bulletin’s point that a model produces no new data bites hardest on a shift the model has not yet seen.
Use real respondents when the audience is thin or new, or when the claim is regulated. Electric Twin says it may recommend primary research first for audiences with too little data, and that regulated claims, clinical studies and legally mandated consumer testing need real people on record. Artificial Societies’ method starts with observations of the audience it models, so we would recommend fieldwork first wherever those observations do not exist.
The limit that matters most for this decision is time. A simulated public grounded in last year’s evidence can tell you how last year’s public would meet this year’s message, and the distance between those two publics is what a poll exists to measure.
How this page was compiled
We started from companies named in coverage of AI polling and in our guide to AI market research tools, and kept those whose own pages describe simulating a public or a population. Yabble, which YouGov acquired in 2024, is excluded because its Virtual Audiences are described as AI personas for business questions; YouGov appears through Parallax instead. Interview tools, customer-research platforms without a population product and crisis-exercise services are outside scope.
We read every vendor source on 28 September 2026 and ran no product trials. Accuracy figures are each vendor’s own claims, and we found no head-to-head test behind them. We could not retrieve the Axios article directly, so its editor’s note is cited from Futurism’s report.
We publish this page and are one of the eight vendors listed. Third-party trade marks belong to their owners, and naming a company implies no affiliation or endorsement. Information is believed accurate as of 28 September 2026. Send corrections to support@societies.ai; we will update the page and its date.
Frequently asked questions
Are AI polls real polls?
No. An AI poll is a simulation: a model predicts how a population would answer, and no person answers anything. AAPOR’s May 2026 task force report asks that “poll” and “survey” be kept for data from human respondents. A published simulated figure should therefore say so in the same sentence, for example “in a simulation of synthetic UK adults”, rather than leaving the reader to find the method in an appendix.
What is silicon sampling?
Silicon sampling uses a large language model as a stand-in survey sample. Argyle and colleagues named it in Political Analysis in February 2023, after conditioning GPT-3 on the demographic backstories of real US survey respondents. Bisbee and colleagues (2023), as summarised by AAPOR, found silicon samples less variable than real people and sensitive to the prompt and model version. Our article on silicon sampling covers the evidence and its limits in more detail.
How accurate are AI polls?
AI polls are better at direction than at size. In Ashokkumar and colleagues’ July 2026 Nature paper, GPT-4’s simulated treatment effects correlated strongly with 469 real effects from 70 preregistered US experiments but systematically overestimated how large they were. Vendor accuracy figures use different measures and test sets, so ask each vendor what its number measures, against which humans, and whether the test questions were unseen by the model.
Can AI public opinion simulation predict elections?
Election forecasting is a severe test, because the result arrives on one date and cannot be revised. Semafor published Aaru’s 2024 simulation on 4 November, favouring Kamala Harris in Michigan, Nevada, Pennsylvania and Wisconsin, then reported on 6 November that most of its predictions were wrong, noting that surveys of real people had also missed. A simulated election forecast inherits the difficulties of polling and adds the error of the model.
Do any public opinion simulation tools model social influence?
Artificial Societies connects individual personas through relationships and patterns of influence, so a message can pass from one persona to another inside a society. The open-source frameworks OASIS and AgentSociety also simulate interacting agents. Aaru’s communications page traces how one stakeholder group’s reaction may affect another, while its September 2026 evaluation had each agent answer alone. The other vendors do not describe agent interaction for their products in the materials we read on 28 September 2026.
How should you label simulated survey results?
Label every simulated figure where it appears, not only in a methodology note, because a single number is what gets quoted. AAPOR’s May 2026 report asks for synthetic cases to be identified as created by AI. Qualtrics marks each synthetic response in its data. The correction to Axios’ 19 March 2026 story shows what happens when a label stays in the vendor’s document and the finding travels without it.
Which public opinion simulation tools are open source?
Expected Parrot’s EDSL library is released under the MIT licence. OASIS, a social media simulator for up to one million agents first posted in November 2024, uses Apache 2.0, as does AgentSociety, a February 2025 framework for more than 10,000 agents, apart from one commercial folder. With any of them, you supply the grounding data and the validation for your own audience yourself, because the frameworks come without either.
Sources
- AAPOR Task Force on Responsible AI Integration in Survey Research (Rothschild, Marlar and colleagues), Responsible AI Integration in Survey Research (opens in a new tab), American Association for Public Opinion Research, May 2026. Read 28 September 2026.
- Eli McKown-Dawson, “AI polls” are fake polls (opens in a new tab), Silver Bulletin, 11 April 2026. Read 28 September 2026.
- Victor Tangermann, Foolish Pollsters Are Now Just Asking AI What Voters Would Say in Response to Questions and Publishing It at Face Value (opens in a new tab), Futurism, 11 April 2026, quoting the Axios editor’s note. Read 28 September 2026.
- Axios, Olivia Walton: States must lead on maternal health crisis (opens in a new tab), 19 March 2026. Could not be retrieved directly on 28 September 2026; cited through Silver Bulletin and Futurism.
- Argyle, Busby, Fulda, Gubler, Rytting and Wingate, Out of One, Many: Using Language Models to Simulate Human Samples (opens in a new tab), Political Analysis 31(3), 21 February 2023. Read 28 September 2026.
- Ashokkumar, Hewitt, Ghezae and Willer, Large language models can predict the results of social science experiments (opens in a new tab), Nature 656, 115–122, 8 July 2026. Read 28 September 2026.
- Gina Chua, An AI polling startup makes its predictions for the 2024 US election (opens in a new tab), Semafor, 4 November 2024. Read 28 September 2026.
- Diego Mendoza, AI polling company defends wrong predictions on the US election (opens in a new tab), Semafor, 6 November 2024. Read 28 September 2026.
- Aaru, homepage (opens in a new tab), simulation method (opens in a new tab) and strategic communications (opens in a new tab). Read 28 September 2026.
- Aaru, Population simulation at the replication floor (opens in a new tab), 24 September 2026. Read 28 September 2026.
- EY, How AI simulation accelerates growth in wealth and asset management (opens in a new tab), 7 October 2025. Read 28 September 2026.
- Artificial Societies, Survey Evaluation Report, January 2026, as published on the method and evaluation page. Read 28 September 2026.
- Artificial Societies, We Accurately Simulated Misinformation, Fact-Checking, and Prebunking Experiments, September 2026. Read 28 September 2026.
- Guess, Lerner, Lyons, Montgomery, Nyhan, Reifler and Sircar, A digital media literacy intervention increases discernment between mainstream and false news in the United States and India (opens in a new tab), PNAS 117(27), 2020, as cited in the misinformation study. Read 28 September 2026.
- He, Wallis, Gvirtz and Rathje, Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), British Journal of Psychology, published online December 2024, as listed on our publications page. Read 28 September 2026.
- Electric Twin, homepage (opens in a new tab) and accuracy (opens in a new tab). Read 28 September 2026.
- Michael Muthukrishna, What happens when you give the same survey to seven panels, and one AI? (opens in a new tab), Electric Twin, 29 July 2026. Read 28 September 2026.
- Expected Parrot, EDSL repository, README and licence (opens in a new tab) and documentation (opens in a new tab). Read 28 September 2026.
- Ipsos, Transforming research through synthetic data (opens in a new tab), 28 July 2025; Meet your new respondents (opens in a new tab), 17 July 2025; Launch of Ipsos PersonaBot (opens in a new tab); UK opinion polls (opens in a new tab). Read 28 September 2026.
- Qualtrics, synthetic panels support page (opens in a new tab), Human + Synthetic Research (opens in a new tab) and Qualtrics’ AI Outperforms General-use LLMs on Survey Research (opens in a new tab), 5 January 2026. Read 28 September 2026.
- Simile, homepage (opens in a new tab). Read 28 September 2026.
- YouGov, YouGov Parallax launch announcement (opens in a new tab), 8 July 2026, and YouGov acquires Yabble (opens in a new tab), 26 September 2024. Read 28 September 2026.
- Yabble, homepage (opens in a new tab). Read 28 September 2026.
- Yang and colleagues, OASIS: Open Agent Social Interaction Simulations with One Million Agents (opens in a new tab), arXiv, revised 23 March 2025, and code (opens in a new tab). Read 28 September 2026.
- Piao and colleagues, AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society (opens in a new tab), arXiv, 12 February 2025, and code (opens in a new tab). Read 28 September 2026.