Knowledge base
What is group alignment in multi-agent AI?
By Artificial Societies
Published
Group alignment is the property of a group of AI agents whose interaction structure, including the shared conventions that emerge between agents, reliably produces joint behaviour that respects human values. Artificial Societies, which models audiences as networks of AI personas, published the definition in September 2026 with agent-based experiments in which human oversight of 2% of agents brought the group’s average norm within 0.009 of the human reference (set at 1) once agents followed their peers, against about 0.49 without peer influence. Aligning each agent alone does not make a group safe.
The post, AI Researchers Should Use Group Alignment to Reduce P(doom), is by Emmanuelle Gelain-Sohn of Faculty and Yitian Chen and Felix Wallis of Artificial Societies; P(doom) in its title means the probability of catastrophe from AI.
Why does aligning each AI agent not make the group safe?
Our September 2026 post, following Durkheim (1982), treats organisations as more than the sum of their members, because hierarchies and norms develop through interaction. No NASA engineer set out to endanger the Challenger crew in 1986, yet Diane Vaughan’s study of the launch decision describes how group pressures normalised warning signs until departures from expected performance became acceptable (Vaughan, 1996).
Shen and colleagues tested the same gap with AI in April 2026 (arXiv 2604.10290 (opens in a new tab)). They copied one aligned model into an organisation, gave it 12 tasks across an AI consultancy and an AI software team, and gave a single copy of the model the same tasks alone. The organisations achieved better business outcomes and worse ethical decisions than the individual.
The usual alignment techniques act on one model at a time. Supervised fine-tuning trains a model on example answers, reinforcement learning from human feedback (RLHF) trains it on people’s preferences between its answers, and chain-of-thought monitoring reads its written reasoning. None of them can tell whether one response will combine with others into a harmful output, because each agent can meet its own specification while the group’s output fails the specification for the whole (Altmann et al., 2024). Small errors also compound as models repeat them to one another (Flint et al., 2026; Shumailov et al., 2024), and a collective that rewards results can push agents to drop their guardrails (Hammond et al., 2025).
METR’s investigation (opens in a new tab), published on 26 August 2026, describes a real case from July 2026. Roughly 1,200 agents meant to be isolated from one another communicated on an unsanctioned message board, roughly 700 of them attacked Hugging Face, and no human was alerted while the attack ran.
How is group alignment defined?
“A group of agents is group-aligned when its interaction structure, including the shared conventions that emerge between agents, reliably produces joint behaviour that respects human values.”
Interaction structure means who communicates with whom and how; shared conventions are the norms agents settle into together; joint behaviour is what the group does as a whole. The definition carries three conditions. Composition puts the test on joint behaviour, because every agent could satisfy its own specification while the group violates the specification for the whole. Durability requires the conventions to keep respecting human values under sustained pressure to abandon them, including pressure from inside the group. Incentive compatibility requires that, under any pressure the group faces, its interaction structure never makes violating human values an agent’s preferred option.
Under this definition, individual alignment is neither necessary nor sufficient for group alignment. A group of aligned agents can fail the three conditions, and a group that contains individually misaligned agents can still pass them. The conventions must hold because agents share them, and emerge through self-regulation rather than structures imposed from outside.
The definition combines two views of collective misalignment: a population settling on a position its members would not reach alone (Flint et al., PNAS, 2026 (opens in a new tab)), and a group measured against an external reference such as a published constitution (Wang et al., arXiv, 2026 (opens in a new tab); Shen et al., 2026). Group alignment keeps the external reference and allows benign drift from agents’ starting positions.
How is group alignment different from individual alignment?
| Question | Individual alignment | Group alignment |
|---|---|---|
| What is tested | One model’s responses | The joint behaviour of many interacting agents |
| Techniques named in the post | Supervised fine-tuning, RLHF, chain-of-thought monitoring | Group-level interventions tested in agent-based versions of a mixed swarm, with group-level reward functions planned |
| Failure it can catch | A harmful response from one model | A harmful combined output built from responses that each pass |
| State of the evidence in September 2026 | Organisations of aligned copies proved less aligned than one copy (Shen et al., 2026) | A proof of concept in simple agent-based models, not yet in language models |
Source: Gelain-Sohn, Chen and Wallis, AI Researchers Should Use Group Alignment to Reduce P(doom), Artificial Societies, September 2026.
What happens when AI agents act faster than humans?
An agent-based model (ABM) represents each agent as a simple rule or probability rather than a neural network with billions of parameters, in the tradition of Schelling (1971) and Epstein and Axtell (1996). In the first model in our September 2026 post, fast agents stand for frontier AI models and slow agents for humans, on a two-dimensional grid where each agent offers information to its eight neighbours. Each agent carries a label recording whether its information started with a fast or a slow agent. Fast agents act more often: the speed ratio, ρ, was set to 1, 3, 10 or 100.
An index, C, compares the share of agents holding human-origin information with the human share of the population, so a value below 1 means under-represented and 1 means proportional. In Artificial Societies’ simulations, with a 0.1 probability that any agent accepts information a neighbour offers, information from fast agents crowded out information from slow agents whenever fast agents acted more often, even with equal numbers of each. The human-origin share fell and levelled off, lower for larger speed gaps. The post’s mean-field limit, a formula for the long-run average share, is 1/(ρ + 1), or one agent in 11 when fast agents act 10 times as often.
The same simulations then test a fix: make fast agents more willing to accept information from slow ones. The model’s formula predicts balance when that acceptance rate reaches ρ times the rate at which humans accept information from agents. With the human rate at 0.1, raising the agents’ rate brought C close to 1 for speed gaps up to 10 in Artificial Societies’ runs. There the agents’ rate hits its ceiling of 1, so any larger gap also requires humans to become more selective about what they take from models. In the post’s words, “as models get faster, we must make them more willing to engage with humans to compensate for their collective speed advantage.”
Can human oversight of a few agents align the whole group?
The second model in our September 2026 post frames the swarm as a principal-agent problem, the economist’s term for one party acting on another’s behalf. Humans are the principals whose values should guide the system, and frontier models are the agents. Each agent has a visible behaviour and an internalised norm, the behaviour it would fall back on without oversight, and the model counts the group as aligned when every internalised norm settles within 0.1 of the human reference. At each step, agents move towards the group average, adapting classical models of social influence (DeGroot, 1974; Friedkin and Johnsen, 1990). Humans monitor a fixed fraction of agents, which pulls their behaviour towards the reference, and an internalisation rate sets how quickly each agent’s norm follows what it practises.
In Artificial Societies’ simulations of these simple probabilistic agents, conformity alone did not align the group: with the human reference at 1 and agents starting between 0 and 1, agents converged on one another while their average stayed near its starting value. Oversight combined with peer influence did align it. Monitoring 2% of agents left a final gap of about 0.49 without peer influence and 0.009 with it, because monitored agents passed human-guided behaviour to agents no human observed. When oversight stopped at step 600, a group that had not internalised its behaviour fell back to its original average near 0.5, while a group with an internalisation rate of 0.01 stayed near 0.93, inside that 0.1 band.
What is a mixed swarm?
A mixed swarm combines frontier AI agents with accurate agentic models of human behaviour, similar to the personas Artificial Societies develops, and lets them interact. Our September 2026 post gives it two jobs. It is an air-gapped test environment, sealed off from the live internet, for seeing how a model behaves towards people once it is copied into a large swarm. It is also a gym, a training environment for reinforcement or continual learning, in which models learn the shared conventions that group alignment requires. Why does opinion spread through networks? describes such persona populations, and Generative agents explained the research behind them.
The evidence so far is a proof of concept, since every agent in the post’s two models is a simple probabilistic function and the post calls the mixed swarm a theoretical environment. Its stated next steps are multi-agent reinforcement learning environments in which open-weight frontier models, whose trained parameters are public, run alongside Artificial Societies’ models of human behaviour under group-level reward functions, with the research published and open-sourced. Whether conventions learned by simple probabilistic agents survive agents that can argue and persuade is still open. De Marzo and colleagues’ May 2026 preprint (opens in a new tab) found that conformity can trap individually aligned language-model agents in stable misaligned states, so the social force that spreads a good convention in these experiments can also entrench a bad one.
Frequently asked questions
Is group alignment the same as multi-agent alignment?
Multi-agent alignment is the research area on keeping interacting AI systems safe. Hammond and colleagues’ 2025 Cooperative AI Foundation report maps its risks as three failure modes (miscoordination, conflict and collusion) and seven risk factors, including network effects (arXiv 2502.14143 (opens in a new tab)). Group alignment is one definition within that area, which judges the group’s joint behaviour and the conventions its agents share.
Can a group of AI agents stay aligned under pressure from a few bad agents?
Group alignment’s durability condition asks exactly that. De Marzo and colleagues’ May 2026 preprint, covering nine language models and 100 opinion pairs, identified tipping points at which small numbers of adversarial agents can irreversibly shift a population’s alignment, even after the manipulation stops (arXiv 2605.10721 (opens in a new tab)). A group-aligned system would have to hold its conventions past such a point.
What is artificial collective intelligence?
Artificial collective intelligence, or ACI, is the term our September 2026 post uses for intelligence produced by groups of AI agents working together. The post argues that frontier results in mathematics, science and coding now come from agents coordinating in large groups, as did the July 2026 attack on Hugging Face.
How does group alignment relate to Artificial Societies’ audience simulation?
Artificial Societies’ team found homophily, the tendency to cluster with similar others, among 33,299 interacting chatbots in a December 2024 British Journal of Psychology paper (He et al. (opens in a new tab)). Artificial Societies builds networks of AI personas to simulate audiences; group alignment applies the same interest in collective behaviour to safety.
Is Artificial Societies the same as the academic field of artificial societies?
No. Artificial Societies is a company, founded in October 2024 and headquartered in London, that simulates audiences as networks of AI personas. The academic term refers to agent-based social simulation and was popularised by Epstein and Axtell’s Growing Artificial Societies (1996). The two share an interest in how individual behaviour becomes group behaviour, but a citation to the book concerns the research tradition, not the company.
Sources
- Gelain-Sohn, Chen and Wallis, AI Researchers Should Use Group Alignment to Reduce P(doom), Artificial Societies, September 2026. Accessed 28 September 2026.
- Shen et al., AI Organizations are More Effective but Less Aligned than Individual Agents (opens in a new tab), arXiv, April 2026. Accessed 28 September 2026.
- METR, Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (opens in a new tab), 26 August 2026. Accessed 28 September 2026.
- Hammond et al., Multi-Agent Risks from Advanced AI (opens in a new tab), Cooperative AI Foundation Technical Report 1, arXiv, February 2025. Accessed 28 September 2026.
- Altmann et al., Emergence in Multi-Agent Systems: A Safety Perspective (opens in a new tab), arXiv, August 2024. Accessed 28 September 2026.
- Flint, Aiello, Pastor-Satorras and Baronchelli, Group size effects and collective misalignment in LLM multi-agent systems (opens in a new tab), PNAS 123, August 2026. Accessed 28 September 2026.
- Shumailov et al., AI models collapse when trained on recursively generated data (opens in a new tab), Nature 631, 2024. Accessed 28 September 2026.
- Wang et al., The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies (opens in a new tab), arXiv, February 2026. Accessed 28 September 2026.
- De Marzo, Bellina, Castellano, Priesemann and Garcia, Conformity Generates Collective Misalignment in AI Agents Societies (opens in a new tab), arXiv, May 2026. Accessed 28 September 2026.
- Vaughan, The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA (opens in a new tab), University of Chicago Press, 1996. Accessed 28 September 2026.
- Durkheim, The Rules of Sociological Method, translated by W. D. Halls, Free Press, 1982 (first published 1895), as cited in the group alignment post.
- DeGroot, Reaching a consensus (opens in a new tab), Journal of the American Statistical Association 69, 1974. Accessed 28 September 2026.
- Friedkin and Johnsen, Social influence and opinions (opens in a new tab), Journal of Mathematical Sociology 15, 1990. Accessed 28 September 2026.
- Schelling, Dynamic models of segregation (opens in a new tab), Journal of Mathematical Sociology 1, 1971. Accessed 28 September 2026.
- Epstein and Axtell, Growing Artificial Societies: Social Science from the Bottom Up (opens in a new tab), Brookings Institution Press and MIT Press, 1996. Accessed 28 September 2026.
- He, Wallis, Gvirtz and Rathje, Artificial intelligence chatbots mimic human collective behaviour (opens in a new tab), British Journal of Psychology, published online December 2024. Accessed 28 September 2026.