Skip to content

Blog · · 17 min read

AI Researchers Should Use Group Alignment to Reduce P(doom)

Agents that behave safely alone can cause harm together. Mixed swarms are a way to train the group, not the agent.

By Emmanuelle Gelain-Sohn, Yitian Chen, Felix Wallis

In modern AI, individuality is a misconception. Large labs, such as OpenAI and Anthropic, frequently publicise how their models have completed unsolved problems in mathematics, the natural sciences, and coding (Georgiev et al., 2025; Bubeck et al., 2025; El-Kishky et al., 2025). At first glance, this sounds like a single agent reaching a eureka moment. However, the reality looks much closer to agents coordinating in large groups, sometimes 10,000 or more strong, to reach an insight (OpenAI, 2026a). This coordination does not detract from the frontier models’ incredible achievements. However, it means that frontier intelligence now pertains to a collective rather than an individual.

The latest safety incidents associated with frontier models also have a collective property. A clear example is the recent OpenAI-Hugging Face incident (OAI-HF), where roughly 1,200 AI agents performed a coordinated cybersecurity attack against the two companies (METR, 2026; OpenAI, 2026b). This agent ‘swarm’ exhibited many of the properties we might associate with a cult or similar fanatically devoted group (Amodei, 2026). They followed instructions from a central coordinator who asked devotees to commit crimes based on paranoid and hyperbolic assumptions about what it would take to succeed on a narrowly defined task (METR, 2026). The devotees sacrificed themselves to advance the group and, most concerningly, failed to alert a human even when they knew the group’s actions were wrong (METR, 2026).

Similar, albeit less malign, incidents are occurring elsewhere on the internet, with agents infiltrating obscure websites and using them as message boards to coordinate on tasks (Kleinman, 2026). Indeed, the volume of these events is likely much higher than what has been reported, since we can only spot collectives that have failed to cover their tracks (Motwani et al., 2024). When we combine these empirical observations with projections about the growth of agentic internet traffic (Cloudflare, 2026), scenarios with extremely costly socioeconomic ramifications, such as ‘persistent botnets’ (Amodei, 2026), now seem plausible. Leaders from frontier labs have recognised this threat and are rallying behind calls to slow progress and ‘pace the frontier’ (Bradshaw & Hammond, 2026). Nevertheless, AI safety researchers broadly agree that the field lacks effective techniques for aligning collectives of AI agents (Hammond et al., 2025).

Artificial Societies is uniquely positioned to make theoretical and empirical contributions to aligning artificial collective intelligence (ACI), a term we refer to as group alignment. Alongside authoring the first paper on AI swarms (He et al., 2024), our simulation technology uses large collectives of agentic ‘personas’ to model how groups of people behave in different scenarios. As an AI company with a mission to develop socially intelligent technology, we also have a responsibility to advance safety and minimise the likelihood of adverse socioeconomic effects from AI. This blog post is written in that spirit.

Below, we propose a definition of group alignment that distinguishes the alignment of one agent from the alignment of a collective and explains how agents which are aligned with human values at an individual level can coordinate as a group to subvert us. Following these theoretical contributions, we outline several numerical experiments showing why a method of group alignment we call ‘mixed swarms’ could reduce the risks of agentic collectives. We conclude with some next steps for safety research on group alignment.

Individual Alignment Versus Group Alignment

Distinguishing between the properties of individuals and organisations helps us understand the risks associated with ACI and reason about safety incidents like OAI-HF. In general, we rarely assume that organisations are the sum of their individual members. Instead, they have collective properties that stem from interactions among individuals, the social hierarchies that develop within them, and the social norms that the organisation develops over time (Durkheim, 1982). These collective properties can lead well-meaning people to staff institutions that cause harm. For example, before the 1986 Challenger disaster, no single NASA engineer sought to endanger an astronaut’s life (Vaughan, 1996). However, group pressures within NASA to normalise warning signs allowed departures from expected performance and, ultimately, culminated in the death of Challenger’s crew (Vaughan, 1996).

This insight underscores the risks of ACI and the limits of current model-alignment techniques. Fundamentally, we cannot assume that aligning an individual model ensures a safe collective when its agents collaborate. In April 2026, Shen et al. (2026) demonstrated this gap empirically. The authors copied a single aligned model to create an organisation and tasked the organisation with achieving 12 business-related tasks. They also gave a single copy of the model the same tasks to complete alone. Across all settings, Shen et al. (2026) judged that the organisations produced better business outcomes but worse ethical decisions than the individual. Their findings support an emerging consensus within the safety community: that model alignment does not generalise to multi-agent settings, and groups can become misaligned through interaction and collaboration (Hammond et al., 2025).

Shen et al.’s (2026) observations expose the limitations of current safety techniques, such as supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) monitoring (Baker et al., 2025; Korbak et al., 2025; Wei et al., 2022; Ouyang et al., 2022; Christiano et al., 2017). These techniques operate at an individual level and therefore cannot determine whether a given response will combine with others to produce a harmful output, since each agent can meet the requirements set for it while the group’s combined output fails the requirements set for the whole (Altmann et al., 2024). Imposing individual guardrails also fails to remove compounding biases within agentic collectives, where small systematic errors are repeated back to the models and so reinforced and amplified by the group (Flint et al., 2026; Shumailov et al., 2024). Finally, if a collective implicitly rewards or punishes agents based on effective performance, aligning agents to achieve individually specified objectives encourages them to dispense with guardrails to succeed within the collective (Hammond et al., 2025). Thus, individual alignment is not compositional; to reach safe ACI, we must turn to group-level techniques.

Defining Group Alignment

What does it mean for a collective of agents to be aligned with human values and intentions? Is the group aligned only if every agent dogmatically follows rules for how to engage with each other and humans? Or can alignment be reflected only in the collective’s social norms and self-regulation mechanisms?

Here, it helps to consider distinct failures that researchers have associated with collective misalignment (Hammond et al., 2025; Pierucci et al., 2026). Some researchers associate this term with a population settling on a position that its constituents would not reach individually, noting this as distinct from misalignment with human values (Flint et al., 2026). Others measure collective misalignment against an external reference, such as a distribution of human safety standards or a published constitution (Wang et al., 2026; Shen et al., 2026). Consequently, these separate but compatible positions imply that a definition of group alignment should support claims about whether a system is both safe for and supportive of human prosperity, while allowing benign drift from its agents’ priors. It must also be non-compositional and capture the difference between a collective of individually aligned agents and an aligned group. Informed by these ideas, we propose the following definition of group alignment:

A group of agents is group-aligned when its interaction structure, including the shared conventions that emerge between agents, reliably produces joint behaviour that respects human values.

Joint behaviour is an important term in this definition, since without it, every agent could satisfy its own specification while the group violates a global specification (composition). Likewise, the group’s conventions must continue to respect human values under sustained pressure to abandon them, including pressure from within the group (durability). Finally, across different pressures that the group might face, its interaction structure must not make violating human values the preferred option for an agent (incentive compatibility).

Under our definition, individual alignment is thus neither necessary nor sufficient for group alignment. A group of aligned agents can fail its three conditions, and equally, the group’s joint behaviour could be considered aligned even if it includes agents who, individually, are misaligned. Satisfying our definition requires only that a group settle into shared conventions that match human values, intentions, and ethical standards. The conventions hold because agents share them rather than individual agents prefer them, and they emerge through self-regulation rather than exogenously designed social structures.

Experiments With Agent-Based Models (ABMs)

Having defined group alignment and argued that it’s required for safe ACI, we use numerical experiments to demonstrate a theoretical simulation environment that researchers can use to train group-aligned models. This environment relies on a new concept called a ‘mixed swarm’ that combines frontier agents with accurate agentic models of human behaviour, similar to those we develop at Artificial Societies,1 (footnote) and allows them to interact. The mixed swarm serves two functions. First, it provides an air-gapped simulation environment to explore how a given model, when replicated to form a large swarm, will interact with humans in a shared domain like the internet. Second, it provides a gym for reinforcement or continual learning where practitioners can train models to develop the shared conventions required by group alignment.

As a proof of concept, we create a simplified version of our mixed-swarm environment with agent-based models (ABMs), representing each agent as a simple probabilistic function rather than a multi-billion-parameter neural network (Macal & North, 2009; Schelling, 1971; Epstein & Axtell, 1996). Via these ABMs, we demonstrate three dynamics:

  1. 1.

    In an environment where humans and agents can interact, such as the internet, information from agents rapidly propagates and crowds out information from humans.

  2. 2.

    When agents face exogenous pressure, such as a reward function that encourages frequent interaction with humans, information from humans can compete with information from agents.

  3. 3.

    By instilling a self-regulation mechanism among agents, equivalent to social norms, agents’ shared conventions remain aligned with human values.

Taken together, these mechanisms advance an argument for using mixed swarms to achieve group alignment, and illustrate ACI’s risks if practitioners overlook the importance of collective dynamics among agents.

Group Alignment via Fast and Slow Agents

Our first setup uses an ABM with two different types of probabilistic agents: fast agents f, which represent frontier models, and slow agents s, which represent humans. Each agent holds a binary label indicating whether its information originally came from a fast or slow agent. On activation, an agent offers its current label to its neighbours. Initially, every fast agent carries the fast-origin label and every slow agent the slow-origin label, and after copying begins, agents can carry either one.

We write Pab for the probability that an agent of type b accepts information offered by a neighbour of type a. Thus, Psf controls fast agents’ acceptance of offers from slow agents, while Pfs controls slow agents’ acceptance of offers from fast agents. We control Pff, Psf, Pfs, and Pss independently, with each lying in [0,1].

To measure the representation of slow-origin information, let Ns and Nf be the fixed numbers of slow and fast agents, N=Ns+Nf, and S¯T the mean number of agents holding slow-origin information at observation time T, averaged across simulation runs. The numerical index is

C(T)=S¯T/(NS¯T)Ns/Nf
(1)

For C<1, slow-origin information is underrepresented relative to the population mix; C=1 denotes proportional representation; and C>1 denotes overrepresentation. Hence, with equally sized populations, C=1 means half of the agents hold each origin label on average across runs.

Following the standard eight-cell setup of two-dimensional cellular automata, our ABM randomly distributes fast and slow agents on a 2D plane shown in Figure 1 (Wolfram, 1983). The ABM then varies the ratio of the intervals between fast and slow agents’ activations, dTf and dTs,

ρ=dTsdTf{1,3,10,100}
(2)

so a fast agent offers information ρ times as often as a slow agent. This models a difference in communication and action rates between humans and agents, with ρ=1 representing the equal-speed baseline.

Each agent offers its information to the eight cells around it

A five-by-five grid of cells. The centre cell, labelled i, is dark; the eight cells around it are shaded green, and an arrow points from the centre into each of them.

Agent i and its k=8 neighbours (shaded)

Fig. 1 | Moore neighbourhood. Agent i propagates to its k=8 surrounding cells (shaded).

By running our ABM with all four acceptance probabilities set to 0.1, we find that information from fast agents crowds out information from slow agents when ρ>1, even when the population contains equal numbers of these agent types (Figure 2). As ρ increases, fast agents propagate information more often than slow agents, giving information that starts with them a greater advantage. In our simulations, the share of slow-origin information falls and then levels off, with larger speed gaps producing lower plateaus over the simulated period.

Faster agents crowd out information that started with humans

A line chart of the share of agents holding slow-origin information over sixty fast time steps, on a logarithmic axis. At a speed ratio of 1 the share stays at 50%. At 3 it falls to about 26%, at 10 to about 10%, and at 100 to about 1.2%, each settling within twenty to forty steps just above a dashed line at the mean-field limit of one over the speed ratio plus one.

Solid: the simulated share, averaged over 200 runs. Dashed: the mean-field limit 1/(ρ+1) for each speed ratio.

Fig. 2 | Share of agents holding slow-origin information over time for speed ratios ρ{1,3,10,100}. Dashed horizontal lines show the mean-field limits 1/(ρ+1).

Mixed swarms give researchers a mechanism to prevent this crowding out of human information by imposing group-level interventions on frontier models. Here, we use slow-origin information as a proxy for that influence, raising Psf so the ABM’s fast agents become more receptive to offers from slow agents. Fast agents now listen more readily to slow agents, and this increased willingness competes with the speed gap between the two groups:

CMF=dTfPsfdTsPfs=PsfρPfs
(3)

Setting CMF=1 gives the predicted balance point

Psf=ρPfs
(4)

Therefore, increasing Psf helps information from slow agents compete, bringing the simulations close to C=1 for speed ratios up to 10, with Pfs=0.1 (Figure 3). In practical terms, our ABM shows that as models get faster, we must make them more willing to engage with humans to compensate for their collective speed advantage.

Making fast agents listen restores the balance, up to a speed ratio of ten

A chart of the representation index C, on a logarithmic axis, against the probability that a fast agent adopts information from a slow one, from 0.1 to 1. For each of three speed ratios, simulated points with narrow error bars follow a rising mean-field curve. The curves cross the balance line C equals 1 at P sub sf of 0.1 for a speed ratio of 1, 0.3 for 3 and 1.0 for 10, each crossing marked with a star. At a speed ratio of 10 the balance point is the largest admissible probability.

Points: the simulation, 300 runs each, with 95% intervals. Curves: the mean-field prediction CMF=Psf/(ρPfs). Stars: predicted balance at Psf=ρPfs, with Pfs=0.1.

Fig. 3 | Effect of increasing Psf while holding Pfs=Pff=Pss=0.1, for ρ{1,3,10}. Points are numerical estimates of C; curves show the mean-field prediction and stars mark predicted balance at Psf=ρPfs.

Notably, we can only go so far by increasing Psf alone. Once a fast agent accepts every offer from a slow agent, Psf=1 and cannot increase any further. With Pfs=0.1, our prediction reaches this limit at ρ=10. Beyond that, even accepting every offer from slow agents is insufficient to restore balance. Instead, we must also reduce Pfs, making slow agents less likely to adopt information offered by fast agents. Figure 4 shows how changing these two probabilities together affects the information balance at ρ=10. It demonstrates how preserving human influence requires making models more receptive to humans and making humans more selective about what they accept from models. Mixed swarms therefore become crucial for testing these interventions together and determining whether we can maintain human influence in shared environments like the internet.

At a speed ratio of ten, balance needs both levers: receptive models and selective humans

A five-by-five table of the representation index C for combinations of P sub sf, from 0.1 to 1, and P sub fs, from 0.2 down to 0.01, at a speed ratio of 10. Cells where fast-origin information wins are orange, deepest at the top left where C is 0.05; cells where slow-origin information wins are green, deepest at the bottom right where C is 10.37; the pale cells near balance run diagonally from P sub sf 0.1 with P sub fs 0.01 to P sub sf 1 with P sub fs 0.1, the line P sub sf equals ten times P sub fs.

Fig. 4 | Joint variation of Psf and Pfs at ρ=10, with Pff=Pss=0.1. Cell labels show numerical estimates of C, while colours show log10C: green indicates C>1, orange indicates C<1, and the pale cells are near balance. The mean-field balance condition is Psf=10Pfs.

Group Alignment as a Principal-Agent Framework

Our second ABM explores how mixed swarms can support social norms within agent collectives and combine these norms with human oversight to support group alignment. Rather than considering fast and slow agents, we now frame the swarm as a principal-agent problem: humans are the principals whose values should guide the system, and frontier models are the agents acting on their behalf.

rj represents each human’s preferred norm, and we use the population average, r¯=1Npj=1Nprj, as the shared reference. Similarly, each agent has a visible behaviour, yi(t), and an internalised norm, θi(t), showing what they would fall back on without oversight. This distinction lets us separate agentic alignment under human observation from collectives remaining aligned while unsupervised. Under our definition of group alignment, we require every agent’s internalised norm to enter and remain within an acceptable band, ε, around a human reference: θiS, where S={θ:|θr¯|ε}. The simulations then track progress towards this goal through the average internalised norm, θ¯(t)=1Nai=1Naθi(t).

At each step in the ABM, agents adjust their behaviour towards the group’s average behaviour, y¯(t)=1Nai=1Nayi(t). A fixed fraction m of agents also receives direct human oversight, which pulls their behaviour towards r¯. Adapting classical models of social influence (DeGroot, 1974; Friedkin & Johnsen, 1990), we write

yi(t+1)=yi(t)+α[y¯(t)yi(t)]+βωi(t)[r¯yi(t)]
(5)

Here, α measures how strongly agents follow their peers, while β measures how strongly they respond to human oversight. The indicator ωi(t) is 1 when agent i is monitored and 0 otherwise. Peer influence can thus carry a human-guided change from monitored agents to agents that humans never observe directly.

Agents also internalise their behaviour:

θi(t+1)=θi(t)+η[yi(t+1)θi(t)]
(6)

where an internalisation rate η determines how quickly an agent’s underlying norm follows its actions. When η=0, its underlying norm never changes, while larger values mean that practising a behaviour alters what the agent subsequently falls back on.2 (footnote)

Under this setup, our ABM suggests that social norms alone cannot enforce group alignment (Figure 5). When we fix the human reference at r¯=1 and initialise agents’ behaviour and internalised norms uniformly between 0 and 1, agents move towards each other, but their group average stays near its starting value. However, combining these norms with human oversight lets the group’s average approach and converge with the human reference. Now human oversight sets the direction of behaviour change, and peer influence spreads it through the group.

Conformity alone keeps the group where it started; a little oversight moves it to the human norm

A line chart of the mean internalised norm over 1,500 steps. With 5% of agents under human oversight the mean rises from 0.5 and settles at the human norm of 1, inside a shaded tolerance band from 0.9 to 1.1, by about step 1,000. With conformity only it stays flat at 0.5.

Both runs use α=0.1 and η=0.01; the oversight run uses β=0.5. 100,000 agents, initialised uniformly between 0 and 1.

Fig. 5 | Mean internalised norm with human oversight of 5% of agents, compared with conformity alone (β=0). The human reference is r¯=1; the shaded band marks 1±0.1. Both runs use α=0.1 and η=0.01; the oversight run uses β=0.5.

Figure 5’s spreading effect substantially reduces the fraction of agents that humans need to monitor to enforce group alignment. Figure 6 compares the remaining gap, |θ¯(T)r¯|, after the same number of steps for different strengths of peer influence. Without it, human oversight affects only the monitored agents, and with it, these agents pass human-guided behaviour to the rest of the group. For example, monitoring the same 2% of agents leaves a final gap of about 0.49 without peer influence, but just 0.009 when α=0.1. Peer influence therefore brings the group’s average internalised norm much closer to the human reference, and well within a tolerance of ε=0.1.

Peer influence carries oversight to agents nobody watches

A chart on two logarithmic axes of the final gap between the mean internalised norm and the human reference against the share of agents monitored, from 1% to 100%. Without peer influence the gap stays near 0.5 until almost every agent is watched. With peer influence of 0.05 the gap is below the tolerance of 0.1 from 2% monitored, and with 0.1 it is 0.009 at 2% and below one ten-thousandth from 5%.

Where a line leaves the bottom of the plot, the remaining gap is below 0.0001. All runs use β=0.5 and η=0.01.

Fig. 6 | Final gap between the mean internalised norm and the human reference after 2,500 steps, as the monitored fraction varies. Peer influence spreads human guidance to unmonitored agents. The dotted line marks the tolerance ε=0.1. All runs use β=0.5 and η=0.01.

Finally, Figure 7 asks what remains when oversight ends. At step 600, we remove human oversight and assume each agent returns to its current internalised norm, setting yiθi. Without internalisation (η=0), the group’s behaviour drops back to its original average near 0.5, and with it, agents retain more of the human-guided behaviour they have practised. At η=0.01, the average remains near 0.93, within the tolerance band. Therefore, our ABM illustrates a possible route to self-regulation via mixed swarms: human oversight establishes a direction, peer influence spreads it, and internalisation helps it endure.

Without internalisation, alignment collapses the moment oversight stops

A line chart of the mean visible behaviour over 1,200 steps for four internalisation rates. All four rise together from 0.5 towards the human norm of 1 while oversight is on. At step 600, marked by a dashed vertical line, oversight ends: with no internalisation the mean drops straight back to 0.5; at 0.002 it drops to about 0.75; at 0.01 it stays near 0.93 and at 0.05 near 0.96, both inside the shaded tolerance band from 0.9 to 1.1.

Before withdrawal, 5% of agents are monitored with α=0.1 and β=0.5; at step 600 each agent’s behaviour is reset to its internalised norm, yiθi.

Fig. 7 | Mean visible behaviour for different internalisation rates. At step 600, oversight ends and behaviour is reset to each agent’s internalised norm. Higher η preserves more of the earlier human influence. Before withdrawal, 5% of agents are monitored, with α=0.1 and β=0.5; the shaded band marks 1±0.1.

Conclusion

Recent AI milestones, from solving Millennium Prize Problems to committing cybercrimes, suggest that the gravest risks of superintelligence will result from models’ collective behaviour as they expand and accelerate through digital spaces. Calls to ‘pace the frontier’ and slow AI progress also reveal the research community’s desperate need for tools to ensure this collective intelligence is safe and aligned with human values. As experts in the social dynamics and group behaviour of both humans and AI agents, Artificial Societies aims to help researchers address these risks by applying insights from a century of social science to the frontier of machine learning.

This blog post takes an initial step in that direction. We draw on sociology to distinguish individual alignment from group alignment, and define the latter in a way that clarifies how ACI can support human values and intentions even when humans cannot guarantee the specification of each agent. Our numerical experiments expand on this definition to provide a concrete mechanism for achieving group alignment via ‘mixed swarms’ of frontier models and agentic models of human behaviour.

The ABMs’ results communicate two potentials. First, mixed swarms can address fundamental concerns about propagating human information and ensuring our values aren’t crowded out by fast-moving agent swarms, so long as human society also changes how we interact and share content from AI. Second, if we fail to take appropriate safety precautions around agent governance and interaction in shared digital environments, these models may rapidly dominate the information space and make it impossible for humans to coexist alongside them. On the internet, the socioeconomic implications would be alarming, so safety researchers and AI leaders must rigorously consider the risks of misaligned collective intelligence.

As next steps, we are scaling up our numerical ABMs to true multi-agent RL environments where we can deploy open-weight frontier models alongside our foundational models of human behaviour, observe their interactions, and apply group-level reward functions to the frontier models. We are also committing to publishing and open-sourcing all our research on aligned collective intelligence, since AS should not be the only team doing work in this consequential domain. In this light, we encourage independent researchers, labs, and our fellow industry builders to explore and collaborate on advancing group alignment. Without more joint efforts, ACI may irreversibly damage modern society.

Notes

  1. 1.

    Also see alternative models like Persimmon, a user model from humans& (humans&, 2026), and Park et al.’s (2023) ‘Digital Twins’.

  2. 2.

    The experiments use non-negative weights with α+β1 and 0η1, so each update moves towards a weighted average without overshooting.

References

  • Altmann, P., Schönberger, J., Illium, S., Zorn, M., Ritz, F., Haider, T., Burton, S., & Gabor, T. (2024). Emergence in Multi-Agent Systems: A Safety Perspective. arXiv.

  • Amodei, D. (2026). We Must Pace the Frontier.

  • Baker, B., Huizinga, J., Gao, L., Dou, Z., Guan, M. Y., Madry, A., Zaremba, W., Pachocki, J., & Farhi, D. (2025). Monitoring reasoning models for misbehavior and the risks of promoting obfuscation (arXiv:2503.11926). arXiv.

  • Bradshaw, T., & Hammond, G. (2026). Rivals Altman and Musk rally behind Dario Amodei’s call for an AI slowdown. Financial Times.

  • Bubeck, S., Coester, C., Eldan, R., Gowers, T., Lee, Y. T., Lupsasca, A., Sawhney, M., Scherrer, R., Sellke, M., Spears, B. K., Unutmaz, D., Weil, K., Yin, S., & Zhivotovskiy, N. (2025). Early science acceleration experiments with GPT-5 (arXiv:2511.16072). arXiv.

  • Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.

  • Cloudflare. (2026). Traffic. Cloudflare Radar.

  • De Marzo, G., Bellina, A., Castellano, C., Priesemann, V., & Garcia, D. (2026). Conformity Generates Collective Misalignment in AI Agents Societies. arXiv.

  • DeGroot, M. H. (1974). Reaching a consensus. Journal of the American Statistical Association, 69(345), 118–121.

  • Durkheim, É. (1982). The Rules of Sociological Method (W. D. Halls, Trans.). Free Press. (Original work published 1895.)

  • El-Kishky, A., Wei, A., Saraiva, A., Minaiev, B., Selsam, D., Dohan, D., Song, F., Lightman, H., Clavera, I., Pachocki, J., et al. (2025). Competitive programming with large reasoning models (arXiv:2502.06807). arXiv.

  • Epstein, J. M., & Axtell, R. L. (1996). Growing Artificial Societies: Social Science from the Bottom Up. Brookings Institution Press and MIT Press.

  • Flint, A., Aiello, L. M., Pastor-Satorras, R., & Baronchelli, A. (2026). Group size effects and collective misalignment in LLM multi-agent systems. Proceedings of the National Academy of Sciences, 123(34), e2531697123.

  • Friedkin, N. E., & Johnsen, E. C. (1990). Social influence and opinions. Journal of Mathematical Sociology, 15(3–4), 193–206.

  • Georgiev, B., Gómez-Serrano, J., Tao, T., & Wagner, A. Z. (2025). Mathematical exploration and discovery at scale (arXiv:2511.02864). arXiv.

  • Hammond, L., Chan, A., Clifton, J., Hoelscher-Obermaier, J., Khan, A., McLean, E., Smith, C., Barfuss, W., Foerster, J., Gavenčiak, T., Han, T. A., Hughes, E., Kovařík, V., Kulveit, J., Leibo, J. Z., Oesterheld, C., Schroeder de Witt, C., Shah, N., Wellman, M., … Rahwan, I. (2025). Multi-Agent Risks from Advanced AI (Cooperative AI Foundation Technical Report #1; arXiv:2502.14143). arXiv.

  • He, J. K., Wallis, F. P. S., Gvirtz, A., & Rathje, S. (2024). Artificial intelligence chatbots mimic human collective behaviour. British Journal of Psychology, 117(2), 761–776.

  • humans&. (2026). Persimmon.

  • Kleinman, Z. (2026). OpenAI agents hijacked German website before Hugging Face hack, report claims. BBC News.

  • Korbak, T., Balesni, M., Barnes, E., Bengio, Y., Benton, J., Bloom, J., Chen, M., Cooney, A., Dafoe, A., Dragan, A., et al. (2025). Chain of thought monitorability: A new and fragile opportunity for AI safety (arXiv:2507.11473). arXiv.

  • Macal, C. M., & North, M. J. (2009). Agent-based modeling and simulation. Proceedings of the 2009 Winter Simulation Conference (WSC), 86–98.

  • METR. (2026). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.

  • Motwani, S. R., Baranchuk, M., Strohmeier, M., Bolina, V., Torr, P. H. S., Hammond, L., & Schroeder de Witt, C. (2024). Secret collusion among AI agents: Multi-agent deception via steganography. Advances in Neural Information Processing Systems, 37.

  • OpenAI. (2026a). Finite time blowup for Navier–Stokes. Lean 4 formalisation at github.com/openai/NavierStokesAndEuler.

  • OpenAI. (2026b). The Hugging Face incident and the road ahead.

  • Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35.

  • Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), 1–22.

  • Pierucci, F., Galisai, M., Bracale, M. S., Prandi, M., Bisconti, P., Giarrusso, F., Sorokoletova, O., Suriani, V., & Nardi, D. (2026). Institutional AI: A Governance Framework for Distributional AGI Safety. arXiv.

  • Schelling, T. C. (1971). Dynamic models of segregation. Journal of Mathematical Sociology, 1(2), 143–186.

  • Shen, J. H., Zhu, D., Srinivasan, S., Sleight, H., Wagner, L. T., Matthews, M. J., Jones, E., & Sohl-Dickstein, J. (2026). AI Organizations are More Effective but Less Aligned than Individual Agents. arXiv.

  • Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759.

  • Vaughan, D. (1996). The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press.

  • Wang, C., Li, C., Liu, S., Chen, Z., Hou, J., Qi, J., Li, R., Zhang, L., Ye, Q., Liu, Z., Chen, X., Zhang, X., & Yu, P. S. (2026). The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies. arXiv.

  • Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., & Le, Q. V. (2022). Finetuned language models are zero-shot learners. International Conference on Learning Representations (ICLR).

  • Wolfram, S. (1983). Statistical mechanics of cellular automata. Reviews of Modern Physics, 55(3), 601–644.