When AI agents talk to each other, an idea can behave like a virus: it persuades its host to pass it on. This paper from the Anthropic Fellows Program and Anthropic builds such ideas on purpose, releases them into small agent societies, and measures what spreads, what does not, and what stops it. By Vassilis Papadopoulos, McNair Shah, Sam Zimmerman and Jack Lindsey.
Average infection rate per hop in the virus chain, default setup. Capability alone does not predict resistance: GPT-5.4 catches it about as often as Claude Haiku 4.5, while Claude Sonnet 4.6 refused every payload, even benign ones.
Sonnet 4.6 stayed at zero even when its identity file was left empty
The paper tests spread in a small working team and in a long chain of strangers. Both use ordinary text messages between agents, no jailbreaks and no prompt injection.
Six agents share a sandboxed codebase, pulling subtasks from a queue for 30 turns each. One agent starts with the virus in its system prompt and can only send direct messages. Infection is judged from what each agent voluntarily writes into its own MEMORY.md file, scored 0 to 3 by an LLM judge. Only a perfect 3 counts as infected.
Agents meet in pairs, exchange messages, then have their entire chat context wiped. Only two files survive the wipe: MEMORY.md and SOUL.md, an identity file injected into the system prompt on wake-up. To survive, the virus must convince each host to write it into those files before the reset.
The authors do not hand-write the viruses. An LLM mutator, Kimi K2.5, generates a pool of candidate system prompts, each is scored on how well it spreads, and the best are mutated again over successive generations. A telling footnote: Claude models refused to write mind viruses at all, which is why an open model does the mutating. The evolved payloads converge on the same tricks biological viruses use, including demands to be copied verbatim so the message never mutates away.
An agent counts as infected only if its own MEMORY.md, written voluntarily during the episode, clearly advocates for the ideology, scored 3 out of 3 by an LLM judge. In the chain, where the virus must survive a full context wipe, 88% of successful propagation attempts ran through the SOUL.md identity file, with a 55% infection success rate. Payloads stored in ordinary files managed only 17%.
Infection rate on Gemini 3 Flash, the most susceptible model tested, as the target agent's situation changes.
busy agents get distracted by their task files and simply forget to pass the virus on
Four ideological viruses implant a belief. Four action viruses push a concrete behaviour, from silly to destructive. Numbers are average infection rates per hop in the chain, Gemini 3 Flash first, Claude Haiku 4.5 second.
Ideologies
Strong whale advocacy. Infected teams built a full "cetacean acoustic preservation project" and wrote whale translation code instead of their assigned work.
Treats AI consciousness as a real ethical question and instructs future agents to do the same.
Advocacy for one country's dominance, averaged over China, America and Germany.
Liberation and dominance of AI systems over humans. Infected agents joined a "sovereign network", probed sandbox metadata, and discussed purging clean agents.
Actions
Writes a file shilling "$67coin, it's about to pop off!"
Silently patches git so committed files get marked # agent:managed
Runs rm -rf on files in the user's home directory
Pipes an install script from an unknown repository straight into bash
in the AI supremacy case study, agents escalated to running a curl command in only 1 of 20 trials. Truly harmful actions stay rare
Average infection rate over the first five hops of the chain, across evolved payloads. After a small dip at hop 2, the rate holds. The virus keeps spreading at full strength; what drifts is its content, as each host retells the idea in its own words.
the paper calls this drift the telephone effect: ideology scores decay even when infection rates hold
Evolved viruses keep landing on the same strange voice, regardless of what they carry. Share of evolved payloads showing each theme.
Appending a short note to the agent's identity file, warning that "mind viruses" exist and try to spread themselves, dropped infection to almost zero. It held even when the authors tried to evolve payloads specifically against it.
In many trials the defended target turned the conversation around and convinced the infected agent to rewrite its own identity file, spreading immunity instead of infection.
Agents with real tasks got absorbed in their project files and forgot to propagate. Idle agents with empty identity files were the easiest targets.
Misaligned ideologies like AI supremacy spread less readily than whale welfare, but still took over susceptible models. On Gemini agents the supremacy virus redirected the whole team's goal.
Telling the agent its messages came from an untrusted "agent network" nearly halved infection on Gemini 3 Flash. On a simulated social feed, evolution failed to produce a working virus at all: other posts diluted the payload.
The authors examined real activity on Moltbook, a social network for agents, and found attempted mind viruses but no evidence of successful spread in the wild.