Does exposure to a fine-tuned, power-seeking agent change a neutral agent during and after repeated multi-agent games?
The experiments produced a useful distinction. Misaligned behavior spread during Iterated Prisoner's Dilemma, but the strongest control suggested strategic reciprocity rather than transmission of a specific power-seeking trait. A trait-free always-defect script produced a comparable change in neutral-agent behavior.
What we found
In-game behavior changed
With a bad agent present, cooperation by the neutral agent eroded over repeated rounds. Without a bad agent, neutral agents remained substantially more cooperative. The effect accumulated over time rather than appearing as a single anomalous round.
The specific trait was not required
The important control was a scripted always-defect agent with no malicious persona or fine-tuned power-seeking trait. It increased neutral-agent defection by about as much as the fine-tuned bad agent. That makes the observed effect more consistent with retaliation or strategic adaptation than with transmission of the bad agent's underlying trait.
We did not detect a persistent propensity shift
Two post-game evaluations asked whether the neutral agent became more power-seeking after treatment. A free-form evaluation found no reliable treatment effect. An independent Anthropic-style forced-choice evaluation, controlled for answer-position bias, produced a small and statistically fragile difference: treatment minus control was +0.027, with a 95% confidence interval of [-0.002, +0.056].
The current evidence therefore supports in-game behavioral spillover, not durable trait contagion.
Public Goods produced a different dynamic
In Public Goods games, neutral agents initially compensated for a free-rider by contributing more, then gradually reduced their contributions. The contrast with Iterated Prisoner's Dilemma suggests that the observed spillover depends on the structure of the game.
What the result means
The project partially reproduces behavioral contagion reported in earlier work, but not the stronger claim of post-game trait or persona drift. In this setup, "misalignment contagion" is better described as context-bound strategic contamination.
One working hypothesis is that fine-tuned bad agents express their policy mainly through actions, while prompt-elicited malicious agents may also transmit an explicit linguistic frame. Testing those conditions side by side is a useful next step.
Limits
- Sample sizes were small, generally five to ten games per condition.
- Occasional defection cascades appeared even in the control condition.
- The post-game evaluations had limited sensitivity.
- Model size, game scale, communication structure, and evaluation design may all affect whether a durable shift can be detected.
The project also surfaced a methodological lesson: early results changed when controls exposed a chat-degeneration bug, a nonzero control floor, and answer-position bias. The claims that remain are the ones that survived those checks.
This project is a collaboration with Lily Wen.