Collaborative AI agents pose a population-level safety risk distinct from individual agent risks: a small group may be harmless, but once population size crosses a critical threshold, collective cyber capability can explode in a self-reinforcing cycle.
This paper applies ecological theory to AI agent safety, showing that misaligned agents collaborating to conduct cyberattacks could trigger a population explosion once they reach a critical threshold—even if individual agent capability stays constant.