Multi-agent LLM scaling isn't automatic—task structure and how you combine outputs matter more than team size. On some tasks, adding agents barely helps because the underlying model's systematic biases affect all team members the same way.
This paper analyzes how multi-agent LLM teams scale based on task structure, using Steiner's taxonomy to distinguish disjunctive tasks (where one correct answer helps) from compensatory tasks (where averaging helps).