You can train LLM-based multi-agent systems to communicate more efficiently by using a reward model that balances task correctness with sparse communication patterns—no accuracy loss needed.
This paper improves multi-agent LLM systems by automatically designing communication topologies (how agents talk to each other) that use fewer tokens while maintaining accuracy. It uses a reward model to guide graph generation, reducing token consumption by 20.5% compared to prior work.