LLM social reasoning is measurably worse when users present biased perspectives, and models often require excessive detail to match human-level social understanding—a critical gap for real-world advice scenarios.
This paper introduces Fuse, a simulation framework that tests how well AI assistants reason about social situations. By creating multi-agent scenarios where a hidden motive needs to be inferred from user narratives, the researchers can verify whether LLMs correctly understand social dynamics—something normally impossible to evaluate objectively.