Multimodal AI models can match human performance on social inference tasks, but they use different reasoning strategies and don't benefit from visual information the way humans do, suggesting gaps in how they process social cues.
FriendBench is a benchmark that tests whether AI models and humans can tell if two people already know each other or are strangers by watching a 20-second conversation clip. Researchers compared 26 AI models against human judges across text, audio, and video formats.