FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models — ThinkLLM