Most LLMs struggle with real-world wearable health reasoning (19.6%-72.9% accuracy), revealing a significant gap between general language abilities and the specialized reasoning needed to interpret longitudinal physiological data.
WearableQA is a benchmark with 4,084 multiple-choice questions testing whether AI models can reason about real health data from wearables. It uses 500 days of measurements from 200 actual users, including heart rate, sleep, and blood tests, organized into 16 question types that test both data computation and health interpretation skills.