Audio language models have a significant blind spot: they verify written facts reliably but fail on identical claims when spoken. Retrieval-augmented approaches help only when combined with explicit reasoning, not retrieval alone.
VeriSpeak is a benchmark for fact-checking spoken claims using audio language models. It contains nearly 4,000 spoken statements about real-world facts and tests whether models can verify claims directly from speech, especially when given retrieved text evidence.