LLMs don't arbitrate evidence rationally—they rely on predictable heuristics like recency bias and text preference, which can cause failures in real-world decision systems that combine multiple information sources.
This paper studies how large language models decide between conflicting evidence from text, numbers, and external tools.