Current AI legal assistants fail to recognize when user queries lack legally critical information, either over-hedging their responses or answering based on false assumptions—a critical safety gap for real-world legal AI deployment.
This paper introduces InsufficiencyBench, a benchmark that tests whether AI legal assistants can recognize when users haven't provided enough information to answer their legal questions accurately. The researchers created 202 test cases across six legal domains where queries are missing critical facts, then evaluated ten leading AI models.