LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.
This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.