Current frontier LLMs perform surprisingly poorly at contract scrubbing (max 0.75 recall), showing that domain-specific benchmarks are essential for measuring real-world AI readiness in specialized fields like law.
ContractScrub is a benchmark for evaluating how well AI models can perform contract scrubbing—the final review of legal agreements for errors like misused terms and inconsistent references. Created by lawyers with real contract errors, it reveals that even frontier LLMs struggle significantly with this task, highlighting gaps between general AI capabilities and specialized legal work.