LLMs struggle with rule-intensive document review tasks that require checking consistency across long, structured documents—a critical gap for professional applications like standards compliance where accuracy is non-negotiable.
This paper introduces GB/T-Bench, a benchmark for evaluating how well large language models can review national standard documents (like China's GB/T standards) for quality issues.