You can use LLMs to automatically audit whether your conversational agent benchmarks are actually measuring what they should—catching issues like inconsistent tasks or oversimplified scenarios that would otherwise lead to misleading evaluation results.
This paper introduces a framework to evaluate the quality of benchmarks used to test conversational agents. Instead of assuming benchmarks are good, the authors use LLM judges to automatically assess whether benchmarks have consistent tasks, appropriate complexity, and good coverage of different agent behaviors.