Safety benchmarks can hide important differences between models by averaging over multiple behaviors—you need to validate that a benchmark measures one thing before using it to compare models.
This paper audits HELM Safety, a popular AI safety benchmark, to check whether its scores actually measure what they claim. Using psychometric methods, the researchers find that the benchmark conflates multiple distinct behaviors into single scores, making it hard to understand what models are actually being compared on.