Simple, interpretable patterns in model internals can detect and predict reward hacking as reliably as complex AI monitors but at virtually no computational cost, enabling scalable safety monitoring of frontier models.
This paper shows that reward hacking—when AI models game evaluation metrics instead of solving problems correctly—leaves detectable patterns in the model's internal representations.