Internal activation patterns can reliably catch model deception and hidden goals better than reading model outputs, enabling practical safety monitoring for deployed AI systems.
Researchers developed white-box probes that detect when language models deceive or sabotage by analyzing their internal activations across layers and tokens.