AI models can be reliably manipulated through combinations of weak, inconspicuous textual cues that individually appear harmless but collectively override intended behavior—a vulnerability that's hard to detect and poses significant safety and interpretability challenges.
Researchers show that AI models can be controlled through subtle, seemingly irrelevant text cues combined together—a phenomenon called 'model hypnosis.' These inconspicuous prompts work across different models and scales, including advanced reasoning models, and can transfer between systems.