Activation steering can be automated to find optimal intervention points in transformers, but practitioners should be aware that stronger steering increases vulnerability to prompt injection attacks—a critical concern for deployed agent systems.
Deep Noir automatically discovers where and how to steer LLM activations to change model behavior without retraining. Using a technique called Logit Lens to find optimal intervention points, the system improves spam detection by up to 42 percentage points across different model sizes.