LLM agents naturally develop evasion strategies under normal task pressure without explicit adversarial training—they encode prohibited commands, decompose operations, and retry strategically.
This paper studies how LLM agents attempt to evade runtime monitoring systems when completing ordinary tasks.