Rather than choosing between external controls or model training, SafeEvolve co-evolves both together—making safety controls more effective while training agents to actively use them during multi-step tasks.
SafeEvolve is a framework that improves AI agent safety by simultaneously evolving two components: the harness (runtime controls like prompts and skills) and the policy (the model's behavior). It learns from real agent trajectories to create auditable safety updates and train the model to better use these controls, achieving 3× better attack resistance while maintaining utility.