A system that inspects AI agent outputs and their creation process to detect sabotage or unsafe behavior before deployment.