Fixing your evaluator is as important as fixing your policy—when agents improve, what they need to be judged on changes, so you need a system to evolve your judge alongside your agent.
VeriFine is a framework that improves AI agents through co-evolving the policy, training data, and evaluation judge. When an agent's performance plateaus, humans help refine the judge by resolving disagreements on tricky cases, then the improved judge guides better training. Tested on driving and robot navigation, it shows continuous improvement as new failure patterns emerge.