AI agents can generate code, but generating code with formal proofs that work together across entire repositories remains an unsolved problem—current best agents only solve 27 of 43 real-world instances.
Vero is a benchmark for evaluating whether AI agents can generate both correct code implementations and machine-checked formal proofs together across real multi-module software repositories.