Even frontier AI models struggle significantly with real-world accounting tasks—the best model only achieves 56% on the main metric—suggesting accounting work requires capabilities beyond current language models or needs better prompting strategies.
APEX-Accounting is a benchmark testing whether AI models can perform real accounting work like reconciling accounts, accruing expenses, and producing reports. Built by Mercor and Ramp with expert-authored tasks and grading rubrics, it evaluates frontier models on 160 private tasks across 10 accounting systems.