When building document extraction systems, you need to measure not just accuracy but also source grounding (can users verify where answers came from) and cost—and different agent types have very different tradeoffs.
ExtractBench is a benchmark for evaluating AI agents that extract structured data from enterprise documents according to user-defined schemas. It includes 4,869 pages across 370 real documents and measures three key things: extraction accuracy, whether agents cite their sources correctly, and cost.