LLM agents are faster and cheaper than humans on office tasks, but they don't yet match human quality—this benchmark lets you measure the cost-quality tradeoff for your use case.
OmegaUse-OfficeVal is a benchmark for testing LLM agents on realistic office tasks (like document editing, spreadsheet work) that take ~2.3 hours of human labor each. It uniquely pairs tasks with economic data—human labor costs and LLM inference costs—so you can directly compare whether AI is cheaper and faster than hiring someone.