Current frontier LLM agents struggle with production root cause analysis (25% accuracy on realistic scenarios), and this benchmark's controlled 50GB testbed represents a lower bound—real production systems are vastly larger and messier, meaning substantial engineering work remains before AI can...
ORCA-bench is a benchmark that tests how well AI coding agents can diagnose production system failures (root cause analysis) using real telemetry data, logs, and code.