Current AI agents show unreliable performance on quantum engineering tasks despite appearing capable—systematic benchmarking is needed to build trustworthy autonomous quantum systems.
This paper introduces Quantum-Harbor, a virtual lab for testing AI agents on quantum engineering tasks, and QIQCBench, a benchmark with 49 expert-designed tasks covering calibration, error correction, and quantum sensing. Testing 17 AI agents reveals significant gaps between claimed capabilities and reliable performance in real quantum operations.