LLM capability depends heavily on the surrounding harness (prompts, tools, orchestration), and frontier models show measurable but inconsistent ability to optimize these components—a skill that will become increasingly important as AI systems become more agentic.
This paper introduces HarnessOpt-Bench, a benchmark for measuring how well large language models can automatically improve AI agent systems by editing their prompts, tools, and control flow.