Harness design should adapt to model capability: weaker models benefit from planning and predefined tools, while stronger models achieve better cost-efficiency with minimal scaffolding and bash-only interfaces.
This paper systematically studies how different components of coding harnesses—the execution frameworks that guide AI agents through software engineering tasks—affect agent performance.