For autonomous ML engineering, a minimal harness giving an LLM direct access to read, write, and bash commands performs as well as complex multi-agent systems—the model itself, not the infrastructure, drives performance.
This paper challenges the complexity of modern ML engineering agents by comparing elaborate multi-agent systems against a simple baseline where an LLM directly accesses code execution tools. The authors find that under equal time budgets, simpler agents perform as well as complex orchestrated systems, suggesting the LLM backbone matters far more than the surrounding machinery.