Current LLM agents show some ability to optimize hyperparameters based on experimental feedback, but struggle with sustained iteration, understanding complex logs, and consistently reaching target performance—revealing gaps between agent reasoning and practical scientific optimization.
AgentHPOBench is a benchmark that tests whether AI agents can optimize machine learning experiments by interpreting results and making informed decisions about hyperparameter changes.