LLM agents struggle to convert their general capabilities into cost-efficient task-specific solutions, but when they succeed, the savings are dramatic—suggesting bottling is a valuable but underdeveloped capability worth improving.
This paper introduces BOTTLED, a benchmark testing whether LLM agents can autonomously create cheaper, task-specific solutions from their general capabilities. Agents receive unlabeled workloads with fixed budgets and must decide their own approach—like training small models or writing programs.