Tool and Agent Use Benchmark 2
Tests models on autonomous tool use and agentic task completion in realistic web and computer interaction scenarios
Models must complete multi-step tasks using tools (web search, code execution, API calls) in realistic scenarios. Evaluates planning, tool selection, error recovery, and goal completion across diverse domains.
Shows open-weight models only. Commercial API models (GPT-4o, Claude, Gemini) are not submitted to the Open LLM Leaderboard — their scores come from provider-reported benchmarks.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 | 98.5% |
| 2 | GPT-5.5 | 98.0% |
| 3 | Qwen3.7 Max | 94.7% |
| 4 | Claude Opus 4.8 | 94.4% |
| 5 | GPT-5.6 Terra | 86.3 |
| 6 | GPT-5.6 Sol | 85.1% |
| 7 | Claude Sonnet 4 | 60.0% |
| 8 | Claude Opus 4 | 59.6% |
| 9 | Claude 3.7 Sonnet | 58.4% |
| 10 | o1 | 50.0% |
| 11 | o4-mini | 49.2% |
| 12 | Claude 3.5 Sonnet | 46.0% |
| 13 | GPT-4o | 42.8% |
| 14 | o3 Mini | 32.4% |