Evaluates instruction-following ability using diverse, complex instructions that test a model's ability to precisely adhere to specified constraints
Tests models on following complex, multi-constraint instructions across diverse task types. Uses automatic evaluation with programmatic and LLM-based verification. More challenging than IFEval due to more complex and varied constraints.
Shows open-weight models only. Commercial API models (GPT-4o, Claude, Gemini) are not submitted to the Open LLM Leaderboard — their scores come from provider-reported benchmarks.
| # | Model | Score |
|---|---|---|
| 1 | Qwen3.8 Max | 82.8% |
| 2 | Grok 4.3 | 81.3% |
| 3 | Qwen3.7 Max | 79.1% |
| 4 | GPT-5.5 | 75.9% |
| 5 | GPT-5.6 Sol | 72.7 |
| 6 | GPT-5.6 Terra | 71.2% |
| 7 | o3 | 69.3% |
| 8 | Claude Fable 5 | 63.5% |
| 9 | Claude Opus 4.8 | 62.2% |
| 10 | Claude Opus 4.5 | 58.0% |
| 11 | Gemini 2.5 Pro | 52.3% |
| 12 | Claude Sonnet 4 | 42.3% |
| 13 | DeepSeek R1 | 38.0% |