A harder variant of MMLU with 10 answer choices instead of 4 and more reasoning-intensive questions, reducing noise from random guessing
12,032 questions across 14 domains, expanded from MMLU's 4-choice to 10-choice format. Questions are filtered for difficulty and augmented with reasoning-heavy problems from STEM sources. Significantly reduces guessing advantage (10% vs 25% random baseline).
Shows open-weight models only. Commercial API models (GPT-4o, Claude, Gemini) are not submitted to the Open LLM Leaderboard — their scores come from provider-reported benchmarks.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 | 92.4% |
| 2 | Claude Opus 5 | 91.6% |
| 3 | Claude Opus 5 | 91.6% |
| 4 | Claude Fable 5 | 91.5% |
| 5 | Gemini 3.7 Flash | 90.1% |
| 6 |
| Claude Opus 4.8 |
| 89.6% |
| 7 | Grok 4.6 | 89.4% |
| 8 | Qwen3.7 Max | 89.3% |
| 9 | Gemini 3.6 Flash | 89.3% |
| 10 | Grok 4.5 | 89.2% |
| 11 | GPT-5.6 Sol | 89.1% |
| 12 | Qwen3.8 Max | 88.6% |
| 13 | GPT-5.5 | 88.1% |
| 14 | Claude Sonnet 5 | 87.5% |
| 15 | GPT-5.6 Terra | 86.7% |
| 16 | DeepSeek V4 Flash Vision Exp | 86.2% |
| 17 | GPT-5.6 Luna | 86.0% |
| 18 | Qwen2.5 32B | 53.4% |
| 19 | Qwen2.5 32B Instruct | 51.9% |
| 20 | Qwen2.5 14B Instruct | 43.4% |
| 21 | Phi 3 medium 128k instruct | 41.2% |
| 22 | gemma 2 27b it | 38.3% |
| 23 | Qwen2.5 7B | 37.4% |
| 24 | Qwen2.5 7B Instruct | 36.5% |
| 25 | gemma 2 9b it | 31.9% |
| 26 | Llama 3.1 8B Instruct | 31.1% |
| 27 | Qwen2.5 Coder 7B Instruct | 26.1% |
| 28 | Qwen2.5 3B Instruct | 25.1% |
| 29 | Meta Llama 3 8B | 24.6% |
| 30 | Llama 3.2 3B Instruct | 24.4% |
| 31 | Qwen2.5 1.5B Instruct | 20.0% |
| 32 | Mistral 7B Instruct v0.2 | 19.1% |
| 33 | phi 2 | 18.1% |
| 34 | gemma 2 2b it | 17.2% |
| 35 | Qwen2 1.5B Instruct | 16.7% |
| 36 | Llama 3.2 3B | 16.5% |
| 37 | Qwen2.5 0.5B Instruct | 7.7% |
| 38 | Llama 3.2 1B Instruct | 7.6% |
| 39 | gpt j 6b | 2.7% |
| 40 | Llama 3.2 1B | 2.3% |
| 41 | distilgpt2 | 2.1% |
| 42 | gpt2 | 1.8% |
| 43 | falcon 7b instruct | 1.7% |
| 44 | gpt2 large | 1.6% |
| 45 | pythia 160m | 1.3% |
| 46 | TinyLlama 1.1B Chat v1.0 | 1.1% |