Continuously updated coding benchmark using new competitive programming problems from LeetCode, AtCoder, and Codeforces to prevent contamination
Collects new competitive programming problems published after training cutoff dates of evaluated models. Problems include code generation, self-repair, code execution prediction, and test output prediction. Automatically refreshed to avoid benchmark contamination.
Shows open-weight models only. Commercial API models (GPT-4o, Claude, Gemini) are not submitted to the Open LLM Leaderboard — their scores come from provider-reported benchmarks.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 | 90.5% |
| 2 | Claude Fable 5 | 89.8% |
| 3 | Claude Opus 5 | 89.0% |
| 4 | Claude Opus 5 | 89.0% |
| 5 | Gemini 3.7 Flash |
| 88.7% |
| 6 | Grok 4.6 | 88.2% |
| 7 | Gemini 3.6 Flash | 88.1% |
| 8 | Qwen3.8 Max | 87.9% |
| 9 | Claude Opus 4.8 | 87.8% |
| 10 | Grok 4.5 | 87.4% |
| 11 | DeepSeek V4 Flash Vision Exp | 87.3% |
| 12 | Qwen3.7 Max | 87.1% |
| 13 | GPT-5.6 Terra | 85.9% |
| 14 | GPT-5.5 | 85.3% |
| 15 | GPT-5.6 Sol | 82.6% |
| 16 | Claude Sonnet 5 | 82.4% |
| 17 | Qwen3 235B A22B | 70.7% |
| 18 | Gemini 2.5 Pro | 70.4% |
| 19 | DeepSeek R1 | 65.9% |