Long Context Retrieval
Tests models on retrieving specific information from very long documents, measuring long-context comprehension and retrieval accuracy
Models must locate and extract specific facts, figures, or passages from long documents (10K–1M tokens). Tests robustness of attention mechanisms and context utilisation at extended lengths.
Shows open-weight models only. Commercial API models (GPT-4o, Claude, Gemini) are not submitted to the Open LLM Leaderboard — their scores come from provider-reported benchmarks.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 | 85.3% |
| 2 | Claude Opus 5.5 | 84.7% |
| 3 | GPT-5.5 | 84.3% |
| 4 | GPT-5.6 Sol | 84.0% |
| 5 | GPT-5.6 Luna | 83.7% |
| 6 | GPT-6 Sol | 83.7% |
| 7 | GPT-6 Luna |
| 83.3% |
| 8 | GPT-5.6 Terra | 83.0% |
| 9 | GPT-6.1 Sol | 83.0% |
| 10 | Claude Sonnet 5.5 | 82.7% |
| 11 | Claude Fable 5 | 82.3% |
| 12 | Claude Sonnet 5 | 82.0% |
| 13 | Gemini 3.7 Flash | 81.7% |
| 14 | GPT-6 Astra | 80.7% |
| 15 | Grok 4.6 | 80.3% |
| 16 | Gemini 3.6 Flash | 80.0% |
| 17 | Gemini 4 Argon | 79.7% |
| 18 | DeepSeek V4 Flash Vision Exp | 79.7% |
| 19 | Claude Opus 5 | 79.3% |
| 20 | Claude Opus 5 | 79.3% |
| 21 | Grok 4.5 | 79.3% |
| 22 | Qwen3.7 Max | 79.0% |
| 23 | Claude Opus 4.8 | 77.7% |
| 24 | Grok 4.7 | 76.7% |
| 25 | Claude Opus 4.7 | 70.3% |
| 26 | o3 | 69.3% |
| 27 | Grok 4 | 68.0% |
| 28 | Gemini 2.5 Pro | 66.0% |
| 29 | Claude Sonnet 4.5 | 66.0% |
| 30 | Claude Opus 4.5 | 65.3% |
| 31 | Claude Sonnet 4 | 65.0% |