Claude Opus 4.8 is the latest and most capable entry in Anthropic's Opus line, refining what was already a strong foundation. It pushes further on code generation quality, agentic task execution, and calibrated honesty — areas where 4.7 was already competitive. The 1M token context window gives it room for large codebases and lengthy documents. The trade-off remains the same: this is Anthropic's heaviest model, so expect higher latency and cost compared to Sonnet or Haiku.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| Humanity's Last Exam | 49.8 | accuracy | 23d ago |
| GPQA Diamond | 93.6 | accuracy | 23d ago |
| LiveCodeBench | 87.8 | accuracy | 23d ago |
| LCR | 77.7 | accuracy | 23d ago |
| IFBench | 62.2 | accuracy | 23d ago |
| TAU2 | 94.4 | accuracy | 23d ago |
| SWE-Bench | 88.6 | accuracy | 23d ago |
| SciCode | 54.4 | accuracy | 23d ago |
| MMLU-Pro | 89.6 | accuracy | 23d ago |