Most capable model in Anthropic's lineup with always-on thinking. Excels at complex reasoning, long-horizon agentic tasks, creative writing, and deep analysis.
State-of-the-art software engineering
Most capable reasoning with always-on thinking
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| LiveCodeBench | 89.8 | accuracy | 23d ago |
| IFBench | 63.5 | accuracy | 23d ago |
| MMLU-Pro | 91.5 | accuracy | 23d ago |
| SWE-Bench | 95.0 | accuracy | 23d ago |
| LCR | 82.3 | accuracy | 23d ago |
| Humanity's Last Exam | 55.5 | accuracy | 23d ago |
| TAU2 | 98.5 | accuracy | 23d ago |
| GPQA Diamond | 92.6 | accuracy |
Best-in-class instruction following
Advanced tool use for agentic workflows
Exceptional creative writing
1M token context window
Excellent multilingual capabilities
Strong multimodal vision understanding
Deep factual knowledge
| 23d ago |
| SciCode | 61.0 | accuracy | 23d ago |