R1 thinks out loud — it works through problems step by step, showing its reasoning chain before arriving at an answer. This makes it particularly transparent on math, logic, and coding tasks, where you can follow (and verify) its work. The trade-off is verbosity: responses are often long, and the model can over-deliberate on simple questions.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| Aider Polyglot | 56.9 | accuracy | 4mo ago |
| AIME 2024 | 79.8 | accuracy | 4mo ago |
| LiveCodeBench | 65.9 | accuracy | 4mo ago |
| SciCode | 4.6 | main_problem_pass@1 | 4mo ago |
| IFBench | 38.0 | prompt_level_loose_accuracy | 4mo ago |
| SWE-Bench | 49.2 | accuracy | 4mo ago |
| AIME 2025 | 87.5 | accuracy | 4mo ago |