GPT-4o is OpenAI's multimodal workhorse — comfortable switching between analyzing images and generating text without missing a beat. It handles a wide range of tasks with consistent reliability, from parsing complex documents to interpreting visual inputs, all within a generous 128k-token context window. It tends to be thorough and articulate, though like most large models, it can occasionally over-explain or hedge.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| Chatbot Arena | 1285.0 | Bradley-Terry Elo | 1y ago |
| AIME 2024 | 12.0 | accuracy | 4mo ago |
| SciCode | 1.5 | main_problem_pass@1 | 4mo ago |
| Humanity's Last Exam | 9.4 | 0-shot | 1y ago |
| MMLU | 88.7 | 5-shot | 2y ago |
| HumanEval | 90.2 | pass@1 | 2y ago |
| MT-Bench | 9.3 | GPT-4-judge | 2y ago |
| GPQA Diamond | 53.6 | 0-shot-CoT | 2y ago |
| MATH | 76.6 | 4-shot | 2y ago |
| SWE-Bench | 33.2 | accuracy | 4mo ago |
| HumanEval+ | 86.6 | pass@1 | 2y ago |
| Aider Polyglot | 23.1 | accuracy | 4mo ago |
| TAU2 | 42.8 | accuracy | 4mo ago |