GPT-5.5 is OpenAI's most capable general-purpose model, designed for tasks requiring sustained multi-step reasoning, complex tool orchestration, and deep coding. It handles a 1M+ token context window, making it suitable for large-scale document analysis and code review. The model excels at agentic workflows where it needs to plan, execute, and iterate across multiple tool calls. Cost and latency are at the high end of OpenAI's lineup — reach for GPT-5.5 Mini or Nano for lighter tasks.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| Humanity's Last Exam | 41.4 | accuracy | 23d ago |
| LiveCodeBench | 85.3 | accuracy | 23d ago |
| LCR | 84.3 | accuracy | 23d ago |
| TAU2 | 98.0 | accuracy | 23d ago |
| GPQA Diamond | 93.6 | accuracy | 23d ago |
| MMLU-Pro | 88.1 | accuracy | 23d ago |
| IFBench | 75.9 | accuracy | 23d ago |
| SciCode | 55.8 | accuracy | 23d ago |