GPT-6.1 Sol is OpenAI's recommended cost-performance model: same $2/$10 pricing as GPT-6 Sol (cached input halved to $0.10), but much stronger, scoring 75.2% on DeepSWE v1.1 versus 68.8% for GPT-6 Sol and 71.4% on OSWorld 2.0 against Astra's 73.5%. It supports reasoning effort from low to max and tool calling through the Responses and Chat Completions APIs. Input is text and images (no audio or video) with a 1.05M-token context window and up to 128K output tokens.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| TerminalBench | 56.1 | accuracy | today |
| SciCode | 54.2 | accuracy | today |
| LCR | 83.0 | accuracy | today |
| Humanity's Last Exam | 52.9 | accuracy | today |
Updated today
Predecessors
GPT-6 Sol