GPT-5.6 Sol starts answering sooner, while Claude Opus 5 streams output faster and recorded fewer server errors across traffic routed through us. For API developers, that makes the initial wait and handling failures central to the choice. Sol’s lower headline prices also become a much smaller advantage after caching.
The short answer
Choose GPT-5.6 Sol for interactive API features where a short wait before the first token matters most. Choose Claude Opus 5 for automated coding jobs where fewer server failures and faster output streaming take priority. Across traffic routed through us, caching left only a small blended-cost difference.
At a glance
GPT-5.6 Sol offers slightly more room for long prompts, while Claude Opus 5 has computer-use support available through our endpoints. In our catalog, Sol has a 1.05-million-token context window versus Opus 5’s 1 million. Both cap output at 128,000 tokens, so Sol’s extra capacity accommodates more source material rather than a longer generated response.
At a glance
| Claude Opus 5 | GPT-5.6 Sol | |
|---|---|---|
| Trained by | Anthropic | OpenAI |
| Released | 2026-07-24 | 2026-07-09 |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Providers on Requesty | anthropic, vertex, bedrock | openai, openai-responses, bedrock, azure |
| Image input | Yes | Yes |
| Reasoning controls | Yes | Yes |
| Prompt caching | Yes | Yes |
| Tool calling | Yes | Yes |
| Computer use | Yes | No |
| Web search | Yes | Yes |
| Structured output | Yes | Yes |
| Open weights | No | No |
| Data retention (first-party endpoint) | 30 days | 30 days |
| Trains on your data | No | No |
| Share of Requesty tokens | 4.1% | 6.7% |
Pricing and caching
Customers paid similar blended rates after caching across traffic routed through us from July 6 to September 7, 2026: $1.80 per million tokens for Opus 5 and $1.71 for Sol. Opus’s higher cache savings narrowed the list-price gap in our routed cost measurements.
Our headline endpoints list Sol at $4 per million input tokens and $20 per million output tokens, versus Opus at $5 and $25. Sol’s promotional rate applies to requests with at most 272,000 input tokens and is available through at least November 21, 2026, according to the OpenAI changelog and eligibility reporting.
Pricing and caching
Our routed customer traffic · 2026-07-06 to 2026-09-07

▶Price breakdown and benchmark costs
▶Endpoint list prices
| Model | Endpoint | Input | Output | Cached input | Context |
|---|---|---|---|---|---|
| Claude Opus 5 | anthropic | $5.00 | $25.00 | $0.500 | 1M |
| Claude Opus 5 | vertex | $5.00 | $25.00 | $0.500 | 1M |
| Claude Opus 5 | bedrock | $5.00 | $25.00 | $0.500 | 1M |
| Claude Opus 5 | vertex (us) | $5.50 | $27.50 | $0.550 | 1M |
| Claude Opus 5 | vertex (eu) | $5.50 | $27.50 | $0.550 | 1M |
| Claude Opus 5 | bedrock (eu-west-3) | $5.50 | $27.50 | $0.550 | 1M |
| Claude Opus 5 | bedrock (eu-central-1) | $5.50 | $27.50 | $0.550 | 1M |
| Claude Opus 5 | bedrock (eu-north-1) | $5.50 | $27.50 | $0.550 | 1M |
| Claude Opus 5 | bedrock (eu-west-1) | $5.50 | $27.50 | $0.550 | 1M |
| GPT-5.6 Sol | openai | $2.50 | $15.00 | $0.250 | 1.1M |
| GPT-5.6 Sol | openai-responses | $2.50 | $15.00 | $0.250 | 1.1M |
| GPT-5.6 Sol | openai | $4.00 | $20.00 | $0.400 | 1.1M |
| GPT-5.6 Sol | bedrock (us-east-1) | $4.40 | $22.00 | $0.440 | 1.1M |
| GPT-5.6 Sol | bedrock (us-east-2) | $4.40 | $22.00 | $0.440 | 1.1M |
| GPT-5.6 Sol | openai-responses | $5.00 | $30.00 | $0.500 | 1.1M |
| GPT-5.6 Sol | azure (eastus2) | $5.00 | $30.00 | $0.500 | 1.1M |
| GPT-5.6 Sol | azure (eastus2) | $5.00 | $30.00 | $0.500 | 1.1M |
| GPT-5.6 Sol | azure (swedencentral) | $5.50 | $33.00 | $0.550 | 1.1M |
| GPT-5.6 Sol | azure (swedencentral) | $5.50 | $33.00 | $0.550 | 1.1M |
Benchmarks and use cases
The models are close on Artificial Analysis’s Coding Index, with a wider gap in Opus’s favor on the Intelligence and Agentic indices. The comparison uses Claude Opus 5 (Adaptive Reasoning, Max Effort) and GPT-5.6 Sol (max), with Intelligence Index version 4.3. Opus’s larger lead is in reasoning and agentic results rather than the coding index.
Benchmark scores
Artificial Analysis · source score units

Latency and throughput
Across traffic routed through us from July 6 to September 7, 2026, Sol had lower median and p95 time to first token. Opus generated output faster once streaming began, with higher median tokens per second. These describe different parts of the wait: when an answer starts appearing and how quickly the remaining text arrives.
Opus also recorded fewer server errors: a 0.1% 5xx rate versus Sol’s 6.8% in our routed traffic measurements. That is a substantial operational difference for jobs that must handle failed requests.
Latency and throughput
Our routed customer traffic · 2026-07-06 to 2026-09-07

▶Endpoint results
| Endpoint | First token | Output speed | Total |
|---|---|---|---|
| vertex (eu) | 1,808 ms | n/a | 7.8 s |
| bedrock (eu-central-1) | 1,809 ms | n/a | 7.4 s |
| bedrock (eu-west-3) | 1,911 ms | n/a | 8.5 s |
| bedrock (eu-west-1) | 2,161 ms | n/a | 10.1 s |
| bedrock (eu-north-1) | 2,337 ms | n/a | 10.8 s |
| bedrock | 2,342 ms | 86 tok/s | 4.9 s |
| vertex | 2,559 ms | n/a | 12.2 s |
| bedrock | 2,833 ms | n/a | 4.7 s |
| anthropic | 2,874 ms | 88 tok/s | 5.4 s |
| vertex (eu) | 3,882 ms | 79 tok/s | 6.6 s |
| vertex (us) | 4,128 ms | 165 tok/s | 5.5 s |
| vertex | 4,136 ms | 75 tok/s | 7.1 s |
| anthropic | 6,212 ms | n/a | 13.2 s |
| bedrock (eu-north-1) | 7,689 ms | 174 tok/s | 9.0 s |
| bedrock (eu-west-3) | 11,352 ms | 113 tok/s | 13.3 s |
| bedrock (eu-central-1) | 18,018 ms | n/a | 18.0 s |
| Endpoint | First token | Output speed | Total |
|---|---|---|---|
| bedrock (us-east-1) | 1,547 ms | n/a | 11.9 s |
| azure (swedencentral) | 1,688 ms | n/a | 9.3 s |
| openai | 1,941 ms | 78 tok/s | 4.8 s |
| azure (eastus2) | 3,745 ms | n/a | 17.3 s |
| openai | 7,496 ms | n/a | 11.2 s |
| openai | 11,548 ms | 79 tok/s | 14.3 s |
| azure (eastus2) | 18,919 ms | 86 tok/s | 21.5 s |
| openai-responses | 20,060 ms | 78 tok/s | 22.9 s |
| azure (swedencentral) | 22,357 ms | 70 tok/s | 25.5 s |
| azure (eastus2) | 23,167 ms | 75 tok/s | 26.0 s |
| openai-responses | 26,290 ms | 110 tok/s | 28.3 s |
| bedrock (us-east-2) | 28,333 ms | 119 tok/s | 30.2 s |
| bedrock (us-east-1) | 28,881 ms | 41 tok/s | 34.2 s |
Posts on X
Here are some posts on X comparing both models.
Independent reviews
- Across 50 identical one-shot prompts, Opus 5 and GPT-5.6 Sol had nearly equal average scores, with Sol tested at medium reasoning effort.Julian Goldie
- GPT-5.6 Sol caught more known code-review issues; Opus 5 at x-high effort produced the cleanest actionable comment stream but had the lowest coverage.coderabbit.ai
- In a same-prompt React Native habit-tracker build, final app quality was roughly tied; Opus 5 used more tokens because of simulator screenshot validation.youtube.com
How each model feels to use
Claude Opus 5
- Users report longer Opus 5 replies than prior Claude models, including explanations and steps they did not request.botmonster.comaiadopters.club
- Some developers report Opus 5 expanding coding tasks, changing designs or adding features beyond the requested scope.botmonster.com
- Some Claude Code users report repeating style preferences during long Opus 5 sessions despite including them in CLAUDE.md.modemguides.comlucadidomenico.studio
GPT-5.6 Sol
- Some practitioners using GPT-5.6 Sol in Codex describe it as literal and predictable, taking fewer unrequested liberties in daily coding.news.ycombinator.com
Recent changes
- As of August 28, 2026: Sol API pricing fell to $4 input and $20 output per million tokens for inputs up to 272,000 tokens, promotional through at least November 21, 2026.developers.openai.comtechnology.org
Which should you choose?
Choose GPT-5.6 Sol for daily coding and interactive features if you want quicker starts and more literal execution of your requests. Practitioners on Hacker News describe switching to Sol in Codex because it takes fewer unrequested liberties. If Claude’s wordiness makes routine exchanges harder to follow, that is a practical reason to try Sol.
Choose Claude Opus 5 for repository-level coding and novel reasoning, where DataCamp’s benchmark synthesis reports strengths. Keep task boundaries explicit: developer reports describe extra explanations and scope expansion that can make it more work to direct.
Our routed traffic measurements show Sol starts sooner, while Opus streams faster and recorded fewer 5xx errors. With coding index results close and blended rates similar after caching, prioritize the model whose working style suits the task.
Frequently asked questions
- Which model fits more source code or documents?
- GPT-5.6 Sol has a 1.05-million-token context window versus Opus 5’s 1 million in our catalog. Both cap generated output at 128,000 tokens.
- How do cached-input prices compare?
- The headline endpoints charge $0.40 per million cached-read tokens for Sol and $0.50 for Opus 5 in our catalog. Across our routed traffic, Opus’s higher cache savings narrowed the blended-cost gap.
- Which model should I choose for code review?
- Choose Sol when finding more known issues matters most. Choose Opus 5 at x-high effort when keeping review comments actionable matters more than coverage. CodeRabbit’s review found that trade-off in its tests.
- How much does Claude Opus 5 cost?
- The canonical endpoint lists Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, with $0.50 per million cached-read tokens in our catalog. Across traffic routed through us, customers paid a $1.80 blended rate per million tokens after cache savings.
- How much does GPT-5.6 Sol cost?
- The canonical endpoint lists GPT-5.6 Sol at $4 per million input tokens and $20 per million output tokens, with $0.40 per million cached-read tokens in our catalog. The promotional rate applies to requests with at most 272,000 input tokens and is available through at least November 21, 2026. In our routed traffic the blended rate was $1.71 per million tokens after caching.
- SEP '26
GPT-6 Astra is the best model you may not be allowed to call: gated access, the EU gap and safety stops
Astra launched 3 September at $10 and $50 per million tokens with the first Critical cybersecurity rating in OpenAI's framework. Three days later, most developers cannot call it, Foundry customers in the EU cannot keep it in region, and the ones who can call it are learning that a safety monitor can end a job mid-run. The benchmarks are one story. Getting it into production is another.
- SEP '26
The list price is now a range: peak hours, promo windows and host floors
In August the frontier rate card stopped being a number. DeepSeek bills V4 Pro at $1.32 input during seven weekday hours and $0.66 the rest of the time. OpenAI cut GPT-5.6 Sol to $4 and $20 but only through 21 November. Gemini 3.8 Flash doubles on 1 January. And the cheapest GPT-4-class price in the market is set by a reseller, not by the lab. Your cost model needs a clock, a calendar and a host column.
- SEP '26
GPT-6 Astra scores 61 on the independent index, the same as Sol, at 2.5x the price
OpenAI launched GPT-6 Astra on 3 September. Artificial Analysis put it at 61 on the intelligence index, identical to GPT-5.6 Sol, with Meta Muse Spark 1.3 ahead at 62, and coding at 67 against Sol's 65, while the API costs 2.5x more per token. The real gains are elsewhere, and they are the ones worth routing for.




