# The State of Production AI

Model share, cost and speed measured from routed production traffic, 2026-06-01 to 2026-08-03.

- Source: https://www.requesty.ai/rankings
- Payload version: 1
- Generated at: 2026-08-03T17:37:39.521355982Z
- Models: 40 across 29 providers
- Blended cost per million tokens: $1.02
- Cache hit rate: 67.9%, removing 61.7% of the bill

## Every published number

| rank | model | vendor | product_line | token_share_pct | spend_share_pct | request_share_pct | token_share_delta_pp | blended_cost_per_mtok_usd | median_session_cost_usd | cost_per_request_usd | cache_hit_rate_pct | cache_savings_pct | median_ttft_ms | p95_ttft_ms | median_tokens_per_sec | tokens_per_request | reasoning_request_pct | stream_request_pct | calls_per_session | session_minutes | error_rate_5xx_pct |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | claude-sonnet-4-6 | anthropic | sonnet | 11.6 | 18.5 | 5.4 | 2.2 | 1.64 | 0.35 | 0.081289 | 62.7 | 57.3 | 1468 | 4179 | 77 | 49621 | 0.7 | 54.8 | 6 | 4.1 | 0.1 |
| 2 | deepseek-v4-flash | deepseek | deepseek | 11.3 | 0.4 | 5.5 | 1.4 | 0.03 | 0.03 | 0.001507 | 87 | 77.6 | 1200 | 7560 | 79 | 47592 | 59.1 | 48 | 4 | 0.9 | 0.5 |
| 3 | claude-opus-4-8 | anthropic | opus | 9.7 | 16.6 | 2.2 | 2 | 1.75 | 1.51 | 0.175132 | 81.6 | 76.4 | 1739 | 5826 | 87 | 100255 | 2.6 | 70.5 | 12 | 6.7 | 0.2 |
| 4 | deepseek-v4-pro | deepseek | deepseek | 6.8 | 0.8 | 2.5 | -3.8 | 0.12 | 0.04 | 0.00749 | 91.3 | 79.6 | 1000 | 3828 | 56 | 61578 | 75.5 | 62.2 | 5 | 19.5 | 0.4 |
| 5 | glm-5.2 | zai | glm | 6.2 | 3.2 | 1.9 | 0 | 0.53 | 0.39 | 0.039754 | 81.9 | 62 | 1936 | 16513 | 105 | 74464 | 35.5 | 83.2 | 13 | 3.5 | 1.1 |
| 6 | claude-sonnet-5 | anthropic | sonnet | 5.5 | 4.4 | 1.5 | 3.4 | 0.83 | 0.58 | 0.070305 | 76.9 | 71.4 | 1363 | 3975 | 99 | 85085 | 7.9 | 60.4 | 10 | 3.1 | 0.1 |
| 7 | gemini-3.5-flash | google | gemini-flash | 4.6 | 4.2 | 3.1 | 0 | 0.95 | 0.06 | 0.032283 | 57.5 | 48.4 | 2030 | 13751 | 489 | 34084 | 54 | 56.3 | 4 | 0.3 | 0.1 |
| 8 | gemini-2.5-flash | google | gemini-flash | 3.8 | 3 | 9.1 | 0.3 | 0.81 |  | 0.007763 | 4.8 | 1.2 | 1573 | 9418 | 3347 | 9549 | 71.5 | 1 | 3 |  | undefined |
| 9 | gpt-5.4 | openai | gpt-5 | 3.5 | 6.6 | 3.3 | 0.5 | 1.9 | 0.27 | 0.047443 | 53.5 | 38 | 989 | 57096 | 102 | 24970 | 9.5 | 56.6 | 7 | 0.7 | 0.2 |
| 10 | claude-haiku-4-5 | anthropic | haiku | 3.3 | 4.1 | 8 | -0.5 | 1.26 | 0.03 | 0.012001 | 36 | 20.7 | 661 | 1823 | 128 | 9507 | 0.1 | 14.6 | 4 | 0.3 | 0.2 |
| 11 | gpt-5.5 | openai | gpt-5 | 3.2 | 7.3 | 1.4 | 0.3 | 2.32 | 0.73 | 0.126336 | 77.6 | 62 | 11289 | 58684 | 215 | 54485 | 68.8 | 55.3 | 9 | 4.7 | 0.4 |
| 12 | gemini-2.5-flash-lite | google | gemini-flash-lite | 2.3 | 0.3 | 11 | -0.1 | 0.13 |  | 0.000641 | 2.1 | 1.3 | 274 | 1230 | 263 | 4894 | 3.2 | 0.8 |  |  | 0.1 |
| 13 | claude-opus-4-7 | anthropic | opus | 2.3 | 4.7 | 0.6 | -2 | 2.11 | 1.98 | 0.200699 | 73.9 | 66.8 | 2060 | 7309 | 108 | 94966 | 1.4 | 72.1 | 13 | 8.6 | 0.2 |
| 14 | deepseek-chat | deepseek | deepseek | 2 | 0.2 | 7.4 | -0.1 | 0.11 |  | 0.000679 | 40.7 | 27.3 | 568 | 1351 | 75 | 6125 | undefined | 0.5 |  |  | 0.1 |
| 15 | gemini-3.1-pro-preview | google | gemini-pro | 1.9 | 4.8 | 2.1 | -1 | 2.62 |  | 0.053455 | 18.1 | 11 | 9886 | 50527 | 118 | 20437 | 85.9 | 14 |  |  | 3.8 |
| 16 | gpt-5.6-luna | openai | gpt-5 | 1.8 | 0.6 | 1 | 0.1 | 0.36 | 0.15 | 0.015541 | 76.2 | 63.3 | 1677 | 6515 | 77 | 42601 | 69.6 | 47.9 | 16 | 2.4 | 0.4 |
| 17 | claude-opus-4-6 | anthropic | opus | 1.7 | 5.2 | 0.8 | -0.3 | 3.02 | 0.78 | 0.149222 | 65.7 | 62 | 2409 | 6349 | 67 | 49355 | 0.7 | 61.9 | 10 | 13.8 | 0.2 |
| 18 | minimax-m3 | minimaxi | minimax | 1.7 | 0.2 | 0.8 | -1.3 | 0.12 | 0.01 | 0.005997 | 91.9 | 66.7 | 1848 | 7580 | 90 | 49052 | 46.9 | 43.4 | 3 | 0.3 | 0.3 |
| 19 | gpt-5.6-sol | openai | gpt-5 | 1.5 | 1.9 | 0.3 | -2.5 | 1.26 | 1.02 | 0.146155 | 87.2 | 75.8 | 1290 | 6836 | 46 | 116184 | 68 | 84.8 | 9 | 4.4 | 0.3 |
| 20 | claude-opus-5 | anthropic | opus | 1.2 | 1.7 | 0.2 | 0.4 | 1.43 | 2.47 | 0.168123 | 88.9 | 80.6 | 1837 | 4601 | 78 | 117943 | 33.4 | 74 | 21 | 9 | 0.6 |
| 21 | gemini-3.1-flash-lite | google | gemini-flash-lite | 1.2 | 0.3 | 5.1 | -0.5 | 0.26 | 0.01 | 0.001457 | 23.4 | 18 | 943 | 3633 | 283 | 5505 | 1 | 3.7 | 5 | 2.7 | undefined |
| 22 | gpt-5.6-terra | openai | gpt-5 | 1.2 | 0.9 | 0.4 | 0.5 | 0.82 | 0.2 | 0.055343 | 81.9 | 68.9 | 1254 | 5371 | 84 | 67126 | 65.6 | 75.1 | 5 | 0.9 | 1.7 |
| 23 | kimi-k3 | moonshot | kimi | 1.1 | 0.8 | 0.4 | 0 | 0.75 | 0.63 | 0.05061 | 90.9 | 74.7 | 3524 | 12147 | 47 | 67607 | 67.1 | 64.3 | 15 | 6.8 | 0.3 |
| 24 | gpt-4o-mini | openai | gpt-mini | 0.9 | 0.1 | 9.1 | 0 | 0.13 |  | 0.00031 | 55.2 | 23.8 | 1145 | 4686 | 90 | 2406 | undefined | 0.4 |  |  | 0.6 |
| 25 | claude-fable-5 | anthropic | fable | 0.9 | 3.5 | 0.2 | 0.3 | 4.06 | 2.68 | 0.353596 | 74.9 | 67.7 | 4041 | 8841 | 80 | 87062 | 5.9 | 65.2 | 12 | 7.5 | 0.2 |
| 26 | kimi-k2.7-code | moonshot | kimi | 0.8 | 0.2 | 0.3 | 0.7 | 0.3 | 0.23 | 0.016767 | 94.2 | 70.7 | 1741 | 6093 | 92 | 55478 | 47.2 | 77.9 | 16 | 2.7 | 0.4 |
| 27 | gemini-3-flash-preview | google | gemini-flash | 0.6 | 0.4 | 1.4 | -0.1 | 0.64 | 0.11 | 0.006478 | 31.6 | 20.1 | 4601 | 18513 | 933 | 10103 | 61.1 | 12.6 | 9 | 8.9 | 0.1 |
| 28 | claude-sonnet-4-5 | anthropic | sonnet | 0.6 | 1.1 | 0.7 | 0.1 | 2.03 | 0.31 | 0.037046 | 53.6 | 49.2 | 1711 | 4186 | 64 | 18229 | 0.4 | 25 | 8 | 1.8 | 0.2 |
| 29 | gpt-5.4-mini | openai | gpt-mini | 0.5 | 0.3 | 1.4 | 0 | 0.63 |  | 0.004912 | 37.7 | 28.1 | 692 | 1869 | 141 | 7798 | 3.9 | 48 | 2 |  | 0.1 |
| 30 | gpt-5.1 | openai | gpt-5 | 0.4 | 0.5 | 1 | 0 | 1.2 | 0.03 | 0.012215 | 22.8 | 17.3 | 533 | 3264 | 118 | 10162 | 14.2 | 59.7 | 4 | 0.8 | 0.3 |
| 31 | minimax-m2.5 | minimaxi | minimax | 0.4 | 0.1 | 0.3 | 0.1 | 0.12 | 0.04 | 0.003738 | 75.5 | 60.6 | 795 | 4083 | 76 | 31467 | 0.1 | 75.6 | 28.5 | 5.3 | 0.9 |
| 32 | gpt-4.1-mini | openai | gpt-mini | 0.3 | 0.1 | 2.5 | -0.3 | 0.31 |  | 0.001024 | 30.5 | 28.3 | 505 | 2045 | 84 | 3279 | undefined | 1.9 |  |  | 1.4 |
| 33 | glm-5.1 | zai | glm | 0.3 | 0.3 | 0.2 | 0 | 0.82 | 0.18 | 0.035904 | 53.5 | 41.9 | 2070 | 12915 | 44 | 44017 | 9 | 86.7 | 9 | 2 | 6.2 |
| 34 | kimi-k2.6 | moonshot | kimi | 0.3 | 0.1 | 0.2 | -0.1 | 0.44 | 0.25 | 0.017674 | 74.4 | 53.5 | 1652 | 8629 | 48 | 40432 | 14.2 | 84.2 | 10 | 6 | 2.6 |
| 35 | grok-4.5 | xai | grok | 0.3 | 0.3 | 0.1 | -0.5 | 0.83 | 0.22 | 0.069957 | 94.3 | 70.5 | 1487 | 5197 | 75 | 84163 | 89.3 | 58.2 | 7.5 | 5.7 | 3 |
| 36 | hy3 | tencent | hunyuan | 0.3 | undefined | 0.1 | -0.3 | 0.03 |  | 0.00154 | 75.6 | 57 | 4155 | 16029 | 61 | 57527 | 60.4 | 71.7 | 2 | 23.8 | 0.4 |
| 37 | mimo-v2.5 | xiaomi | mimo | 0.3 | undefined | 0.1 | -0.1 | 0.02 |  | 0.001691 | 90.5 | 84.6 | 1384 | 5014 | 82 | 71476 | 88.2 | 90.7 |  |  | 1 |
| 38 | nemotron-3-ultra-550b-a55b | nvidia | nemotron | 0.3 | undefined | 0.1 | -0.2 |  |  | 0.000212 | 56.8 | undefined | 2533 | 22143 | 65 | 49025 | 3.1 | 67.2 | 2 | 1 | 9 |
| 39 | gpt-5-mini | openai | gpt-mini | 0.3 | 0.1 | 0.8 | 0.1 | 0.29 |  | 0.002151 | 59.1 | 29.3 | 3137 | 13796 | 106 | 7320 | 79.6 | 26.1 | 2 | 0.4 | 0.3 |
| 40 | deepseek-v4-flash-0731 | deepseek | deepseek | 0.2 | undefined | 0.2 | 4.6 | 0.07 | 0.08 | 0.001713 | 86.9 | 46.7 | 330 | 1401 | 133 | 24992 | 0.1 | 67.8 | 21 | 4.8 | 0.4 |

## Column notes

- `rank`: Position by token share over the window
- `model`: Model version as routed
- `vendor`: Model vendor
- `product_line`: Vendor product line
- `token_share_pct`: Share of all tokens processed over the window
- `spend_share_pct`: Share of all model spend over the window
- `request_share_pct`: Share of all requests over the window
- `token_share_delta_pp`: Week over week change in token share, percentage points
- `blended_cost_per_mtok_usd`: Charged price per million tokens, prompt cache discounts included
- `median_session_cost_usd`: Median spend on one session, trace-id clients only
- `cost_per_request_usd`: Mean spend on one request
- `cache_hit_rate_pct`: Prompt cache hit rate
- `cache_savings_pct`: Share of the bill removed by caching
- `median_ttft_ms`: Median time to first token, streaming responses only
- `p95_ttft_ms`: p95 time to first token, streaming responses only
- `median_tokens_per_sec`: Median output speed, streaming responses only
- `tokens_per_request`: Mean tokens per request
- `reasoning_request_pct`: Share of requests that returned reasoning tokens
- `stream_request_pct`: Share of requests streamed
- `calls_per_session`: Median calls in one session
- `session_minutes`: Median session length in minutes
- `error_rate_5xx_pct`: Server error rate

## How this is measured

- Every figure is a share or a per-unit rate. No request, token, session or spend totals are published, so nothing here sums back into platform volume or revenue.
- Each model version is its own row, because moving from one version to the next is the decision teams make. The product line is reported alongside for grouping.
- A model, provider and week is published only when enough separate organizations used it and no single one dominated it. 2.9% of traffic in this window failed those floors and is excluded, which is why shares do not sum to exactly 100.
- Time to first token and output speed come from streaming responses only, because a non streaming request has no first token event.
- A session is a run of calls sharing one trace id with no gap over 30 minutes. Only clients that send a trace id appear, and the figure is a median because session cost has a long tail.
- Cost per million tokens is what was charged with prompt cache discounts included, which is why it sits under provider list prices.
- An empty cell is a withheld or unsampled metric, not a zero.

## Citation

Requesty, "The State of Production AI", 2026-06-01 to 2026-08-03. https://www.requesty.ai/rankings
