| # | model | served by | context | share of leader | max out |
|---|---|---|---|---|---|
| 01 | xAI Corp. | 2M | - | ||
| 02 | xAI Corp. | 2M | - | ||
| 03 | xAI Corp. | 2M | - | ||
| 04 | xAI Corp. | 2M | - | ||
| 05 | xAI Corp. | 2M | - | ||
| 06 | AWS Bedrock | 1.1M | 128K | ||
| 07 | AWS Bedrock | 1.1M | 128K | ||
| 08 | AWS Bedrock | 1.1M | 128K | ||
| 09 | AWS Bedrock | 1.1M | 128K | ||
| 10 | AWS Bedrock | 1.1M | 128K | ||
| 11 | AWS Bedrock | 1.1M | 128K | ||
| 12 | AWS Bedrock | 1.1M | 128K | ||
| 13 | AWS Bedrock | 1.1M | 128K | ||
| 14 | AWS Bedrock | 1.1M | 128K | ||
| 15 | AWS Bedrock | 1.1M | 128K | ||
| 16 | AWS Bedrock | 1.1M | 128K | ||
| 17 | AWS Bedrock | 1.1M | 128K | ||
| 18 | AWS Bedrock | 1.1M | 128K | ||
| 19 | Microsoft Azure AI | 1.1M | 128K | ||
| 20 | Microsoft Azure AI | 1.1M | 128K | ||
| 21 | Microsoft Azure AI | 1.1M | 128K | ||
| 22 | Microsoft Azure AI | 1.1M | 128K | ||
| 23 | Microsoft Azure AI | 1.1M | 128K | ||
| 24 | Microsoft Azure AI | 1.1M | 128K | ||
| 25 | Microsoft Azure AI | 1.1M | 128K | ||
| 26 | Microsoft Azure AI | 1.1M | 128K | ||
| 27 | Microsoft Azure AI | 1.1M | 128K | ||
| 28 | Microsoft Azure AI | 1.1M | 128K | ||
| 29 | Microsoft Azure AI | 1.1M | 128K | ||
| 30 | Microsoft Azure AI | 1.1M | 128K |
Bars are each context window's share of the largest in this list.
grok-4-1-fast-non-reasoning and grok-4-1-fast-reasoning are effectively tied at the top of this list
One bar per ranked model, best on the left, drawn as a share of the leader. 5 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Longest context 30 ranked by context
A larger context window means more tokens you can fit in a single prompt, useful for whole-codebase analysis, long document Q&A, and agentic workflows. Note: effective quality often degrades past 128K tokens; prompt caching (supported on many models) is usually a better approach for repeated long context than brute-forcing more tokens in every call.
method
Ranked by maximum context window, the total tokens (input plus output) a model accepts in one request. Effective quality often degrades well below the advertised maximum, and most production workloads do better with prompt caching and retrieval than with more tokens per call.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
