leaderboard

30 of 708 by context
Top 30 models ranked by context
#modelserved bycontextshare of leadermax out
01grok-4-1-fast-non-reasoningxAI Corp.2M-
02grok-4-1-fast-reasoningxAI Corp.2M-
03grok-4-fast-non-reasoningxAI Corp.2M-
04grok-4-fastxAI Corp.2M-
05grok-4.2-betaxAI Corp.2M-
06gpt-5.6-terra @us-west-2AWS Bedrock1.1M128K
07gpt-5.6-luna @us-east-2AWS Bedrock1.1M128K
08gpt-5.4 @us-east-2AWS Bedrock1.1M128K
09gpt-5.6-luna @us-east-1AWS Bedrock1.1M128K
10gpt-5.6-luna @us-west-2AWS Bedrock1.1M128K
11gpt-5.4 @us-west-2AWS Bedrock1.1M128K
12gpt-5.4 @us-east-1AWS Bedrock1.1M128K
13gpt-5.6-terra @us-east-1AWS Bedrock1.1M128K
14gpt-5.6-sol @us-east-1AWS Bedrock1.1M128K
15gpt-5.6-sol @us-east-2AWS Bedrock1.1M128K
16gpt-5.5 @us-east-1AWS Bedrock1.1M128K
17gpt-5.5 @us-east-2AWS Bedrock1.1M128K
18gpt-5.6-terra @us-east-2AWS Bedrock1.1M128K
19gpt-5.4 @germanywestcentralMicrosoft Azure AI1.1M128K
20gpt-6-astra @eastus2Microsoft Azure AI1.1M128K
21gpt-5.4Microsoft Azure AI1.1M128K
22openai-responses/gpt-5.6-luna @eastus2Microsoft Azure AI1.1M128K
23gpt-5.6-terra @germanywestcentralMicrosoft Azure AI1.1M128K
24gpt-5.4 @westeuropeMicrosoft Azure AI1.1M128K
25openai-responses/gpt-5.4 @francecentralMicrosoft Azure AI1.1M128K
26gpt-5.6-terra @swedencentralMicrosoft Azure AI1.1M128K
27openai-responses/gpt-5.4 @eastus2Microsoft Azure AI1.1M128K
28gpt-5.6-sol @francecentralMicrosoft Azure AI1.1M128K
29gpt-5.5 @westeuropeMicrosoft Azure AI1.1M128K
30gpt-5.6-luna @francecentralMicrosoft Azure AI1.1M128K

Bars are each context window's share of the largest in this list.

top of list

longest-context
context2Mgrok-4-1-fast-non-reasoning
entries70830 shown
served byxAI Corp.high first

grok-4-1-fast-non-reasoning and grok-4-1-fast-reasoning are effectively tied at the top of this list

distribution

30 entries
med 1.1M2M#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 5 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Longest context 30 ranked by context

A larger context window means more tokens you can fit in a single prompt, useful for whole-codebase analysis, long document Q&A, and agentic workflows. Note: effective quality often degrades past 128K tokens; prompt caching (supported on many models) is usually a better approach for repeated long context than brute-forcing more tokens in every call.

method

Ranked by maximum context window, the total tokens (input plus output) a model accepts in one request. Effective quality often degrades well below the advertised maximum, and most production workloads do better with prompt caching and retrieval than with more tokens per call.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

708 entriesshowing 30sort high firstsource contextleader grok-4-1-fast-non-reasoning 2Mrebuilt daily