leaderboard

30 of 57 by mmlu pro
Top 30 models ranked by mmlu pro
#modelserved bymmlu proshare of leaderin / out per M
01gemini-3-pro-previewGoogle LLC (Gemini API)89.8%$2.00 / $12.00
02claude-opus-4-5Anthropic PBC89.5%$5.00 / $25.00
03gemini-3-flash-previewGoogle LLC (Gemini API)89.0%$0.50 / $3.00
04claude-opus-4-1 @us-east5Google LLC (Vertex AI)88.0%$15.00 / $75.00
05claude-sonnet-4-5Anthropic PBC87.5%$3.00 / $15.00
06gpt-5.2OpenAI Inc.87.4%$1.75 / $14.00
07claude-opus-4 @us-east5Google LLC (Vertex AI)87.3%$15.00 / $75.00
08gpt-5OpenAI Inc.87.1%$1.25 / $10.00
09gpt-5.1OpenAI Inc.87.0%$1.25 / $10.00
10grok-4xAI Corp.86.6%$3.00 / $15.00
11gemini-2.5-proGoogle LLC (Gemini API)86.2%$1.25 / $10.00
12deepseek-v3.2Google LLC (Vertex AI)86.2%$0.56 / $1.68
13GLM-4.7Z.ai85.6%$0.60 / $2.20
14o3OpenAI Inc.85.3%$2.00 / $8.00
15grok-4-fastxAI Corp.85.0%$0.20 / $0.50
16deepseek-r1-turboNovita AI84.9%$0.70 / $2.50
17kimi-k2Google LLC (Vertex AI)84.8%$0.60 / $2.50
18Qwen/Qwen3-235B-A22B-Instruct-2507DeepInfra Inc.84.3%$0.07 / $0.10
19claude-sonnet-4 @us-east5Google LLC (Vertex AI)84.2%$3.00 / $15.00
20o1OpenAI Inc.84.1%$15.00 / $60.00
21deepseek-ai/DeepSeek-V3.1DeepInfra Inc.83.3%$0.30 / $1.00
22o4-miniOpenAI Inc.83.2%$1.10 / $4.40
23gemini-2.5-flashGoogle LLC (Gemini API)83.2%$0.30 / $2.50
24GLM-4.6Z.ai82.9%$0.60 / $2.20
25grok-3-minixAI Corp.82.8%$0.30 / $0.50
26gpt-5-miniOpenAI Inc.82.8%$0.25 / $2.00
27Qwen/Qwen3-235B-A22BDeepInfra Inc.82.8%$0.20 / $0.60
28minimax-m2MiniMax82.0%$0.30 / $1.20
29deepseek-v3-0324Novita AI81.9%$0.40 / $1.30
30zai-org/GLM-4.5-AirDeepInfra Inc.81.5%$0.20 / $1.10

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

knowledge
mmlu pro89.8%gemini-3-pro-preview
entries5730 shown
served byGoogle LLC (Gemini API)high first

gemini-3-pro-preview and claude-opus-4-5 are within a rounding error of each other, so pick on price or latency

distribution

30 entries
med 85.0%89.8%#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best for knowledge 30 ranked by mmlu pro

MMLU Pro tests broad knowledge across academic and professional subjects with harder, reasoning-heavy questions than the original MMLU. A strong score indicates wide, reliable factual coverage.

method

Scores for MMLU Pro come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

57 entriesshowing 30sort high firstsource benchmarkleader gemini-3-pro-preview 89.8%rebuilt daily