| # | model | served by | mmlu pro | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | Google LLC (Gemini API) | 89.8% | $2.00 / $12.00 | ||
| 02 | Anthropic PBC | 89.5% | $5.00 / $25.00 | ||
| 03 | Google LLC (Gemini API) | 89.0% | $0.50 / $3.00 | ||
| 04 | Google LLC (Vertex AI) | 88.0% | $15.00 / $75.00 | ||
| 05 | Anthropic PBC | 87.5% | $3.00 / $15.00 | ||
| 06 | OpenAI Inc. | 87.4% | $1.75 / $14.00 | ||
| 07 | Google LLC (Vertex AI) | 87.3% | $15.00 / $75.00 | ||
| 08 | OpenAI Inc. | 87.1% | $1.25 / $10.00 | ||
| 09 | OpenAI Inc. | 87.0% | $1.25 / $10.00 | ||
| 10 | xAI Corp. | 86.6% | $3.00 / $15.00 | ||
| 11 | Google LLC (Gemini API) | 86.2% | $1.25 / $10.00 | ||
| 12 | Google LLC (Vertex AI) | 86.2% | $0.56 / $1.68 | ||
| 13 | Z.ai | 85.6% | $0.60 / $2.20 | ||
| 14 | OpenAI Inc. | 85.3% | $2.00 / $8.00 | ||
| 15 | xAI Corp. | 85.0% | $0.20 / $0.50 | ||
| 16 | Novita AI | 84.9% | $0.70 / $2.50 | ||
| 17 | Google LLC (Vertex AI) | 84.8% | $0.60 / $2.50 | ||
| 18 | DeepInfra Inc. | 84.3% | $0.07 / $0.10 | ||
| 19 | Google LLC (Vertex AI) | 84.2% | $3.00 / $15.00 | ||
| 20 | OpenAI Inc. | 84.1% | $15.00 / $60.00 | ||
| 21 | DeepInfra Inc. | 83.3% | $0.30 / $1.00 | ||
| 22 | OpenAI Inc. | 83.2% | $1.10 / $4.40 | ||
| 23 | Google LLC (Gemini API) | 83.2% | $0.30 / $2.50 | ||
| 24 | Z.ai | 82.9% | $0.60 / $2.20 | ||
| 25 | xAI Corp. | 82.8% | $0.30 / $0.50 | ||
| 26 | OpenAI Inc. | 82.8% | $0.25 / $2.00 | ||
| 27 | DeepInfra Inc. | 82.8% | $0.20 / $0.60 | ||
| 28 | MiniMax | 82.0% | $0.30 / $1.20 | ||
| 29 | Novita AI | 81.9% | $0.40 / $1.30 | ||
| 30 | DeepInfra Inc. | 81.5% | $0.20 / $1.10 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
gemini-3-pro-preview and claude-opus-4-5 are within a rounding error of each other, so pick on price or latency
One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best for knowledge 30 ranked by mmlu pro
MMLU Pro tests broad knowledge across academic and professional subjects with harder, reasoning-heavy questions than the original MMLU. A strong score indicates wide, reliable factual coverage.
method
Scores for MMLU Pro come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
