| # | model | served by | math index | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | OpenAI Inc. | 99.0 | $1.75 / $14.00 | ||
| 02 | Google LLC (Gemini API) | 97.0 | $0.50 / $3.00 | ||
| 03 | Google LLC (Gemini API) | 95.7 | $2.00 / $12.00 | ||
| 04 | Z.ai | 95.0 | $0.60 / $2.20 | ||
| 05 | Google LLC (Vertex AI) | 94.7 | $0.60 / $2.50 | ||
| 06 | OpenAI Inc. | 94.3 | $1.25 / $10.00 | ||
| 07 | OpenAI Inc. | 94.0 | $1.25 / $10.00 | ||
| 08 | Groq Inc. | 93.4 | $0.15 / $0.75 | ||
| 09 | xAI Corp. | 92.7 | $3.00 / $15.00 | ||
| 10 | Google LLC (Vertex AI) | 92.0 | $0.56 / $1.68 | ||
| 11 | Anthropic PBC | 91.3 | $5.00 / $25.00 | ||
| 12 | NVIDIA | 91.0 | free / free | ||
| 13 | DeepInfra Inc. | 91.0 | $0.07 / $0.10 | ||
| 14 | OpenAI Inc. | 90.7 | $1.10 / $4.40 | ||
| 15 | xAI Corp. | 89.7 | $0.20 / $0.50 | ||
| 16 | OpenAI Inc. | 88.3 | $2.00 / $8.00 | ||
| 17 | Anthropic PBC | 88.0 | $3.00 / $15.00 | ||
| 18 | Google LLC (Gemini API) | 87.7 | $1.25 / $10.00 | ||
| 19 | Z.ai | 86.0 | $0.60 / $2.20 | ||
| 20 | OpenAI Inc. | 85.0 | $0.25 / $2.00 | ||
| 21 | xAI Corp. | 84.7 | $0.30 / $0.50 | ||
| 22 | Anthropic PBC | 83.7 | $1.00 / $5.00 | ||
| 23 | OpenAI Inc. | 83.7 | $0.05 / $0.40 | ||
| 24 | DeepInfra Inc. | 82.0 | $0.20 / $0.60 | ||
| 25 | DeepInfra Inc. | 80.7 | $0.20 / $1.10 | ||
| 26 | Google LLC (Vertex AI) | 80.3 | $15.00 / $75.00 | ||
| 27 | MiniMax | 78.3 | $0.30 / $1.20 | ||
| 28 | Novita AI | 76.0 | $0.70 / $2.50 | ||
| 29 | Google LLC (Vertex AI) | 74.3 | $3.00 / $15.00 | ||
| 30 | Google LLC (Gemini API) | 73.3 | $0.30 / $2.50 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
gpt-5.2 leads gemini-3-flash-preview by 2.0 points
One bar per ranked model, best on the left, drawn as a share of the leader. 15 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best at math 30 ranked by math index
The Math Index aggregates competition and advanced math evaluations (including AIME). These problems require real symbolic reasoning across multiple novel steps, memorization gets a model nowhere.
method
Scores for Math Index come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
