leaderboard

30 of 49 by math index
Top 30 models ranked by math index
#modelserved bymath indexshare of leaderin / out per M
01gpt-5.2OpenAI Inc.99.0$1.75 / $14.00
02gemini-3-flash-previewGoogle LLC (Gemini API)97.0$0.50 / $3.00
03gemini-3-pro-previewGoogle LLC (Gemini API)95.7$2.00 / $12.00
04GLM-4.7Z.ai95.0$0.60 / $2.20
05kimi-k2Google LLC (Vertex AI)94.7$0.60 / $2.50
06gpt-5OpenAI Inc.94.3$1.25 / $10.00
07gpt-5.1OpenAI Inc.94.0$1.25 / $10.00
08gpt-oss-120bGroq Inc.93.4$0.15 / $0.75
09grok-4xAI Corp.92.7$3.00 / $15.00
10deepseek-v3.2Google LLC (Vertex AI)92.0$0.56 / $1.68
11claude-opus-4-5Anthropic PBC91.3$5.00 / $25.00
12nemotron-3-nano-30b-a3bNVIDIA91.0free / free
13Qwen/Qwen3-235B-A22B-Instruct-2507DeepInfra Inc.91.0$0.07 / $0.10
14o4-miniOpenAI Inc.90.7$1.10 / $4.40
15grok-4-fastxAI Corp.89.7$0.20 / $0.50
16o3OpenAI Inc.88.3$2.00 / $8.00
17claude-sonnet-4-5Anthropic PBC88.0$3.00 / $15.00
18gemini-2.5-proGoogle LLC (Gemini API)87.7$1.25 / $10.00
19GLM-4.6Z.ai86.0$0.60 / $2.20
20gpt-5-miniOpenAI Inc.85.0$0.25 / $2.00
21grok-3-minixAI Corp.84.7$0.30 / $0.50
22claude-haiku-4-5Anthropic PBC83.7$1.00 / $5.00
23gpt-5-nanoOpenAI Inc.83.7$0.05 / $0.40
24Qwen/Qwen3-235B-A22BDeepInfra Inc.82.0$0.20 / $0.60
25zai-org/GLM-4.5-AirDeepInfra Inc.80.7$0.20 / $1.10
26claude-opus-4-1 @us-east5Google LLC (Vertex AI)80.3$15.00 / $75.00
27minimax-m2MiniMax78.3$0.30 / $1.20
28deepseek-r1-turboNovita AI76.0$0.70 / $2.50
29claude-sonnet-4 @us-east5Google LLC (Vertex AI)74.3$3.00 / $15.00
30gemini-2.5-flashGoogle LLC (Gemini API)73.3$0.30 / $2.50

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

math
math index99.0gpt-5.2
entries4930 shown
served byOpenAI Inc.high first

gpt-5.2 leads gemini-3-flash-preview by 2.0 points

distribution

30 entries
med 89.099.0#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 15 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best at math 30 ranked by math index

The Math Index aggregates competition and advanced math evaluations (including AIME). These problems require real symbolic reasoning across multiple novel steps, memorization gets a model nowhere.

method

Scores for Math Index come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

49 entriesshowing 30sort high firstsource benchmarkleader gpt-5.2 99.0rebuilt daily