leaderboard

30 of 127 by gpqa diamond
Top 30 models ranked by gpqa diamond
#modelserved bygpqa diamondshare of leaderin / out per M
01gpt-6-astraOpenAI Inc.96.1%$10.00 / $50.00
02gemini-3.8-flash @euGoogle LLC (Vertex AI)95.3%$0.83 / $4.13
03grok-4.6xAI Corp.94.9%$2.00 / $6.00
04gpt-5.6-solOpenAI Inc.94.1%$4.00 / $20.00
05gemini-3.1-pro-previewGoogle LLC (Gemini API)94.1%$2.00 / $12.00
06claude-fable-5.1Anthropic PBC93.7%$10.00 / $50.00
07kimi-k3Moonshot AI93.5%$3.00 / $15.00
08gpt-5.5OpenAI Inc.93.5%$5.00 / $30.00
09qwen3.8-2.4t-a95bTensorX Ltd.93.5%$2.50 / $6.00
10claude-opus-5Anthropic PBC93.2%$5.00 / $25.00
11grok-4.5xAI Corp.93.1%$2.00 / $6.00
12minimax-m3MiniMax92.9%$0.30 / $1.20
13deepseek-v4-pro-0813DeepSeek92.8%$1.32 / $3.96
14gemini-3.6-flashGoogle LLC (Gemini API)92.8%$1.50 / $7.00
15qwen3.8-maxAlibaba Cloud92.7%$2.00 / $6.00
16claude-fable-5Anthropic PBC92.6%$10.00 / $50.00
17gpt-5.6-terraOpenAI Inc.92.5%$2.00 / $12.00
18qwen3.7-maxAlibaba Cloud92.3%$2.50 / $7.50
19qwen3.8-flash-nextTensorX Ltd.92.3%$0.20 / $0.50
20gemini-3.5-flashGoogle LLC (Gemini API)92.1%$1.50 / $9.00
21gemini-3.7-flashGoogle LLC (Vertex AI)92.1%$0.75 / $3.75
22claude-opus-4-8Anthropic PBC92.0%$5.00 / $25.00
23gpt-5.4OpenAI Inc.92.0%$2.50 / $15.00
24glm-5.3Z.ai91.7%$1.40 / $4.40
25gpt-5.3-codexOpenAI Inc.91.5%$1.75 / $14.00
26claude-opus-4-7Anthropic PBC91.4%$5.00 / $25.00
27glm-5.3-flashZ.ai91.2%$0.15 / $0.50
28kimi-k2.6Moonshot AI91.1%$0.95 / $4.00
29claude-sonnet-5Anthropic PBC91.1%$2.00 / $10.00
30gpt-5.6-lunaOpenAI Inc.91.1%$0.20 / $1.20

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

reasoning
gpqa diamond96.1%gpt-6-astra
entries12730 shown
served byOpenAI Inc.high first

gpt-6-astra leads gemini-3.8-flash by 0.8 points

distribution

30 entries
med 92.7%96.1%#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best for reasoning 30 ranked by gpqa diamond

GPQA Diamond is a set of graduate-level science questions written by domain experts and filtered so that PhD students with internet access still struggle. It's the most reliable signal we have for "does this model actually reason" vs "is it pattern-matching training data".

method

Scores for GPQA Diamond come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

127 entriesshowing 30sort high firstsource benchmarkleader gpt-6-astra 96.1%rebuilt daily