| # | model | served by | gpqa diamond | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | OpenAI Inc. | 96.1% | $10.00 / $50.00 | ||
| 02 | Google LLC (Vertex AI) | 95.3% | $0.83 / $4.13 | ||
| 03 | xAI Corp. | 94.9% | $2.00 / $6.00 | ||
| 04 | OpenAI Inc. | 94.1% | $4.00 / $20.00 | ||
| 05 | Google LLC (Gemini API) | 94.1% | $2.00 / $12.00 | ||
| 06 | Anthropic PBC | 93.7% | $10.00 / $50.00 | ||
| 07 | Moonshot AI | 93.5% | $3.00 / $15.00 | ||
| 08 | OpenAI Inc. | 93.5% | $5.00 / $30.00 | ||
| 09 | TensorX Ltd. | 93.5% | $2.50 / $6.00 | ||
| 10 | Anthropic PBC | 93.2% | $5.00 / $25.00 | ||
| 11 | xAI Corp. | 93.1% | $2.00 / $6.00 | ||
| 12 | MiniMax | 92.9% | $0.30 / $1.20 | ||
| 13 | DeepSeek | 92.8% | $1.32 / $3.96 | ||
| 14 | Google LLC (Gemini API) | 92.8% | $1.50 / $7.00 | ||
| 15 | Alibaba Cloud | 92.7% | $2.00 / $6.00 | ||
| 16 | Anthropic PBC | 92.6% | $10.00 / $50.00 | ||
| 17 | OpenAI Inc. | 92.5% | $2.00 / $12.00 | ||
| 18 | Alibaba Cloud | 92.3% | $2.50 / $7.50 | ||
| 19 | TensorX Ltd. | 92.3% | $0.20 / $0.50 | ||
| 20 | Google LLC (Gemini API) | 92.1% | $1.50 / $9.00 | ||
| 21 | Google LLC (Vertex AI) | 92.1% | $0.75 / $3.75 | ||
| 22 | Anthropic PBC | 92.0% | $5.00 / $25.00 | ||
| 23 | OpenAI Inc. | 92.0% | $2.50 / $15.00 | ||
| 24 | Z.ai | 91.7% | $1.40 / $4.40 | ||
| 25 | OpenAI Inc. | 91.5% | $1.75 / $14.00 | ||
| 26 | Anthropic PBC | 91.4% | $5.00 / $25.00 | ||
| 27 | Z.ai | 91.2% | $0.15 / $0.50 | ||
| 28 | Moonshot AI | 91.1% | $0.95 / $4.00 | ||
| 29 | Anthropic PBC | 91.1% | $2.00 / $10.00 | ||
| 30 | OpenAI Inc. | 91.1% | $0.20 / $1.20 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
gpt-6-astra leads gemini-3.8-flash by 0.8 points
One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best for reasoning 30 ranked by gpqa diamond
GPQA Diamond is a set of graduate-level science questions written by domain experts and filtered so that PhD students with internet access still struggle. It's the most reliable signal we have for "does this model actually reason" vs "is it pattern-matching training data".
method
Scores for GPQA Diamond come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
