Best AI models for reasoning
GPQA Diamond is a set of graduate-level science questions written by domain experts and filtered so that PhD students with internet access still struggle. It's the most reliable signal we have for "does this model actually reason" vs "is it pattern-matching training data".
- 馃
gpt-6-astraOpenAI Inc.路$10.00 / $50.00 per 1M96.1%96.1% - 馃
gemini-3.8-flash50% offGoogle LLC (Vertex AI)路$0.75 / $3.75 per 1M95.3%95.3% - 馃grok-4.6xAI Corp.路$2.00 / $6.00 per 1M94.9%94.9%
- 4
gemini-3.1-pro-previewGoogle LLC (Gemini API)路$2.00 / $12.00 per 1M94.1%94.1% - 5
gpt-5.6-solOpenAI Inc.路$4.00 / $20.00 per 1M94.1%94.1% - 6claude-fable-5.1Anthropic PBC路$10.00 / $50.00 per 1M93.7%93.7%
- 7
kimi-k3Moonshot AI路$3.00 / $15.00 per 1M93.5%93.5% - 8
gpt-5.5OpenAI Inc.路$5.00 / $30.00 per 1M93.5%93.5% - 9qwen3.8-2.4t-a95bTensorX Ltd.路$2.50 / $6.00 per 1M93.5%93.5%
- 10claude-opus-5Anthropic PBC路$5.00 / $25.00 per 1M93.2%93.2%
- 11grok-4.5xAI Corp.路$2.00 / $6.00 per 1M93.1%93.1%
- 12minimax-m350% offMiniMax路$0.30 / $1.20 per 1M92.9%92.9%
- 13
deepseek-v4-pro-0813DeepSeek路$1.32 / $3.96 per 1M92.8%92.8% - 14
gemini-3.6-flashGoogle LLC (Gemini API)路$1.50 / $7.00 per 1M92.8%92.8% - 15
qwen3.8-maxAlibaba Cloud路$2.00 / $6.00 per 1M92.7%92.7% - 16claude-fable-5Anthropic PBC路$10.00 / $50.00 per 1M92.6%92.6%
- 17
gpt-5.6-terraOpenAI Inc.路$2.00 / $12.00 per 1M92.5%92.5% - 18
qwen3.7-maxAlibaba Cloud路$2.50 / $7.50 per 1M92.3%92.3% - 19qwen3.8-flash-nextTensorX Ltd.路$0.20 / $0.50 per 1M92.3%92.3%
- 20
gemini-3.5-flashGoogle LLC (Gemini API)路$1.50 / $9.00 per 1M92.1%92.1% - 21
gemini-3.7-flash@eu50% offGoogle LLC (Vertex AI)路$0.83 / $4.13 per 1M92.1%92.1% - 22claude-opus-4-8Anthropic PBC路$5.00 / $25.00 per 1M92.0%92.0%
- 23
gpt-5.4OpenAI Inc.路$2.50 / $15.00 per 1M92.0%92.0% - 24
glm-5.3Z.ai路$1.40 / $4.40 per 1M91.7%91.7% - 25
gpt-5.3-codexOpenAI Inc.路$1.75 / $14.00 per 1M91.5%91.5% - 26claude-opus-4-7Anthropic PBC路$5.00 / $25.00 per 1M91.4%91.4%
- 27
glm-5.3-flashZ.ai路$0.15 / $0.50 per 1M91.2%91.2% - 28
kimi-k2.6Moonshot AI路$0.95 / $4.00 per 1M91.1%91.1% - 29claude-sonnet-5Anthropic PBC路$2.00 / $10.00 per 1M91.1%91.1%
- 30
gpt-5.6-lunaOpenAI Inc.路$0.20 / $1.20 per 1M91.1%91.1%
Explore other rankings
How we rank
Scores for GPQA Diamond come from Artificial Analysis, an independent AI benchmarking service. When a model is available through multiple providers (e.g. Anthropic direct, AWS Bedrock, Google Vertex), we show one canonical entry per model family so the ranking isn't polluted by duplicates. Benchmarks measure specific skills, so always validate on your own workload before committing.
One API for every model on this list
Requesty is OpenAI-compatible and routes to 600+ models. Switch between any of the models above by changing one parameter in your code.
