Best AI models for agentic coding
Terminal-Bench Hard measures how well a model operates as a coding agent in a real terminal, running commands, editing files, and fixing repositories end-to-end. It is the closest proxy to how models perform inside tools like Claude Code, Cursor and Codex.
- 馃
gpt-5.6-solOpenAI Inc.路$4.00 / $20.00 per 1M65.9%65.9% - 馃claude-fable-5Anthropic PBC路$10.00 / $50.00 per 1M62.9%62.9%
- 馃
gpt-5.5OpenAI Inc.路$5.00 / $30.00 per 1M60.6%60.6% - 4claude-opus-4-8Anthropic PBC路$5.00 / $25.00 per 1M58.3%58.3%
- 5
gpt-5.4OpenAI Inc.路$2.50 / $15.00 per 1M57.6%57.6% - 6
gpt-5.6-terraOpenAI Inc.路$2.00 / $12.00 per 1M57.6%57.6% - 7
gemini-3.1-pro-previewGoogle LLC (Gemini API)路$2.00 / $12.00 per 1M53.8%53.8% - 8claude-sonnet-4-6Anthropic PBC路$3.00 / $15.00 per 1M53.0%53.0%
- 9
gpt-5.3-codexOpenAI Responses路$1.75 / $14.00 per 1M53.0%53.0% - 10
gpt-5.4-miniOpenAI Inc.路$0.75 / $4.50 per 1M52.3%52.3% - 11claude-opus-4-7Anthropic PBC路$5.00 / $25.00 per 1M51.5%51.5%
- 12
glm-5.2Z.ai路$1.40 / $4.40 per 1M50.8%50.8% - 13
qwen3.7-maxAlibaba Cloud路$2.50 / $7.50 per 1M50.8%50.8% - 14claude-opus-4-5Anthropic PBC路$5.00 / $25.00 per 1M47.0%47.0%
- 15
gpt-5.2OpenAI Inc.路$1.75 / $14.00 per 1M47.0%47.0% - 16
qwen3.7-plus20% offAlibaba Cloud路$0.32 / $1.28 per 1M47.0%47.0% - 17claude-opus-4-6Anthropic PBC路$5.00 / $25.00 per 1M46.2%46.2%
- 18
deepseek-v4-pro-0424DeepInfra Inc.路$1.30 / $2.60 per 1M46.2%46.2% - 19
gpt-5.1OpenAI Inc.路$1.25 / $10.00 per 1M45.5%45.5% - 20
kimi-k2.7-codeMoonshot AI路$0.95 / $4.00 per 1M44.7%44.7% - 21
kimi-k2.6Moonshot AI路$0.95 / $4.00 per 1M43.9%43.9% - 22
qwen3.6-plusAlibaba Cloud路$0.50 / $3.00 per 1M43.9%43.9% - 23
glm-5.1Z.ai路$1.40 / $4.40 per 1M43.2%43.2% - 24
GLM-5Z.ai路$1.00 / $3.20 per 1M43.2%43.2% - 25
XiaomiMiMo/MiMo-V2.5-ProDeepInfra Inc.路$1.00 / $3.00 per 1M43.2%43.2% - 26minimax-m350% offMiniMax路$0.30 / $1.20 per 1M42.4%42.4%
- 27
gpt-5.4-nanoOpenAI Inc.路$0.20 / $1.25 per 1M42.4%42.4% - 28
gemini-3-pro-previewGoogle LLC (Gemini API)路$2.00 / $12.00 per 1M41.7%41.7% - 29
gemini-3.5-flashGoogle LLC (Gemini API)路$1.50 / $9.00 per 1M40.9%40.9% - 30
qwen/qwen3.5-397b-a17bNovita AI路$0.60 / $3.60 per 1M40.9%40.9%
Explore other rankings
How we rank
Scores for Terminal-Bench Hard come from Artificial Analysis, an independent AI benchmarking service. When a model is available through multiple providers (e.g. Anthropic direct, AWS Bedrock, Google Vertex), we show one canonical entry per model family so the ranking isn't polluted by duplicates. Benchmarks measure specific skills, so always validate on your own workload before committing.
One API for every model on this list
Requesty is OpenAI-compatible and routes to 600+ models. Switch between any of the models above by changing one parameter in your code.
