leaderboard

30 of 95 by terminal-bench hard
Top 30 models ranked by terminal-bench hard
#modelserved byterminal-bench hardshare of leaderin / out per M
01gpt-5.6-solOpenAI Inc.65.9%$4.00 / $20.00
02claude-fable-5Anthropic PBC62.9%$10.00 / $50.00
03gpt-5.5OpenAI Inc.60.6%$5.00 / $30.00
04claude-opus-4-8Anthropic PBC58.3%$5.00 / $25.00
05gpt-5.6-terraOpenAI Inc.57.6%$2.00 / $12.00
06gpt-5.4OpenAI Inc.57.6%$2.50 / $15.00
07gemini-3.1-pro-previewGoogle LLC (Gemini API)53.8%$2.00 / $12.00
08claude-sonnet-4-6Anthropic PBC53.0%$3.00 / $15.00
09gpt-5.3-codexOpenAI Inc.53.0%$1.75 / $14.00
10gpt-5.4-miniOpenAI Inc.52.3%$0.75 / $4.50
11claude-opus-4-7Anthropic PBC51.5%$5.00 / $25.00
12qwen3.7-maxAlibaba Cloud50.8%$2.50 / $7.50
13glm-5.2Z.ai50.8%$1.40 / $4.40
14qwen3.7-plusAlibaba Cloud47.0%$0.32 / $1.28
15claude-opus-4-5Anthropic PBC47.0%$5.00 / $25.00
16gpt-5.2OpenAI Inc.47.0%$1.75 / $14.00
17claude-opus-4-6Anthropic PBC46.2%$5.00 / $25.00
18deepseek-v4-pro-0424DeepInfra Inc.46.2%$1.30 / $2.60
19gpt-5.1OpenAI Inc.45.5%$1.25 / $10.00
20kimi-k2.7-codeMoonshot AI44.7%$0.95 / $4.00
21qwen3.6-plusAlibaba Cloud43.9%$0.50 / $3.00
22kimi-k2.6Moonshot AI43.9%$0.95 / $4.00
23GLM-5Z.ai43.2%$1.00 / $3.20
24glm-5.1Z.ai43.2%$1.40 / $4.40
25XiaomiMiMo/MiMo-V2.5-ProDeepInfra Inc.43.2%$1.00 / $3.00
26minimax-m3MiniMax42.4%$0.30 / $1.20
27gpt-5.4-nanoOpenAI Inc.42.4%$0.20 / $1.25
28gemini-3-pro-previewGoogle LLC (Gemini API)41.7%$2.00 / $12.00
29minimax-m2.7MiniMax39.4%$0.30 / $1.20
30gemini-3.5-flashGoogle LLC (Gemini API)39.4%$1.50 / $9.00

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

agentic-coding
terminal-bench hard65.9%gpt-5.6-sol
entries9530 shown
served byOpenAI Inc.high first

gpt-5.6-sol leads claude-fable-5 by 3.0 points

distribution

30 entries
med 47.0%65.9%#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 3 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best coding agent 30 ranked by terminal-bench hard

Terminal-Bench Hard measures how well a model operates as a coding agent in a real terminal, running commands, editing files, and fixing repositories end-to-end. It is the closest proxy to how models perform inside tools like Claude Code, Cursor and Codex.

method

Scores for Terminal-Bench Hard come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

95 entriesshowing 30sort high firstsource benchmarkleader gpt-5.6-sol 65.9%rebuilt daily