| # | model | served by | terminal-bench hard | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | OpenAI Inc. | 65.9% | $4.00 / $20.00 | ||
| 02 | Anthropic PBC | 62.9% | $10.00 / $50.00 | ||
| 03 | OpenAI Inc. | 60.6% | $5.00 / $30.00 | ||
| 04 | Anthropic PBC | 58.3% | $5.00 / $25.00 | ||
| 05 | OpenAI Inc. | 57.6% | $2.00 / $12.00 | ||
| 06 | OpenAI Inc. | 57.6% | $2.50 / $15.00 | ||
| 07 | Google LLC (Gemini API) | 53.8% | $2.00 / $12.00 | ||
| 08 | Anthropic PBC | 53.0% | $3.00 / $15.00 | ||
| 09 | OpenAI Inc. | 53.0% | $1.75 / $14.00 | ||
| 10 | OpenAI Inc. | 52.3% | $0.75 / $4.50 | ||
| 11 | Anthropic PBC | 51.5% | $5.00 / $25.00 | ||
| 12 | Alibaba Cloud | 50.8% | $2.50 / $7.50 | ||
| 13 | Z.ai | 50.8% | $1.40 / $4.40 | ||
| 14 | Alibaba Cloud | 47.0% | $0.32 / $1.28 | ||
| 15 | Anthropic PBC | 47.0% | $5.00 / $25.00 | ||
| 16 | OpenAI Inc. | 47.0% | $1.75 / $14.00 | ||
| 17 | Anthropic PBC | 46.2% | $5.00 / $25.00 | ||
| 18 | DeepInfra Inc. | 46.2% | $1.30 / $2.60 | ||
| 19 | OpenAI Inc. | 45.5% | $1.25 / $10.00 | ||
| 20 | Moonshot AI | 44.7% | $0.95 / $4.00 | ||
| 21 | Alibaba Cloud | 43.9% | $0.50 / $3.00 | ||
| 22 | Moonshot AI | 43.9% | $0.95 / $4.00 | ||
| 23 | Z.ai | 43.2% | $1.00 / $3.20 | ||
| 24 | Z.ai | 43.2% | $1.40 / $4.40 | ||
| 25 | DeepInfra Inc. | 43.2% | $1.00 / $3.00 | ||
| 26 | MiniMax | 42.4% | $0.30 / $1.20 | ||
| 27 | OpenAI Inc. | 42.4% | $0.20 / $1.25 | ||
| 28 | Google LLC (Gemini API) | 41.7% | $2.00 / $12.00 | ||
| 29 | MiniMax | 39.4% | $0.30 / $1.20 | ||
| 30 | Google LLC (Gemini API) | 39.4% | $1.50 / $9.00 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
gpt-5.6-sol leads claude-fable-5 by 3.0 points
One bar per ranked model, best on the left, drawn as a share of the leader. 3 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best coding agent 30 ranked by terminal-bench hard
Terminal-Bench Hard measures how well a model operates as a coding agent in a real terminal, running commands, editing files, and fixing repositories end-to-end. It is the closest proxy to how models perform inside tools like Claude Code, Cursor and Codex.
method
Scores for Terminal-Bench Hard come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
