| # | model | served by | coding index | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | Anthropic PBC | 81.6 | $10.00 / $50.00 | ||
| 02 | Anthropic PBC | 78.0 | $5.00 / $25.00 | ||
| 03 | OpenAI Inc. | 77.4 | $4.00 / $20.00 | ||
| 04 | OpenAI Inc. | 76.9 | $10.00 / $50.00 | ||
| 05 | xAI Corp. | 76.8 | $2.00 / $6.00 | ||
| 06 | OpenAI Inc. | 76.7 | $2.00 / $12.00 | ||
| 07 | Anthropic PBC | 76.5 | $10.00 / $50.00 | ||
| 08 | Google LLC (Vertex AI) | 76.3 | $0.83 / $4.13 | ||
| 09 | Moonshot AI | 76.2 | $3.00 / $15.00 | ||
| 10 | OpenAI Inc. | 74.9 | $5.00 / $30.00 | ||
| 11 | Z.ai | 74.8 | $1.40 / $4.40 | ||
| 12 | Anthropic PBC | 74.3 | $5.00 / $25.00 | ||
| 13 | Anthropic PBC | 73.6 | $5.00 / $25.00 | ||
| 14 | TensorX Ltd. | 73.1 | $0.20 / $0.50 | ||
| 15 | xAI Corp. | 72.4 | $2.00 / $6.00 | ||
| 16 | TensorX Ltd. | 71.9 | $2.50 / $6.00 | ||
| 17 | Alibaba Cloud | 71.8 | $2.00 / $6.00 | ||
| 18 | Z.ai | 71.5 | $0.15 / $0.50 | ||
| 19 | Anthropic PBC | 71.5 | $2.00 / $10.00 | ||
| 20 | Google LLC (Vertex AI) | 71.5 | $0.75 / $3.75 | ||
| 21 | OpenAI Inc. | 71.4 | $0.20 / $1.20 | ||
| 22 | OpenAI Inc. | 71.1 | $2.50 / $15.00 | ||
| 23 | Google LLC (Gemini API) | 69.2 | $1.50 / $7.00 | ||
| 24 | Fireworks AI | 69.1 | $0.22 / $0.66 | ||
| 25 | DeepSeek | 68.8 | $1.32 / $3.96 | ||
| 26 | Z.ai | 68.8 | $1.40 / $4.40 | ||
| 27 | Google LLC (Gemini API) | 68.8 | $2.00 / $12.00 | ||
| 28 | Alibaba Cloud | 66.0 | $2.50 / $7.50 | ||
| 29 | Anthropic PBC | 63.0 | $3.00 / $15.00 | ||
| 30 | Moonshot AI | 61.8 | $0.95 / $4.00 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
claude-fable-5.1 leads claude-opus-5 by 3.6 points
One bar per ranked model, best on the left, drawn as a share of the leader. 13 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best for coding 30 ranked by coding index
The Coding Index blends multiple coding evaluations: contamination-free code generation (LiveCodeBench), research-level scientific coding (SciCode), and agentic terminal tasks (Terminal-Bench). It is a broader, harder-to-game signal than any single coding benchmark.
How to choose the best AI coding model
The best AI model for coding depends on your use case. For agentic coding tasks (editing files, running commands, fixing repos end-to-end), models with high Terminal-Bench scores perform best inside tools like Claude Code, Cursor, Codex, and Aider. For pure code generation from specifications, LiveCodeBench scores predict quality. For research and scientific computing, SciCode results matter most.
In production, model choice interacts with cost and latency. A slightly lower-ranked model at 10x cheaper per token may be the better default for high-volume autocomplete, while the top-ranked model is reserved for complex multi-file refactors. AI gateways like Requesty let you route between models dynamically: use the best model for hard tasks and a fast, cheap model for simple completions, all through one API.
Access all top coding models through a single OpenAI-compatible API at requesty.ai. Automatic prompt caching saves 40-60% on token costs, and failover ensures your coding agent never stops due to a single provider outage.
method
Scores for Coding Index come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
questions 3
What is the best AI model for coding in 2026?
Which AI model is best for agentic coding (Claude Code, Cursor, Codex)?
How can I access the best coding models through one API?
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
