leaderboard

30 of 75 by coding index
Top 30 models ranked by coding index
#modelserved bycoding indexshare of leaderin / out per M
01claude-fable-5.1Anthropic PBC81.6$10.00 / $50.00
02claude-opus-5Anthropic PBC78.0$5.00 / $25.00
03gpt-5.6-solOpenAI Inc.77.4$4.00 / $20.00
04gpt-6-astraOpenAI Inc.76.9$10.00 / $50.00
05grok-4.6xAI Corp.76.8$2.00 / $6.00
06gpt-5.6-terraOpenAI Inc.76.7$2.00 / $12.00
07claude-fable-5Anthropic PBC76.5$10.00 / $50.00
08gemini-3.8-flash @euGoogle LLC (Vertex AI)76.3$0.83 / $4.13
09kimi-k3Moonshot AI76.2$3.00 / $15.00
10gpt-5.5OpenAI Inc.74.9$5.00 / $30.00
11glm-5.3Z.ai74.8$1.40 / $4.40
12claude-opus-4-8Anthropic PBC74.3$5.00 / $25.00
13claude-opus-4-7Anthropic PBC73.6$5.00 / $25.00
14qwen3.8-flash-nextTensorX Ltd.73.1$0.20 / $0.50
15grok-4.5xAI Corp.72.4$2.00 / $6.00
16qwen3.8-2.4t-a95bTensorX Ltd.71.9$2.50 / $6.00
17qwen3.8-maxAlibaba Cloud71.8$2.00 / $6.00
18glm-5.3-flashZ.ai71.5$0.15 / $0.50
19claude-sonnet-5Anthropic PBC71.5$2.00 / $10.00
20gemini-3.7-flashGoogle LLC (Vertex AI)71.5$0.75 / $3.75
21gpt-5.6-lunaOpenAI Inc.71.4$0.20 / $1.20
22gpt-5.4OpenAI Inc.71.1$2.50 / $15.00
23gemini-3.6-flashGoogle LLC (Gemini API)69.2$1.50 / $7.00
24deepseek-v4-flash-0731Fireworks AI69.1$0.22 / $0.66
25deepseek-v4-pro-0813DeepSeek68.8$1.32 / $3.96
26glm-5.2Z.ai68.8$1.40 / $4.40
27gemini-3.1-pro-previewGoogle LLC (Gemini API)68.8$2.00 / $12.00
28qwen3.7-maxAlibaba Cloud66.0$2.50 / $7.50
29claude-sonnet-4-6Anthropic PBC63.0$3.00 / $15.00
30kimi-k2.6Moonshot AI61.8$0.95 / $4.00

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

coding
coding index81.6claude-fable-5.1
entries7530 shown
served byAnthropic PBChigh first

claude-fable-5.1 leads claude-opus-5 by 3.6 points

distribution

30 entries
med 72.281.6#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 13 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best for coding 30 ranked by coding index

The Coding Index blends multiple coding evaluations: contamination-free code generation (LiveCodeBench), research-level scientific coding (SciCode), and agentic terminal tasks (Terminal-Bench). It is a broader, harder-to-game signal than any single coding benchmark.

How to choose the best AI coding model

The best AI model for coding depends on your use case. For agentic coding tasks (editing files, running commands, fixing repos end-to-end), models with high Terminal-Bench scores perform best inside tools like Claude Code, Cursor, Codex, and Aider. For pure code generation from specifications, LiveCodeBench scores predict quality. For research and scientific computing, SciCode results matter most.

In production, model choice interacts with cost and latency. A slightly lower-ranked model at 10x cheaper per token may be the better default for high-volume autocomplete, while the top-ranked model is reserved for complex multi-file refactors. AI gateways like Requesty let you route between models dynamically: use the best model for hard tasks and a fast, cheap model for simple completions, all through one API.

Access all top coding models through a single OpenAI-compatible API at requesty.ai. Automatic prompt caching saves 40-60% on token costs, and failover ensures your coding agent never stops due to a single provider outage.

method

Scores for Coding Index come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

questions 3

What is the best AI model for coding in 2026?
Based on the Coding Index (combining LiveCodeBench, SciCode, and Terminal-Bench), the top AI coding models in 2026 are Claude Opus 4.8, GPT-5.5, and Gemini 3. Claude Opus 4.8 leads on agentic coding tasks, GPT-5.5 excels at code generation from specs, and Gemini 3 offers strong performance with large context windows for whole-codebase analysis.
Which AI model is best for agentic coding (Claude Code, Cursor, Codex)?
For agentic coding tools that edit files and run terminal commands, Terminal-Bench scores are the most predictive benchmark. Models ranking highest on Terminal-Bench Hard perform best inside Claude Code, Cursor, Codex, and Aider because the benchmark mirrors real coding agent workflows.
How can I access the best coding models through one API?
Requesty provides a single OpenAI-compatible API endpoint that routes to all top coding models (Claude, GPT, Gemini, DeepSeek, and 600+ others). You switch models by changing one parameter. Automatic failover ensures your coding agent keeps working even if one provider goes down. Prompt caching saves 40-60% on repeated code patterns.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

75 entriesshowing 30sort high firstsource benchmarkleader claude-fable-5.1 81.6rebuilt daily