| # | model | served by | τ²-bench | share of leader | in / out per M |
|---|---|---|---|---|---|
| 01 | Z.ai | 99.1% | $1.40 / $4.40 | ||
| 02 | Anthropic PBC | 98.5% | $10.00 / $50.00 | ||
| 03 | Novita AI | 98.5% | $0.20 / $1.15 | ||
| 04 | Z.ai | 98.2% | $1.00 / $3.20 | ||
| 05 | Alibaba Cloud | 97.7% | $0.50 / $3.00 | ||
| 06 | xAI Corp. | 97.7% | $1.25 / $2.50 | ||
| 07 | Z.ai | 97.7% | $1.40 / $4.40 | ||
| 08 | DeepInfra Inc. | 96.2% | $1.30 / $2.60 | ||
| 09 | Moonshot AI | 95.9% | $0.95 / $4.00 | ||
| 10 | Moonshot AI | 95.9% | $0.60 / $3.00 | ||
| 11 | Z.ai | 95.9% | $0.60 / $2.20 | ||
| 12 | Google LLC (Gemini API) | 95.6% | $2.00 / $12.00 | ||
| 13 | Google LLC (Gemini API) | 95.6% | $1.50 / $9.00 | ||
| 14 | MiniMax | 95.3% | $0.30 / $1.20 | ||
| 15 | Alibaba Cloud | 94.7% | $2.50 / $7.50 | ||
| 16 | Anthropic PBC | 94.4% | $5.00 / $25.00 | ||
| 17 | Mistral AI SAS | 94.2% | $1.65 / $8.25 | ||
| 18 | DeepInfra Inc. | 94.2% | $1.00 / $3.00 | ||
| 19 | OpenAI Inc. | 93.9% | $5.00 / $30.00 | ||
| 20 | DeepInfra Inc. | 93.9% | $0.26 / $2.60 | ||
| 21 | Alibaba Cloud | 93.0% | $0.32 / $1.28 | ||
| 22 | Google LLC (Vertex AI) | 93.0% | $0.60 / $2.50 | ||
| 23 | Novita AI | 92.4% | $0.30 / $2.50 | ||
| 24 | Runware Inc. | 92.1% | $0.05 / $0.07 | ||
| 25 | Anthropic PBC | 92.1% | $5.00 / $25.00 | ||
| 26 | Microsoft Azure AI | 92.1% | $1.75 / $14.00 | ||
| 27 | Google LLC (Vertex AI) | 90.6% | $0.56 / $1.68 | ||
| 28 | xAI Corp. | 90.4% | $0.30 / $0.50 | ||
| 29 | Moonshot AI | 90.1% | $0.95 / $4.00 | ||
| 30 | Novita AI | 89.8% | $0.30 / $2.50 |
Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.
glm-5.2 and claude-fable-5 are within a rounding error of each other, so pick on price or latency
One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.
Best for tool use 30 ranked by τ²-bench
τ²-Bench measures multi-turn agentic tool use: calling functions, following policies, and completing realistic tasks over many turns. If you are building agents or tool-calling workflows, this predicts real-world reliability better than single-shot benchmarks.
method
Scores for τ²-Bench come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
