leaderboard

30 of 95 by τ²-bench
Top 30 models ranked by τ²-bench
#modelserved byτ²-benchshare of leaderin / out per M
01glm-5.2Z.ai99.1%$1.40 / $4.40
02claude-fable-5Anthropic PBC98.5%$10.00 / $50.00
03stepfun/step-3.7-flashNovita AI98.5%$0.20 / $1.15
04GLM-5Z.ai98.2%$1.00 / $3.20
05qwen3.6-plusAlibaba Cloud97.7%$0.50 / $3.00
06grok-4.3xAI Corp.97.7%$1.25 / $2.50
07glm-5.1Z.ai97.7%$1.40 / $4.40
08deepseek-v4-pro-0424DeepInfra Inc.96.2%$1.30 / $2.60
09kimi-k2.6Moonshot AI95.9%$0.95 / $4.00
10kimi-k2.5Moonshot AI95.9%$0.60 / $3.00
11GLM-4.7Z.ai95.9%$0.60 / $2.20
12gemini-3.1-pro-previewGoogle LLC (Gemini API)95.6%$2.00 / $12.00
13gemini-3.5-flashGoogle LLC (Gemini API)95.6%$1.50 / $9.00
14minimax-m2.5MiniMax95.3%$0.30 / $1.20
15qwen3.7-maxAlibaba Cloud94.7%$2.50 / $7.50
16claude-opus-4-8Anthropic PBC94.4%$5.00 / $25.00
17mistral-medium-3-5Mistral AI SAS94.2%$1.65 / $8.25
18XiaomiMiMo/MiMo-V2.5-ProDeepInfra Inc.94.2%$1.00 / $3.00
19gpt-5.5OpenAI Inc.93.9%$5.00 / $30.00
20Qwen/Qwen3.5-27BDeepInfra Inc.93.9%$0.26 / $2.60
21qwen3.7-plusAlibaba Cloud93.0%$0.32 / $1.28
22kimi-k2Google LLC (Vertex AI)93.0%$0.60 / $2.50
23inclusionai/ring-2.6-1tNovita AI92.4%$0.30 / $2.50
24qwen3.5-4bRunware Inc.92.1%$0.05 / $0.07
25claude-opus-4-6Anthropic PBC92.1%$5.00 / $25.00
26gpt-5.2-codex @eastus2Microsoft Azure AI92.1%$1.75 / $14.00
27deepseek-v3.2Google LLC (Vertex AI)90.6%$0.56 / $1.68
28grok-3-minixAI Corp.90.4%$0.30 / $0.50
29kimi-k2.7-codeMoonshot AI90.1%$0.95 / $4.00
30inclusionai/ling-2.6-1tNovita AI89.8%$0.30 / $2.50

Bars are each score's share of the leader's, so a short bar is a real gap and not a rounding difference. One row per model family, using the lab's own endpoint where it exists.

top of list

tool-use
τ²-bench99.1%glm-5.2
entries9530 shown
served byZ.aihigh first

glm-5.2 and claude-fable-5 are within a rounding error of each other, so pick on price or latency

distribution

30 entries
med 94.6%99.1%#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. 30 models sit within 10 percent of the leader, so the top of this list is a cluster rather than a winner. The dashed line is the median.

notes

reference

Best for tool use 30 ranked by τ²-bench

τ²-Bench measures multi-turn agentic tool use: calling functions, following policies, and completing realistic tasks over many turns. If you are building agents or tool-calling workflows, this predicts real-world reliability better than single-shot benchmarks.

method

Scores for τ²-Bench come from Artificial Analysis, an independent benchmarking service. When a model is served by several providers (Anthropic direct, AWS Bedrock, Google Vertex), one canonical entry represents the model family so the ranking is not padded with duplicates. Benchmarks measure specific skills: validate on your own workload before committing.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

95 entriesshowing 30sort high firstsource benchmarkleader glm-5.2 99.1%rebuilt daily