| # | list | measures | source | order |
|---|---|---|---|---|
| 01 | Smartest overall | Ranked by Intelligence Index | benchmark | high first |
| 02 | Best for coding | Ranked by Coding Index | benchmark | high first |
| 03 | Best coding agent | Ranked by Terminal-Bench Hard | benchmark | high first |
| 04 | Best for reasoning | Ranked by GPQA Diamond | benchmark | high first |
| 05 | Best at math | Ranked by Math Index | benchmark | high first |
| 06 | Best for tool use | Ranked by τ²-Bench | benchmark | high first |
| 07 | Best for knowledge | Ranked by MMLU Pro | benchmark | high first |
| 08 | Cheapest | Lowest input + output price per 1M tokens | price | low first |
| 09 | Longest context | Max tokens in a single prompt | context | high first |
Every list is rebuilt daily from the same catalog the router reads at request time. Nothing here is hand-ranked. A benchmark measures one narrow skill, so two lists disagreeing is the normal case and not an error.
rankings 9 lists
Each list is built from one published benchmark or from live catalog pricing, rebuilt daily from the same catalog the router uses. 7 of the 9 rank on a benchmark score, and the rest rank on a number the catalog publishes directly: price per million tokens, or context window.
Smartest overall
The Artificial Analysis Intelligence Index combines many independent evaluations (reasoning, coding, math, science, agentic tool use) into a single capability score. It is the best at-a-g...
Best for coding
The Coding Index blends multiple coding evaluations: contamination-free code generation (LiveCodeBench), research-level scientific coding (SciCode), and agentic terminal tasks (Terminal-B...
Best coding agent
Terminal-Bench Hard measures how well a model operates as a coding agent in a real terminal, running commands, editing files, and fixing repositories end-to-end. It is the closest proxy t...
Best for reasoning
GPQA Diamond is a set of graduate-level science questions written by domain experts and filtered so that PhD students with internet access still struggle. It's the most reliable signal we...
Best at math
The Math Index aggregates competition and advanced math evaluations (including AIME). These problems require real symbolic reasoning across multiple novel steps, memorization gets a model...
Best for tool use
τ²-Bench measures multi-turn agentic tool use: calling functions, following policies, and completing realistic tasks over many turns. If you are building agents or tool-calling workflows,...
Best for knowledge
MMLU Pro tests broad knowledge across academic and professional subjects with harder, reasoning-heavy questions than the original MMLU. A strong score indicates wide, reliable factual cov...
Cheapest
Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cos...
Longest context
A larger context window means more tokens you can fit in a single prompt, useful for whole-codebase analysis, long document Q&A, and agentic workflows. Note: effective quality often degra...
one api for every model on these lists
Requesty is OpenAI-compatible and routes to every model ranked here. Change one parameter to switch, and keep the failover, caching and spend controls.
