lists

9 categories
The 9 ranking lists, what each measures, and how it sorts
#listmeasuressourceorder
01Smartest overallRanked by Intelligence Indexbenchmarkhigh first
02Best for codingRanked by Coding Indexbenchmarkhigh first
03Best coding agentRanked by Terminal-Bench Hardbenchmarkhigh first
04Best for reasoningRanked by GPQA Diamondbenchmarkhigh first
05Best at mathRanked by Math Indexbenchmarkhigh first
06Best for tool useRanked by τ²-Benchbenchmarkhigh first
07Best for knowledgeRanked by MMLU Probenchmarkhigh first
08CheapestLowest input + output price per 1M tokenspricelow first
09Longest contextMax tokens in a single promptcontexthigh first

Every list is rebuilt daily from the same catalog the router reads at request time. Nothing here is hand-ranked. A benchmark measures one narrow skill, so two lists disagreeing is the normal case and not an error.

notes

reference

rankings 9 lists

Each list is built from one published benchmark or from live catalog pricing, rebuilt daily from the same catalog the router uses. 7 of the 9 rank on a benchmark score, and the rest rank on a number the catalog publishes directly: price per million tokens, or context window.

Smartest overall

The Artificial Analysis Intelligence Index combines many independent evaluations (reasoning, coding, math, science, agentic tool use) into a single capability score. It is the best at-a-g...

Best for coding

The Coding Index blends multiple coding evaluations: contamination-free code generation (LiveCodeBench), research-level scientific coding (SciCode), and agentic terminal tasks (Terminal-B...

Best coding agent

Terminal-Bench Hard measures how well a model operates as a coding agent in a real terminal, running commands, editing files, and fixing repositories end-to-end. It is the closest proxy t...

Best for reasoning

GPQA Diamond is a set of graduate-level science questions written by domain experts and filtered so that PhD students with internet access still struggle. It's the most reliable signal we...

Best at math

The Math Index aggregates competition and advanced math evaluations (including AIME). These problems require real symbolic reasoning across multiple novel steps, memorization gets a model...

Best for tool use

τ²-Bench measures multi-turn agentic tool use: calling functions, following policies, and completing realistic tasks over many turns. If you are building agents or tool-calling workflows,...

Best for knowledge

MMLU Pro tests broad knowledge across academic and professional subjects with harder, reasoning-heavy questions than the original MMLU. A strong score indicates wide, reliable factual cov...

Cheapest

Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cos...

Longest context

A larger context window means more tokens you can fit in a single prompt, useful for whole-codebase analysis, long document Q&A, and agentic workflows. Note: effective quality often degra...

one api for every model on these lists

Requesty is OpenAI-compatible and routes to every model ranked here. Change one parameter to switch, and keep the failover, caching and spend controls.

9 lists7 benchmark2 catalogrebuilt dailyno hand ranking