| # | model | served by | avg /M | leader / this | in / out per M |
|---|---|---|---|---|---|
| 01 | DeepInfra Inc. | $0.03 | $0.02 / $0.05 | ||
| 02 | Novita AI | $0.05 | $0.05 / $0.05 | ||
| 03 | Novita AI | $0.05 | $0.05 / $0.05 | ||
| 04 | Novita AI | $0.05 | $0.05 / $0.05 | ||
| 05 | DeepInfra Inc. | $0.06 | $0.02 / $0.10 | ||
| 06 | Runware Inc. | $0.06 | $0.05 / $0.07 | ||
| 07 | Novita AI | $0.07 | $0.07 / $0.07 | ||
| 08 | DeepInfra Inc. | $0.08 | $0.06 / $0.11 | ||
| 09 | DeepInfra Inc. | $0.09 | $0.07 / $0.10 | ||
| 10 | Runware Inc. | $0.09 | $0.03 / $0.14 | ||
| 11 | Novita AI | $0.09 | $0.09 / $0.09 | ||
| 12 | DeepInfra Inc. | $0.10 | $0.04 / $0.16 | ||
| 13 | DeepInfra Inc. | $0.10 | $0.07 / $0.14 | ||
| 14 | DeepInfra Inc. | $0.11 | $0.07 / $0.14 | ||
| 15 | DeepInfra Inc. | $0.11 | $0.07 / $0.14 | ||
| 16 | Runware Inc. | $0.11 | $0.09 / $0.13 | ||
| 17 | OpenAI Inc. | $0.11 | $0.02 / $0.20 | ||
| 18 | Runware Inc. | $0.11 | $0.08 / $0.15 | ||
| 19 | DeepInfra Inc. | $0.12 | $0.07 / $0.16 | ||
| 20 | Alibaba Cloud | $0.13 | $0.05 / $0.20 | ||
| 21 | Google LLC (Gemini API) | $0.13 | $0.05 / $0.20 | ||
| 22 | DeepInfra Inc. | $0.13 | $0.05 / $0.20 | ||
| 23 | Fireworks AI | $0.13 | $0.05 / $0.20 | ||
| 24 | Sail Research Co. | $0.14 | $0.09 / $0.18 | ||
| 25 | DeepInfra Inc. | $0.14 | $0.09 / $0.18 | ||
| 26 | DeepInfra Inc. | $0.14 | $0.06 / $0.22 | ||
| 27 | DeepInfra Inc. | $0.15 | $0.10 / $0.20 | ||
| 28 | Nebius AI | $0.15 | $0.06 / $0.24 | ||
| 29 | Doubleword | $0.15 | $0.10 / $0.20 | ||
| 30 | Runware Inc. | $0.16 | $0.07 / $0.25 |
Bars are inverted so the cheapest row is longest. Free tiers are excluded, so every row here is a production-priced option.
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo beats sao10k/l3-8b-stheno-v3.2 by 1.4x
One bar per ranked model, best on the left, drawn as a share of the leader. No other model comes within 10 percent of the leader, so the top of this list is decisive. The dashed line is the median.
Cheapest 30 ranked by avg /M
Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cost matters and you can test model quality on your own workload.
method
Ranked by combined input and output price per million tokens. Models with a $0 tier are excluded, so this list reflects options you can point production traffic at. Pricing syncs from upstream providers and Requesty adds no markup.
Built from the same catalog the router reads at request time, rebuilt daily.
other lists 8
one api for every model on this list
Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.
