Cheapest AI models by price per million tokens
Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cost matters and you can test model quality on your own workload.
- 馃
meta-llama/Meta-Llama-3.1-8B-Instruct-TurboDeepInfra Inc.路$0.02 in / $0.05 out$0.03 avg$0.03 avg - 馃
sao10k/l3-8b-lunarisNovita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 馃
sao10k/l3-8b-stheno-v3.2Novita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 4
meta-llama/llama-3.1-8b-instructNovita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 5
Qwen/Qwen3.5-2BDeepInfra Inc.路$0.02 in / $0.10 out$0.06 avg$0.06 avg - 6qwen3.5-4bRunware Inc.路$0.05 in / $0.07 out$0.06 avg$0.06 avg
- 7
baichuan/baichuan-m2-32bNovita AI路$0.07 in / $0.07 out$0.07 avg$0.07 avg - 8
phi-4:flexDeepInfra Inc.路$0.06 in / $0.11 out$0.08 avg$0.08 avg - 9
Qwen/Qwen3-235B-A22B-Instruct-2507DeepInfra Inc.路$0.07 in / $0.10 out$0.09 avg$0.09 avg - 10gpt-oss-120bRunware Inc.路$0.03 in / $0.14 out$0.09 avg$0.09 avg
- 11
gryphe/mythomax-l2-13bNovita AI路$0.09 in / $0.09 out$0.09 avg$0.09 avg - 12
nvidia/Nemotron-3-Nano-30B-A3B:flexDeepInfra Inc.路$0.04 in / $0.16 out$0.10 avg$0.10 avg - 13
phi-4DeepInfra Inc.路$0.07 in / $0.14 out$0.10 avg$0.10 avg - 14
deepseek-v4-flash-0731:flexDeepInfra Inc.路$0.07 in / $0.14 out$0.11 avg$0.11 avg - 15
deepseek-v4-flash-0424:flexDeepInfra Inc.路$0.07 in / $0.14 out$0.11 avg$0.11 avg - 16qwen3.5-9bRunware Inc.路$0.09 in / $0.13 out$0.11 avg$0.11 avg
- 17
gpt-5-nano:flexOpenAI Inc.路$0.02 in / $0.20 out$0.11 avg$0.11 avg - 18deepseek-v4-flash-0731Runware Inc.路$0.08 in / $0.15 out$0.11 avg$0.11 avg
- 19
Qwen/Qwen2.5-Coder-32B-InstructDeepInfra Inc.路$0.07 in / $0.16 out$0.12 avg$0.12 avg - 20
nvidia/Nemotron-3-Nano-30B-A3BDeepInfra Inc.路$0.05 in / $0.20 out$0.13 avg$0.13 avg - 21
qwen-turboAlibaba Cloud路$0.05 in / $0.20 out$0.13 avg$0.13 avg - 22
gemini-2.5-flash-lite:flexGoogle LLC (Gemini API)路$0.05 in / $0.20 out$0.13 avg$0.13 avg - 23nemotron-lightning-3.5-30b-a3bFireworks AI路$0.05 in / $0.20 out$0.13 avg$0.13 avg
- 24
deepseek-v4-flash-0731DeepInfra Inc.路$0.09 in / $0.18 out$0.14 avg$0.14 avg - 25deepseek-v4-flash-0731Sail Research Co.路$0.09 in / $0.18 out$0.14 avg$0.14 avg
- 26
Qwen/Qwen3-32B:flexDeepInfra Inc.路$0.06 in / $0.22 out$0.14 avg$0.14 avg - 27
deepseek-v4-flash-0424DeepInfra Inc.路$0.10 in / $0.20 out$0.15 avg$0.15 avg - 28deepseek-v4-flash-0424:flexDoubleword路$0.10 in / $0.20 out$0.15 avg$0.15 avg
- 29
nvidia/nemotron-3-nano-omniNebius AI路$0.06 in / $0.24 out$0.15 avg$0.15 avg - 30glm-5.3-flashRunware Inc.路$0.07 in / $0.25 out$0.16 avg$0.16 avg
Explore other rankings
Smartest overall
Ranked by Intelligence Index
Best for coding
Ranked by Coding Index
Best coding agent
Ranked by Terminal-Bench Hard
Best for reasoning
Ranked by GPQA Diamond
Best at math
Ranked by Math Index
Best for tool use
Ranked by 蟿虏-Bench
Best for knowledge
Ranked by MMLU Pro
Longest context
Max tokens in a single prompt
How we rank
Ranked by combined input + output price per million tokens. Models with a $0 tier are excluded so this list reflects production-priced options you can deploy against real traffic. Pricing is synced in real-time from upstream providers. Pay as you go adds 5% on top, or 0% if you bring your own keys.
One API for every model on this list
Requesty is OpenAI-compatible and routes to 600+ models. Switch between any of the models above by changing one parameter in your code.
