Cheapest AI models by price per million tokens
Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cost matters and you can test model quality on your own workload.
- 馃
meta-llama/Meta-Llama-3.1-8B-Instruct-TurboDeepInfra Inc.路$0.02 in / $0.05 out$0.03 avg$0.03 avg - 馃
meta-llama/llama-3.1-8b-instructNovita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 馃
Sao10K/L3-8B-Stheno-v3.2Novita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 4
sao10k/l3-8b-lunarisNovita AI路$0.05 in / $0.05 out$0.05 avg$0.05 avg - 5
Qwen/Qwen3.5-2BDeepInfra Inc.路$0.02 in / $0.10 out$0.06 avg$0.06 avg - 6
baichuan/baichuan-m2-32bNovita AI路$0.07 in / $0.07 out$0.07 avg$0.07 avg - 7
Qwen/Qwen3-235B-A22B-Instruct-2507DeepInfra Inc.路$0.07 in / $0.10 out$0.09 avg$0.09 avg - 8
gryphe/mythomax-l2-13bNovita AI路$0.09 in / $0.09 out$0.09 avg$0.09 avg - 9
phi-4DeepInfra Inc.路$0.07 in / $0.14 out$0.10 avg$0.10 avg - 10
gpt-5-nano:flexOpenAI Inc.路$0.02 in / $0.20 out$0.11 avg$0.11 avg - 11
Qwen/Qwen2.5-Coder-32B-InstructDeepInfra Inc.路$0.07 in / $0.16 out$0.12 avg$0.12 avg - 12
qwen-turboAlibaba Cloud路$0.05 in / $0.20 out$0.13 avg$0.13 avg - 13
nvidia/Nemotron-3-Nano-30B-A3BDeepInfra Inc.路$0.05 in / $0.20 out$0.13 avg$0.13 avg - 14
deepseek-ai/DeepSeek-V4-FlashDeepInfra Inc.路$0.10 in / $0.20 out$0.15 avg$0.15 avg - 15
nvidia/nemotron-3-nano-omniNebius AI路$0.06 in / $0.24 out$0.15 avg$0.15 avg - 16deepseek-v4-flash:flexDoubleword路$0.10 in / $0.20 out$0.15 avg$0.15 avg
- 17
mistralai/mistral-nemoNovita AI路$0.17 in / $0.17 out$0.17 avg$0.17 avg - 18gpt-oss-20bFireworks AI路$0.07 in / $0.30 out$0.18 avg$0.18 avg
- 19
Qwen/Qwen3-32BDeepInfra Inc.路$0.10 in / $0.30 out$0.20 avg$0.20 avg - 20
qwen/qwen3-32bNebius AI路$0.10 in / $0.30 out$0.20 avg$0.20 avg - 21
gemma-3-27b-itNebius AI路$0.10 in / $0.30 out$0.20 avg$0.20 avg - 22
qwen/qwen3-30b-a3b-instruct-2507Nebius AI路$0.10 in / $0.30 out$0.20 avg$0.20 avg - 23
inclusionai/ling-2.6-flashNovita AI路$0.10 in / $0.30 out$0.20 avg$0.20 avg - 24
gemma-4-26B-A4B-itDeepInfra Inc.路$0.07 in / $0.34 out$0.20 avg$0.20 avg - 25
meta-llama/Llama-3.3-70B-Instruct-TurboDeepInfra Inc.路$0.12 in / $0.30 out$0.21 avg$0.21 avg - 26parasail-deepseek-v4-flashParasail路$0.14 in / $0.28 out$0.21 avg$0.21 avg
- 27
deepseek-chatDeepSeek路$0.14 in / $0.28 out$0.21 avg$0.21 avg - 28
deepseek-v4-flashDeepSeek路$0.14 in / $0.28 out$0.21 avg$0.21 avg - 29
deepseek-reasonerDeepSeek路$0.14 in / $0.28 out$0.21 avg$0.21 avg - 30deepseek-v4-flashsference路$0.14 in / $0.28 out$0.21 avg$0.21 avg
Explore other rankings
Smartest overall
Ranked by Intelligence Index
Best for coding
Ranked by Coding Index
Best coding agent
Ranked by Terminal-Bench Hard
Best for reasoning
Ranked by GPQA Diamond
Best at math
Ranked by Math Index
Best for tool use
Ranked by 蟿虏-Bench
Best for knowledge
Ranked by MMLU Pro
Longest context
Max tokens in a single prompt
How we rank
Ranked by combined input + output price per million tokens. Models with a $0 tier are excluded so this list reflects production-priced options you can deploy against real traffic. Pricing is synced in real-time from upstream providers and Requesty charges no markup.
One API for every model on this list
Requesty is OpenAI-compatible and routes to 600+ models. Switch between any of the models above by changing one parameter in your code.
