leaderboard

30 of 696 by avg /M
Top 30 models ranked by avg /M
#modelserved byavg /Mleader / thisin / out per M
01meta-llama/Meta-Llama-3.1-8B-Instruct-TurboDeepInfra Inc.$0.03$0.02 / $0.05
02sao10k/l3-8b-stheno-v3.2Novita AI$0.05$0.05 / $0.05
03sao10k/l3-8b-lunarisNovita AI$0.05$0.05 / $0.05
04meta-llama/llama-3.1-8b-instructNovita AI$0.05$0.05 / $0.05
05Qwen/Qwen3.5-2BDeepInfra Inc.$0.06$0.02 / $0.10
06qwen3.5-4bRunware Inc.$0.06$0.05 / $0.07
07baichuan/baichuan-m2-32bNovita AI$0.07$0.07 / $0.07
08phi-4:flexDeepInfra Inc.$0.08$0.06 / $0.11
09Qwen/Qwen3-235B-A22B-Instruct-2507DeepInfra Inc.$0.09$0.07 / $0.10
10gpt-oss-120bRunware Inc.$0.09$0.03 / $0.14
11gryphe/mythomax-l2-13bNovita AI$0.09$0.09 / $0.09
12nvidia/Nemotron-3-Nano-30B-A3B:flexDeepInfra Inc.$0.10$0.04 / $0.16
13phi-4DeepInfra Inc.$0.10$0.07 / $0.14
14deepseek-v4-flash-0731:flexDeepInfra Inc.$0.11$0.07 / $0.14
15deepseek-v4-flash-0424:flexDeepInfra Inc.$0.11$0.07 / $0.14
16qwen3.5-9bRunware Inc.$0.11$0.09 / $0.13
17gpt-5-nano:flexOpenAI Inc.$0.11$0.02 / $0.20
18deepseek-v4-flash-0731Runware Inc.$0.11$0.08 / $0.15
19Qwen/Qwen2.5-Coder-32B-InstructDeepInfra Inc.$0.12$0.07 / $0.16
20qwen-turboAlibaba Cloud$0.13$0.05 / $0.20
21gemini-2.5-flash-lite:flexGoogle LLC (Gemini API)$0.13$0.05 / $0.20
22nvidia/Nemotron-3-Nano-30B-A3BDeepInfra Inc.$0.13$0.05 / $0.20
23nemotron-lightning-3.5-30b-a3bFireworks AI$0.13$0.05 / $0.20
24deepseek-v4-flash-0731Sail Research Co.$0.14$0.09 / $0.18
25deepseek-v4-flash-0731DeepInfra Inc.$0.14$0.09 / $0.18
26Qwen/Qwen3-32B:flexDeepInfra Inc.$0.14$0.06 / $0.22
27deepseek-v4-flash-0424DeepInfra Inc.$0.15$0.10 / $0.20
28nvidia/nemotron-3-nano-omniNebius AI$0.15$0.06 / $0.24
29deepseek-v4-flash-0424:flexDoubleword$0.15$0.10 / $0.20
30glm-5.3-flashRunware Inc.$0.16$0.07 / $0.25

Bars are inverted so the cheapest row is longest. Free tiers are excluded, so every row here is a production-priced option.

top of list

cheapest
avg /M$0.03meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
entries69630 shown
served byDeepInfra Inc.low first

meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo beats sao10k/l3-8b-stheno-v3.2 by 1.4x

distribution

30 entries
med $0.11$0.03#1#30

One bar per ranked model, best on the left, drawn as a share of the leader. No other model comes within 10 percent of the leader, so the top of this list is decisive. The dashed line is the median.

notes

reference

Cheapest 30 ranked by avg /M

Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cost matters and you can test model quality on your own workload.

method

Ranked by combined input and output price per million tokens. Models with a $0 tier are excluded, so this list reflects options you can point production traffic at. Pricing syncs from upstream providers and Requesty adds no markup.

Built from the same catalog the router reads at request time, rebuilt daily.

one api for every model on this list

Requesty is OpenAI-compatible. Switch between any two models above by changing one parameter, and keep the failover, caching and spend controls.

696 entriesshowing 30sort low firstsource priceleader meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo $0.03rebuilt daily