Requesty

Cheapest AI models by price per million tokens

Ranked by combined input + output price per million tokens (excluding free-tier models). These are production-ready models that punch well above their price point, great defaults when cost matters and you can test model quality on your own workload.

  1. 馃
    DeepInfra Inc. logo
    meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
    DeepInfra Inc.$0.02 in / $0.05 out
    $0.03 avg
  2. 馃
    Novita AI logo
    sao10k/l3-8b-lunaris
    Novita AI$0.05 in / $0.05 out
    $0.05 avg
  3. 馃
    Novita AI logo
    sao10k/l3-8b-stheno-v3.2
    Novita AI$0.05 in / $0.05 out
    $0.05 avg
  4. 4
    Novita AI logo
    meta-llama/llama-3.1-8b-instruct
    Novita AI$0.05 in / $0.05 out
    $0.05 avg
  5. 5
    DeepInfra Inc. logo
    Qwen/Qwen3.5-2B
    DeepInfra Inc.$0.02 in / $0.10 out
    $0.06 avg
  6. 6
    qwen3.5-4b
    Runware Inc.$0.05 in / $0.07 out
    $0.06 avg
  7. 7
    Novita AI logo
    baichuan/baichuan-m2-32b
    Novita AI$0.07 in / $0.07 out
    $0.07 avg
  8. 8
    DeepInfra Inc. logo
    phi-4:flex
    DeepInfra Inc.$0.06 in / $0.11 out
    $0.08 avg
  9. 9
    DeepInfra Inc. logo
    Qwen/Qwen3-235B-A22B-Instruct-2507
    DeepInfra Inc.$0.07 in / $0.10 out
    $0.09 avg
  10. 10
    gpt-oss-120b
    Runware Inc.$0.03 in / $0.14 out
    $0.09 avg
  11. 11
    Novita AI logo
    gryphe/mythomax-l2-13b
    Novita AI$0.09 in / $0.09 out
    $0.09 avg
  12. 12
    DeepInfra Inc. logo
    nvidia/Nemotron-3-Nano-30B-A3B:flex
    DeepInfra Inc.$0.04 in / $0.16 out
    $0.10 avg
  13. 13
    DeepInfra Inc. logo
    phi-4
    DeepInfra Inc.$0.07 in / $0.14 out
    $0.10 avg
  14. 14
    DeepInfra Inc. logo
    deepseek-v4-flash-0731:flex
    DeepInfra Inc.$0.07 in / $0.14 out
    $0.11 avg
  15. 15
    DeepInfra Inc. logo
    deepseek-v4-flash-0424:flex
    DeepInfra Inc.$0.07 in / $0.14 out
    $0.11 avg
  16. 16
    qwen3.5-9b
    Runware Inc.$0.09 in / $0.13 out
    $0.11 avg
  17. 17
    OpenAI Inc. logo
    gpt-5-nano:flex
    OpenAI Inc.$0.02 in / $0.20 out
    $0.11 avg
  18. 18
    deepseek-v4-flash-0731
    Runware Inc.$0.08 in / $0.15 out
    $0.11 avg
  19. 19
    DeepInfra Inc. logo
    Qwen/Qwen2.5-Coder-32B-Instruct
    DeepInfra Inc.$0.07 in / $0.16 out
    $0.12 avg
  20. 20
    DeepInfra Inc. logo
    nvidia/Nemotron-3-Nano-30B-A3B
    DeepInfra Inc.$0.05 in / $0.20 out
    $0.13 avg
  21. 21
    Alibaba Cloud logo
    qwen-turbo
    Alibaba Cloud$0.05 in / $0.20 out
    $0.13 avg
  22. 22
    Google LLC (Gemini API) logo
    gemini-2.5-flash-lite:flex
    Google LLC (Gemini API)$0.05 in / $0.20 out
    $0.13 avg
  23. 23
    nemotron-lightning-3.5-30b-a3b
    Fireworks AI$0.05 in / $0.20 out
    $0.13 avg
  24. 24
    DeepInfra Inc. logo
    deepseek-v4-flash-0731
    DeepInfra Inc.$0.09 in / $0.18 out
    $0.14 avg
  25. 25
    deepseek-v4-flash-0731
    Sail Research Co.$0.09 in / $0.18 out
    $0.14 avg
  26. 26
    DeepInfra Inc. logo
    Qwen/Qwen3-32B:flex
    DeepInfra Inc.$0.06 in / $0.22 out
    $0.14 avg
  27. 27
    DeepInfra Inc. logo
    deepseek-v4-flash-0424
    DeepInfra Inc.$0.10 in / $0.20 out
    $0.15 avg
  28. 28
    deepseek-v4-flash-0424:flex
    Doubleword$0.10 in / $0.20 out
    $0.15 avg
  29. 29
    Nebius AI logo
    nvidia/nemotron-3-nano-omni
    Nebius AI$0.06 in / $0.24 out
    $0.15 avg
  30. 30
    glm-5.3-flash
    Runware Inc.$0.07 in / $0.25 out
    $0.16 avg

How we rank

Ranked by combined input + output price per million tokens. Models with a $0 tier are excluded so this list reflects production-priced options you can deploy against real traffic. Pricing is synced in real-time from upstream providers. Pay as you go adds 5% on top, or 0% if you bring your own keys.

One API for every model on this list

Requesty is OpenAI-compatible and routes to 600+ models. Switch between any of the models above by changing one parameter in your code.