Requesty
Free Tool

LLM Cost Calculator

Compare API costs across 12 popular models. See how much smart routing and caching save at your scale.

Your workload

Optimization levers

63% savings
All GPT 5.6 Sol$4,800/mo
With routing$2,532/mo
Routing + caching$1,772/mo

Cost per model (300,000 requests/month)

DeepSeek V4 FlashDeepSeek
In $/M
$0.44
Out $/M
$1.32
Monthly
$264
Cached
$185
GLM 5.2Z.ai
In $/M
$0.75
Out $/M
$2.40
Monthly
$468
Cached
$328
Gemini 3.7 FlashGoogle
In $/M
$0.75
Out $/M
$3.75
Monthly
$630
Cached
$441
DeepSeek V4 ProDeepSeek
In $/M
$1.32
Out $/M
$3.96
Monthly
$792
Cached
$554
Claude Haiku 4.5Anthropic
In $/M
$1.00
Out $/M
$5.00
Monthly
$840
Cached
$588
GLM 5.3Z.ai
In $/M
$1.40
Out $/M
$4.40
Monthly
$864
Cached
$605
Gemini 3.5 FlashGoogle
In $/M
$1.50
Out $/M
$9.00
Monthly
$1,440
Cached
$1,008
Claude Sonnet 5Anthropic
In $/M
$2.00
Out $/M
$10.00
Monthly
$1,680
Cached
$1,176
Claude Opus 5Anthropic
In $/M
$5.00
Out $/M
$25.00
Monthly
$4,200
Cached
$2,940
GPT 5.6 SolOpenAI
In $/M
$5.00
Out $/M
$30.00
Monthly
$4,800
Cached
$3,360

Get these savings automatically

Requesty routes your requests to the cheapest model that meets your quality bar, caches repeated prompts, and fails over automatically. One API, 600+ models, zero infrastructure.

Frequently asked questions

How much does it cost to use LLM APIs?

LLM API costs vary widely by model. GPT 5.6 Sol costs $5 per million input tokens and $30 per million output tokens. DeepSeek V4 Flash costs $0.44 per million input and $1.32 per million output. Across a production workload that gap is the difference between a four-figure and a five-figure monthly bill. Smart routing through an AI gateway can reduce this by 50 to 80%.

How does smart routing reduce LLM costs?

Smart routing analyzes each request and sends simple tasks to cheap models (like DeepSeek V4 Flash or Gemini 3.7 Flash) and complex tasks to frontier models (like Claude Opus 5 or GPT 5.6 Sol). Since most production traffic is simple, routing 50 to 70% of requests to budget models typically cuts total spend by 50 to 80% with no quality loss on complex tasks.

What is the cheapest LLM API for production use?

As of 2026, DeepSeek V4 Flash ($0.44 per million input, $1.32 per million output), Gemini 3.7 Flash ($0.75 per million input, $3.75 per million output) and GLM 5.2 ($0.75 per million input, $2.40 per million output) are among the cheapest production-quality LLM APIs. For harder work, GLM 5.3 sits at $1.40 per million input. Using an AI gateway lets you access all of these through one API.