LLM Cost Calculator
Compare API costs across 12 popular models. See how much smart routing and caching save at your scale.
Your workload
Optimization levers
Cost per model (300,000 requests/month)
- In $/M
- $0.44
- Out $/M
- $1.32
- Monthly
- $264
- Cached
- $185
- In $/M
- $0.75
- Out $/M
- $2.40
- Monthly
- $468
- Cached
- $328
- In $/M
- $0.75
- Out $/M
- $3.75
- Monthly
- $630
- Cached
- $441
- In $/M
- $1.32
- Out $/M
- $3.96
- Monthly
- $792
- Cached
- $554
- In $/M
- $1.00
- Out $/M
- $5.00
- Monthly
- $840
- Cached
- $588
- In $/M
- $1.40
- Out $/M
- $4.40
- Monthly
- $864
- Cached
- $605
- In $/M
- $1.50
- Out $/M
- $9.00
- Monthly
- $1,440
- Cached
- $1,008
- In $/M
- $2.00
- Out $/M
- $10.00
- Monthly
- $1,680
- Cached
- $1,176
- In $/M
- $5.00
- Out $/M
- $25.00
- Monthly
- $4,200
- Cached
- $2,940
- In $/M
- $5.00
- Out $/M
- $30.00
- Monthly
- $4,800
- Cached
- $3,360
| Model | Provider | Input $/M | Output $/M | Monthly cost | With caching |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | DeepSeek | $0.44 | $1.32 | $264 | $185 |
| GLM 5.2 | Z.ai | $0.75 | $2.40 | $468 | $328 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $630 | $441 | |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | $792 | $554 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $840 | $588 |
| GLM 5.3 | Z.ai | $1.40 | $4.40 | $864 | $605 |
| Gemini 3.5 Flash | $1.50 | $9.00 | $1,440 | $1,008 | |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $1,680 | $1,176 |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $4,200 | $2,940 |
| GPT 5.6 Sol | OpenAI | $5.00 | $30.00 | $4,800 | $3,360 |
Get these savings automatically
Requesty routes your requests to the cheapest model that meets your quality bar, caches repeated prompts, and fails over automatically. One API, 600+ models, zero infrastructure.
Frequently asked questions
How much does it cost to use LLM APIs?
LLM API costs vary widely by model. GPT 5.6 Sol costs $5 per million input tokens and $30 per million output tokens. DeepSeek V4 Flash costs $0.44 per million input and $1.32 per million output. Across a production workload that gap is the difference between a four-figure and a five-figure monthly bill. Smart routing through an AI gateway can reduce this by 50 to 80%.
How does smart routing reduce LLM costs?
Smart routing analyzes each request and sends simple tasks to cheap models (like DeepSeek V4 Flash or Gemini 3.7 Flash) and complex tasks to frontier models (like Claude Opus 5 or GPT 5.6 Sol). Since most production traffic is simple, routing 50 to 70% of requests to budget models typically cuts total spend by 50 to 80% with no quality loss on complex tasks.
What is the cheapest LLM API for production use?
As of 2026, DeepSeek V4 Flash ($0.44 per million input, $1.32 per million output), Gemini 3.7 Flash ($0.75 per million input, $3.75 per million output) and GLM 5.2 ($0.75 per million input, $2.40 per million output) are among the cheapest production-quality LLM APIs. For harder work, GLM 5.3 sits at $1.40 per million input. Using an AI gateway lets you access all of these through one API.
