OFFICIAL PRICES + YOUR ASSUMPTIONS
LLM API Cost Calculator & Model Routing
Estimate monthly AI API cost with cache, retries, failures, fallback traffic, and accepted workflows kept visible as assumptions.
Price source verified 2026-07-27
Modeled monthly cost breakdown
Exact price components for the selected model route. Fallback traffic is included in each token component.
DeepSeek V4 Pro$0.435/M input · $0.003625/M cached · $0.87/M output
Official source Sensitivity
Modeled USD per 1,000 accepted workflows. Click a cell to apply its acceptance and cached-input rates.
Acceptance ↓ / Cache →
0%20%40%50%60%70%80%95%
90%
85%
80%
75%
Current scenario: 50% cached input and 85% acceptance. These are assumptions, not TokenAir observations.
Candidate comparison
Same workload and acceptance assumption; only official model prices change.
| Scenario | Model | Input / 1M | Output / 1M | Modeled monthly cost | Accepted workflows | Cost / accepted | Evidence |
|---|---|---|---|---|---|---|---|
| Selected route | DeepSeek V4 ProDeepSeek | $0.44 | $0.87 | $129.64 | 85,000 | $0.00153 | Official docs |
| Candidate 1 | DeepSeek V4 FlashDeepSeek | $0.14 | $0.28 | $91.77 | 85,000 | $0.00108 | Official docs |
| Candidate 2 | Gemini 3.5 Flash-LiteGoogle Gemini | $0.30 | $2.50 | $152.14 | 85,000 | $0.00179 | Official docs |
| Candidate 3 | Codestral 25.08Mistral AI | $0.30 | $0.90 | $120.60 | 85,000 | $0.00142 | Official docs |
How to read thisThe calculator is real arithmetic over official list prices; it is not a production forecast until you replace every default with measured workload data.
Review methodology Start by measuring accepted workflows, retries and cacheable input in your own logs. Then compare the exported scenario with your invoice.