Monthly API cost
Enter your traffic and token sizes - see the monthly cost for every model, cheapest first.
≈ 70.0M input tokens · 35.0M output tokens per month
| Model | Cheapest via | Cost / month | vs cheapest |
|---|---|---|---|
| gpt-oss-20b | Darkbloom/ 13 | $4.90 | ×1.0 |
| Mistral Small 3 | DeepInfra | $6.30 | ×1.3 |
| DeepSeek V4 Flash 0731 | OpenInference/ 26 | $6.30 | ×1.3 |
| Qwen3.7 Flash | Alibaba | $6.65 | ×1.4 |
| Command R7B (12-2024) | Cohere | $7.88 | ×1.6 |
| gpt-oss-120b | AkashML/ 20 | $8.05 | ×1.6 |
| GPT-5 Nano | OpenAI/ 2 | $8.75 | ×1.8 |
| Phi 4 | DeepInfra | $9.80 | ×2.0 |
| Qwen3 30B A3B Instruct 2507 | StreamLake/ 5 | $10.12 | ×2.1 |
| Qwen3.5-9B | Darkbloom/ 6 | $10.15 | ×2.1 |
| Gemini 2.5 Flash Lite | Google AI Studio/ 2 | $10.50 | ×2.1 |
| Gemma 4 26B A4B | Darkbloom/ 11 | $10.64 | ×2.2 |
| Gemma 3 27B | DeepInfra/ 4 | $11.20 | ×2.3 |
| Mistral Small 3.2 24B | DeepInfra/ 3 | $12.25 | ×2.5 |
| Granite 4.2 8B | CoreWeave/ 2 | $12.25 | ×2.5 |
Live prices, updated 2026-09-15. Each model is priced at its cheapest provider for your workload; cached input uses each provider's cache-read rate. Excludes batch discounts and free tiers.