Cost engine
What does your AI stack cost?
Pick what you're building, set your scale - get the cheapest viable stack, layer by layer, priced at the best provider for your workload.
Your scale
Traffic profilesets the GPU utilisation assumed for self-hosting
Recommended
cheapest viable stack for your workload
≈ $6.70 /mo
- ·$0.12/mo below the next viable combination
- ·no dedicated GPU needed (self-host pays off above ~2.8M req/mo)
- ·every model shown fits your 5K-token context per request
swap any layer below ↓
Build your stack
- LLM API$6.50
- Embeddings$0.08compare all ↗· initial index amortised over 12 months + query embeddings
- Vector DB$0.12
- Object storage$<0.01compare all ↗· medium
- Observability$0compare all ↗· free tier assumed
Total, your workload Confidence: medium≈ $6.70 /mo
Assumptions used
· · · · · · prices refreshed 2026-09-15
Where the bill goes
LLM API · input tokens60%
LLM API · output tokens37%
Vector DB2%
Embeddings1%
How to cut it
- ·cut output to ~300 tokens/request-$1.19/mo(18%)
- ·raise cache hit to 80%-$1.09/mo(16%)
Live prices, updated 2026-09-15. A model, not a quote - each layer is priced at its cheapest provider; embeddings amortise the initial index over 12 months. Confirm on each provider's site.