Cost engine

What does your AI stack cost?

Pick what you're building, set your scale - get the cheapest viable stack, layer by layer, priced at the best provider for your workload.

Your scale
Traffic profilesets the GPU utilisation assumed for self-hosting
Recommended
cheapest viable stack for your workload
≈ $6.70 /mo
  • ·$0.12/mo below the next viable combination
  • ·no dedicated GPU needed (self-host pays off above ~2.8M req/mo)
  • ·every model shown fits your 5K-token context per request
swap any layer below ↓
Build your stack
  • LLM API$6.50
    or self-host: 1× A40 at ~55% GPU use $358/mo· pays off above ~2.8M requests/mo
  • Embeddings$0.08
    compare all ↗· initial index amortised over 12 months + query embeddings
  • Vector DB$0.12
    compare all ↗· 1 vector per ~512-token chunk · 1 query per request· medium
  • Object storage$<0.01
  • Observability$0
    compare all ↗· free tier assumed
Total, your workload Confidence: medium≈ $6.70 /mo
Assumptions used

· · · · · · prices refreshed 2026-09-15

Where the bill goes
LLM API · input tokens60%
LLM API · output tokens37%
Vector DB2%
Embeddings1%
How to cut it
  • ·cut output to ~300 tokens/request-$1.19/mo(18%)
  • ·raise cache hit to 80%-$1.09/mo(16%)

Live prices, updated 2026-09-15. A model, not a quote - each layer is priced at its cheapest provider; embeddings amortise the initial index over 12 months. Confirm on each provider's site.

Concepts on this page