How much does a RAG app cost to run?
A retrieval-augmented app touches four paid layers. Here's the cheapest viable combination for a reference workload, with a note on each.
The generation call. Input dominates (retrieved context), so cached-input pricing and a cheap-but-capable model matter most.
◈ or self-host: 1× A40 at ~55% GPU use $358/mo· pays off above ~2.8M requests/mo
One vector per chunk at index time, plus a small embedding per query. The initial index is a one-off; ongoing cost is tiny.
Stores and searches the vectors. Serverless bills storage + queries; dedicated bills compute time - the crossover is around a few million vectors.
The raw documents sit in object storage. Cheap; the swing factor is egress if you re-index often.
Tracing and eval. Langfuse and Helicone have generous free tiers - $0 until real volume.
Live prices, updated 2026-09-15. A model, not a quote - each layer is priced at its cheapest provider; embeddings amortise the initial index over 12 months. Confirm on each provider's site.