← LLM API prices

Gemma 4 31B

Google

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

In
$0.09
Out
$0.34
Blended
$0.15
Where to buy this model
12 access routes
RouteInOutBlendedvs
DeepInfra cheapestfp4$0.09$0.34$0.15-
CoreWeave fp4$0.10$0.34$0.16+7%
Venice bf16$0.12$0.36$0.18+20%
Chutes fp4$0.12$0.37$0.18+20%
Crusoe $0.14$0.40$0.21+40%
Friendli $0.14$0.40$0.21+40%
Novita bf16$0.14$0.40$0.21+40%
Parasail fp8$0.15$0.40$0.21+40%
Together $0.39$0.97$0.53+253%
SambaNova $0.38$1.15$0.57+280%
ModelRun fp4$0.75$1.00$0.81+440%
SiliconFlow fp8$0.75$1.00$0.81+440%

official API · cloud platform · inference host · USD per 1M tokens. Same model, different providers - price and quantization vary.

Ctx
262K
max output
16K
cache read /1M
$0.05
Family
gemma
Modalities
image · text · video
Features
json · reasoning · tools

Live pricing, updated 2026-09-15.

Concepts on this page