Gemma 4 26B A4B
GoogleGemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
In
$0.09
Out
$0.30
Blended
$0.14
Where to buy this model
11 access routes| Route | In | Out | Blended | vs |
|---|---|---|---|---|
| ● Darkbloom cheapest | $0.04 | $0.22 | $0.09 | - |
| ● DekaLLM bf16 | $0.06 | $0.33 | $0.13 | +44% |
| $0.07 | $0.34 | $0.14 | +56% | |
| ● NextBit bf16 | $0.09 | $0.30 | $0.14 | +56% |
| ● Cloudflare | $0.10 | $0.30 | $0.15 | +67% |
| ● Makora | $0.10 | $0.34 | $0.16 | +78% |
| ● Venice bf16 | $0.13 | $0.40 | $0.20 | +122% |
| $0.13 | $0.40 | $0.20 | +122% | |
| $0.13 | $0.40 | $0.20 | +122% | |
| ● SiliconFlow fp8 | $0.14 | $0.40 | $0.21 | +133% |
| $0.15 | $0.60 | $0.26 | +189% |
● official API · ● cloud platform · ● inference host · USD per 1M tokens. Same model, different providers - price and quantization vary.
- Ctx
- 262K
- max output
- 236K
- cache read /1M
- $0.05
- Family
- gemma
- Modalities
- image · text · video
- Features
- json · reasoning · tools
Live pricing, updated 2026-09-15.