← LLM API prices

GLM 5.3 Flash

Zhipu

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

In
$0.15
Out
$0.50
Blended
$0.24
Where to buy this model
27 access routes
RouteInOutBlendedvs
DeepInfra cheapestfp4$0.07$0.25$0.12-
Relace $0.09$0.30$0.14+17%
Morph fp8$0.10$0.35$0.16+33%
Wafer $0.10$0.35$0.16+33%
StreamLake fp8$0.11$0.37$0.18+50%
GMICloud fp8$0.11$0.38$0.18+50%
Reka fp8$0.13$0.44$0.21+75%
Novita fp8$0.13$0.44$0.21+75%
Makora $0.14$0.47$0.22+83%
Crusoe fp4$0.15$0.50$0.24+100%
CoreWeave fp8$0.15$0.50$0.24+100%
Sail Research fp8$0.15$0.50$0.24+100%
AtlasCloud fp8$0.15$0.50$0.24+100%
Fireworks $0.15$0.50$0.24+100%
Phala fp8$0.15$0.50$0.24+100%
Friendli $0.15$0.50$0.24+100%
SiliconFlow fp8$0.15$0.50$0.24+100%
DigitalOcean $0.15$0.50$0.24+100%
Together $0.15$0.50$0.24+100%
Parasail fp8$0.15$0.50$0.24+100%
BaseTen fp8$0.15$0.50$0.24+100%
Venice $0.15$0.50$0.24+100%
Io Net fp8$0.15$0.50$0.24+100%
Cloudflare $0.15$0.50$0.24+100%
Z.AI fp8$0.15$0.50$0.24+100%
NextBit fp8$0.18$0.59$0.28+133%
Modal fp8$0.45$1.50$0.71+492%

official API · cloud platform · inference host · USD per 1M tokens. Same model, different providers - price and quantization vary.

Ctx
1.31072M
max output
131K
cache read /1M
$0.03
Family
glm
Modalities
text · image · video
Features
json · reasoning · tools

Live pricing, updated 2026-09-15.

Concepts on this page