Skip to content
Dashboard

GLM 5.3 Flash

GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.

Input and output price
Prices from: Input $0.07, Output $0.24, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'zai/glm-5.3-flash',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by Z.AI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M2.2 s210 tps
$0.70/M
$2.20/M
Read$0.13/M
digitalocean logo
09/02/2026
1M0.2 s330 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.12/M
baseten logo
blackbox logo
deepinfra logo
+12
08/18/2026
1M0.5 s209 tps
$2.10/M
$6.60/M
Read$0.21/M
alibaba logo
baseten logo
fireworks logo
06/23/2026
1M0.2 s501 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.11/M
alibaba logo
baseten logo
blackbox logo
+15
06/16/2026
205K1.4 s36 tps
$1.40/M
$4.40/M
Read$0.26/M
deepinfra logo
novita logo
zai logo
04/07/2026
203K0.9 s121 tps
$1/M
$3.20/M
Read$0.20/M
bedrock logo
novita logo
zai logo
02/12/2026

Your use is subject to Z.AI's Terms & Privacy Policies.