Skip to content
Dashboard

GLM 5.3 Flash

GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.

Input and output price
Prices from: Input $0.07, Output $0.24, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'zai/glm-5.3-flash',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Regional Inference
Free Tier
Release Date
Z.AI
50% off
Legal:TermsPrivacy
1M131K2.6 s47 tps
$0.08/M
$0.25/M
Read$0.02/M
+1
08/26/2026
1M131K0.8 s132 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
US
08/26/2026
Novita AI
50% off
Legal:TermsPrivacy
1M128K1.0 s68 tps
$0.08/M
$0.25/M
Read$0.02/M
+1
08/26/2026
GMICloud
50% off
Legal:TermsPrivacy
1M1M5.2 s50 tps
$0.08/M
$0.25/M
Read$0.02/M
+1
08/26/2026
DeepInfra
50% off
Legal:TermsPrivacy
1M1M0.8 s43 tps
$0.08/M
$0.25/M
Read$0.02/M
+1
08/26/2026
1M1M0.9 s79 tps
$0.10/M
$0.40/M
Read$0.01/M
+1
08/26/2026
1M1M0.7 s16 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M1M0.5 s141 tps
$0.45/M
$1.50/M
Read$0.09/M
+1
08/26/2026
1M1M1.1 s104 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M1M1.0 s43 tps
$0.15/M
$0.42/M
Read$0.01/M
+1
08/26/2026
1M1M1.5 s139 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M131K1.3 s122 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M1M1.0 s73 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M131K1.8 s106 tps
$0.15/M
$0.50/M
Read$0.03/M
+1
08/26/2026
1M131K
$0.07/M
$0.24/M
Read$0.01/M
+1
08/26/2026
Runware
50% off
Legal:TermsPrivacy
131K131K0.5 s167 tps
$0.08/M
$0.25/M
Read$0.02/M
+1
08/26/2026
1M131K0.5 s258 tps
$0.10/M
$0.40/M
Read$0.03/M
+1
08/26/2026

Copy link to headingPlayground

Try out GLM 5.3 Flash by Z.AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

zai logo
zai logo

GLM 5.3 Flash

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Copy link to headingMore models by Z.AI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M2.2 s206 tps
$0.70/M
$2.20/M
Read$0.13/M
digitalocean logo
09/02/2026
1M0.1 s276 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.12/M
baseten logo
blackbox logo
deepinfra logo
+12
08/18/2026
1M0.5 s186 tps
$2.10/M
$6.60/M
Read$0.21/M
alibaba logo
baseten logo
fireworks logo
06/23/2026
1M0.2 s679 tps
$0.70/M+1 more
$2.20/M+1 more
Read$0.11/M
alibaba logo
baseten logo
blackbox logo
+15
06/16/2026
205K1.6 s62 tps
$1.40/M
$4.40/M
Read$0.26/M
deepinfra logo
novita logo
zai logo
04/07/2026
203K0.8 s131 tps
$1/M
$3.20/M
Read$0.20/M
bedrock logo
novita logo
zai logo
02/12/2026

Your use is subject to Z.AI's Terms & Privacy Policies.