Skip to content
Dashboard

GLM 5.3 Flash

GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.

Input and output price
Prices from: Input $0.07, Output $0.24, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'zai/glm-5.3-flash',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Your use is subject to Z.AI's Terms & Privacy Policies.