Skip to content
Dashboard

Qwen 3.8 Flash

Qwen3.8-Flash is Qwen’s fast, cost-efficient multimodal model, combining advanced reasoning and generation with a native 1M-token context window. Built for coding, agentic workflows, and visual understanding, it handles large codebases, long documents, charts, videos, and desktop applications.

Input and output price
Input $0.16, Output $0.47, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'alibaba/qwen3.8-flash',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by Alibaba Cloud

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
991K1.1 s59 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+2
alibaba logo
09/01/2026
1M0.3 s107 tps
$0.10/M
$0.40/M
Read$0.01/M
Write$0.63/M
+1
alibaba logo
deepinfra logo
morph logo
+2
08/14/2026
1M3.6 s127 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+1
alibaba logo
fireworks logo
08/02/2026
991K1.3 s76 tps
$0.03/M+2 more
$0.13/M+2 more
Read$0.006/M
Write$0.04/M
+2
alibaba logo
07/28/2026
1M1.9 s58 tps
$0.40/M+1 more
$1.60/M+1 more
Read$0.08/M
Write$0.50/M
+2
alibaba logo
06/02/2026
262K0.6 s56 tps
$0.20/M
$0.88/M
Read$0.11/M
alibaba logo
deepinfra logo
09/23/2025

Your use is subject to Alibaba Cloud's Terms & Privacy Policies.