Skip to content
Dashboard

Qwen 3.8 Flash Next

Qwen3.8-Flash-Next is Qwen’s experimental open-weight multimodal language model, pairing 125B parameters with just 6B activated for efficient reasoning and generation. With native 262K-token context extensible to 1M, vision support, and an architecture optimized for lower long-context latency, it is built for demanding agentic coding, tool use, visual reasoning, and multimodal automation.

Input and output price
Input $0.12, Output $0.40, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'alibaba/qwen3.8-flash-next',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by Alibaba Cloud

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
991K1.3 s44 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+2
alibaba logo
09/01/2026
991K1.3 s100 tps
$0.16/M
$0.47/M
Read$0.02/M
Write$0.20/M
+2
alibaba logo
08/26/2026
1M0.5 s120 tps
$0.10/M
$0.40/M
Read$0.01/M
Write$0.63/M
+1
alibaba logo
deepinfra logo
morph logo
+2
08/14/2026
1M3.7 s133 tps
$2/M
$6/M
Read$0.25/M
Write$2.50/M
+1
alibaba logo
fireworks logo
08/02/2026
991K2.1 s84 tps
$0.03/M+2 more
$0.13/M+2 more
Read$0.006/M
Write$0.04/M
+2
alibaba logo
07/28/2026
1M1.8 s58 tps
$0.40/M+1 more
$1.60/M+1 more
Read$0.08/M
Write$0.50/M
+2
alibaba logo
06/02/2026

Your use is subject to Alibaba Cloud's Terms & Privacy Policies.