Back
AlibabaAugust 26, 20262 sources

Alibaba Releases Low-Cost Qwen3.8-Flash-Next, a 125B MoE Previewing Qwen4

AI Analysis

Released August 26, Qwen3.8-Flash-Next is a sparse MoE that activates about 6B of its 125B parameters per token, delivering strong price-to-performance for coding and office workloads. It ships with open weights, a native 262K context window extendable to one million tokens, and multimodal support — and Alibaba frames it as an experimental preview of the forthcoming Qwen4 architecture. The headline claim is efficiency: training reportedly requires roughly one-ninth the cost of comparable models.

The launch drew immediate ecosystem support. NVIDIA congratulated Alibaba and announced Day 0 fine-tuning via NeMo AutoModel and NeMo RL plus recipes for SGLang, vLLM, and LightLLM; the model also landed in OpenCode Go. Alibaba's own account touted '125B/6B, 1M context, multimodal.' Meta's Muse-Glimmer-30B and a Qwen 3.8-27B variant simultaneously appeared on Amazon SageMaker JumpStart, underscoring how fast open weights now propagate across clouds.

Competitively, it sharpens the cost war among open-weight providers — a theme running through this week alongside Gemini 3.7 Flash's aggressive pricing and GLM-5.3's open-weight release. Analysts at AI Business cautioned enterprises to look past headline token prices to real-world metrics like accuracy, latency, and tool-use reliability. The 2% share bump suggests investors reward the cost story, but the durability of Alibaba's efficiency claims — and whether Qwen4 proper delivers on the preview — remains the thing to watch.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog