Back
AlibabaAugust 26, 20262 sources

Alibaba unveils cost-efficient Qwen3.8-Flash-Next, an open-weight Qwen4 preview

AI Analysis

Alibaba released Qwen3.8-Flash-Next, positioned as an early preview of its forthcoming Qwen4 architecture. The multimodal Mixture-of-Experts model carries 125 billion parameters and, per Alibaba, delivers competitive performance against Anthropic's Opus 4.6 and DeepSeek's V4-Flash while cutting training and inference costs to about one-ninth those of its predecessor, with notable gains in coding and office-task capabilities. Crucially, the weights are open-sourced on Hugging Face and ModelScope — a contrast to Alibaba's simultaneous decision to close the weights on its Wan3.0 video model.

The local-LLM community reacted enthusiastically. r/LocalLLaMA hosted multiple threads, including a 1,098-upvote 'tomorrow' hype post and a 724-upvote day-0 unsloth support announcement, while NVIDIA published a guide to running the model on GB300 NVL72 hardware for agentic coding. LlamaIndex's Jerry Liu reported that Qwen 3.8 led 20+ open-weight models on his ExtractBench document-extraction benchmark, calling it 'the pack' leader ahead of kimi-k3.

The release intensifies pressure on the value tier, arriving the same week as Gemini 3.7 Flash, Grok 4.6, and GLM-5.3-Flash. Alibaba's stock rose on the news as investors read the cost focus as a play for enterprise coding and office workloads. Notably, Thomson Reuters recently built its Thomson-1 model by 'realigning' Alibaba's open-source Qwen to reduce reliance on Anthropic's Claude — evidence that open Qwen weights are becoming a foundation for enterprises seeking to escape proprietary-model costs. Watch how Qwen4 proper builds on this preview and whether the one-ninth cost claim holds under independent benchmarking.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog