Back
AlibabaSeptember 18, 20263 sources

Alibaba Releases Qwen3.8-Omni-Flash with 1M-Token Context and Up to 98% Cheaper Audio

AI Analysis

Qwen3.8-Omni-Flash is Alibaba's most aggressive push yet on omnimodal price-performance. It handles text, images, audio and video natively, carries a 1-million-token context window, and reports a roughly 25% average benchmark gain over its predecessor Qwen3.5-Omni-Plus across 30 evaluations. The headline is economic: audio input pricing drops more than 98% and video-audio input 93%, a move that reframes real-time multimodal apps from premium to commodity.

On capability, the model supports speech recognition across 74 languages and generation in 29, and Alibaba highlighted improved agentic performance on video tasks — a nod to agents that watch and act on live streams. Unlike much of Qwen's lineup, this release is API-only with no open weights for self-hosting, a notable retreat from the open-weight momentum Alibaba otherwise rode this week.

Competitively, the pricing cut is a direct shot at OpenAI's and Google's multimodal tiers and at incumbent transcription APIs, arriving the same week xAI shipped Grok Voice Transcribe 2.0 at $0.10/hour. Together they signal a price war in speech and multimodal I/O. Qwen3.8-Omni-Flash also lands alongside sibling releases — Qwen-Image-2.1 and Qwen3.8-LiveTranslate — underscoring a rapid-fire cadence.

The caveat developers flagged is the closed-weights posture: after building goodwill on open releases, an API-only omni model concentrates control back with Alibaba and raises data-residency questions, sharpened by the geopolitical friction over a US Federal Register site briefly running Qwen-powered search. Watch whether the pricing holds past introductory windows and whether independent benchmarks confirm the 25% claim.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog