Alibaba's Qwen ships 1M-token Omni-Flash with 98% audio-cost cut as US government use draws scrutiny

Alibaba released Qwen3.8-Omni-Flash, its next-generation omni-modal model capable of processing text, images, audio, and video within a single model and workflow, with agentic tool use. It supports a 1-million-token context window and posts an average score improvement of over 25% across 29 benchmarks. The most aggressive move is on price: API costs for audio input were cut by more than 98%, and audio-video input costs by more than 93% — an economics play that undercuts rivals on multimodal inference.
The cost reduction is the story's competitive core. By collapsing audio pricing, Alibaba is targeting voice and multimodal agent workloads where per-token audio costs have been a barrier, positioning Qwen as the low-cost omni-modal option against Google's Omni and OpenAI's multimodal offerings. The 1M-token context also puts it in the same tier as Moonshot's Kimi K3, both landing on the market the same week and intensifying the long-context race.
But the launch is shadowed by a national-security controversy. Reporting revealed the Federal Register, run by the National Archives, had a search feature powered by Qwen live as of September 17 — days after the FBI accused Alibaba of imitating Anthropic's work. The juxtaposition of a Chinese model quietly running on a US government website amid IP-theft accusations drew sharp reaction over supply-chain and provenance risks.
Separately, Alibaba also open-sourced its Damo Radar model for detecting nearly 150 abdominal conditions from CT scans, underscoring a broad Qwen-and-beyond release cadence. Watch whether US agencies review third-party model provenance in procurement, and whether the 98% audio-cost cut forces competitors to respond on pricing.