Alibaba slashes Qwen-Audio 3.1 voice API prices by up to 95%

Alibaba cut the price of its Qwen-Audio 3.1 voice APIs in a single release:
- speech recognition down 95%;
- text-to-speech down 70%;
- real-time voice down 85%.
WION described it as cutting the price of AI voice tools by up to 95% in one move, and it turns voice AI, until now a premium tier at many Western providers, into a near-commodity input for developers building voice agents.
The cuts sit within a broader full-stack strategy Alibaba laid out at its Apsara Conference:
- Qwen 4, planned at 5 to 10 trillion parameters, is in training.
- The in-house Zhenwu V900 accelerator, rated at three times its predecessor's performance, enters mass production in Q1 2027.
- Data-center capacity is set to expand beyond 20GW by 2032.
- New cloud regions are planned for Finland, the Netherlands and Turkey.
Owning chips, models and cloud gives Alibaba room to price aggressively in ways pure model labs cannot.
Competitively, this is the Chinese price war extending from text into speech. It lands the same week DeepSeek cut V4.1-Flash pricing 70%. It pressures Google's new Gemini 3.8 Flash TTS, OpenAI's realtime voice APIs and ElevenLabs-style specialists. Open-source builders are already running Qwen3-TTS on SageMaker via vLLM-Omni, so Alibaba now competes on both open weights and cheap hosted endpoints.
The questions to watch:
- whether quality and latency hold up at scale;
- whether the new European regions ease compliance concerns for Western customers;
- whether competitors match the prices or differentiate on voice cloning and expressivity instead.