Alibaba releases Qwen Audio 3.1 and cuts audio AI prices by 95 percent
3 sources · Google News · The Decoder · news.aibase.com- Boom: Alibaba slashed AI audio service prices by up to 95 percent alongside the launch
- Boom: Qwen Audio 3.1 includes five models covering ASR, TTS, and real-time interaction
- Boom: ASR-Next model adds multi-speaker identification with timestamps and emotion detection
- Boom: ASR model automatically removes filler words and improves dialect recognition
The story in full
Alibaba's Qwen team released Qwen-Audio-3.1 on September 23, 2026, a lineup of five models covering automatic speech recognition, text-to-speech, and real-time interaction. Alongside the release, Alibaba cut prices on its AI audio services by up to 95 percent.
The ASR model improves multilingual and dialect recognition and removes filler words automatically. A variant called ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise. The TTS component handles multilingual speech synthesis.
Analysis
370 wordsOn September 23, 2026, Alibaba's Qwen team released Qwen-Audio-3.1, a set of five models spanning automatic speech recognition, text-to-speech synthesis, and real-time voice interaction. The ASR model improves recognition across multiple languages and dialects while automatically stripping filler words and repetitions from transcripts. A more advanced variant, ASR-Next, adds multi-speaker identification with timestamps and can detect emotions, ambient sounds, and machine noise. Alongside the model launch, Alibaba cut prices on its AI audio services by up to 95 percent.
The scale of the price cut is the detail that elevates this beyond a routine product update. A 95 percent reduction compresses what developers and businesses pay for audio AI services dramatically, which could accelerate adoption across industries that had treated the cost as a barrier, including transcription services, call center automation, and accessibility tooling. The breadth of the lineup, covering recognition, synthesis, and real-time interaction in one release, also signals that Alibaba is positioning Qwen Audio as a full-stack audio AI platform rather than a single-purpose tool. The genuine open question is whether the capability improvements match the pricing aggression, or whether the price cut is primarily a market-share play in a space where several competitors are already established.
None of the three camps have published reactions to this story yet. The Pro-AI camp would typically treat a major price cut of this kind as evidence that AI capabilities are becoming widely accessible, lowering the barrier for developers and smaller companies to build voice-enabled products. The Anti-AI camp would likely raise concerns about the displacement of human roles in transcription, voice acting, and call center work, especially as emotion detection and speaker identification expand what automated systems can do. The Middle Ground camp would probably acknowledge the practical benefits while pressing for transparency about accuracy rates, data privacy in multilingual contexts, and the conditions under which emotion and ambient sound detection is deployed.
The argument that would most directly settle the capability debate is independent benchmarking of Qwen-Audio-3.1 against competing ASR and TTS systems on multilingual and dialect-heavy datasets. On the commercial side, whether other audio AI providers respond with their own price cuts in the weeks following September 23 would indicate how much competitive pressure the launch actually creates.
Where do you stand?
Add your take
0 reader votesSign in with Google to pick a side and post. Your vote moves the story's Doom / Boom score.
Sources
3 articles from 3 outlets- Google NewsAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent - the-decoder.com
- The DecoderAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
- news.aibase.comAlibaba Qwen Launches Qwen-Audio-3.1: Five Speech Models Released Simultaneously, ASR Prices Drop by 95%

