Alibaba Ships Qwen-Audio 3.1, Cuts Voice API Prices By Up To 95%

Alibaba's five-model Qwen-Audio 3.1 stack undercuts ElevenLabs and OpenAI Realtime pricing, adds ASR-Next speaker ID and TTS-Next single-pass mixed audio.

Alibaba Ships Qwen-Audio 3.1, Cuts Voice API Prices By Up To 95%

Alibaba Cloud has shipped Qwen-Audio 3.1, a five-model voice stack that slashes API pricing by up to 95% and adds new "Next" models for both speech recognition and speech synthesis. Released on September 24, 2026, the upgrade is Alibaba's most aggressive push yet to undercut ElevenLabs, OpenAI Realtime and Google's voice APIs — and it lands the same week China's biggest cloud provider unveiled the Zhenwu V900 AI chip.

The Five Models

The stack now includes an upgraded ASR that strips filler words and repetitions across dozens of languages and dialects; ASR-Next, which layers speaker identification, timestamps, emotion detection and ambient sound classification on top; TTS with text-prompt-driven emotion, speed and style; TTS-Next, generating "voice, sound effects and background audio in a single pass"; and the upgraded Realtime, which listens while speaking and can be interrupted mid-sentence. A separate Realtime Plus variant released on September 20 pushes the context window to 262,144 tokens for long-running voice agents that need to keep session state.

Pricing Shock

The headline is on the invoice. Alibaba is cutting ASR pricing by up to 95%, TTS by around 70%, and Realtime by roughly 85%. On a per-minute basis, this puts Qwen-Audio's ASR at a fraction of ElevenLabs' equivalent tier — although ElevenLabs still covers 74 languages against TTS-Next's 16. For enterprises building call-center agents or accessibility apps, Alibaba is betting that the price gap will pull workloads away from Western incumbents faster than the language gap can push them back.

Qwen-Audio 3.1 real-time voice agent architecture diagram

Where It Slots In

Qwen-Audio 3.1 is the voice complement to Alibaba's broader agent stack — Qwen3.8-Omni-Flash for text/vision/audio agents, and AgentCore for enterprise deployment. Combined, Alibaba is packaging voice agents that can be spun up entirely on its own infrastructure, in Chinese and English, at price points designed for high-volume conversational deployment.

What This Signals

Alibaba's play mirrors what Chinese labs have already done to text LLM pricing: race the cost curve down while Western labs focus on capability. With DeepSeek, Xiaomi MiMo and Alibaba all pushing multimodal pricing near zero, the pressure now sits on OpenAI's Realtime and Google's Gemini Live tiers to defend margins.

Reporting based on coverage from AI Weekly, Superpower Daily, AlphaSignal and Data Studios.

Category: Natural Language Processing

Tags: AI Alibaba Natural Language Processing China Generative AI AI Agents

Related Articles