Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings
The Decoder reported on July 21, 2026 that Alibaba's Qwen-Audio-3.0-TTS-Plus took first place in Artificial Analysis' Speech Arena leaderboard for provider voices. The model supports 16 languages, style control through natural language or tags such as [angry] and [giggles], and improved handling of noisy or echo-heavy reference recordings for voice cloning. For AI characters and VTubers, it expands options for multilingual and expressive voice generation, although its speed remains a weakness compared with some rivals.

The interesting part for AI characters is not only the voice quality, but the ability to steer nonverbal cues like [giggles]. The real craft will be deciding how finely operators shape those emotional reactions.