Qwen-Audio-3.1 logo

Qwen-Audio-3.1

Visit

Alibaba's Qwen-Audio-3.1 stack adds ASR-Next and TTS-Next, with TTS topping the Artificial Analysis TTS leaderboard and API prices cut up to 95%.

Share:
View alternatives

Qwen-Audio-3.1

Qwen-Audio-3.1 is Alibaba's 2026 refresh of the Qwen audio family, announced from the Qwen team account on 2026-09-23. It is not one checkpoint but a five-model stack. The existing ASR, TTS and Realtime products get upgrades, and two new models join them: ASR-Next for audio understanding and TTS-Next for audio creation.

The headline for most teams is price. Alibaba announced cuts of roughly 70% on TTS, 85% on Realtime and up to 95% on ASR across the family, which matters for anyone paying per minute of speech rather than per seat.

What ships in the stack

Model Role API id on Qwen Cloud
ASR Multilingual and dialect transcription, with native polishing that strips fillers and repetitions qwen-audio-3.1-asr-flash-filetrans
ASR-Next Multi-speaker ASR with speaker labels, timestamps and aligned transcripts, plus emotion and sound captioning announced, API coming soon
TTS Instruction-controlled synthesis with cross-lingual voice transfer qwen-audio-3.1-tts-flash
TTS-Next Unified LM plus diffusion pass for voice, sound effects and background audio announced, API coming soon
Realtime Full-duplex speech model with interruption handling and empathetic pacing qwen-audio-3.1-realtime-plus

The TTS model in detail

The official technical overview describes Qwen-Audio-3.1-TTS as a production-oriented synthesis system built on a 12.5 Hz low-frame-rate speech tokenizer and a five-stage training pipeline: separate language-model and flow-matching pretraining, joint training with high-quality data annealing, language-model reinforcement learning, flow-matching robustness training, and a final reinforcement-learning pass.

Control comes from free-style natural-language instructions plus 86 inline tags for phrase and word level direction, including non-verbal events such as laughter, breathing and coughing. Coverage is 16 languages, seven of them new in this release, and 20 Chinese dialect regions. One-pass long-form synthesis reaches 3 minutes, output runs up to 48 kHz, and the model is built to keep working from noisy, reverberant or unclear reference speech without a separate denoising step. Alibaba states it ranks first on the independent Artificial Analysis Text-to-Speech leaderboard.

Pricing

Rates below are from the Qwen Cloud model pages, checked on 2026-09-23.

Model Input Output Rate limit
ASR-Flash-Filetrans $0.15 / 1M tokens $0.47 / 1M tokens 600 RPM
TTS-Flash $0.23 / 1M tokens $1.87 / 1M tokens 180 RPM
Realtime-Plus published on the Qwen Cloud pricing pages published on the Qwen Cloud pricing pages see the model page

Limitations

  • Two of the five models are announcement-only. ASR-Next and TTS-Next had no public API id or price at check time, and the team said more API access is coming.
  • No open weights in this revision. The 3.1 audio family is hosted only. Open audio weights from Alibaba stop at the earlier Qwen-Omni generation.
  • Dialect support is China-centric. The 20 dialect regions are Chinese; other coverage follows the 16-language list.
  • The leaderboard claim is vendor-cited. The Artificial Analysis position is reported by Alibaba, so verify it against the live board for your own language pair.

FAQ

Is Qwen-Audio-3.1 open source?

No. The 3.1 audio family runs as a hosted API on Qwen Cloud. For an open-weight TTS model in a comparable tier, look at Breeze TTS 2.

What does ASR-Next add over plain ASR?

Speaker labels, timestamps and aligned transcripts, plus emotion, ambient and machine-sound understanding for sound captioning, event localization and audio question answering.

Can I run it locally?

Not from these checkpoints. sanoTTS solves a different problem entirely: a sub-3M-parameter model that runs in a browser or on an ESP32 board.

Alternatives

Comments

No comments yet. Be the first to comment!