Qwen-Audio-3.1
Qwen-Audio-3.1 is Alibaba's 2026 refresh of the Qwen audio family, announced from the Qwen team account on 2026-09-23. It is not one checkpoint but a five-model stack. The existing ASR, TTS and Realtime products get upgrades, and two new models join them: ASR-Next for audio understanding and TTS-Next for audio creation.
The headline for most teams is price. Alibaba announced cuts of roughly 70% on TTS, 85% on Realtime and up to 95% on ASR across the family, which matters for anyone paying per minute of speech rather than per seat.
What ships in the stack
| Model | Role | API id on Qwen Cloud |
|---|---|---|
| ASR | Multilingual and dialect transcription, with native polishing that strips fillers and repetitions | qwen-audio-3.1-asr-flash-filetrans |
| ASR-Next | Multi-speaker ASR with speaker labels, timestamps and aligned transcripts, plus emotion and sound captioning | announced, API coming soon |
| TTS | Instruction-controlled synthesis with cross-lingual voice transfer | qwen-audio-3.1-tts-flash |
| TTS-Next | Unified LM plus diffusion pass for voice, sound effects and background audio | announced, API coming soon |
| Realtime | Full-duplex speech model with interruption handling and empathetic pacing | qwen-audio-3.1-realtime-plus |
The TTS model in detail
The official technical overview describes Qwen-Audio-3.1-TTS as a production-oriented synthesis system built on a 12.5 Hz low-frame-rate speech tokenizer and a five-stage training pipeline: separate language-model and flow-matching pretraining, joint training with high-quality data annealing, language-model reinforcement learning, flow-matching robustness training, and a final reinforcement-learning pass.
Control comes from free-style natural-language instructions plus 86 inline tags for phrase and word level direction, including non-verbal events such as laughter, breathing and coughing. Coverage is 16 languages, seven of them new in this release, and 20 Chinese dialect regions. One-pass long-form synthesis reaches 3 minutes, output runs up to 48 kHz, and the model is built to keep working from noisy, reverberant or unclear reference speech without a separate denoising step. Alibaba states it ranks first on the independent Artificial Analysis Text-to-Speech leaderboard.
Pricing
Rates below are from the Qwen Cloud model pages, checked on 2026-09-23.
| Model | Input | Output | Rate limit |
|---|---|---|---|
| ASR-Flash-Filetrans | $0.15 / 1M tokens | $0.47 / 1M tokens | 600 RPM |
| TTS-Flash | $0.23 / 1M tokens | $1.87 / 1M tokens | 180 RPM |
| Realtime-Plus | published on the Qwen Cloud pricing pages | published on the Qwen Cloud pricing pages | see the model page |
Limitations
- Two of the five models are announcement-only. ASR-Next and TTS-Next had no public API id or price at check time, and the team said more API access is coming.
- No open weights in this revision. The 3.1 audio family is hosted only. Open audio weights from Alibaba stop at the earlier Qwen-Omni generation.
- Dialect support is China-centric. The 20 dialect regions are Chinese; other coverage follows the 16-language list.
- The leaderboard claim is vendor-cited. The Artificial Analysis position is reported by Alibaba, so verify it against the live board for your own language pair.
FAQ
Is Qwen-Audio-3.1 open source?
No. The 3.1 audio family runs as a hosted API on Qwen Cloud. For an open-weight TTS model in a comparable tier, look at Breeze TTS 2.
What does ASR-Next add over plain ASR?
Speaker labels, timestamps and aligned transcripts, plus emotion, ambient and machine-sound understanding for sound captioning, event localization and audio question answering.
Can I run it locally?
Not from these checkpoints. sanoTTS solves a different problem entirely: a sub-3M-parameter model that runs in a browser or on an ESP32 board.
Alternatives
- Breeze TTS 2: open-weight multilingual TTS with reference-free voice design.
- ElevenLabs Turbo v2.5: hosted TTS from a dedicated voice vendor.
- Deepgram Nova-2: hosted ASR with broad language coverage.
- Qwen3.8-Omni-Flash: the omni-modal sibling when you want audio, video and text in one model instead of a pipeline of specialists.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Obsidian CLI + Codex: Turn Your Vault into an Agent Knowledge Engine
Obsidian CLI gives Codex and other agents a searchable, auditable, link-aware interface to a local Vault, using real cases and reproducible workflows.

Grok Bot and Hermes Bot: one person finally gets a think tank and a secretariat
Grok Bot now ships with Cursor Pro+. Hermes Bot runs on a VPS. They are not smarter chat boxes. The think tank advises, the secretariat executes, and you still make the call.
Hook OpenCode Go into Codex on Windows. Do Not Open a Second Toolkit.
A ChatGPT-signed Codex desktop app still shows mostly GPT in the picker. On Windows, enable only OpenCode Go and the Grok, GLM, Kimi, DeepSeek, and MiniMax models you already pay for appear in the same selector. Keys stay local. Native GPT stays put.