Qwen-Audio 3.1: five speech models, up to 95% cheaper
TL;DR
- Alibaba's Qwen team released Qwen-Audio-3.1, five models for speech recognition, text-to-speech and real-time voice, The Decoder reported on 23 September.
- TTS prices fell about 70 percent, real-time roughly 85 percent and ASR up to 95 percent.
- The file transcription model is listed on Qwen Cloud at $0.15 per 1M input tokens and $0.47 per 1M output tokens.