Qwen-Audio 3.1: five speech models, up to 95% cheaper

TL;DR

  • Alibaba's Qwen team released Qwen-Audio-3.1, five models for speech recognition, text-to-speech and real-time voice, The Decoder reported on 23 September.
  • TTS prices fell about 70 percent, real-time roughly 85 percent and ASR up to 95 percent.
  • The file transcription model is listed on Qwen Cloud at $0.15 per 1M input tokens and $0.47 per 1M output tokens.

Read the full story

Sign in with your email to read AI News. It’s free.

Share this story

Explain like I’m 15