works on

From the 2 of 9 linked papers with an AI index.

collaborators

9 papers

eess.AS2026

Qwen-Audio-VAE Technical Report

Ziyue Jiang, Dake Guo, Zekai Zhang +11

Qwen-Audio-VAE is a low‑bitrate, fast‑encoding continuous audio autoencoder that produces compact latent representations for scalable text‑to‑audio generation, using a causal encod…

cs.SD2026

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24

Qwen-Music is a large language model‑based system that generates high‑fidelity songs with vocals from text prompts or re‑imagines existing tracks, using a semantic token representa…

cs.SD2026

Qwen3-TTS Technical Report

Hangrui Hu, Xinfa Zhu, Ting He +13

In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3…

cs.SD2025

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Yujia Xiao, Liumeng Xue, Lei He +8

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into…

eess.AS2025

KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction

Kangxiang Xia, Xinfa Zhu, Jixun Yao +3

We introduce KALL-E, a novel autoregressive (AR) language model for text-to-speech (TTS) synthesis that operates by predicting the next distribution of continuous speech frames. Un…

cs.SD2025

Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis

Wenjie Tian, Xinfa Zhu, Hanke Xie +3

Recent progress in text-to-speech (TTS) has achieved impressive naturalness and flexibility, especially with the development of large language model (LLM)-based approaches. However…