From the 2 of 8 linked papers with an AI index.
8 papers
Qwen-Audio-VAE Technical Report
Ziyue Jiang, Dake Guo, Zekai Zhang +11
Qwen-Audio-VAE is a low‑bitrate, fast‑encoding continuous audio autoencoder that produces compact latent representations for scalable text‑to‑audio generation, using a causal encod…
Qwen-Music Technical Report
Jin Xu, Kangdi Wang, Ruibin Yuan +24
Qwen-Music is a large language model‑based system that generates high‑fidelity songs with vocals from text prompts or re‑imagines existing tracks, using a semantic token representa…
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
Bingshen Mu, Xian Shi, Xiong Wang +3
Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The m…
Qwen3-TTS Technical Report
Hangrui Hu, Xinfa Zhu, Ting He +13
In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3…
Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
He Wang, Linhan Ma, Dake Guo +4
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…