audio autoencoder 1cover song generation 1fast encoding 1large language models 1large-scale training 1low-bitrate compression 1music generation 1semantic tokenization 1text-to-audio generation 1text-to-music 1transformer architecture 1
From the 2 of 12 linked papers with an AI index.
1 citations · 1 across the 7 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…
cs.CL2025
Qwen2.5-Omni Technical Report
Jin Xu, Zhifang Guo, Jinzheng He +11
In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously gene…