audio autoencoder 1cover song generation 1fast encoding 1large language models 1large-scale training 1low-bitrate compression 1music generation 1semantic tokenization 1text-to-audio generation 1text-to-music 1transformer architecture 1
From the 2 of 12 linked papers with an AI index.
1 citations · 1 across the 7 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2026
Qwen-Audio-VAE Technical Report
Ziyue Jiang, Dake Guo, Zekai Zhang +11
Qwen-Audio-VAE is a low‑bitrate, fast‑encoding continuous audio autoencoder that produces compact latent representations for scalable text‑to‑audio generation, using a causal encod…
eess.AS2025
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
Shengpeng Ji, Tianle Liang, Yangzhuo Li +11
End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' convers…