audio autoencoder 1fast encoding 1large-scale training 1low-bitrate compression 1text-to-audio generation 1transformer architecture 1
From the 1 of 5 linked papers with an AI index.
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2026
DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios
Bingshen Mu, Xian Shi, Xiong Wang +6
Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who s…
eess.AS2026
Qwen-Audio-VAE Technical Report
Ziyue Jiang, Dake Guo, Zekai Zhang +11
Qwen-Audio-VAE is a low‑bitrate, fast‑encoding continuous audio autoencoder that produces compact latent representations for scalable text‑to‑audio generation, using a causal encod…