Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Yi Luo, Rongzhi Gu, Jixun Yao
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with highe…
eess.AS2024
Gull: A Generative Multifunctional Audio Codec
Yi Luo, Jianwei Yu, Hangting Chen +2
We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of task…
eess.AS2024
The Sound Demixing Challenge 2023 $\unicode{x2013}$ Cinematic Demixing Track
Stefan Uhlich, Giorgio Fabbro, Masato Hirano +14
This paper summarizes the cinematic demixing (CDX) track of the Sound Demixing Challenge 2023 (SDX'23). We provide a comprehensive summary of the challenge setup, detailing the str…