2 papers
cs.SD2025
SpectroStream: A Versatile Neural Codec for General Audio
Yunpeng Li, Kehang Han, Brian McWilliams +2
We propose SpectroStream, a full-band multi-channel neural audio codec. Successor to the well-established SoundStream, SpectroStream extends its capability beyond 24 kHz monophonic…
eess.AS2025
MAD Speech: Measures of Acoustic Diversity of Speech
Matthieu Futeral, Andrea Agostinelli, Marco Tagliasacchi +2
Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the e…