1 citations · 1 across the 8 of their papers we have counts for
11 papers
ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding
Luca Della Libera, Cem Subakan, Mirco Ravanelli
Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging,…
Exploring Token-Space Manipulation in Latent Audio Tokenizers
Francesco Paissan, Luca Della Libera, Mirco Ravanelli +1
Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it diffic…
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
Jingyi Li, Luca Della Libera, Mirco Ravanelli +2
Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it…
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
Luca Della Libera, Cem Subakan, Mirco Ravanelli
Large language models show that simple autoregressive training can yield scalable and coherent generation, but extending this paradigm to speech remains challenging due to the enta…
How to Choose a Reinforcement-Learning Algorithm
Fabian Bongratz, Vladimir Golkov, Lukas Mautner +5
The field of reinforcement learning offers a large variety of concepts and methods to tackle sequential decision-making problems. This variety has become so large that choosing an…
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
Pooneh Mousavi, Jarod Duret, Salah Zaiem +4
Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paraling…