1 citations · 1 across the 13 of their papers we have counts for
15 papers · 1 filter
Exploring Token-Space Manipulation in Latent Audio Tokenizers
Francesco Paissan, Luca Della Libera, Mirco Ravanelli +1
Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it diffic…
DASB - Discrete Audio and Speech Benchmark
Pooneh Mousavi, Jarod Duret, Darius Petermann +5
Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both gen…
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
Jihoon Jeong, Pooneh Mousavi, Mirco Ravanelli +1
Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. I…
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
Jingyi Li, Luca Della Libera, Mirco Ravanelli +2
Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it…
Toward Faithful Explanations in Acoustic Anomaly Detection
Maab Elrashid, Anthony Deschênes, Cem Subakan +3
Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency.…
Discrete Audio Tokens: More Than a Survey!
Pooneh Mousavi, Gallil Maimon, Adel Moumen +18
Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and infere…