activity
20172026
most citedSpeechBrain: A General-Purpose Speech Toolkit

514 citations · 548 across the 45 of their papers we have counts for

collaborators
Showing cs.SDShow all

21 papers · 1 filter

cs.SD2026

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

Luca Della Libera, Cem Subakan, Mirco Ravanelli

Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging,…

cs.SD2026

Exploring Token-Space Manipulation in Latent Audio Tokenizers

Francesco Paissan, Luca Della Libera, Mirco Ravanelli +1

Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it diffic…

cs.SD2026

Listen First, Then Answer: Timestamp-Grounded Speech Reasoning

Jihoon Jeong, Pooneh Mousavi, Mirco Ravanelli +1

Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. I…

cs.SD2026

LL-SDR: Low-Latency Speech enhancement through Discrete Representations

Jingyi Li, Luca Della Libera, Mirco Ravanelli +2

Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it…

cs.SD2026

Toward Faithful Explanations in Acoustic Anomaly Detection

Maab Elrashid, Anthony Deschênes, Cem Subakan +3

Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency.…

cs.SD2025

Virtual Consistency for Audio Editing

Matthieu Cervera, Francesco Paissan, Mirco Ravanelli +1

Free-form, text-based audio editing remains a persistent challenge, despite progress in inversion-based neural methods. Current approaches rely on slow inversion procedures, limiti…