activity
20172026
most citedSpeechBrain: A General-Purpose Speech Toolkit

514 citations · 518 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization

Luca Della Libera, Cem Subakan, Mirco Ravanelli

Neural audio codecs are at the core of modern conversational speech technologies, converting continuous speech into sequences of discrete tokens that can be processed by LLMs. Howe…

cs.LG2025

Sample Compression for Self Certified Continual Learning

Jacob Comeau, Mathieu Bazinet, Pascal Germain +1

Continual learning algorithms aim to learn from a sequence of tasks. In order to avoid catastrophic forgetting, most existing approaches rely on heuristics and do not provide compu…

cs.LG2025

FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

Luca Della Libera, Francesco Paissan, Cem Subakan +1

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored a…

cs.LG2024

Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming

Shubham Gupta, Isaac Neri Gomez-Sarmiento, Faez Amjed Mezdari +2

We propose a novel approach for humming transcription that combines a CNN-based architecture with a dynamic programming-based post-processing algorithm, utilizing the recently intr…

cs.LG2018

Learning the Base Distribution in Implicit Generative Models

Cem Subakan, Oluwasanmi Koyejo, Paris Smaragdis

Popular generative model learning methods such as Generative Adversarial Networks (GANs), and Variational Autoencoders (VAE) enforce the latent representation to follow simple dist…