1 citations · 1 across the 3 of their papers we have counts for
9 papers
Factorized RVQ-GAN For Disentangled Speech Tokenization
Sameer Khurana, Dominik Klement, Antoine Laurent +13
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…
Task-Aware Unified Source Separation
Kohei Saijo, Janek Ebbers, François G. Germain +2
Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or…
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
Kohei Saijo, Gordon Wichern, François G. Germain +2
Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
Kohei Saijo, Gordon Wichern, François G. Germain +2
Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…
Sound Event Bounding Boxes
Janek Ebbers, Francois G. Germain, Gordon Wichern +1
Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence con…
Why does music source separation benefit from cacophony?
Chang-Bin Jeon, Gordon Wichern, François G. Germain +1
In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These ran…