activity
20182022
most citedOn Psychoacoustically Weighted Cost Functions Towards Resource-Efficient Deep Neural Networks for Speech Denoising

3 citations · 8 across the 6 of their papers we have counts for

collaborators

9 papers

eess.AS2022

The Potential of Neural Speech Synthesis-based Data Augmentation for Personalized Speech Enhancement

Anastasia Kuznetsova, Aswin Sivaraman, Minje Kim

With the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic me…

cs.SD2021

Adapting Speech Separation to Real-World Meetings Using Mixture Invariant Training

Aswin Sivaraman, Scott Wisdom, Hakan Erdogan +1

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models in the sense that it does not require ground-…

eess.AS2021

Zero-Shot Personalized Speech Enhancement through Speaker-Informed Model Selection

Aswin Sivaraman, Minje Kim

This paper presents a novel zero-shot learning approach towards personalized speech enhancement through the use of a sparsely active ensemble model. Optimizing speech denoising sys…

eess.AS20212 cited

Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification

Aswin Sivaraman, Sunwoo Kim, Minje Kim

Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If t…

cs.CL2021

Detecting Extraneous Content in Podcasts

Sravana Reddy, Yongze Yu, Aasish Pappu +3

Podcast episodes often contain material extraneous to the main content, such as advertisements, interleaved within the audio and the written descriptions. We present classifiers th…

eess.AS20201 cited

Sparse Mixture of Local Experts for Efficient Speech Enhancement

Aswin Sivaraman, Minje Kim

In this paper, we investigate a deep learning approach for speech denoising through an efficient ensemble of specialist neural networks. By splitting up the speech denoising task i…