17 citations · 18 across the 3 of their papers we have counts for
4 papers
Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems
Yangyang Xia, Buye Xu, Anurag Kumar
Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-worl…
A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement
Tyler Vuong, Yangyang Xia, Richard M. Stern
We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a…
Learnable Spectro-temporal Receptive Fields for Robust Voice Type Discrimination
Tyler Vuong, Yangyang Xia, Richard Stern
Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording…
Weighted Speech Distortion Losses for Neural-network-based Real-time Speech Enhancement
Yangyang Xia, Sebastian Braun, Chandan K. A. Reddy +3
This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channe…