2 citations · 3 across the 4 of their papers we have counts for
8 papers
Learnable Front Ends Based on Temporal Modulation for Music Tagging
Yinghao Ma, Richard M. Stern
While end-to-end systems are becoming popular in auditory signal processing including automatic music tagging, models using raw audio as input needs a large amount of data and comp…
A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement
Tyler Vuong, Yangyang Xia, Richard M. Stern
We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a…
Learnable Spectro-temporal Receptive Fields for Robust Voice Type Discrimination
Tyler Vuong, Yangyang Xia, Richard Stern
Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording…
Non causal deep learning based dereverberation
Jorge Wuth, Richard M. Stern, Nestor Becerra Yoma
In this paper we demonstrate the effectiveness of non-causal context for mitigating the effects of reverberation in deep-learning-based automatic speech recognition (ASR) systems.…
On combining features for single-channel robust speech recognition in reverberant environments
José Novoa, Josué Fredes, Jorge Wuth +3
This paper addresses the combination of complementary parallel speech recognition systems to reduce the error rate of speech recognition systems operating in real highly-reverberan…
Weighted delay-and-sum beamforming guided by visual tracking for human-robot interaction
José Novoa, Rodrigo Mahu, Alejandro Díaz +3
This paper describes the integration of weighted delay-and-sum beamforming with speech source localization using image processing and robot head visual servoing for source tracking…