3 citations · 8 across the 6 of their papers we have counts for
9 papers
Spatial Processing Front-End For Distant ASR Exploiting Self-Attention Channel Combinator
Dushyant Sharma, Rong Gong, James Fosburgh +3
We present a novel multi-channel front-end based on channel shortening with theWeighted Prediction Error (WPE) method followed by a fixed MVDR beamformer used in combination with a…
Self-Attention Channel Combinator Frontend for End-to-End Multichannel Far-field Speech Recognition
Rong Gong, Carl Quillen, Dushyant Sharma +3
When a sufficiently large far-field training data is presented, jointly optimizing a multichannel frontend and an end-to-end (E2E) Automatic Speech Recognition (ASR) backend shows…
A Simple Fusion of Deep and Shallow Learning for Acoustic Scene Classification
Eduardo Fonseca, Rong Gong, Xavier Serra
In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data drive…
Towards an efficient deep learning model for musical onset detection
Rong Gong, Xavier Serra
In this paper, we propose an efficient and reproducible deep learning model for musical onset detection (MOD). We first review the state-of-the-art deep learning models for MOD, an…
Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions
Rong Gong, Xavier Serra
In this paper, we tackle the singing voice phoneme segmentation problem in the singing training scenario by using language-independent information -- onset and prior coarse duratio…
Identification of potential Music Information Retrieval technologies for computer-aided jingju singing training
Rong Gong, Xavier Serra
Music Information Retrieval (MIR) technologies have been proven useful in assisting western classical singing training. Jingju (also known as Beijing or Peking opera) singing is di…