activity
20162022
most citedTime-domain speaker extraction network

36 citations · 105 across the 24 of their papers we have counts for

collaborators

31 papers

cs.SD2022

I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization

Dianwen Ng, Jia Qi Yip, Tanmay Surana +6

Noise robustness in keyword spotting remains a challenge as many models fail to overcome the heavy influence of noises, causing the deterioration of the quality of feature embeddin…

cs.CL2022

Self-critical Sequence Training for Automatic Speech Recognition

Chen Chen, Yuchen Hu, Nana Hou +3

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that m…

cs.SD20229 cited

Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning

Chen Chen, Nana Hou, Yuchen Hu +3

Automated Audio captioning (AAC) is a cross-modal task that generates natural language to describe the content of input audio. Most prior works usually extract single-modality acou…

cs.SD2022

Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information

Heqing Zou, Yuke Si, Chen Chen +2

Speech Emotion Recognition (SER) aims to help the machine to understand human's subjective emotion from only audio information. However, extracting and utilizing comprehensive in-d…

cs.SD2022

Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data

Chen Chen, Nana Hou, Yuchen Hu +2

Noise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in…

cs.SD2022

Estimation of speaker age and height from speech signal using bi-encoder transformer mixture model

Tarun Gupta, Duc-Tuan Truong, Tran The Anh +1

The estimation of speaker characteristics such as age and height is a challenging task, having numerous applications in voice forensic analysis. In this work, we propose a bi-encod…