36 citations · 112 across the 37 of their papers we have counts for
19 papers · 1 filter
Codec Data Augmentation for Time-domain Heart Sound Classification
Ansh Mishra, Jia Qi Yip, Eng Siong Chng
Heart auscultations are a low-cost and effective way of detecting valvular heart diseases early, which can save lives. Nevertheless, it has been difficult to scale this screening m…
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition
Yuchen Hu, Chen Chen, Ruizhe Li +2
Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing…
Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
Yuchen Hu, Chen Chen, Heqing Zou +2
Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they woul…
Probabilistic Back-ends for Online Speaker Recognition and Clustering
Alexey Sholokhov, Nikita Kuzmin, Kong Aik Lee +1
This paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-end…
Learning Speaker Representation with Semi-supervised Learning approach for Speaker Profiling
Shangeth Rajaa, Pham Van Tung, Chng Eng Siong
Speaker profiling, which aims to estimate speaker characteristics such as age and height, has a wide range of applications inforensics, recommendation systems, etc. In this work, w…
A Unified Speaker Adaptation Approach for ASR
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung +3
Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mism…