activity
20162023
most citedTime-domain speaker extraction network

36 citations · 112 across the 37 of their papers we have counts for

collaborators
Showing eess.ASShow all

19 papers · 1 filter

eess.AS2023

Codec Data Augmentation for Time-domain Heart Sound Classification

Ansh Mishra, Jia Qi Yip, Eng Siong Chng

Heart auscultations are a low-cost and effective way of detecting valvular heart diseases early, which can save lives. Nevertheless, it has been difficult to scale this screening m…

eess.AS2023

MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition

Yuchen Hu, Chen Chen, Ruizhe Li +2

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing…

eess.AS2023★ 1 cited

Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation

Yuchen Hu, Chen Chen, Heqing Zou +2

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they woul…

eess.AS2023

Probabilistic Back-ends for Online Speaker Recognition and Clustering

Alexey Sholokhov, Nikita Kuzmin, Kong Aik Lee +1

This paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-end…

eess.AS2021★ 2 cited

Learning Speaker Representation with Semi-supervised Learning approach for Speaker Profiling

Shangeth Rajaa, Pham Van Tung, Chng Eng Siong

Speaker profiling, which aims to estimate speaker characteristics such as age and height, has a wide range of applications inforensics, recommendation systems, etc. In this work, w…

eess.AS2021

A Unified Speaker Adaptation Approach for ASR

Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung +3

Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mism…