activity
20172022
most citedThe Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap

27 citations · 61 across the 10 of their papers we have counts for

collaborators

14 papers

eess.AS2022

Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition

Zili Huang, Zhuo Chen, Naoyuki Kanda +6

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…

eess.AS20222 cited

Adapting self-supervised models to multi-talker speech recognition using speaker embeddings

Zili Huang, Desh Raj, Paola García +1

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-t…

cs.CL2022

SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning

Tzu-hsun Feng, Annie Dong, Ching-Feng Yeh +11

We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge buil…

eess.AS2022

Investigating self-supervised learning for speech enhancement and separation

Zili Huang, Shinji Watanabe, Shu-wen Yang +2

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target spe…

cs.CL20223 cited

SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang +14

Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years. In speech, a model pre-trained by self-supervi…

eess.AS20212 cited

Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker

Maokui He, Desh Raj, Zili Huang +3

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fi…