activity
20182022
most citedMinimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition

16 citations · 101 across the 22 of their papers we have counts for

collaborators
Showing eess.ASShow all

12 papers · 1 filter

eess.AS2022

The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

Naijun Zheng, Na Li, Xixin Wu +6

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in…

eess.AS2021

Towards Robust Speaker Verification with Target Speaker Enhancement

Chunlei Zhang, Meng Yu, Chao Weng +1

This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embeddin…

eess.AS2021

TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation

Helin Wang, Bo Wu, Lianwu Chen +7

In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose…

eess.AS20202 cited

Self-supervised Text-independent Speaker Verification using Prototypical Momentum Contrastive Learning

Wei Xia, Chunlei Zhang, Chao Weng +2

In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a moment…

eess.AS20204 cited

Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization

Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +4

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which…

eess.AS2020

Replay and Synthetic Speech Detection with Res2net Architecture

Xu Li, Na Li, Chao Weng +4

Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-cal…