activity
20182022
most citedRecent Developments on ESPnet Toolkit Boosted by Conformer

40 citations · 83 across the 14 of their papers we have counts for

collaborators

16 papers

cs.SD20223 cited

TESSP: Text-Enhanced Self-Supervised Speech Pre-training

Zhuoyuan Yao, Shuo Ren, Sanyuan Chen +3

Self-supervised speech pre-training empowers the model with the contextual structure inherent in the speech signal while self-supervised text pre-training empowers the model with l…

eess.AS2022

Distinguishable Speaker Anonymization based on Formant and Fundamental Frequency Scaling

Jixun Yao, Qing Wang, Yi Lei +4

Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy conc…

eess.AS2022

Preserving background sound in noise-robust voice conversion via multi-task learning

Jixun Yao, Yi Lei, Qing Wang +6

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research…

cs.SD20221 cited

MFCCA:Multi-Frame Cross-Channel attention for multi-speaker ASR in Multi-party meeting scenario

Fan Yu, Shiliang Zhang, Pengcheng Guo +4

Recently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel…

eess.AS20225 cited

NWPU-ASLP System for the VoicePrivacy 2022 Challenge

Jixun Yao, Qing Wang, Li Zhang +3

This paper presents the NWPU-ASLP speaker anonymization system for VoicePrivacy 2022 Challenge. Our submission does not involve additional Automatic Speaker Verification (ASV) mode…

cs.SD2022

Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

Fan Yu, Shiliang Zhang, Pengcheng Guo +13

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologie…