12 citations · 38 across the 13 of their papers we have counts for
15 papers · 1 filter
VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion
Disong Wang, Shan Yang, Dan Su +3
Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speec…
The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge
Naijun Zheng, Na Li, Xixin Wu +6
This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in…
DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion
Songxiang Liu, Yuewen Cao, Dan Su +1
Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and exp…
FastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulation
Songxiang Liu, Yuewen Cao, Na Hu +2
This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than r…
Replay and Synthetic Speech Detection with Res2net Architecture
Xu Li, Na Li, Chao Weng +4
Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-cal…
Distortionless Multi-Channel Target Speech Enhancement for Overlapped Speech Recognition
Bo Wu, Meng Yu, Lianwu Chen +4
Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measur…