158 citations · 674 across the 63 of their papers we have counts for
11 papers · 1 filter
Rep2wav: Noise Robust text-to-speech Using self-supervised representations
Qiushi Zhu, Yu Gu, Rilin Chen +4
Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from rea…
Eeg2vec: Self-Supervised Electroencephalographic Representation Learning
Qiushi Zhu, Xiaoying Zhao, Jie Zhang +3
Recently, many efforts have been made to explore how the brain processes speech using electroencephalographic (EEG) signals, where deep learning-based approaches were shown to be a…
CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen +5
Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo +4
For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach…
BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker Conditions
Jie Zhang, Qing-Tian Xu, Qiu-Shi Zhu +1
Time-domain single-channel speech enhancement (SE) still remains challenging to extract the target speaker without any prior information on multi-talker conditions. It has been sho…
Speech Enhancement with Multi-granularity Vector Quantization
Xiao-Ying Zhao, Qiu-Shi Zhu, Jie Zhang
With advances in deep learning, neural network based speech enhancement (SE) has developed rapidly in the last decade. Meanwhile, the self-supervised pre-trained model and vector q…