7 papers
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
Qiushi Zhu, Jie Zhang, Yu Gu +2
Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-fi…
Rep2wav: Noise Robust text-to-speech Using self-supervised representations
Qiushi Zhu, Yu Gu, Rilin Chen +4
Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from rea…
Eeg2vec: Self-Supervised Electroencephalographic Representation Learning
Qiushi Zhu, Xiaoying Zhao, Jie Zhang +3
Recently, many efforts have been made to explore how the brain processes speech using electroencephalographic (EEG) signals, where deep learning-based approaches were shown to be a…
BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker Conditions
Jie Zhang, Qing-Tian Xu, Qiu-Shi Zhu +1
Time-domain single-channel speech enhancement (SE) still remains challenging to extract the target speaker without any prior information on multi-talker conditions. It has been sho…
Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
Yuchen Hu, Ruizhe Li, Chen Chen +3
Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-in…
Speech Enhancement with Multi-granularity Vector Quantization
Xiao-Ying Zhao, Qiu-Shi Zhu, Jie Zhang
With advances in deep learning, neural network based speech enhancement (SE) has developed rapidly in the last decade. Meanwhile, the self-supervised pre-trained model and vector q…