activity
20222024
most citedVATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

36 citations · 109 across the 17 of their papers we have counts for

collaborators

17 papers

eess.AS2024

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

Qiushi Zhu, Jie Zhang, Yu Gu +2

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-fi…

eess.AS2023★ 1 cited

Rep2wav: Noise Robust text-to-speech Using self-supervised representations

Qiushi Zhu, Yu Gu, Rilin Chen +4

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from rea…

eess.AS2023

Noise-aware Speech Enhancement using Diffusion Probabilistic Model

Yuchen Hu, Chen Chen, Ruizhe Li +2

With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However…

eess.AS2023★ 1 cited

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

Yuchen Hu, Ruizhe Li, Chen Chen +3

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most exis…

eess.AS2023★ 2 cited

Eeg2vec: Self-Supervised Electroencephalographic Representation Learning

Qiushi Zhu, Xiaoying Zhao, Jie Zhang +3

Recently, many efforts have been made to explore how the brain processes speech using electroencephalographic (EEG) signals, where deep learning-based approaches were shown to be a…

eess.AS2023

BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker Conditions

Jie Zhang, Qing-Tian Xu, Qiu-Shi Zhu +1

Time-domain single-channel speech enhancement (SE) still remains challenging to extract the target speaker without any prior information on multi-talker conditions. It has been sho…