21 citations · 37 across the 4 of their papers we have counts for
4 papers
speech and noise dual-stream spectrogram refine network with speech distortion loss for robust speech recognition
Haoyu Lu, Nan Li, Tongtong Song +4
In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. T…
UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
Haoyu Lu, Yuqi Huo, Guoxing Yang +4
Large-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream…
LGDN: Language-Guided Denoising Network for Video-Language Modeling
Haoyu Lu, Mingyu Ding, Nanyi Fei +2
Video-language modeling has attracted much attention with the rapid growth of web videos. Most existing methods assume that the video frames and text description are semantically c…
Multimodal foundation models are better simulators of the human brain
Haoyu Lu, Qiongyi Zhou, Nanyi Fei +8
Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI).…