32 citations · 43 across the 5 of their papers we have counts for
4 papers · 1 filter
Revisiting Pre-training in Audio-Visual Learning
Ruoxuan Feng, Wenke Xia, Di Hu
Pre-training technique has gained tremendous success in enhancing model performance on various tasks, but found to perform worse than training from scratch in some uni-modal situat…
Learning in Audio-visual Context: A Review, Analysis, and New Perspective
Yake Wei, Di Hu, Yapeng Tian +1
Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developin…
Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction
Yingzi Fan, Longfei Han, Yue Zhang +3
Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the aud…
Class-aware Sounding Objects Localization via Audiovisual Correspondence
Di Hu, Yake Wei, Rui Qian +3
Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achie…