2.2k citations · 6.2k across the 56 of their papers we have counts for
8 papers · 1 filter
CgT-GAN: CLIP-guided Text GAN for Image Captioning
Jiarui Yu, Haoran Li, Yanbin Hao +3
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated…
Bi-directional Distribution Alignment for Transductive Zero-Shot Learning
Zhicai Wang, Yanbin Hao, Tingting Mu +3
It is well-known that zero-shot learning (ZSL) can suffer severely from the problem of domain shift, where the true and learned data distributions for the unseen classes do not mat…
Attention in Attention: Modeling Context Correlation for Efficient Video Classification
Yanbin Hao, Shuo Wang, Pei Cao +4
Attention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current rese…
Group Contextualization for Video Recognition
Yanbin Hao, Hao Zhang, Chong-Wah Ngo +1
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational uni…
A-FMI: Learning Attributions from Deep Networks via Feature Map Importance
An Zhang, Xiang Wang, Chengfang Fang +3
Gradient-based attribution methods can aid in the understanding of convolutional neural networks (CNNs). However, the redundancy of attribution features and the gradient saturation…
Learning to Compose and Reason with Language Tree Structures for Visual Grounding
Richang Hong, Daqing Liu, Xiaoyu Mo +2
Grounding natural language in images, such as localizing "the black dog on the left of the tree", is one of the core problems in artificial intelligence, as it needs to comprehend…