45 citations · 124 across the 5 of their papers we have counts for
6 papers · 1 filter
Inception Transformer
Chenyang Si, Weihao Yu, Pan Zhou +3
Recent studies show that Transformer has strong capability of building long-range dependencies, yet is incompetent in capturing high frequencies that predominantly convey local inf…
Mugs: A Multi-Granular Self-Supervised Learning Framework
Pan Zhou, Yichen Zhou, Chenyang Si +3
In self-supervised learning, multi-granular features are heavily desired though rarely investigated, as different downstream tasks (e.g., general and fine-grained classification) o…
Refiner: Refining Self-attention for Vision Transformers
Daquan Zhou, Yujun Shi, Bingyi Kang +6
Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most…
Heterogeneous Graph Learning for Visual Commonsense Reasoning
Weijiang Yu, Jingwen Zhou, Weihao Yu +2
Visual commonsense reasoning task aims at leading the research field into solving cognition-level reasoning with the ability of predicting correct answers and meanwhile providing c…
Knowledge-Embedded Routing Network for Scene Graph Generation
Tianshui Chen, Weihao Yu, Riquan Chen +1
To understand a scene in depth not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since t…
Deep Reasoning with Knowledge Graph for Social Relationship Understanding
Zhouxia Wang, Tianshui Chen, Jimmy Ren +3
Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…