72 citations · 134 across the 6 of their papers we have counts for
8 papers
Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation
Xingning Dong, Tian Gan, Xuemeng Song +3
Scene Graph Generation, which generally follows a regular encoder-decoder pipeline, aims to first encode the visual contents within the given image and then parse them into a compa…
MERIt: Meta-Path Guided Contrastive Learning for Logical Reasoning
Fangkai Jiao, Yangyang Guo, Xuemeng Song +1
Logical reasoning is of vital importance to natural language understanding. Previous studies either employ graph-based models to incorporate prior knowledge about logical relations…
Hierarchical Deep Residual Reasoning for Temporal Moment Localization
Ziyang Ma, Xianjing Han, Xuemeng Song +2
Temporal Moment Localization (TML) in untrimmed videos is a challenging task in the field of multimedia, which aims at localizing the start and end points of the activity in the vi…
Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos
Zongmeng Zhang, Xianjing Han, Xuemeng Song +2
This paper focuses on tackling the problem of temporal language localization in videos, which aims to identify the start and end points of a moment described by a natural language…
Answer Questions with Right Image Regions: A Visual Attention Regularization Approach
Yibing Liu, Yangyang Guo, Jianhua Yin +3
Visual attention in Visual Question Answering (VQA) targets at locating the right image regions regarding the answer prediction, offering a powerful technique to promote multi-moda…
Market2Dish: Health-aware Food Recommendation
Wenjie Wang, Ling-yu Duan, Hao Jiang +3
With the rising incidence of some diseases, such as obesity and diabetes, a healthy diet is arousing increasing attention. However, most existing food-related research efforts focu…