84 citations · 116 across the 3 of their papers we have counts for
9 papers
From General to Specific: Informative Scene Graph Generation via Balance Adjustment
Yuyu Guo, Lianli Gao, Xuanhan Wang +5
The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene und…
Feature Space Targeted Attacks by Statistic Alignment
Lianli Gao, Yaya Cheng, Qilong Zhang +2
By adding human-imperceptible perturbations to images, DNNs can be easily fooled. As one of the mainstream methods, feature space targeted attacks perturb images by modulating thei…
Universal Weighting Metric Learning for Cross-Modal Matching
Jiwei Wei, Xing Xu, Yang Yang +3
Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial…
Cooperative Cross-Stream Network for Discriminative Action Representation
Jingran Zhang, Fumin Shen, Xing Xu +1
Spatial and temporal stream model has gained great success in video action recognition. Most existing works pay more attention to designing effective features fusion methods, which…
Temporal Reasoning Graph for Activity Recognition
Jingran Zhang, Fumin Shen, Xing Xu +1
Despite great success has been achieved in activity analysis, it still has many challenges. Most existing work in activity recognition pay more attention to design efficient archit…
Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
Tan Wang, Xing Xu, Yang Yang +3
A major challenge in matching images and text is that they have intrinsically different data distributions and feature representations. Most existing approaches are based either on…