119 citations · 249 across the 8 of their papers we have counts for
4 papers · 1 filter
Counterfactual Samples Synthesizing for Robust Visual Question Answering
Long Chen, Xin Yan, Jun Xiao +3
Despite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the trai…
Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao +3
Scene graphs -- objects as nodes and visual relationships as edges -- describe the whereabouts and interactions of the things and stuff in an image for comprehensive scene understa…
Graph-Theoretic Spatiotemporal Context Modeling for Video Saliency Detection
Lina Wei, Fangfang Wang, Xi Li +2
As an important and challenging problem in computer vision, video saliency detection is typically cast as a spatiotemporal context modeling problem over consecutive frames. As a re…
Video Question Answering via Attribute-Augmented Attention Network Learning
Yunan Ye, Zhou Zhao, Yimeng Li +3
Video Question Answering is a challenging problem in visual information retrieval, which provides the answer to the referenced video content according to the question. However, the…