16 citations · 18 across the 3 of their papers we have counts for
3 papers
cs.CL2022★ 1 cited
Equivariant and Invariant Grounding for Video Question Answering
Yicong Li, Xiang Wang, Junbin Xiao +1
Video Question Answering (VideoQA) is the task of answering the natural language questions about a video. Producing an answer requires understanding the interplay across visual sce…
cs.CV2022★ 1 cited
Video Graph Transformer for Video Question Answering
Junbin Xiao, Pan Zhou, Tat-Seng Chua +1
This paper proposes a Video Graph Transformer (VGT) model for Video Quetion Answering (VideoQA). VGT's uniqueness are two-fold: 1) it designs a dynamic graph transformer module whi…
cs.CV2021★ 16 cited
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering
Junbin Xiao, Angela Yao, Zhiyuan Liu +3
Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been foc…