22 citations · 40 across the 4 of their papers we have counts for
4 papers
Redundancy-aware Transformer for Video Question Answering
Yicong Li, Xun Yang, An Zhang +3
This paper identifies two kinds of redundancy in the current VideoQA paradigm. Specifically, the current video encoders tend to holistically embed all video clues at different gran…
The XPRESS Challenge: Xray Projectomic Reconstruction -- Extracting Segmentation with Skeletons
Tri Nguyen, Mukul Narwani, Mark Larson +9
The wiring and connectivity of neurons form a structural basis for the function of the nervous system. Advances in volume electron microscopy (EM) and image segmentation have enabl…
Equivariant and Invariant Grounding for Video Question Answering
Yicong Li, Xiang Wang, Junbin Xiao +1
Video Question Answering (VideoQA) is the task of answering the natural language questions about a video. Producing an answer requires understanding the interplay across visual sce…
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering
Junbin Xiao, Angela Yao, Zhiyuan Liu +3
Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been foc…