2.2k citations · 3.1k across the 15 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 22 cited
Redundancy-aware Transformer for Video Question Answering
Yicong Li, Xun Yang, An Zhang +3
This paper identifies two kinds of redundancy in the current VideoQA paradigm. Specifically, the current video encoders tend to holistically embed all video clues at different gran…
cs.CV2023
Discovering Spatio-Temporal Rationales for Video Question Answering
Yicong Li, Junbin Xiao, Chun Feng +2
This paper strives to solve complex video question answering (VideoQA) which features long video containing multiple objects and events at different time. To tackle the challenge,…
cs.CV2022★ 4 cited
Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation
Zhenguang Liu, Runyang Feng, Haoming Chen +4
Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequen…