19 citations · 19 across the 3 of their papers we have counts for
3 papers
cs.CV2024
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Guo Chen, Yicheng Liu, Yifei Huang +6
Most existing video understanding benchmarks for multimodal large language models (MLLMs) focus only on short videos. The limited number of benchmarks for long video understanding…
cs.CV2024
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
Baoqi Pei, Guo Chen, Jilan Xu +8
In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Buildi…
cs.CV2024★ 19 cited
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen, Yifei Huang, Jilan Xu +7
Understanding videos is one of the fundamental directions in computer vision research, with extensive efforts dedicated to exploring various architectures such as RNN, 3D CNN, and…