86 citations · 100 across the 6 of their papers we have counts for
7 papers · 1 filter
Memory-and-Anticipation Transformer for Online Action Understanding
Jiahao Wang, Guo Chen, Yifei Huang +2
Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal…
VideoLLM: Modeling Video Sequence with Large Language Models
Guo Chen, Yin-Dong Zheng, Jiahao Wang +8
With the exponential growth of video data, there is an urgent need for automated technology to analyze and comprehend video content. However, existing video understanding models ar…
MRSN: Multi-Relation Support Network for Video Action Detection
Yin-Dong Zheng, Guo Chen, Minglei Yuan +1
Action detection is a challenging video understanding task, requiring modeling spatio-temporal and interaction relations. Current methods usually model actor-actor and actor-contex…
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Guo Chen, Sen Xing, Zhe Chen +18
In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, includin…
Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 2022
Yin-Dong Zheng, Guo Chen, Jiahao Wang +2
Capturing the state changes of interacting objects is a key technology for understanding human-object interactions. This technical report describes our method using heterogeneous b…
Uncertainty-based Network for Few-shot Image Classification
Minglei Yuan, Qian Xu, Chunhao Cai +3
The transductive inference is an effective technique in the few-shot learning task, where query sets update prototypes to improve themselves. However, these methods optimize the mo…