93 citations · 108 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 93 cited
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Yi Wang, Kunchang Li, Yizhuo Li +14
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on…
cs.CV2022★ 14 cited
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Guo Chen, Sen Xing, Zhe Chen +18
In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, includin…
cs.CV2021★ 1 cited
MPN: Multimodal Parallel Network for Audio-Visual Event Localization
Jiashuo Yu, Ying Cheng, Rui Feng
Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained vid…