32 citations · 48 across the 2 of their papers we have counts for
1 paper · 2 filters
Yi Wang, Kunchang Li, Xinhao Li +17
We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialog…