93 citations · 107 across the 3 of their papers we have counts for
4 papers
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Yi Wang, Kunchang Li, Yizhuo Li +14
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on…
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Guo Chen, Sen Xing, Zhe Chen +18
In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, includin…
Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Yicong Hong, Zun Wang, Qi Wu +1
Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. The fundamen…
Non-Standard Primordial Clocks from Dynamical Mass in Alternative to Inflation Scenarios
Yi Wang, Zun Wang, Yuhang Zhu
In the primordial universe, oscillations of heavy fields can be considered as standard clocks to measure the expansion or contraction history of the universe. Those standard clocks…