1 paper
Xitong Yang, Haoqi Fan, Lorenzo Torresani +2
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We…