1 paper · 1 filter
Xinlin Wang, Yujiao Xiang, Yuheng Zhou +11
Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA…