1 paper
Javier Izquierdo, Aygul Zagidullina
I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We expl…