2 papers
cs.CV2026
Recurrent Video Masked Autoencoders
Daniel Zoran, Nikhil Parthasarathy, Yi Yang +3
We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of vide…
cs.CV2025
Self-supervised video pretraining yields robust and more human-aligned visual representations
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira +1
Humans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, sta…