activity
20242026
most citedHow to Spin an Object: First, Get the Shape Right

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CV2026

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson +5

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different vis…

cs.CV2026

Recurrent Video Masked Autoencoders

Daniel Zoran, Nikhil Parthasarathy, Yi Yang +3

We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of vide…

cs.CV20261 cited

How to Spin an Object: First, Get the Shape Right

Rishabh Kabra, Drew A. Hudson, Sjoerd van Steenkiste +2

Image-to-3D models increasingly rely on hierarchical generation to disentangle geometry and texture. However, the design choices underlying these two-stage models--particularly the…

cs.CV2025

LayerLock: Non-collapsing Representation Learning with Progressive Freezing

Goker Erdogan, Nikhil Parthasarathy, Catalin Ionescu +5

We introduce LayerLock, a simple yet effective approach for self-supervised visual representation learning, that gradually transitions from pixel to latent prediction through progr…

cs.CV2025

Scaling 4D Representations

João Carreira, Dilara Gokay, Michael King +32

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x20…

cs.CV2024

Moving Off-the-Grid: Scene-Grounded Video Representations

Sjoerd van Steenkiste, Daniel Zoran, Yi Yang +13

Current vision models typically maintain a fixed correspondence between their representation structure and image space. Each layer comprises a set of tokens arranged "on-the-grid,"…