collaborators

7 papers

cs.CV2026

Video Generation Models are General-Purpose Vision Learners

Letian Wang, Chuhan Zhang, Rishabh Kabra +9

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a genera…

cs.CV2026

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson +5

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different vis…

cs.CV2026

Frozen Forecasting: A Unified Evaluation

Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera +7

Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "co…

cs.CV2026

How to Spin an Object: First, Get the Shape Right

Rishabh Kabra, Drew A. Hudson, Sjoerd van Steenkiste +2

Image-to-3D models increasingly rely on hierarchical generation to disentangle geometry and texture. However, the design choices underlying these two-stage models--particularly the…

cs.CV2026

OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts

Shiting Xiao, Rishabh Kabra, Yuhang Li +3

The ability to segment objects based on open-ended language prompts remains a critical challenge, requiring models to ground textual semantics into precise spatial masks while hand…

cs.CV2025

Scaling 4D Representations

João Carreira, Dilara Gokay, Michael King +32

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x20…