activity
20152026
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations · 3.5k across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2026

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9

Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusi…

cs.CV2026

TAPNext++: What's Next for Tracking Any Point (TAP)?

Sebastian Jung, Artem Zholus, Martin Sundermeyer +6

Tracking-Any-Point (TAP) models aim to track any point through a video which is a crucial task in AR/XR and robotics applications. The recently introduced TAPNext approach proposes…

cs.CV2026

Forecasting Motion in the Wild

Neerja Thakkar, Shiry Ginosar, Jacob Walker +3

Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectori…

cs.CV2025

Frozen Forecasting: A Unified Evaluation

Jacob C Walker, Pedro Vélez, Luisa Polania Cabrera +7

Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "co…

cs.CV2025

Direct Motion Models for Assessing Generated Videos

Kelsey Allen, Carl Doersch, Guangyao Zhou +9

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular…

cs.CV2025

TAPNext: Tracking Any Point (TAP) as Next Token Prediction

Artem Zholus, Carl Doersch, Yi Yang +7

Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods…