4 citations · 7 across the 18 of their papers we have counts for
16 papers · 1 filter
SeeSE3: Emergence of 3D Space in Vision Features
Caroline Chen, Sayna Ebrahimi, Fedor Kitashov +4
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D aw…
Gen4U: Unifying Video Generation and Understanding via Diffusion
Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models ov…
Perception Test 2025: Challenge Summary and a Unified VQA Extension
Joseph Heyward, Nikhil Parthasarathy, Tyler Zhu +5
The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benc…
Unique Lives, Shared World: Learning from Single-Life Videos
Tengda Han, Sayna Ebrahimi, Dilara Gokay +8
We introduce the "single-life" learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual. We leverage the multiple viewp…
Dynamic Reflections: Probing Video Representations with Text Alignment
Tyler Zhu, Tengda Han, Leonidas Guibas +2
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encod…
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
Artem Zholus, Carl Doersch, Yi Yang +7
Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods…