2 papers
cs.CV2026
SeeSE3: Emergence of 3D Space in Vision Features
Caroline Chen, Sayna Ebrahimi, Fedor Kitashov +4
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D aw…
cs.CV2026
Gen4U: Unifying Video Generation and Understanding via Diffusion
Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models ov…