14 papers
Unified Video Dense Prediction from Disjoint Data
Yihong Sun, Seoung Wug Oh, Jiahui Huang +2
Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, doma…
Efficient Tracking and Understanding Object Transformations
Yihong Sun, Bharath Hariharan
Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensive. TubeletGraph recently sh…
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
Gene Chou, Charles Herrmann, Kyle Genova +6
We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produc…
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
Wenxuan Peng, Bharath Hariharan, Hadar Averbuch-Elor
Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repe…
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
Aditya Chetan, Eric Cai, Peeyush Kushwaha +5
The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-graine…
ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
Chia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang +5
Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific…