9 papers
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Seokju Cho, Ryo Hachiuma, Abhishek Badki +8
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-aug…
4DP-QA: Scalable QA for 4D Perception in Vision Language Models
Seokju Cho, Abhishek Badki, Hang Su +5
Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself…
Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification
Inès Hyeonsu Kim, Woojeong Jin, Soowon Son +6
Person re-identification (Re-ID) often faces challenges due to variations in human poses and camera viewpoints, which significantly affect the appearance of individuals across imag…
AnthroTAP: Learning Point Tracking with Real-World Motion
Inès Hyeonsu Kim, Seokju Cho, Jahyeok Koo +5
Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly synthetic$\unicode{x2014}$the only source currently feasi…
MV-TAP: Tracking Any Point in Multi-View Videos
Jahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim +6
Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications…
Seurat: From Moving Points to Depth
Seokju Cho, Jiahui Huang, Seungryong Kim +1
Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However,…