20 papers
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Soowon Son, Honggyu An, Jisu Nam +7
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Seokju Cho, Ryo Hachiuma, Abhishek Badki +8
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-aug…
4DP-QA: Scalable QA for 4D Perception in Vision Language Models
Seokju Cho, Abhishek Badki, Hang Su +5
Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself…
Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text
Jaeyeong Kim, Ines Kim, Jahyeok Koo +1
We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent ambiguity of language, generat…
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
Minkyung Kwon, Jinhyeok Choi, Youngjin Shin +3
We present MORPHOS, a novel autoregressive framework that generates dynamic 3D assets from videos across diverse representations, including meshes, 3D Gaussians, and radiance field…
Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
Mungyeom Kim, Minkyeong Jeon, Honggyu An +10
Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame,…