17 papers
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian +4
Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent…
RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
Ehsan Ahmadi, Hunter Schofield, Behzad Khamidehi +5
Supervised open-loop training has been widely adopted for training traffic simulation models; however, it fails to capture the inherently dynamic, multi-agent interactions common i…
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
David Huang, Guile Wu, Chengjie Huang +2
Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effecti…
Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation
Shiyao Qian, Yuan Ren, Dongfeng Bai +1
Simulation is essential for autonomous driving, yet current frameworks often model vehicles as rigid assets and fail to capture part-level articulation. With perception algorithms…
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
Guile Wu, David Huang, Dongfeng Bai +1
Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily f…
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
Pan Wang, Yang Liu, Guile Wu +23
4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of sp…