6 papers
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian +4
Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent…
RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning
Ehsan Ahmadi, Hunter Schofield, Behzad Khamidehi +5
Supervised open-loop training has been widely adopted for training traffic simulation models; however, it fails to capture the inherently dynamic, multi-agent interactions common i…
STaR: Scalable Task-Conditioned Retrieval for Long-Horizon Multimodal Robot Memory
Mingfeng Yuan, Hao Zhang, Mahan Mohammadi +3
Mobile robots are often deployed over long durations in diverse open, dynamic scenes, including indoor setting such as warehouses and manufacturing facilities, and outdoor settings…
Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving
Hunter Schofield, Mohammed Elmahgiubi, Kasra Rezaee +1
World models have become increasingly popular in acting as learned traffic simulators. Recent work has explored replacing traditional traffic simulators with world models for polic…
HIPPo: Harnessing Image-to-3D Priors for Model-free Zero-shot 6D Pose Estimation
Yibo Liu, Zhaodong Jiang, Binbin Xu +7
This work focuses on model-free zero-shot 6D object pose estimation for robotics applications. While existing methods can estimate the precise 6D pose of objects, they heavily rely…
L-PR: Exploiting LiDAR Fiducial Marker for Unordered Low Overlap Multiview Point Cloud Registration
Yibo Liu, Jinjun Shan, Amaldev Haridevan +1
Point cloud registration is a prerequisite for many applications in computer vision and robotics. Most existing methods focus on pairwise registration of two point clouds with high…