5 papers
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
Zhanpeng Luo, Ce Zhang, Silong Yong +6
Multi-modal Large Language Models (MLLMs) have demonstrated strong capabilities in general-purpose perception and reasoning, but they still struggle with tasks that require spatial…
Unifying Deep Predicate Invention with Pre-trained Foundation Models
Qianwei Wang, Bowen Li, Zhanpeng Luo +6
Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture…
Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather
Zhijian He, Feifei Liu, Yuwei Li +4
Multi-modal 3D object detection is important for reliable perception in robotics and autonomous driving. However, its effectiveness remains limited under adverse weather conditions…
Instant4D: 4D Gaussian Splatting in Minutes
Zhanpeng Luo, Haoxi Ran, Li Lu
Dynamic view synthesis has seen significant advances, yet reconstructing scenes from uncalibrated, casual video remains challenging due to slow optimization and complex parameter e…
Imagine with the Teacher: Complete Shape in a Multi-View Distillation Way
Zhanpeng Luo, Linna Wang, Guangwu Qian +1
Point cloud completion aims to recover the completed 3D shape of an object from its partial observation caused by occlusion, sensor's limitation, noise, etc. When some key semantic…