5 papers
UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models
Lin Sun, Zhiwei Guan, Conglin Wang +7
Mainstream Fast-Slow dual system vision-language-action models decouple a high-frequency action expert from a low-frequency vision-language model for efficiency, yet they face a fu…
Energy-based Compositional Diffusion Planning
Tao Sun, Utkarsh Aashu Mishra, Jiaxin Lu +2
Compositional diffusion planners aim to solve long-horizon robotic tasks using short training trajectories. Yet, current approaches often rely on the heuristic stitching of local p…
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
Yue Pan, Tao Sun, Liyuan Zhu +4
Point cloud registration aligns multiple unposed point clouds into a common reference frame and is a core step for 3D reconstruction and robot localization without initial guess. I…
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
Xianzheng Ma, Tao Sun, Shuai Chen +7
Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on te…
Rectified Point Flow: Generic Point Cloud Pose Estimation
Tao Sun, Liyuan Zhu, Shengyu Huang +2
We introduce Rectified Point Flow, a unified parameterization that formulates pairwise point cloud registration and multi-part shape assembly as a single conditional generative pro…