4 papers
FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
Yanjia Huang, Shuo Liu, Sheng Liu +4
Long-horizon robot manipulation tasks remain challenging for Vision-Language-Action (VLA) policies due to drift and exposure bias, often denoise the entire trajectory with fixed hy…
Diffusion Models for Robotic Manipulation: A Survey
Rosa Wolf, Yitian Shi, Sheng Liu +1
Diffusion generative models have demonstrated remarkable success in visual domains such as image and video generation. They have also recently emerged as a promising approach in ro…
VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility
Yitian Shi, Di Wen, Guanqi Chen +5
We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveragi…
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
Sicheng Wang, Sheng Liu, Weiheng Wang +2
Embodied intelligence seamlessly integrates vision, language, and action.~However, most multimodal robotic models rely on massive fine-tuning, incurring high time and hardware cost…