5 papers
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Ruiteng Zhao, Zhengshen Zhang, Yue Su +6
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both ali…
NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving
Jiahui Li, Jiawei Sun, Zixiang Ren +7
Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compact scene tokens for downstre…
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
Ruiteng Zhao, Wenshuo Wang, Yicheng Ma +4
Force sensing is a crucial modality for Vision-Language-Action (VLA) frameworks, as it enables fine-grained perception and dexterous manipulation in contact-rich tasks. We present…
3D Affordance Keypoint Detection for Robotic Manipulation
Zhiyang Liu, Ruiteng Zhao, Lei Zhou +6
This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts' functionality. The propo…
DexGrasp-Diffusion: Diffusion-based Unified Functional Grasp Synthesis Method for Multi-Dexterous Robotic Hands
Zhengshen Zhang, Lei Zhou, Chenchen Liu +6
The versatility and adaptability of human grasping catalyze advancing dexterous robotic manipulation. While significant strides have been made in dexterous grasp generation, curren…