10 papers
Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control
Jun Chen, Erdent Bao, Wenlong Dong +7
Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step…
GCNGrasp-VP: Affordance-Guided View Planning for Efficient Task-Oriented Grasping
Zanjia Tong, Wenlong Dong, Chengjie Zhang +1
Task-oriented grasping performance degrades significantly when object views suffer from occlusions. Existing task-oriented grasping methods typically assume task-relevant regions a…
VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models
Dehao Huang, Aoxiang Gu, Chengjie Zhang +5
Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level signal for monitoring manipulation in open-world environments and supportin…
Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations
Chao Tang, Jiacheng Xu, Haofei Lu +4
Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and t…
Easy-IIL: Reducing Human Operational Burden in Interactive Imitation Learning via Assistant Experts
Chengjie Zhang, Chao Tang, Wenlong Dong +3
Interactive Imitation Learning (IIL) typically relies on extensive human involvement for both offline demonstration and online interaction. Prior work primarily focuses on reducing…
Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning
Yufan Mao, Hanjing Ye, Wenlong Dong +2
Navigating complex environments requires robots to effectively store observations as memories and leverage them to answer human queries about spatial locations, which is a critical…