13 papers · 1 filter
VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation
Jiayi Chen, Wenlong Dong, Yan Huang +5
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous c…
Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control
Jun Chen, Erdent Bao, Erdemt Bao +8
Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step…
GCNGrasp-VP: Affordance-Guided View Planning for Efficient Task-Oriented Grasping
Zanjia Tong, Wenlong Dong, Chengjie Zhang +1
Task-oriented grasping performance degrades significantly when object views suffer from occlusions. Existing task-oriented grasping methods typically assume task-relevant regions a…
VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models
Dehao Huang, Aoxiang Gu, Chengjie Zhang +5
Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level signal for monitoring manipulation in open-world environments and supportin…
Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations
Chao Tang, Jiacheng Xu, Haofei Lu +4
Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and t…
Easy-IIL: Reducing Human Operational Burden in Interactive Imitation Learning via Assistant Experts
Chengjie Zhang, Chao Tang, Wenlong Dong +3
Interactive Imitation Learning (IIL) typically relies on extensive human involvement for both offline demonstration and online interaction. Prior work primarily focuses on reducing…