4 papers
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
Sijin Chen, Kaixuan Jiang, Haixin Shi +6
We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it…
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
Haosheng Li, Weixin Mao, Zihan Lan +6
Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. H…
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
Zihan Lan, Weixin Mao, Haosheng Li +4
In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally an…
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
Haosheng Li, Weixin Mao, Weipeng Deng +7
Multi-hand semantic grasp generation aims to generate feasible and semantically appropriate grasp poses for different robotic hands based on natural language instructions. Although…