4 papers
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
Quang Nguyen, Tri Le, Huy Nguyen +5
Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language com…
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao +8
Humanoid loco-manipulation, which integrates whole-body locomotion with dexterous manipulation, remains a fundamental challenge in robotics. Beyond whole-body coordination and bala…
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
Nghia Nguyen, Minh Nhat Vu, Tung D. Ta +4
Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is wide…
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
Huy Hoang Nguyen, An Vuong, Anh Nguyen +2
Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with clu…