2 papers
cs.RO2025
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
Yandu Chen, Kefan Gu, Yuqing Wen +3
Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose em…
cs.CV2024
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
Wenrui Li, Wei Han, Yandu Chen +4
Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains une…