9 citations · 10 across the 7 of their papers we have counts for
7 papers
Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping
Tao Chen, Lizheng Liu, Jiaxu Wang +4
Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similar…
U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
Linzhi Wu, Aoran Mei, Xiyue Wang +2
Diffusion-based methods have been acknowledged as a powerful paradigm for end-to-end visuomotor control in robotics. Most existing approaches adopt a Diffusion Policy in U-Net arch…
SO-DETR: Leveraging Dual-Domain Features and Knowledge Distillation for Small Object Detection
Huaxiang Zhang, Hao Zhang, Aoran Mei +2
Detection Transformer-based methods have achieved significant advancements in general object detection. However, challenges remain in effectively detecting small objects. One key d…
UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery
Huaxiang Zhang, Kai Liu, Zhongxue Gan +1
Unmanned aerial vehicle object detection (UAV-OD) has been widely used in various scenarios. However, most existing UAV-OD algorithms rely on manually designed components, which re…
ReplanVLM: Replanning Robotic Tasks with Visual Language Models
Aoran Mei, Guo-Niu Zhu, Huaxiang Zhang +1
Large language models (LLMs) have gained increasing popularity in robotic task planning due to their exceptional abilities in text analytics and generation, as well as their broad…
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
Huaxiang Zhang, Yaojia Mu, Guo-Niu Zhu +1
Accurate visual understanding is imperative for advancing autonomous systems and intelligent robots. Despite the powerful capabilities of vision-language models (VLMs) in processin…