activity
20242026
most citedUAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery

9 citations · 10 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AI2026

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

Tao Chen, Lizheng Liu, Jiaxu Wang +4

Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similar…

cs.RO2025

U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation

Linzhi Wu, Aoran Mei, Xiyue Wang +2

Diffusion-based methods have been acknowledged as a powerful paradigm for end-to-end visuomotor control in robotics. Most existing approaches adopt a Diffusion Policy in U-Net arch…

cs.CV2025★ 1 cited

SO-DETR: Leveraging Dual-Domain Features and Knowledge Distillation for Small Object Detection

Huaxiang Zhang, Hao Zhang, Aoran Mei +2

Detection Transformer-based methods have achieved significant advancements in general object detection. However, challenges remain in effectively detecting small objects. One key d…

cs.CV2025★ 9 cited

UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery

Huaxiang Zhang, Kai Liu, Zhongxue Gan +1

Unmanned aerial vehicle object detection (UAV-OD) has been widely used in various scenarios. However, most existing UAV-OD algorithms rely on manually designed components, which re…

cs.RO2024

ReplanVLM: Replanning Robotic Tasks with Visual Language Models

Aoran Mei, Guo-Niu Zhu, Huaxiang Zhang +1

Large language models (LLMs) have gained increasing popularity in robotic task planning due to their exceptional abilities in text analytics and generation, as well as their broad…

cs.CV2024

InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding

Huaxiang Zhang, Yaojia Mu, Guo-Niu Zhu +1

Accurate visual understanding is imperative for advancing autonomous systems and intelligent robots. Despite the powerful capabilities of vision-language models (VLMs) in processin…