3 papers
cs.CV2026
Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
Siqi Lu, Wanying Xu, Yongbin Zheng +3
Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from…
cs.CV2025
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
Jianhang Yao, Yongbin Zheng, Siqi Lu +2
To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize…
cs.RO2025
HRT1: One-Shot Human-to-Robot Trajectory Transfer for Mobile Manipulation
Sai Haneesh Allu, Jishnu Jaykumar P, Ninad Khargonkar +3
We introduce a novel system for human-to-robot trajectory transfer that enables robots to manipulate objects by learning from human demonstration videos. The system consists of fou…