From the 1 of 7 linked papers with an AI index.
7 papers
Event-RGB Adaptive Tracking for Nighttime Highway Perception
Haidong Wang, Hengxing Cai, Wanlei Li +2
The paper introduces a Joint Event‑RGB Adaptive Tracking (JEAT) framework that fuses asynchronous event camera data with RGB images using an adaptive extended Kalman filter to impr…
AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions
Hengxing Cai, Yijie Rao, Ligang Huang +7
Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making…
SA-GCS: Semantic-Aware Gaussian Curriculum Scheduling for UAV Vision-Language Navigation
Hengxing Cai, Jinhan Dong, Yijie Rao +8
Unmanned Aerial Vehicle (UAV) Vision-Language Navigation (VLN) aims to enable agents to accurately localize targets and plan flight paths in complex environments based on natural l…
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
Mingjun Xu, Jinhan Dong, Jue Hou +5
Multimodal document retrieval systems enable information access across text, images, and layouts, benefiting various domains like document-based question answering, report analysis…
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
Hengxing Cai, Jinhan Dong, Jingjun Tan +7
Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing…
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
Mingjun Xu, Zehui Wang, Hengxing Cai +1
Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass tex…