works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

Event-RGB Adaptive Tracking for Nighttime Highway Perception

Haidong Wang, Hengxing Cai, Wanlei Li +2

The paper introduces a Joint Event‑RGB Adaptive Tracking (JEAT) framework that fuses asynchronous event camera data with RGB images using an adaptive extended Kalman filter to impr…

cs.CL2026

AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions

Hengxing Cai, Yijie Rao, Ligang Huang +7

Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making…

cs.CL2025

SA-GCS: Semantic-Aware Gaussian Curriculum Scheduling for UAV Vision-Language Navigation

Hengxing Cai, Jinhan Dong, Yijie Rao +8

Unmanned Aerial Vehicle (UAV) Vision-Language Navigation (VLN) aims to enable agents to accurately localize targets and plan flight paths in complex environments based on natural l…

cs.AI2025

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval

Mingjun Xu, Jinhan Dong, Jue Hou +5

Multimodal document retrieval systems enable information access across text, images, and layouts, benefiting various domains like document-based question answering, report analysis…

cs.CL2025

FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models

Hengxing Cai, Jinhan Dong, Jingjun Tan +7

Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing…

cs.IR2025

A Multi-Granularity Retrieval Framework for Visually-Rich Documents

Mingjun Xu, Zehui Wang, Hengxing Cai +1

Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass tex…