collaborators

10 papers

cs.CV2026

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

Shenghong Yi, Lin Zhang, Muzian Li +6

Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-…

cs.CV2026

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Haoyu Zhang, Shuoxun Zhang, Peng Ye +5

Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme sc…

cs.CV2026

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

Haoxian Zhou, Chuanzhi Xu, Langyi Chen +5

Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brig…

cs.CV2026

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

Langyi Chen, Chuanzhi Xu, Haoxian Zhou +6

Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at scale because specialized sens…

cs.CV2026

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

Mengzhao Wang, Yanli Ji, Wangmeng Zuo +2

Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeated…

cs.CV2026

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

Haoyu Zhang, Jingyi Zhou, Peng Ye +4

With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot ability. However, due to th…