4 papers
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
Jiahui Cui, Yan Zhao, Kan Wei +4
Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone…
MotionScape: A Motion-Stratified UAV Video Benchmark for World Modeling and Future Video Generation
Zile Guo, Zhan Chen, Xiaoxuan Liu +5
Unmanned aerial vehicles (UAVs) are increasingly crucial for low-altitude autonomy and complex environment understanding. World models enable UAVs to anticipate how future states m…
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
Zhan Chen, Zile Guo, Enze Zhu +4
Video prediction is plagued by a fundamental trilemma: achieving high-resolution and perceptual quality typically comes at the cost of real-time speed, hindering its use in latency…
UNetMamba: An Efficient UNet-Like Mamba for Semantic Segmentation of High-Resolution Remote Sensing Images
Enze Zhu, Zhan Chen, Dingkai Wang +3
Semantic segmentation of high-resolution remote sensing images is vital in downstream applications such as land-cover mapping, urban planning and disaster assessment.Existing Trans…