activity
20242026
collaborators

7 papers

cs.CV2026

SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation

Wen Jiang, Kangyao Huang, Li Wang +9

UAVs play an important role in applications such as autonomous exploration, disaster response, and infrastructure inspection. However, UAV VLN in complex 3D environments remains ch…

cs.RO2026

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments

Wen Jiang, Hanfang Liang, Li Wang +9

Recent advances in multimodal large models have significantly improved UAV vision-language navigation (UAV-VLN) by enhancing high-level perception and reasoning. However, existing…

cs.CV2026

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models

Feihong Yan, Shaoyu Liu, Haixuan Wang +4

Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents direct generation at higher r…

cs.CV2026

EventFlash: Towards Efficient MLLMs for Event-Based Vision

Shaoyu Liu, Jianing Li, Guanghui Zhao +4

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, curre…

cs.CV2025

LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration

Wen Jiang, Li Wang, Kangyao Huang +6

Unmanned aerial vehicles (UAVs) are crucial tools for post-disaster search and rescue, facing challenges such as high information density, rapid changes in viewpoint, and dynamic s…

cs.CV2025

EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs

Shaoyu Liu, Jianing Li, Guanghui Zhao +2

Multimodal large language models (MLLMs) have made significant advancements in event-based vision, yet the comprehensive evaluation of their capabilities within a unified benchmark…