4 papers
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…
AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
Ruipu Wu, Yige Zhang, Jinyu Chen +5
Aerial Vision-and-Language Navigation (VLN) is an emerging task that enables Unmanned Aerial Vehicles (UAVs) to navigate outdoor environments using natural language instructions an…
SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps
Haojun Xu, Jiaqi Xiang, Wu Wei +5
A typical human strategy for giving navigation guidance is to sketch route maps based on the environmental layout. Inspired by this, we introduce Sketch map-based visual Navigation…
"Hi AirStar, Guide Me to the Badminton Court."
Ziqin Wang, Jinyu Chen, Xiangyi Zheng +3
Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly appro…