Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
Pengyiang Liu, Zhongyue Shi, Hongye Hao +7
Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across m…
cs.CV2025
AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
Ruipu Wu, Yige Zhang, Jinyu Chen +5
Aerial Vision-and-Language Navigation (VLN) is an emerging task that enables Unmanned Aerial Vehicles (UAVs) to navigate outdoor environments using natural language instructions an…