3 papers
cs.CV2026
GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision
Dianyu Wang, Peirong Zhang, Xuyang Li +2
Recent multimodal large language models (MLLMs) have shown strong cross-modal understanding and coordinate generation abilities in visual grounding. However, transferring these abi…
cs.CV2026
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
Zile Guo, Zhan Chen, Enze Zhu +4
Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. Fo…
cs.CV2025
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
Zhan Chen, Zile Guo, Enze Zhu +4
Video prediction is plagued by a fundamental trilemma: achieving high-resolution and perceptual quality typically comes at the cost of real-time speed, hindering its use in latency…