collaborators

7 papers

cs.CV2026

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

Xin Zhang, Haochen Wang, Yikang Zhou +2

The paper presents CycleGRPO, a reinforcement learning framework that lets a multimodal language model generate region captions and then use those captions to re‑localize the regio…

cs.CV2026

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

Dengxian Gong, Yuanzheng Wu, Haobo Yuan +11

This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure t…

cs.CV2026

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

Weisong Liu, Haochen Wang, Kuan Gao +8

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to co…

cs.CV2026

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Jiahao Meng, Yue Tan, Qi Xu +12

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…

cs.RO2026

Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction

Hao Zhou, Lu Qi, Jason Li +5

Trajectory prediction is critical for autonomous driving, enabling safe and efficient planning in dense, dynamic traffic. Most existing methods optimize prediction accuracy under f…

cs.CV2025

AirSim360: A Panoramic Simulation Platform within Drone View

Xian Ge, Yuling Pan, Yuhang Zhang +12

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data…