7 papers
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
Xin Zhang, Haochen Wang, Yikang Zhou +2
The paper presents CycleGRPO, a reinforcement learning framework that lets a multimodal language model generate region captions and then use those captions to re‑localize the regio…
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
Dengxian Gong, Yuanzheng Wu, Haobo Yuan +11
This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure t…
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
Weisong Liu, Haochen Wang, Kuan Gao +8
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to co…
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Jiahao Meng, Yue Tan, Qi Xu +12
Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…
Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction
Hao Zhou, Lu Qi, Jason Li +5
Trajectory prediction is critical for autonomous driving, enabling safe and efficient planning in dense, dynamic traffic. Most existing methods optimize prediction accuracy under f…
AirSim360: A Panoramic Simulation Platform within Drone View
Xian Ge, Yuling Pan, Yuhang Zhang +12
The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data…