3 papers
cs.CV2026
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding
Yifan Wang, Peiming Li, Shiyu Li +5
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…
cs.CV2026
GeoStream: Toward Precise Camera Controlled Streaming Video Generation
Yizhou Zhao, Yifan Wang, Xiaoyuan Wang +11
Accurate interactive camera control is essential for video-based world models, but most existing approaches learn camera motion implicitly, leading to inaccurate control under out-…
cs.AI2025
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
Yubo Zhao, Qi Wu, Yifan Wang +2
This paper advances motion agents empowered by large language models (LLMs) toward autonomous navigation in dynamic and cluttered environments, significantly surpassing first and r…