activity
20242026
collaborators

5 papers

cs.CV2026

Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding

Yifan Wang, Peiming Li, Shiyu Li +5

While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…

cs.CV2026

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

Yun Wang, Junbin Xiao, Han Lyu +6

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence…

cs.CV2026

GeoStream: Toward Precise Camera Controlled Streaming Video Generation

Yizhou Zhao, Yifan Wang, Xiaoyuan Wang +11

Accurate interactive camera control is essential for video-based world models, but most existing approaches learn camera motion implicitly, leading to inaccurate control under out-…

cs.AI2025

Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning

Yubo Zhao, Qi Wu, Yifan Wang +2

This paper advances motion agents empowered by large language models (LLMs) toward autonomous navigation in dynamic and cluttered environments, significantly surpassing first and r…

cs.CV2024

Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

Qi Wu, Yubo Zhao, Yifan Wang +3

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these cha…