activity
20242026
collaborators

6 papers

cs.CV2026

Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models

Yijie Qian, Juncheng Wang, Chao Xu +6

As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality…

cs.CV2026

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

Yuxiang Feng, Juncheng Wang, Chao Xu +7

Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial design, but it remains challe…

cs.CV2026

NEWTON: Agentic Planning for Physically Grounded Video Generation

Yuxiang Feng, Juncheng Wang, Chao Xu +7

Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We…

cs.CV2025

Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation

Yijie Qian, Juncheng Wang, Yuxiang Feng +7

Current state-of-the-art paradigms predominantly treat Text-to-Motion (T2M) generation as a direct translation problem, mapping symbolic language directly to continuous poses. Whil…

cs.CV2025

Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving

Yu Yang, Jianbiao Mei, Yukai Ma +5

World models envision potential future states based on various ego actions. They embed extensive knowledge about the driving environment, facilitating safe and scalable autonomous…

cs.CV2024

Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding

Hanchen Tai, Qingdong He, Jiangning Zhang +6

Open-vocabulary 3D scene understanding presents a significant challenge in the field. Recent works have sought to transfer knowledge embedded in vision-language models from 2D to 3…