4 papers
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Yijie Qian, Juncheng Wang, Chao Xu +6
As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality…
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
Yuxiang Feng, Juncheng Wang, Chao Xu +7
Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial design, but it remains challe…
NEWTON: Agentic Planning for Physically Grounded Video Generation
Yuxiang Feng, Juncheng Wang, Chao Xu +7
Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We…
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
Huihan Wang, Zhiwen Yang, Hui Zhang +3
Synthesizing high-quality dynamic medical videos remains a significant challenge due to the need for modeling both spatial consistency and temporal dynamics. Existing Transformer-b…