13 papers
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
Leyi Wu, Yifan Zhao, Jinjie Zhang +11
Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essent…
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Tianshuo Xu, Yichen Xie, Depu Meng +5
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity p…
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
Bo Jiang, Depu Meng, Yihan Hu +3
Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedie…
SpectralSplat: Appearance-Disentangled Feed-Forward Gaussian Splatting for Driving Scenes
Quentin Herau, Tianshuo Xu, Depu Meng +5
Feed-forward 3D Gaussian Splatting methods have achieved impressive reconstruction quality for autonomous driving scenes, yet they entangle scene geometry with transient appearance…
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
Tianshuo Xu, Zhifei Chen, Leyi Wu +5
Recent advances in video generation have shown promise for generating future scenarios, critical for planning and control in autonomous driving and embodied intelligence. However,…
Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics
Tianshuo Xu, Zhifei Chen, Leyi Wu +2
The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controlla…