8 papers
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
Peng Yun, Shouwang Huang, Hao Li +3
Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models strugg…
Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation
Hao Li
Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows that the \emph{starting noise}…
PhysInOne: Visual Physics Learning and Reasoning in One Suite
Siyuan Zhou, Hejun Wang, Hu Cheng +36
We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to mere…
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
Luozheng Qin, Jia Gong, Qian Qiao +6
Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than under…
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
Angen Ye, Boyuan Wang, Chaojun Ni +21
World-Action Models (WAM) initialized from pre-trained video generation backbones have demonstrated remarkable potential for robot policy learning. However, existing approaches fac…
Steering Video Diffusion Transformers with Massive Activations
Xianhang Cheng, Yujian Zheng, Zhenyu Xie +2
Despite rapid progress in video diffusion transformers, how their internal model signals can be leveraged with minimal overhead to enhance video generation quality remains underexp…