4 papers
AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation
Yulu Wu, Jiujun Cheng, Haowen Wang +5
Recent advances in Vision-Language-Action (VLA) and world-model methods have improved generalization in tasks such as robotic manipulation and object interaction. However, Successf…
A Mechanistic View on Video Generation as World Models: State and Dynamics
Luozhou Wang, Zhifei Chen, Yihua Du +11
Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateles…
AnimateAnywhere: Rouse the Background in Human Image Animation
Xiaoyu Liu, Mingshuai Yao, Yabo Zhang +5
Human image animation aims to generate human videos of given characters and backgrounds that adhere to the desired pose sequence. However, existing methods focus more on human acti…
DiffuEraser: A Diffusion Model for Video Inpainting
Xiaowen Li, Haolan Xue, Peiran Ren +1
Recent video inpainting algorithms integrate flow-based pixel propagation with transformer-based generation to leverage optical flow for restoring textures and objects using inform…