9 papers
RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
Ruicheng Zhang, Guangyu Chen, Zunnan Xu +5
Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through ima…
Identity-Consistent Video Generation under Large Facial-Angle Variations
Bin Hu, Zipeng Qi, Guoxi Huang +6
Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of…
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
Ronghui Li, Zhongyuan Hu, Li Siyao +6
Although existing 3D dance generation methods perform well in controlled scenarios, they often struggle to generalize in the wild. When conditioned on unseen music, existing method…
MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
Xiangzuo Wu, Chengwei Ren, Jun Zhou +2
Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view…
Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
Zihao Liu, Zunnan Xu, Shi Shu +4
This work presents Controllable Layer Decomposition (CLD), a method for achieving fine-grained and controllable multi-layer separation of raster images. In practical workflows, des…
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
Ruicheng Zhang, Jun Zhou, Zunnan Xu +5
Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally ex…