6 papers
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
Hong Li, Minqi Meng, Yanjun Liang +10
Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarcity of high-quality PBR data…
PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion
Yichen Yang, Hong Li, Haodong Zhu +4
Existing autoregressive (AR) methods for generating artist-designed meshes struggle to balance global structural consistency with high-fidelity local details, and are susceptible t…
CKT-WAM: Parameter-Efficient Context Knowledge Transfer Between World Action Models
Yuhua Jiang, Yijun Guo, Hongbing Yang +7
World action models (WAMs) provide a powerful generative framework for embodied control, yet transferring knowledge across heterogeneous WAMs remains challenging due to mismatched…
Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
Zangwei Zheng, Xiangyu Peng, Yuxuan Lou +30
Video generation models have achieved remarkable progress in the past year. The quality of AI video continues to improve, but at the cost of larger model size, increased data quant…
MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
Guojun Lei, Chi Wang, Yikai Wang +3
Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are pre…
UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
Guojun Lei, Rong Zhang, Chi Wang +4
We propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video…