10 papers
DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement
Kai Wang, Ziheng Ouyang, Xuying Zhang +2
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches…
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
Jingbo Gong, Yikai Wang, Yushi Lan +6
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formu…
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Yiyang Fu, Chubin Zhang, Shukai Gong +7
It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-Language-Action (VLA) models…
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
Yupeng Zhou, Lianghua Huang, Zhifan Wu +7
In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key ch…
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
Ziheng Ouyang, Yiren Song, Yaoli Liu +4
Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this pap…
AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
Shihao Zhu, Bohan Cao, Ziheng Ouyang +3
Recent diffusion model research focuses on generating identity-consistent images from a reference photo, but they struggle to accurately control age while preserving identity, and…