6 papers
MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models
Yongshun Zhang, Zhongyi Fan, Yonghang Zhang +6
In recent years, large-scale generative models for visual content (\textit{e.g.,} images, videos, and 3D objects/scenes) have made remarkable progress. However, training large-scal…
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
Wenxu Qian, Chaoyue Wang, Hou Peng +3
Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-trai…
Semantix: An Energy Guided Sampler for Semantic Style Transfer
Huiang He, Minghui Hu, Chuanxia Zheng +2
Recent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additional…
MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models
Jing Zhao, Heliang Zheng, Chaoyue Wang +3
Large-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models gen…
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
Wenjie Xuan, Yufei Xu, Shanshan Zhao +4
ControlNet excels at creating content that closely matches precise contours in user-provided masks. However, when these masks contain noise, as a frequent occurrence with non-exper…
One-shot Generative Domain Adaptation in 3D GANs
Ziqiang Li, Yi Wu, Chaoyue Wang +2
3D-aware image generation necessitates extensive training data to ensure stable training and mitigate the risk of overfitting. This paper first considers a novel task known as One-…