5 papers
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
Ziheng Ouyang, Yiren Song, Yaoli Liu +4
Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this pap…
OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
Yaoli Liu, Ziheng Ouyang, Shengtao Lou +1
Reference-guided image generation has progressed rapidly, yet current diffusion models still struggle to preserve fine-grained visual details when refining a generated image using…
AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
Shihao Zhu, Bohan Cao, Ziheng Ouyang +3
Recent diffusion model research focuses on generating identity-consistent images from a reference photo, but they struggle to accurately control age while preserving identity, and…
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
Yupeng Zhou, Zhen Li, Ziheng Ouyang +8
Encoding videos into discrete tokens could align with text tokens to facilitate concise and unified multi-modal LLMs, yet introducing significant spatiotemporal compression compare…
K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs
Ziheng Ouyang, Zhen Li, Qibin Hou
Recent studies have explored combining different LoRAs to jointly generate learned style and content. However, existing methods either fail to effectively preserve both the origina…