2 citations · 2 across the 7 of their papers we have counts for
7 papers · 1 filter
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
Liangbing Zhao, Le Zhuo, Sayak Paul +2
Instruction-based image editing has achieved remarkable success in semantic alignment, yet state-of-the-art models frequently fail to render physically plausible results when editi…
HPSv3: Towards Wide-Spectrum Human Preference Score
Yuhang Ma, Yunhao Shui, Xiaoshi Wu +2
Evaluating text-to-image generation models requires alignment with human perception, yet existing human-centric metrics are constrained by limited data coverage, suboptimal feature…
Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models
Fu-Yun Wang, Yunhao Shui, Jingtan Piao +2
Diffusion models have made substantial advances in image generation, yet models trained on large, unfiltered datasets often yield outputs misaligned with human preferences. Numerou…
GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec Avatars
Keqiang Sun, Amin Jourabloo, Riddhish Bhalodia +9
Photo-realistic and controllable 3D avatars are crucial for various applications such as virtual and mixed reality (VR/MR), telepresence, gaming, and film production. Traditional m…
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
Xiaoshi Wu, Yiming Hao, Manyuan Zhang +5
Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an a…
ECNet: Effective Controllable Text-to-Image Diffusion Models
Sicheng Li, Keqiang Sun, Zhixin Lai +5
The conditional text-to-image diffusion models have garnered significant attention in recent years. However, the precision of these models is often compromised mainly for two reaso…