1 paper
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alig…