1 paper · 1 filter
Jinjie Shen, Wei Deng, Xian Hu +2
Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the enti…