1 paper
Xin Jin, Huanqia Cai, Zhen Li +9
Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic sc…