1 paper
Dorcas Chia Ern Chua, Karen Myn Hui Lee, Jia Yue Tan +9
Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural…