4 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Xuanchang Zhang, Wei Xiong, Lichang Chen +3
In this paper, we study format biases in reinforcement learning from human feedback (RLHF). We observe that many widely-used preference models, including human evaluators, GPT-4, a…