1 paper · 1 filter
Yurong Chen, Yu He, Michael I. Jordan +1
Standard methods for aligning large language models with human preferences learn from pairwise comparisons among sampled candidate responses and regularize toward a reference polic…