1 paper · 1 filter
Giyeong Oh, Junghyun Lee, Jaehyun Park +3
Modern LLMs inherit strong priors from web-scale pretraining, which can limit the headroom of post-training data-selection strategies. While Active Preference Learning (APL) seeks…