Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
Xiaobo Wang, Tong Wu, Min Tang +3
Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation…
cs.CL2025
In-Context Editing: Learning Knowledge from Self-Induced Distributions
Siyuan Qi, Bangcheng Yang, Kailin Jiang +5
In scenarios where language models must incorporate new information efficiently without extensive retraining, traditional fine-tuning methods are prone to overfitting, degraded gen…