Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Rui Yang, Ruomeng Ding, Yong Lin +2
Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning f…
cs.CL2024
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
Hanning Zhang, Shizhe Diao, Yong Lin +6
Large language models (LLMs) have revolutionized numerous domains with their impressive performance but still face their challenges. A predominant issue is the propensity for these…