2 papers
cs.LG2025
Probabilistic Uncertain Reward Model
Wangtao Sun, Xiang Cheng, Xing Yu +5
Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…
cs.LG2025
Shuttle Between the Instructions and the Parameters of Large Language Models
Wangtao Sun, Haotian Xu, Huanxuan Liao +5
The interaction with Large Language Models (LLMs) through instructions has been extensively investigated in the research community. While instructions have been widely used as the…