1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Yanshi Li, Shaopan Xiong, Gengru Chen +6
Reinforcement Learning (RL) has proven highly effective in aligning Large Language Models (LLMs) with human preferences. Typical RL methods optimize under an overall sequence rewar…