1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Yuan Sun, Navid Salami Pargoo, Peter J. Jin +1
Reinforcement Learning from Human Feedback (RLHF) is popular in large language models (LLMs), whereas traditional Reinforcement Learning (RL) often falls short. Current autonomous…