31 citations · 31 across the 1 of their papers we have counts for
1 paper
Yuan Sun, Navid Salami Pargoo, Peter J. Jin +1
Reinforcement Learning from Human Feedback (RLHF) is popular in large language models (LLMs), whereas traditional Reinforcement Learning (RL) often falls short. Current autonomous…