1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Jian Hu, Xibin Wu, Wei Shen +12
Large Language Models (LLMs) fine-tuned via Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning with Verifiable Rewards (RLVR) significantly improve the al…