21 citations · 33 across the 24 of their papers we have counts for
1 paper · 1 filter
Daocheng Fu, Rong Wu, Yu Yang +7
Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on…