1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Beining Wang, Weihang Su, Hongtao Tian +8
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become…