2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Corby Rosset, Ching-An Cheng, Arindam Mitra +3
This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach…