1 paper · 1 filter
Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3
Iterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks (Y…