1 paper
Sen Yang, Leyang Cui, Deng Cai +3
Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating resp…