1 paper · 1 filter
Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…