20 citations · 30 across the 8 of their papers we have counts for
1 paper · 2 filters
Yuanhang Luo, Yeheng Ge, Ruijian Han +1
In this work, we study the learning theory of reward modeling with pairwise comparison data using deep neural networks. We establish a novel non-asymptotic regret bound for deep re…