Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences
Junghyun Lee, Minju Hong, Kwang-Sung Jun +2
We consider the problem of regularized best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized t…
cs.LG2024
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
Jiseok Chae, Chulhee Yun, Donghwan Kim
In minimax optimization, the extragradient (EG) method has been extensively studied because it outperforms the gradient descent-ascent method in convex-concave (C-C) problems. Yet,…