1 paper
Junghyun Lee, Minju Hong, Kwang-Sung Jun +2
We consider the problem of regularized best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized t…