1 paper · 1 filter
Oswin So, Brian Karrer, Chuchu Fan +2
Computation methods for solving entropy-regularized reward optimization -- a class of problems widely used for fine-tuning generative models -- have advanced rapidly. Among those,…