3 papers
cs.LG2025
Large-Scale Auto-bidding with Nash Equilibrium Constraints
Zhiyu Mou, Miao Xu, Rongquan Bai +4
Auto-bidding has become a cornerstone of modern online advertising platforms, enabling many advertisers to automate bidding at scale and optimize campaign performance. However, pre…
cs.LG2025
Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Han Qi, Haochen Yang, Qiaosheng Zhang +1
We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contributi…
math.OC2023
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
Juno Kim, Kakei Yamamoto, Kazusato Oko +2
In this paper, we extend mean-field Langevin dynamics to minimax optimization over probability distributions for the first time with symmetric and provably convergent updates. We p…