2 papers
cs.LG2025
Asymptotically-Optimal Gaussian Bandits with Side Observations
Alexia Atsidakou, Orestis Papadigenopoulos, Constantine Caramanis +2
We study the problem of Gaussian bandits with general side information, as first introduced by Wu, Szepesvari, and Gyorgy. In this setting, the play of an arm reveals information a…
cs.LG2025
InfoPO: On Mutual Information Maximization for Large Language Model Alignment
Teng Xiao, Zhen Ge, Sujay Sanghavi +5
We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in…