collaborators

5 papers

cs.CL2025

Judging with Confidence: Calibrating Autoraters to Preference Distributions

Zhuohang Li, Xiaowei Li, Chengyu Huang +11

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…

cs.LG2025

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6

Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typica…

cs.LG2025

Representation Learning via Non-Contrastive Mutual Information

Zhaohan Daniel Guo, Bernardo Avila Pires, Khimya Khetarpal +2

Labeling data is often very time consuming and expensive, leaving us with a majority of unlabeled data. Self-supervised representation learning methods such as SimCLR (Chen et al.,…

cs.LG2025

Ordering-based Conditions for Global Convergence of Policy Gradient Methods

Jincheng Mei, Bo Dai, Alekh Agarwal +3

We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the pol…

cs.LG2025

Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates

Jincheng Mei, Bo Dai, Alekh Agarwal +4

We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learnin…