3 papers
math.OC2025
Last-Iterate Convergence of Adaptive Riemannian Gradient Descent for Equilibrium Computation
Yang Cai, Michael I. Jordan, Tianyi Lin +2
Equilibrium computation on Riemannian manifolds provides a unifying framework for numerous problems in machine learning and data analytics. One of the simplest yet most fundamental…
cs.LG2025
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
Yixin Liu, Argyris Oikonomou, Weiqiang Zheng +2
Many alignment methods, including reinforcement learning from human feedback (RLHF), rely on the Bradley-Terry reward assumption, which is not always sufficient to capture the full…
cs.LG2025
Provable Partially Observable Reinforcement Learning with Privileged Information
Yang Cai, Xiangyu Liu, Argyris Oikonomou +1
Partial observability of the underlying states generally presents significant challenges for reinforcement learning (RL). In practice, certain \emph{privileged information}, e.g.,…