7 papers
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
Zhengyu Hu, Zheyuan Xiao, Linxin Song +10
LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the high…
On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits
Leo Maynard-Zhang, Zhihan Xiong, Kevin Jamieson +1
We study the fixed-budget best-arm identification (BAI) problem in non-stationary linear bandits. Concretely, given a fixed time budget , finite arm set $\mathcal{…
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
Zhengyu Hu, Jieyu Zhang, Zhihan Xiong +3
Despite the remarkable success of Large Language Models (LLMs), evaluating their outputs' quality regarding preference remains a critical challenge. While existing works usually le…
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
Avinandan Bose, Zhihan Xiong, Yuejie Chi +3
Personalizing large language models (LLMs) to accommodate diverse user preferences is essential for enhancing alignment and user satisfaction. Traditional reinforcement learning fr…
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
Avinandan Bose, Zhihan Xiong, Aadirupa Saha +2
Reinforcement Learning from Human Feedback (RLHF) is currently the leading approach for aligning large language models with human preferences. Typically, these models rely on exten…
Offline congestion games: How feedback type affects data coverage requirement
Haozhe Jiang, Qiwen Cui, Zhihan Xiong +2
This paper investigates when one can efficiently recover an approximate Nash Equilibrium (NE) in offline congestion games. The existing dataset coverage assumption in offline gener…