3 papers
cs.LG2025
Provably Efficient Online RLHF with One-Pass Reward Modeling
Long-Fei Li, Yu-Yang Qian, Peng Zhao +1
Reinforcement Learning from Human Feedback (RLHF) has shown remarkable success in aligning Large Language Models (LLMs) with human preferences. Traditional RLHF methods rely on a f…
cs.LG2025
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
Long-Fei Li, Yu-Jie Zhang, Peng Zhao +1
We study a new class of MDPs that employs multinomial logit (MNL) function approximation to ensure valid probability distributions over the state space. Despite its significant ben…
cs.LG2024
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
Long-Fei Li, Peng Zhao, Zhi-Hua Zhou
We study episodic linear mixture MDPs with the unknown transition and adversarial rewards under full-information feedback, employing dynamic regret as the performance measure. We s…