Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Avoiding scaling in RLHF through Preference-based Exploration
Mingyu Chen, Yiding Chen, Wen Sun +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for large language model (LLM) alignment. This paper studies the setting of online RLHF and foc…
cs.LG2025
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Kianté Brantley, Mingyu Chen, Zhaolin Gao +4
Reinforcement learning (RL) has emerged as a powerful tool for fine-tuning large language models (LLMs) to improve complex reasoning abilities. However, state-of-the-art policy opt…
cs.LG2024
State-free Reinforcement Learning
Mingyu Chen, Aldo Pacchiano, Xuezhou Zhang
In this work, we study the \textit{state-free RL} problem, where the algorithm does not have the states information before interacting with the environment. Specifically, denote th…