Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
Xuerui Su, Yue Wang, Jinhua Zhu +4
With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and a…
cs.LG2018
Finite Sample Analysis of the GTD Policy Evaluation Algorithms in Markov Setting
Yue Wang, Wei Chen, Yuting Liu +2
In reinforcement learning (RL) , one of the key components is policy evaluation, which aims to estimate the value function (i.e., expected long-term accumulated reward) of a policy…
cs.LG2018
Target Transfer Q-Learning and Its Convergence Analysis
Yue Wang, Qi Meng, Wei Cheng +3
Q-learning is one of the most popular methods in Reinforcement Learning (RL). Transfer Learning aims to utilize the learned knowledge from source tasks to help new tasks to improve…