2 papers
cs.LG2024
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
Renye Yan, Yaozhong Gan, You Wu +4
The imbalance of exploration and exploitation has long been a significant challenge in reinforcement learning. In policy optimization, excessive reliance on exploration reduces lea…
cs.LG2024
Transductive Off-policy Proximal Policy Optimization
Yaozhong Gan, Renye Yan, Xiaoyang Tan +2
Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due to its inherent on-policy nature…