Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data
Yi Zhao, Aidan Scannell, Wenshuai Zhao +7
Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online…
cs.LG2026
Bounded Ratio Reinforcement Learning
Yunke Ao, Le Chen, Bruce D. Lee +5
Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However…