2 papers
cs.LG2026
Bounded Ratio Reinforcement Learning
Yunke Ao, Le Chen, Bruce D. Lee +5
Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However…
math.OC2024
Contextual Bilevel Reinforcement Learning for Incentive Alignment
Vinzenz Thoma, Barna Pasztor, Andreas Krause +2
The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce C…