3 papers
cs.AI2025
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
S. R. Eshwar, Aniruddha Mukherjee, Kintan Saha +4
In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here…
cs.LG2025
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
S. R. Eshwar
Time-inconsistent preferences, where agents favor smaller-sooner over larger-later rewards, are a key feature of human and animal decision-making. Quasi-Hyperbolic (QH) discounting…
cs.LG2024
Reinforcement Learning with Quasi-Hyperbolic Discounting
S. R. Eshwar, Mayank Motwani, Nibedita Roy +1
Reinforcement learning has traditionally been studied with exponential discounting or the average reward setup, mainly due to their mathematical tractability. However, such framewo…