Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Monotone and Conservative Policy Iteration Beyond the Tabular Case
S. R. Eshwar, Gugan Thoppe, Ananyabrata Barua +2
We introduce Reliable Policy Iteration (RPI) and Conservative RPI (CRPI), variants of Policy Iteration (PI) and Conservative PI (CPI), that retain tabular guarantees under function…
cs.LG2025
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
S. R. Eshwar
Time-inconsistent preferences, where agents favor smaller-sooner over larger-later rewards, are a key feature of human and animal decision-making. Quasi-Hyperbolic (QH) discounting…
cs.LG2024
Reinforcement Learning with Quasi-Hyperbolic Discounting
S. R. Eshwar, Mayank Motwani, Nibedita Roy +1
Reinforcement learning has traditionally been studied with exponential discounting or the average reward setup, mainly due to their mathematical tractability. However, such framewo…