1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 1 cited
On learning history based policies for controlling Markov decision processes
Gandharv Patil, Aditya Mahajan, Doina Precup
Reinforcementlearning(RL)folkloresuggeststhathistory-basedfunctionapproximationmethods,suchas recurrent neural nets or history-based state abstraction, perform better than their me…
cs.LG2021
Variance Penalized On-Policy and Off-Policy Actor-Critic
Arushi Jain, Gandharv Patil, Ayush Jain +2
Reinforcement learning algorithms are typically geared towards optimizing the expected return of an agent. However, in many practical applications, low variance in the return is de…