22 citations · 41 across the 3 of their papers we have counts for
5 papers
Towards Safe Policy Improvement for Non-Stationary MDPs
Yash Chandak, Scott M. Jordan, Georgios Theocharous +2
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that a…
Evaluating the Performance of Reinforcement Learning Algorithms
Scott M. Jordan, Yash Chandak, Daniel Cohen +2
Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results ar…
Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Philip S. Thomas, Scott M. Jordan, Yash Chandak +2
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the…
Learning Action Representations for Reinforcement Learning
Yash Chandak, Georgios Theocharous, James Kostas +2
Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the str…
Distributed Evaluations: Ending Neural Point Metrics
Daniel Cohen, Scott M. Jordan, W. Bruce Croft
With the rise of neural models across the field of information retrieval, numerous publications have incrementally pushed the envelope of performance for a multitude of IR tasks. H…