2 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.LG2023★ 2 cited
Behavior Alignment via Reward Function Optimization
Dhawal Gupta, Yash Chandak, Scott M. Jordan +2
Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identif…
cs.LG2023
Coagent Networks: Generalized and Scaled
James E. Kostas, Scott M. Jordan, Yash Chandak +5
Coagent networks for reinforcement learning (RL) [Thomas and Barto, 2011] provide a powerful and flexible framework for deriving principled learning rules for arbitrary stochastic…
cs.AI2019★ 2 cited
Soft Options Critic
Elita Lobo, Scott Jordan
The option-critic architecture (Bacon, Harb, and Precup 2017) and several variants have successfully demonstrated the use of the options framework proposed by Sutton et al (Sutton,…