activity
20142024
most citedProximal Reinforcement Learning: A New Theory of Sequential Decision Making in Primal-Dual Spaces

46 citations · 61 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20241 cited

Position: Benchmarking is Limited in Reinforcement Learning Research

Scott M. Jordan, Adam White, Bruno Castro da Silva +2

Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an e…

cs.LG20232 cited

Behavior Alignment via Reward Function Optimization

Dhawal Gupta, Yash Chandak, Scott M. Jordan +2

Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identif…

cs.LG20231 cited

Learning Fair Representations with High-Confidence Guarantees

Yuhong Luo, Austin Hoag, Philip S. Thomas

Representation learning is increasingly employed to generate representations that are predictive across multiple downstream tasks. The development of representation learning algori…

cs.LG2023

Coagent Networks: Generalized and Scaled

James E. Kostas, Scott M. Jordan, Yash Chandak +5

Coagent networks for reinforcement learning (RL) [Thomas and Barto, 2011] provide a powerful and flexible framework for deriving principled learning rules for arbitrary stochastic…

cs.LG2023

Optimization using Parallel Gradient Evaluations on Multiple Parameters

Yash Chandak, Shiv Shankar, Venkata Gandikota +2

We propose a first-order method for convex optimization, where instead of being restricted to the gradient from a single parameter, gradients from multiple parameters can be used d…

cs.LG20233 cited

Off-Policy Evaluation for Action-Dependent Non-Stationary Environments

Yash Chandak, Shiv Shankar, Nathaniel D. Bastian +3

Methods for sequential decision-making are often built upon a foundational assumption that the underlying decision process is stationary. This limits the application of such method…