activity
20172022
most citedFRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback

13 citations · 32 across the 7 of their papers we have counts for

collaborators

9 papers

cs.MA2022

Shaping Advice in Deep Reinforcement Learning

Baicen Xiao, Bhaskar Ramasubramanian, Radha Poovendran

Reinforcement learning involves agents interacting with an environment to complete tasks. When rewards provided by the environment are sparse, agents may not receive immediate feed…

cs.MA20224 cited

Agent-Temporal Attention for Reward Redistribution in Episodic Multi-Agent Reinforcement Learning

Baicen Xiao, Bhaskar Ramasubramanian, Radha Poovendran

This paper considers multi-agent reinforcement learning (MARL) tasks where agents receive a shared global reward at the end of an episode. The delayed nature of this reward affects…

cs.LG20213 cited

Shaping Advice in Deep Multi-Agent Reinforcement Learning

Baicen Xiao, Bhaskar Ramasubramanian, Radha Poovendran

Multi-agent reinforcement learning involves multiple agents interacting with each other and a shared environment to complete tasks. When rewards provided by the environment are spa…

eess.SY2020

Safety-Critical Online Control with Adversarial Disturbances

Bhaskar Ramasubramanian, Baicen Xiao, Linda Bushnell +1

This paper studies the control of safety-critical dynamical systems in the presence of adversarial disturbances. We seek to synthesize state-feedback controllers to minimize a cost…

cs.AI202013 cited

FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback

Baicen Xiao, Qifan Lu, Bhaskar Ramasubramanian +3

Reinforcement learning has been successful in training autonomous agents to accomplish goals in complex environments. Although this has been adapted to multiple settings, including…

cs.LG20191 cited

Potential-Based Advice for Stochastic Policy Learning

Baicen Xiao, Bhaskar Ramasubramanian, Andrew Clark +3

This paper augments the reward received by a reinforcement learning agent with potential functions in order to help the agent learn (possibly stochastic) optimal policies. We show…