230 citations · 483 across the 20 of their papers we have counts for
9 papers · 1 filter
Policy Gradient for Reinforcement Learning with General Utilities
Navdeep Kumar, Kaixin Wang, Kfir Levy +1
In Reinforcement Learning (RL), the goal of agents is to discover an optimal policy that maximizes the expected cumulative rewards. This objective may also be viewed as finding a p…
Actor-Critic based Improper Reinforcement Learning
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan +1
We consider an improper reinforcement learning setting where a learner is given base controllers for an unknown Markov decision process, and wishes to combine them optimally to…
Analysis of Stochastic Processes through Replay Buffers
Shirli Di Castro Shashua, Shie Mannor, Dotan Di-Castro
Replay buffers are a key component in many reinforcement learning schemes. Yet, their theoretical properties are not fully understood. In this paper we analyze a system where a sto…
Outlier Robust Online Learning
Jiashi Feng, Huan Xu, Shie Mannor
We consider the problem of learning from noisy data in practical settings where the size of data is too large to store on a single machine. More challenging, the data coming from t…
Adaptive Lambda Least-Squares Temporal Difference Learning
Timothy A. Mann, Hugo Penedones, Shie Mannor +1
Temporal Difference learning or TD() is a fundamental algorithm in the field of reinforcement learning. However, setting TD's parameter, which controls the timescale of TD u…
Supervised Learning for Optimal Power Flow as a Real-Time Proxy
Raphael Canyasse, Gal Dalal, Shie Mannor
In this work we design and compare different supervised learning algorithms to compute the cost of Alternating Current Optimal Power Flow (ACOPF). The motivation for quick calculat…