activity
20162023
most citedEPOpt: Learning Robust Neural Network Policies Using Model Ensembles

144 citations · 292 across the 30 of their papers we have counts for

collaborators
Showing 2017Show all

11 papers · 1 filter

cs.LG2017

Rate of Change Analysis for Interestingness Measures

Nandan Sudarsanam, Nishanth Kumar, Abhishek Sharma +1

The use of Association Rule Mining techniques in diverse contexts and domains has resulted in the creation of numerous interestingness measures. This, in turn, has motivated resear…

cs.LG2017

Efficient-UCBV: An Almost Optimal Algorithm using Variance Estimates

Subhojyoti Mukherjee, K. P. Naveen, Nandan Sudarsanam +1

We propose a novel variant of the UCB algorithm (referred to as Efficient-UCB-Variance (EUCBV)) for minimizing cumulative regret in the stochastic multi-armed bandit (MAB) setting.…

cs.LG2017

Shared Learning : Enhancing Reinforcement in -Ensembles

Rakesh R Menon, Balaraman Ravindran

Deep Reinforcement Learning has been able to achieve amazing successes in a variety of domains from video games to continuous control by trying to maximize the cumulative reward. H…

cs.LG2017

RAIL: Risk-Averse Imitation Learning

Anirban Santara, Abhishek Naik, Balaraman Ravindran +4

Imitation learning algorithms learn viable policies by imitating an expert's behavior when reward signals are not available. Generative Adversarial Imitation Learning (GAIL) is a s…

cs.LG2017★ 16 cited

Learning to Factor Policies and Action-Value Functions: Factored Action Space Representations for Deep Reinforcement learning

Sahil Sharma, Aravind Suresh, Rahul Ramesh +1

Deep Reinforcement Learning (DRL) methods have performed well in an increasing numbering of high-dimensional visual decision making domains. Among all such visual decision making p…

cs.LG2017

Learning to Mix n-Step Returns: Generalizing lambda-Returns for Deep Reinforcement Learning

Sahil Sharma, Girish Raguvir J, Srivatsan Ramesh +1

Reinforcement Learning (RL) can model complex behavior policies for goal-directed sequential decision making tasks. A hallmark of RL algorithms is Temporal Difference (TD) learning…