activity
20152022
most citedBenchmarking Batch Deep Reinforcement Learning Algorithms

159 citations · 566 across the 24 of their papers we have counts for

collaborators

43 papers

cs.LG20221 cited

Multi-Task Off-Policy Learning from Bandit Feedback

Joey Hong, Branislav Kveton, Sumeet Katariya +2

Many practical applications, such as recommender systems and learning to rank, involve solving multiple similar tasks. One example is learning of recommendation policies for users…

cs.LG20221 cited

Operator Splitting Value Iteration

Amin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh +1

We introduce new planning and reinforcement learning algorithms for discounted MDPs that utilize an approximate model of the environment to accelerate the convergence of the value…

cs.LG2022

Collaborative Multi-agent Stochastic Linear Bandits

Ahmadreza Moradipari, Mohammad Ghavamzadeh, Mahnoosh Alizadeh

We study a collaborative multi-agent stochastic linear bandit setting, where agents that form a network communicate locally to minimize their overall regret. In this setting, e…

cs.LG2022

Multi-Environment Meta-Learning in Stochastic Linear Bandits

Ahmadreza Moradipari, Mohammad Ghavamzadeh, Taha Rajabzadeh +2

In this work we investigate meta-learning (or learning-to-learn) approaches in multi-task linear stochastic bandit problems that can originate from multiple environments. Inspired…

cs.LG2022

Deep Hierarchy in Bandits

Joey Hong, Branislav Kveton, Sumeet Katariya +2

Mean rewards of actions are often correlated. The form of these correlations may be complex and unknown a priori, such as the preferences of a user for recommended products and the…

cs.LG2021

Adaptive Sampling for Minimax Fair Classification

Shubhanshu Shekhar, Greg Fields, Mohammad Ghavamzadeh +1

Machine learning models trained on uncurated datasets can often end up adversely affecting inputs belonging to underrepresented groups. To address this issue, we consider the probl…