activity
20182022
most citedRisk-Sensitive Bayesian Games for Multi-Agent Reinforcement Learning under Policy Uncertainty

1 citations · 2 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG20221 cited

Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration and Planning

Reda Ouhamma, Debabrota Basu, Odalric-Ambrym Maillard

We study the problem of episodic reinforcement learning in continuous state-action spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewa…

cs.LG20221 cited

Risk-Sensitive Bayesian Games for Multi-Agent Reinforcement Learning under Policy Uncertainty

Hannes Eriksson, Debabrota Basu, Mina Alibeigi +1

In stochastic games with incomplete information, the uncertainty is evoked by the lack of knowledge about a player's own and the other players' types, i.e. the utility function and…

cs.LG2020

Inferential Induction: A Novel Framework for Bayesian Reinforcement Learning

Hannes Eriksson, Emilio Jorge, Christos Dimitrakakis +2

Bayesian reinforcement learning (BRL) offers a decision-theoretic solution for reinforcement learning. While "model-based" BRL algorithms have focused either on maintaining a poste…

cs.LG2019

Near-optimal Bayesian Solution For Unknown Discrete Markov Decision Process

Aristide Tossou, Christos Dimitrakakis, Debabrota Basu

We tackle the problem of acting in an unknown finite and discrete Markov Decision Process (MDP) for which the expected shortest path from any state to any other state is bounded by…

cs.LG2019

Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

Aristide Tossou, Debabrota Basu, Christos Dimitrakakis

We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence int…

cs.LG2019

Differential Privacy for Multi-armed Bandits: What Is It and What Is Its Cost?

Debabrota Basu, Christos Dimitrakakis, Aristide Tossou

Based on differential privacy (DP) framework, we introduce and unify privacy definitions for the multi-armed bandit algorithms. We represent the framework with a unified graphical…