activity
20182021
most citedEnsemble Bootstrapping for Q-Learning

1 citations · 1 across the 1 of their papers we have counts for

collaborators

7 papers

cs.LG20211 cited

Ensemble Bootstrapping for Q-Learning

Oren Peer, Chen Tessler, Nadav Merlis +1

Q-learning (QL), a common reinforcement learning algorithm, suffers from over-estimation bias due to the maximization term in the optimal Bellman operator. This bias may lead to su…

cs.LG2021

Confidence-Budget Matching for Sequential Budgeted Learning

Yonathan Efroni, Nadav Merlis, Aadirupa Saha +1

A core element in decision-making under uncertainty is the feedback on the quality of the performed actions. However, in many applications, such feedback is restricted. For example…

cs.LG2020

Reinforcement Learning with Trajectory Feedback

Yonathan Efroni, Nadav Merlis, Shie Mannor

The standard feedback model of reinforcement learning requires revealing the reward of every visited state-action pair. However, in practice, it is often the case that such frequen…

cs.LG2020

Tight Lower Bounds for Combinatorial Multi-Armed Bandits

Nadav Merlis, Shie Mannor

The Combinatorial Multi-Armed Bandit problem is a sequential decision-making problem in which an agent selects a set of arms on each round, observes feedback for each of these arms…

cs.LG2019

Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies

Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh +1

State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Mar…

cs.LG2019

Batch-Size Independent Regret Bounds for the Combinatorial Multi-Armed Bandit Problem

Nadav Merlis, Shie Mannor

We consider the combinatorial multi-armed bandit (CMAB) problem, where the reward function is nonlinear. In this setting, the agent chooses a batch of arms on each round and receiv…