1 citations · 1 across the 1 of their papers we have counts for
7 papers
Ensemble Bootstrapping for Q-Learning
Oren Peer, Chen Tessler, Nadav Merlis +1
Q-learning (QL), a common reinforcement learning algorithm, suffers from over-estimation bias due to the maximization term in the optimal Bellman operator. This bias may lead to su…
Confidence-Budget Matching for Sequential Budgeted Learning
Yonathan Efroni, Nadav Merlis, Aadirupa Saha +1
A core element in decision-making under uncertainty is the feedback on the quality of the performed actions. However, in many applications, such feedback is restricted. For example…
Reinforcement Learning with Trajectory Feedback
Yonathan Efroni, Nadav Merlis, Shie Mannor
The standard feedback model of reinforcement learning requires revealing the reward of every visited state-action pair. However, in practice, it is often the case that such frequen…
Tight Lower Bounds for Combinatorial Multi-Armed Bandits
Nadav Merlis, Shie Mannor
The Combinatorial Multi-Armed Bandit problem is a sequential decision-making problem in which an agent selects a set of arms on each round, observes feedback for each of these arms…
Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh +1
State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Mar…
Batch-Size Independent Regret Bounds for the Combinatorial Multi-Armed Bandit Problem
Nadav Merlis, Shie Mannor
We consider the combinatorial multi-armed bandit (CMAB) problem, where the reward function is nonlinear. In this setting, the agent chooses a batch of arms on each round and receiv…