activity
20182026
most citedEnsemble Bootstrapping for Q-Learning

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG2026

Adaptive Bandit Algorithms for Contextual Matching Markets

Shiyun Lin, Simon Mauras, Vianney Perchet +1

We study bandit learning in matching markets, where players and arms constitute the two market sides, and the players' utilities are linear in the arm contexts. In each round, new…

cs.LG2026

Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching

Nadav Merlis

We study tabular reinforcement learning problems with multiple steps of lookahead information. Before acting, the learner observes steps of future transition and reward real…

cs.LG2025

Online Linear Regression with Paid Stochastic Features

Nadav Merlis, Kyoungseok Jang, Nicolò Cesa-Bianchi

We study an online linear regression setting in which the observed feature vectors are corrupted by noise and the learner can pay to reduce the noise level. In practice, this may h…

cs.LG2024

Reinforcement Learning with Lookahead Information

Nadav Merlis

We study reinforcement learning (RL) problems in which agents observe the reward or transition realizations at their current state before deciding which action to take. Such observ…

cs.LG2024

On Bits and Bandits: Quantifying the Regret-Information Trade-off

Itai Shufaro, Nadav Merlis, Nir Weinberger +1

In many sequential decision problems, an agent performs a repeated task. He then suffers regret and obtains information that he may use in the following rounds. However, sometimes…

cs.LG2024

The Value of Reward Lookahead in Reinforcement Learning

Nadav Merlis, Dorian Baudry, Vianney Perchet

In reinforcement learning (RL), agents sequentially interact with changing environments while aiming to maximize the obtained rewards. Usually, rewards are observed only after acti…