1 citations · 1 across the 9 of their papers we have counts for
14 papers · 1 filter
Adaptive Bandit Algorithms for Contextual Matching Markets
Shiyun Lin, Simon Mauras, Vianney Perchet +1
We study bandit learning in matching markets, where players and arms constitute the two market sides, and the players' utilities are linear in the arm contexts. In each round, new…
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
Nadav Merlis
We study tabular reinforcement learning problems with multiple steps of lookahead information. Before acting, the learner observes steps of future transition and reward real…
Online Linear Regression with Paid Stochastic Features
Nadav Merlis, Kyoungseok Jang, Nicolò Cesa-Bianchi
We study an online linear regression setting in which the observed feature vectors are corrupted by noise and the learner can pay to reduce the noise level. In practice, this may h…
Reinforcement Learning with Lookahead Information
Nadav Merlis
We study reinforcement learning (RL) problems in which agents observe the reward or transition realizations at their current state before deciding which action to take. Such observ…
On Bits and Bandits: Quantifying the Regret-Information Trade-off
Itai Shufaro, Nadav Merlis, Nir Weinberger +1
In many sequential decision problems, an agent performs a repeated task. He then suffers regret and obtains information that he may use in the following rounds. However, sometimes…
The Value of Reward Lookahead in Reinforcement Learning
Nadav Merlis, Dorian Baudry, Vianney Perchet
In reinforcement learning (RL), agents sequentially interact with changing environments while aiming to maximize the obtained rewards. Usually, rewards are observed only after acti…