activity
20242026
collaborators

9 papers

cs.LG2026

Adaptive Bandit Algorithms for Contextual Matching Markets

Shiyun Lin, Simon Mauras, Vianney Perchet +1

We study bandit learning in matching markets, where players and arms constitute the two market sides, and the players' utilities are linear in the arm contexts. In each round, new…

stat.ML2026

On the Hardness of Reinforcement Learning with Transition Look-Ahead

Corentin Pla, Hugo Richard, Marc Abeille +2

We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before decidi…

cs.LG2026

Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching

Nadav Merlis

We study tabular reinforcement learning problems with multiple steps of lookahead information. Before acting, the learner observes steps of future transition and reward real…

cs.LG2025

Online Linear Regression with Paid Stochastic Features

Nadav Merlis, Kyoungseok Jang, Nicolò Cesa-Bianchi

We study an online linear regression setting in which the observed feature vectors are corrupted by noise and the learner can pay to reduce the noise level. In practice, this may h…

cs.GT2025

Stable Matching with Ties: Approximation Ratios and Learning

Shiyun Lin, Simon Mauras, Nadav Merlis +1

We study matching markets with ties, where workers on one side of the market may have tied preferences over jobs, determined by their matching utilities. Unlike classical two-sided…

cs.LG2025

On Bits and Bandits: Quantifying the Regret-Information Trade-off

Itai Shufaro, Nadav Merlis, Nir Weinberger +1

In many sequential decision problems, an agent performs a repeated task. He then suffers regret and obtains information that he may use in the following rounds. However, sometimes…