9 papers · 1 filter
Minimax PAC Bounds for Learning in Exogenous Contextual MDPs
Corentin Pla, Hugo Richard, Marc Abeille +1
We study PAC learning in tabular discounted Markov decision processes with exogenous i.i.d. contexts, with discount factor , finite state space , action space $\mat…
Instance-dependent Stochastic Lipschitz bandit
Marius Potfer, Vianney Perchet
We study the Lipschitz bandit problem, where a learner sequentially maximizes an unknown Lipschitz function over a domain using noisy pointwise ev…
Do Not Trust The Auctioneer: Learning to Bid in Feedback-Manipulated Auctions
Luigi Foscari, Matilde Tullii, Vianney Perchet
Shilling is the use of artificial bids to make competition appear stronger and push prices upward. We study repeated first-price auctions in which shilling affects feedback but not…
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
Pierre Perrault, Vianney Perchet, Michal Valko
We investigate stochastic combinatorial semi-bandits, where the entire joint distribution of outcomes impacts the complexity of the problem instance (unlike in the standard bandits…
Learning in Prophet Inequalities with Noisy Observations
Jung-hun Kim, Vianney Perchet
We study the prophet inequality, a fundamental problem in online decision-making and optimal stopping, in a practical setting where rewards are observed only through noisy realizat…
On the Hardness of Reinforcement Learning with Transition Look-Ahead
Corentin Pla, Hugo Richard, Marc Abeille +2
We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before decidi…