6 papers
Safe Online Bid Optimization with Return on Investment and Budget Constraints
Matteo Castiglioni, Alessandro Nuara, Giulia Romano +3
In online marketing, the advertisers aim to balance achieving high volumes and high profitability. The companies' business units address this tradeoff by maximizing the volumes whi…
Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning
Davide Salaorni, Vincenzo De Paola, Samuele Delpero +9
In recent years, \emph{Reinforcement Learning} (RL) has made remarkable progress, achieving superhuman performance in a wide range of simulated environments. As research moves towa…
A Reinforcement Learning Approach for Optimal Control in Microgrids
Davide Salaorni, Federico Bianchi, Francesco Trovò +1
The increasing integration of renewable energy sources (RESs) is transforming traditional power grid networks, which require new approaches for managing decentralized energy produc…
Sliding-Window Thompson Sampling for Non-Stationary Settings
Marco Fiandri, Alberto Maria Metelli, Francesco Trovò
Non-stationary multi-armed bandits (NS-MABs) model sequential decision-making problems in which the expected rewards of a set of actions, a.k.a.~arms, evolve over time. In this pap…
Thompson Sampling-like Algorithms for Stochastic Rising Bandits
Marco Fiandri, Alberto Maria Metelli, Francesco Trovò
Stochastic rising rested bandit (SRRB) is a setting where the arms' expected rewards increase as they are pulled. It models scenarios in which the performances of the different opt…
Rising Rested Bandits: Lower Bounds and Efficient Algorithms
Marco Fiandri, Alberto Maria Metelli, Francesco Trov`o
This paper is in the field of stochastic Multi-Armed Bandits (MABs), i.e. those sequential selection techniques able to learn online using only the feedback given by the chosen opt…