Bradley-Terry model 1incomparability 1multi-objective reward modeling 1preference learning 1sample complexity 1
From the 1 of 14 linked papers with an AI index.
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
Gianmarco Genalti, Marco Mussi, Nicola Gatti +3
Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolve…
stat.ML2025
A Refined Analysis of UCBVI
Simone Drago, Marco Mussi, Alberto Maria Metelli
In this work, we provide a refined analysis of the UCBVI algorithm (Azar et al., 2017), improving both the bonus terms and the regret analysis. Additionally, we compare our version…
stat.ML2024
Open Problem: Tight Bounds for Kernelized Multi-Armed Bandits with Bernoulli Rewards
Marco Mussi, Simone Drago, Alberto Maria Metelli
We consider Kernelized Bandits (KBs) to optimize a function belonging to the Reproducing Kernel Hilbert Space (RKHS) . Mainstream…