Bradley-Terry model 1incomparability 1multi-objective reward modeling 1preference learning 1sample complexity 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability
Simone Drago, Marco Mussi, Leonardo Bianconi +1
The paper extends preference‑based reinforcement learning by allowing human experts to label trajectory pairs as incomparable, and introduces a Bradley‑Terry‑inspired rationality m…
cs.LG2025
Generalized Kernelized Bandits: A Novel Self-Normalized Bernstein-Like Dimension-Free Inequality and Regret Bounds
Alberto Maria Metelli, Simone Drago, Marco Mussi
We study the regret minimization problem in the novel setting of generalized kernelized bandits (GKBs), where we optimize an unknown function belonging to a reproducing kerne…
stat.ML2025
A Refined Analysis of UCBVI
Simone Drago, Marco Mussi, Alberto Maria Metelli
In this work, we provide a refined analysis of the UCBVI algorithm (Azar et al., 2017), improving both the bonus terms and the regret analysis. Additionally, we compare our version…