3 papers
cs.LG2023
Importance-Weighted Offline Learning Done Right
Germano Gabbianelli, Gergely Neu, Matteo Papini
We study the problem of offline policy optimization in stochastic contextual bandit problems, where the goal is to learn a near-optimal policy based on a dataset of decision data c…
cs.LG2023
Offline Primal-Dual Reinforcement Learning for Linear MDPs
Germano Gabbianelli, Gergely Neu, Nneka Okolo +1
Offline Reinforcement Learning (RL) aims to learn a near-optimal policy from a fixed dataset of transitions collected by another policy. This problem has attracted a lot of attenti…
cs.LG2022
Online Learning with Off-Policy Feedback
Germano Gabbianelli, Matteo Papini, Gergely Neu
We study the problem of online learning in adversarial bandit problems under a partial observability model called off-policy feedback. In this sequential decision making problem, t…