157 citations · 165 across the 9 of their papers we have counts for
9 papers
Adversarial Contextual Bandits Go Kernelized
Gergely Neu, Julia Olkhovskaya, Sattar Vakili
We study a generalization of the problem of online learning in adversarial linear contextual bandits by incorporating loss functions that belong to a reproducing kernel Hilbert spa…
Importance-Weighted Offline Learning Done Right
Germano Gabbianelli, Gergely Neu, Matteo Papini
We study the problem of offline policy optimization in stochastic contextual bandit problems, where the goal is to learn a near-optimal policy based on a dataset of decision data c…
First- and Second-Order Bounds for Adversarial Linear Contextual Bandits
Julia Olkhovskaya, Jack Mayo, Tim van Erven +2
We consider the adversarial linear contextual bandit setting, which allows for the loss functions associated with each of arms to change over time without restriction. Assuming…
Offline Primal-Dual Reinforcement Learning for Linear MDPs
Germano Gabbianelli, Gergely Neu, Nneka Okolo +1
Offline Reinforcement Learning (RL) aims to learn a near-optimal policy from a fixed dataset of transitions collected by another policy. This problem has attracted a lot of attenti…
Optimistic Planning by Regularized Dynamic Programming
Antoine Moulin, Gergely Neu
We propose a new method for optimistic planning in infinite-horizon discounted Markov decision processes based on the idea of adding regularization to the updates of an otherwise s…
Online Learning with Off-Policy Feedback
Germano Gabbianelli, Matteo Papini, Gergely Neu
We study the problem of online learning in adversarial bandit problems under a partial observability model called off-policy feedback. In this sequential decision making problem, t…