3 citations · 5 across the 4 of their papers we have counts for
4 papers · 1 filter
Sequential Counterfactual Risk Minimization
Houssam Zenati, Eustache Diemert, Matthieu Martin +2
Counterfactual Risk Minimization (CRM) is a framework for dealing with the logged bandit feedback problem, where the goal is to improve a logging policy using offline data. In this…
Efficient Kernel UCB for Contextual Bandits
Houssam Zenati, Alberto Bietti, Eustache Diemert +3
In this paper, we tackle the computational efficiency of kernelized UCB algorithms in contextual bandits. While standard methods require a O(CT^3) complexity where T is the horizon…
Zeroth-order non-convex learning via hierarchical dual averaging
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization - i.e., learning processes where, at each stage, the optimizer is facing an unkn…
Online non-convex optimization with imperfect feedback
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes - or otherwise constructs - an inexact model for the lo…