2 citations · 2 across the 2 of their papers we have counts for
3 papers
Stochastic Online Linear Regression: the Forward Algorithm to Replace Ridge
Reda Ouhamma, Odalric Maillard, Vianney Perchet
We consider the problem of online linear regression in the stochastic setting. We derive high probability regret bounds for online ridge regression and the forward algorithm. This…
Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits
Reda Ouhamma, Rémy Degenne, Pierre Gaillard +1
In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of e…
Learning Value Functions in Deep Policy Gradients using Residual Variance
Yannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard +1
Policy gradient algorithms have proven to be successful in diverse decision making and control tasks. However, these methods suffer from high sample complexity and instability issu…