activity
20202025
most citedStochastic Online Linear Regression: the Forward Algorithm to Replace Ridge

2 citations · 3 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design

Andreas Schlaginhaufen, Reda Ouhamma, Maryam Kamgarpour

We study reinforcement learning from human feedback in general Markov decision processes, where agents learn from trajectory-level preference comparisons. A central challenge in th…

cs.LG20221 cited

Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration and Planning

Reda Ouhamma, Debabrota Basu, Odalric-Ambrym Maillard

We study the problem of episodic reinforcement learning in continuous state-action spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewa…

cs.LG20212 cited

Stochastic Online Linear Regression: the Forward Algorithm to Replace Ridge

Reda Ouhamma, Odalric Maillard, Vianney Perchet

We consider the problem of online linear regression in the stochastic setting. We derive high probability regret bounds for online ridge regression and the forward algorithm. This…

cs.LG2021

Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits

Reda Ouhamma, Rémy Degenne, Pierre Gaillard +1

In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of e…

cs.LG2020

Learning Value Functions in Deep Policy Gradients using Residual Variance

Yannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard +1

Policy gradient algorithms have proven to be successful in diverse decision making and control tasks. However, these methods suffer from high sample complexity and instability issu…