18 citations · 42 across the 6 of their papers we have counts for
12 papers
Scalable Representation Learning in Linear Contextual Bandits with Constant Regret Guarantees
Andrea Tirinzoni, Matteo Papini, Ahmed Touati +2
We study the problem of representation learning in stochastic contextual linear bandits. While the primary concern in this domain is usually to find realizable representations (i.e…
Learning One Representation to Optimize All Rewards
Ahmed Touati, Yann Ollivier
We introduce the forward-backward (FB) representation of the dynamics of a reward-free Markov decision process. It provides explicit near-optimal policies for any reward specified…
Sharp Analysis of Smoothed Bellman Error Embedding
Ahmed Touati, Pascal Vincent
The \textit{Smoothed Bellman Error Embedding} algorithm~\citep{dai2018sbeed}, known as SBEED, was proposed as a provably convergent reinforcement learning algorithm with general no…
TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
Joshua Romoff, Peter Henderson, David Kanaa +4
We investigate whether Jacobi preconditioning, accounting for the bootstrap term in temporal difference (TD) learning, can help boost performance of adaptive optimizers. Our method…
Zooming for Efficient Model-Free Reinforcement Learning in Metric Spaces
Ahmed Touati, Adrien Ali Taiga, Marc G. Bellemare
Despite the wealth of research into provably efficient reinforcement learning algorithms, most works focus on tabular representation and thus struggle to handle exponentially or in…
Stable Policy Optimization via Off-Policy Divergence Regularization
Ahmed Touati, Amy Zhang, Joelle Pineau +1
Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO) are among the most successful policy gradient approaches in deep reinforcement learning (RL). While t…