activity
20172022
most citedImproving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning

2 citations · 4 across the 4 of their papers we have counts for

collaborators

12 papers

cs.LG20212 cited

Improving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning

Bogdan Mazoure, Paul Mineiro, Pavithra Srinath +3

We study session-based recommendation scenarios where we want to recommend items to users during sequential interactions to improve their long-term utility. Optimizing a long-term…

cs.LG2020

A Theoretical Analysis of Catastrophic Forgetting through the NTK Overlap Matrix

Thang Doan, Mehdi Bennani, Bogdan Mazoure +2

Continual learning (CL) is a setting in which an agent has to learn from an incoming stream of data during its entire lifetime. Although major advances have been made in the field,…

cs.LG2020

Deep Reinforcement and InfoMax Learning

Bogdan Mazoure, Remi Tachet des Combes, Thang Doan +2

We begin with the hypothesis that a model-free agent whose representations are predictive of properties of future states (beyond expected rewards) will be more capable of solving a…

cs.LG2020

Representation of Reinforcement Learning Policies in Reproducing Kernel Hilbert Spaces

Bogdan Mazoure, Thang Doan, Tianyu Li +4

We propose a general framework for policy representation for reinforcement learning tasks. This framework involves finding a low-dimensional embedding of the policy on a reproducin…

cs.AI20191 cited

Efficient Planning under Partial Observability with Unnormalized Q Functions and Spectral Learning

Tianyu Li, Bogdan Mazoure, Doina Precup +1

Learning and planning in partially-observable domains is one of the most difficult problems in reinforcement learning. Traditional methods consider these two problems as independen…

cs.LG2019

Attraction-Repulsion Actor-Critic for Continuous Control Reinforcement Learning

Thang Doan, Bogdan Mazoure, Moloud Abdar +3

Continuous control tasks in reinforcement learning are important because they provide an important framework for learning in high-dimensional state spaces with deceptive rewards, w…