Publications (9)
Improving a Proportional Integral Controller with Reinforcement Learning on a Throttle Valve Benchmark
Paul Daoudi, Bojan Mavkov, Bogdan Robu +4
This paper presents a learning-based control strategy for non-linear throttle valves with an asymmetric hysteresis, leading to a near-optimal controller without requiring any prior…
Building a User Foundation Model for the Open Web
Solal Vernier, Ivan Can Arisoy, Merwan Barlier +1
The paper introduces a self‑supervised transformer model trained on fragmented web browsing histories to create user representations that improve click prediction and bidding perfo…
Differentially Private Policy Gradient
Alexandre Rio, Merwan Barlier, Igor Colin
Motivated by the increasing deployment of reinforcement learning in the real world, involving a large consumption of personal data, we introduce a differentially private (DP) polic…
A Conservative Approach for Few-Shot Transfer in Off-Dynamics Reinforcement Learning
Paul Daoudi, Christophe Prieur, Bogdan Robu +2
Off-dynamics Reinforcement Learning (ODRL) seeks to transfer a policy from a source environment to a target environment characterized by distinct yet similar dynamics. In this cont…
Adaptive Sample Sharing for Multi Agent Linear Bandits
Hamza Cherkaoui, Merwan Barlier, Igor Colin
The multi-agent linear bandit setting is a well-known setting for which designing efficient collaboration between agents remains challenging. This paper studies the impact of data…
Differentially Private Deep Model-Based Reinforcement Learning
Alexandre Rio, Merwan Barlier, Igor Colin +1
We address private deep offline reinforcement learning (RL), where the goal is to train a policy on standard control tasks that is differentially private (DP) with respect to indiv…
PromptPack: Scaling LLM Annotation Agents for Online Recommendation
Sebastian Koralewski, Merwan Barlier, Yulia Stolin +1
Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single-call LLM annotation agent y…
Price of Safety in Linear Best Arm Identification
Xuedong Shang, Igor Colin, Merwan Barlier +1
We introduce the safe best-arm identification framework with linear feedback, where the agent is subject to some stage-wise safety constraint that linearly depends on an unknown pa…
Enhancing Reinforcement Learning Agents with Local Guides
Paul Daoudi, Bogdan Robu, Christophe Prieur +2
This paper addresses the problem of integrating local guide policies into a Reinforcement Learning agent. For this, we show how to adapt existing algorithms to this setting before…