papers

Publications (9)

eess.SY2024

Improving a Proportional Integral Controller with Reinforcement Learning on a Throttle Valve Benchmark

Paul Daoudi, Bojan Mavkov, Bogdan Robu +4

This paper presents a learning-based control strategy for non-linear throttle valves with an asymmetric hysteresis, leading to a near-optimal controller without requiring any prior…

cs.LG2026

Building a User Foundation Model for the Open Web

Solal Vernier, Ivan Can Arisoy, Merwan Barlier +1

The paper introduces a self‑supervised transformer model trained on fragmented web browsing histories to create user representations that improve click prediction and bidding perfo…

#user modeling#self-supervised learning#transformer encoder#real-time bidding
cs.LG2025

Differentially Private Policy Gradient

Alexandre Rio, Merwan Barlier, Igor Colin

Motivated by the increasing deployment of reinforcement learning in the real world, involving a large consumption of personal data, we introduce a differentially private (DP) polic…

cs.LG2024

A Conservative Approach for Few-Shot Transfer in Off-Dynamics Reinforcement Learning

Paul Daoudi, Christophe Prieur, Bogdan Robu +2

Off-dynamics Reinforcement Learning (ODRL) seeks to transfer a policy from a source environment to a target environment characterized by distinct yet similar dynamics. In this cont…

cs.LG2025

Adaptive Sample Sharing for Multi Agent Linear Bandits

Hamza Cherkaoui, Merwan Barlier, Igor Colin

The multi-agent linear bandit setting is a well-known setting for which designing efficient collaboration between agents remains challenging. This paper studies the impact of data…

cs.LG2024

Differentially Private Deep Model-Based Reinforcement Learning

Alexandre Rio, Merwan Barlier, Igor Colin +1

We address private deep offline reinforcement learning (RL), where the goal is to train a policy on standard control tasks that is differentially private (DP) with respect to indiv…

cs.AI2026

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

Sebastian Koralewski, Merwan Barlier, Yulia Stolin +1

Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single-call LLM annotation agent y…

stat.ML2023

Price of Safety in Linear Best Arm Identification

Xuedong Shang, Igor Colin, Merwan Barlier +1

We introduce the safe best-arm identification framework with linear feedback, where the agent is subject to some stage-wise safety constraint that linearly depends on an unknown pa…

cs.LG2024

Enhancing Reinforcement Learning Agents with Local Guides

Paul Daoudi, Bogdan Robu, Christophe Prieur +2

This paper addresses the problem of integrating local guide policies into a Reinforcement Learning agent. For this, we show how to adapt existing algorithms to this setting before…