3 papers
cs.LG2025
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
Aleksei Arzhantsev, Otmane Sakhi, Flavian Vasile
Reinforcement learning (RL) is central to improving reasoning in large language models (LLMs) but typically requires ground-truth rewards. Test-Time Reinforcement Learning (TTRL) r…
cs.LG2025
Offline Contextual Bandit with Counterfactual Sample Identification
Alexandre Gilotte, Otmane Sakhi, Imad Aouali +1
In production systems, contextual bandit approaches often rely on direct reward models that take both action and context as input. However, these models can suffer from confounding…
stat.ML2025
Non-Linear Counterfactual Aggregate Optimization
Benjamin Heymann, Otmane Sakhi
We consider the problem of directly optimizing a non-linear function of an outcome, where this outcome itself is the sum of many small contributions. The non-linearity of the funct…