Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think
Otmane Sakhi, Aleksei Arzhantsev, Imad Aouali +1
Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged or asynchronous, so updates…
cs.LG2025
Offline Contextual Bandit with Counterfactual Sample Identification
Alexandre Gilotte, Otmane Sakhi, Imad Aouali +1
In production systems, contextual bandit approaches often rely on direct reward models that take both action and context as input. However, these models can suffer from confounding…