3 papers
cs.LG2026
Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think
Otmane Sakhi, Aleksei Arzhantsev, Imad Aouali +1
Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged or asynchronous, so updates…
cs.IR2026
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
Imad Aouali, Flavian Vasile, Otmane Sakhi +2
LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet existing evaluations often reduc…
cs.LG2025
Offline Contextual Bandit with Counterfactual Sample Identification
Alexandre Gilotte, Otmane Sakhi, Imad Aouali +1
In production systems, contextual bandit approaches often rely on direct reward models that take both action and context as input. However, these models can suffer from confounding…