3 papers
cs.LG2026
Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think
Otmane Sakhi, Aleksei Arzhantsev, Imad Aouali +1
Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged or asynchronous, so updates…
cs.LG2026
Self-Consistency via Marginal Sharpening
Aleksei Arzhantsev, Otmane Sakhi, Nicolas Chopin
Inference-time sampling can elicit strong reasoning abilities from language models without additional training. Existing power-sampling methods do so by sharpening the distribution…
cs.LG2025
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
Aleksei Arzhantsev, Otmane Sakhi, Flavian Vasile
Reinforcement learning (RL) is central to improving reasoning in large language models (LLMs) but typically requires ground-truth rewards. Test-Time Reinforcement Learning (TTRL) r…