2 papers
cs.LG2025
Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories
Jhouben Cuesta-Ramirez, Samuel Beaussant, Mehdi Mounsif
Large Language Models (LLMs) trained via Reinforcement Learning (RL) have recently achieved impressive results on reasoning benchmarks. Yet, growing evidence shows that these model…
cs.LG2025
Scaling Algorithm Distillation for Continuous Control with Mamba
Samuel Beaussant, Mehdi Mounsif
Algorithm Distillation (AD) was recently proposed as a new approach to perform In-Context Reinforcement Learning (ICRL) by modeling across-episodic training histories autoregressiv…