2 papers
cs.LG2026
OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation
Chenhao Qiu, Dawei Li, Yechao Zhang +2
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score…
cs.LG2025
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
Kefan Song, Amir Moeini, Peng Wang +4
Reinforcement learning (RL) is a framework for solving sequential decision-making problems. In this work, we demonstrate that, surprisingly, RL emerges during the inference time of…