Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Libra: Efficient Resource Management for Agentic RL Post-Training
Kaiwen Chen, Xin Tan, Jingzong Li +1
Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the rollout stage generate…
cs.LG2025
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
Junhao Shen, Haiteng Zhao, Yuzhe Gu +7
Enhancing large vision-language models (LVLMs) with visual slow-thinking reasoning is crucial for solving complex multimodal tasks. However, since LVLMs are mainly trained with vis…