7 papers
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning
Zetian Hu, Shunyu Liu, Junjie Zhang +4
Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on…
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
Junjie Zhang, Guozheng Ma, Shunyu Liu +5
Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs). However, existing step-level…
PVI: Plug-in Visual Injection for Vision-Language-Action Models
Zezhou Zhang, Songxin Zhang, Xiao Xiong +8
VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for…
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
Junjie Zhang, Guozheng Ma, Shunyu Liu +6
Reinforcement Learning with Verifiable Rewards~(RLVR) has emerged as a powerful learn-to-reason paradigm for large reasoning models to tackle complex tasks. However, the current RL…
L0: Reinforcement Learning to Become General Agents
Junjie Zhang, Jingyi Xi, Zhuoyang Song +7
Training large language models (LLMs) to act as autonomous agents for multi-turn, long-horizon tasks remains significant challenges in scalability and training efficiency. To addre…
Supervised Optimism Correction: Be Confident When LLMs Are Sure
Junjie Zhang, Rushuai Yang, Shunyu Liu +5
In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…