7 papers
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
Feng Zhang, Xinhong Ma, Ziqiang Dong +5
Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…
Reinforcing VLAs in Task-Agnostic World Models
Yucen Wang, Rui Yu, Fengming Zhang +5
Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly…
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
Zhiyuan Fan, Wenwei Jin, Feng Zhang +4
Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactions, thus enabling adaptation…
Vehicle-as-Prompt: A Unified Deep Reinforcement Learning Framework for Heterogeneous Fleet Vehicle Routing Problem
Shihong Huang, Shengjie Wang, Lei Gao +4
Unlike traditional homogeneous routing problems, the Heterogeneous Fleet Vehicle Routing Problem (HFVRP) involves heterogeneous fixed costs, variable travel costs, and capacity con…
ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
Feng Zhang, Zezhong Tan, Xinhong Ma +5
To address the limited capability expansion and low sample efficiency of Reinforcement Learning (RL), recent methods have integrated ''hints'' into post-training, which are prefix…
Towards Flash Thinking via Decoupled Advantage Policy Optimization
Zezhong Tan, Hang Gao, Xinhong Ma +2
Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although exi…