collaborators

7 papers

cs.LG2026

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

Feng Zhang, Xinhong Ma, Ziqiang Dong +5

Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…

cs.AI2026

Reinforcing VLAs in Task-Agnostic World Models

Yucen Wang, Rui Yu, Fengming Zhang +5

Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly…

cs.AI2026

Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents

Zhiyuan Fan, Wenwei Jin, Feng Zhang +4

Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactions, thus enabling adaptation…

cs.LG2026

Vehicle-as-Prompt: A Unified Deep Reinforcement Learning Framework for Heterogeneous Fleet Vehicle Routing Problem

Shihong Huang, Shengjie Wang, Lei Gao +4

Unlike traditional homogeneous routing problems, the Heterogeneous Fleet Vehicle Routing Problem (HFVRP) involves heterogeneous fixed costs, variable travel costs, and capacity con…

cs.CV2026

ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning

Feng Zhang, Zezhong Tan, Xinhong Ma +5

To address the limited capability expansion and low sample efficiency of Reinforcement Learning (RL), recent methods have integrated ''hints'' into post-training, which are prefix…

cs.AI2025

Towards Flash Thinking via Decoupled Advantage Policy Optimization

Zezhong Tan, Hang Gao, Xinhong Ma +2

Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although exi…