most citedVAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.AI2025

Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models

Yuhua Jiang, Jiawei Huang, Yufeng Yuan +4

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer f…

cs.AI2025

Truncated Proximal Policy Optimization

Tiantian Fan, Lingjun Liu, Yu Yue +20

Recently, test-time scaling Large Language Models (LLMs) have demonstrated exceptional reasoning capabilities across scientific and professional tasks by generating long chains-of-…

cs.CL2025

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier

Yuhua Jiang, Yuwen Xiong, Yufeng Yuan +5

Large Language Models (LLMs) have demonstrated impressive capabilities in complex reasoning tasks, yet they still struggle to reliably verify the correctness of their own outputs.…

cs.AI20251 cited

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Yu Yue, Yufeng Yuan, Qiying Yu +24

We present VAPO, Value-based Augmented Proximal Policy Optimization framework for reasoning models., a novel framework tailored for reasoning models within the value-based paradigm…

cs.LG2025

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Wenyuan Xu, Xiaochen Zuo, Chao Xin +3

Reinforcement Learning from Human Feedback (RLHF) has emerged as a important paradigm for aligning large language models (LLMs) with human preferences during post-training. This fr…

cs.LG2025

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Wei Shen, Guanlin Liu, Zheng Wu +5

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences. While recent research has focused on algorithmic improvement…