1 citations · 1 across the 31 of their papers we have counts for
5 papers · 1 filter
APPO: Agentic Procedural Policy Optimization
Xucong Wang, Ziyu Ma, Yong Wang +5
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…
DEvo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
Ru Zhang, Renda Li, Ziyu Ma +4
Reinforcement learning (RL) has demonstrated potential for enhancing reasoning in large language models (LLMs). However, effective RL training, which requires medium-difficulty tra…
AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting
Renda Li, Hailang Huang, Fei Wei +3
Reinforcement learning (RL) has demonstrated considerable potential for enhancing reasoning in large language models (LLMs). However, existing methods suffer from Gradient Starvati…
Tree Search for LLM Agent Reinforcement Learning
Yuxiang Ji, Ziyu Ma, Yong Wang +3
Recent advances in reinforcement learning (RL) have significantly enhanced the agentic capabilities of large language models (LLMs). In long-term and multi-turn agent tasks, existi…
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
Xiangxiang Chu, Hailang Huang, Xiao Zhang +2
Reinforcement Learning (RL) can directly enhance the reasoning capabilities of large language models without extensive reliance on Supervised Fine-Tuning (SFT). In this work, we re…