most citedAgentic Reinforcement Learning with Implicit Step Rewards

1 citations · 1 across the 6 of their papers we have counts for

collaborators

11 papers

cs.SE2025

Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search

Guochang Li, Yuchen Liu, Zhen Qin +7

Repository-level software engineering tasks require large language models (LLMs) to efficiently navigate and extract information from complex codebases through multi-turn tool inte…

cs.CL20251 cited

Agentic Reinforcement Learning with Implicit Step Rewards

Xiaoqian Liu, Ke Wang, Yuchuan Wu +4

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, spa…

cs.CL2025

CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization

Xinge Ye, Rui Wang, Yuchuan Wu +4

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles w…

cs.CL2025

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

Haonan Zhang, Run Luo, Xiong Liu +10

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, exis…

cs.CL2025

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence

Guiyang Hou, Xing Gao, Yuchuan Wu +8

Recently, Large Language Models (LLMs) have made significant progress in IQ-related domains that require careful thinking, such as mathematics and coding. However, enhancing LLMs'…

cs.CL2025

Supervised Optimism Correction: Be Confident When LLMs Are Sure

Junjie Zhang, Rushuai Yang, Shunyu Liu +5

In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…