most citedAgentic Reinforcement Learning with Implicit Step Rewards

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding

Xiaofeng Wang, Kakam Chong, Shuai Xiao +9

Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Curre…

cs.CL2026

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

Pinyi Zhang, Ting-En Lin, Yuchuan Wu +7

Personalized alignment of large language models seeks to adapt responses to individual user preferences, typically via reinforcement learning. A key challenge is obtaining accurate…

cs.CL2025

MOA: Multi-Objective Alignment for Role-Playing Agents

Chonghua Liao, Ke Wang, Yuchuan Wu +3

Role-playing agents (RPAs) require balancing multiple objectives, such as instruction following, persona consistency, and stylistic fidelity, which are not always perfectly aligned…

cs.CL20251 cited

Agentic Reinforcement Learning with Implicit Step Rewards

Xiaoqian Liu, Ke Wang, Yuchuan Wu +4

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, spa…

cs.CL2025

CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization

Xinge Ye, Rui Wang, Yuchuan Wu +4

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles w…

cs.CL2025

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

Haonan Zhang, Run Luo, Xiong Liu +10

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, exis…