2 papers
cs.CL2026
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
Xia Zeng, Yihan Chen, Luhui Liu +3
We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The agent must follow a multi-stage St…
cs.CL2026
Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
Jianghao Su, Xia Zeng, Luhui Liu +3
Large Language Models (LLMs) empowered with Tool-Integrated Reasoning (TIR) can iteratively plan, call external tools, and integrate returned information to solve complex, long-hor…